Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data processing automation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Data Mining for Understanding and Impriving Decision-Making Affecting Ground Delay Programs

The continuous growth in the demand for air transportation results in an imbalance between airspace capacity and traffic demand. The airspace capacity of a region depends on the ability of the system to maintain safe separation between aircraft in the region. In addition to growing demand, the airspace capacity is severely limited by convective weather. During such conditions, traffic managers at the FAA's Air Traffic Control System Command Center (ATCSCC) and dispatchers at various Airlines' Operations Center (AOC) collaborate to mitigate the demand-capacity imbalance caused by weather. The end result is the implementation of a set of Traffic Flow Management (TFM) initiatives such as ground delay programs, reroute advisories, flow metering, and ground stops. Data Mining is the automated process of analyzing large sets of data and then extracting patterns in the data. Data mining tools are capable of predicting behaviors and future trends, allowing an organization to benefit from past experience in making knowledge-driven decisions. The work reported in this paper is focused on ground delay programs. Data mining algorithms have the potential to develop associations between weather patterns and the corresponding ground delay program responses. If successful, they can be used to improve and standardize TFM decision resulting in better predictability of traffic flows on days with reliable weather forecasts. The approach here seeks to develop a set of data mining and machine learning models and apply them to historical archives of weather observations and forecasts and TFM initiatives to determine the extent to which the theory can predict and explain the observed traffic flow behaviors.

data mining↗

Informing forest carbon inventories under the Paris Agreement using ground-based forest monitoring data

Human interactions with forests have shaped Earth's climate for millennia and will continue to do so as we target net-zero emission goals. Accurately characterizing these climate impacts requires making reliable forest carbon data available for forest monitoring and planning. Here, we develop a semi-automated process for submitting forest carbon measurements from the largest relevant scientific database to the International Panel on Climate Change's Emission Factor Database, which currently has sparse forest carbon data. Building this bridge from scientific research to international policy is an important step towards managing forests in a net-zero motivated future. Humans have been influencing Earth's climate via transformative impacts on forests for millennia, and forests are now recognized as critical to climate change mitigation under the Paris Agreement. The efficacy of climate change mitigation planning and reporting depends on quality data on forest carbon (C) stocks and changes. The Emission Factor Database (EFDB) of the International Panel on Climate Change (IPCC) is intended to be a definitive source for such data, but needs comprehensive and well-documented data to be so. To facilitate submission of forest C estimates from scientific studies to EFDB, we develop and document a process for semi-automated data submission from the Global Forest C database (ForC v4.0), which is the largest compilation of ground-based forest C estimates. We then assess the data currently available through ForC and provide recommendations for improving forest data collection, analysis, and reporting. As of September 2024, ForC contained ~19,286 records potentially relevant to EFDB, 1068 of which had been submitted and posted to EFDB. These represented 19% of the total EFDB records for forest land. Records were unevenly distributed across variables and geographic regions. ForC records (37%) reviewed could not be submitted because the original publication lacked required information. In the future, ground-based forest C estimates should target gaps in the record, and studies should ensure that they report all information necessary for inclusion in EFDB. Given that climate change is rapidly impacting the world's forests, timely reporting of recent estimates will be critical to accurate forest C inventories.

54 ENVIRONMENTAL SCIENCES↗

In Situ Transmission Electron Microscopy: Signal processing challenges and examples

Transmission electron microscopy (TEM) is a powerful tool for imaging material structure and characterizing material chemistry. Recent advances in data collection technology for TEM have enabled high-volume and high-resolution data collection at a microsecond frame rate. Here, taking advantage of these advances in data collection rates requires the development and application of data processing tools, including image analysis, feature extraction, and streaming data processing techniques. In this article, we highlight a few areas in materials science that have benefited from combining signal processing and statistical analysis with data collection capabilities in TEM and present a future outlook on opportunities of integrating signal processing with automated TEM data analysis.

36 MATERIALS SCIENCE↗

Situational Lightning Climatologies for Central Florida: Phase V

The AMU added three years of data to the POR from the previous work resulting in a 22-year POR for the warm season months from 1989-2010. In addition to the flow regime stratification, moisture and stability stratifications were added to separate more active from less active lighting days within the same flow regime. The parameters used for moisture and stability stratifications were PWAT and TI which were derived from sounding data at four Florida radiosonde sites. Lightning data consisted of NLDN CG lightning flashes within 30 NM of each airfield. The AMU increased the number of airfields from nine to thirty-six which included the SLF, CCAFS, PAFB and thirty-three airfields across Florida. The NWS MLB requested the AMU calculate lightning climatologies for additional airfields that they support as a backup to NWS TBW which was then expanded to include airfields supported by NWS JAX and NWS MFL. The updated climatologies of lightning probabilities are based on revised synoptic-scale flow regimes over the Florida peninsula (Lambert 2007) for 5-, 10-, 20- and 30-NM radius range rings around the thirty-six airfields in 1-, 3- and 6-hour increments. The lightning, flow regime, moisture and stability data were processed in S-PLUS software using scripts written by the AMU to automate much of the data processing. The S-PLUS data files were exported to Excel to allow the files to be combined in Excel Workbooks for easier data handling and to create the tables and charts for the Gill. The AMU revised the Gill developed in the previous phase (Bauman 2009) with the new data and provided users with an updated HTML tool to display and manipulate the data and corresponding charts. The tool can be used with most web browsers and is computer operating system independent. The AMU delivered two Gills - one with just the PWAT stratification and one with both the PWAT and TI stratifications due to insufficient data in some of the PWATITI stratification combinations. This will allow the forecasters to choose a moisture-only or moisture/stability stratification depending on the flow regime and available data.

Bauman, William H., III↗

Automated Reduction of Data from Images and Holograms

Laser techniques are widely used for the diagnostics of aerodynamic flow and particle fields. The storage capability of holograms has made this technique an even more powerful. Over 60 researchers in the field of holography, particle sizing and image processing convened to discuss these topics. The research program of ten government laboratories, several universities, industry and foreign countries were presented. A number of papers on holographic interferometry with applications to fluid mechanics were given. Several papers on combustion and particle sizing, speckle velocimetry and speckle interferometry were given. A session on image processing and automated fringe data reduction techniques and the type of facilities for fringe reduction was held.

Lee, G.↗

Modern Scientific Data Governance Framework

Science has entered the era of Big Data with new challenges related to data governance, stewardship, and management. The existing data governance practices must catch up to ensure proper data management. Existing data governance policies and stewardship best practices tend to be disconnected from operational data management practices and enforcement and mainly exist in well-meaning documents or reports. These governance policies are, at best, partially implemented and rarely monitored or audited. In addition, existing governance policies keep adding additional data management steps that require a human, ‘a data steward’, in the loop, and the cost of data management can no longer scale proportionately with the current and future increased data volume and complexity. The goal for developing an updated data governance framework is to modernize scientific data governance to the reality of Big data and align it with the current technology trends such as cloud computing and AI. The goals of this framework are two folds. One is to ensure thoroughness that the governance adequately covers the entire data life cycle. Two, provide a practical approach that offers a consistent and repeatable process for different projects. Three core principles ground this framework. First, focus on just enough governance and prevent data governance from becoming a roadblock toward the scientific process. Remove any unnecessary processes and steps. Second, automate data management steps where possible. Actively remove steps that require ‘human in the loop’ within the management process to be efficient and scale with increasing data. Third, all the processes should continually be optimized using quantified metrics to streamline the monitoring and auditing workflows.

Rahul Ramachandran↗

Quantification of Operational Risk Using A Data Mining

What is Data Mining? - Data Mining is the process of finding actionable information hidden in raw data. - Data Mining helps find hidden patterns, trends, and important relationships often buried in a sea of data - Typically, automated software tools based on advanced statistical analysis and data modeling technology can be utilized to automate the data mining process

Perera, J. Sebastian↗

SETI prototype system for NASA's Sky Survey microwave observing project - A progress report

Two complementary search strategies, a Targeted Search and a Sky Survey, are part of NASA's SETI microwave observing project scheduled to begin in October of 1992. The current progress in the development of hardware and software elements of the JPL Sky Survey data processing system are presented. While the Targeted Search stresses sensitivity allowing the detection of either continuous or pulsed signals over the 1-3 GHz frequency range, the Sky Survey gives up sensitivity to survey the 99 percent of the sky that is not covered by the Targeted Search. The Sky Survey spans a larger frequency range from 1-10 GHz. The two searches will deploy special-purpose digital signal processing equipment designed and built to automate the observing and data processing activities. A two-million channel digital wideband spectrum analyzer and a signal processor system will serve as a prototype for the SETI Sky Survey processor. The design will permit future expansion to meet the SETI requirement that the processor concurrently search for left and right circularly polarized signals.

Klein, M. J.↗

Data management

The following tasks were prioritized: software acquisition management plan; space station flight data system architectural study; space station user data system interface; automation of software development process; automation of software testing; distributed data base management; ADA (automated data acquisition) evaluation and transition and planning; network operating system software; fault tolerant computer validation methodology for onboard data management system; systems integration; artificial intelligence/expert systems; space station data network concept; space station standard interface protocols; space station data networks systems; integrated software development facility; and language trade studies.

Love, G.↗

NASA GeneLab RNASeq Consensus Pipeline: A Nextflow Implementation

The NASA GeneLab project (genelab.nasa.gov) seeks to accelerate space biology research through cataloging and democratizing omics data. Since raw omics data is largely inaccessible to non-bioinformaticians, GeneLab works with the scientific community to develop standard processing pipelines to generate and publish processed data. Unlike raw data, processed data has greater immediate value to a wide range of users with varying technical backgrounds and computational capabilities. Standardizing processing workflows is essential to match the pace of raw data generation, ensure reproducibility, and enable standardized processed data for comparison across datasets. Previously, GeneLab developed a standardized pipeline for processing RNAseq data, referred to as the ‘GeneLab RNAseq Consensus Pipeline (RCP)’, in collaboration with GeneLab’s Analysis Working Groups. The work presented here is a Nextflow implementation of GeneLab’s RCP that automates and accelerates data processing of RNASeq datasets hosted on GeneLab. In addition to the core data processing, the workflow also includes staging of GeneLab raw data and a robust verification and validation (V&V) program that runs after each processing step to identify errors in real-time, stop additional downstream computation, and preserve computational resources. The workflow, including the staging and V&V functionality, is open source for others to reuse and modify at https://github.com/nasa/GeneLab_Data_Processing/tree/master/RNAseq.

Jonathan Dejesus Oribello↗

NASA GeneLab RNASeq Consensus Pipeline: A Nextflow Implementation

The NASA GeneLab project (genelab.nasa.gov) seeks to accelerate space biology research through cataloging and democratizing omics data. Since raw omics data is largely inaccessible to non-bioinformaticians, GeneLab works with the scientific community to develop standard processing pipelines to generate and publish processed data. Unlike raw data, processed data has greater immediate value to a wide range of users with varying technical backgrounds and computational capabilities. Standardizing processing workflows is essential to match the pace of raw data generation, ensure reproducibility, and enable standardized processed data for comparison across datasets. Previously, GeneLab developed a standardized pipeline for processing RNAseq data, referred to as the ‘GeneLab RNAseq Consensus Pipeline (RCP)’, in collaboration with GeneLab’s Analysis Working Groups. The work presented here is a Nextflow implementation of GeneLab’s RCP that automates and accelerates data processing of RNASeq datasets hosted on GeneLab. In addition to the core data processing, the workflow also includes staging of GeneLab raw data and a robust verification and validation (V&V) program that runs after each processing step to identify errors in real-time, stop additional downstream computation, and preserve computational resources. The workflow, including the staging and V&V functionality, is open source for others to reuse and modify at https://github.com/nasa/GeneLab_Data_Processing/tree/master/RNAseq.

Jonathan D Oribello↗

TPSAS-NF1676L-33992-DND

The CERES Science Team integrates and fuses observations from 6 CERES instruments aboard the Terra, Aqua, S-NPP, and NOAA-20 missions with data from more than 20 other unique data sources. Following the November 2017 launch of CERES Flight Model 6 (FM6) onboard NOAA20, CERES has now amassed over 80 instrument-years of valuable Earth radiation budget data. The rapidly growing volume of CERES data coupled with the introduction of new data products alongside improvements to existing science algorithms fosters the requirement for faster, more flexible, and scalable data production and orchestration. New virtualized, cloud-centric compute hardware hosted by the NASA Langley Research Center’s (LaRC) Atmospheric Sciences Data Center (ASDC) provides an ideal environment for these ever-increasing data production demands for CERES. This poster discusses updates to the implementation of the CERES Data Management Team’s (DMT) CERES AuTomAted job Loading sYSTem (CATALYST), a custom data processing workflow engine for CERES, to use on-demand computing resources to perform automated CERES data production processing in a Linux-based container environment. Linux containers provide CERES the flexibility to build multiple production environments in containers tailored for specific workloads and allow effortless provisioning of resources based on the CERES Science Team’s data production requirements.

Thomas N. Hillyer↗

Graphical Language for Data Processing

A graphical language for processing data allows processing elements to be connected with virtual wires that represent data flows between processing modules. The processing of complex data, such as lidar data, requires many different algorithms to be applied. The purpose of this innovation is to automate the processing of complex data, such as LIDAR, without the need for complex scripting and programming languages. The system consists of a set of user-interface components that allow the user to drag and drop various algorithmic and processing components onto a process graph. By working graphically, the user can completely visualize the process flow and create complex diagrams. This innovation supports the nesting of graphs, such that a graph can be included in another graph as a single step for processing. In addition to the user interface components, the system includes a set of .NET classes that represent the graph internally. These classes provide the internal system representation of the graphical user interface. The system includes a graph execution component that reads the internal representation of the graph (as described above) and executes that graph. The execution of the graph follows the interpreted model of execution in that each node is traversed and executed from the original internal representation. In addition, there are components that allow external code elements, such as algorithms, to be easily integrated into the system, thus making the system infinitely expandable.

Alphonso, Keith↗

Artificial intelligence issues related to automated computing operations

Large data processing installations represent target systems for effective applications of artificial intelligence (AI) constructs. The system organization of a large data processing facility at the NASA Marshall Space Flight Center is presented. The methodology and the issues which are related to AI application to automated operations within a large-scale computing facility are described. Problems to be addressed and initial goals are outlined.

Hornfeck, William A.↗

Oscilloscope Data Push Program

Data acquisition (DAQ) is a complex and costly process. Creating DAQ systems for analyzing a system requires expensive electronics and a dedicated team of engineers for support, posing a challenge for users who readily need data. This project is a proof of concept to create a temporary or one-off DAQ system using equipment commonly available to every team. We aim to automate the data acquisition process from the Rohde \& Schwarz RTO 1044 oscilloscope, convert the acquired binary data into floating point values, and store the results in a CSV file format. By developing a Python program to handle these tasks, we seek to reduce the manual effort involved in data collection, significantly increasing efficiency.

Osei-Tutu, Jason↗

The Atmospheric Data Acquisition And Interpolation Process For Center-TRACON Automation System

The Center-TRACON Automation System (CTAS), an advanced new air traffic automation program, requires knowledge of spatial and temporal atmospheric conditions such as the wind speed and direction, the temperature and the pressure in order to accurately predict aircraft trajectories. Real-time atmospheric data is available in a grid format so that CTAS must interpolate between the grid points to estimate the atmospheric parameter values. The atmospheric data grid is generally not in the same coordinate system as that used by CTAS so that coordinate conversions are required. Both the interpolation and coordinate conversion processes can introduce errors into the atmospheric data and reduce interpolation accuracy. More accurate algorithms may be computationally expensive or may require a prohibitively large amount of data storage capacity so that trade-offs must be made between accuracy and the available computational and data storage resources. The atmospheric data acquisition and processing employed by CTAS will be outlined in this report. The effects of atmospheric data processing on CTAS trajectory prediction will also be analyzed, and several examples of the trajectory prediction process will be given.

Jardin, M. R.↗

Challenges and Opportunities for Electric Utility Modeling and Asset Valuation Frameworks: Case Study on Valuing New Pumped Storage Hydropower

Asset valuation by electric utilities is becoming increasingly difficult in the rapidly changing electric sector. Rapid deployment of variable generation and inverter-based storage systems along with uncertain demand growth, climate, policies, and other factors create a challenging environment for understanding the value proposition of a new potential asset. This report describes an effort between the Tennessee Valley Authority (TVA) and three U.S. Department of Energy laboratories to perform a detailed review of utility modeling and analysis practices for asset valuation and identify challenges and opportunities for advancing its methods into the future. It focuses on a case study of new potential pumped storage hydropower (PSH) because of growing interest in new PSH capacity to provide energy balancing, firm capacity, and a range of ancillary services. Staff from the DOE labs conducted systematic interviews about current practices in capacity expansion modeling, production-cost modeling, hydrological modeling, and transmission stability modeling while also discussing how scenario analysis is conducted and how models and data are integrated. The effort resulted in a set of model, integration, and scenario recommendations that could be valuable to TVA, other utilities, system operators, and other stakeholders conducting integrated grid analysis. Individual model recommendations suggest exploring computational tradeoffs with detail and resolution across spatiotemporal structure, supply- and demand-side details, transmission overlays, market interactions, and ancillary services. Automated processes to pass data between models and conduct larger scenario suites could also enhance valuation practices by enabling a more consistent study of asset value across a broader range of uncertain future grid conditions where PSH could be particularly valuable. TVA and other industry stakeholders can learn from and adapt applied research-grade methods developed by DOE laboratories and other research institutions to improve decision making and accelerate progress towards a reliable, economic, sustainable energy system.

13 HYDRO ENERGY↗