Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data-driven learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Observations on the Application of Machine Learning Techniques to Aviation Operations

There is an increasing interest in applying methods based on Machine Learning Techniques (MLT) to problems in aviation operations. The current interest is based on developments in Cloud Computing, the availability of open software and the success of MLT in automation, consumer behavior and finance involving large database. Historically aviation operations have been analyzed using physics-based models and provide information for making operational decisions. This talk describes issues to be addressed in applying either model-driven or data-driven methods. Aviation operations involving many decision makers, multiple objectives, poor or unavailable physics-based models and a rich historical database are prime candidates for analysis using data-driven methods. The issues are illustrated by a detailed example and summary of current research in the area. The application of MLT to aviation operations falls into two categories: (a) based on the lack of a physics-based model, MLT is the favored approach and (b) marginal difference between regression methods using physics-based models and MLT. Further research is needed in the selection of MLT to critical aviation operations. As always, the best approach depends on the task, the physical understanding of the problem and the quality and quantity of the available data.

Sridhar, Banavar↗

Application of Machine Learning Techniques to Aviation Operations: A Case Study

There is an increasing interest in applying methods based on Machine Learning Techniques (MLT) to problems in aviation operations. The current interest is based on developments in Cloud Computing, the availability of open software and the success of MLT in automation, consumer behavior and finance involving large database. Historically aviation operations have been analyzed using physics-based models and provide information for making operational decisions. This talk describes issues to be addressed in applying either model-driven or data-driven methods. Aviation operations involving many decision makers, multiple objectives, poor or unavailable physics-based models and a rich historical database are prime candidates for analysis using data-driven methods. The issues are illustrated by a detailed example and summary of current research in the area. The application of MLT to aviation operations falls into two categories 58; (a) based on the lack of a physics-based model, MLT is the favored approach and (b) marginal difference between regression methods using physics-based models and MLT. Further research is needed in the selection of MLT to critical aviation operations. As always, the best approach depends on the task, the physical understanding of the problem and the quality and quantity of the available data.

Sridhar, Banavar↗

Application of Machine Learning Techniques to Aviation Operations: NASA Case Studies

There is an increasing interest in applying methods based on Machine Learning Techniques(MLT) to problems in aviation operations. The current interest is based on developments in Cloud Computing, the availability of open software and the success of MLT in automation, consumer behavior and finance involving large database. Historically aviation operations have been analyzed using physics-based models and provide information for making operational decisions. This talk describes issues to be addressed in applying either model-driven or data-driven methods. Aviation operations involving many decision makers, multiple objectives, poor or unavailable physics-based models and a rich historical database are prime candidates for analysis using data-driven methods. The issues relating to data, feature selection and validation of the models are illustrated by examining case studies of the application of MLT to problems in air traffic management at NASA. Further research is needed in the application of MLT to critical aviation operations. As always, the best approach depends on the task, the physical understanding of the problem and the quality and quantity of the available data.

Sridhar, Banavar↗

Arm DevSummit Keynote: Environment for Data Engineering in Virtual Reality: Ethical Considerations

Currently, the US Government is going through a large-scale data Transformation effort, where the GSA playbook is guiding all agencies to make their processes data-driven with the help of emerging technologies. To address concerns about algorithm sharing, AI adoption, and vendor lock-in, the NASA Langley Research Center Digital Transformation Group has developed the Environment for Data Engineering in Virtual Reality (EnDEVR), a data science ecosystem that allows users to command and investigate customizable data analyses from a VR environment. The system has been evaluated in the Oculus Rift S and Quest environments, two popular VR systems powered by the ARM architecture. Current and future development will employ several AI capabilities to guide research and automation within the environment and to facilitate algorithm sharing. In this talk, we will discuss the results of an initial ethical investigation and recommended considerations for the use of AI within the system.

Artificial Intelligence↗

Mapping Global Forest Age from Forest Inventories, Biomass and Climate Data

Forest age can determine the capacity of a forest to uptake carbon from the atmosphere. However, a lack of global diagnostics that reflect the forest stage and associated disturbance regimes hampers the quantification of age-related differences in forest carbon dynamics. This study provides a new global distribution of forest age circa 2010, estimated using a machine learning approach trained with more than 40 000 plots using forest inventory, biomass and climate data. First, an evaluation against the plot-level measurements of forest age reveals that the data-driven method has a relatively good predictive capacity of classifying old-growth vs. non-old-growth (precision = 0.81 and 0.99 for old-growth and non-old-growth, respectively) forests and estimating corresponding forest age estimates (NSE = 0.6 – Nash–Sutcliffe efficiency – and RMSE = 50 years – root-mean-square error). However, there are systematic biases of overestimation in young- and underestimation in old-forest stands, respectively. Globally, we find a large variability in forest age with the old-growth forests in the tropical regions of Amazon and Congo, young forests in China, and intermediate stands in Europe. Furthermore, we find that the regions with high rates of deforestation or forest degradation (e.g. the arc of deforestation in the Amazon) are composed mainly of younger stands. Assessment of forest age in the climate space shows that the old forests are either in cold and dry regions or warm and wet regions, while young–intermediate forests span a large climatic gradient. Finally, comparing the presented forest age estimates with a series of regional products reveals differences rooted in different approaches and different in situ observations and global-scale products. Despite showing robustness in cross-validation results, additional methodological insights on further developments should as much as possible harmonize data across the different approaches. The forest age dataset presented here provides additional insights into the global distribution of forest age to better understand the global dynamics in the forest water and carbon cycles. The forest age datasets are openly available at https://doi.org/10.17871/ForestAgeBGI.2021 (Besnard et al., 2021).

Simon Besnard↗

A New Machine Learning Based Analysis for Improving Satellite Retrieved Atmospheric Composition Data: OMI SO2 as an Example

Despite recent progress, satellite retrievals of anthropogenic SO2 still suffer from relatively low signal-tonoise ratios. In this study, we demonstrate a new machine learning data analysis method to improve the quality of satellite SO2 products. In the absence of large ground-truth datasets for SO2, we start from SO2 slant column densities (SCDs) retrieved from the Ozone Monitoring Instrument (OMI) using a data-driven, physically based algorithm and calculate the ratio between the SCD and the root mean square (rms) of the fitting residuals for each pixel. To build the training data, we select presumably clean pixels with small SCD / rms ratios (SRRs) and set their target SCDs to zero. For polluted pixels with relatively large SRRs, we set the target to the original retrieved SCDs. We then train neural networks (NNs) to reproduce the target SCDs using predictors including SRRs for individual pixels, solar zenith, viewing zenith and phase angles, scene reflectivity, and O3 column amounts, as well as the monthly mean SRRs. For data analysis, we employ two NNs: (1) one trained daily to produce analyzed SO2 SCDs for polluted pixels each day and (2) the other trained once every month to produce analyzed SCDs for less polluted pixels for the entire month. Test results for 2005 show that our method can significantly reduce noise and artifacts over background regions. Over polluted areas, the monthly mean NN-analyzed and original SCDs generally agree to within ±15 %, indicating that our method can retain SO2 signals in the original retrievals except for large volcanic eruptions. This is further confirmed by running both the NN-analyzed and original SCDs through a topdown emission algorithm to estimate the annual SO2 emissions for ∼ 500 anthropogenic sources, with the two datasets yielding similar results. We also explore two alternative approaches to the NN-based analysis method. In one, we employ a simple linear interpolation model to analyze the original SCD retrievals. In the other, we develop a PCA–NN algorithm that uses OMI measured radiances, transformed and dimension-reduced with a principal component analysis (PCA) technique, as inputs to NNs for SO2 SCD retrievals. While the linear model and the PCA–NN algorithm can reduce retrieval noise, they both underestimate SO2 over polluted areas. Overall, the results presented here demonstrate that our new data analysis method can significantly improve the quality of existing OMI SO2 retrievals. The method can potentially be adapted for other sensors and/or species and enhance the value of satellite data in air quality research and applications.

Can Li↗

A Robust Machine Learning Schema for Developing, Maintaining, and Disseminating Machine Learning Models

Recent advances in the development of machine learning (ML) algorithms have enabled the creation of predictive models that can improve decision making, decrease computational cost, and improve efficiency in a variety of fields. As an organization begins to develop and implement such models, the data used in the training, validation, and testing of ML models, the model parameters, and the use cases or limitations of the models must be properly stored to ensure models are both fully traceable and used correctly. In the context of predicting material behavior, advances in computationally intense, physics-based modeling of material behavior at various length scales and the emergence of Integrated Computational Materials Engineering (ICME) have driven the need for developing data-driven surrogate models of the physics-based simulation tools using ML techniques. Surrogate model development allows for accurate material behavior prediction at a fraction of the cost of its physics-based counterpart, allowing for multiscale simulations of real-world applications, further enabling the ability to design fit-for-purpose materials for a reasonable computational investment. However, training such models requires extensive data, and thus, effective data management is necessary to reach the full potential that ML can offer to material design and ICME. This paper proposes a generalized, robust schema that allows organizations to store both real (experimental) and virtual (simulation) data used to train ML models and the defining model parameters and architectures within the Granta MI Platform. The developed schema allows for various types of data inputs and outputs, including single point values, time-series data, and images that can be used in the prediction of material behavior, while following outlined best practices for effective data management. An effective schema for ML data and models can help prevent the recreation of virtual/real training data and surrogate models, help reduce the time to create new models similar to existing ones by offering a starting point in the hyperparameter determination stages, minimize resources devoted to verification and validation (V&V) and certification of models, and ensure that data and surrogate models are not misused due to full traceability of both the data and ML model. It also allows organizations access to models that have already been developed, such that they can be used in the design of new materials, enabling the overall goals of ICME.

Brandon L. Hearley↗

Validation of Machine Learning Algorithms for Hyperspectral Inversion of Common Water Quality Indicators

The upcoming transition to a diverse suite hyperspectral airborne and orbiting optical sensors will provide an unprecedented opportunity to measure inland water quality characteristics at a fidelity not previously achievable. This presentation will assess prototype deep learning models trained on synthetic hyperspectral data and validated with collocated in-situ measurements. Synthesized data is becoming increasingly popular for use in data-driven approaches to complex problems, and can compliment real data to increase performance on complex and unusual phenomenon, reduce or test bias, and experiment to demonstrate explainability. We will present insights from hyperspectral inversions of Chlorophyl-a, Phycocyanin, and concentration of non-algal particles using selected orbiting and airborne sensors over diverse, optically complex aquatic scenarios. We analyze how various optical water types affect fidelity of results and where improvements can be made as we prototype for globally operational water quality algorithms which can be leveraged by upcoming hyperspectral missions such as the Surface Biology and Geology (SBG) mission.

Surface Biology and Geology (SBG)↗

Capturing Complex Multivariate Time Series Interactions to Detect High-Risk Adverse Events During Flight

The reduction of aviation safety metrics below target thresholds continue to drive down the number of aviation fatalities and accidents. To meet future safety demands, sustained efforts by aviation agencies promoting safety assurance processes and systems have prompted ongoing research on identifying and mitigating in-flight risks. With the projected increase in passenger load factor and rollout of more autonomous systems into the national airspace, the need to detect high-risk events in-time or ahead-of-time is becoming increasingly crucial. New anomaly detection and precursor identification algorithms will need to scale to different airframes, levels of autonomy, and system complexity. While the pervasiveness of deep learning has resulted in the development of performant anomaly detection methods, these sophisticated models currently suffer from low end-user interpretability. Building off our previous work on identifying adverse events in multivariate flight data during descent, we propose a data-driven approach for detecting in-flight adverse events caused by the complex interplay of flight variables. Our approach utilizes ordinal patterns of important aircraft stability variables (e.g., airspeed and descent rate) to capture multivariate flight dynamics that can be used to predict the onset of unstable approaches, a high-risk adverse event that can occur during approach. Through the use of ordinal patterns, we aim to create more interpretable detection models of in-flight adverse events that can be translated to future autonomous systems without difficulty. Our analysis shows the presence of distinct ordinal pattern distributions that can be used to predict unstable approaches 1 minute ahead of time with an accuracy of 0.69 and a recall of 0.73 and 30 seconds ahead with an accuracy of 0.70 and a recall of 0.86.

Risk detection↗

Capturing Multivariate Time Series Interactions to Detect High‑Risk Instability During Approach

The reduction of aviation safety metrics below target thresholds continue to drive down the number of aviation fatalities and accidents. To meet future safety demands, sustained efforts by aviation agencies promoting safety assurance processes and systems have prompted ongoing research on identifying and mitigating in-flight risks. With the projected increase in passenger load factor and rollout of more autonomous systems into the national airspace, the need to detect high-risk events in-time or ahead-of-time is becoming increasingly crucial. New anomaly detection and precursor identification algorithms will need to scale to different airframes, levels of autonomy, and system complexity. While the pervasiveness of deep learning has resulted in the development of performant anomaly detection methods, these sophisticated models currently suffer from low end-user interpretability. Building off our previous work on identifying adverse events in multivariate flight data during descent, we propose a data-driven approach for detecting in-flight adverse events caused by the complex interplay of flight variables. Our approach utilizes ordinal patterns of important aircraft stability variables (e.g., airspeed and descent rate) to capture multivariate flight dynamics that can be used to predict the onset of unstable approaches, a high-risk adverse event that can occur during approach. Through the use of ordinal patterns, we aim to create more interpretable detection models of in-flight adverse events that can be translated to future autonomous systems without difficulty. Our analysis shows the presence of distinct ordinal pattern distributions that can be used to predict unstable approaches 1 minute ahead of time with an accuracy of 0.69 and a recall of 0.73 and 30 seconds ahead with an accuracy of 0.70 and a recall of 0.86.

Risk detection↗

Leveraging Human Performance Data to Change the Narrative that People are the Safety Problem

The study of errors and failure has a long and productive history in the behavioral sciences. By studying how systems fail, we rule out various mechanisms for how those systems might work, thereby refining our theories of how they actually work. Human performance, however, includes more than errors; human performance comprises both failures and successes. A systematic bias to collect and analyze data only on error affects the decisions we make as a community by promoting the narrative that “people are the safety problem.” This narrative manifests in both obvious and subtle ways in the design of systems intended for human use. When the only safety data that are available are about human failure, then “data-driven” designs can only consider that humans fail. Changing this narrative will depend on new data and new ways to examine data – specifically, data on the processes by which human create and contribute to safety. An alternate narrative is that people represent a primary source of safety, through their capability to anticipate, monitor for, respond to, and learn from expected and unexpected change. This presentation will describe research efforts to expand the range of safety-relevant events to include not just rare safety failures but frequent safety successes. These efforts include use of data from both operations and simulations to develop methods and metrics for learning from structured observation, self-report, and system data.

Jon Holbrook↗

NASA’s Digital Information Platform to Accelerate the Transformation of the National Airspace System

In order to accelerate the digital transformation of airspace operations, a foundational framework and infrastructure for providing sustainable, data-driven, and cohesive decision-making digital services for both traditional and emergent air vehicles is being developed. The reference implementation of Digital Information Platform builds an ecosystem for the aviation community by providing access to a secure and trusted source of aviation data and services. Several key features and services have been implemented to enable secure data sharing, communication, and service registration on the Platform. The technical approach used to implement these features is presented here. NASA-developed integrated aviation data and machine learning based prediction services to optimize airspace operations are available on the Platform. These services are being evaluated in an operational environment by flight operators and the real-world benefits are being captured. The Platform fosters collaboration among industry and researchers to develop complex aviation services and the aim is to make it publicly accessible for consumption by the aviation community.

Digital Transformation↗

NASA’s Digital Information Platform to Accelerate the Transformation of the National Airspace System

In order to accelerate the digital transformation of airspace operations, a foundational framework and infrastructure for providing sustainable, data-driven, and cohesive decision-making digital services for both traditional and emergent air vehicles is being developed. The reference implementation of Digital Information Platform builds an ecosystem for the aviation community by providing access to a secure and trusted source of aviation data and services. Several key features and services have been implemented to enable secure data sharing, communication, and service registration on the Platform. The technical approach used to implement these features is presented here. NASA-developed integrated aviation data and machine learning based prediction services to optimize airspace operations are available on the Platform. These services are being evaluated in an operational environment by flight operators and the real-world benefits are being captured. The Platform fosters collaboration among industry and researchers to develop complex aviation services and the aim is to make it publicly accessible for consumption by the aviation community.

Digital Transformation↗

An Advanced Open-Source Platform for Air Quality Analysis, Visualization, and Prediction

Ambient air pollution is the largest environmental health risk factor, leading to several million premature deaths globally per year. The challenge of combating poor air quality is exacerbated by growing urban populations, changing emissions, and a warming climate. While there have been many advances monitoring and modeling of atmospheric composition, reflected in the dramatic increase in archived Earth Observations, there is no single measurement or method that alone can provide an accurate depiction of the entire atmosphere. The rapidly growing collections of observational and modeling data require us to be smarter about what data to include, and how such data is used. In recent years, NASA has invested significantly in advancing the concepts for Analytics Collaborative Framework (ACF) [5] and New Observing Strategies (NOS) [4] to tackle our software infrastructure need for harmonized data management and dynamic acquisition of diverse measurements for on-demand, interactive, multivariate analysis, and access [3]. It is not enough to have a big data, standalone analytics solution; it is critical that we start integrating data from remote sensing, modeling, and in-situ networks in a harmonized manner that enables timely and data-driven decision-making for air quality management. This work presents the design and development of an Air Quality Analytics Collaborative Framework (AQ ACF), as part of NASA’s Advanced Information Systems Technology (AIST) effort, to establish a data, machine-learning, and numerically driven platform for air quality analysis, visualization, and prediction.

Liu, Qian↗

Development of Digital Twin Technologies for Climate Projections

Climate projections are increasingly needed for adaptation, climate resilience and related decision making. However, existing projections have systematic biases, are limited in scope, and are not readily available for most potential users. While the ideal of an observational data-driven ‘digital twin’ for climate is initially attractive, there is only a very limited set of climate data available with which to train such a tool. Nonetheless, we are confident that there is a role for ‘digital twin technologies’ in removing biases, increasing computational efficiency, expanding scenarios and data accessibility.

digital twins↗

FluxSat: Long-term Earth Science Data Record (ESDR) for Terrestrial Gross Primary Production (GPP) based on satellite data calibrated with eddy covariance data

Gross primary production (GPP), the amount of carbon dioxide (CO 2 ) assimilated by plants through photosynthesis, is one of the most variable and uncertain components of the global carbon cycle. Global GPP has been estimated with a number of process-based models, data-driven, and hybrid approaches. Dynamic global vegetation models (DGVMs), driven by observed environmental changes, are used for global carbon budget assessments and long-term (climate) prediction. Benchmarking these and other models globally with data-driven GPP estimates is critical for understanding the land sink and ensuring accurate forecasts of the carbon cycle. In addition, global data-driven GPP estimates are crucial for studies of interannual variability, including trends that are linked to mechanisms with large uncertainties, such as the indirect CO 2 fertilization effect related to greening. In response to a community need for a GPP data set that well captures spatio-temporal variability, we developed FluxSat, a data-driven approach that optimizes the use of satellite reflectance data from the NASA MODerate-resolution Imaging Spectroradiometer (MODIS) on the Terra and Aqua satellites, calibrated using ground-based eddy covariance (EC) data. We are enhancing (spatially, higher resolution) and extending FluxSat (in time, with additional sensors) to create a high quality long term GPP Earth System Data Record (ESDR) for use in model benchmarking, carbon cycle modeling, and studies of trends and interannual variability. Our team’s objectives are to: 1. Update and document the current MODIS FluxSat GPP (daily, 0.05o and 0.5o resolutions) products with latest available MODIS and EC data sets; 2. Extend FluxSat GPP record forward in time with the Visible Infrared Imaging Radiometer Suite (VIIRS) on operational weather satellites going forward; 3. Extend FluxSat GPP record backward in time using the Advanced Very High Resolution Radiometer (AVHRR) on weather satellites dating back to 1981; 4. Provide higher spatial resolution MODIS and VIIRS GPP (0.0083o). 5. Thoroughly evaluate all FluxSat products with independent data; and 6. Create a homogenized long-term GPP record spanning 40+ years. We will discuss plans for this long-term data set that is supported through the NASA Making Earth System Data Records for Use in Research Environments (MEaSUREs) program.

gross Primary Production↗

First Impressions of Data Driven Printmaking: Preliminary Outcomes of an Art and Science Collaboration

Katie Baldwin (University of Alabama Huntsville) and Helen Parache (NASA Marshall Space Flight Center) collaborated on a pilot project with two upper division printmaking classes at the UAH Art Department in Spring 2023. Presentations and guided discussions focused on what artists and scientists can learn from each other, drawing from the larger discourse, and an introduction to artists that incorporate scientific data in their work. Parache and Baldwin facilitated a series of activities where students engaged with data to develop content. Each session ended with creating a plan with each student for a data-driven project utilizing printmaking which was shared in a talk open to the public. This pilot collaboration brought up themes of accessibility, social relevance, emotional truth, and creative solutions. Next steps include repeating the workshop with tweaks in Fall 2023 and strengthening relationships between artists and scientists co-located in Huntsville, Alabama.

data-driven art↗

A Machine Learning Approach to Determine Surface Radiative Fluxes based on CERES Observations

The Clouds and Earth’s Radiant Energy System (CERES) projects provides satellite-based observations of the radiative fluxes and clouds systems. CERES climate quality data products typically take several months of calibration and validation before release to the public. An alternative data product, Fast Longwave and Shortwave radiative Flux (FLASHFlux), was created to provide data to the applied sciences and educational users. FLASHFlux provides Top-of-Atmosphere radiative fluxes, Clouds properties, and parameterized surface radiative fluxes within four days for footprint (Level 2) data. We investigate the use of Artificial Neural Network (ANN) using MODerate resolution Imaging Spectroradiometer (MODIS) derived clouds properties and meteorology from the Global Assimilation and Meteorology Office (GMAO) scaled to the CERES footprint from the CERES Clouds Radiative Swath (CRS) data product to compute surface radiative fluxes. We test ANN produce fluxes against surface fluxes produced from the Fu-Liou model used in CRS and the Langley Parameterized Shortwave Algorithm (LPSA) and Langley Parameterized Longwave Algorithm (LPLA) used in FLASHFlux. We also validated each model with ground-based observations. Furthermore, we investigate Leave-One-Feature-Out Importance (LOFO) to evaluate the significance of each feature in our training and provide insight for future models. Advances in machine learning, along with increases in computational capabilities and available data allow us to estimate effects of unresolved processes in our climate without direct modeling. This work evaluates the ability to create accurate data-driven models to supplement or replace current models that estimate surface radiative fluxes.

Climatology↗