Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data-driven learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

105 records · Page 6

A Robust Machine Learning Schema for Developing, Maintaining, and Disseminating Machine Learning Models

Recent advances in the development of machine learning (ML) algorithms have enabled the creation of predictive models that can improve decision making, decrease computational cost, and improve efficiency in a variety of fields. As an organization begins to develop and implement such models, the data used in the training, validation, and testing of ML models, the model parameters, and the use cases or limitations of the models must be properly stored to ensure models are both fully traceable and used correctly. In the context of predicting material behavior, advances in computationally intense, physics-based modeling of material behavior at various length scales and the emergence of Integrated Computational Materials Engineering (ICME) have driven the need for developing data-driven surrogate models of the physics-based simulation tools using ML techniques. Surrogate model development allows for accurate material behavior prediction at a fraction of the cost of its physics-based counterpart, allowing for multiscale simulations of real-world applications, further enabling the ability to design fit-for-purpose materials for a reasonable computational investment. However, training such models requires extensive data, and thus, effective data management is necessary to reach the full potential that ML can offer to material design and ICME. This paper proposes a generalized, robust schema that allows organizations to store both real (experimental) and virtual (simulation) data used to train ML models and the defining model parameters and architectures within the Granta MI Platform. The developed schema allows for various types of data inputs and outputs, including single point values, time-series data, and images that can be used in the prediction of material behavior, while following outlined best practices for effective data management. An effective schema for ML data and models can help prevent the recreation of virtual/real training data and surrogate models, help reduce the time to create new models similar to existing ones by offering a starting point in the hyperparameter determination stages, minimize resources devoted to verification and validation (V&V) and certification of models, and ensure that data and surrogate models are not misused due to full traceability of both the data and ML model. It also allows organizations access to models that have already been developed, such that they can be used in the design of new materials, enabling the overall goals of ICME.

Brandon L. Hearley↗

Validation of Machine Learning Algorithms for Hyperspectral Inversion of Common Water Quality Indicators

The upcoming transition to a diverse suite hyperspectral airborne and orbiting optical sensors will provide an unprecedented opportunity to measure inland water quality characteristics at a fidelity not previously achievable. This presentation will assess prototype deep learning models trained on synthetic hyperspectral data and validated with collocated in-situ measurements. Synthesized data is becoming increasingly popular for use in data-driven approaches to complex problems, and can compliment real data to increase performance on complex and unusual phenomenon, reduce or test bias, and experiment to demonstrate explainability. We will present insights from hyperspectral inversions of Chlorophyl-a, Phycocyanin, and concentration of non-algal particles using selected orbiting and airborne sensors over diverse, optically complex aquatic scenarios. We analyze how various optical water types affect fidelity of results and where improvements can be made as we prototype for globally operational water quality algorithms which can be leveraged by upcoming hyperspectral missions such as the Surface Biology and Geology (SBG) mission.

Surface Biology and Geology (SBG)↗

Capturing Complex Multivariate Time Series Interactions to Detect High-Risk Adverse Events During Flight

The reduction of aviation safety metrics below target thresholds continue to drive down the number of aviation fatalities and accidents. To meet future safety demands, sustained efforts by aviation agencies promoting safety assurance processes and systems have prompted ongoing research on identifying and mitigating in-flight risks. With the projected increase in passenger load factor and rollout of more autonomous systems into the national airspace, the need to detect high-risk events in-time or ahead-of-time is becoming increasingly crucial. New anomaly detection and precursor identification algorithms will need to scale to different airframes, levels of autonomy, and system complexity. While the pervasiveness of deep learning has resulted in the development of performant anomaly detection methods, these sophisticated models currently suffer from low end-user interpretability. Building off our previous work on identifying adverse events in multivariate flight data during descent, we propose a data-driven approach for detecting in-flight adverse events caused by the complex interplay of flight variables. Our approach utilizes ordinal patterns of important aircraft stability variables (e.g., airspeed and descent rate) to capture multivariate flight dynamics that can be used to predict the onset of unstable approaches, a high-risk adverse event that can occur during approach. Through the use of ordinal patterns, we aim to create more interpretable detection models of in-flight adverse events that can be translated to future autonomous systems without difficulty. Our analysis shows the presence of distinct ordinal pattern distributions that can be used to predict unstable approaches 1 minute ahead of time with an accuracy of 0.69 and a recall of 0.73 and 30 seconds ahead with an accuracy of 0.70 and a recall of 0.86.

Risk detection↗

Capturing Multivariate Time Series Interactions to Detect High‑Risk Instability During Approach

The reduction of aviation safety metrics below target thresholds continue to drive down the number of aviation fatalities and accidents. To meet future safety demands, sustained efforts by aviation agencies promoting safety assurance processes and systems have prompted ongoing research on identifying and mitigating in-flight risks. With the projected increase in passenger load factor and rollout of more autonomous systems into the national airspace, the need to detect high-risk events in-time or ahead-of-time is becoming increasingly crucial. New anomaly detection and precursor identification algorithms will need to scale to different airframes, levels of autonomy, and system complexity. While the pervasiveness of deep learning has resulted in the development of performant anomaly detection methods, these sophisticated models currently suffer from low end-user interpretability. Building off our previous work on identifying adverse events in multivariate flight data during descent, we propose a data-driven approach for detecting in-flight adverse events caused by the complex interplay of flight variables. Our approach utilizes ordinal patterns of important aircraft stability variables (e.g., airspeed and descent rate) to capture multivariate flight dynamics that can be used to predict the onset of unstable approaches, a high-risk adverse event that can occur during approach. Through the use of ordinal patterns, we aim to create more interpretable detection models of in-flight adverse events that can be translated to future autonomous systems without difficulty. Our analysis shows the presence of distinct ordinal pattern distributions that can be used to predict unstable approaches 1 minute ahead of time with an accuracy of 0.69 and a recall of 0.73 and 30 seconds ahead with an accuracy of 0.70 and a recall of 0.86.

Risk detection↗

Leveraging Human Performance Data to Change the Narrative that People are the Safety Problem

The study of errors and failure has a long and productive history in the behavioral sciences. By studying how systems fail, we rule out various mechanisms for how those systems might work, thereby refining our theories of how they actually work. Human performance, however, includes more than errors; human performance comprises both failures and successes. A systematic bias to collect and analyze data only on error affects the decisions we make as a community by promoting the narrative that “people are the safety problem.” This narrative manifests in both obvious and subtle ways in the design of systems intended for human use. When the only safety data that are available are about human failure, then “data-driven” designs can only consider that humans fail. Changing this narrative will depend on new data and new ways to examine data – specifically, data on the processes by which human create and contribute to safety. An alternate narrative is that people represent a primary source of safety, through their capability to anticipate, monitor for, respond to, and learn from expected and unexpected change. This presentation will describe research efforts to expand the range of safety-relevant events to include not just rare safety failures but frequent safety successes. These efforts include use of data from both operations and simulations to develop methods and metrics for learning from structured observation, self-report, and system data.

Jon Holbrook↗

NASA’s Digital Information Platform to Accelerate the Transformation of the National Airspace System

In order to accelerate the digital transformation of airspace operations, a foundational framework and infrastructure for providing sustainable, data-driven, and cohesive decision-making digital services for both traditional and emergent air vehicles is being developed. The reference implementation of Digital Information Platform builds an ecosystem for the aviation community by providing access to a secure and trusted source of aviation data and services. Several key features and services have been implemented to enable secure data sharing, communication, and service registration on the Platform. The technical approach used to implement these features is presented here. NASA-developed integrated aviation data and machine learning based prediction services to optimize airspace operations are available on the Platform. These services are being evaluated in an operational environment by flight operators and the real-world benefits are being captured. The Platform fosters collaboration among industry and researchers to develop complex aviation services and the aim is to make it publicly accessible for consumption by the aviation community.

Digital Transformation↗

NASA’s Digital Information Platform to Accelerate the Transformation of the National Airspace System

In order to accelerate the digital transformation of airspace operations, a foundational framework and infrastructure for providing sustainable, data-driven, and cohesive decision-making digital services for both traditional and emergent air vehicles is being developed. The reference implementation of Digital Information Platform builds an ecosystem for the aviation community by providing access to a secure and trusted source of aviation data and services. Several key features and services have been implemented to enable secure data sharing, communication, and service registration on the Platform. The technical approach used to implement these features is presented here. NASA-developed integrated aviation data and machine learning based prediction services to optimize airspace operations are available on the Platform. These services are being evaluated in an operational environment by flight operators and the real-world benefits are being captured. The Platform fosters collaboration among industry and researchers to develop complex aviation services and the aim is to make it publicly accessible for consumption by the aviation community.

Digital Transformation↗

An Advanced Open-Source Platform for Air Quality Analysis, Visualization, and Prediction

Ambient air pollution is the largest environmental health risk factor, leading to several million premature deaths globally per year. The challenge of combating poor air quality is exacerbated by growing urban populations, changing emissions, and a warming climate. While there have been many advances monitoring and modeling of atmospheric composition, reflected in the dramatic increase in archived Earth Observations, there is no single measurement or method that alone can provide an accurate depiction of the entire atmosphere. The rapidly growing collections of observational and modeling data require us to be smarter about what data to include, and how such data is used. In recent years, NASA has invested significantly in advancing the concepts for Analytics Collaborative Framework (ACF) [5] and New Observing Strategies (NOS) [4] to tackle our software infrastructure need for harmonized data management and dynamic acquisition of diverse measurements for on-demand, interactive, multivariate analysis, and access [3]. It is not enough to have a big data, standalone analytics solution; it is critical that we start integrating data from remote sensing, modeling, and in-situ networks in a harmonized manner that enables timely and data-driven decision-making for air quality management. This work presents the design and development of an Air Quality Analytics Collaborative Framework (AQ ACF), as part of NASA’s Advanced Information Systems Technology (AIST) effort, to establish a data, machine-learning, and numerically driven platform for air quality analysis, visualization, and prediction.

Liu, Qian↗

Development of Digital Twin Technologies for Climate Projections

Climate projections are increasingly needed for adaptation, climate resilience and related decision making. However, existing projections have systematic biases, are limited in scope, and are not readily available for most potential users. While the ideal of an observational data-driven ‘digital twin’ for climate is initially attractive, there is only a very limited set of climate data available with which to train such a tool. Nonetheless, we are confident that there is a role for ‘digital twin technologies’ in removing biases, increasing computational efficiency, expanding scenarios and data accessibility.

digital twins↗

FluxSat: Long-term Earth Science Data Record (ESDR) for Terrestrial Gross Primary Production (GPP) based on satellite data calibrated with eddy covariance data

Gross primary production (GPP), the amount of carbon dioxide (CO 2 ) assimilated by plants through photosynthesis, is one of the most variable and uncertain components of the global carbon cycle. Global GPP has been estimated with a number of process-based models, data-driven, and hybrid approaches. Dynamic global vegetation models (DGVMs), driven by observed environmental changes, are used for global carbon budget assessments and long-term (climate) prediction. Benchmarking these and other models globally with data-driven GPP estimates is critical for understanding the land sink and ensuring accurate forecasts of the carbon cycle. In addition, global data-driven GPP estimates are crucial for studies of interannual variability, including trends that are linked to mechanisms with large uncertainties, such as the indirect CO 2 fertilization effect related to greening. In response to a community need for a GPP data set that well captures spatio-temporal variability, we developed FluxSat, a data-driven approach that optimizes the use of satellite reflectance data from the NASA MODerate-resolution Imaging Spectroradiometer (MODIS) on the Terra and Aqua satellites, calibrated using ground-based eddy covariance (EC) data. We are enhancing (spatially, higher resolution) and extending FluxSat (in time, with additional sensors) to create a high quality long term GPP Earth System Data Record (ESDR) for use in model benchmarking, carbon cycle modeling, and studies of trends and interannual variability. Our team’s objectives are to: 1. Update and document the current MODIS FluxSat GPP (daily, 0.05o and 0.5o resolutions) products with latest available MODIS and EC data sets; 2. Extend FluxSat GPP record forward in time with the Visible Infrared Imaging Radiometer Suite (VIIRS) on operational weather satellites going forward; 3. Extend FluxSat GPP record backward in time using the Advanced Very High Resolution Radiometer (AVHRR) on weather satellites dating back to 1981; 4. Provide higher spatial resolution MODIS and VIIRS GPP (0.0083o). 5. Thoroughly evaluate all FluxSat products with independent data; and 6. Create a homogenized long-term GPP record spanning 40+ years. We will discuss plans for this long-term data set that is supported through the NASA Making Earth System Data Records for Use in Research Environments (MEaSUREs) program.

gross Primary Production↗

First Impressions of Data Driven Printmaking: Preliminary Outcomes of an Art and Science Collaboration

Katie Baldwin (University of Alabama Huntsville) and Helen Parache (NASA Marshall Space Flight Center) collaborated on a pilot project with two upper division printmaking classes at the UAH Art Department in Spring 2023. Presentations and guided discussions focused on what artists and scientists can learn from each other, drawing from the larger discourse, and an introduction to artists that incorporate scientific data in their work. Parache and Baldwin facilitated a series of activities where students engaged with data to develop content. Each session ended with creating a plan with each student for a data-driven project utilizing printmaking which was shared in a talk open to the public. This pilot collaboration brought up themes of accessibility, social relevance, emotional truth, and creative solutions. Next steps include repeating the workshop with tweaks in Fall 2023 and strengthening relationships between artists and scientists co-located in Huntsville, Alabama.

data-driven art↗

A Machine Learning Approach to Determine Surface Radiative Fluxes based on CERES Observations

The Clouds and Earth’s Radiant Energy System (CERES) projects provides satellite-based observations of the radiative fluxes and clouds systems. CERES climate quality data products typically take several months of calibration and validation before release to the public. An alternative data product, Fast Longwave and Shortwave radiative Flux (FLASHFlux), was created to provide data to the applied sciences and educational users. FLASHFlux provides Top-of-Atmosphere radiative fluxes, Clouds properties, and parameterized surface radiative fluxes within four days for footprint (Level 2) data. We investigate the use of Artificial Neural Network (ANN) using MODerate resolution Imaging Spectroradiometer (MODIS) derived clouds properties and meteorology from the Global Assimilation and Meteorology Office (GMAO) scaled to the CERES footprint from the CERES Clouds Radiative Swath (CRS) data product to compute surface radiative fluxes. We test ANN produce fluxes against surface fluxes produced from the Fu-Liou model used in CRS and the Langley Parameterized Shortwave Algorithm (LPSA) and Langley Parameterized Longwave Algorithm (LPLA) used in FLASHFlux. We also validated each model with ground-based observations. Furthermore, we investigate Leave-One-Feature-Out Importance (LOFO) to evaluate the significance of each feature in our training and provide insight for future models. Advances in machine learning, along with increases in computational capabilities and available data allow us to estimate effects of unresolved processes in our climate without direct modeling. This work evaluates the ability to create accurate data-driven models to supplement or replace current models that estimate surface radiative fluxes.

Climatology↗

Planning Bias: Planning as a Source of Sampling Bias

Many data-driven planning methods are trained on data generated by planners. It is well known that many statistical learning methods are sensitive to sampling bias, and yet there has been little or no attention to planning as a sampling method and its role in introducing sampling bias into planner-generated training data. Recently, it has been demonstrated that A**,* in the presence of problems with variable heuristic error, prefers some solutions over other equally cost-optimal solutions. But, as we discuss in this paper, mitigation may not be as simple as resolving arbitrary tie-breaking by sampling from ties uniformly at random. In this paper, we formalize an intuition of planning bias. We focus on problems which output a single solution. Diverse planning only complicates the problem by generalizing it to bias in the set of sets; we show how it is subject to bias in the single solution. We make some useful observations about deterministic algorithms in contrast to non-deterministic algorithms. We explain how information entropy may be a good way to measure planning bias, and discuss some issues in evaluating practical approaches to measurement. We address the intuition that uniform random tiebreaking should mitigate bias; and sketch a novel approach to constructing an appropriate random distribution for duplicate detection during forward search for unbiased A*. Finally, we suggest directions for future work.

Planning Scheduling Algorithms↗

Stochastic Verification by Analysis for Autonomous Systems Management Architecture (ASMA)

The Gateway Vehicle Systems Manager (VSM) is the top-level of a distributed, hierarchical software control system. VSM is data-driven and will make decisions related to mission, fault, resource management and vehicle control. These attributes combined with a high degree of autonomy make it susceptible to emergent behavior. In order to achieve the high level of confidence needed in this critical system, the VSM team has developed a multifaceted verification strategy employing traditional verification techniques, simulation, model checking, and runtime verification. Individual algorithms are verified using conventional testing and model checking using assume-guarantee contracts. A discrete event-based simulation approach is being developed to verify timelines. This presentation describes an enhancement to the verification approach using analysis to enhance system robustness by detecting and resolving the potential for emergent behavior. The verification by analysis employs a Software in the Loop (SITL) environment with real flight software executing on emulated processors, simulations of vehicle subsystems, flight dynamics, and human inputs. Since the possible input space and configuration data set are too large for exhaustive testing, a Monte Carlo approach is used to cover feasible scenarios, augmented with corner cases and known higher-risk scenarios. A key problem in using Monte Carlo-based system verification is evaluating test results to ensure that system behavior is correct. The presentation describes the approach the VSM team uses to monitor behavior for compliance with predetermined boundaries and to identify anomalous behavior for further analysis. This presentation describes the multi-level systems approach to verification, and the simulation-based layer that covers the feasible state space: 1. Overview of the Gateway VSM 2. Special challenges due to heterogeneous, hierarchical architecture 3. Modeling and simulation environment using flight software and system simulations 4. Developing input sets to ensure state-space coverage 5. Developing model and data configuration sets to ensure model coverage 6. Interpreting results without predetermined outcomes 7. Lessons learned and future work

Verification and Validation↗

Development of a Safety Hazards Risk Assessment Tool for Uncrewed Aircraft System Traffic Management during Preflight Planning

Tremendous growth in the uncrewed and remotely piloted vehicle market is expected in low-altitude, uncontrolled airspace, resulting in potential decreases in safety without systems that support monitoring, assessing, and mitigating risk. At NASA, the System-Wide Safety (SWS) project has been developing a suite of data-driven tools to predict hazards so that the potential risks that these hazards pose can be mitigated. Services to predict various hazards have been developed, including battery capacity, proximity to static obstacles, population risks, global positioning system signal strength, radio frequency spectrum interference risk, and vertiport congestion. These services can monitor hazards along a flight path and if any risks posed by these hazards exceed a threshold, the uncrewed aircraft system (UAS) fleet manager can be alerted to mitigate the risk by modifying the flight path, changing the scheduled departure or arrival times, and/or diverting the vehicle to an alternate vertiport. These services were originally developed to monitor and assess risks during flight, but they have been adapted to assess hazard risks prior to departure so that a fleet manager can evaluate the potential risks for a fleet of UAS along their planned flight paths. These services have been integrated into a prototype tool called the Supplemental Data Service Provider-Consolidated Dashboard (SDSP-CD), developed at NASA Ames Research Center. The tool consists of a dashboard which provides a comprehensive overview for a number of risks and a map display that shows the details of the hazards along each flight’s path. Based on the findings from three previous studies, the SDSP-CD has been updated with new design elements and functions. In this paper, we describe lessons learned from the previous studies, changes made to the interface, and the feedback received during a follow-up usability study. Overall, participants reported that there is a substantial benefit of having a fleet manager use a consolidated dashboard to assess hazards for the vehicles in their fleet and to provide situational awareness to potential risks so that they can be mitigated prior to flight. Once the SDSP-CD matures, it will need to be integrated into flight and mission planning tools. Some initial thoughts on how this integration should be accomplished are shared in this paper. Finally, the functional differences between preflight vs. in-flight risk assessment and the differences in fleet manager vs. UAS pilot roles that may require different information and user interactions are discussed.

preflight↗