Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data-driven algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23

Predicting Arrival and Departure Runway Assignments with Machine Learning

Runway assignments at major airports are made by air traffic controllers subject to various constraints, and to achieve various objectives. In this research, we describe our efforts training machine learning (ML) models to predict both departure and arrival runway assignments using an entirely data-driven approach. This approach is compared to existing rule-based approaches developed in previous research using input from Subject Matter Experts. The models have features derived from various FAA data feeds, and leverage multiple machine learning algorithms. Results for models trained for nine major U.S. airports are described and compared to one another across various important dimensions. Particular attention was paid to developing a repeatable framework for training these models so the approach could be scaled to other airports, and to developing models that are useful in a real-time environment. In addition, the models were designed to be functional in a real-time environment to support NASA’s ATD-2 project, as part of an ML-powered shadow system to compare against the performance of the fielded system.

machine learning↗

Predicting Arrival and Departure Runway Assignments with Machine Learning

Runway assignments at major airports are made by air traffic controllers subject to various constraints, and to achieve various objectives. In this research, we describe our efforts training machine learning (ML) models to predict both departure and arrival runway assignments using an entirely data-driven approach. This approach is compared to existing rule-based approaches developed in previous research using input from Subject Matter Experts. The models have features derived from various FAA data feeds, and leverage multiple machine learning algorithms. Results for models trained for nine major U.S. airports are described and compared to one another across various important dimensions. Particular attention was paid to developing a repeatable framework for training these models so the approach could be scaled to other airports, and to developing models that are useful in a real-time environment. In addition, the models were designed to be functional in a real-time environment to support NASA’s ATD-2 project, as part of an ML-powered shadow system to compare against the performance of the fielded system.

machine learning↗

A Robust Machine Learning Schema for Developing, Maintaining, and Disseminating Machine Learning Models

Recent advances in the development of machine learning (ML) algorithms have enabled the creation of predictive models that can improve decision making, decrease computational cost, and improve efficiency in a variety of fields. As an organization begins to develop and implement such models, the data used in the training, validation, and testing of ML models, the model parameters, and the use cases or limitations of the models must be properly stored to ensure models are both fully traceable and used correctly. In the context of predicting material behavior, advances in computationally intense, physics-based modeling of material behavior at various length scales and the emergence of Integrated Computational Materials Engineering (ICME) have driven the need for developing data-driven surrogate models of the physics-based simulation tools using ML techniques. Surrogate model development allows for accurate material behavior prediction at a fraction of the cost of its physics-based counterpart, allowing for multiscale simulations of real-world applications, further enabling the ability to design fit-for-purpose materials for a reasonable computational investment. However, training such models requires extensive data, and thus, effective data management is necessary to reach the full potential that ML can offer to material design and ICME. This paper proposes a generalized, robust schema that allows organizations to store both real (experimental) and virtual (simulation) data used to train ML models and the defining model parameters and architectures within the Granta MI Platform. The developed schema allows for various types of data inputs and outputs, including single point values, time-series data, and images that can be used in the prediction of material behavior, while following outlined best practices for effective data management. An effective schema for ML data and models can help prevent the recreation of virtual/real training data and surrogate models, help reduce the time to create new models similar to existing ones by offering a starting point in the hyperparameter determination stages, minimize resources devoted to verification and validation (V&V) and certification of models, and ensure that data and surrogate models are not misused due to full traceability of both the data and ML model. It also allows organizations access to models that have already been developed, such that they can be used in the design of new materials, enabling the overall goals of ICME.

Brandon L. Hearley↗

Validation of Machine Learning Algorithms for Hyperspectral Inversion of Common Water Quality Indicators

The upcoming transition to a diverse suite hyperspectral airborne and orbiting optical sensors will provide an unprecedented opportunity to measure inland water quality characteristics at a fidelity not previously achievable. This presentation will assess prototype deep learning models trained on synthetic hyperspectral data and validated with collocated in-situ measurements. Synthesized data is becoming increasingly popular for use in data-driven approaches to complex problems, and can compliment real data to increase performance on complex and unusual phenomenon, reduce or test bias, and experiment to demonstrate explainability. We will present insights from hyperspectral inversions of Chlorophyl-a, Phycocyanin, and concentration of non-algal particles using selected orbiting and airborne sensors over diverse, optically complex aquatic scenarios. We analyze how various optical water types affect fidelity of results and where improvements can be made as we prototype for globally operational water quality algorithms which can be leveraged by upcoming hyperspectral missions such as the Surface Biology and Geology (SBG) mission.

Surface Biology and Geology (SBG)↗

Predicting Fiber Failure of Plain Weave Fabric with Recursive Multiscale Micromechanics

Recent advances in the development of machine learning (ML) algorithms have enabled the creation of predictive models that can improve decision making, decrease computational cost, and improve efficiency in a variety of fields. As an organization begins to develop and implement such models, the data used in the training, validation, and testing of machine learning models, the model parameters, and the use cases or limitations of the models must be properly stored to ensure models are both fully traceable and used correctly. In the context of predicting material behavior, advances in computationally intense, physics-based, modeling of material behavior at various length scales, and the emergence of Integrated Computational Materials Engineering (ICME) have driven the need for developing data-driven surrogate models of the physics-based simulation tools using machine learning (ML) techniques. Surrogate model development allows for accurate material behavior prediction at a fraction of the cost of its physics-based counterpart, allowing for multiscale simulations of real-world applications, further enabling the ability to design fit-for-purpose materials for a reasonable computational investment. However, training such models requires extensive data, and thus effective data management is necessary to reach the full potential that ML can offer to material design and ICME. This paper proposes a generalized, robust schema that allows organizations to store both real (experimental) and virtual (simulation) data used to train machine learning models and the defining model parameters and architectures. The developed schema allows for various types of data inputs and outputs, including single point values, time-series data, and images that can be used in for various types of machine learning models while following outlined best practices for effective data management. An effective schema for machine learning data and models can help prevent the recreation of virtual/real training data and surrogate models, can help reduce the time to create new models similar to existing ones by offering a starting point in the hyperparameter determination stages, minimize resources devoted to verification and validation (V&V) and certification of models, and ensure that data and surrogate models are not misused due to full traceability of both the data and ML model. It also allows organizations access to models that have already been developed, such that they can be used in the design of new materials, enabling the overall goals of ICME.

Failure↗

Capturing Complex Multivariate Time Series Interactions to Detect High-Risk Adverse Events During Flight

The reduction of aviation safety metrics below target thresholds continue to drive down the number of aviation fatalities and accidents. To meet future safety demands, sustained efforts by aviation agencies promoting safety assurance processes and systems have prompted ongoing research on identifying and mitigating in-flight risks. With the projected increase in passenger load factor and rollout of more autonomous systems into the national airspace, the need to detect high-risk events in-time or ahead-of-time is becoming increasingly crucial. New anomaly detection and precursor identification algorithms will need to scale to different airframes, levels of autonomy, and system complexity. While the pervasiveness of deep learning has resulted in the development of performant anomaly detection methods, these sophisticated models currently suffer from low end-user interpretability. Building off our previous work on identifying adverse events in multivariate flight data during descent, we propose a data-driven approach for detecting in-flight adverse events caused by the complex interplay of flight variables. Our approach utilizes ordinal patterns of important aircraft stability variables (e.g., airspeed and descent rate) to capture multivariate flight dynamics that can be used to predict the onset of unstable approaches, a high-risk adverse event that can occur during approach. Through the use of ordinal patterns, we aim to create more interpretable detection models of in-flight adverse events that can be translated to future autonomous systems without difficulty. Our analysis shows the presence of distinct ordinal pattern distributions that can be used to predict unstable approaches 1 minute ahead of time with an accuracy of 0.69 and a recall of 0.73 and 30 seconds ahead with an accuracy of 0.70 and a recall of 0.86.

Risk detection↗

Capturing Multivariate Time Series Interactions to Detect High‑Risk Instability During Approach

The reduction of aviation safety metrics below target thresholds continue to drive down the number of aviation fatalities and accidents. To meet future safety demands, sustained efforts by aviation agencies promoting safety assurance processes and systems have prompted ongoing research on identifying and mitigating in-flight risks. With the projected increase in passenger load factor and rollout of more autonomous systems into the national airspace, the need to detect high-risk events in-time or ahead-of-time is becoming increasingly crucial. New anomaly detection and precursor identification algorithms will need to scale to different airframes, levels of autonomy, and system complexity. While the pervasiveness of deep learning has resulted in the development of performant anomaly detection methods, these sophisticated models currently suffer from low end-user interpretability. Building off our previous work on identifying adverse events in multivariate flight data during descent, we propose a data-driven approach for detecting in-flight adverse events caused by the complex interplay of flight variables. Our approach utilizes ordinal patterns of important aircraft stability variables (e.g., airspeed and descent rate) to capture multivariate flight dynamics that can be used to predict the onset of unstable approaches, a high-risk adverse event that can occur during approach. Through the use of ordinal patterns, we aim to create more interpretable detection models of in-flight adverse events that can be translated to future autonomous systems without difficulty. Our analysis shows the presence of distinct ordinal pattern distributions that can be used to predict unstable approaches 1 minute ahead of time with an accuracy of 0.69 and a recall of 0.73 and 30 seconds ahead with an accuracy of 0.70 and a recall of 0.86.

Risk detection↗

A Machine Learning Approach to Determine Surface Radiative Fluxes based on CERES Observations

The Clouds and Earth’s Radiant Energy System (CERES) projects provides satellite-based observations of the radiative fluxes and clouds systems. CERES climate quality data products typically take several months of calibration and validation before release to the public. An alternative data product, Fast Longwave and Shortwave radiative Flux (FLASHFlux), was created to provide data to the applied sciences and educational users. FLASHFlux provides Top-of-Atmosphere radiative fluxes, Clouds properties, and parameterized surface radiative fluxes within four days for footprint (Level 2) data. We investigate the use of Artificial Neural Network (ANN) using MODerate resolution Imaging Spectroradiometer (MODIS) derived clouds properties and meteorology from the Global Assimilation and Meteorology Office (GMAO) scaled to the CERES footprint from the CERES Clouds Radiative Swath (CRS) data product to compute surface radiative fluxes. We test ANN produce fluxes against surface fluxes produced from the Fu-Liou model used in CRS and the Langley Parameterized Shortwave Algorithm (LPSA) and Langley Parameterized Longwave Algorithm (LPLA) used in FLASHFlux. We also validated each model with ground-based observations. Furthermore, we investigate Leave-One-Feature-Out Importance (LOFO) to evaluate the significance of each feature in our training and provide insight for future models. Advances in machine learning, along with increases in computational capabilities and available data allow us to estimate effects of unresolved processes in our climate without direct modeling. This work evaluates the ability to create accurate data-driven models to supplement or replace current models that estimate surface radiative fluxes.

Climatology↗

Planning Bias: Planning as a Source of Sampling Bias

Many data-driven planning methods are trained on data generated by planners. It is well known that many statistical learning methods are sensitive to sampling bias, and yet there has been little or no attention to planning as a sampling method and its role in introducing sampling bias into planner-generated training data. Recently, it has been demonstrated that A**,* in the presence of problems with variable heuristic error, prefers some solutions over other equally cost-optimal solutions. But, as we discuss in this paper, mitigation may not be as simple as resolving arbitrary tie-breaking by sampling from ties uniformly at random. In this paper, we formalize an intuition of planning bias. We focus on problems which output a single solution. Diverse planning only complicates the problem by generalizing it to bias in the set of sets; we show how it is subject to bias in the single solution. We make some useful observations about deterministic algorithms in contrast to non-deterministic algorithms. We explain how information entropy may be a good way to measure planning bias, and discuss some issues in evaluating practical approaches to measurement. We address the intuition that uniform random tiebreaking should mitigate bias; and sketch a novel approach to constructing an appropriate random distribution for duplicate detection during forward search for unbiased A*. Finally, we suggest directions for future work.

Planning Scheduling Algorithms↗

Damage Detection of a Pressure Vessel with Smart Sensing and Deep Learning

Structural Health Monitoring plays a crucial role in ensuring the safety and reliability of critical infrastructure, including pressure vessels involved in various applications. This research reports the damage detection of a pressure box employed in space habitat that operates in harsh environment where both structural failure and bolt joint loosening may occur. These failure modes are extremely hard to model based on first principles. We explore proper sensing mechanism and the associated inverse analysis algorithm that can elucidate the health condition of the pressure box. It is identified that piezoelectric impedance based active interrogation can provide necessary information for damage detection in such a system. Concurrently, deep learning technique leveraging spatial convolutional neural network is synthesized to analyze the raw data acquired and identify different types of damage. By training the deep learning model on a dataset of healthy and various damage scenarios, we can achieve high accuracy in identifying the presence of damage and its type. This research provides a data-driven methodology for structural damage detection using deep learning and has the potential to be extended to various systems with different failure modes.

Yang Zhang↗

Assessment of Quantum ML Applicability for Climate Actions: Comparison of the Variational Quantum Classifier and the Quantum Support Vector Classifier with Classical ML Models

Climate change refers to significant and long-term alterations in the Earth’s climate patterns, typically resulting from human activities that increase greenhouse gas emissions. Addressing climate change is not merely an option but a necessity, demanding creative solutions and efforts from individuals, researchers, communities, and governments. Despite the capabilities of machine learning (ML) with data-driven solutions promising to combat climate change-related problems, they face challenges stemming from traditional computational methods and prolonged training times, impeding their practical utility. Recent strides in quantum computing have permeated diverse domains, spanning from manufacturing engineering and pharmaceutical discovery to the latest frontier of detecting climate anomalies. With the potential to substantially reduce time and computational complexity, quantum computing shows promise in addressing climate change impacts. Its distinctive features will enable the concurrent exploration of expansive solution spaces, making it well-suited for analyzing extensive climate datasets, simulating intricate climate models, optimizing resource allocation, and discerning patterns in climate data for mitigation and adaptation endeavors. This study explores the potential of using Quantum machine learning (QML) techniques on climate and weather data obtained from NASA Giovannis. We used two QML algorithms, the Quantum Support Vector Classifier (QSVC) and the Variational Quantum Classifier (VQC) models, using the IBM Qiskit ML 0.7.2 ecosystem. We used an actual 127-Qubit IBM Quantum Computer (IBM 127-qubit Eagle) in this study. The methodology and results sections describe the experiences gained from applying and evaluating quantum ML results on climate and weather data obtained from NASA satellites as a novel practical application of quantum computing.

Earth Observational Data↗

Geothermal Operational Optimization with Machine Learning

The Geothermal Operational Optimization with Machine Learning (GOOML) project has developed a generic and extensible component-based system modeling framework to study complex geothermal fields using a data-driven approach. Through building a digital twin of a geothermal steam field with the GOOML modeling framework, operators can analyze historical and forecasted power production, explore possible steam field configurations, and optimize real world operations, all in a cost-effective digital environment. The GOOML modeling software is based on a historical data-assimilation framework that uses first-principal thermodynamics to model steam field components using historical data, and a forecast framework that uses machine-learning-driven models of steam field components to predict future operations. This modeling framework creates countless new opportunities for digital exploration of steam field design and operations. To date, digital twins have been developed for several steam fields in New Zealand and the United States. These digital twins have been validated by comparing hindcast predictions against historical production data. Field design and operations have been explored using genetic optimization and reinforcement learning. Initial results show compelling and often surprising opportunities for improved design and operation of fields with 2 to 5 percent improvements in annual energy production. GOOML is driving a step-change in geothermal operations by applying state-of-the-art machine learning algorithms, comprehensive data analytics, and a first-of-its-kind intelligent geothermal systems model.

40 EE - Geothermal Technologies Office (EE-4G)↗

Data-Driven Adaptive Damping Controller for Wind Power Plants with Doubly-Fed Induction Generators: Preprint

This paper presents an adaptive damping controller for wind power plants in which the turbines are equipped with doubly-fed induction generators. The controller is designed to respond to an input control signal that is triggered according to the system operating conditions. A processing unit continuously estimates the electromechanical modes of oscillation based on real-time streaming data acquired from a phasor measurement unit that is strategically positioned on the grid. The decision to trigger (or not trigger) the control signal is automatic, based on the relative damping of the dominant mode. The modes are estimated using the dynamic mode decomposition algorithm with time-delay embedding. Numerical simulations performed on the two-area system demonstrate that the proposed controller enhances the rotor angle stability for both small-signal and large disturbances, and is adaptive to changing grid conditions.

61 RADIATION PROTECTION AND DOSIMETRY↗

Data Curation for Machine Learning Applied to Geothermal Power Plant Operational Data for GOOML: Geothermal Operational Optimization with Machine Learning: Preprint

Geothermal Operational Optimization with Machine Learning (GOOML) is a transferable and extensible component-based geothermal asset modeling framework that considers complex steamfield relationships and identifies optimization prospects using a data-driven approach to physics-guided, data-centric machine learning. This framework has been used to develop digital twins that provide steamfield operators with operational environments to analyze and understand historical and forecasted power production, explore new steamfield configuration possibilities, and seek optimal asset management in real world applications. To create, test, and apply the GOOML framework, diverse time-series datasets spanning multiple years were sourced from various geothermal power plant components within several complex real-world geothermal operations. These operations are based in the United States and New Zealand and include a variety of technologies, end-uses and configurations, collectively covering nearly all relevant operating conditions for modern geothermal fields. Datasets were acquired from multiple sources to ensure that machine learning experiments generalized properly to various operating conditions. It was found that the data varied in quality, format, and completeness. To ensure consistency between the various datasets, a standardized data curation process was developed to reliably streamline data preparation. This paper will discuss best practices as learned from the GOOML data curation process which takes the following steps: 1) acquisition of large quantities of data from power plant operators, 2) digestion of data to gain an initial understanding of what is included, 3) data transformation, which includes converting the data into a standardized machine-readable format so that they can be visualized, quality checked, and cleaned, 4) quality assurance and quality control, involving identification of significant data gaps and apparent anomalies through mapping of data features to real world componentry via the GOOML historical model, followed by discussion with modelers and power plant operators to identify additional data needs and to resolve issues, 5) use in machine learning algorithms, and 6) repetition of steps one through five until all data needs are met and data are deemed suitable for producing trustworthy modeling results which may be disseminated, ideally along with the curated dataset. This iterative process is focused on improving the quality of the data rather than tuning machine learning model parameters and supports a shift towards data-centric AI as a means to improving real-world applicability of geothermal machine learning projects.

access↗

GOOML - Finding Optimization Opportunities for Geothermal Operations: Preprint

Geothermal Operational Optimization with Machine Learning (GOOML) is a transferable and extensible component-based geothermal asset modeling framework that considers complex steamfield relationships and identifies optimization prospects using a data-driven approach. We have used this framework to develop digital twins that provide steamfield operators with an operational environment to analyze and understand historical and forecasted power production, explore new steamfield configuration possibilities, and seek optimal asset management for real world applications. The GOOML modeling software is built on a generic component-based systems framework that allows for both historical and forecast analysis. A GOOML model can perform historical data-assimilation using first-principal thermodynamics to create a meaningful data model. Historical production data can then be coupled with a forecast framework to train machine-learning models of steamfield components to predict future outputs. This modeling environment enables digital exploration of steamfield design configurations and operational scenarios. GOOML digital twins have been developed for steamfields in New Zealand and the United States representing differing power generation and field conditions. These digital twins have been validated by comparing hindcast predictions against historical production data. Reinforcement learning experiments were conducted to demonstrate the ability to programmatically explore the operations space using machine learning agents. Our initial results are compelling; two to five percent increases in annual energy production were demonstrated by the GOOML models with no additional infrastructure build required. GOOML offers a new approach to geothermal operations by applying state-of-the-art machine learning algorithms, comprehensive data analytics, and interaction with digital twins. Through application of these tools, operators will realize greater availability and higher net generation which will increase the cost effectiveness of geothermal energy projects.

access↗

Adaptable Data Driven Model Predictive Control for Heat Pipe Microreactors

To establish a technical basis for self-regulating microreactors, a model predictive control (MPC) system is investigated to proactively respond to anomalies and disturbances in anticipation of potential deviations from operating setpoints. Due to the difficulty of developing a physics-based surrogate model that can accurately match plant data in various operating conditions, machine learning algorithms are used in MPC, which allow for learning from both simulation and operation data, thus efficiently describing the targeted transient with arbitrary accuracy. However, one of the biggest concerns in applying ML algorithms like artificial neural networks (ANNs) is that the predictive capabilities of ANN are limited by training data. If there are gaps between the training and target domain, the accuracy of an ANN can degrade significantly when it is used to predict unseen data. To improve the predictive capability of ANN and enable a confident use of data-driven MPCs outside the training data, this study proposes an adaptive data-driven MPC framework. The system will monitor the discrepancy between plant responses and surrogate predictions, fine-tune the ANN-based surrogate when a large discrepancy is detected, and continue MPC operation with updated surrogates. The framework is demonstrated on a point kinetic model for microreactors. The hyperparameters of the update strategy, including layers to update, error thresholds, learning rate discount, and number of data points used for fine-tuning, are optimized so the simulated microreactor is able to follow changes in setpoint with the smallest of deviations.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Extreme Risk Mitigation in Reinforcement Learning using Extreme Value Theory

Risk-sensitive reinforcement learning (RL) has garnered significant attention in recent years due to the growing interest in deploying RL agents in real-world scenarios. A critical aspect of risk awareness involves modelling highly rare risk events (rewards) that could potentially lead to catastrophic outcomes. These infrequent occurrences present a formidable challenge for data-driven methods aiming to capture such risky events accurately. While risk-aware RL techniques do exist, they suffer from high variance estimation due to the inherent data scarcity. Our work proposes to enhance the resilience of RL agents when faced with very rare and risky events by focusing on refining the predictions of the extreme values predicted by the state-action value distribution. To achieve this, we formulate the extreme values of the state-action value function distribution as parameterized distributions, drawing inspiration from the principles of extreme value theory (EVT). We propose an extreme value theory based actor-critic approach, namely, Extreme Valued Actor-Critic (EVAC) which effectively addresses the issue of infrequent occurrence by leveraging EVT-based parameterization. Importantly, we theoretically demonstrate the advantages of employing these parameterized distributions in contrast to other risk-averse algorithms. Our evaluations show that the proposed method outperforms other risk averse RL algorithms on a diverse range of benchmark tasks, each encompassing distinct risk scenarios.

Wang, Yu↗

A Distributed Model Identification Algorithm for Multi-Agent Systems: Preprint

In this study, we investigate agent-based approach for system model identification with emphasis on power distribution system applications. Departing from conventional practices of relying on historical data for offline model identification, we adopt online update approach utilizing real-time data by employing the latest data points for gradient computation. This methodology offers advantages including a large reduction in the communication network's bandwidth requirements by minimizing the data exchanged at each iteration and enabling the model to adapt in real-time to disturbances. Furthermore, we extend our model identification process from linear frameworks to more complex non-linear convex models. This extension is validated through numerical studies demonstrating improved control performance for a synthetic IEEE test case.

data-driven control↗