Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Traditional Machine Learning Models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Unifying Quantum Materials Modeling and Experiments: The Role of Machine Learning Interatomic Potentials

Computational experiments have emerged as a powerful complement to traditional experiments in the design of new materials. The development of machine learning (ML) and deep learning techniques, combined with database construction and data mining, has significantly enhanced traditional quantum mechanical methods. This synergy enables the rapid development of structure-property relationships. In this talk, I will discuss our recent efforts in applying Machine Learning Interatomic Potentials (MLIAPs) to accelerate materials modeling across various material classes and challenging applications where traditional methods fall short. First, I will highlight the success of MLIAPs in accurately modeling the melting behavior of complex materials. Our results demonstrate high fidelity with experimental observations and also with calculated reference melting temperatures. In the second application, I will discuss how MLIAPs are trained and applied to elucidate the interplay between segregation tendencies and surface reconstructions in CuNi alloys under oxidizing conditions. A key factor in the success of these MLIAP applications is the design of minimalistic yet flexible datasets along with a computational framework for training MLIAPs.

Saidi, Wissam↗

28 NREL Stratus - Enabling Workflows to Fuse Data Streams, Modeling, Simulation, and Machine Learning: Preprint

Integrating cloud services into advanced computing facilities provides significant new capabilities over focusing solely on traditional high performance computing (HPC) workloads. This brings complementary capabilities as well as enabling new focused roles for HPC. They are especially potent for workflows that fuse data streams, modeling and simulation ('modsim') and machine learning. A key challenge to adopting a hybrid edge-cloud-HPC model is to align optimal capability, data, and user intent on the right resources for each step in a workflow.?The NREL Stratus service provides a basis for this: Stratus layers capabilities needed to make?cloud services accessible to a lab-based scientific community on commercial offerings, and; currently supports upwards of 200 projects ranging from IOT integration to traditional modeling and simulation. This provides a real-world inventory of scientific workflow elements. A growing knowledge base enables placing these elements appropriately between the edge, cloud, and traditional HPC. This paper outlines a vision via reference architecture and the application of that architecture in a typical workflow highlighting multiple components: sensor data intake, cleaning and transforming (edge/cloud suitable); generation of synthetic data through modsim, computationally heavy ML training and hyperparameter optimization (HPC suitable), and; inference and deployment (cloud ideal). Every step in such a workflow involves a cost-benefit analysis regarding the data movement, computational efficiency, availability, latency, and resource capabilities. The reference architecture and examples outlined allow for understanding new opportunities in the context of emerging workflows that combine IOT, cloud, and HPC to bolster scientific productivity.

AI↗

Comparison of Entry Descent and Landing Aerodynamic Databases with Uncertainty Quantification Developed Using Machine Learning Techniques

When developing the aerodynamic databases for use in trajectory simulations, it is important to develop a system of metrics to qualify which aerodynamic models are best to use. Since aerodynamics are just one input into trajectory simulations, the results of these simulations do not reflect on the quality of the aerodynamic database used. This means that aerodynamic database comparisons must be done offline. While traditional metrics that focus on mean/nominal predictions are a good first step, more robust estimates of the prediction interval become important as more focused uncertainty models are developed. We explore the limitations of evaluating aerodynamic models based purely on nominal-centered response surfaces. Before elaborating and evaluating metrics based on distributed models, the value of evaluating prediction interval and confidence interval are discussed to conclude that prediction intervals are more relevant to the use of trajectory analysis. Several metrics to evaluate the prediction interval are introduced with a focus on the standard calibration metric. Finally, we compare candidate models using both mean and distributed metrics. A finalized candidate model developed using state of the art machine learning methods is compared to a baseline model developed using traditional aerodynamic database modeling techniques.

Aerodynamic Database↗

Stable Machine‐Learning Parameterization of Subgrid Processes in a Comprehensive Atmospheric Model Learned From Embedded Convection‐Permitting Simulations

Modern climate projections often suffer from inadequate spatial and temporal resolution due to computational limitations, resulting in inaccurate representations of sub-grid processes. A promising technique to address this is the multiscale modeling framework (MMF), which embeds a kilometer-resolution cloud-resolving model (CRM) within each atmospheric column of a host climate model to replace traditional convection and cloud parameterizations. Machine learning offers a unique opportunity to make MMF more accessible by emulating the embedded CRM and reducing its substantial computational cost. Although many studies have demonstrated proof-of-concept success of achieving stable hybrid simulations, it remains a challenge to achieve near operational-level success with real geography and comprehensive variable emulation that includes, for example, explicit cloud condensate coupling. In this study, we present a stable hybrid model capable of integrating for at least 5 years with near operational-level complexity, including coarse-grid geography, seasonality, explicit cloud condensate and wind predictions, and land coupling. Our model demonstrates skillful online performance, achieving a 5-year zonal mean tropospheric temperature bias within 2 K, water vapor bias within 1 g/kg, and a precipitation root mean square error of 0.96 mm/day. Key factors contributing to our online performance include an expressive U-Net architecture and physical thermodynamic constraints for microphysics. With microphysical constraints mitigating unrealistic cloud formation, our work is the first to demonstrate realistic multi-year cloud condensate climatology under the MMF framework. Despite these advances, online diagnostics reveal persistent biases in certain regions, highlighting the need for innovative strategies to further optimize online performance.

Hu, Zeyuan [NVIDIA Corporation, Santa Clara, CA (U↗

ClimGen: Learning the Forcing-Response Relationship in Climate System

Solar Radiation Management (SRM) is emerging as a potential geoengineering strategy to address the anthropogenic impact on climate, but its effective implementation requires an iterative and large ensemble of highly accurate and efficient climate projections. Traditional climate projections rely on executing computationally demanding and time-consuming numerical climate models. Recent advances in machine learning (ML) aim to enhance these approaches by emulating traditional methods. In this work, we propose a novel framework for directly learning the relationship between solar radiation flux at the top of the atmosphere and the corresponding surface temperature response. To evaluate the feasibility of this direct ML-based projection, we developed a dataset using an intermediate complexity model, incorporating a comprehensive suite of different forcing patterns and evaluation metrics to rigorously assess the ML model’s performance. We introduce a Conditional Denoising Diffusion Probabilistic Model (cDDPM) for this task, which demonstrates encouraging skill in representing climate statistics under previously unseen forcing patterns. This approach provides a promising pathway for direct climate projections by accurately learning the forcing-response relationship, with a wide range of applications in impact mitigation, emissions policy design, and SRM strategies.

Chen, Tse-Chun [BATTELLE (PACIFIC NW LAB)] (ORCID:↗

A systematic review of machine learning in groundwater monitoring

With increasing concerns about water scarcity, groundwater has become crucial since this resource provides most of the freshwater needs. However, various human and natural activities often contaminate the groundwater, making it unsuitable for use. Over the years, scientists and engineers have used many methods to predict and track groundwater contamination as part of environmental monitoring. Consequently, there is an urgent need for improved methods, particularly in the face of increasing contamination. Machine learning has sometimes been used to monitor groundwater, air quality, and climate. Traditional methods must be improved due to the complexity and large amount of environmental data. This includes using hybrid models that combine traditional and new techniques. Despite the use of machine learning in many scientific areas, there is a lack of comprehensive reviews focusing on its use in environmental monitoring, especially groundwater monitoring. We aim to fill this gap by exploring machine-learning applications in groundwater monitoring. We discuss relevant methods, their limitations, and future potential. We summarize research on automating data processing and model training using groundwater sensor data. Our research underscores the transformative potential of machine learning to revolutionize long-term groundwater monitoring and contamination detection, providing valuable insights for future research and practical applications.

AI/ML↗

Evolving Multi-hazard Machine Learning Modeling for Advanced Risk-Informed Infrastructure Resilience Assessment

The socioeconomic impacts of pipeline incidents have escalated over the past three decades, revealing the limitation of traditional risk modeling methods when applied to extensive pipeline networks. This research aims to develop machine learning (ML) models that effectively identify, rank, and predict the diverse hazards and socioeconomic consequences associated with pipeline incidents. Utilizing historical data on pipeline incidents alongside weather and oceanographic data from the 1980s onward, the Houston metropolitan area serves as a testbed for the proposed methodologies. The research segments the combined datasets into three consecutive periods, demonstrating the efficacy of the updated model in predicting future events, particularly concerning precipitation rate data. Despite the challenges posed by a relatively limited dataset, local-level ML modeling offers valuable insights into the spatial and temporal dynamics of multiple hazards that contribute to pipeline incidents. These findings hold significant implications for future research, particularly in understanding and mitigating risks in various locations across the Gulf Coast and other coastal regions.

42 ENGINEERING↗

Enabling Interoperability in Earth System Digital Twins (ESDT): Integrating Observations, Models, and AI for Actionable Insights Through NASA'S Intelligent Systems Technology Program

NASA’s Intelligent Systems Technology Program (IST) is driving a paradigm shift in Earth science through the development of Earth System Digital Twins (ESDT). These integrated information systems create a dynamic "digital replica" of the Earth by harmonizing continuous, multi-source observations with high-fidelity models and state-of-the-art artificial intelligence (AI) that enable “What now?”, “What next?”, and “What if?” scenario building. These scenarios are reflected in NASA IST’s series of ESDTs, from the Coastal Zone Digital Twin that integrates complex data on the current state of the Chesapeake Bay to the Terrestrial Environmental Rapid-Replication and Assimilation Hydrometeorological (TerraHydro) AI-based ESDT that forecasts water movement across Earth’s surface, to the Agriculture Land Information System (AgLIS) which can be used to assess optimal planting dates and crop yield estimates. By bridging the gap between vast data archives and actionable insights, these projects enable a system-of-systems approach to understanding complex, interacting Earth processes. This poster will highlight recent innovations and future directions from NASA’s ESDT initiatives: Continuous Data Assimilation & Multi-Source Fusion. A core requirement of the ESDT work is the transition from static models to dynamic "living" replicas. This involves creating frameworks for the continual assimilation of near-real-time data from uncoordinated, heterogeneous sources, including satellite observations and airborne assets, and ground-based Internet of Things (IoT) sensors. These systems link design, operational status, and environmental data, ensuring the digital twin accurately reflects the current state of the physical Earth system. High-Fidelity Hybrid Modeling & Computational Acceleration to enable interactive "what-if" explorations, programs are moving beyond traditional, slow physical solvers by developing fast surrogate machine learning models and Deep Generative Models (DGMs). These hybrid approaches use neural networks to emulate complex physics, such as cloud feedback or ocean dynamics, at a fraction of the original computing cost, often leveraging advanced hardware like Graphics Processing Units (GPUs) to achieve the necessary scale. Federated Ecosystems & Interoperable Frameworks rather than building isolated tools, NASA IST is moving toward federated ESDTs and reusable analytic collaborative frameworks. This theme focuses on interoperability standards and common ontologies that allow specialized digital twins to interact and share data. This system-of-systems architecture supports multi-discipline investigations, such as analyzing how upstream watershed changes impact downstream urban flooding or how wildfire emissions affect regional air quality. By leveraging these advancements, ESDTs empower researchers and decision-makers to conduct real-time analysis and run complex hypothetical scenarios, ultimately improving our understanding of Earth’s evolving systems and informing critical real-world applications.

Earth System↗

Large-scale scenarios of electric vehicle charging with a data-driven model of control

Transportation electrification is forecast to bring millions of new electric vehicles to roads worldwide this decade. Planning to support those vehicles depends on detailed scenarios of their electricity demand in both uncontrolled and controlled or smart charging scenarios. In this work, we present a novel modeling approach to enable rapid generation of demand estimates that represent the impact of controlled charging for large-scale scenarios with millions of individual drivers. To model the effect of load modulation control on aggregate charging profiles, we propose a novel machine learning approach that replaces traditional optimization approaches. We demonstrate its performance modeling workplace charging control under a range of electricity rate schedules, achieving small errors (2.5%–4.5%) while accelerating computations by more than 4000 times. To generate the uncontrolled charging demand for scenarios with residential, workplace, and public charging we use statistical representations of a large data set of real charging sessions. We demonstrate the methodology by generating diverse sets of scenarios for California's charging demand in 2030 which consider multiple charging segments and controls, each run locally in under 50 s. We further demonstrate support for rate design by modeling the large-scale impact of a new, custom rate schedule for workplace charging.

33 ADVANCED PROPULSION SYSTEMS↗

Identification of new marker genes from plant single‐cell RNA‐seq data using interpretable machine learning methods

Summary An essential step in the analysis of single‐cell RNA sequencing data is to classify cells into specific cell types using marker genes. In this study, we have developed a machine learning pipeline called single‐cell predictive marker (SPmarker) to identify novel cell‐type marker genes in the Arabidopsis root. Unlike traditional approaches, our method uses interpretable machine learning models to select marker genes. We have demonstrated that our method can: assign cell types based on cells that were labelled using published methods; project cell types identified by trajectory analysis from one data set to other data sets; and assign cell types based on internal GFP markers. Using SPmarker, we have identified hundreds of new marker genes that were not identified before. As compared to known marker genes, the new marker genes have more orthologous genes identifiable in the corresponding rice single‐cell clusters. The new root hair marker genes also include 172 genes with orthologs expressed in root hair cells in five non‐ Arabidopsis species, which expands the number of marker genes for this cell type by 35–154%. Our results represent a new approach to identifying cell‐type marker genes from scRNA‐seq data and pave the way for cross‐species mapping of scRNA‐seq data in plants.

54 ENVIRONMENTAL SCIENCES↗

Glass Design Using Machine Learning Property Models with Prediction Uncertainties: Nuclear Waste Glass Formulation

The United States Department of Energy is responsible for managing the legacy nuclear waste stored in underground tanks at the Hanford Site. The waste will be separately vitrified as low-activity waste and high-level waste fractions. Waste glass formulation algorithms have been traditionally developed using partial quadratic mixture property-composition models. Recently, machine learning (ML) techniques have been used to predict glass properties and discover new glass materials for nuclear waste vitrification, and these advancements can be utilized to improve waste glass composition design. In this proof-of-principle study, ML algorithms such as Gaussian process regression (GPR) were used to interpolate glass properties (e.g., viscosity, electrical conductivity, chemical durability). After selecting appropriate sets of GPR hyper-parameters for each property, an optimization program was developed to formulate glass compositions to maximize waste loading while simultaneously satisfying property within constraints. The results of the ML-based waste loadings and glass compositions were compared to those obtained using the traditional methods. Comparing to the previous glass design framework, the ML-based optimization methods offer improved glass designs and a streamlined approach to generation of optimally designed data and near real-time updates.

glass formulation, machine learning, constraints, ↗

Improving vertical detail in simulated temperature and humidity data using machine learning

Atmospheric models used for weather forecasting and climate predictions discretise the atmosphere onto a vertical grid. There are however atmospheric phenomena that occur on scales smaller than the thickness of those model layers. The formation of low-level clouds due to temperature inversions is an example. This leads to atmospheric models underestimating, or even missing, these clouds and their radiative effects. Using radiosonde observations as training data, a machine learning model is used to improve the vertical detail of modelled profiles of temperature and specific humidity. In addition, a physics-informed machine learning model is developed and compared to the traditional approach; showing improvements in the cloud fraction profiles calculated from its predictions. The vertically enhanced profiles also improve the representation of layers of convective inhibition and anomalous refractivity gradients. This work facilitates targeted improvements to the representation of certain atmospheric processes without the burden of increased memory and computational cost from increasing vertical resolution throughout the whole model.

54 ENVIRONMENTAL SCIENCES↗

Explainable machine learning for incipient anomaly detection in compact molten salt heat exchanger with overlapping feature distributions

High-temperature molten salt-cooled reactors (MSCRs) are a promising next-generation nuclear technology option, offering efficient power conversion and inherent safety features. However, the reliability of these systems depends on the robust operation of heat exchangers (HXs), which are susceptible to failure due to temperature gradients and channel plugging caused by fluid freezing. Conventional monitoring methods, relying on inlet and outlet measurements, lack the spatial resolution needed to detect early-stage faults. We propose a novel design of a compact salt-to-salt matrix-type HX design consisting of interleaved arrays of parallel tubes, with integrated synthetic fiber optic distributed temperature sensing (DTS) to enable localized detection of incipient faults. To evaluate performance of this design, we generate high-fidelity synthetic data using heat transfer computational modeling to simulate channel plugging, and introduce sensor noise for realistic modeling of measurements. The dataset comprises of 97% normal operation and 3% anomaly cases, with each anomaly class representing 1% of the data. These early anomalies result in overlapping temperature profiles between normal and faulty channels, producing a non-separable dataset that challenges traditional classification techniques. We benchmark eight supervised machine learning (ML) models and demonstrate that XGBoost achieves the highest performance. To improve transparency, we develop an explainability framework combining Shapley values and partially ordered sets (POSETs) to quantify and structurally analyze feature importance. This approach identifies both dominant predictors and ambiguous feature relationships, enhancing trust and interpretability. Our results highlight the potential of combining DTS and explainable ML with intelligent feature selection to improve predictive maintenance and ensure operational resilience in advanced nuclear systems.

Prantikos, Konstantinos [Argonne National Laborato↗

OPF-Learn: An Open-Source Framework for Creating Representative AC Optimal Power Flow Datasets: Preprint

Increasing levels of renewable generation motivate a growing interest in data-driven approaches for AC optimal power flow (AC OPF) to manage uncertainty. However, a lack of disciplined dataset creation and benchmarking prohibits useful comparison between approaches in the literature. To instigate confidence, models must be able to reliably predict solutions across a wide range of operating conditions. This paper develops the OPF-Learn package for Julia and Python which uses a computationally efficient approach to create representative datasets that span a wide spectrum of the AC OPF feasible region. Load profiles are uniformly sampled from a convex set that contains the AC OPF feasible set. For each infeasible point found, the convex set is reduced using infeasibility certificates, found by utilizing properties of a relaxed formulation. The framework is shown to generate datasets which are more representative of the entire feasible space versus traditional techniques seen in the literature, improving machine learning model performance.

dataset↗

Performance on HPC Platforms Is Possible Without C++

Computing at large scales has become extremely challenging due to increasing heterogeneity in both hardware and software. More and more scientific workflows must tackle a range of scales and use machine learning and AI intertwined with more traditional numerical modeling methods, placing more demands on computational platforms. These constraints indicate a need to fundamentally rethink the way computational science is done and the tools that are needed to enable these complex workflows. The current set of C++-based solutions may not suffice, and relying exclusively upon C++ may not be the best option, especially because several newer languages and boutique solutions offer more robust design features to tackle the challenges of heterogeneity. In June 2023, we held a mini symposium that explored the use of newer languages and heterogeneity solutions that are not tied to C++ and that offer options beyond template metaprogramming and Parallel. For for performance and portability. In conclusion, we describe some of the presentations and discussion from the mini symposium in this article.

97 MATHEMATICS AND COMPUTING↗

Predicting battery capacity from impedance at varying temperature and state of charge using machine learning

Prediction of battery health from electrochemical impedance spectroscopy (EIS) data can enable rapid measurement of battery state in real-world applications without using additional sensors or time-consuming performance measurements. However, deconvoluting the effect of capacity, state of charge, and temperature on EIS response is complicated analytically. Here, various machine-learning models, such as linear, Gaussian process, random forest, and artificial neural network regression, are utilized to predict capacity from EIS using hundreds of capacity, direct current (DC) resistance, and EIS measurements recorded under varying conditions of health, temperature, and state of charge (SOC). Several feature extraction and selection methods from traditional electrochemical analysis and statistical modeling are explored using machine-learning pipelines. EIS data from just two frequencies can accurately predict capacity, and interrogation shows that the optimal set of frequencies is not usually intuitive. Best results are achieved with an ensemble model, which predicts battery capacity with a mean absolute error of 1.9% on data from unobserved cells.

25 ENERGY STORAGE↗

Popnet : computer vision based deep learning model for forecasting gridded population

Here, this study introduces Popnet, a deep learning model for forecasting 1 km-gridded populations, integrating U-Net, ConvLSTM, a Spatial Autocorrelation module and deep ensemble methods. Using spatial variables and population data from 2000 to 2020, Popnet predicts South Korea’s population trends by age groups (under 14, 15-64 and over 65) up to 2040. In validation, it outperforms traditional machine learning and state-of-the-art computer vision models. The output of this model discovered significant polarisation: population growth in urban areas, especially the capital region, and severe depopulation in rural areas. Popnet is a robust tool for offering significant insights to policymakers and related stakeholders about the detailed future population, which allows them to establish detailed, localised planning and resource allocations.

computer vision↗

Interpretable machine learning models classify minerals via spectroscopy

Developing methods to identify mineral species confidently and rapidly from Raman spectral analysis is critical to numerous fields. Traditionally, analysis relies on pattern matching the Raman spectrum of an unknown dataset with a supporting library of well-characterized spectral data, which may prove difficult for environmental samples that are poorly crystalline or phase mixtures. Here, we developed interpretable machine learning models that can classify uranium minerals by secondary oxyanion chemistry and other physicochemical properties based solely on Raman spectra. This new ML method produces a mineral profile of physical and chemical properties for an unknown sample and can rapidly classify or identify unknown minerals from Raman data, without the need for an exact pattern match in a spectral library. Training models are validated by 1. Strong correlation of high confidence model regions with published spectroscopic assignments and 2. Correct classification of a mineral not present in training data. Training data are from the Compendium of Uranium Raman and Infrared Experimental Spectra and available crystallographic information files within the open-source Smart Spectral Matching scientific framework. Physically meaningful classifier models can rapidly identify key structural and chemical information about unknown uranium minerals and the overall methodology is broadly applicable for mineral phases.

Machine learning↗