Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine learning for science”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Unsupervised Anomaly Detection in High-Dimensional Flight Data Using Convolutional Variational Auto-Encoder

The modern National Airspace System (NAS) is an extremely safe system and the aviation industry has experienced a steady decrease in fatalities over the years. This can be attributed to both improved flight critical systems with redundant hardware and software protections, as well as an increased focus on active monitoring and response to real time and historically identified vulnerabilities by implementing more resilient procedures and protocols. The main approach for identifying vulnerabilities in operations leverages domain expertise using knowledge about how the system should behave within the expected tolerances to known safety margins. This approach works well when the system has a well-defined operating condition. However, the operations in the NAS can be highly complex with various nuances that render it difficult to clearly pre-define all known safety vulnerabilities. With the advancement of data science and machine learning techniques, the potential to automatically identify emerging vulnerabilities in the observed operations has become more practical in recent years. The state-of-the-art anomaly detection approaches in aerospace data usually rely on supervised or semi-supervised learning. However, in many real-world problems such as flight safety, creating labels for the data requires huge amount of effort and is largely impractical. To address this challenge, we developed a Convolutional Variational Auto-Encoder (CVAE), which is an unsupervised learning approach for anomaly detection in high-dimensional heterogeneous time-series data. We validate performance of CVAE compared to the state-of-the-art supervised learning approach as well as unsupervised clustering-based approach using KMeans++ and kernel-based approach using One-Class Support Vector Machine (OC-SVM) on Yahoo!'s benchmark time series anomaly detection data. Finally, we showcase performance of CVAE on a case study of identifying anomalies in the first 60 seconds of commercial flights' take-offs using Flight Operational Quality Assurance (FOQA) data.

Memarzadeh, Milad↗

Unsupervised Anomaly Detection in High-Dimensional Flight Data Using Convolutional Variational Auto-Encoder

The modern National Airspace System (NAS) is an extremely safe system. The industry has experienced a steady decrease in fatalities over the years. This can be contributed to both improved flight critical systems with redundant hardware and software protections as well as an increased focus on active monitoring and response to real time and historically identified vulnerabilities by implementing more resilient procedures and protocols. The main practice for identifying vulnerabilities in operations leverages domain expertise using knowledge about how the system should behave with the expected tolerances to known safety margins. This approach works well when the system has a well-defined operating condition. However, the operations in the NAS can be highly complex with various nuances that render it difficult to clearly pre-define all known safety vulnerabilities. With the advancement of data science and machine learning techniques, the potential to automatically identify emerging vulnerabilities in the observed operations has become more practical in recent years. The state-of-the-art anomaly detection approaches in aerospace data usually rely on supervised or semi-supervised learning. However, in many real-world problems such as flight safety creating labels for the data requires huge amount of efforts and is largely expensive. As a result, in this article, we develop a Convolutional Variational Auto-Encoder (CVAE), an unsupervised learning approach for anomaly detection in high-dimensional heterogeneous time-series data. We validate performance of CVAE compared to the state-of-the-art supervised learning approach (as an upper bound) as well as an supervised clustering based on K-Means (as a lower bound) on Yahoo!'s benchmark time series anomaly detection data. Finally, we showcase performance of CVAE on a case study of identifying anomalies in the first 60 seconds of commercial flights' take-offs using Flight Operational Quality Assurance (FOQA) data.

Milad Memarzadeh↗

Super-Resolution from Space: Using MERRA-2 and MAIAC Satellite Imagery to Produce Daily Continuous 1 km PM2.5 Estimates

PM2.5 measurements from ground stations are the gold standard when available, but the expense and coverage of such stations limits widespread monitoring. Having accurate PM2.5 estimates outside the range of these stations is important for monitoring this crucial aspect of air quality. The goal of this project is to produce daily 1 km continuous PM2.5 estimates for the contiguous US relying primarily on satellite-derived data sources. This is important because models based on such data can be more easily expanded outside the study area and produce global estimates as well. The temporal availability of such data products is often weekly/daily, unlike land-use products with are often available at a yearly or worse temporal resolution. To achieve our goal, we use a couple of different deep neural network architectures to produce PM2.5 measurements at 10 km and 1 km resolution. We use two model architectures, a UNET-like model and a GAN-based model. We train both models using MERRA-2 data and MAIAC AOD data scaled to 10 km and 1 km for the two different prediction resolutions. MERRA-2 imagery is data rich with a wide range of geospatial variables at 50 km and has long historical availability (beginning in 1980). We’re also using higher spatial resolution MAIAC data at 1 km to provide finer resolution spatial context. This essentially leverages the spatial resolution of MAIAC data and the “wider” information of MERRA-2 data to predict PM2.5. For the target data we’re using a modeled 1 km PM2.5 dataset produced by Harvard to pre-train our models and then fine-tune our models using ground station measurements. Not only are our results comparable with the performance of the Harvard dataset, but can be generalized to any area or time where MERRA-2 and MAIAC data is available.

satellite imagery↗

Squeezing Every Last 'Bit' of Information from Enceladus Mass Spectrometry

Potential opportunities to return to Enceladus in Discovery and Flagship class missions inspire development of next-generation instruments and creative approaches to sample collection, sample analysis, and data analysis and transmission strategies. Mass spectrometers (MS) are ideally suited to future Enceladus missions due to their analytical power in identifying a range of molecular and ionic compositions – including complex organics – and potentially astrobiologically-important features such as isotope ratios, chirality, and enantiomeric excess. However, long communication delays from Enceladus and limited bandwidth limits the data transmission from these higher-data-volume instruments, likely delaying mission-related response to new data. We explore the utility of data science and machine learning (ML) on isotope ratio (IR)MS data collected from laboratory analogs of Enceladus to: 1) process data quickly for rapid ground-based analyses, 2) understand if compositional and biosignature information could be extracted from IRMS data, and 3) evaluate whether onboard ML techniques could improve sample analysis, cadence, and transmission prioritization. Laboratory analogs analyzed isotopes of volatile CO2 that interacted with seawaters of varying composition, and include both abiotic and biotic (microbially-influenced) experiments. Enceladus’s alkaline oceans promote speciation of carbon into multiple forms (e.g., H2CO3 / CO2, HCO3-, and CO32-), each of which could be isotopically fractionated by abiotic or biotic reactions. Large (>2‰) changes in carbon isotopes (δ13C) are observed from some biotic experiments inoculated with complex microbial ecosystems relative to the abiotic seawaters. ML training and classification suggests that microbial samples can be distinguished from abiotic samples, yet that a broad range of microbial experiments are necessary to train ML models to cover a range of complexities including disequilibria, and isotopic and compositional fractionation.

geochemistry↗

Deep Nonparametric Estimation of Operators between Infinite Dimensional Spaces

Learning operators between infinitely dimensional spaces is an important learning task arising in machine learning, imaging science, mathematical modeling and simulations, etc. This paper studies the nonparametric estimation of Lipschitz operators using deep neural networks. Non-asymptotic upper bounds are derived for the generalization error of the empirical risk minimizer over a properly chosen network class. Under the assumption that the target operator exhibits a low dimensional structure, our error bounds decay as the training sample size increases, with an attractive fast rate depending on the intrinsic dimension in our estimation. Our assumptions cover most scenarios in real applications and our results give rise to fast rates by exploiting low dimensional structures of data in operator estimation. We also investigate the influence of network structures (e.g., network width, depth, and sparsity) on the generalization error of the neural network estimator and propose a general suggestion on the choice of network structures to maximize the learning efficiency quantitatively.

97 MATHEMATICS AND COMPUTING↗

Final Report (October 2024): University of Tennessee, Knoxville (UTK) contribution to: FusMatML: Machine Learning Atomistic Modeling for Fusion Materials Collaborative Project led by Dr. Aidan Thompson, Sandia National Laboratory

The rapid growth of the field of Machine Learning Inter-Atomic Potentials (MLIAP) has lead to a profusion of methods, all of which have some similarity to each other, but each also restricted to particular design choices, often arrived at in a rather ad hoc fashion. Beyond anecdotal evidence, and some benchmarking studies on specific problems, little progress has been made in developing design principles for MLIAPs. The goal of this project is to use machine learning, data science, and uncertainty quantification methods to optimize the design choices for MLIAP.

Density functional theory, Helium and Hydrogen↗

Advancing Geothermal Research: Fiscal Year 2025 Accomplishments Report

This is a summary of geothermal work done at the National Renewable Energy Laboratory (NREL) in Fiscal Year 2025. This year brought increased attention to the geothermal industry and NREL's geothermal research portfolio. With more than 70 active projects, NREL research spanned the areas of resource exploration and characterization; conventional and next-generation geothermal technologies; subsurface thermal energy storage; heating and cooling; co-production of geothermal with critical minerals and oil and gas; modeling and analysis leveraging expertise in data science and machine learning; and more.

15 GEOTHERMAL ENERGY↗

Utilizing Gamma Signals to Find Optimal Uranium Wells

Uranium is a very important resource when using nuclear power. Only a small fraction of the uranium we use in the US is domestically sourced. Our goal is to effectively find and mine uranium in a way that is generally accurate and not difficult. Using computer science and machine learning, we want to automate a reasoning system that geologists use to analyze where the uranium ore bodies are. Presenting a general explanation of the goals and impacts of this project for the High School Intern showcase.

58 - GEOSCIENCES↗

Advancing Geothermal Research: Fiscal Year 2025 Accomplishments Report

This is a summary of geothermal work done at the National Laboratory of the Rockies in Fiscal Year 2025. This year brought increased attention to the geothermal industry and NLR's geothermal research portfolio. With more than 70 active projects, NLR research spanned the areas of resource exploration and characterization; conventional and next-generation geothermal technologies; subsurface thermal energy storage; heating and cooling; co-production of geothermal with critical minerals and oil and gas; modeling and analysis leveraging expertise in data science and machine learning; and more.

15 GEOTHERMAL ENERGY↗

Science Autonomy for Ocean Worlds Astrobiology: A Perspective

Astrobiology missions to ocean worlds in our solar system must overcome both scientific and technological challenges due to extreme temperature and radiation conditions, long communication times, and limited bandwidth. While such tools could not replace ground-based analysis by science and engineering teams, machine learning algorithms could enhance the science return of these missions through development of autonomous science capabilities. Examples of science autonomy include onboard data analysis and subsequent instrument optimization, data prioritization (for transmission), and real-time decision-making based on data analysis. Similar advances could be made to develop streamlined data processing software for rapid ground-based analyses. Here we discuss several ways machine learning and autonomy could be used for astrobiology missions, including landing site selection, prioritization and targeting of samples, classification of “features” (e.g., proposed biosignatures) and novelties (uncharacterized, “new” features, which may be of most interest to agnostic astrobiological investigations), and data transmission.

ocean worlds↗

MIDAS: Modeling Individual Differences using Advanced Statistics

This research explores novel methods for extracting relevant information from EEG data to characterize individual differences in cognitive processing. Our approach combines expertise in machine learning, statistics, and cognitive science, advancing the state-of-the art in all three domains. Specifically, by using cognitive science expertise to interpret results and inform algorithm development, we have developed a generalizable and interpretable machine learning method that can accurately predict individual differences in cognition. The output of the machine learning method revealed surprising features of the EEG data that, when interpreted by the cognitive science experts, provided novel insights to the underlying cognitive task. Additionally, the outputs of the statistical methods show promise as a principled approach to quickly find regions within the EEG data where individual differences lie, thereby supporting cognitive science analysis and informing machine learning models. This work lays methodological ground work for applying the large body of cognitive science literature on individual differences to high consequence mission applications.

97 MATHEMATICS AND COMPUTING↗

The Machine Learning Showroom: Presentation to OCIO Data Science Summit

Artificial Intelligence/Machine Learning (AI/ML) has become an indispensable tool for descriptive, predictive and prescriptive analytics. Demand for AI/ML models at NASA is outpacing Data Scientist staff. The AI/ML Showroom is an effort to empower NASA professionals to evaluate AI/ML solutions for their problems in a scalable self-help manner, relying on coding examples, reference use cases, digital assistant guides, jam sessions, video training, and pre-configured cloud resources.

machine learning↗

From the Knowledge-based Digital Platform (KbDP) Concept for Advanced Air Mobility Research to a Preliminary Prototype

Advanced Air Mobility (AAM) encompasses a range of innovative operational and technological changes to aviation (electric aircraft, increasingly automated aircraft, increasingly automated airspace operations, etc.) that are transforming aviation’s role in everyday movement of people and goods. There are multiple associated concepts and use cases for AAM, all interrelated, including small Unmanned Aircraft System (UAS) Traffic Management (UTM), Upper-Class E Traffic Management (ETM), Extensible Traffic Management (xTM), Regional Air Mobility (RAM), and Urban Air Mobility (UAM). These AAM operations must integrate with traditional Air Traffic Management (ATM) operations, as well as non-aviation modes of transportation and logistics. National Aeronautics and Space Administration (NASA) is spearheading an innovative digital engineering approach to integrate, communicate, and facilitate the research of multi-modal transportation systems. The Knowledge-based Digital Platform (KbDP) is a concept being developed that ties the workflows of Project Managers (PM), Principal Investigators (PI), and System Engineers together across organizational boundaries. It does so through the management of an information database defined by mathematical, data science, and system engineering principles. Machine Learning (ML) algorithms play a key role in this concept by extracting meaningful knowledge from the information database, which the human user leverages to greatly improve the efficiency and effectiveness of their research. Expected benefits of this concept include improved technology transfers from research to production, improved research portfolio investments, and research outcomes that are more integrated with all aspects of the multi-modal transportation problem. The preliminary KbDP prototype has been realized using UAM as a pathfinder use case and developed by a team of system engineer, software developer, data scientist, and interns.

Systems Engineering↗

Learning macroscopic internal variables and history dependence from microscopic models

This paper concerns the study of history dependent phenomena in heterogeneous materials in a two-scale setting where the material is specified at a fine microscopic scale of heterogeneities that is much smaller than the coarse macroscopic scale of application. Here, we specifically study a polycrystalline medium where each grain is governed by crystal plasticity while the solid is subjected to macroscopic dynamic loads. The theory of homogenization allows us to solve the macroscale problem directly with a constitutive relation that is defined implicitly by the solution of the microscale problem. However, the homogenization leads to a highly complex history dependence at the macroscale, one that can be quite different from that at the microscale. In this paper, we examine the use of machine-learning, and especially deep neural networks, to harness data generated by repeatedly solving the finer scale model to: (i) gain insights into the history dependence and the macroscopic internal variables that govern the overall response; and (ii) to create a computationally efficient surrogate of its solution operator, that can directly be used at the coarser scale with no further modeling. We do so by introducing a recurrent neural operator (RNO), and show that: (i) the architecture and the learned internal variables can provide insight into the physics of the macroscopic problem; and (ii) that the RNO can provide multiscale, specifically FE 2 , accuracy at a cost comparable to a conventional empirical constitutive relation.

36 MATERIALS SCIENCE↗

Cu–Ni Oxidation Mechanism Unveiled: A Machine Learning-Accelerated First-Principles and in Situ TEM Study

Here, the development of accurate methods for determining how alloy surfaces spontaneously restructure under reactive and corrosive environments is a key, long-standing, grand challenge in materials science. Using machine learning-accelerated density functional theory and rare-event methods, in conjunction with in situ environmental transmission electron microscopy (ETEM), we examine the interplay between surface reconstructions and preferential segregation tendencies of CuNi(100) surfaces under oxidation conditions. Our modeling approach predicts that oxygen-induced Ni segregation in CuNi alloys favors Cu(100)-O c(2 × 2) reconstruction and destabilizes the Cu(100)-O (2√2 × √2)R45° missing row reconstruction (MRR). In situ ETEM experiments validate these predictions and show Ni segregation followed by NiO nucleation and growth in regions without MRR, with secondary nucleation and growth of Cu 2 O in MRR regions. Our approach based on combining disparate computational components and in situ ETEM provides a holistic description of the oxidation mechanism in CuNi, which applies to other alloy systems.

36 MATERIALS SCIENCE↗

Special Issue: Geostatistics and Machine Learning

Abstract Recent years have seen a steady growth in the number of papers that apply machine learning methods to problems in the earth sciences. Although they have different origins, machine learning and geostatistics share concepts and methods. For example, the kriging formalism can be cast in the machine learning framework of Gaussian process regression. Machine learning, with its focus on algorithms and ability to seek, identify, and exploit hidden structures in big data sets, is providing new tools for exploration and prediction in the earth sciences. Geostatistics, on the other hand, offers interpretable models of spatial (and spatiotemporal) dependence. This special issue on Geostatistics and Machine Learning aims to investigate applications of machine learning methods as well as hybrid approaches combining machine learning and geostatistics which advance our understanding and predictive ability of spatial processes.

58 GEOSCIENCES↗

MS25: Materials Science-Focused Benchmark Data Set for Machine Learning Interatomic Potentials

Here, we present MS25, a benchmark data set for evaluating machine learning interatomic potentials (MLIPs) across diverse materials-relevant systems including MgO surfaces, liquid water, zeolites, a catalytic Pt surface reaction, high-entropy alloys (HEAs), and disordered Zr-oxides. Five MLIP architectures (MACE, NequIP, Allegro, MTP, and Torch-ANI) are trained and tested, focusing not only on traditional metrics (energies, forces, and stresses) but also explicitly validating derived physical observables such as lattice constants, volumes, and reaction barriers. We find that most models reach comparable accuracy on standard error metrics across the simple systems, although equivariant MLIPs offer 1.5–2× improvements over nonequivariant MLIPs in energy and force error for structurally complex or compositionally disordered environments such as HEAs and Zr–O systems. Our analysis highlights that low errors in energy and force predictions do not guarantee reliable observables, emphasizing the necessity of explicit validation. We demonstrate limitations in cross-framework transferability, as models trained on one zeolite framework (CHA) fail to reliably generalize to predictions of structurally distinct frameworks (e.g., MFI). Size-extensive tests show some dependence on system size for MgO, resulting from forced periodicity. The HEA and Zr–O data sets are identified as challenging tests for future benchmarks and MLIP model architecture developments as they show significant differentiation in error between MLIP architectures and are still relatively difficult at 1000 training images. Moving forward, we recommend that benchmarking efforts shift their focus from marginal accuracy improvements in energy and force errors toward identifying and understanding model failure modes, rigorously assessing transferability, and evaluating how their errors affect observable predictions. For researchers looking to choose an MLIP architecture, we suggest selecting equivariant MLIP architectures if the complexity of the system is a challenge. For simple materials problems, auxiliary features such as integration with molecular dynamics engines, trade-offs between computational data set generation cost vs MLIP inference speed, and framework integration may play a more important decision factor than small differences in error metrics that are unlikely to matter for production-level research.

chemical structure↗