Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Machine Learning Models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25

On the Prediction of Aerosol-Cloud Interactions Within a Data-Driven Framework

Aerosol-cloud interactions (ACI) pose the largest uncertainty for climate projection. Among many challenges of understanding ACI, the question of whether ACI can be deterministically predicted has not been explicitly answered. Here we attempt to answer this question by predicting cloud droplet number concentration N c from aerosol number concentration N a and ambient conditions using a data-driven framework. We use aerosol properties, vertical velocity fluctuations, and meteorological states from the ACTIVATE field observations (2020–2022) as predictors to estimate N c . We show that the campaign-wide N c can be successfully predicted using machine learning models despite the strongly nonlinear and multi-scale nature of ACI. However, the observation-trained machine learning model fails to predict N c in individual cases while it successfully predicts N c of randomly selected data points that cover a broad spatiotemporal scale. This suggests that, within a data-driven framework, the N c prediction is uncertain at fine spatiotemporal scales.

54 ENVIRONMENTAL SCIENCES↗

Discovery of hydrogen storage molecules using large language models and machine learning

Accelerating the discovery of new molecules with targeted properties is a central challenge in molecular design. In this contribution, we present an AI-driven molecular discovery framework that integrates Large Language Models (LLMs) for generative molecular design with Machine Learning (ML)-based screening to identify novel Liquid Organic Hydrogen Carrier (LOHC) candidates. Using the developed framework, LOHC molecules were systematically generated, evaluated, and refined iteratively, combining LLM-guided molecular generation and ML-predicted hydrogenation enthalpies (Δ H ), under physicochemical property constraints such as optimal melting points (MP), desired hydrogen storage capacity (wt% H 2 ), and synthetic accessibility (SA) scores. This approach enabled the discovery of 42 new LOHC candidates in two distinct campaigns, one seeded with experimentally known and another with previously computationally identified LOHCs, respectively. Although we began with different numbers of starting molecules (31 vs . 7 seed molecules), both runs yielded a comparable number of viable candidates, suggesting an influence of chemically intuitive seed molecule selection for success. Selected LOHC molecules, such as 3-methyl pyridine, 1-ethylnapthalene, 1,1-diphenylethane, and benzofuran, were experimentally tested and compared with benchmark LOHCs (toluene and 9-ethylcarbazole) for hydrogenation using a series of commercial supported metal catalysts. The order of conversion into fully hydrogenated products at 200 °C was 3-methyl pyridine (100%) > 9-ethyl carbazole (86.4%) > 2,3-benzofuran (74%) > 1,1-diphenylethane (66.9%) > 1-ethylnapthalene (66.7%) > toluene (57%), further validating the AI-guided molecular design. This study demonstrates promise of LLM-driven molecular design in conjunction with ML-based screening for accelerated discovery and design of molecules.

Harb, Hassan [Argonne National Laboratory (ANL), A↗

XPlacer Machine Learning Guided Data Placement

XPlacer Machine Learning Guided Data Placement uses machine learning model to assist programmers to select the data placement advises for application running on HPC system with Nvidia GPUs. XPlacer Machine Learning Guided Data Placement consists three components: (1) 4 revised benchmarks from Rodinia benchmarks; (2) compiler script and execution script to generate variants, and collect training data for the machine learning model; (3) the scripts to post-process the collected data and generate the machine leaning model.

Liao, Chunhua↗

SOMAS: a platform for data-driven material discovery in redox flow battery development

Abstract Aqueous organic redox flow batteries offer an environmentally benign, tunable, and safe route to large-scale energy storage. The energy density is one of the key performance parameters of organic redox flow batteries, which critically depends on the solubility of the redox-active molecule in water. Prediction of aqueous solubility remains a challenge in chemistry. Recently, machine learning models have been developed for molecular properties prediction in chemistry and material science. The fidelity of a machine learning model critically depends on the diversity, accuracy, and abundancy of the training datasets. We build a comprehensive open access organic molecular database “Solubility of Organic Molecules in Aqueous Solution” (SOMAS) containing about 12,000 molecules that covers wider chemical and solubility regimes suitable for aqueous organic redox flow battery development efforts. In addition to experimental solubility, we also provide eight distinctive quantum descriptors including optimized geometry derived from high-throughput density functional theory calculations along with six molecular descriptors for each molecule. SOMAS builds a critical foundation for future efforts in artificial intelligence-based solubility prediction models.

25 ENERGY STORAGE↗

Evaluation of global terrestrial evapotranspiration using state-of-the-art approaches in remote sensing, machine learning and land surface modeling

Evapotranspiration (ET) is critical in linking global water, carbon and energy cycles. However, direct measurement of global terrestrial ET is not feasible. Here, we first reviewed the basic theory and state-of-the-art approaches for estimating global terrestrial ET, including remote-sensing-based physical models, machine-learning algorithms and land surface models (LSMs). We then utilized 4 remote-sensing-based physical models, 2 machine-learning algorithms and 14 LSMs to analyze the spatial and temporal variations in global terrestrial ET. The results showed that the ensemble means of annual global terrestrial ET estimated by these three categories of approaches agreed well, with values ranging from 589.6 mm/yr (6.56×10^4 cu.km/yr) to 617.1 mm/yr (6.87×10^4 cu.km/yr). For the period from 1982 to 2011, both the ensembles of remote-sensing-based physical models and machine-learning algorithms suggested increasing trends in global terrestrial ET (0.62 mm/sq.yr with a significance level of p<0.05 and 0.38 mm yr−2 with a significance level of p<0.05, respectively). In contrast, the ensemble mean of the LSMs showed no statistically significant change (0.23 mm/sq.yr, p>0.05), although many of the individual LSMs reproduced an increasing trend. Nevertheless, all 20 models used in this study showed that anthropogenic Earth greening had a positive role in increasing terrestrial ET. The concurrent small interannual variability, i.e., relative stability, found in all estimates of global terrestrial ET, suggests that a potential planetary boundary exists in regulating global terrestrial ET, with the value of this boundary being around 600 mm/yr. Uncertainties among approaches were identified in specific regions, particularly in the Amazon Basin and arid/semiarid regions. Improvements in parameterizing water stress and canopy dynamics, the utilization of new available satellite retrievals and deep-learning methods, and model–data fusion will advance our predictive understanding of global terrestrial ET.

surface modeling↗

Experimental Setup and Learning-Based AI Model for Developing Accurate PV Inverter Models [Slides]

The integration of power electronics-based interfaces presents challenges due to the absence of detailed models and the high computational complexity. Generic models used in system studies lack accuracy in capturing converter dynamics. This paper proposes a data-driven approach developed from experimental setup data. This approach enhances accuracy in photovoltaic inverter modeling. We used two types of PV inverters in the experiment. The recorded experimental data undergo processing through a machine learning model. Results from the model trained through machine learning is also presented.

14 SOLAR ENERGY↗

The value of human data annotation for machine learning based anomaly detection in environmental systems

Anomaly detection is the process of identifying unexpected data samples in datasets. Automated anomaly detection is either performed using supervised machine learning models, which require a labelled dataset for their calibration, or unsupervised models, which do not require labels. While academic research has produced a vast array of tools and machine learning models for automated anomaly detection, the research community focused on environmental systems still lacks a comparative analysis that is simultaneously comprehensive, objective, and systematic. This knowledge gap is addressed for the first time in this study, where 15 different supervised and unsupervised anomaly detection models are evaluated on 5 different environmental datasets from engineered and natural aquatic systems. To this end, anomaly detection performance, labelling efforts, as well as the impact of model and algorithm tuning are taken into account. As a result, our analysis reveals the relative strengths and weaknesses of the different approaches in an objective manner without bias for any particular paradigm in machine learning. Most importantly, our results show that expert-based data annotation is extremely valuable for anomaly detection based on machine learning.

54 ENVIRONMENTAL SCIENCES↗

High–Resolution Maps of Near–Surface Permafrost for Three Watersheds on the Seward Peninsula, Alaska Derived From Machine Learning

Permafrost soils are a critical component of the global carbon cycle and are locally important because they regulate the hydrologic flux from uplands to rivers. Furthermore, degradation of permafrost soils causes land surface subsidence, damaging infrastructure that is crucial for local communities. Regional and hemispherical maps of permafrost are too coarse to resolve distributions at a scale relevant to assessments of infrastructure stability or to illuminate geomorphic impacts of permafrost thaw. Here we train machine learning models to generate meter–scale maps of near–surface permafrost for three watersheds in the discontinuous permafrost region. The models were trained using ground truth determinations of near–surface permafrost presence from measurements of soil temperature and electrical resistivity. We trained three classifiers: extremely randomized trees (ERTr), support vector machines (SVM), and an artificial neural network (ANN). Model uncertainty was determined using k–fold cross validation, and the modeled extents of near–surface permafrost were compared to the observed extents at each site. At–a–site near–surface permafrost distributions predicted by the ERTr produced the highest accuracy (70%–90%). However, the transferability of the ERTr to the sites outside of the training data set was poor, with accuracies ranging from 50% to 77%. The SVM and ANN models had lower accuracies for at–a–site prediction (70%–83%), yet they had greater accuracy when transferred to the non–training site (62%–78%). These models demonstrate the potential for integrating high–resolution spatial data and machine learning models to develop maps of near–surface permafrost extent at resolutions fine enough to assess infrastructure vulnerability and landscape morphology influenced by permafrost thaw.

54 ENVIRONMENTAL SCIENCES↗

In-process monitoring and prediction of droplet quality in droplet-on-demand liquid metal jetting additive manufacturing using machine learning

Abstract In droplet-on-demand liquid metal jetting (DoD-LMJ) additive manufacturing, complex physical interactions govern the droplet characteristics, such as size, velocity, and shape. These droplet characteristics, in turn, determine the functional quality of the printed parts. Hence, to ensure repeatable and reliable part quality it is necessary to monitor and control the droplet characteristics. Existing approaches for in-situ monitoring of droplet behavior in DoD-LMJ rely on high-speed imaging sensors. The resulting high volume of droplet images acquired is computationally demanding to analyze and hinders real-time control of the process. To overcome this challenge, the objective of this work is to use time series data acquired from an in-process millimeter-wave sensor for predicting the size, velocity, and shape characteristics of droplets in DoD-LMJ process. As opposed to high-speed imaging, this sensor produces data-efficient time series signatures that allows rapid, real-time process monitoring. We devise machine learning models that use the millimeter-wave sensor data to predict the droplet characteristics. Specifically, we developed multilayer perceptron-based non-linear autoregressive models to predict the size and velocity of droplets. Likewise, a supervised machine learning model was trained to classify the droplet shape using the frequency spectrum information contained in the millimeter-wave sensor signatures. High-speed imaging data served as ground truth for model training and validation. These models captured the droplet characteristics with a statistical fidelity exceeding 90%, and vastly outperformed conventional statistical modeling approaches. Thus, this work achieves a practically viable sensing approach for real-time quality monitoring of the DoD-LMJ process, in lieu of the existing data-intensive image-based techniques.

Gaikwad, Aniruddha (ORCID:0000000285642621)↗

A Fusion of Geothermal and InSAR Data with Machine Learning for Enhanced Deformation Forecasting at the Geysers

The Geysers geothermal field in California is experiencing land subsidence due to the seismic and geothermal activities taking place. This poses a risk not only to the underlying infrastructure but also to the groundwater level which would reduce the water availability for the local community. Because of this, it is crucial to monitor and assess the surface deformation occurring and adjust geothermal operations accordingly. In this study, we examine the correlation between the geothermal injection and production rates as well as the seismic activity in the area, and we show the high correlation between the injection rate and the number of earthquakes. This motivates the use of this data in a machine learning model that would predict future deformation maps. First, we build a model that uses interferometric synthetic aperture radar (InSAR) images that have been processed and turned into a deformation time series using LiCSBAS, an open-source InSAR time series package, and evaluate the performance against a linear baseline model. The model includes both convolutional neural network (CNN) layers as well as long short-term memory (LSTM) layers and is able to improve upon the baseline model based on a mean squared error metric. Then, after getting preprocessed, we incorporate the geothermal data by adding them as additional inputs to the model. This new model was able to outperform both the baseline and the previous version of the model that uses only InSAR data, motivating the use of machine learning models as well as geothermal data in assessing and predicting future deformation at The Geysers as part of hazard mitigation models which would then be used as fundamental tools for informed decision making when it comes to adjusting geothermal operations.

Yazbeck, Joe (ORCID:0000000302235260)↗

Physics-informed machine learning of the correlation functions in bulk fluids

The Ornstein–Zernike (OZ) equation is the fundamental equation for pair correlation function computations in the modern integral equation theory for liquids. In this work, machine learning models, notably physics-informed neural networks and physics-informed neural operator networks, are explored to solve the OZ equation. The physics-informed machine learning models demonstrate great accuracy and high efficiency in solving the forward and inverse OZ problems of various bulk fluids. The results highlight the significant potential of physics-informed machine learning for applications in thermodynamic state theory.

97 MATHEMATICS AND COMPUTING↗

Residential Demand Flexibility: Modeling Occupant Behavior using Sociodemographic Predictors

Demand flexibility (DF) has the potential to increase the saturation of renewables in the grid and reduce operating costs for both utilities and customers. However, less than 8% of U.S. residential electric customers are enrolled in DF programs. A major research gap on this topic is an uneven understanding of behavioral drivers of electricity use and DF program participation at the household level. In this study, we employ machine learning models to predict residential occupant behavior in activities relevant to DF. We model occupants' extensive decisions (i.e., choice of action) and intensive behaviors (i.e., amount of time spent) during peak and off-peak time periods using the publicly available American Time Use Survey, which includes activities data for approximately 200,000 respondents. In our machine learning models, predictions for both extensive and intensive behavior fell within a +/-20% error margin at the aggregate level. We identify 13 key sociodemographic predictors of DF-related intensive behavior using LASSO inference and beta coefficient ranking. However, these top predictors differ by activity, suggesting potential scope for differential user targeting for DF events and technologies during program design. This work also contributes to understanding when and who might adopt these DF technologies based on their daily routine activities.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

A machine learning degradation model for electrochemical capacitors operated at high temperature

Electrochemical capacitors (ECs) have only recently been considered as an alternative power source for telemetry sensors of drilling equipment for geothermal or oil and gas exploration. The lifecycle analysis and modelling of ECs is underrepresented in literature in comparison to other storage devices e.g. Li-ion batteries. This paper investigates the degradation of ECs when cycled outside the manufacturer-specified operating temperature envelope and proposes a machine learning-based approach for modelling the degradation. Experimental results show that end of life, defined as a 30% decrease in capacitance, occurs at 1,000 cycles when the environmental temperature exceeds the maximum operating temperature by 30%. The life-cycle test data is then used as an input to a Gaussian process regression (GPR) algorithm to predict the capacitance fade trend. The GPR is validated on a total of nine commercial cells from two different manufacturers, achieving an average root mean squared percent error of less than 2% and a mean calibration score of 93% when referenced to a 95% confidence interval. The model can be utilized to determine the EC degradation rate at a range of operating temperature values.

42 ENGINEERING↗

Predicting U.S. federal fleet electric vehicle charging patterns using internal combustion engine vehicle fueling transaction statistics

Utilizing fueling transactions from internal combustion engine vehicles (ICEVs), the authors estimated how frequently midday public charging would be required for U.S. federal fleet battery electric vehicles (BEVs). Fueling transaction summary statistics are more widely available than trip-level telematics data, making this methodology more accessible and transferable to other researchers and fleet managers considering BEV replacements. For example, readers can easily apply a linear model using only the count of back-to-back fueling events at gas stations over 57 straight-line miles apart to predict days exceeding range. This linear regression predicted binned days exceeding 250 miles at 80% accuracy on a hold-out test set from the same fleet as the training data and 66 % accuracy on a new fleet displaying different driving behaviors. The authors additionally provide linear equations for days exceeding 200 and 300 miles as alternative range estimates to account for differences in BEV range and temperature impacts. Beyond the single-feature linear models which readers can apply, the authors tuned and trained other machine learning models on a variety of fueling transaction statistics including consecutive transaction distances, transaction distance from garage, estimated miles traveled from fuel economy and fuel quantity, and transaction periodicity. Utilizing a subset of 1678 light-duty federal fleet vehicles which contained daily vehicle miles traveled (VMT) in addition to fueling statistics, the authors determined which fueling transaction statistics were most relevant in predicting driving days exceeding 250 miles (an approximation of BEV rated driving range). In support of the U.S. federal fleet transition to zero-emission vehicles (ZEVs), the authors used these statistics and machine learning models to predict the frequency of BEV midday charging. After training models on the subset with VMT, the authors predicted days exceeding rated range for 112,902 light-duty vehicles operating in similar circumstances in the federal fleet using a Support Vector Regressor (SVR). In conclusion, they then used the projections as part of the ZEV Planning and Charging (ZPAC) tool to identify optimal candidates for BEVs for the federal fleet. An anonymized version of ZPAC is included in the supplementary materials.

25 ENERGY STORAGE↗

Nearest-Neighbor Machine Learning Feature Selection for Interpretation of Microbial Molecular Signatures from Isotope Ratio Mass Spectrometry Data

Mass spectrometry (MS) promises to be a powerful tool for potential biosignature detection during astrobiological missions on ocean worlds in our solar system. Accurate and generalizable machine learning methods could enhance science return on investment by predicting seawater chemistry and classifying isotopic biosignatures, either as a signature consistent with microbial life (biotic) or as a novelty (unclassified/unique). However, machine learning models are likely to be complex and involve interactions between MS features, making biosignatures difficult to interpret. Feature selection methods provide biological and chemical context that help interpret the mechanisms of machine learning models, but these methods also need the ability to detect complex interactions. Previously, we developed a machine learning feature selection algorithm called nearest-neighbor projected distance regression (NPDR) that has the ability to identify important model features that involve complex interactions and automatically reduce correlation and the dimensionality in a high-dimensional variable space. The standard distance metrics used in NPDR – Manhattan and Euclidean – assume the multivariate data are isotropic, which is often violated in real data due to differences in the covariance between variables. Thus, we extend NPDR to include a random forest distance, and other anisotropic distance metrics, for computing nearest neighbors. We also augment the isotope-ratio MS data with time-series features from the raw MS signal to improve biotic classification. We test NPDR on our novel experimental ocean world seawater analog MS data. We measure isotope fractionations of volatile CO 2 that could be measured in exospheres or plumes. Samples include baseline abiotic conditions using a range of possible seawater chemistry consistent with Europa and Enceladus, and biotic samples that include microbes in these seawaters. We use penalized NPDR with random forest proximity to identify interpretable microbial molecular signatures. We compare features with random forest importance, and we train a classifier that discriminates between biotic and abiotic samples with high accuracy. These ML-trained ocean-world analog MS data could be used to assist in identifying biosignatures during future missions.

geochemistry↗

Evaluation of global terrestrial evapotranspiration using state-of-the-art approaches in remote sensing, machine learning and land surface modeling

Abstract. Evapotranspiration (ET) is critical in linking global water, carbon and energy cycles. However, direct measurement of global terrestrial ET is not feasible. Here, we first reviewed the basic theory and state-of-the-art approaches for estimating global terrestrial ET, including remote-sensing-based physical models, machine-learning algorithms and land surface models (LSMs). We then utilized 4 remote-sensing-based physical models, 2 machine-learning algorithms and 14 LSMs to analyze the spatial and temporal variations in global terrestrial ET. The results showed that the ensemble means of annual global terrestrial ET estimated by these three categories of approaches agreed well, with values ranging from 589.6 mm yr-1 (6.56×104 km3 yr-1) to 617.1 mm yr-1 (6.87×10 4 km3 yr -1 ). For the period from 1982 to 2011, both the ensembles of remote-sensing-based physical models and machine-learning algorithms suggested increasing trends in global terrestrial ET (0.62 mm yr -2 with a significance level of p<0.05 and 0.38 mm yr -2 with a significance level of p<0.05, respectively). In contrast, the ensemble mean of the LSMs showed no statistically significant change (0.23 mm yr -2 , p>0.05), although many of the individual LSMs reproduced an increasing trend. Nevertheless, all 20 models used in this study showed that anthropogenic Earth greening had a positive role in increasing terrestrial ET. The concurrent small interannual variability, i.e., relative stability, found in all estimates of global terrestrial ET, suggests that a potential planetary boundary exists in regulating global terrestrial ET, with the value of this boundary being around 600 mm yr -1 . Uncertainties among approaches were identified in specific regions, particularly in the Amazon Basin and arid/semiarid regions. Improvements in parameterizing water stress and canopy dynamics, the utilization of new available satellite retrievals and deep-learning methods, and model–data fusion will advance our predictive understanding of global terrestrial ET.

58 GEOSCIENCES↗

ChemoGraph: Interactive Visual Exploration of the Chemical Space

Exploratory analysis of the chemical space is an important task in the field of cheminformatics. For example, in drug discovery research, chemists investigate sets of thousands of chemical compounds in order to identify novel yet structurally similar synthetic compounds to replace natural products. Manually exploring the chemical space inhabited by all possible molecules and chemical compounds is impractical, and therefore presents a challenge. To fill this gap, we present ChemoGraph, a novel visual analytics technique for interactively exploring related chemicals. In ChemoGraph, we formalize a chemical space as a hypergraph and apply novel machine learning models to compute related chemical compounds. It uses a database to find related compounds from a known space and a machine learning model to generate new ones, which helps enlarge the known space. Moreover, ChemoGraph highlights interactive features that support users in viewing, comparing, and organizing computationally identified related chemicals. With a drug discovery usage scenario and initial expert feedback from a case study, we demonstrate the usefulness of ChemoGraph.

chemical space exploration↗

Multifidelity computing for coupling full and reduced order models

Hybrid physics-machine learning models are increasingly being used in simulations of transport processes. Many complex multiphysics systems relevant to scientific and engineering applications include multiple spatiotemporal scales and comprise a multifidelity problem sharing an interface between various formulations or heterogeneous computational entities. To this end, we present a robust hybrid analysis and modeling approach combining a physics-based full order model (FOM) and a data-driven reduced order model (ROM) to form the building blocks of an integrated approach among mixed fidelity descriptions toward predictive digital twin technologies. At the interface, we introduce a long short-term memory network to bridge these high and low-fidelity models in various forms of interfacial error correction or prolongation. The proposed interface learning approaches are tested as a new way to address ROM-FOM coupling problems solving nonlinear advection-diffusion flow situations with a bifidelity setup that captures the essence of a broad class of transport processes.

59 BASIC BIOLOGICAL SCIENCES↗