Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “empirical machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Physics-informed hybrid modeling methodology for building infiltration

Infiltration is responsible for one-third to one-half of the space conditioning load of a typical residential home, but the modeling of infiltration for building energy modeling is either represented by over-simplified equations or dependent on over-generalized rules of thumb. Here, this paper develops a physics-informed data-driven methodology for modeling infiltration using building-specific empirical measurements. The developed hybrid methodology combines machine-learning categorization and grey-box sub-modeling to improve the accuracy and generalization of commonly used grey-box infiltration models. The developed methodology excels at predicting infiltration by improving the ability to predict infiltration under unseen environmental conditions using machine learning algorithms with physical significance. In a case study conducted using the iUnit, a modular studio apartment experimental test facility located at the National Renewable Energy Laboratory, we use empirical airtightness measurements to fit an infiltration model using the developed methodology. We find that the developed methodology can improve the overall model accuracy by 43% and improve extrapolation by 38%, compared with the model based on the common grey-box infiltration equation. We also notice that the selected features can improve the performance of a pure machine-learning model, indicating that our methodology identifies the features with the most physical significance to infiltration modeling.

97 MATHEMATICS AND COMPUTING↗

Accuracy, transferability, and computational efficiency of interatomic potentials for simulations of carbon under extreme conditions

Large-scale atomistic molecular dynamics (MD) simulations provide an exceptional opportunity to advance the fundamental understanding of carbon under extreme conditions of high pressures and temperatures. However, the fidelity of these simulations depends heavily on the accuracy of classical interatomic potentials governing the dynamics of many-atom systems. Here, this study critically assesses several popular empirical potentials for carbon, as well as machine learning interatomic potentials (MLIPs), in their ability to simulate a range of physical properties at high pressures and temperatures, including the diamond equation of state, its melting line, shock Hugoniot, uniaxial compressions, and the structure of liquid carbon. Empirical potentials fail to accurately predict the behavior of carbon under high pressure–temperature conditions. In contrast, MLIPs demonstrate quantum accuracy, with Spectral Neighbor Analysis Potential (SNAP) and atomic cluster expansion (ACE) being the most accurate in reproducing the density functional theory results. ACE displays remarkable transferability despite not being specifically trained for extreme conditions. Furthermore, ACE and SNAP exhibit superior computational performance on graphics processing unit-based systems in billion atom MD simulations, with SNAP emerging as the fastest. In addition to offering practical guidance in selecting an interatomic potential with a fine balance of accuracy, transferability, and computational efficiency, this work also highlights transformative opportunities for groundbreaking scientific discoveries facilitated by quantum-accurate MD simulations with MLIPs on emerging exascale supercomputers.

36 MATERIALS SCIENCE↗

Exploring Classification of Topological Priors With Machine Learning for Feature Extraction

In many scientific endeavors, increasingly abstract representations of data allow for new interpretive methodologies and conceptualization of phenomena. For example, moving from raw imaged pixels to segmented and reconstructed objects allows researchers new insights and means to direct their studies toward relevant areas. Thus, the development of new and improved methods for segmentation remains an active area of research. With advances in machine learning and neural networks, scientists have been focused on employing deep neural networks such as U-Net to obtain pixel-level segmentations, namely, defining associations between pixels and corresponding/referent objects and gathering those objects afterward. Topological analysis, such as the use of the Morse-Smale complex to encode regions of uniform gradient flow behavior, offers an alternative approach: first, create geometric priors, and then apply machine learning to classify. This approach is empirically motivated since phenomena of interest often appear as subsets of topological priors in many applications. Using topological elements not only reduces the learning space but also introduces the ability to use learnable geometries and connectivity to aid the classification of the segmentation target. Here, in this article, we describe an approach to creating learnable topological elements, explore the application of ML techniques to classification tasks in a number of areas, and demonstrate this approach as a viable alternative to pixel-level classification, with similar accuracy, improved execution time, and requiring marginal training data.

97 MATHEMATICS AND COMPUTING↗

A Fortran–Python interface for integrating machine learning parameterization into earth system models

Abstract. Parameterizations in earth system models (ESMs) are subject to biases and uncertainties arising from subjective empirical assumptions and incomplete understanding of the underlying physical processes. Recently, the growing representational capability of machine learning (ML) in solving complex problems has spawned immense interests in climate science applications. Specifically, ML-based parameterizations have been developed to represent convection, radiation, and microphysics processes in ESMs by learning from observations or high-resolution simulations, which have the potential to improve the accuracies and alleviate the uncertainties. Previous works have developed some surrogate models for these processes using ML. These surrogate models need to be coupled with the dynamical core of ESMs to investigate the effectiveness and their performance in a coupled system. In this study, we present a novel Fortran–Python interface designed to seamlessly integrate ML parameterizations into ESMs. This interface showcases high versatility by supporting popular ML frameworks like PyTorch, TensorFlow, and scikit-learn. We demonstrate the interface's modularity and reusability through two cases: an ML trigger function for convection parameterization and an ML wildfire model. We conduct a comprehensive evaluation of memory usage and computational overhead resulting from the integration of Python codes into the Fortran ESMs. By leveraging this flexible interface, ML parameterizations can be effectively developed, tested, and integrated into ESMs.

54 ENVIRONMENTAL SCIENCES↗

Designing complex concentrated alloys with quantum machine learning and language modeling

Designing novel complex concentrated alloys (CCAs) is an essential topic in materials science. However, due to the complicated high-dimensional component-property relationship, tuning material properties by researchers’ experience is challenging, even when guided by physical or empirical rules. Here, we adopt quantum computing (QC) technology and machine learning models to provide a proof-of-concept application of QC in physical metallurgy. We propose a quantum support vector machine (QSVM) model to predict single-phase CCAs. We show that fine-tuned quantum kernels with entanglement deliver promising performance, with a maximum accuracy of 89.4%. The QSVM model is then used to identify 1,741 lightweight CCAs jointly with a new text-mining-based method. Meanwhile, we devise a controllable approach to study the effect of noise on model performance and find that the noise level needs to be minimized for high-performance QSVM models. Finally, this study provides a practical and general approach to designing CCAs based on quantum technologies.

36 MATERIALS SCIENCE↗

Modeling of the cold electron plasma density for radiation belt physics

This review focusses strictly on existing plasma density models, including ionospheric source models, empirical density models, physics-based and machine-learning density models. This review is framed in the context of radiation belt physics and space weather codes. The review is limited to the most commonly used models or to models recently developed and promising. A great variety of conditions is considered such as the magnetic local time variation, geomagnetic conditions, ionospheric source regions, radial and latitudinal dependence, and collisional vs. collisionless conditions. These models can serve to complement satellite observations of the electron plasma density when data are lacking, are for most of them commonly used in radiation belt physics simulations and can improve our understanding of the plasmasphere dynamics.

79 ASTRONOMY AND ASTROPHYSICS↗

Machine-Learning Assisted Identification of Battery Life Models

Predictive battery life models are commonly utilized to extrapolate degradation trends observed during accelerated aging tests for simulation of degradation in real-world applications. Thus, fitting accelerated aging data as accurately as possible and with low uncertainty is crucial for making believable projections of battery lifetime, but it is challenging to identify algebraic expressions that accurately fit multivariate degradation trends. A review of models published in literature reveal some common expressions for fitting calendar aging data, which is only dependent on temperature and state-of-charge, but no consistency across many models for fitting cycle aging data, indicating the need for a statistically rigorous data driven approach for developing empirical models. This talk will describe a machine-learning assisted method for identification of predictive battery life models utilizing bilevel optimization and symbolic regression. Bilevel optimization with cross-validation is used to statistically determine cell- and stress-dependent model parameters, while symbolic regression identifies both linear and multiplicative candidate expressions to predict stress-dependent degradation rates by selecting low-order subsets of features from a generated feature library. Because model expressions are identified empirically, it is crucial to ensure resulting models behave according to physical expectations, so the stability of models for interpolation or extrapolation is interrogated qualitatively through simulation and quantitatively through cross-validation and uncertainty quantification via bootstrap resampling. This model identification approach substantially improves upon models identified purely using expert judgement in terms of both accuracy and uncertainty. Model simulation and validation is then conducted by deriving a state-equation form of the predictive model, enabling simulation of battery aging under dynamic stresses. This enables validation of the predictive battery model on lab-based tests with varying conditions or on drive-cycle or application-cycle testing protocols. Parameter uncertainty can be carried forward into model simulation, giving lifetime estimates and confidence windows for cell- or system-level lifetime. The financial impact of battery model uncertainty can be estimated by incorporating uncertainty into a technoeconomic model.

battery↗

Multi-reward reinforcement learning based development of inter-atomic potential models for silica

Abstract Silica is an abundant and technologically attractive material. Due to the structural complexities of silica polymorphs coupled with subtle differences in Si–O bonding characteristics, the development of accurate models to predict the structure, energetics and properties of silica polymorphs remain challenging. Current models for silica range from computationally efficient Buckingham formalisms (BKS, CHIK, Soules) to reactive (ReaxFF) and more recent machine-learned potentials that are flexible but computationally costly. Here, we introduce an improved formalism and parameterization of BKS model via a multireward reinforcement learning (RL) using an experimental training dataset. Our model concurrently captures the structure, energetics, density, equation of state, and elastic constants of quartz (equilibrium) as well as 20 other metastable silica polymorphs. We also assess its ability in capturing amorphous properties and highlight the limitations of the BKS-type functional forms in simultaneously capturing crystal and amorphous properties. We demonstrate ways to improve model flexibility and introduce a flexible formalism, machine-learned ML-BKS, that outperforms existing empirical models and is on-par with the recently developed 50 to 100 times more expensive Gaussian approximation potential (GAP) in capturing the experimental structure and properties of silica polymorphs and amorphous silica.

36 MATERIALS SCIENCE↗

Moisture availability mediates the relationship between terrestrial gross primary production and solar-induced chlorophyll fluorescence: Insights from global-scale variations

Effective use of solar-induced chlorophyll fluorescence (SIF) to estimate and monitor gross primary production (GPP) in terrestrial ecosystems requires a comprehensive understanding and quantification of the relationship between SIF and GPP. To date, this understanding is incomplete and somewhat controversial in the literature. Here we derived the GPP/SIF ratio from multiple data sources as a diagnostic metric to explore its global-scale patterns of spatial variation and potential climatic dependence. We found that the growing season GPP/SIF ratio varied substantially across global land surfaces, with the highest ratios consistently found in boreal regions. Spatial variation in GPP/SIF was strongly modulated by climate variables. The most striking pattern was a consistent decrease in GPP/SIF from cold-and-wet climates to hot-and-dry climates. We propose that the reduction in GPP/SIF with decreasing moisture availability may be related to stomatal responses to aridity. Furthermore, we show that GPP/SIF can be empirically modeled from climate variables using a machine learning (random forest) framework, which can improve the modeling of ecosystem production and quantify its uncertainty in global terrestrial biosphere models. Finally, our results point to the need for targeted field and experimental studies to better understand the patterns observed and to improve the modeling of the relationship between SIF and GPP over broad scales.

59 BASIC BIOLOGICAL SCIENCES↗

Updated ASME design correlations and qualification plan for powder bed fusion 316H stainless steel

This report provides an update on the Advanced Materials and Manufacturing Technologies (AMMT) program effort to qualify Laser-Powder Bed Fusion (L-PBF) 316H stainless steel for use with the ASME Boiler & Pressure Vessel Code Section III, Division 5 rules. The report summarizes progress in testing and characterizing L-PBF material at elevated temperatures by providing preliminary design data for L-PBF 316H and by comparing the elevated temperature performance of the L-PBF material to wrought and conventional fusion welded 316H. The report then updates the initial AMMT qualification plan for L-PBF 316H, originally developed in 2023, to update the accelerated qualification strategy adopted in that plan to account for the new high temperature test data. The report also explores a few methods for further accelerating the qualification process using machine learning techniques to supplement the more conventional, empirical analysis methods typically used by ASME to correlate and extrapolate time-dependent material test data.

36 MATERIALS SCIENCE↗

Loss Landscape Analysis for Reliable Quantized ML Models for Scientific Sensing

In this paper, we propose a method to perform empirical analysis of the loss landscape of machine learning (ML) models. The method is applied to two ML models for scientific sensing, which necessitates quantization to be deployed and are subject to noise and perturbations due to experimental conditions. Our method allows assessing the robustness of ML models to such effects as a function of quantization precision and under different regularization techniques -- two crucial concerns that remained underexplored so far. By investigating the interplay between performance, efficiency, and robustness by means of loss landscape analysis, we both established a strong correlation between gently-shaped landscapes and robustness to input and weight perturbations and observed other intriguing and non-obvious phenomena. Our method allows a systematic exploration of such trade-offs a priori, i.e., without training and testing multiple models, leading to more efficient development workflows. This work also highlights the importance of incorporating robustness into the Pareto optimization of ML models, enabling more reliable and adaptive scientific sensing systems.

Baldi, Tommaso [Pisa, Scuola Normale Superiore]↗

Homogeneous ice nucleation in an ab initio machine-learning model of water

Molecular simulations have provided valuable insight into the microscopic mechanisms underlying homogeneous ice nucleation. While empirical models have been used extensively to study this phenomenon, simulations based on first-principles calculations have so far proven prohibitively expensive. Here, we circumvent this difficulty by using an efficient machine-learning model trained on density-functional theory energies and forces. We compute nucleation rates at atmospheric pressure, over a broad range of supercoolings, using the seeding technique and systems of up to hundreds of thousands of atoms simulated with ab initio accuracy. The key quantity provided by the seeding technique is the size of the critical cluster (i.e., a size such that the cluster has equal probabilities of growing or melting at the given supersaturation), which is used together with the equations of classical nucleation theory to compute nucleation rates. We find that nucleation rates for our model at moderate supercoolings are in good agreement with experimental measurements within the error of our calculation. We also study the impact of properties such as the thermodynamic driving force, interfacial free energy, and stacking disorder on the calculated rates.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Thermophysical properties of FLiBe using moment tensor potentials

Fluoride salts are prospective materials for applications in some next-generation nuclear reactors and their thermophysical properties at various conditions are of interest. Experimental measurement of the properties of these salts is often difficult and, in some cases, unfeasible due to challenges from high temperatures, impurity control, and corrosivity. Therefore, accurate theoretical methods are needed for fluoride salt property prediction. In this work, we used moment tensor potentials (MTP) to approximate the potential energy surface of eutectic FLiBe (66.6% LiF – 33.3% BeF2) predicted by the ab initio (DFT-D3) method. Here, we then used the developed potential and molecular dynamics to obtain several thermophysical properties of FLiBe, including radial distribution functions, density, self-diffusion coefficients, thermal expansion, specific heat capacity, bulk modulus, viscosity, and thermal conductivity. Our results show that the MTP potential approximates the potential energy surface accurately and the overall approach yields very good agreement with experimental values. The converged fitting can be obtained with less than 600 configurations generated from DFT calculations, which data can be generated in just 1200 core hours on today's typical processors. The MTP potential is faster than many machine learning potentials and about one order of magnitude slower than widely used empirical molten salt potentials such as Tosi/Fumi.

36 MATERIALS SCIENCE↗

Green AI: Insights Into Deep Learning's Looming Energy Efficiency Crisis

As demands grow to integrate artificial intelligence into every aspect of industry, commerce, and life, deep learning's exploding energy cost has become a looming crisis, making AI systems a salient energy-efficiency challenge. One might expect that doubling a neural network's size would halve its error rate, or at least allow it to achieve greater performance given the same amount of time and energy. I will present clear and substantial scientific evidence which indicates that not only is this intuition wildly wrong, but that neural networks scale so poorly that to increase deep learning performance by only a small fraction can easily require an order of magnitude or more increase in computational resources and energy. Further, the marginal trade-off price of to increase model performance rapidly explodes as performance targets are increased. To address this challenge, I will provide a toolkit of techniques that can be applied today to mitigate the inefficiency of modern deep learning. And, I will conclude by illuminating a practical path forward towards efficient, Green AI.

artificial intelligence↗

Comparing machine learning and interpolation methods for loop-level calculations

The need to approximate functions is ubiquitous in science, either due to empirical constraints or high computational cost of accessing the function. In high-energy physics, the precise computation of the scattering cross-section of a process requires the evaluation of computationally intensive integrals. A wide variety of methods in machine learning have been used to tackle this problem, but often the motivation of using one method over another is lacking. Comparing these methods is typically highly dependent on the problem at hand, so we specify to the case where we can evaluate the function a large number of times, after which quick and accurate evaluation can take place. We consider four interpolation and three machine learning techniques and compare their performance on three toy functions, the four-point scalar Passarino-Veltman D_0 D 0 function, and the two-loop self-energy master integral M. We find that in low dimensions (d = 3), traditional interpolation techniques like the Radial Basis Function perform very well, but in higher dimensions (d=5, 6, 9) we find that multi-layer perceptrons (a.k.a neural networks) do not suffer as much from the curse of dimensionality and provide the fastest and most accurate predictions.

97 MATHEMATICS AND COMPUTING↗

Degradation and Modeling of Large-Format Commercial Lithium-Ion Cells as a Function of Chemistry, Design, and Aging Conditions

Demand for large-format (>10 Ah) lithium-ion batteries has increased substantially in recent years, due to the growth of both electric vehicle and stationary energy storage markets. The economics of these applications is sensitive to the lifetime of the batteries, and end-of-life can either be due to energy or power limitations. Despite this, there is little information from cell manufacturers on the sensitivity of cell degradation to environmental conditions or battery use. This work reports accelerated aging test data from four commercial large-format lithium-ion batteries from three manufacturers, with varying design (thickness, casings, ...), chemistry (lithium-iron-phosphate (LFP) or lithium-nickel-manganese-cobalt-oxide positive electrodes (NMC), with graphite (Gr) negative electrodes), and capacity (50 to 250 Amp hours). The tested LFP|Gr cell is found to be relatively insensitive to cycling conditions like temperature or voltage window, while NMC|Gr cells have varying sensitivity. Degradation trends are further investigated by training predictive models: simple polynomial trend lines, a semi-empirical reduced-order model, and an empirical reduced-order model identified using machine-learning based on symbolic regression. Calendar and cycle life are simulated over a variety of conditions to directly compare the various batteries. Cell size and thickness are found to substantially impact sensitivity to temperature during cycle aging, while electrode chemistry impacts depth-of-discharge sensitivity. Real-world battery lifetime is evaluated by simulating residential energy storage and commercial frequency containment reserve systems in several U.S. climate regions. Predicted lifetime across cell types varies from 7 years to 20+ years, though all cells are predicted to have at least 10 year life in certain conditions.

battery lifetime↗

A Dataset of 3D Structural and Simulated Transport Properties of Complex Porous Media

Physical processes that occur within porous materials have wide-ranging applications including - but not limited to - carbon sequestration, battery technology, membranes, oil and gas, geothermal energy, nuclear waste disposal, water resource management. The equations that describe these physical processes have been studied extensively; however, approximating them numerically requires immense computational resources due to the complex behavior that arises from the geometrically-intricate solid boundary conditions in porous materials. Here, we introduce a new dataset of unprecedented scale and breadth, DRP-372: a catalog of 3D geometries, simulation results, and structural properties of samples hosted on the Digital Rocks Portal. The dataset includes 1736 flow and electrical simulation results on 217 samples, which required more than 500 core years of computation. This data can be used for many purposes, such as constructing empirical models, validating new simulation codes, and developing machine learning algorithms that closely match the extensive purely-physical simulation. This article offers a detailed description of the contents of the dataset including the data collection, simulation schemes, and data validation.

3D images↗

Dataset, Code, and Models for Training Deep Learning Potentials for Low Temperature Plasma-Surface Interactions

This repository contains datasets, training scripts, and finished models, and test simulations used in the development of DeepREBO— a machine-learned interatomic potential trained to emulate the REBO2 empirical potential. The data was generated to study deep potential development for simulations of plasma-surface interactions. It uses an active learning framework, starting from a minimal dataset and iteratively expanding it. Included are those generated datasets, the trained models, and simulations used to evaluate the performance of the training process. This resource supports reproducibility and provides a reference framework for training deep potentials in plasma-surface interaction studies.

active learning↗