Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “active machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Performance Evaluation of Gray-box and Machine Learning Models of a Thermal Energy Storage System with Active Insulation

An interior partition wall integrated with active thermal storage and a dynamic insulation system was built and then installed in an office building in Oak Ridge, Tennessee, TN. This smart wall, termed the Empower Wall, was equipped with embedded pipes in the building envelope core component and an additional pipe network enclosing rigid insulation to switch on and off the active insulation dynamically. The performance of the wall's contribution to cooling load reduction under different parameters has been investigated in previous publications. Aiming to be deployed into model predictive control and other optimization methods, simplified and reliable models for the developed wall and the room accommodating it are required. They are needed to characterize the properties and thermal response of both Empower Wall and building envelope, which form an essential component for accurate indoor temperature or cooling/heating demand prediction. In this study, simplified gray-box and regression models as well as machine learning model were developed and the performance of them were compared and analyzed.

Cui, Borui↗

On-the-fly closed-loop materials discovery via Bayesian active learning

Active learning—the field of machine learning (ML) dedicated to optimal experiment design—has played a part in science as far back as the 18th century when Laplace used it to guide his discovery of celestial mechanics. In this work, we focus a closed-loop, active learning-driven autonomous system on another major challenge, the discovery of advanced materials against the exceedingly complex synthesis-processes-structure-property landscape. We demonstrate an autonomous materials discovery methodology for functional inorganic compounds which allow scientists to fail smarter, learn faster, and spend less resources in their studies, while simultaneously improving trust in scientific results and machine learning tools. This robot science enables science-over-the-network, reducing the economic impact of scientists being physically separated from their labs. The real-time closed-loop, autonomous system for materials exploration and optimization (CAMEO) is implemented at the synchrotron beamline to accelerate the interconnected tasks of phase mapping and property optimization, with each cycle taking seconds to minutes. We also demonstrate an embodiment of human-machine interaction, where human-in-the-loop is called to play a contributing role within each cycle. This work has resulted in the discovery of a novel epitaxial nanocomposite phase-change memory material.

36 MATERIALS SCIENCE↗

Screening Cu-Zeolites for Methane Activation Using Curriculum-Based Training

Machine learning (ML), when used synergistically with atomistic simulations, has recently emerged as a powerful tool for accelerated catalyst discovery. However, the application of these techniques has been limited by the lack of interpretable and transferable ML models. In this work, we propose a curriculum-based training (CBT) philosophy to systematically develop reactive machine learning potentials (rMLPs) for high-throughput screening of zeolite catalysts. Our CBT approach combines several different types of calculations to gradually teach the ML model about the relevant regions of the reactive potential energy surface. The resulting rMLPs are accurate, transferable, and interpretable. We further demonstrate the effectiveness of this approach by exhaustively screening thousands of [CuOCu] 2+ sites across hundreds of Cu-zeolites for the industrially relevant methane activation reaction. Specifically, this large-scale analysis of the entire International Zeolite Association (IZA) database identifies a set of previously unexplored zeolites (i.e., MEI, ATN, EWO, and CAS) that show the highest ensemble-averaged rates for [CuOCu] 2+ -catalyzed methane activation. We believe that this CBT philosophy can be generally applied to other zeolite-catalyzed reactions and, subsequently, to other types of heterogeneous catalysts. Thus, this represents an important step toward overcoming the long-standing barriers within the computational heterogeneous catalysis community.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Predicting the Activity and Selectivity of Bimetallic Metal Catalysts for Ethanol Reforming using Machine Learning

Machine learning is ideally suited for the pattern detection in large uniform datasets, but consistent experimental datasets on catalyst studies are often small. Here we demonstrate how a combination of machine learning and first-principles calculations can be used to extract knowledge from a relatively small set of experimental data. The approach is based on combining a complex machine-learning model trained on an extensive computational library of transition-state energies with simple linear regression models of experimental catalytic activities and selectivities from the literature. Using the combined model, we identify the key C–C bond scission reactions involved in ethanol reforming and perform a computational screening for ethanol reforming on monolayer bimetallic catalysts with architectures TM-Pt-Pt(111) and Pt-TM-Pt(111) (TM = 3d transition metals). The model also predicts four promising catalyst compositions for future experimental studies. In conclusion, the approach is not limited to ethanol reforming but is of general use for the interpretation of experimental observations as well as for the computational discovery of novel catalytic materials.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Machine learning interatomic potential for silicon-nitride (Si 3 N 4 ) by active learning

Silicon nitride (Si 3 N 4 ) is an extensively used material in the automotive, aerospace, and semiconductor industries. However, its widespread use is in contrast to the scarce availability of reliable interatomic potentials that can be employed to study various aspects of this material on an atomistic scale, particularly its amorphous phase. In this work, we developed a machine learning interatomic potential, using an efficient active learning technique, combined with the Gaussian approximation potential (GAP) method. Our strategy is based on using an inexpensive empirical potential to generate an initial dataset of atomic configurations, for which energies and forces were recalculated with density functional theory (DFT); thereafter, a GAP was trained on these data and an iterative re-training algorithm was used to improve it by learning on-the-fly. When compared to DFT, our potential yielded a mean absolute error of 8 meV/atom in energy calculations for a variety of liquid and amorphous structures and a speed-up of molecular dynamics simulations by 3–4 orders of magnitude, while achieving a first-rate agreement with experimental results. Our potential is publicly available in an open-access repository.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Physics-Guided Continual Learning for Predicting Emerging Aqueous Organic Redox Flow Battery Material Performance

Aqueous organic redox flow batteries (AORFBs) have gained popularity in renewable energy storage due to their low cost, environmental friendliness and scalability. The rapid discovery of aqueous soluble organic (ASO) redox-active materials necessitates efficient machine learning surrogates for predicting battery performance. The physics-guided continual learning (PGCL) method proposed in this study can incrementally learn data from new ASO electrolytes while addressing catastrophic forgetting issues in conventional machine learning. Using a AORFB database with a thousand potential materials generated by a 780 $\text{cm}^2$ interdigitated cell model, PGCL incorporates AORFB physics to optimize the continual learning task formation and training strategies to retain previously learned battery material knowledge. Finally, the trained PGCL demonstrates its capability in assessing emerging ASO materials within the established parameter space when evaluated with the dihydroxyphenazine isomers.

25 ENERGY STORAGE↗

Elucidating and predicting the dynamic evolution of water and land systems due to natural and energy-related forcings

Focal Area(s): 3. Insight gleaned from complex data (both observed and simulated) using AI, big data analytics, and other advanced methods, including explainable AI and physics- or knowledge-guided AI; & 1. Data acquisition and assimilation enabled by machine learning, AI, and advanced methods including experimental/network design/optimization, unsupervised learning (including deep learning), and hardware-related efforts involving AI (e.g., edge computing). Science Challenge: Interactions between water, land, and energy systems are complex and occur on a variety of scales, ranging from local to basinal to regional. Accurately predicting the behavior of ground water and surface water systems for 5-10 years and beyond requires an understanding of the current system and the ability to model both the natural system at scale and human-induced forcings related to energy and other activities. Artificial intelligence and machine learning (AI/ML) combined with modern compilation and integration efforts for U.S. groundwater and surface water systems present potential solutions to bolstering detailed physics-based models of these systems. Big data tied with ML and physics-based modeling can drive breakthroughs in understanding the earth system, but research is often impeded by data access (e.g., privacy issues), quality, formats, gaps, multi-source, multi-scale, integration, and spatiotemporal challenges. Effective integration of real data and simulated (synthetic) data that fill gaps is critical. Overcoming these complex data and model integration challenges will enable a transformational approach to acquiring enhanced understanding of environmental systems.

54 ENVIRONMENTAL SCIENCES↗

Residential Demand Flexibility: Modeling Occupant Behavior using Sociodemographic Predictors

Demand flexibility (DF) has the potential to increase the saturation of renewables in the grid and reduce operating costs for both utilities and customers. However, less than 8% of U.S. residential electric customers are enrolled in DF programs. A major research gap on this topic is an uneven understanding of behavioral drivers of electricity use and DF program participation at the household level. In this study, we employ machine learning models to predict residential occupant behavior in activities relevant to DF. We model occupants' extensive decisions (i.e., choice of action) and intensive behaviors (i.e., amount of time spent) during peak and off-peak time periods using the publicly available American Time Use Survey, which includes activities data for approximately 200,000 respondents. In our machine learning models, predictions for both extensive and intensive behavior fell within a +/-20% error margin at the aggregate level. We identify 13 key sociodemographic predictors of DF-related intensive behavior using LASSO inference and beta coefficient ranking. However, these top predictors differ by activity, suggesting potential scope for differential user targeting for DF events and technologies during program design. This work also contributes to understanding when and who might adopt these DF technologies based on their daily routine activities.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Fingerprinting Interactions between Proteins and Ligands for Facilitating Machine Learning in Drug Discovery

Molecular recognition is fundamental in biology, underpinning intricate processes through specific protein–ligand interactions. This understanding is pivotal in drug discovery, yet traditional experimental methods face limitations in exploring the vast chemical space. Computational approaches, notably quantitative structure–activity/property relationship analysis, have gained prominence. Molecular fingerprints encode molecular structures and serve as property profiles, which are essential in drug discovery. While two-dimensional (2D) fingerprints are commonly used, three-dimensional (3D) structural interaction fingerprints offer enhanced structural features specific to target proteins. Machine learning models trained on interaction fingerprints enable precise binding prediction. Recent focus has shifted to structure-based predictive modeling, with machine-learning scoring functions excelling due to feature engineering guided by key interactions. Notably, 3D interaction fingerprints are gaining ground due to their robustness. Various structural interaction fingerprints have been developed and used in drug discovery, each with unique capabilities. This review recapitulates the developed structural interaction fingerprints and provides two case studies to illustrate the power of interaction fingerprint-driven machine learning. The first elucidates structure–activity relationships in β2 adrenoceptor ligands, demonstrating the ability to differentiate agonists and antagonists. The second employs a retrosynthesis-based pre-trained molecular representation to predict protein–ligand dissociation rates, offering insights into binding kinetics. Despite remarkable progress, challenges persist in interpreting complex machine learning models built on 3D fingerprints, emphasizing the need for strategies to make predictions interpretable. Binding site plasticity and induced fit effects pose additional complexities. Interaction fingerprints are promising but require continued research to harness their full potential.

3D structural interaction fingerprints↗

Machine Learning with Gradient-Based Optimization of Nuclear Waste Vitrification with Uncertainties and Constraints

Gekko is an optimization suite in Python that solves optimization problems involving mixed-integer, nonlinear, and differential equations. The purpose of this study is to integrate common Machine Learning (ML) algorithms such as Gaussian Process Regression (GPR), support vector regression (SVR), and artificial neural network (ANN) models into Gekko to solve data based optimization problems. Uncertainty quantification (UQ) is used alongside ML for better decision making. These methods include ensemble methods, model-specific methods, conformal predictions, and the delta method. An optimization problem involving nuclear waste vitrification is presented to demonstrate the benefit of ML in this field. ML models are compared against the current partial quadratic mixture (PQM) model in an optimization problem in Gekko. GPR with conformal uncertainty was chosen as the best substitute model as it had a lower mean squared error of 0.0025 compared to 0.018 and more confidently predicted a higher waste loading of 37.5 wt% compared to 34 wt%. The example problem shows that these tools can be used in similar industry settings where easier use and better performance is needed over classical approaches. Future works with these tools include expanding them with other regression models and UQ methods, and exploration into other optimization problems or dynamic control.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Adaptive Interface-PINNs (AdaI-PINNs): An Efficient Physics-Informed Neural Networks Framework for Interface Problems

Here, we present an efficient physics-informed neural networks (PINNs) framework, termed Adaptive Interface-PINNs (AdaI-PINNs), to improve the modeling of interface problems with discontinuous coefficients and/or interfacial jumps. This framework is an enhanced version of its predecessor, Interface PINNs or I-PINNs (Sarma et al.; https://doi.org/10.1016/j.cma.2024.117135), which involves domain decomposition and assignment of different predefined activation functions to the neural networks in each subdomain across a sharp interface, while keeping all other parameters of the neural networks identical. In AdaI-PINNs, the activation functions vary solely in their slopes, which are trained along with the other parameters of the neural networks. This makes the AdaI-PINNs framework fully automated without requiring preset activation functions. Comparative studies on one-dimensional, two-dimensional, and three-dimensional benchmark elliptic interface problems reveal that AdaI-PINNs outperform I-PINNs, reducing computational costs by 2-6 times while producing similar or better accuracy.

97 MATHEMATICS AND COMPUTING↗

Progress of Gas Injection EOR Surveillance in the Bakken Unconventional Play—Technical Review and Machine Learning Study

Although considerable laboratory and modeling activities were performed to investigate the enhanced oil recovery (EOR) mechanisms and potential in unconventional reservoirs, only limited research has been reported to investigate actual EOR implementations and their surveillance in fields. Eleven EOR pilot tests that used CO2, rich gas, surfactant, water, etc., have been conducted in the Bakken unconventional play since 2008. Gas injection was involved in eight of these pilots with huff ‘n’ puff, flooding, and injectivity operations. Surveillance data, including daily production/injection rates, bottomhole injection pressure, gas composition, well logs, and tracer testing, were collected from these tests to generate time-series plots or analytics that can inform operators of downhole conditions. A technical review showed that pressure buildup, conformance issues, and timely gas breakthrough detection were some of the main challenges because of the interconnected fractures between injection and offset wells. The latest operation of co-injecting gas, water, and surfactant through the same injection well showed that these challenges could be mitigated by careful EOR design and continuous reservoir monitoring. Reservoir simulation and machine learning were then conducted for operators to rapidly predict EOR performance and take control actions to improve EOR outcomes in unconventional reservoirs.

Energy & Fuels↗

Low-index mesoscopic surface reconstructions of Au surfaces using Bayesian force fields

Metal surfaces have long been known to reconstruct, significantly influencing their structural and catalytic properties. Many key mechanistic aspects of these subtle transformations remain poorly understood due to limitations of previous simulation approaches. Using active learning of Bayesian machine-learned force fields trained from ab initio calculations, we enable large-scale molecular dynamics simulations to describe the thermodynamics and time evolution of the low-index mesoscopic surface reconstructions of Au (e.g., the Au(111)-‘Herringbone,’ Au(110)-(1 × 2)-‘Missing-Row,’ and Au(100)-‘Quasi-Hexagonal’ reconstructions). This capability yields direct atomistic understanding of the dynamic emergence of these surface states from their initial facets, providing previously inaccessible information such as nucleation kinetics and a complete mechanistic interpretation of reconstruction under the effects of strain and local deviations from the original stoichiometry. We successfully reproduce previous experimental observations of reconstructions on pristine surfaces and provide quantitative predictions of the emergence of spinodal decomposition and localized reconstruction in response to strain at non-ideal stoichiometries. A unified mechanistic explanation is presented of the kinetic and thermodynamic factors driving surface reconstruction. Furthermore, we study surface reconstructions on Au nanoparticles, where characteristic (111) and (100) reconstructions spontaneously appear on a variety of high-symmetry particle morphologies.

36 MATERIALS SCIENCE↗

FAIR AI models in high energy physics

Abstract The findable, accessible, interoperable, and reusable (FAIR) data principles provide a framework for examining, evaluating, and improving how data is shared to facilitate scientific discovery. Generalizing these principles to research software and other digital products is an active area of research. Machine learning models—algorithms that have been trained on data without being explicitly programmed—and more generally, artificial intelligence (AI) models, are an important target for this because of the ever-increasing pace with which AI is transforming scientific domains, such as experimental high energy physics (HEP). In this paper, we propose a practical definition of FAIR principles for AI models in HEP and describe a template for the application of these principles. We demonstrate the template’s use with an example AI model applied to HEP, in which a graph neural network is used to identify Higgs bosons decaying to two bottom quarks. We report on the robustness of this FAIR AI model, its portability across hardware architectures and software frameworks, and its interpretability.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Pathway-based analyses of gene expression profiles at low doses of ionizing radiation

Radiation exposure poses a significant threat to human health. Emerging research indicates that even low-dose radiation once believed to be safe, may have harmful effects. This perception has spurred a growing interest in investigating the potential risks associated with low-dose radiation exposure across various scenarios. To comprehensively explore the health consequences of low-dose radiation, our study employs a robust statistical framework that examines whether specific groups of genes, belonging to known pathways, exhibit coordinated expression patterns that align with the radiation levels. Notably, our findings reveal the existence of intricate yet consistent signatures that reflect the molecular response to radiation exposure, distinguishing between low-dose and high-dose radiation. Moreover, we leverage a pathway-constrained variational autoencoder to capture the nonlinear interactions within gene expression data. By comparing these two analytical approaches, our study aims to gain valuable insights into the impact of low-dose radiation on gene expression patterns, identify pathways that are differentially affected, and harness the potential of machine learning to uncover hidden activity within biological networks. This comparative analysis contributes to a deeper understanding of the molecular consequences of low-dose radiation exposure.

63 RADIATION, THERMAL, AND OTHER ENVIRON. POLLUTAN↗

Machine Learning Models for Network Traffic Classification in Programmable Logic

Network traffic classification via machine learning on network packet payloads has emerged as an active area of research for network security due to the high accuracy machine learning models have achieved in classifying payloads. For effective deployment as part of network security, these machine learning models must not only classify malicious packet payloads accurately, they must also identify anomalous payloads and perform inference at speeds generally faster than 10,000 packets per second to be effective. This work explores the in- ference speeds and accuracy of several neural network models implemented in programmable logic on various field programmable gate arrays (FPGA) including the Xilinx VC1902 and Xilinx Zynq Ultrascale+. This work also presents the design and performance of both an autoencoder and variational autoencoder programmed on the FPGA for identifying anomalous packet payloads. The performance benefits of the FPGA implementation for this type of packet payload inspection driven by machine learning are compared against graphics processing unit (GPU) inference implementations run on two state-of-the-art datacenter GPU devices, the NVIDIA V100 and A100. The model accuracy difference between the FPGA and GPU implementations was found to be 4% or less while the Xilinx VC1902 outperformed both the NVIDIA V100 and A100 for inference speeds on all the models explored except the variational autoencoder.

97 MATHEMATICS AND COMPUTING↗

Active learning of ternary alloy structures and energies

Abstract Machine learning models with uncertainty quantification have recently emerged as attractive tools to accelerate the navigation of catalyst design spaces in a data-efficient manner. Here, we combine active learning with a dropout graph convolutional network (dGCN) as a surrogate model to explore the complex materials space of high-entropy alloys (HEAs). We train the dGCN on the formation energies of disordered binary alloy structures in the Pd-Pt-Sn ternary alloy system and improve predictions on ternary structures by performing reduced optimization of the formation free energy, the target property that determines HEA stability, over ensembles of ternary structures constructed based on two coordinate systems: (a) a physics-informed ternary composition space, and (b) data-driven coordinates discovered by the Diffusion Maps manifold learning scheme. Both reduced optimization techniques improve predictions of the formation free energy in the ternary alloy space with a significantly reduced number of DFT calculations compared to a high-fidelity model. The physics-based scheme converges to the target property in a manner akin to a depth-first strategy, whereas the data-driven scheme appears more akin to a breadth-first approach. Both sampling schemes, coupled with our acquisition function, successfully exploit a database of DFT-calculated binary alloy structures and energies, augmented with a relatively small number of ternary alloy calculations, to identify stable ternary HEA compositions and structures. This generalized framework can be extended to incorporate more complex bulk and surface structural motifs, and the results demonstrate that significant dimensionality reduction is possible in thermodynamic sampling problems when suitable active learning schemes are employed.

Chemistry↗