Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “active machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Low latency optical-based mode tracking with machine learning deployed on FPGAs on a tokamak

Active feedback control in magnetic confinement fusion devices is desirable to mitigate plasma instabilities and enable robust operation. Optical high-speed cameras provide a powerful, non-invasive diagnostic and can be suitable for these applications. Here, in this study, we process high-speed camera data, at rates exceeding 100 kfps, on in situ field-programmable gate array (FPGA) hardware to track magnetohydrodynamic (MHD) mode evolution and generate control signals in real time. Our system utilizes a convolutional neural network (CNN) model, which predicts the n = 1 MHD mode amplitude and phase using camera images with better accuracy than other tested non-deep-learning-based methods. By implementing this model directly within the standard FPGA readout hardware of the high-speed camera diagnostic, our mode tracking system achieves a total trigger-to-output latency of 17.6 μs and a throughput of up to 120 kfps. This study at the High Beta Tokamak-Extended Pulse (HBT-EP) experiment demonstrates an FPGA-based high-speed camera data acquisition and processing system, enabling application in real-time machine-learning-based tokamak diagnostic and control as well as potential applications in other scientific domains.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

A comparison of machine learning methods to classify radioactive elements using prompt-gamma-ray neutron activation data

The detection of illicit radiological materials is critical to establishing a robust second line of defence in nuclear security. Neutron-capture prompt-gamma activation analysis (PGAA) can be used to detect multiple radioactive materials across the entire Periodic Table. However, long detection times and a high rate of false positives pose a significant hindrance in the deployment of PGAA-based systems to identify the presence of illicit substances in nuclear forensics. In the present work, six different machine-learning algorithms were developed to classify radioactive elements based on the PGAA energy spectra. The model performance was evaluated using standard classification metrics and trend curves with an emphasis on comparing the effectiveness of algorithms that are best suited for classifying imbalanced datasets. We analyse the classification performance based on Precision, Recall, F1-score, Specificity, Confusion matrix, ROC-AUC curves, and Geometric Mean Score (GMS) measures. The tree-based algorithms (Decision Trees, Random Forest and AdaBoost) have consistently outperformed Support Vector Machine and K-Nearest Neighbours. Based on the results presented, AdaBoost is the preferred classifier to analyse data containing PGAA spectral information due to the high recall and minimal false negatives reported in the minority class.

97 MATHEMATICS AND COMPUTING↗

Improving the accessibility and transferability of machine learning algorithms for identification of animals in camera trap images: MLWIC2

Motion-activated wildlife cameras (or “camera traps”) are frequently used to remotely and noninvasively observe animals. The vast number of images collected from camera trap projects has prompted some biologists to employ machine learning algorithms to automatically recognize species in these images, or at least filter-out images that do not contain animals. These approaches are often limited by model transferability, as a model trained to recognize species from one location might not work as well for the same species in different locations. Furthermore, these methods often require advanced computational skills, making them inaccessible to many biologists. We used 3 million camera trap images from 18 studies in 10 states across the United States of America to train two deep neural networks, one that recognizes 58 species, the “species model,” and one that determines if an image is empty or if it contains an animal, the “empty-animal model.” Our species model and empty-animal model had accuracies of 96.8% and 97.3%, respectively. Furthermore, the models performed well on some out-of-sample datasets, as the species model had 91% accuracy on species from Canada (accuracy range 36%–91% across all out-of-sample datasets) and the empty-animal model achieved an accuracy of 91%–94% on out-of-sample datasets from different continents. Our software addresses some of the limitations of using machine learning to classify images from camera traps. By including many species from several locations, our species model is potentially applicable to many camera trap studies in North America. We also found that our empty-animal model can facilitate removal of images without animals globally. We provide the trained models in an R package (MLWIC2: Machine Learning for Wildlife Image Classification in R), which contains Shiny Applications that allow scientists with minimal programming experience to use trained models and train new models in six neural network architectures with varying depths.

59 BASIC BIOLOGICAL SCIENCES↗

Machine Learning‐Augmented Molecular Dynamics Simulations (MD) Reveal Insights Into the Disconnect Between Affinity and Activation of ZTP Riboswitch Ligands

Abstract The challenge of targeting RNA with small molecules necessitates a better understanding of RNA–ligand interaction mechanisms. However, the dynamic nature of nucleic acids, their ligand‐induced stabilization, and how conformational changes influence gene expression pose significant difficulties for experimental investigation. This work employs a combination of computational and experimental methods to address these challenges. By integrating structure‐informed design, crystallography, and machine learning‐augmented all‐atom molecular dynamics simulations (MD), we synthesized, biophysically and biochemically characterized, and studied the dissociation of a library of small molecule activators of the 5‐aminoimidazole–4–carboxamide ribonucleotide triphosphate (ZTP) riboswitch, a ligand‐binding RNA motif that regulates bacterial gene expression. We uncovered key interaction mechanisms, revealing valuable insights into the role of ligand binding kinetics on riboswitch activation. Further, we established that ligand on‐rates determine activation potency as opposed to binding affinity and elucidated RNA structural differences, which provide mechanistic insights into the interplay of RNA structure on riboswitch activation.

Chemistry↗

Infusion of AI/ML Technology into Operational NASA Data Systems

NASA has been developing a variety of Artificial Intelligence / Machine Learning technologies related to Earth Observations. In most cases, the full value of such a technology is realized when it is infused into an operational system. NASA’s Earth Science Data Systems program has been formulating repeatable methods to execute technology infusion. These efforts include the Advancing Collaborative Connections for Earth System Science (ACCESS) program, a Technology Infusion Playbook, and an assemblage of working groups investigating methods for infusion collaboration, community development, and capacity building. ESDS has also been executing a pathfinder activity to infuse a machine-learning-driven recommender of science keywords for Earth Observation datasets, which is intended to be used for metadata curation in the Earth Observation System Data and Information System.

C Lynnes↗

Datasets for Custom-trained Machine-learning Interatomic Potentials: Nitric Acid Aqueous Solution

This dataset was generated using an iterative active learning strategy with the ArcaNN software package (https://github.com/arcann-chem/arcann_training) to train machine-learning interatomic potentials (MLIPs) for aqueous nitric acid. Each active-learning cycle consisted of three stages: (1) training, (2) exploration, and (3) labeling. The initial training set comprised approximately 800 randomly selected configurations from a previous study by Lewis et al. (https://doi.org/10.1021/jp205510q), which investigated nitric acid solutions at 2, 3, 4, and 5 mol/L. For all configurations, single-point calculations of atomic forces and total energies were performed at the quantum density functional theory BLYP-D2 and PBE-D3 levels of theory using the CP2K Quickstep module. Valence electrons were treated explicitly, while core electrons on all atoms were represented by norm-conserving Goedecker–Teter–Hutter (GTH) pseudopotentials. Long-range dispersion interactions were accounted for using Grimme dispersion corrections. Wave functions were expanded in a mixed Gaussian-and-plane-wave scheme using TZV2P-MOLOPT basis sets for all elements and an 800 Ry auxiliary plane-wave cutoff for the electron density. Self-consistent field convergence was accelerated using orbital transformation and Direct Inversion in the Iterative Subspace, with a convergence threshold of 10^{-6}. All single-point calculations were carried out in periodic orthorhombic cells whose dimensions match those of the molecular configurations sampled from earlier trajectories. The CELL_REF keyword in CP2K was used to define a fixed reference cell, ensuring consistency in the reference data used for MLIP training, particularly when cell fluctuations are present in NpT simulations. The resulting high-fidelity energies and forces constitute the ground-truth labels used to train the MLIPs contained in this dataset.

Dinpajooh, Mohammadhasan [Pacific Northwest Nation↗

Recent Progress on Surface Water Quality Models Utilizing Machine Learning Techniques

Surface waterbodies are heavily exposed to pollutants caused by natural disasters and human activities. Empowering sensor technologies in water quality monitoring, sufficient measurements have become available to develop machine learning (ML) models. Numerous ML models have quickly been adopted to predict water quality indicators in various surface waterbodies. This paper reviews 78 recent articles from 2022 to October 2024, categorizing water quality models utilizing ML into three groups: Point-to-Point (P2P), which estimates the current target value based on other measurements at the same time point; Sequence-to-Point (S2P), which utilizes previous time series data to predict the target value at one time point ahead; and Sequence-to-Sequence (S2S), which uses previous time series data to forecast sequential target values in the future. The ML models used in each group are classified and compared according to water quality indicators, data availability, and model performance. Widely used strategies for improving performance, including feature engineering, hyperparameter tuning, and transfer learning, are recognized and described to enhance model effectiveness. The interpretability limitations of ML applications are discussed. This review provides a perspective on emerging ML for surface water quality models.

machine learning (ML)↗

Probe microscopy is all you need *

We pose that microscopy offers an ideal real-world experimental environment for the development and deployment of active Bayesian and reinforcement learning methods. Indeed, the tremendous progress achieved by machine learning (ML) and artificial intelligence over the last decade has been largely achieved via the utilization of static data sets, from the paradigmatic MNIST to the bespoke corpora of text and image data used to train large models such as GPT3, DALL·E and others. However, it is now recognized that continuous, minute improvements to state-of-the-art do not necessarily translate to advances in real-world applications. We argue that a promising pathway for the development of ML methods is via the route of domain-specific deployable algorithms in areas such as electron and scanning probe microscopy and chemical imaging. This will benefit both fundamental physical studies and serve as a test bed for more complex autonomous systems such as robotics and manufacturing. Favorable environment characteristics of scanning and electron microscopy include low risk, extensive availability of domain-specific priors and rewards, relatively small effects of exogenous variables, and often the presence of both upstream first principles as well as downstream learnable physical models for both statics and dynamics. Recent developments in programmable interfaces, edge computing, and access to application programming interfaces (APIs) facilitating microscope control, all render the deployment of ML codes on operational microscopes straightforward. We discuss these considerations and hope that these arguments will lead to create novel set of development targets for the ML community by accelerating both real world ML applications and scientific progress.

47 OTHER INSTRUMENTATION↗

Custom-trained Machine-learning Interatomic Potentials: ZnCl2 Aqueous Solution

This dataset was generated using an iterative active-learning strategy implemented in the ArcaNN software package (https://github.com/arcann-chem/arcann_training) to train machine-learning interatomic potentials for aqueous ZnCl2 solutions. Each active-learning cycle consisted of three stages: training, exploration, and labeling. The initial training set combined configurations generated in this work from enhanced-sampling ab initio molecular dynamics simulations with configurations from a previously reported neural-network-potential study of aqueous ZnCl2. The enhanced-sampling ab initio molecular dynamics simulations involved Zn–Cl separation and the chloride coordination number around Zn²? as collective variables. These configurations served as the seed dataset. Subsequent active-learning cycles expanded the training set by identifying and labeling configurations that were poorly represented by the current models, thereby improving coverage of ion-association states and changes in local coordination and charge-state environments relevant to the solution free-energy landscape. For all selected configurations, single-point calculations of the total energies and atomic forces were performed within density functional theory using the CP2K Quickstep module. Reference calculations employed the revPBE-D3 and r2SCAN exchange-correlation functionals. Motivated by recent work on aqueous Zn²?, the main revPBE calculations omitted D3 dispersion contributions involving Zn²?, while retaining the D3 correction for water and chloride. For comparison, fully dispersion-corrected revPBE-D3 reference calculations were also performed, with D3 applied to all species, including Zn²?. Valence electrons were treated explicitly, while core electrons were represented using norm-conserving Goedecker–Teter–Hutter pseudopotentials. The wave functions were expanded using the mixed Gaussian-and-plane-wave scheme with TZV2P-MOLOPT basis sets for all elements and a 600 Ry auxiliary plane-wave cutoff for the electron density. Self-consistent-field convergence was accelerated using the orbital-transformation and Direct Inversion in the Iterative Subspace algorithms, with a convergence threshold of 10?6. All single-point calculations were performed in periodic orthorhombic cells. The CELL_REF keyword in CP2K was used to define a fixed reference cell with a box length of 25 Å. This treatment ensured a consistent reference for configurations extracted from NpT trajectories with fluctuating cell dimensions. The resulting DFT energies and atomic forces constitute the ground-truth labels used to train the MLIPs. The resulting MLIP was trained for aqueous ZnCl2 solutions spanning concentrations from 0 to 30 molal and a broad pH range, from strongly acidic to strongly basic conditions. Representative examples of configurations included in the MLIP training dataset are provided below. These include 1) Representative configurations from the dataset labeled at the revPBE-D3 level, with D3 dispersion interactions involving Zn2+ excluded (revPBE-wo-D3). 2) Representative configurations from the dataset labeled at the fully dispersion-corrected revPBE-D3 level, with D3 interactions applied to all species, including Zn2+ (revPBE-D3). 3) Representative configurations from the dataset labeled at the r2SCAN level of theory (r2SCAN).

Dinpajooh, Mohammadhasan [Pacific Northwest Nation↗

Generalized Brønsted‐Evans‐Polanyi Relationships for Reactions on Metal Surfaces from Machine Learning

Abstract Brønsted‐Evans‐Polanyi (BEP) relationships, i. e., a linear scaling between reaction and activation energies, lie at the core of computational design of heterogeneous catalysts. However, BEPs are not general and often require reparameterization for each class of reactions. Here we construct generalized BEPs (gBEPs), which can predict activation energies for a diverse dataset of reactions of C, O, N and H containing molecules on metal surfaces. In a first step we develop a set of descriptors based on scaling relationships that can capture the change in chemical identity of reactants during the reaction. Subsequently, we use the reaction energy, these descriptors and a single descriptor for the surface structure to parameterize machine learning based regression approaches for the prediction of activation energies. The best approach we developed shows a Mean Absolute Error (MAE) of 0.11 eV for the training set (80 % of the data set) and 0.23 eV for the test set (20 % of the data set). The methodology presented here allows to calculate activation energies within fractions of seconds on a typical personal computer and due to its generality, accuracy and simplicity in application it might prove to be useful in transition metal catalyst design.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Uncertainty-aware molecular dynamics from Bayesian active learning for phase transformations and thermal transport in SiC

Abstract Machine learning interatomic force fields are promising for combining high computational efficiency and accuracy in modeling quantum interactions and simulating atomistic dynamics. Active learning methods have been recently developed to train force fields efficiently and automatically. Among them, Bayesian active learning utilizes principled uncertainty quantification to make data acquisition decisions. In this work, we present a general Bayesian active learning workflow, where the force field is constructed from a sparse Gaussian process regression model based on atomic cluster expansion descriptors. To circumvent the high computational cost of the sparse Gaussian process uncertainty calculation, we formulate a high-performance approximate mapping of the uncertainty and demonstrate a speedup of several orders of magnitude. We demonstrate the autonomous active learning workflow by training a Bayesian force field model for silicon carbide (SiC) polymorphs in only a few days of computer time and show that pressure-induced phase transformations are accurately captured. The resulting model exhibits close agreement with both ab initio calculations and experimental measurements, and outperforms existing empirical models on vibrational and thermal properties. The active learning workflow readily generalizes to a wide range of material systems and accelerates their computational understanding.

36 MATERIALS SCIENCE↗

Prediction of O and OH Adsorption on Transition Metal Oxide Surfaces from Bulk Descriptors

In the search for stable and active catalysts, density functional theory and machine learning (ML) based models can accelerate the screening of materials. While stability is conveniently addressed on the bulk level of computation, the modelling of catalytic activity requires expensive surface simulations. Here, in this work, we develop models for the surface adsorption energy of O and OH intermediates across a consistent and extensive dataset of pure transition metal oxide surfaces. We show that adsorption energies across metal oxidation states of +2 to +6 are well captured from the metal-oxygen bond strength extracted from the bulk level calculation. Specifically, we calculate the integrated crystal orbital Hamiltonian population (ICOHP) of the metal-oxygen bond in the bulk oxide and employ a simple normalization scheme to obtain a strong correlation with adsorption energetics. By combining our ICOHP descriptor with non DFT features in a Gaussian Process regression (GPR) model, we achieve high model accuracy with mean absolute errors of 0.166 and 0.219 eV for OH and O adsorption, respectively. By targeting the O-OH adsorption energy difference with our GPR model, we predict the the oxygen evolution reaction (OER) activity from bulk descriptors only. Furthermore, we utilize the strong correlation between the COHP and metal oxygen bond lengths to rapidly predict adsorption energetics and catalytic activity from the optimized bulk geometry. Our approach can enable an efficient search for active catalysts by eliminating the need for surface calculations in the initial screening phase.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Active learning for robust, high-complexity reactive atomistic simulations

Machine learned reactive force fields based on polynomial expansions have been shown to be highly effective for describing simulations involving reactive materials. Nevertheless, the highly flexible nature of these models can give rise to a large number of candidate parameters for complicated systems. In these cases, reliable parameterization requires a well-formed training set, which can be difficult to achieve through standard iterative fitting methods. In this paper, we present an active learning approach based on cluster analysis and inspired by Shannon information theory to enable semi-automated generation of informative training sets and robust machine learned force fields. The use of this tool is demonstrated for development of a model based on linear combinations of Chebyshev polynomials explicitly describing up to four-body interactions, for a chemically and structurally diverse system of C/O under extreme conditions. We show that this flexible training database management approach enables development of models exhibiting excellent agreement with Kohn–Sham density functional theory in terms of structure, dynamics, and speciation.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

How Silica Surface Chemistry Modulates Interfacial Water: Insights from Machine Learning Molecular Dynamics

Controlling water structure and dynamics at silica interfaces are central to a wide range of technologies, including protective oxide layers for solar water splitting and nanoporous membranes. In this work, we develop a machine learning interatomic potential, trained via active learning, to achieve ab initio accuracy for water confined between hydroxylated silica surfaces over a range of silanol coverages and slit widths. We find that partially hydroxylated surfaces (50 and 75% OH) support stronger water−surface hydrogen bonding and more extended interfacial density profiles than fully hydroxylated (100% OH) surfaces, indicating that increasing OH coverage does not necessarily strengthen interfacial hydrogenbond networks. Translational diffusion decreases approximately linearly with slit width and OH coverage, whereas rotational dynamics respond nonlinearly. In particular, at the smallest slit width of 5 Å, 75% OH coverage produces an enhanced local tetrahedral ordered interfacial network that strongly suppresses reorientation, while 100% coverage yields a crowded, disordered interfacial layer that also hinders rotation. In contrast, the 50% OH coverage is sufficiently sparse that it does not markedly alter water structure or dynamics under confinement. These results show that coupled control of pore size and surface chemistry enables nonlinear tuning of interfacial water structure and transport, providing a design strategy for optimizing porous silica for either enhanced interfacial stability and controlled reactivity or rapid and selective transport.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Role of Local Structure in the Enhanced Dynamics of Deformed Glasses

External stress can accelerate molecular mobility of amorphous solids by several orders of magnitude. The changes in mobility are commonly interpreted through the Eyring model, which invokes an empirical activation volume. Here, we analyze constant-stress molecular dynamics simulations and propose a structure-dependent Eyring model, connecting activation volume to a machine-learned field, softness. We show that stress has a heterogeneous effect on the mobility that depends on local structure through softness. The barrier impeding relaxation reduces more for well-packed particles, which explains the narrower distribution of relaxation time observed under stress.

74 ATOMIC AND MOLECULAR PHYSICS↗