Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “active machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Bayesian inference analysis of jet quenching using inclusive jet and hadron suppression measurements

The JETSCAPE Collaboration reports a new determination of the jet transport parameter $\hat{q}$ in the quark-gluon plasma (QGP) using Bayesian inference, incorporating all available inclusive hadron and jet yield suppression data measured in heavy-ion collisions at the BNL Relativistic Heavy Ion Collider (RHIC) and the CERN Large Hadron Collider (LHC). This multi-observable analysis extends the previously published JETSCAPE Bayesian inference determination of $\hat{q}$, which was based solely on a selection of inclusive hadron suppression data. jetscape is a modular framework incorporating detailed dynamical models of QGP formation and evolution, and jet propagation and interaction in the QGP. Virtuality-dependent partonic energy loss in the QGP is modeled as a thermalized weakly coupled plasma, with parameters determined from Bayesian calibration using soft-sector observables. This Bayesian calibration of $\hat{q}$ utilizes active learning, a machine-learning approach, for efficient exploitation of computing resources. The experimental data included in this analysis span a broad range in collision energy and centrality, and in transverse momentum. In order to explore the systematic dependence of the extracted parameter posterior distributions, several different calibrations are reported, based on combined jet and hadron data; on jet or hadron data separately; and on restricted kinematic or centrality ranges of the jet and hadron data. Tension is observed in comparison of these variations, providing new insights into the physics of jet transport in the QGP and its theoretical formulation.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

An intelligent Data Delivery Service for and beyond the ATLAS experiment

The intelligent Data Delivery Service (iDDS) has been developed to cope with the huge increase of computing and storage resource usage in the coming LHC data taking. It has been designed to intelligently orchestrate workflows and data management systems, decoupling data pre-processing, delivery, and primary processing in large scale workflows. It is an experiment-agnostic service that has been deployed to serve data carousel (orchestrating efficient processing of tape-resident data), machine learning hyperparameter optimization, active learning, and other complex multi-stage workflows defined via DAG (Directed Acyclic Graph), CWL (Common Workflow Language) and other descriptions, including a growing number of analysis workflows. We will at first introduce some deployed use cases in a summary. Then we will focus on new improvements and use cases under developments in ATLAS, Rubin Observatory and sPHENIX, together with future efforts.

97 MATHEMATICS AND COMPUTING↗

Performance Analysis of an Optimization Algorithm for Metamaterial Design on the Integrated High-Performance Computing and Quantum Systems

Optimizing metamaterials with complex geometries is a big challenge. Although an active learning algorithm, combining machine learning (ML), quantum computing, and optical simulation, has emerged as an efficient optimization tool, it still faces difficulties in optimizing complex structures that have potentially high performance. In this work, we comprehensively analyze the performance of an optimization algorithm for metamaterial design on the integrated HPC and quantum systems. We demonstrate significant time advantages through message-passing interface (MPI) parallelization on the high-performance computing (HPC) system showing approximately 54% faster ML tasks and 67 times faster optical simulation against serial workloads. Furthermore, we analyze the performance of a quantum algorithm designed for optimization, which runs with various quantum simulators on a local computer or HPC-quantum system. Results showcase ~24 times speedup when executing the optimization algorithm on the HPC-quantum hybrid system. This study paves a way to optimize complex metamaterials using the integrated HPC-quantum system.

Kim, Seongmin↗

Dial

A key step in almost all scientific endeavors is answering the question: Given this data I already collected, what new data do I expect will yield the most useful information toward my scientific objective? The area of (sequential) experimental design has long been investigating answers to this question, but in recent years techniques from the machine learning subfield of active learning are increasingly applied. Researchers need a simple software tool for active learning applied to experimental design that can easily integrate into their existing workflows. This computer code, Dial, provides a microservice in ORNL's INTERSECT ecosystem for active learning applied to experimental design. By being part of the INTERSECT ecosystem, Dial is simple to integrate into any INTERSECT-based workflow. Dial provides multiple backend options, where a backend is an implementation of a specific active learning method. Users can select the backend that performs best for their application. Developers can also add new backends as needed. At its core, Dial receives a set of pre-existing measurements and input parameter bounds and then recommends one or more new sets of parameters to measure. Dial also includes interfaces to other microservices in the INTERSECT ecosystem so that it can be incorporated into INTERSECT campaigns. Dial provides a simple, yet powerful interface to convert automated INTERSECT workflows into autonomous workflows that adapt based on the results that are obtained. A shared microservice for active learning prevents duplicated effort by each application team implementing its own adaptive design of experiments tool.

Drane, Lance [Oak Ridge National Laboratory (ORNL)↗

Probing Active Sites in Cu x Pd y Cluster Catalysts by Machine-Learning-Assisted X-ray Absorption Spectroscopy

Size-selected clusters are important model catalysts because of their narrow size and compositional distributions, as well as enhanced activity and selectivity in many reactions. Still, their structure-activity relationships are, in general, elusive. The main reason is the difficulty in identifying and quantitatively characterizing the catalytic active site in the clusters when it is confined within subnanometric dimensions and under the continuous structural changes the clusters can undergo in reaction conditions. Using machine learning approaches for analysis of the operando X-ray absorption near-edge structure spectra, we obtained accurate speciation of the Cu x Pd y cluster types during the propane oxidation reaction and the structural information about each type. As a result, we elucidated the information about active species and relative roles of Cu and Pd in the clusters.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Atomic-scale identification of active sites of oxygen reduction nanocatalysts

Heterogeneous nanocatalysts play a crucial role in both the chemical and energy industries. Despite substantial advancements in theoretical, computational and experimental studies, identifying their active sites remains a major challenge. Here we utilize atomic electron tomography to determine the three-dimensional atomic structure of PtNi and Mo-doped PtNi nanocatalysts for the electrochemical oxygen reduction reaction. We then employ the experimental atomic structures as input to first-principles-trained machine learning to identify the active sites of the nanocatalysts. Through the analysis of the structure–activity relationships, we formulate an equation termed the local environment descriptor, which balances the strain and ligand effects to provide physical and chemical insights into active sites in the oxygen reduction reaction. The ability to determine the three-dimensional atomic structure and chemical composition of realistic nanoparticles, combined with machine learning, could transform our fundamental understanding of the active sites of catalysts and guide the rational design of optimal nanocatalysts.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Machine Learning-Guided Identification of PET Hydrolases from Natural Diversity

The enzymatic depolymerization of poly(ethylene terephthalate) (PET) is emerging as a leading chemical recycling technology for waste polyester. As part of this endeavor, new candidate enzymes identified from natural diversity can serve as useful starting points for enzyme evolution and engineering. In this study, we improved upon HMM searches by applying an iterative machine learning strategy to identify 400 putative PET-degrading enzymes (PET hydrolases) from naturally occurring homologs. Using high-throughput (HTP) experimental techniques, we successfully expressed and purified >200 enzyme candidates and assayed them for PET hydrolysis activity as a function of pH, temperature, and substrate crystallinity. From this library, we discovered 91 previously unknown PET hydrolases, 35 of which retain activity at pH 4.5 on crystalline material, which are conditions relevant to developing more efficient commercial processes. Notably, four enzymes showed equal to or higher activity than LCC-ICCG, a benchmark PET hydrolase, at this challenging condition in our screening assay, and 11 of which have pH optima <7. Using these data, we identified regions of PETases statistically correlated to activity at lower pH. We additionally investigated the effect of condition-specific activity data on trained machine learning predictors and found a precision (putative hit rate) improvement of up to 30% compared to a Hidden Markov Model alone. Our findings show that by pointing enzyme discovery toward conditions of interest with multiple rounds of experimental and machine learning, we can discover large sets of active enzymes and explore factors associated with activity at those conditions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

High temperature oxidation of corrosion resistant alloys from machine learning

Parabolic rate constants, k p , were collected from published reports and calculated from corrosion product data (sample mass gain or corrosion product thickness) and tabulated for 75 alloys exposed to temperatures between ~800 and 2000 K (~500–1700 °C; 900–3000°F). Data were collected for environments including lab air, ambient and supercritical carbon dioxide, supercritical water, and steam. Materials studied include low- and high-Cr ferritic and austenitic steels, nickel superalloys, and aluminide materials. A combination of Arrhenius analysis, simple linear regression, supervised and unsupervised machine learning methods were used to investigate the relations between composition and oxidation kinetics. The supervised machine learning techniques produced the lowest mean standard errors. The most significant elements controlling oxidation kinetics were Ni, Cr, Al, and Fe, with Mo and Co composition also found to be significant features. The activation energies produced from the machine learning analysis were in the correct distributions for the diffusion constants for the oxide scales expected to dominate in each class.

Materials Science↗

Learning protocols for the fast and efficient control of active matter

Exact analytic calculation shows that optimal control protocols for passive molecular systems often involve rapid variations and discontinuities. However, similar analytic baselines are not generally available for active-matter systems, because it is more difficult to treat active systems exactly. Here we use machine learning to derive efficient control protocols for active-matter systems, and find that they are characterized by sharp features similar to those seen in passive systems. We show that it is possible to learn protocols that effect fast and efficient state-to-state transformations in simulation models of active particles by encoding the protocol in the form of a neural network. We use evolutionary methods to identify protocols that take active particles from one steady state to another, as quickly as possible or with as little energy expended as possible. Our results show that protocols identified by a flexible neural-network ansatz, which allows the optimization of multiple control parameters and the emergence of sharp features, are more efficient than protocols derived recently by constrained analytical methods. Our learning scheme is straightforward to use in experiment, suggesting a way of designing protocols for the efficient manipulation of active matter in the laboratory.

74 ATOMIC AND MOLECULAR PHYSICS↗

Rational design of heterogeneous single-site catalysts via surface organometallic chemistry

Single-site heterogeneous catalysts offer an attractive route to unite the molecular precision of homogeneous catalysis with the durability and practical advantages of solids. Surface organometallic chemistry (SOMC) provides a particularly powerful strategy for this purpose by grafting molecular precursors onto tailored surfaces and converting support functionalities into ligand environments for isolated metal centers. As a result, SOMC brings the language and logic of coordination chemistry to heterogeneous catalysis, where the support becomes an integral part of the active site coordination sphere. This Review surveys recent progress in the rational design of SOMC-derived single-site catalysts, with emphasis on synthetic routes, post synthetic transformations, and the deliberate tuning of catalytic behavior through metal-support interactions. Discussions are made on how support identity, hydroxyl topology, acidity, and redox activity shape the geometry, electronic structure, and oxidation state of supported metal sites, as well as how these factors determine activity, selectivity, and stability. We also examine a central limitation of these systems: despite their molecularly informed design, supported single sites often exist as structurally distributed ensembles rather than uniform species, particularly on amorphous supports. This site heterogeneity, along with catalyst dynamics under operating conditions, remains a major barrier to definitive structure-activity relationships. Therefore, emerging approaches that combine advanced characterization, first-principles modeling, ensemble kinetics, and machine learning to resolve active-site structure and guide catalyst development are highlighted. Together, these advances position SOMC as a versatile coordination chemistry framework for the predictive design of heterogeneous catalysts with well-defined molecularly tailored active sites.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Uncovering electronic and geometric descriptors of chemical activity for metal alloys and oxides using unsupervised machine learning

Here, we show that unsupervised machine learning (ML) using principal component analysis (PCA) provides a straightforward pathway for developing accurate and interpretable electronic-structure descriptors of the chemical and catalytic properties of materials. We demonstrate the approach by finding chemisorption descriptors for metal alloys and surface oxygens on metals and metal oxides. In both cases, the principal component (PC) descriptors yield ML models that predict the material’s chemical properties with competitive accuracy compared to ML models built using established descriptors. Importantly, interpreting the electronic-structure patterns captured by each PC descriptor via signal reconstruction suggests potential design motifs for future electronic-structure descriptor design and allows us to identify links between a material’s geometric and catalytic properties. Ultimately, we show that the unsupervised ML approach provides a route to find electronic-structure descriptors of the catalytic properties of materials that readily connect to geometric structure and composition.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A novel methodology for gamma-ray spectra dataset procurement over varying standoff distances and source activities

The adoption of machine learning approaches for gamma-ray spectroscopy has received considerable attention in the literature. Many studies have investigated the deployment of various algorithm architectures to a specific task. However, little attention has been afforded to the development of the datasets leveraged to train the models. Such training datasets typically span a set of environmental or detector parameters to encompass a problem space of interest to a user. Variations in these measurement parameters will also induce fluctuations in the detector response, including expected pile-up and ground scatter effects. Fundamental to this work is the understanding that 1) the underlying spectral shape varies as the measurement parameters change and 2) the statistical uncertainties associated with two spectra impact their level of similarity. While previous studies attribute some arbitrary discretization to the measurement parameters for the generation of their synthetic training data, this work introduces a principled methodology for efficient spectral-based discretization of a problem space. A signal-to-noise ratio (SNR) respective spectral comparison measure and a Gaussian Process Regression (GPR) model are used to predict the spectral similarity across a range of measurement parameters. This innovative approach effectively showcased its capability by dividing a problem space, ranging from 5 cm to 100 cm standoff distances and 5 μCi–100 μCi of 137 Cs, into three unique combinations of measurement parameters. The findings from this work will aid in creating more robust datasets, which incorporate many possible measurement scenarios, reduce the number of required experimental test set measurements, and possibly enable experimental training data collection for gamma-ray spectroscopy.

data science↗

The Deeper, Wider, Faster programme: exploring stellar flare activity with deep, fast cadenced DECam imaging via machine learning

ABSTRACT We present our 500 pc distance-limited study of stellar flares using the Dark Energy Camera as part of the Deeper, Wider, Faster programme. The data were collected via continuous 20-s cadence g-band imaging and we identify 19 914 sources with precise distances from Gaia DR2 within 12, ∼3 deg2, fields over a range of Galactic latitudes. An average of ∼74 min is spent on each field per visit. All light curves were accessed through a novel unsupervised machine learning techniques designed for anomaly detection. We identify 96 flare events occurring across 80 stars, the majority of which are M dwarfs. Integrated flare energies range from ∼1031–1037 erg, with a proportional relationship existing between increased flare energy with increased distance from the Galactic plane, representative of stellar age leading to declining yet more energetic flare events. In agreement with previous studies we observe an increase in flaring fraction from M0 to M6 spectral types. Furthermore, we find a decrease in the flaring fraction of stars as vertical distance from the galactic plane is increased, with a steep decline present around ∼100 pc. We find that $\sim 70{{\ \rm per\ cent}}$ of identified flares occur on short time-scales of <8 min. Finally, we present our associated flare rates, finding a volumetric rate of 2.9 ± 0.3 × 10−6 flares pc−3 h−1.

Webb, S.↗

Objective Phenotyping of Root System Architecture Using Image Augmentation and Machine Learning in Alfalfa (Medicago sativa L.)

Active breeding programs specifically for root system architecture (RSA) phenotypes remain rare; however, breeding for branch and taproot types in the perennial crop alfalfa is ongoing. Phenotyping in this and other crops for active RSA breeding has mostly used visual scoring of specific traits or subjective classification into different root types. While image-based methods have been developed, translation to applied breeding is limited. This research is aimed at developing and comparing image-based RSA phenotyping methods using machine and deep learning algorithms for objective classification of 617 root images from mature alfalfa plants collected from the field to support the ongoing breeding efforts. Our results show that unsupervised machine learning tends to incorrectly classify roots into a normal distribution with most lines predicted as the intermediate root type. Encouragingly, random forest and TensorFlow-based neural networks can classify the root types into branch-type, taproot-type, and an intermediate taproot-branch type with 86% accuracy. With image augmentation, the prediction accuracy was improved to 97%. Coupling the predicted root type with its prediction probability will give breeders a confidence level for better decisions to advance the best and exclude the worst lines from their breeding program. This machine and deep learning approach enables accurate classification of the RSA phenotypes for genomic breeding of climate-resilient alfalfa.

59 BASIC BIOLOGICAL SCIENCES↗

AL4GAP: Active learning workflow for generating DFT-SCAN accurate machine-learning potentials for combinatorial molten salt mixtures

Machine learning interatomic potentials have emerged as a powerful tool for bypassing the spatiotemporal limitations of ab initio simulations, but major challenges remain in their efficient parameterization. We present AL4GAP, an ensemble active learning software workflow for generating multicomposition Gaussian approximation potentials (GAP) for arbitrary molten salt mixtures. The workflow capabilities include: (1) setting up user-defined combinatorial chemical spaces of charge neutral mixtures of arbitrary molten mixtures spanning 11 cations (Li, Na, K, Rb, Cs, Mg, Ca, Sr, Ba and two heavy species, Nd, and Th) and 4 anions (F, Cl, Br, and I), (2) configurational sampling using low-cost empirical parameterizations, (3) active learning for down-selecting configurational samples for single point density functional theory calculations at the level of Strongly Constrained and Appropriately Normed (SCAN) exchange-correlation functional, and (4) Bayesian optimization for hyperparameter tuning of two-body and many-body GAP models. Here, we apply the AL4GAP workflow to showcase high throughput generation of five independent GAP models for multicomposition binary-mixture melts, each of increasing complexity with respect to charge valency and electronic structure, namely: LiCl–KCl, NaCl–CaCl 2 , KCl–NdCl 3 , CaCl 2 –NdCl 3 , and KCl–ThCl 4 . Our results indicate that GAP models can accurately predict structure for diverse molten salt mixture with density functional theory (DFT)-SCAN accuracy, capturing the intermediate range ordering characteristic of the multivalent cationic melts.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Machine Learning Using Open Data Sources for Detection of Nuclear Proliferation Activities (U)

In FY2020, Savannah River National Laboratory (SRNL) in collaboration with the Sanghani Center for Artificial Intelligence and Data Analytics (SCAIDA) at Virginia Polytechnic Institute and State University (VT) and funded by the Department of Energy’s (DOE) Defense Nuclear Nonproliferation Research and Development, began developing a demonstration prototype system that uses multiple machine learning and data analytic methods on large-scale open data sources to identify new, developing, and/or undeclared nuclear programs. Using the announcement in May 2018 of the proposed Savannah River Plutonium Processing Facility (SRPPF) as a test subject, the goal of this 2-year project is to forecast the SRPPF using only data prior to May 2018. The project work is split into a preliminary prototype development for the first year with an initial evaluation of viability followed by the second year of development to create an integrated prototype system and more extensive performance evaluation. This report documents the results of the preliminary-phase tasks.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Explainable machine learning for hydrogen diffusion in metals and random binary alloys

Hydrogen diffusion in metals and alloys plays an important role in the discovery of new materials for fuel cell and energy storage technology. While analytic models use hand-selected features that have clear physical ties to hydrogen diffusion, they often lack accuracy when making quantitative predictions. Machine learning models are capable of making accurate predictions, but their inner workings are obscured, rendering it unclear which physical features are truly important. To develop interpretable machine learning models to predict the activation energies of hydrogen diffusion in metals and random binary alloys, we create a database for physical and chemical properties of the species and use it to fit six machine learning models. Our models achieve root-mean-squared errors between 98–119 meV on the testing data and accurately predict that elemental Ru has a large activation energy, while elemental Cr and Fe have small activation energies. By analyzing the feature importances of these fitted models, we identify relevant physical properties for predicting hydrogen diffusivity. While metrics for measuring the individual feature importances for machine learning models exist, correlations between the features lead to disagreement between models and limit the conclusions that can be drawn. Instead grouped feature importance, formed by combining the features via their correlations, agree across the six models and reveal that the two groups containing the packing factor and electronic specific heat are particularly significant for predicting hydrogen diffusion in metals and random binary alloys. In conclusion, this framework allows us to interpret machine learning models and enables rapid screening of new materials with the desired rates of hydrogen diffusion.

36 MATERIALS SCIENCE↗