Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “active machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Experimental discovery of structure–property relationships in ferroelectric materials via active learning

Emergent functionalities of structural and topological defects in ferroelectric materials underpin an extremely broad spectrum of applications ranging from domain wall electronics to high dielectric and electromechanical responses. Many of these functionalities have been discovered and quantified via local scanning probe microscopy methods. However, the search has until now been based on either trial and error, or using auxiliary information such as the topography or domain wall structure to identify potential objects of interest on the basis of the intuition of operator or pre-existing hypotheses, with subsequent manual exploration. Here we report the development and implementation of a machine learning framework that actively discovers relationships between local domain structure and polarization-switching characteristics in ferroelectric materials encoded in the hysteresis loop. The hysteresis loops and their scalar descriptors such as nucleation bias, coercive bias and the hysteresis loop area (or more complex functionals of hysteresis loop shape) and corresponding uncertainties are used to guide the discovery of these relationships via automated piezoresponse force microscopy and spectroscopy experiments. As such, this approach combines the power of machine learning methods to learn the correlative relationships between high-dimensional data, as well as human-based physics insights encoded into the acquisition function. For ferroelectric materials, this automated workflow demonstrates that the discovery path and sampling points of on- and off-field hysteresis loops are largely different, indicating that on- and off-field hysteresis loops are dominated by different mechanisms. Here, the proposed approach is universal and can be applied to a broad range of modern imaging and spectroscopy methods ranging from other scanning probe microscopy modalities to electron microscopy and chemical imaging.

36 MATERIALS SCIENCE↗

Unsupervised learning for identifying events in active target experiments

This article presents novel applications of unsupervised machine learning methods to the problem of event separation in an active target detector, the Active-Target Time Projection Chamber (AT-TPC). The overarching goal is to group similar events in the early stages of the data analysis, thereby improving efficiency by limiting the computationally expensive processing of unnecessary events. The application of unsupervised clustering algorithms to the analysis of two-dimensional projections of particle tracks from a resonant proton scattering experiment on 46 Ar is introduced. We explore the performance of autoencoder neural networks and a pre-trained VGG16 Simonyan and Zisserman (2015) convolutional neural network. We study clustering performance on both data from a simulated 46 Ar experiment, and real events from the AT-TPC detector. We find that a -means algorithm applied to simulated data in the VGG16 latent space forms almost perfect clusters. Additionally, the VGG16+-means approach finds high purity clusters of proton events for real experimental data. Here, we also explore the application of clustering the latent space of autoencoder neural networks for event separation. While these networks show strong performance, they suffer from high variability in their results.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Considerations for using Privacy Preserving Machine Learning Techniques for Safeguards

In international nuclear safeguards, the International Atomic Energy Agency (IAEA) is tasked with inspecting and verifying nuclear facilities and their activities. Data analytics and machine learning to support inspections require large amounts of data that nuclear facility operators may consider proprietary or sensitive, so the IAEA may not have full access. Allowing computation over private data without compromising its security therefore has value for safeguards inspections and analysis. Privacy-preserving machine learning (PPML) consists of security-focused techniques that allow data analytics and machine learning algorithms to run on sensitive data without revealing it. This includes ideas like homomorphic encryption (HE), secure multiparty computation (SMPC), and secure enclaves. HE allows algorithms and mathematical operations to be conducted directly on the encrypted data instead of first decrypting it. With SMPC, multiple entities collaboratively compute over distributed data such that no party is able to directly view any others’ original data. Secure enclaves allow computation to take place in a separate and heavily blocked-off section of a CPU. Techniques like these allow for several potential use cases in which the security of data is essential. With SMPC, machine learning models can be trained over the input data from multiple entities, resulting in a model that all users can benefit from without leaking the input data from any particular entity. With SMPC or a zero-knowledge proof (ZKP), an algorithm returning some single answer or truth value can be run on someone else’s data without ever needing to see that data, potentially allowing for verification or proof of some underlying question. HE can allow for outsourcing computation on data to a hostile or untrusted environment. Although most of the research in this field resides within the health and financial domains, tools from PPML may have similar applications in nuclear safeguards. Allowing the IAEA to compute over proprietary information, such as process models and raw sensor data using PPML techniques, provides the baseline for running complex analytics without needing direct unencrypted access to the underlying data, maintaining its privacy. Important limitations to consider for these techniques include the efficiency and level of security required. The security of HE and SMPC come at the cost of speed—the significant amount of overhead means that algorithms implemented in these protocols and encryption schemes are slower than when run on plaintext. Additionally, several important parameters determine what techniques or protocols are used based on the security requirements. SMPC protocols may need to be selected for resistance against a party that attempts to deviate from the protocol to distort the result or gain access to additional information, and a protocol secure against these attacks may further increase the overhead of the algorithm.

97 MATHEMATICS AND COMPUTING↗

Potentially Underestimated Gas Flaring Activities—A New Approach to Detect Combustion Using Machine Learning and NASA’s Black Marble Product Suite

Monitoring changes in greenhouse gas (GHG) emission is critical for assessing climate mitigation efforts towards the Paris Agreement goal. A crucial aspect of science-based GHG monitoring is to provide objective information for quality assurance and uncertainty assessment of the reported emissions. Emission estimates from combustion events (gas flaring and biomass burning) are often calculated based on activity data (AD) from satellite observations, such as those detected from the visible infrared imaging radiometer suite (VIIRS) onboard the Suomi-NPP and NOAA-20 satellites. These estimates are often incorporated into carbon models for calculating emissions and removals. Consequently, errors and uncertainties associated with AD propagate into these models and impact emission estimates. Deriving uncertainty of AD is therefore crucial for transparency of emission estimates but remains a challenge due to the lack of evaluation data or alternate estimates. This work proposes a new approach using machine learning (ML) for combustion detection from NASA's Black Marble product suite and explores the assessment of potential uncertainties through comparison with existing detections. We jointly characterize combustion using thermal and light emission signals, with the latter improving detection of probable weaker combustion with less distinct thermal signatures. Being methodologically independent, the differences in ML-derived estimates with existing approaches can indicate the potential uncertainties in detection. The approach was applied to detect gas flares over the Eagle Ford Shale, Texas. We analyzed the spatio-temporal variations in detections and found that approximately 79.04% and 72.14% of the light emission-based detections are missed by ML-derived detections from VIIRS thermal bands and existing datasets, respectively. This improvement in combustion detection and scope for uncertainty assessment is essential for comprehensive monitoring of resulting emissions and we discuss the steps for extending this globally.

gas flaring↗

Biosynthesis of bioprivileged, linear molecules via novel carboligase reactions

Over the award period, we made progress on the three aims. We screened twenty-five carboligases for activity coupling twenty-one possible -keto acids (Aim 1). The carboligases were selected across a diverse set of protein sequences. Using Q-Exactive UHPLC-MS, we tested a total of 210 coupled products per enzyme and generated a dataset of 5250 enzyme-substrate activity relationships. We identified multiple enzymes that had activity for synthesizing suberic acid and heptanoic acid (Aim 2). We built a random forest model for predicting the activity of each enzyme toward substrates on which it was not tested using the data from Aim 1. Finally, we evaluated growth defects that occurred due to expression of different carboligases in E. coli (Aim 3). We were able to identify specific metabolites and putative pathways that, when supplemented in the media, recovered the growth defect associated with the presence of specific carboligases. We are in the process of publishing two manuscript describing the methods for high-throughput screening of enzyme promiscuity, using machine learning to predict activity on untested substrates, and enzyme activity data we collected. This project has produced enabling data for biosynthesis of a range of new-to-nature compounds to support biomanufacturing.

60 APPLIED LIFE SCIENCES↗

SIM_EXPLORE: Software for Directed Exploration of Complex Systems

Physics-based numerical simulation codes are widely used in science and engineering to model complex systems that would be infeasible to study otherwise. While such codes may provide the highest- fidelity representation of system behavior, they are often so slow to run that insight into the system is limited. Trying to understand the effects of inputs on outputs by conducting an exhaustive grid-based sweep over the input parameter space is simply too time-consuming. An alternative approach called "directed exploration" has been developed to harvest information from numerical simulators more efficiently. The basic idea is to employ active learning and supervised machine learning to choose cleverly at each step which simulation trials to run next based on the results of previous trials. SIM_EXPLORE is a new computer program that uses directed exploration to explore efficiently complex systems represented by numerical simulations. The software sequentially identifies and runs simulation trials that it believes will be most informative given the results of previous trials. The results of new trials are incorporated into the software's model of the system behavior. The updated model is then used to pick the next round of new trials. This process, implemented as a closed-loop system wrapped around existing simulation code, provides a means to improve the speed and efficiency with which a set of simulations can yield scientifically useful results. The software focuses on the case in which the feedback from the simulation trials is binary-valued, i.e., the learner is only informed of the success or failure of the simulation trial to produce a desired output. The software offers a number of choices for the supervised learning algorithm (the method used to model the system behavior given the results so far) and a number of choices for the active learning strategy (the method used to choose which new simulation trials to run given the current behavior model). The software also makes use of the LEGION distributed computing framework to leverage the power of a set of compute nodes. The approach has been demonstrated on a planetary science application in which numerical simulations are used to study the formation of asteroid families.

Burl, Michael↗

Machine learning and process-based modeling of spatiotemporal changes in active layer thickness across Alaska

Permafrost degradation poses a growing threat to infrastructure stability and ecosystem resilience in the rapidly warming Arctic. We investigated the spatiotemporal dynamics of active layer thickness (ALT) across Alaska by integrating field observations, environmental datasets, a physically based Stefan model, and machine learning (ML) techniques. Using weather projections from the Coupled Model Intercomparison Project Phase 6 under two Shared Socioeconomic Pathways (SSP 2-4.5 and SSP 5-8.5), we assessed ALT sensitivity to projected future weather conditions. The random forest (RF) model outperformed the Stefan approach in predicting ALT on the training dataset (R² = 0.84 vs. 0.53) but demonstrated lower generalizability on the test dataset (R² = 0.24 vs. 0.54). The root mean square error (RMSE) for the RF model for training and testing ranged from 14 to 22 cm, compared to 17 and 18 cm for the Stefan model. Variable importance analysis revealed that mean annual temperature and slope angle were the strongest predictors of ALT, accounting for 19% and 18% of the variance, respectively, followed by sediment transport index (14%) and stream power index (11%). Comparative analysis of baseline ALT predictions showed the Stefan model tended to project a thicker active layer (mean ± SD: 65 ± 16 cm), compared to the RF model (mean ± SD: 59 ± 8.8) cm). Both models indicated a latitudinal gradient in ALT, with shallower depths at higher latitudes. Projected ALT increases by 2100 were estimated at 3.3 ± 2.2 cm under SSP 2-4.5 and 5.9 ± 4.0 cm under SSP 5-8.5 for the ML model, whereas the Stefan model projected substantially larger increases of 13 ± 2.6 cm (SSP 2-4.5) and 28 ± 4.4 cm (SSP5-8.5). Spatial analysis showed the greatest ALT increases in northern Alaska, with relatively smaller changes in southern regions. These findings highlight the complex, multifactorial nature of ALT dynamics and the value of hybrid modeling approaches. As rising temperatures accelerate permafrost thaw, changes in ALT can disrupt ecosystems, damage infrastructures, and enhance the release of stored soil carbon, highlighting the urgent need for improved predictive capabilities to inform adaptation strategies in the Arctic.

Climate sciences↗

AutoTandemML: Active Learning Enhanced Tandem Neural Networks for Inverse Design Problems

Inverse design in science and engineering involves determining optimal design parameters that achieve desired performance outcomes, a process often hindered by the complexity and high dimensionality of design spaces, leading to significant computational costs. To tackle this challenge, we propose a novel hybrid approach that combines active learning with Tandem Neural Networks to enhance the efficiency and effectiveness of solving inverse design problems. Active learning allows to selectively sample the most informative data points, reducing the required dataset size without compromising accuracy. We investigate this approach using three benchmark problems: airfoil inverse design, photonic surface inverse design, and scalar boundary condition reconstruction in diffusion partial differential equations. We demonstrate that integrating active learning with Tandem Neural Networks outperforms standard approaches across the benchmark suite, achieving better accuracy with fewer training samples.

97 MATHEMATICS AND COMPUTING↗

An autonomous laboratory for the accelerated synthesis of novel materials

To close the gap between the rates of computational screening and experimental realization of novel materials, we introduce the A-Lab, an autonomous laboratory for the solid-state synthesis of inorganic powders. This platform uses computations, historical data from the literature, machine learning (ML) and active learning to plan and interpret the outcomes of experiments performed using robotics. Over 17 days of continuous operation, the A-Lab realized 41 novel compounds from a set of 58 targets including a variety of oxides and phosphates that were identified using large-scale ab initio phase-stability data from the Materials Project and Google DeepMind. Synthesis recipes were proposed by natural-language models trained on the literature and optimized using an active-learning approach grounded in thermodynamics. Analysis of the failed syntheses provides direct and actionable suggestions to improve current techniques for materials screening and synthesis design. The high success rate demonstrates the effectiveness of artificial-intelligence-driven platforms for autonomous materials discovery and motivates further integration of computations, historical knowledge and robotics.

36 MATERIALS SCIENCE↗

Bayesian inference analysis of jet quenching using inclusive jet and hadron suppression measurements

The JETSCAPE Collaboration reports a new determination of the jet transport parameter $\hat{q}$ in the quark-gluon plasma (QGP) using Bayesian inference, incorporating all available inclusive hadron and jet yield suppression data measured in heavy-ion collisions at the BNL Relativistic Heavy Ion Collider (RHIC) and the CERN Large Hadron Collider (LHC). This multi-observable analysis extends the previously published JETSCAPE Bayesian inference determination of $\hat{q}$, which was based solely on a selection of inclusive hadron suppression data. jetscape is a modular framework incorporating detailed dynamical models of QGP formation and evolution, and jet propagation and interaction in the QGP. Virtuality-dependent partonic energy loss in the QGP is modeled as a thermalized weakly coupled plasma, with parameters determined from Bayesian calibration using soft-sector observables. This Bayesian calibration of $\hat{q}$ utilizes active learning, a machine-learning approach, for efficient exploitation of computing resources. The experimental data included in this analysis span a broad range in collision energy and centrality, and in transverse momentum. In order to explore the systematic dependence of the extracted parameter posterior distributions, several different calibrations are reported, based on combined jet and hadron data; on jet or hadron data separately; and on restricted kinematic or centrality ranges of the jet and hadron data. Tension is observed in comparison of these variations, providing new insights into the physics of jet transport in the QGP and its theoretical formulation.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

An intelligent Data Delivery Service for and beyond the ATLAS experiment

The intelligent Data Delivery Service (iDDS) has been developed to cope with the huge increase of computing and storage resource usage in the coming LHC data taking. It has been designed to intelligently orchestrate workflows and data management systems, decoupling data pre-processing, delivery, and primary processing in large scale workflows. It is an experiment-agnostic service that has been deployed to serve data carousel (orchestrating efficient processing of tape-resident data), machine learning hyperparameter optimization, active learning, and other complex multi-stage workflows defined via DAG (Directed Acyclic Graph), CWL (Common Workflow Language) and other descriptions, including a growing number of analysis workflows. We will at first introduce some deployed use cases in a summary. Then we will focus on new improvements and use cases under developments in ATLAS, Rubin Observatory and sPHENIX, together with future efforts.

97 MATHEMATICS AND COMPUTING↗

Performance Analysis of an Optimization Algorithm for Metamaterial Design on the Integrated High-Performance Computing and Quantum Systems

Optimizing metamaterials with complex geometries is a big challenge. Although an active learning algorithm, combining machine learning (ML), quantum computing, and optical simulation, has emerged as an efficient optimization tool, it still faces difficulties in optimizing complex structures that have potentially high performance. In this work, we comprehensively analyze the performance of an optimization algorithm for metamaterial design on the integrated HPC and quantum systems. We demonstrate significant time advantages through message-passing interface (MPI) parallelization on the high-performance computing (HPC) system showing approximately 54% faster ML tasks and 67 times faster optical simulation against serial workloads. Furthermore, we analyze the performance of a quantum algorithm designed for optimization, which runs with various quantum simulators on a local computer or HPC-quantum system. Results showcase ~24 times speedup when executing the optimization algorithm on the HPC-quantum hybrid system. This study paves a way to optimize complex metamaterials using the integrated HPC-quantum system.

Kim, Seongmin↗

Dial

A key step in almost all scientific endeavors is answering the question: Given this data I already collected, what new data do I expect will yield the most useful information toward my scientific objective? The area of (sequential) experimental design has long been investigating answers to this question, but in recent years techniques from the machine learning subfield of active learning are increasingly applied. Researchers need a simple software tool for active learning applied to experimental design that can easily integrate into their existing workflows. This computer code, Dial, provides a microservice in ORNL's INTERSECT ecosystem for active learning applied to experimental design. By being part of the INTERSECT ecosystem, Dial is simple to integrate into any INTERSECT-based workflow. Dial provides multiple backend options, where a backend is an implementation of a specific active learning method. Users can select the backend that performs best for their application. Developers can also add new backends as needed. At its core, Dial receives a set of pre-existing measurements and input parameter bounds and then recommends one or more new sets of parameters to measure. Dial also includes interfaces to other microservices in the INTERSECT ecosystem so that it can be incorporated into INTERSECT campaigns. Dial provides a simple, yet powerful interface to convert automated INTERSECT workflows into autonomous workflows that adapt based on the results that are obtained. A shared microservice for active learning prevents duplicated effort by each application team implementing its own adaptive design of experiments tool.

Drane, Lance [Oak Ridge National Laboratory (ORNL)↗

Measuring Constraint-Set Utility for Partitional Clustering Algorithms

Clustering with constraints is an active area of machine learning and data mining research. Previous empirical work has convincingly shown that adding constraints to clustering improves the performance of a variety of algorithms. However, in most of these experiments, results are averaged over different randomly chosen constraint sets from a given set of labels, thereby masking interesting properties of individual sets. We demonstrate that constraint sets vary significantly in how useful they are for constrained clustering; some constraint sets can actually decrease algorithm performance. We create two quantitative measures, informativeness and coherence, that can be used to identify useful constraint sets. We show that these measures can also help explain differences in performance for four particular constrained clustering algorithms.

constraints↗

Probing Active Sites in Cu x Pd y Cluster Catalysts by Machine-Learning-Assisted X-ray Absorption Spectroscopy

Size-selected clusters are important model catalysts because of their narrow size and compositional distributions, as well as enhanced activity and selectivity in many reactions. Still, their structure-activity relationships are, in general, elusive. The main reason is the difficulty in identifying and quantitatively characterizing the catalytic active site in the clusters when it is confined within subnanometric dimensions and under the continuous structural changes the clusters can undergo in reaction conditions. Using machine learning approaches for analysis of the operando X-ray absorption near-edge structure spectra, we obtained accurate speciation of the Cu x Pd y cluster types during the propane oxidation reaction and the structural information about each type. As a result, we elucidated the information about active species and relative roles of Cu and Pd in the clusters.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗