Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “active machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Low Activity Waste Glass Optimization with Property Models from Machine Learning, Part 2: Experimental Validation and Active Learning

The United States Department of Energy is responsible for managing legacy nuclear waste stored in underground tanks at the Hanford Site. To treat the waste, it is planned as the current baseline to separately vitrify low-activity waste (LAW) and high-level waste fractions. Previously, machine learning (ML) based glass property models (e.g., chemical durability, viscosity, electrical conductivity and SO3 solubility) were developed with prediction uncertainties. A waste glass optimization approach was then established to enable the capability of using these ML models in LAW glass formulation. In this study, the previous ML models were first experimentally validated, and the results were incorporated back into the database to update the ML models. The updated models and formulations showed increased waste loading while reducing the failure rate, demonstrating improved predictive accuracy, reduced uncertainties, and the effectiveness of active learning in guiding high-dimensional, nonlinear LAW glass design. This represents the first experimental validation of ML based LAW glass formulation, with practical benefits such as higher waste loading, shorter mission duration, and lower operational risk.

Lu, Xiaonan (ORCID:0000000179708148)↗

Learning Extended Finite State Machines

We present an active learning algorithm for inferring extended finite state machines (EFSM)s, combining data flow and control behavior. Key to our learning technique is a novel learning model based on so-called tree queries. The learning algorithm uses the tree queries to infer symbolic data constraints on parameters, e.g., sequence numbers, time stamps, identifiers, or even simple arithmetic. We describe sufficient conditions for the properties that the symbolic constraints provided by a tree query in general must have to be usable in our learning model. We have evaluated our algorithm in a black-box scenario, where tree queries are realized through (black-box) testing. Our case studies include connection establishment in TCP and a priority queue from the Java Class Library.

Register Automata↗

Understanding Heating in Active Region Cores through Machine Learning. I. Numerical Modeling and Predicted Observables

To adequately constrain the frequency of energy deposition in active region cores in the solar corona, systematic comparisons between detailed models and observational data are needed. In this paper, we describe a pipeline for forward modeling active region emission using magnetic field extrapolations and field-aligned hydrodynamic models. We use this pipeline to predict time-dependent emission from active region NOAA 1158 for low-, intermediate-, and high-frequency nanoflares. In each pixel of our predicted multi-wavelength, time-dependent images, we compute two commonly used diagnostics: the emission measure slope and the time lag. We find that signatures of the heating frequency persist in both of these diagnostics. In particular, our results show that the distribution of emission measure slopes narrows and the mean decreases with decreasing heating frequency and that the range of emission measure slopes is consistent with past observational and modeling work. Furthermore, we find that the time lag becomes increasingly spatially coherent with decreasing heating frequency while the distribution of time lags across the whole active region becomes more broad with increasing heating frequency. In a follow-up paper, we train a random forest classifier on these predicted diagnostics and use this model to classify real observations of NOAA 1158 in terms of the underlying heating frequency.

UV radiation↗

Active learning using hybrid surrogate tool life modeling for machining process optimization

Here, this paper describes an active learning approach for part-to-part iterative machining process optimization using a hybrid surrogate tool life model. A probabilistic interpolating tool life model is developed by combining the empirical Taylor-type tool life equation and the model fit error. The probabilistic tool life model is then used to calculate the machining cost per part distribution. The optimal machining parameters are selected using an expected improvement in machining cost per part criterion. The method is validated numerically using experimental results; the results show a median convergence error of 2.2% after three tests over 400 simulations. The method is validated experimentally on two industrial applications for Ti-6Al-4V roughing resulting in a cost per part reduction greater than 23% after two tests. The described method is a robust solution for rapid convergence to optimal machining parameters in an industrial production environment.

Active learning↗

AN AUTOMATED MACHINE LEARNING-GENETIC ALGORITHM FRAMEWORK WITH ACTIVE LEARNING FOR DESIGN OPTIMIZATION

The use of machine learning (ML)-based surrogate models is a promising technique to significantly accelerate simulation-driven design optimization of internal combustion (IC) engines, due to the high computational cost of running computational fluid dynamics (CFD) simulations. However, training the ML models requires hyperparameter selection, which is often done using trial-and-error and domain expertise. Another challenge is that the data required to train these models are often unknown a priori. In this work, we present an automated hyperparameter selection technique coupled with an active learning approach to address these challenges. The technique presented in this study involves the use of a Bayesian approach to optimize the hyperparameters of the base learners that make up a super learner model. In addition to performing hyperparameter optimization (HPO), an active learning approach is employed, where the process of data generation using simulations, ML training, and surrogate optimization is performed repeatedly to refine the solution in the vicinity of the predicted optimum. The proposed approach is applied to the optimization of a compression ignition engine with control parameters relating to fuel injection, in-cylinder flow, and thermodynamic conditions. It is demonstrated that by automatically selecting the best values of the hyperparameters, a 1.6% improvement in merit value is obtained, compared to an improvement of 1.0% with default hyperparameters. Overall, the framework introduced in this study reduces the need for technical expertise in training ML models for optimization while also reducing the number of simulations needed for performing surrogate-based design optimization.

Owoyele, Opeoluwa↗

Using Active Learning to Rapidly Develop Machine Learned Diffusion Coefficients of CO 2 Conversion Reagents in Metal–Organic Frameworks

Here, we used a combined molecular dynamics/active learning (AL) approach to create machine learning models that can predict the diffusion coefficient of epichlorohydrin and chloropropene carbonate, the reactant and product of a common CO 2 cycloaddition reaction, in metal–organic frameworks (MOFs). Nanoporous MOFs are effective catalysts for the cycloaddition of CO 2 to epoxides. The diffusion rates within nanoporous catalysts can control the rate of reaction as the reactants and products must diffuse to the active sites within the MOF and then out of the nanoporous material for reusability. However, the diffusion process is routinely ignored when searching for new materials in catalytic applications. Here we verified improvement during the AL process by consistently tracking metrics on the same groups of MOFs to ensure consistency. Metal identity was found to have little impact on diffusion rates, while structural features like pore limiting diameter act as a threshold where a minimum value is needed for high diffusion rates. We identified the MOFs with the highest epichlorohydrin and chloropropene carbonate diffusion coefficients which can be used for further studies of reaction energetics.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Large-scale genomic analyses with machine learning uncover predictive patterns associated with fungal phytopathogenic lifestyles and traits

Abstract Invasive plant pathogenic fungi have a global impact, with devastating economic and environmental effects on crops and forests. Biosurveillance, a critical component of threat mitigation, requires risk prediction based on fungal lifestyles and traits. Recent studies have revealed distinct genomic patterns associated with specific groups of plant pathogenic fungi. We sought to establish whether these phytopathogenic genomic patterns hold across diverse taxonomic and ecological groups from the Ascomycota and Basidiomycota, and furthermore, if those patterns can be used in a predictive capacity for biosurveillance. Using a supervised machine learning approach that integrates phylogenetic and genomic data, we analyzed 387 fungal genomes to test a proof-of-concept for the use of genomic signatures in predicting fungal phytopathogenic lifestyles and traits during biosurveillance activities. Our machine learning feature sets were derived from genome annotation data of carbohydrate-active enzymes (CAZymes), peptidases, secondary metabolite clusters (SMCs), transporters, and transcription factors. We found that machine learning could successfully predict fungal lifestyles and traits across taxonomic groups, with the best predictive performance coming from feature sets comprising CAZyme, peptidase, and SMC data. While phylogeny was an important component in most predictions, the inclusion of genomic data improved prediction performance for every lifestyle and trait tested. Plant pathogenicity was one of the best-predicted traits, showing the promise of predictive genomics for biosurveillance applications. Furthermore, our machine learning approach revealed expansions in the number of genes from specific CAZyme and peptidase families in the genomes of plant pathogens compared to non-phytopathogenic genomes (saprotrophs, endo- and ectomycorrhizal fungi). Such genomic feature profiles give insight into the evolution of fungal phytopathogenicity and could be useful to predict the risks of unknown fungi in future biosurveillance activities.

59 BASIC BIOLOGICAL SCIENCES↗

Active meta-learning for predicting and selecting perovskite crystallization experiments

Autonomous experimentation systems use algorithms and data from prior experiments to select and perform new experiments in order to meet a specified objective. In most experimental chemistry situations, there is a limited set of prior historical data available, and acquiring new data may be expensive and time consuming, which places constraints on machine learning methods. Active learning methods prioritize new experiment selection by using machine learning model uncertainty and predicted outcomes. Meta-learning methods attempt to construct models that can learn quickly with a limited set of data for a new task. Here in this paper, we applied the model-agnostic meta-learning (MAML) model and the Probabilistic LATent model for Incorporating Priors and Uncertainty in few-Shot learning (PLATIPUS) approach, which extends MAML to active learning, to the problem of halide perovskite growth by inverse temperature crystallization. Using a dataset of 1870 reactions conducted using 19 different organoammonium lead iodide systems, we determined the optimal strategies for incorporating historical data into active and meta-learning models to predict reaction compositions that result in crystals. We then evaluated the best three algorithms (PLATIPUS and active-learning k-nearest neighbor and decision tree algorithms) with four new chemical systems in experimental laboratory tests. With a fixed budget of 20 experiments, PLATIPUS makes superior predictions of reaction outcomes compared to other active-learning algorithms and a random baseline.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

DOE FAIR Surrogate Benchmarks Supporting AI and Simulation Research (SBI Surrogate Benchmark Initiative) (Final Report)

Computational Science is being revolutionized by integrating AI and simulation and, in particular, by deep learning surrogate models that can replace all or part of traditional large‐scale HPC computations. Such surrogates can achieve remarkable performance improvements, as much as several orders of magnitude, and save both compute time and energy. The Surrogate Benchmark Initiative (SBI) project creates a community repository and FAIR (Findable, Accessible, Interoperable, and Reusable) data ecosystem for HPC application surrogate benchmarks. The SBI team comes from Argonne National Laboratory (ANL), Indiana University (IU), Rutgers University, the University of Tennessee, Knoxville (UTK), and the University of Virginia(UVA). SBI repositories include data, code, and all relevant collateral artifacts, that the science and engineering community needs to use and reuse these data sets and surrogates. SBI repositories generate active research from both participants in SBI and the broader AI and domain science communities. This project develops surrogates that use several different neural nets to learn and quickly infer the results of simulations and data systems and capture them as surrogate benchmarks with a rich set of metadata, covering. Data; Model; Metrics specification; Machine specification; Science, Speed, Power Results, We research FAIR metadata for these benchmarks. We develop application surrogate examples as benchmarks across many fields (ANL, UTK, IU, UVA). We also study non Surrogate benchmarks that have many common features and similar issues regarding FAIRness. We work with MLCommons (UVA, UTK), which is a major machine learning benchmarking activity where we get metadata ontologies, software, and benchmarks, benchmarks have datasets, models, and metadata, and they need a technical framework developed by UTK and Rutgers and deployed by UVA. We study features of Surrogates, including performance, training set size, and uncertainty quantification (Rutgers, UVA and IU).

97 MATHEMATICS AND COMPUTING↗

FAIR Surrogate Benchmarks Supporting AI and Simulation Research (Final Report)

Computational Science is being revolutionized by integrating AI and simulation and, in particular, by deep learning surrogate models that can replace all or part of traditional large‐scale HPC computations. Such surrogates can achieve remarkable performance improvements, as much as several orders of magnitude, and save both compute time and energy. The Surrogate Benchmark Initiative (SBI) project creates a community repository and FAIR (Findable, Accessible, Interoperable, and Reusable) data ecosystem for HPC application surrogate benchmarks. The SBI team comes from Argonne National Laboratory (ANL), Indiana University (IU), Rutgers University, the University of Tennessee, Knoxville (UTK), and the University of Virginia (UVA). SBI repositories include data, code, and all relevant collateral artifacts that the science and engineering community need to use and reuse these data sets and surrogates. SBI repositories generate active research from both the participants in SBI and the broad community of AI and domain scientists. This project develops surrogates that use several different neural nets to learn and quickly infer the results of simulations and data systems and captures them as surrogate benchmarks with a rich set of metadata covering: Data; Model; Metrics specification; Machine specification; and Science, Speed, and Power Results. We research FAIR metadata for these benchmarks. We develop application surrogate examples as benchmarks across many fields (ANL, UTK, IU, UVA). We also study non-Surrogate benchmarks that have many common features and similar issues as regards FAIRness. We work with MLCommons (UVA, UTK), which is a major machine learning benchmarking activity where we get metadata ontologies, software, and benchmarks, Benchmarks have datasets, models, and metadata and they need a technical framework developed by UTK and Rutgers and deployed by UVA. We study features of Surrogates including performance, training set size, and uncertainty quantification (Rutgers, UVA and IU).

97 MATHEMATICS AND COMPUTING↗

Anomaly Detection, Active Learning, Precursor Identification,and Human Knowledge for Autonomous System Safety

The project Autonomy Teaming and TRajectories for ComplexTrusted Operational Reliability (ATTRACTOR) researched and developed Artificial Intelligence with application to multi-Unmanned Aerial Systems (UAS) missions. Such missions, like other complex systems-of-systems, are likely to have previously-unknown, safety relevant anomalies occur due to many possible factors including system failures or degradations, emergent behavior, changes in the environment in which the systems operate, changes in the way the systems are operated. We discuss the application of anomaly detection, active learning, and precursor identification to identify such anomalies and the conditions under which they are more likely to appear. We demonstrate results on simulated multi-UAS missions that show promise to be applied to real missions.

machine learning↗

Building a machine learning surrogate model for wildfire activities within a global Earth system model

Abstract. Wildfire is an important ecosystem process, influencing land biogeophysical and biogeochemical dynamics and atmospheric composition. Fire-driven loss of vegetation cover, for example, directly modifies the surface energy budget as a consequence of changing albedo, surface roughness, and partitioning of sensible and latent heat fluxes. Carbon dioxide and methane emitted by fires contribute to a positive atmospheric forcing, whereas emissions of carbonaceous aerosols may contribute to surface cooling. Process-based modeling of wildfires in Earth system land models is challenging due to limited understanding of human, climate, and ecosystem controls on fire counts, fire size, and burned area. Integration of mechanistic wildfire models within Earth system models requires careful parameter calibration, which is computationally expensive and subject to equifinality. To explore alternative approaches, we present a deep neural network (DNN) scheme that surrogates the process-based wildfire model with the Energy Exascale Earth System Model (E3SM) interface. The DNN wildfire model accurately simulates observed burned area with over 90 % higher accuracy with a large reduction in parameterization time compared with the current process-based wildfire model. The surrogate wildfire model successfully captured the observed monthly regional burned area during validation period 2011 to 2015 (coefficient of determination, R2=0.93). Since the DNN wildfire model has the same input and output requirements as the E3SM process-based wildfire model, our results demonstrate the applicability of machine learning for high accuracy and efficient large-scale land model development and predictions.

58 GEOSCIENCES↗

Product Consistency Test Results for the LAW ML1 Glasses

This report summarizes the chemical analysis of Product Consistency Test (PCT) leachates received from Pacific Northwest National Laboratory (PNNL). The leachates are from a series of quenched simulated nuclear waste glasses designated Low-Activity Waste Machine Learning (LAW ML1) glasses that were designed and fabricated at PNNL. The reported data will be used in the development, validation, and implementation of enhanced property/composition models for waste glass vitrification at Hanford. The elemental release for the study glasses is reported as normalized concentration NCi. NCi of several elements was computed for both the target and measured glass compositions, which were similar, resulting in no significant differences. Several of the glasses exhibited NC B , NC Na , and/or NC Si values that were greater than the Waste Treatment Plant (WTP) low-activity waste constraint of 4 g/L. All reference glasses included with the study glasses had measurements that fell within the expected ranges.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Phase Stability Through Machine Learning

Understanding the phase stability of a chemical system constitutes the foundation of materials science. Knowledge of the equilibrium state of a system under arbitrary thermodynamic conditions provides valuable information about the types of phases that are likely to be synthesized and how to get there. Accessing the phase diagram in a materials system provides one with the information necessary to design materials and microstructures with optimal properties. While the materials science community has long been focused on exploiting this knowledge to navigate the materials space, recent advances in machine learning (ML) and artificial intelligence (AI) have provided the community with novel ways of interrogating the materials thermodynamics space. Furthermore, this work presents some of the most recent advances in ML/AI applied to phase stability and thermodynamics of materials. Prof. John Morral always had a passion for understanding and teaching the fundamental characteristics of phase diagrams. This review is written to honor his memory.

36 MATERIALS SCIENCE↗

Composition Measurements of the LAW ML1 Glasses

This report provides the results from the chemical analyses of the glass compositions of the quenched Low Activity Waste Machine Learning study glasses, a series of simulated nuclear waste glasses designed and fabricated at Pacific Northwest National Laboratory. These data will be used in the development, validation, and implementation of enhanced property/composition models for waste glass vitrification at Hanford. Chemical analyses were performed on a representative sample of each of the quenched glasses to allow for comparisons with target compositions. The relative differences between the target and measured concentrations of F - for one glass, P 2 O 5 for two glasses, SO 3 for three glasses, and ZrO 2 for two of the glasses were greater than 10%. These results can be used in further characterization of this series of glasses, including the normalization of Product Consistency Test results.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

GLUE Code: A framework handling communication and interfaces between scales

Many scientific applications are inherently multiscale in nature. Such complex physical phenomena often require simultaneous execution and coordination of simulations spanning multiple time and length scales. This is possible by combining expensive small-scale simulations (such as molecular dynamics simulations) with larger scale simulations (such continuum limit/hydro solvers) to allow for considerably larger systems using task and data parallelism. However, the granularity of the tasks can be very large and often leads to load imbalance. Traditionally, we use approximations to streamline the computation of the more costly interactions and this introduces trade-offs between simulation cost and accuracy. In recent years, the available computational power and the advances in machine learning have made computing these scale-bridging interactions and multiscale simulations more feasible. One driving application has been plasma modeling in inertial confinement fusion (ICF), which is fundamentally multiscale in nature. This requires deep understanding of how to extrapolate microscopic information into macroscopically relevant scales. For example, in ICF one needs an accurate understanding of the connection between experimental observables and the underlying microphysics. The properties of the larger scales are often affected by the microscale behavior incorporated usually into the equations of state and ionic and electronic transport coefficients (Liboff, 1959; Rinderknecht et al., 2014; Rosenberg et al., 2015; Ross et al., 2017). Instead of incorporating this information using reliable molecular dynamics (MD) simulations, one often needs to use theoretical models, due to the inability of MD to reach engineering scales (Glosli et al., 2007; Marinak et al., 1998). One approach to resolve this issue is by coupling two MD simulations of different scales via force interpolation, e.g., the AdResS method (Krekeler et al., 2018; Nagarajan et al., 2013). Another approach, which we will pursue in the scope of this work, is by enabling scale bridging between MD simulations and meso/macro-scale models through the development and support of application programming interfaces that these different applications can interact through.

54 ENVIRONMENTAL SCIENCES↗

Identifying recharge sources and their impacts on a North Central New Mexico shallow aquifer using unsupervised machine learning

In this article, shallow aquifers are important but highly variable resources in arid to semi-arid regions. Limited shallow aquifer volume results in high sensitivity to recharge fluctuations, which can impact the local fauna and flora, and transport of contaminants in the aquifer or vadose zone. Aquifer response to external forcing (e.g., precipitation) is usually solved by estimating aquifer parameters and running physics-based models to match known fluctuations of hydraulic head. However, this technique is time and computationally expensive. Furthermore, high aquifer complexity decreases precision in physics-based models. Alternatively supervised machine learning is used to predict aquifer dynamics. However, these techniques rely on input data and struggle to interpret aquifer response for missing sources (i.e., snowpack data). To counter these problems, we propose an unsupervised machine learning technique (NMFk) to estimate the impact of different sources on aquifer recharge. NMFk is used to understand the influence of external forcing on shallow aquifer recharge in the Pajarito Plateau (Los Alamos, NM, USA). The results show how NMFk can be used to reduce the data dimension in a complex field dataset to three recharge signals that cause fluctuations within the field data. Here, the source signals are interpreted as rainfall, snowmelt, and a delayed aquifer response to the previous two signals. These results evidence how heterogeneous aquifers delimited by canyons incised into the Pajarito Plateau respond in similar ways to the source signals identified by NMFk. Furthermore, results show the importance of the local geology where faults act as sinks, and anthropogenic disturbances can facilitate infiltration amplifying the interpreted signal.

54 ENVIRONMENTAL SCIENCES↗