Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine learning and data science”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Data-centric framework for crystal structure identification in atomistic simulations using machine learning

Atomic-level modeling performed at large scales enables the investigation of mesoscale materials properties with atom-by-atom resolution. The spatial complexity of such cross-scale simulations renders them unsuitable for simple human visual inspection. Instead, specialized structure characterization techniques are required to aid interpretation. These have historically been challenging to construct, requiring significant intuition and effort. Here we propose an alternative framework for a fundamental structural characterization task: classifying atoms according to the crystal structure to which they belong. Our approach is data-centric and favors the employment of Machine Learning over heuristic rules of classification. A group of data-science tools and simple local descriptors of atomic structure are employed together with an efficient synthetic training set. We also introduce the first standard and publicly available benchmark data set for evaluation of algorithms for crystal-structure classification. Further, it is demonstrated that our data-centric framework outperforms all of the most popular heuristic methods—especially at high temperatures when lattices are the most distorted—while introducing a systematic route for generalization to new crystal structures. Moreover, through the use of outlier detection algorithms our approach is capable of discerning between amorphous atomic motifs (i.e., noncrystalline phases) and unknown crystal structures, making it uniquely suited for exploratory materials synthesis simulations.

36 MATERIALS SCIENCE↗

Stiff neural ordinary differential equations

Neural Ordinary Differential Equations (ODEs) are a promising approach to learn dynamical models from time-series data in science and engineering applications. This work aims at learning neural ODEs for stiff systems, which are usually raised from chemical kinetic modeling in chemical and biological systems. We first show the challenges of learning neural ODEs in the classical stiff ODE systems of Robertson’s problem and propose techniques to mitigate the challenges associated with scale separations in stiff systems. We then present successful demonstrations in stiff systems of Robertson’s problem and an air pollution problem. The demonstrations show that the usage of deep networks with rectified activations, proper scaling of the network outputs as well as loss functions, and stabilized gradient calculations are the key techniques enabling the learning of stiff neural ODEs. The success of learning stiff neural ODEs opens up possibilities of using neural ODEs in applications with widely varying time-scales, such as chemical dynamics in energy conversion, environmental engineering, and life sciences.

97 MATHEMATICS AND COMPUTING↗

Uncertainty quantification in multivariable regression for material property prediction with Bayesian neural networks

With the increased use of data-driven approaches and machine learning-based methods in material science, the importance of reliable uncertainty quantification (UQ) of the predicted variables for informed decision-making cannot be overstated. UQ in material property prediction poses unique challenges, including multi-scale and multi-physics nature of materials, intricate interactions between numerous factors, limited availability of large curated datasets, etc. In this work, we introduce a physics-informed Bayesian Neural Networks (BNNs) approach for UQ, which integrates knowledge from governing laws in materials to guide the models toward physically consistent predictions. To evaluate the approach, we present case studies for predicting the creep rupture life of steel alloys. Experimental validation with three datasets of creep tests demonstrates that this method produces point predictions and uncertainty estimations that are competitive or exceed the performance of conventional UQ methods such as Gaussian Process Regression. Additionally, we evaluate the suitability of employing UQ in an active learning scenario and report competitive performance. The most promising framework for creep life prediction is BNNs based on Markov Chain Monte Carlo approximation of the posterior distribution of network parameters, as it provided more reliable results in comparison to BNNs based on variational inference approximation or related NNs with probabilistic outputs.

36 MATERIALS SCIENCE↗

Summary Report of the SOS26 Workshop held March 11-14, 2024

The SOS26 workshop was organized by Oak Ridge National Laboratory (ORNL) and held March 11-14, 2024, at Cocoa Beach, Florida. The SOS is a workshop organized annually, with a focus on distributed high performance computing (HPC). The technical program of the workshop was developed jointly by Sandia National Laboratories (SNL), ORNL and the Swiss National Supercomputing Center (CSCS). The 2024 SOS26 workshop theme was "Versatile HPC for the evolving and expanding needs of science" and had seven technical sessions covering HPC, data and machine learning (ML) topics. Each session consisted of four or five presentations, followed by a panel discussion. This report documents the workshop proceedings covering all the technical sessions.

97 MATHEMATICS AND COMPUTING↗

Special Issue: Geostatistics and Machine Learning

Abstract Recent years have seen a steady growth in the number of papers that apply machine learning methods to problems in the earth sciences. Although they have different origins, machine learning and geostatistics share concepts and methods. For example, the kriging formalism can be cast in the machine learning framework of Gaussian process regression. Machine learning, with its focus on algorithms and ability to seek, identify, and exploit hidden structures in big data sets, is providing new tools for exploration and prediction in the earth sciences. Geostatistics, on the other hand, offers interpretable models of spatial (and spatiotemporal) dependence. This special issue on Geostatistics and Machine Learning aims to investigate applications of machine learning methods as well as hybrid approaches combining machine learning and geostatistics which advance our understanding and predictive ability of spatial processes.

58 GEOSCIENCES↗

Comparison of machine learning techniques to optimize the analysis of plutonium surrogate material via a portable LIBS device

The utilization of machine learning techniques has become commonplace in the analysis of optical emission spectra. These methods are often limited to variants of principal components analysis (PCA), partial-least squares (PLS), and artificial neural networks (ANNs). A plethora of other techniques exist and are well established in the world of data science, yet are seldom investigated for their use in spectroscopic problems. In this study, machine learning techniques were used to analyze optical emission spectra of laser-induced plasma from ceria pellets doped with silicon in order to predict silicon content. Additionally, a boosted regression ensemble model was created, and its predictive accuracy was compared to that of traditional PCA, PLS, and ANN regression models. Boosted regression tree ensembles yielded fits with R-squared (R2) values as high as 0.964 and mean-squared errors of prediction (MSEPs) as low as 0.074, providing the most accurate predictive model. Neural networks performed with slightly lower R2 values and higher MSEPs compared to the ensemble methods, thus indicating susceptibility to overfitting.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Deep learning symmetries and their Lie groups, algebras, and subalgebras from first principles

Abstract We design a deep-learning algorithm for the discovery and identification of the continuous group of symmetries present in a labeled dataset. We use fully connected neural networks to model the symmetry transformations and the corresponding generators. The constructed loss functions ensure that the applied transformations are symmetries and the corresponding set of generators forms a closed (sub)algebra. Our procedure is validated with several examples illustrating different types of conserved quantities preserved by symmetry. In the process of deriving the full set of symmetries, we analyze the complete subgroup structure of the rotation groups SO (2), SO (3), and SO (4), and of the Lorentz group S O ( 1 , 3 ) . Other examples include squeeze mapping, piecewise discontinuous labels, and SO (10), demonstrating that our method is completely general, with many possible applications in physics and data science. Our study also opens the door for using a machine learning approach in the mathematical study of Lie groups and their properties.

97 MATHEMATICS AND COMPUTING↗

Metal hydride composition-derived parameters as machine learning features for material design and H 2 storage

Though hydrogen is a promising energy carrier for a green future, many challenges persist. One is the difficulty in engineering storage solutions, with metal hydrides being a leading contender among solid-state strategies. To facilitate efficient searching of candidate materials, ridge regression, simple decision trees, random forest ensembles, and gradient boosting ensembles were employed to predict the energy of formation, with the random forest ensemble resulting in the lowest test set error. First, two public databases, Materials Project and HydPark, were searched for metal hydrides. Feature engineering was performed before the models were developed, resulting in electronegativity, density, atomic density, d-character, f-character, band gap, hydrogen weight fraction, magnetization, temperature, and pressure being retained. The models were then benchmarked by the lowest test error before a random forest ensemble was used to populate entries missing energy of formation. Furthermore, all were then scored by hydrogen storage capacity and energy of formation suitability. Readily available features including several derived from only the chemical formula which were found to be highly predictive. and so are promising for high-throughput screening of arbitrary novel hydride formulations and blends for thermodynamic feasibility.

25 ENERGY STORAGE↗

Machine learned synthesizability predictions aided by density functional theory

Abstract A grand challenge of materials science is predicting synthesis pathways for novel compounds. Data-driven approaches have made significant progress in predicting a compound’s synthesizability; however, some recent attempts ignore phase stability information. Here, we combine thermodynamic stability calculated using density functional theory with composition-based features to train a machine learning model that predicts a material’s synthesizability. Our model predicts the synthesizability of ternary 1:1:1 compositions in the half-Heusler structure, achieving a cross-validated precision of 0.82 and recall of 0.82. Our model shows improvement in predicting non-half-Heuslers compared to a previous study’s model, and identifies 121 synthesizable candidates out of 4141 unreported ternary compositions. More notably, 39 stable compositions are predicted unsynthesizable while 62 unstable compositions are predicted synthesizable; these findings otherwise cannot be made using density functional theory stability alone. This study presents a new approach for accurately predicting synthesizability, and identifies new half-Heuslers for experimental synthesis.

Lee, Andrew (ORCID:0000000153014295)↗

Predicting Sintering Window of Binder Jet Additively Manufactured Parts Using a Coupled Data Analytics and CALPHAD Approach

Batch-to-batch variation in powder compositions for binder jet additive manufacturing (BJAM) can significantly deter defining an “ideal” sintering window for a given alloy. One way to overcome the problem is by running sintering experiments at various temperatures for each batch of the powder. However, such an approach increases the time required to achieve large-scale production of parts. The predictive capabilities of computational thermodynamic tools like CALPHAD can be leveraged to overcome the challenge, especially for binder jet additive manufacturing, since the process occurs under near-equilibrium conditions. However, calculating the sintering window using CALPHAD can be computationally expensive, considering many possible feedstock compositions within “specification”. Here, we generate high throughput CALPHAD data for nickel-based superalloys to develop machine learning models to predict the sintering window rapidly. The predictive capability of the models has been validated using published results on BJAM of Inconel 718 and 625. Further, validated models are lightweight and can be deployed in an industrial setting to get sintering window in an accelerated manner.

36 MATERIALS SCIENCE↗

A novel methodology for gamma-ray spectra dataset procurement over varying standoff distances and source activities

The adoption of machine learning approaches for gamma-ray spectroscopy has received considerable attention in the literature. Many studies have investigated the deployment of various algorithm architectures to a specific task. However, little attention has been afforded to the development of the datasets leveraged to train the models. Such training datasets typically span a set of environmental or detector parameters to encompass a problem space of interest to a user. Variations in these measurement parameters will also induce fluctuations in the detector response, including expected pile-up and ground scatter effects. Fundamental to this work is the understanding that 1) the underlying spectral shape varies as the measurement parameters change and 2) the statistical uncertainties associated with two spectra impact their level of similarity. While previous studies attribute some arbitrary discretization to the measurement parameters for the generation of their synthetic training data, this work introduces a principled methodology for efficient spectral-based discretization of a problem space. A signal-to-noise ratio (SNR) respective spectral comparison measure and a Gaussian Process Regression (GPR) model are used to predict the spectral similarity across a range of measurement parameters. This innovative approach effectively showcased its capability by dividing a problem space, ranging from 5 cm to 100 cm standoff distances and 5 μCi–100 μCi of 137 Cs, into three unique combinations of measurement parameters. The findings from this work will aid in creating more robust datasets, which incorporate many possible measurement scenarios, reduce the number of required experimental test set measurements, and possibly enable experimental training data collection for gamma-ray spectroscopy.

data science↗

Prediction of Distributed River Sediment Respiration Rates Using Community-Generated Data and Machine Learning

River sediment microbial respiration is a key indicator of ecosystem functioning and the biogeochemical fluxes across this critical zone link surface and subsurface waters. As such, there is tremendous interest in measuring and mapping these respiration rates. Respiration observations are expensive and labor intensive; there is limited data available to the community. An open science, collaborative initiative is collecting samples for respiration rate analysis and multi-scale metadata; this evolving data set is being used for making machine learning (ML) predictions at unsampled sites to help inform continued community engagement. However, it is a challenge to find an optimum configuration for ML models to work with this feature-rich (i.e., 100+ possible input variables) data set. Here, we present results from a two-tiered approach to managing the analysis of this complex data set: (a) a stacked ensemble of models that automatically optimizes hyperparameters and manages the training of many models and (b) feature permutation importance to detect the most important features in the models. The major elements of this workflow are modular, portable, open, and cloud-based thus making this implementation a potential template for other applications. The models developed here predict that sediment organic matter chemistry is one of the most important features for predicting sediment respiration rate. Other larger-scale, important features fall into the categories of climatic, ecological, geological, and fluvial settings. Leveraging these larger-scale features to generate data-driven estimates of river sediment respiration rates reveals spatially consistent but heterogeneous patterns across the river network of the Columbia River Basin.

54 ENVIRONMENTAL SCIENCES↗

Towards physics-informed explainable machine learning and causal models for materials research

From emergent material descriptions to estimation of properties stemming from structures to optimization of process parameters for achieving best performance – all key facets of materials science and related fields have experienced tremendous growth with the introduction of data-driven models. This gradual progression goes at par with developments of machine learning workflows, from purely data-driven shallow models to those that are well-capable in encoding more complex graphs, symbolic representations, invariances, and positional embeddings. Furthermore, this perspective aims at summarizing strategic aspects of such transitions while providing insights into the requirements of bringing in explainable, interpretable predictive models, and causal learning to aid in materials design and discovery. Although the focus remains on a variety of functional materials by providing a handful of case studies, the applications of such integrated methodologies are universal to facilitate fundamental understandings of materials physics while enabling autonomous experiments.

36 MATERIALS SCIENCE↗

Monotonic Gaussian Process for Physics-Constrained Machine Learning With Materials Science Applications

Physics-constrained machine learning is emerging as an important topic in the field of machine learning for physics. One of the most significant advantages of incorporating physics constraints into machine learning methods is that the resulting model requires significantly less data to train. By incorporating physical rules into the machine learning formulation itself, the predictions are expected to be physically plausible. Gaussian process (GP) is perhaps one of the most common methods in machine learning for small datasets. In this paper, we investigate the possibility of constraining a GP formulation with monotonicity on three different material datasets, where one experimental and two computational datasets are used. The monotonic GP is compared against the regular GP, where a significant reduction in the posterior variance is observed. The monotonic GP is strictly monotonic in the interpolation regime, but in the extrapolation regime, the monotonic effect starts fading away as one goes beyond the training dataset. Imposing monotonicity on the GP comes at a small accuracy cost, compared to the regular GP. The monotonic GP is perhaps most useful in applications where data are scarce and noisy, and monotonicity is supported by strong physical evidence.

36 MATERIALS SCIENCE↗

A Machine Learning Correction Model of the Winter Clear-Sky Temperature Bias over the Arctic Sea Ice in Atmospheric Reanalyses

Atmospheric reanalyses are widely used to estimate the past atmospheric near-surface state over sea ice. They provide boundary conditions for sea ice and ocean numerical simulations and relevant information for studying polar variability and anthropogenic climate change. Previous research revealed the existence of large near-surface temperature biases (mostly warm) over the Arctic sea ice in the current generation of atmospheric reanalyses, which is linked to a poor representation of the snow over the sea ice and the stably stratified boundary layer in the forecast models used to produce the reanalyses. These errors can compromise the employment of reanalysis products in support of polar research. Here, we train a fully connected neural network that learns from remote sensing infrared temperature observations to correct the existing generation of uncoupled atmospheric reanalyses (ERA5, JRA-55) based on a set of sea ice and atmospheric predictors, which are themselves reanalysis products. The advantages of the proposed correction scheme over previous calibration attempts are the consideration of the synoptic weather and cloud state, compatibility of the predictors with the mechanism responsible for the bias, and a self-emerging seasonality and multidecadal trend consistent with the declining sea ice state in the Arctic. The correction leads on average to a 27% temperature bias reduction for ERA5 and 7% for JRA-55 if compared to independent in situ observations from the MOSAiC campaign (respectively, 32% and 10% under clear-sky conditions). These improvements can be beneficial for forced sea ice and ocean simulations, which rely on reanalyses surface fields as boundary conditions.

54 ENVIRONMENTAL SCIENCES↗

Status of the Atlas of Neutron Resonances [Slides]

The Atlas of Neutron Resonances is the most comprehensive compilation of neutron resonances, thermal cross sections, resonance integrals and Maxwellian averaged cross sections generally available. For decades, the Atlas was carefully curated and maintained by Dr. Said Mughabghab who sadly passed on during the summer of 2018 after publishing the 2018 edition of the Atlas . We are continuing the development of this important compendium. To a large extent, the Atlas book is generated from a series of text files given in a single purpose domain-specific format. Therefore, we developed a software API and began the systematic assessment of the Atlas files. With this work past, we are now focusing on new efforts to expand the quality and scope of the Atlas . Current and recently completed projects include a cross comparison of the Atlas bibliography with Nuclear Science References and the EXFOR data library, a better determination of average resonance parameters, and using machine learning to assess the correctness of the spin group assignments of resonances tabulated in the Atlas .

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

7th World Congress on Integrated Computational Materials Engineering (ICME 2023) (Final Technical Report)

Integrated Computational Materials Engineering (ICME) has received international attention due to its potential to shorten product development time, while lowering cost and improving design and manufacturing outcomes. ICME is an approach to designing materials solutions for specific applications that use computer modeling programs to predict the behavior of materials and integrate this information into the overall materials, processing, and manufacturing design cycle. The 7th World Congress on Integrated Computational Materials Engineering (ICME 2023) was held in Orlando, Florida from May 21–25, 2023 with the goal to convene stakeholders from across all areas of modeling and simulation, experimental specialization, and design, as well as from across academia, government, and industry, to address ICME tools and techniques and their integration, as well as to examine their application in engineering. This atmosphere facilitated rich interactions between the experimentalists, modelers, and computational and design, from academia, government, and industry, to discuss ICME tools and techniques and their application in engineering.

36 MATERIALS SCIENCE↗