Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine learning and data science”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Closing the Gap between FAIR Data Repositories and Hierarchical Data Formats

Many in the scientific community, particularly in publicly funded research, are pushing to adhere to more accessible data standards to maximize the findability, accessibility, interoperability, and reusability (FAIR) of scientific data, especially with the growing prevalence of machine learning augmented research. Online FAIR data repositories, such as the Open Science Framework (OSF), help facilitate the adoption of these standards by providing frameworks for storage, access, search, APIs, and other features that create organized hubs of scientific data. However, the wider acceptance of such repositories is hindered by the lack of support of hierarchical data formats, such as Technical Data Management Streaming (TDMS) and Hierarchical Data Format 5 (HDF5), that many researchers rely on to organize their datasets. Various tools and strategies should be used to allow hierarchical data formats, FAIR data repositories, and scientific organizations to work more seamlessly together. A pilot project at Los Alamos National Laboratory (LANL) addresses the disconnect between them by integrating the OSF FAIR data repository with hierarchical data renderers, extending support for additional file types in their framework. The multifaceted interactive renderer displays a tree of metadata alongside a table and plot of the data channels in the file. This allows users to quickly and efficiently load large and complex data files directly in the OSF webapp. Users who are browsing files can quickly and intuitively see the files in the way they or their colleagues structured the hierarchical form and immediately grasp their contents. This solution helps bridge the gap between hierarchical data storage techniques and FAIR data repositories, making both of them more viable options for scientific institutions like LANL which have been put off by the lack of integration between them.

97 MATHEMATICS AND COMPUTING↗

Evaluating 239 Pu(n,f) cross sections via machine learning using experimental data, covariances, and measurement features

In this paper, the neutron-induced 239 Pu fission cross section, 239 Pu(n,f), is evaluated from 1–20 MeV using experimental data and associated covariances while also considering information on the measurement, termed features here. For instance, methods to determine the background, sample backing material, or impurities in the sample, are explicitly taken into account in the evaluation process. To this end, outliers in the experimental data are identified with a modified version of the Hybrid Robust Support Vector Machine. In a second step, two machine learning methods (logistic regression with elastic net regularization and random forest regression with SHAP feature importance metric) are used to highlight measurement features that are common among many of the outlying data points. Based on this analysis, penalty uncertainties are added to the experimental covariances of outlying data points that have outlier measurement features and are put through the generalized-least-squares evaluation. The resulting evaluated mean values and covariances differ distinctly from those data evaluated without the penalty uncertainties. These results highlight that certain measurement features should be more closely examined.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Cross-property deep transfer learning framework for enhanced predictive analytics on small materials data

Abstract Artificial intelligence (AI) and machine learning (ML) have been increasingly used in materials science to build predictive models and accelerate discovery. For selected properties, availability of large databases has also facilitated application of deep learning (DL) and transfer learning (TL). However, unavailability of large datasets for a majority of properties prohibits widespread application of DL/TL. We present a cross-property deep-transfer-learning framework that leverages models trained on large datasets to build models on small datasets of different properties. We test the proposed framework on 39 computational and two experimental datasets and find that the TL models with only elemental fractions as input outperform ML/DL models trained from scratch even when they are allowed to use physical attributes as input, for 27/39 (≈ 69%) computational and both the experimental datasets. We believe that the proposed framework can be widely useful to tackle the small data challenge in applying AI/ML in materials science.

36 MATERIALS SCIENCE↗

Computational synthesis of a new generation of 2D-based perovskite quantum materials

Perovskite-based optoelectronic devices have emerged as a promising energy source due to their potential for scalable production. This study introduces “perovskene,” a novel class of 2D materials derived from the ABC3-like perovskites, synthesized via a data-driven, high-throughput computational strategy. We harness machine learning and multitarget deep neural networks to systematically investigate the structure–property relations, paving the way for targeted material design and optimization in fields such as renewable energy, electronics, and catalysis. The characterization of over 1500 synthesized structures shows that more than 500 structures are stable, revealing properties such as ultra-low work function and large magnetic moment, underscoring the potential for advanced technological applications.

2D materials↗

An Open Combinatorial Diffraction Dataset Including Consensus Human and Machine Learning Labels with Quantified Uncertainty for Training New Machine Learning Models

Modern machine learning and autonomous experimentation schemes in materials science rely on accurate analysis of the data ingested by these models. Unfortunately, accurate analysis of the underlying data can be difficult, even for domain experts, complicating the training of the models intended to drive experiments. This is especially true when the goal is to identify the presence of weak signatures in diffraction or spectroscopic datasets. In this work, we examine a set of as-obtained diffraction data that track the phase transition from monoclinic to tetragonal in a Nb-doped VO2 film as a function of temperature and dopant concentration. We then task a set of domain experts and a set of machine learning experts with identifying which phase is present in each diffraction pattern manually and algorithmically, respectively; in both cases, the labels can vary dramatically, especially at the phase boundaries. We use the mode of the labels and the Shannon entropy as a method to capture, preserve and propagate consensus labels and their variance. Further we use the expert labels as a benchmark and demonstrate the use of Shannon entropy weighted scoring to test the performance of machine learning generated labels. Finally, we propose a material data challenge centered around generating improved labeling algorithms. This real-world dataset curated with expert labels can act as test bed for new algorithms. The raw data, annotations and code used in this study are all available online at data.gov and the interested reader is encouraged to replicate and improve the existing models

97 MATHEMATICS AND COMPUTING↗

A general spatial-temporal framework for short-term building temperature forecasting at arbitrary locations with crowdsourcing weather data

Weather forecasting has been a critical component to predict and control building energy consumption for better building energy management. Without accessibility to other data sources, the onsite observed temperatures or the airport temperatures are used in forecast models. In this paper, we present a novel approach by utilizing the crowdsourcing weather data from neighboring personal weather stations (PWS) to improve the weather forecast accuracy around buildings using a general spatial-temporal modeling framework. The final forecast is based on the ensemble of local forecasts for the target location using neighboring PWSs. Our approach is distinguished from existing literature in various aspects. First, we leverage the crowdsourcing weather data from PWS in addition to public data sources. In this way, the data is at much finer time resolution (e.g., at 5-minute frequency) and spatial resolution (e.g., arbitrary location vs grid). Second, our proposed model incorporates spatial-temporal correlation information of weather variables between the target building and a set of neighboring PWSs so that underlying correlations can be effectively captured to improve forecasting performance. Here, we demonstrate the performance of the proposed framework by comparing to the benchmark models on temperature forecasting for a building located at an arbitrary location at San Antonio, Texas, USA. In general, the proposed model framework equipped with machine learning technique such as Random Forest can improve forecasting by 50% compares with persistent model and has 90% chance to outperform airport forecast in short-term forecasting. In a real-time setting, the proposed model framework can provide more accurate temperature forecasting results compared with using airport temperature forecast for most forecast horizon. Moreover, we analyze the sensitivity of model parameters to gain insights on how crowdsourcing data from the neighboring personal weather stations impacts forecasting performance. Finally, we implement our model in other cities such as Syracuse and Chicago to test the model's performance in different landforms and climate types.

54 ENVIRONMENTAL SCIENCES↗

Cross-Cutting Software Solutions in Support of Experimental Analysis Challenges at National Scattering Facilities

Rapid improvements in instrumentation at neutron and light source combined with new capabilities in computing, data analytics, and machine learning offer unprecedented scientific capabilities to address fundamental and applied science questions in a wide area of subjects from quantum materials to biological systems. In order to make these capabilities accessible to the broader scientific community, significant developments are needed related to workflow spanning from experimental control to advanced materials simulations as well as the underpinning scientific software. In this chapter, we explore these needs and recent developments at the Spallation Neutron Source (SNS) at Oak Ridge National Laboratory. A central part of the workflow at the SNS is the data reduction framework Mantid which is developed as a partnership between European and US neutron-scattering facilities.

Proffen, Thomas↗

Forecasting influenza activity using machine-learned mobility map

Human mobility is a primary driver of infectious disease spread. However, existing data is limited in availability, coverage, granularity, and timeliness. Data-driven forecasts of disease dynamics are crucial for decision-making by health officials and private citizens alike. In this work, we focus on a machine-learned anonymized mobility map (hereon referred to as AMM) aggregated over hundreds of millions of smartphones and evaluate its utility in forecasting epidemics. We factor AMM into a metapopulation model to retrospectively forecast influenza in the USA and Australia. We show that the AMM model performs on-par with those based on commuter surveys, which are sparsely available and expensive. We also compare it with gravity and radiation based models of mobility, and find that the radiation model’s performance is quite similar to AMM and commuter flows. Additionally, we demonstrate our model’s ability to predict disease spread even across state boundaries. Our work contributes towards developing timely infectious disease forecasting at a global scale using human mobility datasets expanding their applications in the area of infectious disease epidemiology.

60 APPLIED LIFE SCIENCES↗

Automation is all you need: Faster Earth system models with AI/ML

Focal Area: Data acquisition and assimilation enabled by machine learning (ML), artificial intelligence (AI) and advanced methods. Science Challenge: Tropical cyclones can in- duce extreme water cycle events through dramatic precipitation and storm surge. More reliable models of intensity will translate into better prediction of the impact of extreme events in large scale Earth systems simulations. We demonstrate and describe AI/ML methodologies for rapid assimilation of new, in situ data products.

54 ENVIRONMENTAL SCIENCES↗

Machine learning in materials research: Developments over the last decade and challenges for the future

The number of studies that apply machine learning (ML) to materials science has been growing at a rate of approximately 1.67 times per year over the past decade. In this review, I examine this growth in various contexts. First, I present an analysis of the most commonly used tools (software, databases, materials science methods, and ML methods) used within papers that apply ML to materials science. The analysis demonstrates that despite the growth of deep learning techniques, the use of classical machine learning is still dominant as a whole. It also demonstrates how new research can effectively build upon past research, particular in the domain of ML models trained on density functional theory calculation data. Next, I present the progression of best scores as a function of time on the matbench materials science benchmark for formation enthalpy prediction. In particular, a dramatic improvement of 7 times reduction in error is obtained when progressing from feature-based methods that use conventional ML (random forest, support vector regression, etc.) to the use of graph neural network techniques. Finally, I provide views on future challenges and opportunities, focusing on data size and complexity, extrapolation, interpretation, access, and relevance.

36 MATERIALS SCIENCE↗

Spectral Data Fusion From Handheld Laser-Induced Breakdown Spectroscopy (LIBS) and X-ray Fluorescence (XRF) Analyzers for Improved Detection of Cerium in a Simulated Dispersal Accident

Here, this work implements a mid-level data fusion methodology on spectral data from handheld X-ray fluorescence and laser-induced breakdown spectroscopy analyzers to quantify plutonium surrogate (CeO 2 ) contamination in soil samples for the first time. Spectral data from each analyzer were used independently to train supervised machine learning regressions to predict Ce concentration. Fused features from both data sets were then used to train the same models, comparing prediction performance by evaluating model precision and sensitivity. Fusing principal component scores from the two sensors yielded an order of magnitude improvement in precision and sensitivity of predictions made with an artificial neural network, compared to predictions made by models trained on independent sensor data. As a result, a boosted ensemble trained on the fused spectral features yielded an ideal predictor with root-mean-squared error on the order of 10 –6 and calculated limit of detection order 10 –5 wt %.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Elucidating Abnormal Grain Growth in Thermomagnetic Processed Materials with Transfer Learning and Reinforcement Learning

The goal of this research program is to establish the mechanism governing local grain boundary motion, which is needed to design and process desirable microstructures for better performance, by identifying the relative contributions of grain boundary (GB) energy and mobility to grain growth. Classical models for grain growth assume that the primary mechanism for reducing the total interfacial energy is area reduction and that GB restructuring is not significant. This assumption implies that grain growth is locally driven by curvature. However, recent experimental observations using new non-destructive 3D x-ray diffraction microscopy techniques (3D-XRM) reveal that classic descriptors (i.e., curvature, number of neighbors, grain size) do not predict real grain growth. Instead, local GB motion appears to be governed by its energy relative to its neighbors such that low-energy boundaries replace those of higher energy. However, simulations that incorporate GB energy anisotropy still fail to reproduce these observations. These discrepancies suggest that the common assumption for grain growth theory must be re-examined to predict and, thus, control microstructure evolution in real polycrystals. A significant challenge to testing this assumption is due to anisotropic GB mobility. Mobility may cause abnormal grain growth or affect the final grain shapes or growth rate but its true contributions are unknown because it is difficult to measure. For example, observations in Fe have found that grains associated with high energy and high mobility boundaries tend to experience abnormal grain growth, whereas abnormal grain growth is associated with low energy and high mobility boundaries in alumina. As mobility and energy both control GB motion, it is challenging to isolate the local driving forces necessary to test the common assumption that the primary mechanism is area reduction. The novelty of this work is the use of machine learning tools to capture GB mobility and energy from 3D-XRM measurements in polycrystals to test the common assumption used in grain growth models. Machine learning can capture high-order correlations in dynamic systems like those found in the evolving GB topology. The PIs have developed a physics-regularized interpretable machine learning microstructure evolution (PRIMME) model that accurately replicates the grain growth behavior of its trained data set.

36 MATERIALS SCIENCE↗

Accelerating the rate of discovery: toward high-repetition-rate HED science

As high-intensity short-pulse lasers that can operate at high-repetition-rate (HRR) (>10 H z ) come online around the world, the high energy density (HED) science they enable will experience a radical paradigm shift. The >10 3 increase in shot rate over today's shot-per-hour drivers translates into dramatically faster data acquisition, more experiments, and the ability to exploit machine learning, and thus the potential to significantly accelerate the advancement of HED science. A wide range of HED experiments, from opacity investigations to secondary source generation to plasma nuclear physics, will benefit from the increased statistics, precision, and exploration of phase space. Besides increasing the rate at which scientific experiments can be performed, HRR also allows for the rapid delivery of optimal experiments supported by simulations and modeling augmented by close coupling to empirical data. To fully realize such an HRR framework, numerous subsystems must be developed and brought together, including feedback laser control loops, high-throughput targetry and diagnostics, cognitive simulation, enhanced HED codes, and advanced data analytics. This paper describes the vision for an integrated HRR laser experimental HED system and outlines some of the major considerations and challenges for realizing it.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Accelerating scientific discoveries through data-driven innovations

Developing artificial intelligence (AI) and machine learning (ML) methods that can accelerate scientific discoveries and advance science has become one of the important research directions for the AI/ML research community. It has been gaining increasing attention from researchers in diverse scientific areas, including biomedical science, materials science, climate science, physics, chemistry, and many others. Data-driven AI/ML innovations to enable reliable predictions and optimal decision making for scientific discoveries face several critical challenges, among which are high system complexity, large search space, incomplete knowledge, and small data, all of which demand novel strategies to effectively address them. Meeting these challenges and thereby accelerating scientific discoveries and industrial innovations, calls for research that can take full advantage of the latest advances in AI/ML to integrate data-driven techniques with scientific knowledge and is able to execute them in modern high-performance computing (HPC) environments at scale. This Patterns special collection "Accelerating scientific discoveries through data-driven innovations" features articles that showcase the promising roles of AI/ML and data-driven modeling in accelerating scientific discoveries and may inspire the next wave of data-driven innovations in various scientific domains.

97 MATHEMATICS AND COMPUTING↗

Data-Driven Atomic Physics: Harnessing Machine Learning and High-Repetition-Rate Experiments for Laser-driven HED

High-energy-density plasma experiments are central to progress in atomic physics, fusion energy, and national security science, but they have traditionally been constrained by slow data collection and manual, time-intensive analysis. This project targeted that bottleneck by enabling high-repetition-rate experiments to produce and interpret much larger volumes of data quickly enough to guide experiments while they run.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗