Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine learning and data science”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Closing the Gap between FAIR Data Repositories and Hierarchical Data Formats

Many in the scientific community, particularly in publicly funded research, are pushing to adhere to more accessible data standards to maximize the findability, accessibility, interoperability, and reusability (FAIR) of scientific data, especially with the growing prevalence of machine learning augmented research. Online FAIR data repositories, such as the Open Science Framework (OSF), help facilitate the adoption of these standards by providing frameworks for storage, access, search, APIs, and other features that create organized hubs of scientific data. However, the wider acceptance of such repositories is hindered by the lack of support of hierarchical data formats, such as Technical Data Management Streaming (TDMS) and Hierarchical Data Format 5 (HDF5), that many researchers rely on to organize their datasets. Various tools and strategies should be used to allow hierarchical data formats, FAIR data repositories, and scientific organizations to work more seamlessly together. A pilot project at Los Alamos National Laboratory (LANL) addresses the disconnect between them by integrating the OSF FAIR data repository with hierarchical data renderers, extending support for additional file types in their framework. The multifaceted interactive renderer displays a tree of metadata alongside a table and plot of the data channels in the file. This allows users to quickly and efficiently load large and complex data files directly in the OSF webapp. Users who are browsing files can quickly and intuitively see the files in the way they or their colleagues structured the hierarchical form and immediately grasp their contents. This solution helps bridge the gap between hierarchical data storage techniques and FAIR data repositories, making both of them more viable options for scientific institutions like LANL which have been put off by the lack of integration between them.

97 MATHEMATICS AND COMPUTING↗

Evaluating 239 Pu(n,f) cross sections via machine learning using experimental data, covariances, and measurement features

In this paper, the neutron-induced 239 Pu fission cross section, 239 Pu(n,f), is evaluated from 1–20 MeV using experimental data and associated covariances while also considering information on the measurement, termed features here. For instance, methods to determine the background, sample backing material, or impurities in the sample, are explicitly taken into account in the evaluation process. To this end, outliers in the experimental data are identified with a modified version of the Hybrid Robust Support Vector Machine. In a second step, two machine learning methods (logistic regression with elastic net regularization and random forest regression with SHAP feature importance metric) are used to highlight measurement features that are common among many of the outlying data points. Based on this analysis, penalty uncertainties are added to the experimental covariances of outlying data points that have outlier measurement features and are put through the generalized-least-squares evaluation. The resulting evaluated mean values and covariances differ distinctly from those data evaluated without the penalty uncertainties. These results highlight that certain measurement features should be more closely examined.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Cross-property deep transfer learning framework for enhanced predictive analytics on small materials data

Abstract Artificial intelligence (AI) and machine learning (ML) have been increasingly used in materials science to build predictive models and accelerate discovery. For selected properties, availability of large databases has also facilitated application of deep learning (DL) and transfer learning (TL). However, unavailability of large datasets for a majority of properties prohibits widespread application of DL/TL. We present a cross-property deep-transfer-learning framework that leverages models trained on large datasets to build models on small datasets of different properties. We test the proposed framework on 39 computational and two experimental datasets and find that the TL models with only elemental fractions as input outperform ML/DL models trained from scratch even when they are allowed to use physical attributes as input, for 27/39 (≈ 69%) computational and both the experimental datasets. We believe that the proposed framework can be widely useful to tackle the small data challenge in applying AI/ML in materials science.

36 MATERIALS SCIENCE↗

Computational synthesis of a new generation of 2D-based perovskite quantum materials

Perovskite-based optoelectronic devices have emerged as a promising energy source due to their potential for scalable production. This study introduces “perovskene,” a novel class of 2D materials derived from the ABC3-like perovskites, synthesized via a data-driven, high-throughput computational strategy. We harness machine learning and multitarget deep neural networks to systematically investigate the structure–property relations, paving the way for targeted material design and optimization in fields such as renewable energy, electronics, and catalysis. The characterization of over 1500 synthesized structures shows that more than 500 structures are stable, revealing properties such as ultra-low work function and large magnetic moment, underscoring the potential for advanced technological applications.

2D materials↗

An Open Combinatorial Diffraction Dataset Including Consensus Human and Machine Learning Labels with Quantified Uncertainty for Training New Machine Learning Models

Modern machine learning and autonomous experimentation schemes in materials science rely on accurate analysis of the data ingested by these models. Unfortunately, accurate analysis of the underlying data can be difficult, even for domain experts, complicating the training of the models intended to drive experiments. This is especially true when the goal is to identify the presence of weak signatures in diffraction or spectroscopic datasets. In this work, we examine a set of as-obtained diffraction data that track the phase transition from monoclinic to tetragonal in a Nb-doped VO2 film as a function of temperature and dopant concentration. We then task a set of domain experts and a set of machine learning experts with identifying which phase is present in each diffraction pattern manually and algorithmically, respectively; in both cases, the labels can vary dramatically, especially at the phase boundaries. We use the mode of the labels and the Shannon entropy as a method to capture, preserve and propagate consensus labels and their variance. Further we use the expert labels as a benchmark and demonstrate the use of Shannon entropy weighted scoring to test the performance of machine learning generated labels. Finally, we propose a material data challenge centered around generating improved labeling algorithms. This real-world dataset curated with expert labels can act as test bed for new algorithms. The raw data, annotations and code used in this study are all available online at data.gov and the interested reader is encouraged to replicate and improve the existing models

97 MATHEMATICS AND COMPUTING↗

Phenomena Portal: Large- Scale Visual Exploration of Atmospheric Phenomena

The Earth science community is experiencing a high influx of remote sensing data due to recent advancements in sensor technology. This enables the community to extend their research on a larger scale than ever before. Unfortunately, traditional data processing techniques do not scale well to these new, high volume data sources. State-of-the-art machine learning (ML) pipelines have been proven to overcome these burdens in various other fields but are underexploited within the physical sciences community. Moreover, ML is reliant on labeled data, which is currently sparsely available, owing to the fact that ML adoption is still in the early stages within the Earth and atmospheric science communities. To address these issues, we developed the Phenomena Portal, a visual exploration tool that uses ML to detect various atmospheric phenomena on a global scale. This allows the Earth and atmospheric science communities to view trends of occurrences of phenomena, identify potential relationships between them, and analyze spatiotemporal patterns over time. These detections can also serve as initial labeled data for ML research pertaining to the respective phenomena. The tool also incorporates feedback from subject matter experts to further improve the model detection accuracy, thereby facilitating human-in-the-loop. This presentation will provide an overview of the ML model development and cloud deployment. We also discuss the capabilities of the user interface for displaying the detections.

Muthukumaran Ramasubramanian↗

A general spatial-temporal framework for short-term building temperature forecasting at arbitrary locations with crowdsourcing weather data

Weather forecasting has been a critical component to predict and control building energy consumption for better building energy management. Without accessibility to other data sources, the onsite observed temperatures or the airport temperatures are used in forecast models. In this paper, we present a novel approach by utilizing the crowdsourcing weather data from neighboring personal weather stations (PWS) to improve the weather forecast accuracy around buildings using a general spatial-temporal modeling framework. The final forecast is based on the ensemble of local forecasts for the target location using neighboring PWSs. Our approach is distinguished from existing literature in various aspects. First, we leverage the crowdsourcing weather data from PWS in addition to public data sources. In this way, the data is at much finer time resolution (e.g., at 5-minute frequency) and spatial resolution (e.g., arbitrary location vs grid). Second, our proposed model incorporates spatial-temporal correlation information of weather variables between the target building and a set of neighboring PWSs so that underlying correlations can be effectively captured to improve forecasting performance. Here, we demonstrate the performance of the proposed framework by comparing to the benchmark models on temperature forecasting for a building located at an arbitrary location at San Antonio, Texas, USA. In general, the proposed model framework equipped with machine learning technique such as Random Forest can improve forecasting by 50% compares with persistent model and has 90% chance to outperform airport forecast in short-term forecasting. In a real-time setting, the proposed model framework can provide more accurate temperature forecasting results compared with using airport temperature forecast for most forecast horizon. Moreover, we analyze the sensitivity of model parameters to gain insights on how crowdsourcing data from the neighboring personal weather stations impacts forecasting performance. Finally, we implement our model in other cities such as Syracuse and Chicago to test the model's performance in different landforms and climate types.

54 ENVIRONMENTAL SCIENCES↗

Cross-Cutting Software Solutions in Support of Experimental Analysis Challenges at National Scattering Facilities

Rapid improvements in instrumentation at neutron and light source combined with new capabilities in computing, data analytics, and machine learning offer unprecedented scientific capabilities to address fundamental and applied science questions in a wide area of subjects from quantum materials to biological systems. In order to make these capabilities accessible to the broader scientific community, significant developments are needed related to workflow spanning from experimental control to advanced materials simulations as well as the underpinning scientific software. In this chapter, we explore these needs and recent developments at the Spallation Neutron Source (SNS) at Oak Ridge National Laboratory. A central part of the workflow at the SNS is the data reduction framework Mantid which is developed as a partnership between European and US neutron-scattering facilities.

Proffen, Thomas↗

Forecasting influenza activity using machine-learned mobility map

Human mobility is a primary driver of infectious disease spread. However, existing data is limited in availability, coverage, granularity, and timeliness. Data-driven forecasts of disease dynamics are crucial for decision-making by health officials and private citizens alike. In this work, we focus on a machine-learned anonymized mobility map (hereon referred to as AMM) aggregated over hundreds of millions of smartphones and evaluate its utility in forecasting epidemics. We factor AMM into a metapopulation model to retrospectively forecast influenza in the USA and Australia. We show that the AMM model performs on-par with those based on commuter surveys, which are sparsely available and expensive. We also compare it with gravity and radiation based models of mobility, and find that the radiation model’s performance is quite similar to AMM and commuter flows. Additionally, we demonstrate our model’s ability to predict disease spread even across state boundaries. Our work contributes towards developing timely infectious disease forecasting at a global scale using human mobility datasets expanding their applications in the area of infectious disease epidemiology.

60 APPLIED LIFE SCIENCES↗

Automation is all you need: Faster Earth system models with AI/ML

Focal Area: Data acquisition and assimilation enabled by machine learning (ML), artificial intelligence (AI) and advanced methods. Science Challenge: Tropical cyclones can in- duce extreme water cycle events through dramatic precipitation and storm surge. More reliable models of intensity will translate into better prediction of the impact of extreme events in large scale Earth systems simulations. We demonstrate and describe AI/ML methodologies for rapid assimilation of new, in situ data products.

54 ENVIRONMENTAL SCIENCES↗

Science Autonomy for Ocean Worlds Astrobiology: A Perspective

Astrobiology missions to ocean worlds in our solar system must overcome both scientific and technological challenges due to extreme temperature and radiation conditions, long communication times, and limited bandwidth. While such tools could not replace ground-based analysis by science and engineering teams, machine learning algorithms could enhance the science return of these missions through development of autonomous science capabilities. Examples of science autonomy include onboard data analysis and subsequent instrument optimization, data prioritization (for transmission), and real-time decision-making based on data analysis. Similar advances could be made to develop streamlined data processing software for rapid ground-based analyses. Here we discuss several ways machine learning and autonomy could be used for astrobiology missions, including landing site selection, prioritization and targeting of samples, classification of “features” (e.g., proposed biosignatures) and novelties (uncharacterized, “new” features, which may be of most interest to agnostic astrobiological investigations), and data transmission.

ocean worlds↗

Machine learning in materials research: Developments over the last decade and challenges for the future

The number of studies that apply machine learning (ML) to materials science has been growing at a rate of approximately 1.67 times per year over the past decade. In this review, I examine this growth in various contexts. First, I present an analysis of the most commonly used tools (software, databases, materials science methods, and ML methods) used within papers that apply ML to materials science. The analysis demonstrates that despite the growth of deep learning techniques, the use of classical machine learning is still dominant as a whole. It also demonstrates how new research can effectively build upon past research, particular in the domain of ML models trained on density functional theory calculation data. Next, I present the progression of best scores as a function of time on the matbench materials science benchmark for formation enthalpy prediction. In particular, a dramatic improvement of 7 times reduction in error is obtained when progressing from feature-based methods that use conventional ML (random forest, support vector regression, etc.) to the use of graph neural network techniques. Finally, I provide views on future challenges and opportunities, focusing on data size and complexity, extrapolation, interpretation, access, and relevance.

36 MATERIALS SCIENCE↗

Developing a Machine-Learning-Based Processing Framework for Twitter and Other Crowdsourced Data

Crowdsourced data streams such as Twitter and other social media are important sources of real-time and historical global information for Earth science applications. At the NASA Goddard Earth Sciences Data and Information Services Center (GES DISC), we have been exploring the Twitter data stream for its potential in augmenting the validation program of NASA's Global Precipitation Measurement (GPM) mission. To realize this potential, we need to increase the information density and enhance the quality of filtered precipitation tweets. We have implemented various components of a machine learning (ML)-based processing infrastructure for crowdsourced data that outputs, in this instance, useful and usable information derived from precipitation tweets. We have test enriched the Twitter stream with higher quality active tweets from those knowingly contributing to our effort and from existing crowdsourced programs (e.g., mPING, CoCoRaHS). We have experimented with various algorithms for processing tweets, including Naà ve Bayes, Convolutional Neural Network (CNN), Hierarchical Attention Network (HAN), and semi-supervised learning (with tri-training). Our current work focuses on (1) automated review of Earth science-related publications to determine relationships between discipline research needs and ML algorithms; (2) investigating Sequential Generative Adversarial Network (SeqGAN) for processing precipitation tweets for anomaly detection; and (3) managing crowdsourced data in a way that is compatible with existing NASA satellite data archives and using the data for ML applications. Key results include (1) network visualization of NLP-processed publications in various Earth science disciplines; (2) difference between GPM-linked, generated tweets and collected actual tweets that is small for GPM-determined light to moderate rain cases and high for GPM-determined heavy rain cases; and (3) identification of MongoDB for storing raw tweets and Zarr format for gridded tweets (compatible with GPM data). Our results have taken us a step closer to an operational ML-based tweet processing infrastructure and have already demonstrated that tweet-derived precipitation information is potentially useful for validation of Earth science satellite data.

Teng, William↗

Spectral Data Fusion From Handheld Laser-Induced Breakdown Spectroscopy (LIBS) and X-ray Fluorescence (XRF) Analyzers for Improved Detection of Cerium in a Simulated Dispersal Accident

Here, this work implements a mid-level data fusion methodology on spectral data from handheld X-ray fluorescence and laser-induced breakdown spectroscopy analyzers to quantify plutonium surrogate (CeO 2 ) contamination in soil samples for the first time. Spectral data from each analyzer were used independently to train supervised machine learning regressions to predict Ce concentration. Fused features from both data sets were then used to train the same models, comparing prediction performance by evaluating model precision and sensitivity. Fusing principal component scores from the two sensors yielded an order of magnitude improvement in precision and sensitivity of predictions made with an artificial neural network, compared to predictions made by models trained on independent sensor data. As a result, a boosted ensemble trained on the fused spectral features yielded an ideal predictor with root-mean-squared error on the order of 10 –6 and calculated limit of detection order 10 –5 wt %.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Elucidating Abnormal Grain Growth in Thermomagnetic Processed Materials with Transfer Learning and Reinforcement Learning

The goal of this research program is to establish the mechanism governing local grain boundary motion, which is needed to design and process desirable microstructures for better performance, by identifying the relative contributions of grain boundary (GB) energy and mobility to grain growth. Classical models for grain growth assume that the primary mechanism for reducing the total interfacial energy is area reduction and that GB restructuring is not significant. This assumption implies that grain growth is locally driven by curvature. However, recent experimental observations using new non-destructive 3D x-ray diffraction microscopy techniques (3D-XRM) reveal that classic descriptors (i.e., curvature, number of neighbors, grain size) do not predict real grain growth. Instead, local GB motion appears to be governed by its energy relative to its neighbors such that low-energy boundaries replace those of higher energy. However, simulations that incorporate GB energy anisotropy still fail to reproduce these observations. These discrepancies suggest that the common assumption for grain growth theory must be re-examined to predict and, thus, control microstructure evolution in real polycrystals. A significant challenge to testing this assumption is due to anisotropic GB mobility. Mobility may cause abnormal grain growth or affect the final grain shapes or growth rate but its true contributions are unknown because it is difficult to measure. For example, observations in Fe have found that grains associated with high energy and high mobility boundaries tend to experience abnormal grain growth, whereas abnormal grain growth is associated with low energy and high mobility boundaries in alumina. As mobility and energy both control GB motion, it is challenging to isolate the local driving forces necessary to test the common assumption that the primary mechanism is area reduction. The novelty of this work is the use of machine learning tools to capture GB mobility and energy from 3D-XRM measurements in polycrystals to test the common assumption used in grain growth models. Machine learning can capture high-order correlations in dynamic systems like those found in the evolving GB topology. The PIs have developed a physics-regularized interpretable machine learning microstructure evolution (PRIMME) model that accurately replicates the grain growth behavior of its trained data set.

36 MATERIALS SCIENCE↗