Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine learning and data science”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Data-centric machine learning in quantum information science

Abstract We propose a series of data-centric heuristics for improving the performance of machine learning systems when applied to problems in quantum information science. In particular, we consider how systematic engineering of training sets can significantly enhance the accuracy of pre-trained neural networks used for quantum state reconstruction without altering the underlying architecture. We find that it is not always optimal to engineer training sets to exactly match the expected distribution of a target scenario, and instead, performance can be further improved by biasing the training set to be slightly more mixed than the target. This is due to the heterogeneity in the number of free variables required to describe states of different purity, and as a result, overall accuracy of the network improves when training sets of a fixed size focus on states with the least constrained free variables. For further clarity, we also include a ‘toy model’ demonstration of how spurious correlations can inadvertently enter synthetic data sets used for training, how the performance of systems trained with these correlations can degrade dramatically, and how the inclusion of even relatively few counterexamples can effectively remedy such problems.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Data-driven materials research enabled by natural language processing and information extraction

Given the emergence of data science and machine learning throughout all aspects of society, but particularly in the scientific domain, there is increased importance placed on obtaining data. Data in materials science are particularly heterogeneous, based on the significant range in materials classes that are explored and the variety of materials properties that are of interest. This leads to data that range many orders of magnitude, and these data may manifest as numerical text or image-based information, which requires quantitative interpretation. The ability to automatically consume and codify the scientific literature across domains - enabled by techniques adapted from the field of natural language processing - therefore has immense potential to unlock and generate the rich datasets necessary for data science and machine learning. This review focuses on the progress and practices of natural language processing and text mining of materials science literature and highlights opportunities for extracting additional information beyond text contained in figures and tables in articles. Here, we discuss and provide examples for several reasons for the pursuit of natural language processing for materials, including data compilation, hypothesis development, and understanding the trends within and across fields. Current and emerging natural language processing methods along with their applications to materials science are detailed. We, then, discuss natural language processing and data challenges within the materials science domain where future directions may prove valuable.

36 MATERIALS SCIENCE↗

CAML: Commutative Algebra Machine Learning─A Case Study on Protein–Ligand Binding Affinity Prediction

Recently, Suwayyid and Wei introduced commutative algebra as an emerging paradigm for machine learning and data science. In this work, we propose commutative algebra machine learning (CAML) for the prediction of protein−ligand binding affinities. Specifically, we apply persistent Stanley−Reisner theory, a key concept in combinatorial commutative algebra, to the affinity predictions of protein−ligand binding and metalloprotein−ligand binding. We present three new algorithms, i.e., element-specific commutative algebra, category-specific commutative algebra, and commutative algebra on bipartite complexes, to tackle the complexity of data involved in (metallo) protein−ligand complexes. We show that the proposed CAML outperforms other state-of-theart methods in (metallo) protein−ligand binding affinity predictions, indicating the great potential of commutative algebra learning.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Resource frugal optimizer for quantum machine learning

Quantum-enhanced data science, also known as quantum machine learning (QML), is of growing interest as an application of near-term quantum computers. Variational QML algorithms have the potential to solve practical problems on real hardware, particularly when involving quantum data. However, training these algorithms can be challenging and calls for tailored optimization procedures. Specifically, QML applications can require a large shot-count overhead due to the large datasets involved. In this work, we advocate for simultaneous random sampling over both the dataset as well as the measurement operators that define the loss function. We consider a highly general loss function that encompasses many QML applications, and we show how to construct an unbiased estimator of its gradient. This allows us to propose a shot-frugal gradient descent optimizer called Refoqus (REsource Frugal Optimizer for QUantum Stochastic gradient descent). Our numerics indicate that Refoqus can save several orders of magnitude in shot cost, even relative to optimizers that sample over measurement operators alone.

97 MATHEMATICS AND COMPUTING↗

Ultrafast radiographic imaging and tracking: An overview of instruments, methods, data, and applications

Ultrafast radiographic imaging and tracking (U-RadIT) use state-of-the-art ionizing particle and light sources to experimentally study sub-nanosecond transients or dynamic processes in physics, chemistry, biology, geology, materials science and other fields. These processes are fundamental to modern technologies and applications, such as nuclear fusion energy, advanced manufacturing, communication, and green transportation, which often involve one mole or more atoms and elementary particles, and thus are challenging to compute by using the first principles of quantum physics or other forward models. One of the central problems in U-RadIT is to optimize information yield through, e.g. high-luminosity X-ray and particle sources, efficient imaging and tracking detectors, novel methods to collect data, and large-bandwidth online and offline data processing, regulated by the underlying physics, statistics, and computing power. We review and highlight recent progress in: (a.) Detectors such as high-speed complementary metal-oxide semiconductor (CMOS) cameras, hybrid pixelated array detectors integrated with Timepix4 and other application-specific integrated circuits (ASICs), and digital photon detectors; (b.) U-RadIT modalities such as dynamic phase contrast imaging, dynamic diffractive imaging, and four-dimensional (4D) particle tracking; (c.) U-RadIT data and algorithms such as neural networks and machine learning, and (d.) Applications in ultrafast dynamic material science using XFELs, synchrotrons and laser-driven sources. Hardware-centric approaches to U-RadIT optimization are constrained by detector material properties, low signal-to-noise ratio, high cost and long development cycles of critical hardware components such as ASICs. Interpretation of experimental data, including comparisons with forward models, is frequently hindered by sparse measurements, model and measurement uncertainties, and noise. Alternatively, U-RadIT make increasing use of data science and machine learning algorithms, including experimental implementations of compressed sensing. Machine learning and artificial intelligence approaches, refined by physics and materials information, may also contribute significantly to data interpretation, uncertainty quantification and U-RadIT optimization.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Noise-Resilient Quantum Machine Learning for Stability Assessment of Power Systems

Transient stability assessment (TSA) is a cornerstone for resilient operations of todays interconnected power grids. This paper is a confluence of quantum computing, data science and machine learning to potentially address the power system TSA issue. Here, we devise a quantum TSA (QTSA) method to enable scalable and efficient data-driven transient stability prediction for bulk power systems, which is the first attempt to tackle the TSA issue with quantum computing. Our contributions are three-fold: 1) A high expressibility, low-depth (HELD) quantum circuit is designed for accurate and noise-resilient TSA; 2) A quantum natural gradient descent algorithm is developed for efficient HELD circuit training; 3) A systematical analysis on QTSAs performance under various quantum factors is per-formed. QTSA underpins a foundation of quantum-enabled and data-driven power grid stability analytics. It renders the intractable TSA straightforward and effortless in the Hilbert space, and therefore provides stability information for power system operations. Extensive experiments on quantum simulators and real quantum computers verify the accuracy, noise-resilience, scalability and universality of QTSA.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Curifactory: A research experiment manager

Curifactory is a command line tool and framework for organizing Python experiment code, configuration parameters, and results. It is an opinionated and lightweight approach to workflow management infrastructure and is primarily intended to support researchers conducting experiments on one machine. This software was developed to support the reproducibility of results for several data science projects in the Nuclear Nonproliferation Division at Oak Ridge National Laboratory. Curifactory is intended to be a general framework and is not specific to machine learning or data science. It can aid in any field in which experiments are primarily computation-based studies and can be implemented in Python (e.g., high-energy physics, astronomy, computational chemistry). Here, the design emphasizes the automated caching of intermediate data analysis artifacts to speed up development involving computationally intensive tasks. It also allows for data provenance and experiment reproduction. Individual experiment runs are tracked through logs and their output reports, and entire copies of a run with all cached data and metadata can be exported for others to run using Curifactory on another machine. Curifactory experiments can either be integrated into a project from the beginning or can be written on top of an existing codebase without needing significant modification. A few important views of the Curifactory library can be seen in Figure 1.

97 MATHEMATICS AND COMPUTING↗

MIDAS: Modeling Individual Differences using Advanced Statistics

This research explores novel methods for extracting relevant information from EEG data to characterize individual differences in cognitive processing. Our approach combines expertise in machine learning, statistics, and cognitive science, advancing the state-of-the art in all three domains. Specifically, by using cognitive science expertise to interpret results and inform algorithm development, we have developed a generalizable and interpretable machine learning method that can accurately predict individual differences in cognition. The output of the machine learning method revealed surprising features of the EEG data that, when interpreted by the cognitive science experts, provided novel insights to the underlying cognitive task. Additionally, the outputs of the statistical methods show promise as a principled approach to quickly find regions within the EEG data where individual differences lie, thereby supporting cognitive science analysis and informing machine learning models. This work lays methodological ground work for applying the large body of cognitive science literature on individual differences to high consequence mission applications.

97 MATHEMATICS AND COMPUTING↗

Data as a Key Resource in Catalysis: A Community Account

The deployment of artificial intelligence (AI) is transforming the scientific fields central to interdisciplinary catalysis research. By enabling more effective use of data, AI (including simpler machine learning and data science tools) holds great promise for accelerating discoveries. However, progress has so far been modest, largely due to the lack of standardized, machine-readable, and openly shared catalysis data. This perspective, accounting for community insights emerging at conferences, analyses the underlying reasons for these challenges and proposes solutions to a future whereFAIR data management becomes an integral part of research in catalysis. In the short-term, we deem that mandatory FAIR data depositing prior to scientific publications along with consensualized top-down guidelines on data sharing powered by ease-to-use tools can make the necessary step change happen to catalyse data as key resource in our community.

36 - MATERIALS SCIENCE↗

GridDS: Data Science Toolkit for Energy Grid Data

According to the U.S. Energy Information Administration (EIA), the demand for energy is expected to increase 50% by the year 20501. While energy standards, such as the Institute of Electrical and Electronics Engineers (IEEE) Standard 1547, (Basso 2015) and monitoring with wide area management systems (WAMS) (Liu 2017, Zhou 2016) have enabled large scale data collection and storage, the application of this data in mitigating costs associated with increased consumer demand is an ongoing focus for energy research. This ubiquitous data collection presents a promising opportunity for machine learning and data science to improve efficiency of distributed energy resources (DERs). The GridDS software toolkit is designed to leverage advanced metering infrastructure (AMI), outage management systems data (OMS), Supervisory control Data Acquisition (SCADA), and geographic information systems (GIS) to forecast future energy demands and detect incipient grid failures. GridDS is a python software library designed to be modular and generalizable to data recorded by DERs. In adapting to disparate datasets recorded by various WAMS, GridDS provides a range of unique functionality not presently implemented in current WAMS which have highly specific software infrastructure by design. GridDS functionality ranges from data specification and preparation, to training and validation for state of the art machine learning, to interactive data visualization. For data intake, GridDS combines: Pandera: a library for creating data specifications. TimeScaleDB: a postgresSQL database infrastructure for efficient storage of timeseries data. Dataset class: A custom dataset class / interface that ensures modularity between a range of synthetic and live recorded datasets. Is

Ladd, Alexander↗

Realizing the data-driven, computational discovery of metal-organic framework catalysts

Metal-organic frameworks (MOFs) have been widely investigated for challenging catalytic transformations due to their well-defined structures and high degree of synthetic tunability. These features, at least in principle, make MOFs ideally suited for a computational approach towards catalyst design and discovery. Nonetheless, the widespread use of data science and machine learning to accelerate the discovery of MOF catalysts has yet to be substantially realized. In this review, we provide an overview of recent work that sets the stage for future high-throughput computational screening and machine learning studies involving MOF catalysts. This is followed by a discussion of several challenges currently facing the broad adoption of data-centric approaches in MOF computational catalysis, and we share possible solutions that can help propel the field forward.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Expanded analysis of machine learning models for nuclear transient identification using TPOT

Industries around the world are becoming more and more data driven. The nuclear field is no exception with several different applications being proposed. One popular area of research is the use of machine learning in transient detection. This paper seeks to build upon a previous study which made use of the AutoML package TPOT to train traditional machine learning models to classify transient events occurring with a reactor. Synthetic data was once again collected using a GPWR reactor simulator. Data on 12 different events was collected using 15 different initial conditions. Here, a dataset consisting of over 100,000 data points was compiled and used to train 7 different machine learning models using a pre-defined TPOT dictionary with 12 different preprocessing techniques. Three of the trained models were able to produce validation results in the 90s with the expanded dataset. Once the models were trained, it was possible to look into where during the simulation, misclassifications occurred. Using these three models, analysis was done to determine if TPOT could be used to train models that were effective if important features were missing. The results from this were positive with the newly trained models scoring close to the original models. Finally, to conclude this study, the three high performing models were retrained using different random states to see if there was any major variation when different states were used.

42 ENGINEERING↗

Towards High-Throughput Computation of Phase-and Defect Diagrams

The past decade has seen immense advances in our understanding of defect thermodynamics, and the use of machine learning and data science approaches has played a critical role in these advances [1–14]. In the area of grain boundaries (GBs), a particular focus has been placed on the effects of alloying – namely, GB solute segregation or more broadly, GB alloying [15–25], which has been observed and catalogued across a vast range of systems [26–50]. The impacts of solute segregation to GBs are numerous, and can range from negative effects such as embrittlement – for example, due to impurities [51–53], during irradiation [54–61], or during heat treatment [62–65] – to positive effects such as the stabilization against grain growth [66–69], thus enabling the design of nanocrystalline alloys with access to an enhanced range of functional and mechanical properties, and the reduction of embrittlement through the segregation of GB strengthening solutes [49,70–79].

36 MATERIALS SCIENCE↗

Subseasonal Representation and Predictability of North American Weather Regimes Using Cluster Analysis

Abstract This study focuses on assessing the representation and predictability of North American weather regimes, which are persistent large-scale atmospheric patterns, in a set of initialized subseasonal reforecasts created using the Community Earth System Model, version 2 (CESM2). The k -means clustering was used to extract four key North American (10°–70°N, 150°–40°W) weather regimes within ERA5 reanalysis, which were used to interpret CESM2 subseasonal forecast performance. Results show that CESM2 can recreate the climatology of the four main North American weather regimes with skill but exhibits biases during later lead times with overoccurrence of the West Coast high regime and underoccurrence of the Greenland high and Alaskan ridge regimes. Overall, the West Coast high and Pacific trough regimes exhibited higher predictability within CESM2, partly related to El Niño. Despite biases, several reforecasts were skillful and exhibited high predictability during later lead times, which could be partly attributed to skillful representation of the atmosphere from the tropics to extratropics upstream of North America. The high predictability at the subseasonal time scale of these case-study examples was manifested as an “ensemble realignment,” in which most ensemble members agreed on a prediction despite ensemble trajectory dispersion during earlier lead times. Weather regimes were also shown to project distinct temperature and precipitation anomalies across North America that largely agree with observational products. This study further demonstrates that unsupervised learning methods can be used to uncover sources and limits of subseasonal predictability, along with systematic biases present in numerical prediction systems. Significance Statement North American weather regimes are large-scale atmospheric patterns that can persist for several days. Their skillful subseasonal (2 weeks or greater) prediction can provide valuable lead time to prepare for temperature and precipitation anomalies that can stress energy and water resources. The purpose of this study was to assess the climatological representation and subseasonal predictability of four key North American weather regimes using a research subseasonal prediction system and clustering analysis. We found that the Pacific trough and West Coast high regimes exhibited higher predictability than other regimes and that skillful representation of conditions across the tropics and extratropics can increase predictability during later lead times. Future work will quantify causal pathways associated with high predictability.

58 GEOSCIENCES↗

MS25: Materials Science-Focused Benchmark Data Set for Machine Learning Interatomic Potentials

Here, we present MS25, a benchmark data set for evaluating machine learning interatomic potentials (MLIPs) across diverse materials-relevant systems including MgO surfaces, liquid water, zeolites, a catalytic Pt surface reaction, high-entropy alloys (HEAs), and disordered Zr-oxides. Five MLIP architectures (MACE, NequIP, Allegro, MTP, and Torch-ANI) are trained and tested, focusing not only on traditional metrics (energies, forces, and stresses) but also explicitly validating derived physical observables such as lattice constants, volumes, and reaction barriers. We find that most models reach comparable accuracy on standard error metrics across the simple systems, although equivariant MLIPs offer 1.5–2× improvements over nonequivariant MLIPs in energy and force error for structurally complex or compositionally disordered environments such as HEAs and Zr–O systems. Our analysis highlights that low errors in energy and force predictions do not guarantee reliable observables, emphasizing the necessity of explicit validation. We demonstrate limitations in cross-framework transferability, as models trained on one zeolite framework (CHA) fail to reliably generalize to predictions of structurally distinct frameworks (e.g., MFI). Size-extensive tests show some dependence on system size for MgO, resulting from forced periodicity. The HEA and Zr–O data sets are identified as challenging tests for future benchmarks and MLIP model architecture developments as they show significant differentiation in error between MLIP architectures and are still relatively difficult at 1000 training images. Moving forward, we recommend that benchmarking efforts shift their focus from marginal accuracy improvements in energy and force errors toward identifying and understanding model failure modes, rigorously assessing transferability, and evaluating how their errors affect observable predictions. For researchers looking to choose an MLIP architecture, we suggest selecting equivariant MLIP architectures if the complexity of the system is a challenge. For simple materials problems, auxiliary features such as integration with molecular dynamics engines, trade-offs between computational data set generation cost vs MLIP inference speed, and framework integration may play a more important decision factor than small differences in error metrics that are unlikely to matter for production-level research.

chemical structure↗

Computational optimal transport for molecular spectra: The fully continuous case

Computational optimal transport is used to analyze the difference between pairs of continuous molecular spectra. It is demonstrated that transport distances which are derived from this approach may be a more appropriate measure of the difference between two continuous spectra than more familiar measures of distance under many common circumstances. Associated with the transport distances is the transport map which provides a detailed analysis of the difference between two molecular spectra and is a key component of our study of quantitative differences between two continuous spectra. The use of optimal transport for comparing molecular spectra is developed in detail here with a set of model spectra, so that the discussion is self-contained. The difference between the transport distance and more common definitions of distance is elucidated for some well-chosen examples and it is shown where transport distances may be very useful alternatives to standard definitions of distance. The transport distance between a theoretical and experimental electronic absorption spectrum for SO 2 is studied and it is shown how the theoretical spectrum can be modified to fit the experimental spectrum better adjusting the theoretical band origin and the resolution of the theoretical spectrum. In conclusion, this analysis includes the calculation of transport maps between the theoretical and experimental spectra suggesting future applications of the methodology.

74 ATOMIC AND MOLECULAR PHYSICS↗

Database schema design: Energy Flexibility Environmental Tradeoffs Tool

The Energy Flexibility-Environment Tradeoff Toolset is designed using Streamlit framework, with Python as the programming language. Data storage is facilitated through the use of SQLite. Streamlit is an open-source Python framework for machine learning and data science teams, and SQLite is the most used database engine.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗