Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine learning for science”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21

Atomic cluster expansion potential for large scale simulations of hydrocarbons under shock compression

We present an Atomic Cluster Expansion (ACE) machine learned potential developed for high-fidelity atomistic simulations of hydrocarbons, targeting pressures and temperatures near and above supercritical fluid regimes for molecular fluids. A diverse set of stoichiometries were covered in training, including 1:0 (pure carbon), 1:4 (methane), and 1:1 (benzene), and rich bonding environments sampled at supercritical temperatures, hydrogen rich, reactive mixtures where metastable stoichiometries arise, including 1:2 (ethylene) and 1:3 (ethane). A high-fidelity training database was constructed by performing large-scale quantum molecular dynamic simulations [density functional theory (DFT) MD] of diamond, graphite, methane, and benzene. A novel approach to selecting structures from DFT MD is also presented, which allows for the rapid selection of unique DFT MD frames from complex trajectories. Comparisons to DFT and experimental data demonstrate that the presented ACE potential accurately reproduces isotherms, carbon melting curves, radial distribution functions, and shock Hugoniots for carbon and hydrocarbon systems for pressures up to 100 GPa and temperatures up to 6000 K for hydrocarbon systems and up to 9000 K for pure carbon systems. This work delivers a potential that can be used for accurate, large-scale simulations of shocked hydrocarbons and demonstrates a methodology for fitting and validating machine learning interatomic potentials to complex molecular environments, which can be applied to energetic materials in future works.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Machine learning to alleviate Hubbard-model sign problems

Lattice Monte Carlo calculations of interacting systems on nonbipartite lattices exhibit an oscillatory imaginary phase known as the phase or sign problem, even at zero chemical potential. One method to alleviate the sign problem is to analytically continue the integration region of the state variables into the complex plane via holomorphic flow equations. For asymptotically large flow times, the state variables approach manifolds of constant imaginary phase known as Lefschetz thimbles. Furthermore, flowing such variables and calculating the ensuing Jacobian is a computationally demanding procedure. In this paper, we demonstrate that neural networks can be trained to parametrize suitable manifolds for this class of sign problem and drastically reduce the computational cost for different severely afflicted small volume systems. In particular, we apply our method to the Hubbard model on the triangle and tetrahedron, both of which are nonbipartite. At strong interaction strengths and modest temperatures, the tetrahedron suffers from a severe sign problem that cannot be overcome with standard reweighting techniques, while it quickly yields to our method. We benchmark our results with exact calculations and comment on future directions of this work.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Using principal component analysis to distinguish different dynamic phases in superconducting vortex matter

Vortices in type-II superconductors driven over random disorder are known to exhibit a remarkable variety of distinct nonequilibrium dynamical phases that arise owing to the competition between vortex-vortex interactions, the quenched disorder, and the drive. These include pinned states, elastic flows, plastic or disordered flows, and dynamically reordered moving crystal or moving smectic states. The plastic flow phases can be particularly difficult to characterize since the flows are strongly disordered. Here, we perform principal component analysis (PCA) on the positions and velocities of vortex matter moving over random disorder for different disorder strengths and drives. We find that PCA can distinguish the known dynamic phases as well as or better than previous measures based on transport signatures or topological defect densities. In addition, PCA recognizes distinct plastic flow regimes, a slowly changing channel flow and a moving amorphous fluid flow, that do not produce distinct signatures in the standard measurements. In conclusion, our results suggest that this position and velocity-based PCA approach could be used to characterize dynamic phases in a broader class of systems that exhibit depinning and nonequilibrium phase transitions.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

SymProp: Scaling Sparse Symmetric Tucker Decomposition via Symmetry Propagation

Sparse symmetric tensors are an important class of tensors, and their decompositions serve as powerful tools for revealing low-rank structures. This paper introduces SymProp, a novel approach for scaling sparse symmetric Tucker decomposition by propagating symmetry through intermediate computations. SymProp optimizes two key computational kernels: Sparse Symmetric Tensor Times Same Matrix chain (S3 TTMc) for Higher-Order Orthogonal Iteration (HOOI) and Sparse Symmetric Tensor Times Same Matrix chain Times Core (S3 TTMcTC) for Higher-Order QR Iteration (HOQRI). Our method employs a metaprogramming-based index iteration approach to efficiently handle the upper triangular parts of intermediate dense symmetric tensors. SymProp achieves up to 50.9× speedup over SPLATT and up to 360.8× over Compressed Sparse Symmetric (CSS) format on the S3 TTMc operation. Moreover, our S3 TTMc and S3 TTMcTC implementations support tensor orders four levels higher than state-of-the-art methods. Our HOQRI demonstrates superior scalability and up to a 33.6× speedup over optimized HOOI. By enabling more scalable Tucker decompositions for higher orders, decomposition ranks, and dimension sizes, SymProp opens new possibilities for analyzing complex hypergraph structures in fields such as network science, data mining, and machine learning.

Li, Zecheng [North Carolina State University]↗

Generalized and Mechanistic PV Module Performance Prediction From Computer Vision and Machine Learning on Electroluminescence Images

Electroluminescence (EL) imaging of photovoltiac (PV) modules offers high-speed, high-resolution information about device performance, affording opportunities for greater insight and efficiency in module characterization across manufacturing, research and development, and power plant operations and management. Predicting module electrical properties from EL image features is a critical step toward these applications. In this article, we demonstrate quantification of both generalized and performance mechanism-specific EL image features, using pixel intensity-based and machine learning classification algorithms. From EL image features, we build predictive models for PV module power and series resistance, using time-series current-voltage (I-V) and EL data obtained stepwise on five brands of modules spanning three Si cell types through two accelerated exposures: damp heat (DH) (85°C/85% RH) and thermal cycling (TC) (IEC 61215). Overall, 195 pairs of EL images and I-V characteristics were analyzed, yielding 11700 individual PV cell images. A convolutional neural network was built to classify cells by the severity of busbar corrosion with high accuracy (95%). Generalized power predictive models estimated the maximum power of PV modules from EL images with high confidence and an adjusted-R 2 of 0.88, across all module brands and cell types in extended DH and TC exposures. Mechanistic degradation prediction was demonstrated by quantification of busbar corrosion in EL images of three module brands in DH, and subsequent modeling of series resistance using these mechanism-specific EL image features. For modules exhibiting busbar corrosion, we demonstrated series resistance predictive models with adjusted-R 2 of up to 0.73.

14 SOLAR ENERGY↗

HAM: Hotspot-Aware Manager for Improving Communications with 3D-Stacked Memory

merging High-Performance Computing (HPC) workloads, such as graph analytics, machine learning, and big data science, are data-intensive. Data-intensive workloads usually present fine-grained memory accesses with limited or no data locality, and thus incur frequent cache misses and low utilization of memory bandwidth. 3D-stacked memory devices such as Hybrid Memory Cube (HMC) and High Bandwidth Memory (HBM) can provide significantly higher bandwidth than conventional memory modules. However, the traditional interfaces and optimization methods for JEDEC DDR devices do not allow to fully exploit the potential performance of 3D-stacked memory with the massive amount of irregular memory accesses of data-intensive applications. In this paper, we propose a novel Hotspot-Aware Manager (HAM) infrastructure for 3D-stacked memory devices capable of optimizing memory access streams via request aggregation, hotspot detection, and in-memory prefetching. %and an associated hotspot-aware page policy. We present the HAM design and implementation, and simulate it on a system using RISC-V embedded cores with attached HMC devices. We extensively evaluate HAM with over 12 benchmarks and applications representing diverse irregular memory access patterns. The results show that, on average, HAM reduces redundant requests by 37.51\% and increases the prefetch buffer hit rate by 4.2 times, compared to a baseline streaming prefetcher. On the selected benchmark set, HAM provides performance gains of 21.81\% in average (up to 34.28\%) and power savings of 35.07\% over a standard 3D-stacked memory.

Wang, Xi↗

Generalized Canonical Polyadic Tensor Decomposition

Tensor decomposition is a fundamental unsupervised machine learning method in data science, with applications including network analysis and sensor data processing. This work develops a generalized canonical polyadic (GCP) low-rank tensor decomposition that allows other loss functions besides squared error. For instance, we can use logistic loss or Kullback--Leibler divergence, enabling tensor decomposition for binary or count data. We present a variety of statistically motivated loss functions for various scenarios. We provide a generalized framework for computing gradients and handling missing data that enables the use of standard optimization methods for fitting the model. Furthermore, we demonstrate the flexibility of the GCP decomposition on several real-world examples including interactions in a social network, neural activity in a mouse, and monthly rainfall measurements in India.

97 MATHEMATICS AND COMPUTING↗

Data used in Wainwright, H.M. et al. 2021, “Watershed zonation through hillslope clustering for tractably quantifying above- and belowground watershed heterogeneity and functions”

This data package contains spatial data layers and processing scripts used in Wainwright, H.M. et al. 2021, “Watershed zonation approach for tractably quantifying above-and- belowground watershed heterogeneity and functions”. The purpose of the data and paper is to develop a watershed zonation approach for characterizing watershed organization and function in a tractable manner by applying clustering methods to multiple spatial data layers. The data package contains the geotiff files of spatial data layers, and the processed data values corresponding to the figures in the paper. The Data_description file describe each file in details. The spatial data sets (geotiff) are included in the zip files.

54 ENVIRONMENTAL SCIENCES↗

A Mitigation Strategy for the Prediction Inconsistency of Neural Phase Pickers

Neural phase pickers—neural networks designed and trained to pick seismic phase arrivals—have proven to be a powerful tool for developing earthquake catalogs. However, these pickers suffer from prediction inconsistency in which the results they produce change, sometimes substantially, even under a small perturbation to the input waveform. This problem has not been addressed by the developers and users of these pickers. In this study, we show how prediction inconsistency can negatively affect the completeness of earthquake catalogs developed using neural phase pickers. Further, we show that simply using a small step size for the sliding window when processing continuous waveform data and aggregating the results significantly mitigates this problem. We also highlight the importance of training datasets for increasing the consistency and other performance metrics.

58 GEOSCIENCES↗

Facilitating Machine Learning Collaborations Between Labs, Universities, And Industry

It is clear from numerous recent community reports, papers, and proposals that machine learning is of tremendous interest for particle accelerator applications. The quickly evolving landscape continues to grow in both the breadth and depth of applications including physics modeling, anomaly detection, controls, diagnostics, and analysis. Consequently, laboratories, universities, and companies across the globe have established dedicated machine learning (ML) and data science efforts aiming to make use of these new state-of-the-art tools. The current funding environment in the U.S. is structured in a way that supports specific application spaces rather than larger collaboration on community software. Here, we discuss the existing collaboration bottlenecks and how a shift in the funding environment, and how we develop collaborative tools, can help fuel the next wave of ML advancements for particle accelerators.

Edelen, J.P.↗

Improving the representation of isopycnal mixing in E3SM (Final Report)

The DOE-supported Energy Exascale Earth System Model (E3SM) is a major attempt by the Department of Energy to develop a new Earth System Model using unstructured grids, which allow for the model to concentrate resolution where it is most needed and avoid numerical artifacts experienced by previous Earth System Models (ESMs) using regular grids in which the lateral resolution becomes extremely fine as the resolution shrinks at the poles. This effort required rewriting many of the algorithms previously developed for regular grids. One algorithm that was problematic in the first version of E3SM was the representation of mixing due to turbulent ocean eddies with scales smaller than the model grid. These mesoscale eddies are the primary way by which tracers are stirred along density surfaces in the ocean interior. In particular, the process of isopycnal mixing exchanges fresher, more oxygenated waters from polar regions with saltier, nutrient-rich and oxygen poor waters from tropical regions. Previous work in our group has shown that this mechanism is vitally important for bringing oxygen into poorly ventilated tropical regions (Gnanadesikan et al., 2012, 2013) and plays a significant role in determining the uptake of anthropogenic carbon dioxide (Gnanadesikan et al., 2015). However, in E3SMv1 this process was turned off, as turning it on caused the model to become unstable. Additionally, the rate of isopycnal mixing is directly proportional to a mixing coefficient $A_{Redi}$ whose value varies from less than 400 m 2 /s to 2000 m 2 /s across contemporary climate models. The main thrusts of this proposal therefore were to: 1.) Improve the numerics of isopycnal mixing in E3SM; 2.) Identify ways in which climate models are sensitive to the isopycnal mixing coefficient; and 3.) Develop better representations of the isopycnal mixing coefficient.

54 ENVIRONMENTAL SCIENCES↗

Fission with Exotic Nuclei (Abbreviated Report)

Nuclear fission is a key mechanism involved in the synthesis of heavy elements in the Cosmos and is the primary explanation for the stability of superheavy elements. Nevertheless, our knowledge of fission remains extremely fragmented. Most experiments have been conducted only on a tiny number of stable actinide nuclei and are often incomplete, leading to gaps in our basic understanding of the process. For many radioactive isotopes, basic fission data such as the charge or mass distribution of the fragments is unknown. These gaps cannot always be filled by simulation alone. Common fission models contain too many free parameters and lack predictive power. In contrast, the fundamental theory of fission under development at LLNL is much more predictive, but its current computational cost is too high to be used extensively for data evaluations. A unique window of opportunity to resolve these limitations has recently opened: the U.S. nuclear science community is ramping up major experimental programs at the Facility for Rare Isotope Beams (FRIB, the DOE flagship facility in low-energy nuclear science), and techniques from machine learning have shown great potential to simplify the use of a fundamental, quantum-mechanical theory of fission. This project has two components. On the experimental side, we acquired and deployed at the HIGS facility a new dual Frisch-Grid ionization chamber to measure correlated fragment-mass, kinetic energy, and angular distributions of fission fragments from induced fission. This new device was used to perform measurements of charge, mass and total kinetic energy of fission fragments in the photofission of 238 U and eight gamma-ray beam energies between 6.2 and 13 MeV, which allowed extracting high-precision independent yields for this reaction. The device was also used to perform measurements of the same quantities in the neutron-induced fission of 234 U with monoenergetic beams of energy between 5 and 8 MeV. In parallel, we collaborated with a team at Commissariat à l’énergie atomique et aux énergies alternatives (CEA) to perform a series of measurements of fission yields in inverse kinematics for the two isotopes of 236 U and 240 Pu. The experiment took place at the Grand Accélérateur National d’Ions Lourds in France in June 2023. The deployment of the VAMOS spectrometer with a new array called PISTA allowed determining the excitation energy of the fissioning system within 1 Mega-electronvolts. The second component of the project involved using deep neural networks to build fast and reliable emulators of our current fission models. In an invited paper published in Frontier in Physics, we showed that autoencoders could successfully compress nuclear wavefunctions in nuclear density functional theory. We achieved a dimensionality reduction of the order of two orders of magnitude while keeping the error in the total energy to less than 0.01%. In a second paper submitted to Physical Review Letters in June 2023 with our collaborators at CEA, we showed that variational autoencoders can learn the collective degrees of freedom driving the fission process.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

OLCF’s Advanced Computing Ecosystem (ACE): FY25 Update for Ongoing Efforts

The advent of widespread use of artificial intelligence (AI) and machine learning (ML) models in science, coupled with fast data production rates of scientific instruments strain the traditional batch-oriented high-performance computing (HPC) environment. As scientific exploration continues to require more data and faster processing and analysis, new emerging technologies and capabilities to enable cross-facility and time-sensitive workflows are required for seamless integration of HPC and experimental facilities. The Advanced Computing Ecosystem (ACE) is a strategic initiative within the Oak Ridge Leadership Computing Facility (OLCF) established in 2024 to support the development of cutting-edge technologies to advance computational research and infrastructure at OLCF and across the Department of Energy (DOE). Several DOE initiatives are spearheading the evolution of the scientific landscape by blurring facility boundaries and connecting the user facilities to advance scientific capabilities and ensure energy dominance. The DOE Integrated Research Infrastructure (IRI) program is one example that is laying a foundation to support complex cross-facility workflows. The IRI program aims to integrate diverse computational resources, data infrastructures, and scientific instruments to facilitate collaboration and accelerate scientific discovery. The Interconnected Science Ecosystem (INTERSECT) initiative at Oak Ridge National Laboratory (ORNL) is another example that aims to revolutionize scientific research through AI-driven, interconnected autonomous laboratories and research facilities. Finally, the American Science Cloud (AmSC), recently announced in the “One Big Beautiful Bill”, aims to leverage prior infrastructure efforts of the IRI and automation and AI efforts of INTERSECT (and others) to build a federated, AI-augmented AmSC platform to unify the DOE’s computing, experimental, and data resources to catalyze scientific innovation.

97 MATHEMATICS AND COMPUTING↗

Scientific Discovery with Physics-Informed System Identification (Abbreviated Report)

My fellowship research focused on making physics-based simulations faster and more useful through machine learning. Many problems in science and engineering are governed by partial differential equations, but high-fidelity simulations are often too expensive to run repeatedly. I worked on improving Latent Space Dynamics Identification (LaSDI), a reduced-order modeling framework that compresses large simulation data sets into a smaller representation and then learns how that representation evolves over time. The motivation was to develop reduced models that remain accurate for more challenging systems, especially when predictions must remain reliable over long time intervals or when the underlying dynamics are more complicated than standard methods can easily handle. I also contributed to related work on Quandary, a high-performance software effort for simulation and control of open quantum systems, before focusing primarily on Latent Space Dynamics Identification methods. The main outcomes of the fellowship were two new algorithms (both of which were published), Rollout-LaSDI and Higher-Order LaSDI, together with supporting work on multi-stage Latent Space Dynamics Identification. Rollout-LaSDI improved long-term prediction by training the model to stay accurate over extended time horizons, and Higher-Order LaSDI broadened the method so it could model systems with higher-order time dynamics. My contributions to multistage Latent Space Dynamics Identification also helped show that its later training stages could be simplified without losing effectiveness, and that this behavior held across different model architectures and training strategies. Taken together, these advances improved the accuracy, flexibility, and practical value of reduced-order modeling tools for computational science.

97 MATHEMATICS AND COMPUTING↗

Knowledge Discovery Process: Case Study of RNAV Adherence of Radar Track Data

This talk is an introduction to the knowledge discovery process, beginning with: identifying the problem, choosing data sources, matching the appropriate machine learning tools, and reviewing the results. The overview will be given in the context of an ongoing study that is assessing RNAV adherence of commercial aircraft in the national airspace.

Machine Learning↗

MLtool: Universal Supervised Machine Learning Tool to Model Tabulated Data

Machine Learning (ML) is a subfield of Artificial Intelligence that gives computers the ability to learn from past data without being explicitly programmed. The predictive capabilities of ML models have already been used to facilitate several scientific breakthroughs. However, the practical application of ML is often limited due to the gaps in technical knowledge of its users. The common issue faced by many scientific researchers is the inability to choose the appropriate ML pipelines that are needed to treat real-world data, which is often sparse and noisy. To solve this problem, we have developed an automated Machine Learning tool (MLtool) that includes a set of ML algorithms and approaches to aid scientific researchers. The current version of MLtool is implemented as an object-oriented Python code that is easily extensible. It includes 44 different regression algorithms used to model data. MLtool helps users select the best model for their data, based on the scoring metrics used. Besides regression algorithms, MLtool also includes a suite of pre- and post-processing techniques such as missing value imputation, categorical variable encoding, input feature normalization, uncertainty quantification, exploratory data analysis (EDA), etc. MLtool was tested on several publicly available multi-dimensional data sets and was found capable of making accurate predictions.

Machine learning↗

Find My Astronaut Photo

Explore the source record for details and available documents.

Astronaut Photography↗