Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “distributed machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Modeling the distribution of the endangered Jemez Mountains salamander (Plethodon neomexicanus) in relation to geology, topography, and climate

The Jemez Mountains salamander (Plethodon neomexicanus; hereafter JMS) is an endangered salamander restricted to the Jemez Mountains in north-central New Mexico, United States. This strictly terrestrial and lungless species requires moist surface conditions for activities such as mating and foraging. Threats to its current habitat include fire suppression and ensuing severe fires, changes in forest composition, habitat fragmentation, and climate change. Forest composition changes resulting from reduced fire frequency and increased tree density suggest that its current aboveground habitat does not mirror its historically successful habitat regime. However, because of its limited habitat area and underground behavior, we hypothesized that geology and topography might play a significant role in the current distribution of the salamander. We modeled the distribution of the JMS using a machine learning algorithm to assess how geology, topography, and climate variables influence its distribution. The best habitat suitability model indicates that geology type and maximum winter temperature (November to March) were most important in predicting the distribution of the salamander (23.5% and 50.3% permutation importance, respectively). Minimum winter temperature was also an important variable (21.4%), suggesting this also plays a role in salamander habitat. Our habitat suitability map reveals low uncertainty in model predictions, and we found slight discrepancies between the designated critical habitat and the most suitable areas for the JMS. Because geological features are important to its distribution, we recommend that geological and topographical data are considered, both during survey design and in the description of localities of JMS records once detected.

59 BASIC BIOLOGICAL SCIENCES↗

Power quality disturbances diagnosis: A 2D densely connected convolutional network framework

The fast and accurate diagnosis of power quality disturbances (PQD) aids in avoiding shutdowns and unnecessary procedures, concerning electric energy distribution systems. As such, a number of techniques have been tested and applied in order to reach this objective. Majority of the techniques applied are two-step based. On the first step, power quality disturbances features are extracted. Second step, considering features extracted, disturbance classification is implemented. Recently, relevant literature has presented data-driven signal processing-based approaches, as deep convolutional neural networks (DCNN), which can implement both processing steps while providing automated recognition of patterns and outliers in data. However, not considered by state-of-art, power quality disturbances are evolving in nature, while all possible regularities might not be represented in the dataset. In this work a 2 Dimension Densely Connected Convolutional Network (2D-DenseNet) framework is presented. Further, a case study with synthetic disturbance events are analyzed. Easy-to-implement formulation, built on the 2D-DenseNet, without hard-to-design parameters, highlight potential aspects for real-life implementation.

42 ENGINEERING↗

Robust Medium-Voltage Distribution System State Estimation using Multi-Source Data

Due to the lack of sufficient online measurements for distribution system observability, pseudo-measurements from short-term load or distributed renewable energy resources (DERs) forecasting are used. However, the accuracy of them is low and thus significantly limits the performance of distribution system state estimation (DSSE). In this paper, a robust DSSE that integrates multi-source measurement data is proposed. Specifically, the historical low-voltage (LV) side smart meters are used to forecast load and DERs injections via the support vector machine (SVM) with optimally tuned parameters. By contrast, the online smart meters at LV side are utilized to derive equivalent power injections at the MV/LV transformers, yielding more accurate pseudo-measurements compared to the forecasted injections. Furthermore, to deal with bad data caused by communication loss, instrumental errors and cyber attacks, robust DSSE that relies on generalized maximum-likelihood (GM)-estimation criterion is developed. The projection statistics are developed to adjust the weights of each measurement, leading to better balance between pseudo- and real-time measurements. Numerical results conducted on modified IEEE 33-bus system with DG integration demonstrate the effectiveness and robustness of the proposed method.

distribution system state estimation↗

iDDS: intelligent distributed dispatch and scheduling for workflow orchestration

The intelligent distributed dispatch and scheduling (iDDS) service is a versatile workflow orchestration system designed for large-scale, distributed scientific computing. iDDS extends traditional workload and data management by integrating data-aware execution, conditional logic, and programmable workflows, enabling automation of complex and dynamic processing pipelines. Originally developed for the ATLAS experiment at the large hadron collider, iDDS has evolved into an experiment-agnostic platform that supports both template-driven workflows and a Function-as-a-Task model for Python-based orchestration. This paper presents the architecture and core components of iDDS, highlighting its scalability, modular message-driven design, and integration with systems such as PanDA and Rucio. We demonstrate its versatility through real-world use cases: fine-grained tape resource optimization for ATLAS, orchestration of large Directed Acyclic Graph (DAG) workflows for the Rubin Observatory, distributed hyperparameter optimization for machine learning applications, active learning for physics analyses, and AI-assisted detector design at the electron–ion collider. By unifying workload scheduling, data movement, and adaptive decision-making, iDDS reduces operational overhead and enables reproducible, high-throughput workflows across heterogeneous infrastructures. We conclude with current challenges and future directions, including interactive, cloud-native, and serverless workflow support.

97 MATHEMATICS AND COMPUTING↗

High-throughput virtual laboratory for drug discovery using massive datasets

Time-to-solution for structure-based screening of massive chemical databases for COVID-19 drug discovery has been decreased by an order of magnitude, and a virtual laboratory has been deployed at scale on up to 27,612 GPUs on the Summit supercomputer, allowing an average molecular docking of 19,028 compounds per second. Over one billion compounds were docked to two SARS-CoV-2 protein structures with full optimization of ligand position and 20 poses per docking, each in under 24 hours. GPU acceleration and high-throughput optimizations of the docking program produced 350× mean speedup over the CPU version (50× speedup per node). GPU acceleration of both feature calculation for machine-learning based scoring and distributed database queries reduced processing of the 2.4 TB output by orders of magnitude. The resulting 50× speedup for the full pipeline reduces an initial 43 day runtime to 21 hours per protein for providing high-scoring compounds to experimental collaborators for validation assays.

97 MATHEMATICS AND COMPUTING↗

Integrated Approach to Ancillary PV Component Reliability Assessment (Final Report)

In this project, we have established a nondestructive, generalized methodology that (1) fuses rich field data with advanced ML for proactive reliability forecasting, (2) dramatically reduces experimental iterations via synthetic dataset generation, and (3) achieves unprecedented regression precision in both anomaly detection and component-level degradation assessment—paving the way for truly predictive maintenance of grid-tied PV inverters under diverse outdoor conditions.

14 SOLAR ENERGY↗

A 1 km soil moisture dataset over eastern CONUS generated by assimilating SMAP data into the Noah-MP land surface model

An improved fine-scale soil moisture (SM) dataset at 1 km grid spacing, covering much of the eastern continental US, was generated by assimilating 9 km Soil Moisture Active Passive (SMAP) SM data into the v4.0.1 Noah-MP land surface model. With 12 ensemble members, the assimilation was carried out using the ensemble Kalman filter algorithm within NASA's Land Information System. The SM analysis for 2016 was fully validated against in situ observations from four different networks and compared with four other existing datasets. Results indicate that this SM analysis surpasses other datasets in top-layer SM distribution, including a machine-learning-based product, despite all SM estimates being less heterogeneous than observed. The analysis of anomalous errors suggests that large similarity in intrinsic errors is likely due to overlapping data sources among the selected SM datasets. More detailed evaluations were performed over two geographic areas. The observations collected by the Atmospheric Radiation Measurement facility in Oklahoma suggest that soil temperature and surface heat fluxes are concurrently simulated with good accuracy. Investigation into the 2016 southeastern US drought response further indicates drier conditions and higher evapotranspiration estimates compared to GLEAMv4.1. Notably, large errors are associated with grids having clay soil textures, underscoring the need for refined model treatments for specific soil types to further improve SM estimates. The dataset is publicly available on Zenodo at https://doi.org/10.5281/zenodo.14370563 (Tai et al., 2024).

Tai, Sheng-Lun [Pacific Northwest National Laborat↗

Quantitative Insight to Fission Gas Pores Distribution in Irradiated Annular U-10Zr Metallic Fuel Using Machine Learning

Metallic fuels, particularly U-10Zr and its performance in reactor irradiation conditions, have been thoroughly investigated and are a promising candidate for next-generation sodium-cooled fast spectrum nuclear reactors. Irradiation in reactors can lead to the formation of fission gas and increased pore formation which can significantly impact fuel performance. Due to the large number of pores and various phases formed in metallic fuel during irradiation, a quantitative description of fission gas pores as a function of irradiation conditions is not yet available, undermining the fidelity of fuel performance modeling to support fuel qualification. It has been difficult to clearly detect pore boundaries and distinguish matrix phases from fission gas pores using optical microscopy by using simple threshold methods working with low magnification images. The pre-trained deep learning model for fission gas pore detection was applied to ~10,260 high magnification scanning electron microscopy images. The model increased the accuracy of fission gas pore segmentation to obtain statistical features, which cannot be processed manually. A pre-trained decision tree model was used to classify pores as isolated or connected pores, providing new insight into the correlation between the movement of lanthanides, solid fission products, and the radial temperature gradient developed in fuel irradiation conditions. This paper emphasizes the potential that artificial intelligence-based machine learning models have to accelerate qualification and support nuclear fuel development.

36 MATERIALS SCIENCE↗

PROTEUS: Machine Learning Driven Resilience for Extreme-scale Systems

The objective of this project is to design, develop, and evaluate scalable software to enhance resilience, data checkpointing, program restart, and analysis. The proposed tasks are to 1) develop scalable machine learning techniques to learn temporal change patterns in a scalable and in-situ manner, and to minimize data movement and maximize learning locally closest to data; 2) design a concise data representation and indexing mechanism to capture the distribution of changes in data that can guarantee point-wise user-defined tolerable errors while reducing the data storage requirements by an order of magnitude or more; 3) develop data reduction techniques as library modules; 4) exploit local SSD for minimizing data movement in storage hierarchy; 5) develop anomaly detection algorithms that can predict corruptions based on learning of emerging patterns; 6) develop software libraries to be incorporated within widely used data formats and APIs; and 7) evaluate the proposed software using DOE scientific applications. The outcomes of the proposed work are to satisfy many synergistic data reduction and resilience requirements for large-scale data intensive applications executed on extreme-scale computing systems. The developed mechanism for error-bound data approximation is directly applicable to existing scientific applications. Through machine learning from historical events and change distribution, this work will enable anomaly detection for DOE computer facility.

97 MATHEMATICS AND COMPUTING↗

Smart sensor for online situational awareness in power grids

Waveforms in power grids typically reveal a certain pattern with specific features and peculiarities driven by the system operating conditions, internal and external uncertainties, etc. This prompts an observation of different types of waveforms at the measurement points (substations). An innovative next-generation smart sensor technology includes a measurement unit embedded with sophisticated analytics for power grid online surveillance and situational awareness. The smart sensor brings additional levels of smartness into the existing phasor measurement units (PMUs) and intelligent electronic devices (IEDs). It unlocks the full potential of advanced signal processing and machine learning for online power grid monitoring in a distributed paradigm. Within the smart sensor are several interconnected units for signal acquisition, feature extraction, machine learning-based event detection, and a suite of multiple measurement algorithms where the best-fit algorithm is selected in real-time based on the detected operating condition. Embedding such analytics within the sensors and closer to where the data is generated, the distributed intelligence mechanism mitigates the potential risks to communication failures and latencies, as well as malicious cyber threats, which would otherwise compromise the trustworthiness of the end-use applications in distant control centers. The smart sensor achieves a promising classification accuracy on multiple classes of prevailing conditions in the power grid and accordingly improves the measurement quality across the power grid.

Dehghanian, Payman↗

Noise Robustness and Experimental Demonstration of a Quantum Generative Adversarial Network for Continuous Distributions

Abstract The potential advantage of machine learning in quantum computers is a topic of intense discussion in the literature. Theoretical, numerical, and experimental explorations will most likely be required to understand its power. There have been different algorithms proposed to exploit the probabilistic nature of variational quantum circuits for generative modeling. In this paper, a hybrid architecture for quantum generative adversarial networks (QGANs) is employed and their robustness in the presence of noise is studied. A simple way of adding different types of noise to the quantum generator circuit is devised, and the noisy hybrid QGANs (HQGANs) are simulated numerically to learn continuous probability distributions, and to show that the performance of HQGANs remains unaffected. The effect of different parameters on the training time is also investigated to reduce the computational scaling of the algorithm and simplify its deployment on a quantum computer. The training on Rigetti's Aspen‐4‐2Q‐A quantum processing unit is then performed, and the results from the training are presented. The authors' results pave the way for experimental exploration of different quantum machine learning algorithms on noisy intermediate‐scale quantum devices.

Anand, Abhinav↗

Chapter 7: Learning Stable Local Volt/Var Controllers in Distribution Grids

This chapter describes a framework to synthesize provably stable local Volt/Var controllers for distributed energy resources (DERs) in power distribution grids (DGs). The goal is to control the reactive power injections of DERs to improve the system performance as quantified by a generic optimal reactive power flow (ORPF) problem. To achieve this, we jointly design for each DER the control function, which prescribes the reactive power update rule, and the equilibrium function, which approximates the ORPF solutions from local measurements of voltages and powers. We provide conditions on the equilibrium functions and the control parameters ensuring the stability of the closed-loop system. In particular, we discuss the trade-offs between each set of conditions accounting for practical considerations, like fully exploiting the DERs' generation capabilities and reducing the optimality gap. These conditions are then translated into learning constraints on the neural networks' parameters that are enforced in the training phase. We validate our framework with numerical simulations on the IEEE 37-bus network and through a comparison with an optimized version of standard piece wise linear control rules.

closed-loop asymptotic stability↗

Understanding and control of Zener pinning via phase field and ensemble learning

Zener pinning refers to the dispersion of fine particles which influences grain size distribution via movement of grain boundaries in a polycrystalline material. Grain size distribution in polycrystals has a significant impact on their properties including physical, chemical, mechanical, and optical to name a few. We explore the use of Phase-field modeling and machine-learning techniques to understand and improve the control of grain size distribution via Zener pinning in polycrystalline materials. We develop a machine learning model that determines the relative importance of various parameters to exercise microstructure control via Zener pinning. Our workflow combines high-throughput phase-field simulations and machine learning to address the computational bottlenecks associated with large-scale simulations as well as identify features necessary for microstructure control in polycrystals. A random forest (RF) regression model was developed to predict grain sizes based on five Phase-field model parameters, achieving an average prediction error of 0.72 nm for the training data and 1.44 nm for the test data. The importance of the input parameters is analyzed using the SHapley Additive exPlanations (SHAP) approach which reveals that diffusivity, volume fraction, and particle diameter are the most important parameters in determining the final grain size. These findings will allow us to select the best second-phase particles, optimize grain size distributions and thus design microstructures with the desired properties. The developed method is a highly versatile and generalizable approach that can be used to assess the combined effects of individual features in the presence of multiple variables.

36 MATERIALS SCIENCE↗

Advances in machine-learning-based sampling motivated by lattice quantum chromodynamics

Sampling from known probability distributions is a ubiquitous task in computational science, underlying calculations in domains from linguistics to biology and physics. Generative machine-learning (ML) models have emerged as a promising tool in this space, building on the success of this approach in applications such as image, text, and audio generation. Often, however, generative tasks in scientific domains have unique structures and features—such as complex symmetries and the requirement of exactness guarantees—that present both challenges and opportunities for ML. This Perspective outlines the advances in ML-based sampling motivated by lattice quantum field theory, in particular for the theory of quantum chromodynamics. Enabling calculations of the structure and interactions of matter from our most fundamental understanding of particle physics, lattice quantum chromodynamics is one of the main consumers of open-science supercomputing worldwide. Here, the design of ML algorithms for this application faces profound challenges, including the necessity of scaling custom ML architectures to the largest supercomputers, but also promises immense benefits, and is spurring a wave of development in ML-based sampling more broadly. In lattice field theory, if this approach can realize its early promise it will be a transformative step towards first-principles physics calculations in particle, nuclear and condensed matter physics that are intractable with traditional approaches.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Source Analysis of Ozone Pollution in Liaoyuan City’s Atmosphere Based on Machine Learning Models and HYSPLIT Clustering Method

Firstly, this study investigates the spatiotemporal distribution characteristics of the ozone (O 3 ) pollution in Liaoyuan City using monitoring data from 2015 to 2024. Then, three machine learning models (ML)—random forest (RF), support vector machine (SVM), and artificial neural network (ANN)—are employed to quantify the influence of meteorological and non-meteorological factors on O 3 concentrations. Finally, the HYSPLIT clustering method and CMAQ model are utilized to analyze inter-regional transport characteristics, identifying the causes of O 3 pollution. The results indicate that O 3 pollution in Liaoyuan exhibits a distinct seasonal pattern, with the highest concentrations found in spring and summer, peaking in the afternoon. Among the three ML models, the random forest model demonstrates the best predictive performance (R 2 = 0.9043). Feature importance identifies NO 2 as the primary driving factor, followed by meteorological conditions in the second quarter and land surface characteristics. Furthermore, regional transport significantly contributes to O 3 pollution, with approximately 80% of air mass trajectories in heavily polluted episodes originating from adjacent industrial areas and the sea. The combined effects of transboundary precursors and O 3 transport with local emissions and meteorological conditions further increase the O 3 pollution level. This study highlights the need to strengthen coordinated NO X and VOCs emission reductions and enhance regional joint prevention and control strategies in China.

HYSPLIT clustering↗

Using automated machine learning for the upscaling of gross primary productivity

Estimating gross primary productivity (GPP) over space and time is fundamental for understanding the response of the terrestrial biosphere to climate change. Eddy covariance flux towers provide in situ estimates of GPP at the ecosystem scale, but their sparse geographical distribution limits larger-scale inference. Machine learning (ML) techniques have been used to address this problem by extrapolating local GPP measurements over space using satellite remote sensing data. However, the accuracy of the regression model can be affected by uncertainties introduced by model selection, parameterization, and choice of explanatory features, among others. Recent advances in automated ML (AutoML) provide a novel automated way to select and synthesize different ML models. In this work, we explore the potential of AutoML by training three major AutoML frameworks on eddy covariance measurements of GPP at 243 globally distributed sites. We compared their ability to predict GPP and its spatial and temporal variability based on different sets of remote sensing explanatory variables. Explanatory variables from only Moderate Resolution Imaging Spectroradiometer (MODIS) surface reflectance data and photosynthetically active radiation explained over 70 % of the monthly variability in GPP, while satellite-derived proxies for canopy structure, photosynthetic activity, environmental stressors, and meteorological variables from reanalysis (ERA5-Land) further improved the frameworks' predictive ability. We found that the AutoML framework Auto-sklearn consistently outperformed other AutoML frameworks as well as a classical random forest regressor in predicting GPP but with small performance differences, reaching an r 2 of up to 0.75. We deployed the best-performing framework to generate global wall-to-wall maps highlighting GPP patterns in good agreement with satellite-derived reference data. This research benchmarks the application of AutoML in GPP estimation and assesses its potential and limitations in quantifying global photosynthetic activity.

54 ENVIRONMENTAL SCIENCES↗

High temperature oxidation of corrosion resistant alloys from machine learning

Parabolic rate constants, k p , were collected from published reports and calculated from corrosion product data (sample mass gain or corrosion product thickness) and tabulated for 75 alloys exposed to temperatures between ~800 and 2000 K (~500–1700 °C; 900–3000°F). Data were collected for environments including lab air, ambient and supercritical carbon dioxide, supercritical water, and steam. Materials studied include low- and high-Cr ferritic and austenitic steels, nickel superalloys, and aluminide materials. A combination of Arrhenius analysis, simple linear regression, supervised and unsupervised machine learning methods were used to investigate the relations between composition and oxidation kinetics. The supervised machine learning techniques produced the lowest mean standard errors. The most significant elements controlling oxidation kinetics were Ni, Cr, Al, and Fe, with Mo and Co composition also found to be significant features. The activation energies produced from the machine learning analysis were in the correct distributions for the diffusion constants for the oxide scales expected to dominate in each class.

Materials Science↗

Prediction of Distributed River Sediment Respiration Rates Using Community-Generated Data and Machine Learning

River sediment microbial respiration is a key indicator of ecosystem functioning and the biogeochemical fluxes across this critical zone link surface and subsurface waters. As such, there is tremendous interest in measuring and mapping these respiration rates. Respiration observations are expensive and labor intensive; there is limited data available to the community. An open science, collaborative initiative is collecting samples for respiration rate analysis and multi-scale metadata; this evolving data set is being used for making machine learning (ML) predictions at unsampled sites to help inform continued community engagement. However, it is a challenge to find an optimum configuration for ML models to work with this feature-rich (i.e., 100+ possible input variables) data set. Here, we present results from a two-tiered approach to managing the analysis of this complex data set: (a) a stacked ensemble of models that automatically optimizes hyperparameters and manages the training of many models and (b) feature permutation importance to detect the most important features in the models. The major elements of this workflow are modular, portable, open, and cloud-based thus making this implementation a potential template for other applications. The models developed here predict that sediment organic matter chemistry is one of the most important features for predicting sediment respiration rate. Other larger-scale, important features fall into the categories of climatic, ecological, geological, and fluvial settings. Leveraging these larger-scale features to generate data-driven estimates of river sediment respiration rates reveals spatially consistent but heterogeneous patterns across the river network of the Columbia River Basin.

54 ENVIRONMENTAL SCIENCES↗