Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “distributed machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Event Cause Analysis in Distribution Networks using Synchro Waveform Measurements

This paper presents a machine learning method for event cause analysis to enhance situational awareness in distribution networks. The data streams are captured using time-synchronized high sampling rates synchro waveform measurement units (SWMU). The proposed method is formulated based on a machine learning method, the convolutional neural network (CNN). This method is capable of capturing the spatiotemporal feature of the measurements effectively and perform the event cause analysis. Several events are considered in this paper to encompass a range of possible events in real distribution networks, including capacitor bank switching, transformer energization, fault, and high impedance fault (HIF). The dataset for our study is generated using the real time digital simulator (RTDS) to simulate real-world events. The event cause analysis is performed using only one cycle of the voltage waveforms after the event is detected. The simulation results show the effectiveness of the proposed machine learning-based method compared to the state-of-the-art classifiers.

Niazazari, Iman↗

1.2.4.404 - Data-Driven Approach for Hydropower Plant Controller Prototyping Using Remote Hardware in the Loop (DR-HIL)

Real-time prototyping of hydropower plant controls is important for reducing the cost and the risk of field deployment. This project will 1) collect design and operational data from actual hydro plants and 2) use a physics-informed machine learning approach for real-time emulation of hydropower plants, including hydro turbine and hydrodynamics. The data-driven models will be interfaced with digital real-time simulation at NREL’s Flatirons campus for hardware-in-the-loop (HIL) testing of the governor hardware device or controller-HIL (CHIL). The proposed approach will also establish the connectivity based remote CHIL testing capability using real-time data streams from an actual hydro plant. This integrated hydro-plant emulation with CHIL will be used to prototype hydro-governor controls and eventually provide an opportunity to test hydropower integrated with various technologies (e.g. conventional and renewable generation, energy conversion, etc.) as HIL.

controls prototyping↗

Multi-Fidelity Learning for Distribution System Voltage Probabilistic Analysis with High Penetration of PVs

This paper proposes a multi-fidelity learning approach for distribution voltage probabilistic analysis with high penetration of PVs. Unlike the existing machine learning-based approaches that require a large number of high fidelity data to achieve satisfactory results, our approach strategically leverage massive low fidelity data from inaccurate model simulations and limited high fidelity historical data. The key idea is to use low-fidelity data to establish an initial model and then the high-fidelity data to calibrate and correct the constructed low-fidelity model. This allows us to fuse low- and high-fidelity data, yielding a high fidelity prediction model. Results obtained from a realistic feeder in US with 80% penetration of PVs show that the proposed approach can achieve a similar accuracy to the one with a large number of high fidelity data. This significantly highlights the advantages of the proposed method as compared to existing data-hungry machine learning methods. Different levels of fidelity data and their impacts are also investigated.

distribution system↗

Day-Ahead Probabilistic Forecasting of Net-Load and Demand Response Potentials with High Penetration of Behind-the-Meter Solar-plus-Storage

The goal of this project is to develop advanced methods for day-ahead net-load forecasting, by leveraging the state-of-the-art machine learning techniques. The developed models produce both point and probabilistic forecasts for a variety of use cases, and are versatile to work with different types of data sets. The innovation lies in the novel design of the architectures, leveraging the most recent advances in machine learning that have not been explored in power systems, accompanied by techniques in the broader artificial intelligence fields such as fuzzy systems. This project has achieved the following accomplishments: (1) preprocessing of over 10 data sets covering varying geographical regions, time horizons, and system levels, which form a robust foundation for training and evaluating forecasting models across a wide range of realistic grid scenarios; (2) development of an interactive web app that enables exploratory analysis of load and generation data, and supports better understanding of data trends, anomalies, and correlations, facilitating model development and stakeholder engagement; (3) implementation of over 10 benchmark models for point and probabilistic forecasting, which include a mix of conventional machine learning methods and state-of-the-art deep learning approaches, providing a comprehensive baseline for performance comparison and validation of the proposed models; (4) development of a fuzzy system based gradient boosting model, tailored for small (less than 3 years) data sets, which achieves a mean absolute percentage error (MAPE) of 4% for point forecasting and a 20% improvement in average pinball loss for probabilistic forecasting; (5) development of a Transformer (a state-of-the-art deep learning architecture) based neural network model, tailored for large (3 years or more) data sets, which achieves a MAPE of 2% for point forecasting and a 20% improvement in average pinball loss for probabilistic forecasting; (6) development of a methodology for quantifying DR potential, and extensions of the previous models for multi-target forecasting of net load and DR potential, which achieve a MAPE of 10% for DR potential.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Quantum-Assisted Variational Segmentation for Image-to-Image Wildfire Detection Using Satellite Data

The quantum computing community has been searching for suitable applications to demonstrate the potential of near-term quantum devices. Quantum machine learning is a potential candidate, particularly using models that cannot be efficiently simulated with classical computers [1, 2]. This work focuses on a transition phase of quantum computers where the quantum machine learning model is still simulable classically but projected not to be simulable as the size of the model grows. Ultimately quantum computers may have advantages for high-dimensional real-world problems. Due to the limited number of qubits in current noisy intermediate-scale quantum (NISQ) devices, the direct application of quantum computers in high dimensional data is not feasible. To remedy this problem, an encoder-decoder architecture can be utilized. The encoder model would transform the high-dimensional data into a compact representation, to a level that small quantum computers can be used today (or in the near future), and the decoder would take the quantum processed outputs back to the high-dimensional space. Addressing the two challenges of quantum machine learning, this work investigates a hybrid supervised generative model with a quantum Ising Born machine embedded as the latent distribution. The model contains four main parts (Figure 1.a.): (1) a U-NET architecture responsible for learning segmentation flow, (2) a Prior network responsible for learning an encoded latent distribution of the input data, (3) a Born machine which represents the latent distribution, and (4) a Posterior network in charge of learning the joint encoded latent distribution of inputs and target data. The initial model, proposed by [3], is optimized by (1) maximizing the overlap of the prior and posterior latent distributions, and (2) minimizing the segmentation loss. The proposed model is designed to be investigated in a simulation environment applied to the real-world application of wildfire segmentation. Specifically, the model is designed to solve the patchy wildfire segmentations of Moderate Resolution Imaging Spectroradiometer (MODIS) by taking the MODIS observations and using Visible Infrared Imaging Radiometer Suite’s (VIIRS) consistent wildfire product as the target. The model solves patchy wildfire segmentations and provides insight into the epistemic errors sourced from model variation. The model utilizes the Born machine as a QUBO solver to represent the latent space as a Bernoulli distribution. The proposed configuration allows the variational segmentation model to leverage the true quantum probabilistic nature and derive a more expressive latent configuration, increasing the model performance in describing wildfire segmentations. The quantum probabilistic information of the Born machine is directly incorporated in the Kullback-Leibler divergence loss in the prior and posterior distributions, forcing the Bernoulli latent distribution to maximize the overlap of input and joint input-target distributions. The proposed model is then trained and compared with a baseline only consisting of direct Bernoulli latent distribution with no Born machine representing the latent space. The models are evaluated based on the segmentation metrics, such as precision, recall, intersect of union, with uncertainty boundaries accounting for the stochastic nature of the model. Our findings show that even in low latent-dimensional space (due to the limit in computational power of the classical quantum simulator), we are able to effectively capture the latent representation and hence the model performs better than the baseline. The findings are a projection for scaling the model into higher dimensional latent space with the Born machine surpassing the baseline performance. Figure 1. Sub-figure (a) demonstrates the architecture for the training phase. The model consists of a Prior and Posterior network that encode inputs and joint input-target data into compact representations, respectively. The Born machine represents the latent distribution, and the U-NET branch learns the segmentation patterns of the data. The stochasticity is introduced to the U-NET through its last layer to create meaningful but stochastic segmentations. Sub-figure (b) represents the inference phase where the model takes the stochastic behavior from the prior network and injects that into the U-NET. Each attempt of inference will generate different but similar segmentations from the same distribution of the wildfire event. REFERENCES [1] Coyle, B., Mills, D., Danos, V., & Kashefi, E. (2020). The Born supremacy: quantum advantage and training of an Ising Born machine. npj Quantum Information, 6(1), 1-11. [2] Liu, J. G., & Wang, L. (2018). Differentiable learning of quantum circuit born machines. Physical Review A, 98(6), 062324. [3] Kohl, S., Romera-Paredes, B., Meyer, C., De Fauw, J., Ledsam, J. R., Maier-Hein, K., ... & Ronneberger, O. (2018). A probabilistic u-net for segmentation of ambiguous images. Advances in neural information processing systems, 31.

quantum machine learning↗

Applying Machine Learning and Bayesian Inference to Identify and Locate Moving Anthropogenic Sources Using Distributed Acoustic Sensing Data

Distributed acoustic sensing (DAS) systems, which use existing telecommunication fibers, offer high‐resolution capabilities ideal for recording anthropogenic sources. However, the complexity of urban environments and the large amount of data recorded by DAS require automated methods to efficiently detect and categorize anthropogenic sources. Here, we evaluate how well three machine learning models (k‐nearest neighbor [k‐NN], convolutional neural networks, and recurrent‐convolutional neural networks) can identify various anthropogenic sources recorded by DAS. Our findings reveal that both k‐NN and neural network methods perform well in high signal‐to‐noise ratio (SNR) settings. However, their accuracy decreases at SNRs <4. We also use Kalman filtering, a form of Bayesian inference, on backprojected locations of these sources to recover locations that generally fall within standard smartphone Global Positioning System errors. By combining machine learning and Kalman filter results, we calculate a multidimensional model of moving anthropogenic sources. These results demonstrate the potential of DAS data in urban seismology for accurately identifying and locating such sources. Depending on the research objectives, these sources can be further studied or filtered out to improve the quality of seismic data for earthquake studies. Such methods provide a valuable tool for urban seismology and seismic hazard analysis.

Luckie, Thomas William [Sandia National Laboratori↗

Modeling the distribution of the endangered Jemez Mountains salamander (Plethodon neomexicanus) in relation to geology, topography, and climate

The Jemez Mountains salamander (Plethodon neomexicanus; hereafter JMS) is an endangered salamander restricted to the Jemez Mountains in north-central New Mexico, United States. This strictly terrestrial and lungless species requires moist surface conditions for activities such as mating and foraging. Threats to its current habitat include fire suppression and ensuing severe fires, changes in forest composition, habitat fragmentation, and climate change. Forest composition changes resulting from reduced fire frequency and increased tree density suggest that its current aboveground habitat does not mirror its historically successful habitat regime. However, because of its limited habitat area and underground behavior, we hypothesized that geology and topography might play a significant role in the current distribution of the salamander. We modeled the distribution of the JMS using a machine learning algorithm to assess how geology, topography, and climate variables influence its distribution. The best habitat suitability model indicates that geology type and maximum winter temperature (November to March) were most important in predicting the distribution of the salamander (23.5% and 50.3% permutation importance, respectively). Minimum winter temperature was also an important variable (21.4%), suggesting this also plays a role in salamander habitat. Our habitat suitability map reveals low uncertainty in model predictions, and we found slight discrepancies between the designated critical habitat and the most suitable areas for the JMS. Because geological features are important to its distribution, we recommend that geological and topographical data are considered, both during survey design and in the description of localities of JMS records once detected.

59 BASIC BIOLOGICAL SCIENCES↗

Power quality disturbances diagnosis: A 2D densely connected convolutional network framework

The fast and accurate diagnosis of power quality disturbances (PQD) aids in avoiding shutdowns and unnecessary procedures, concerning electric energy distribution systems. As such, a number of techniques have been tested and applied in order to reach this objective. Majority of the techniques applied are two-step based. On the first step, power quality disturbances features are extracted. Second step, considering features extracted, disturbance classification is implemented. Recently, relevant literature has presented data-driven signal processing-based approaches, as deep convolutional neural networks (DCNN), which can implement both processing steps while providing automated recognition of patterns and outliers in data. However, not considered by state-of-art, power quality disturbances are evolving in nature, while all possible regularities might not be represented in the dataset. In this work a 2 Dimension Densely Connected Convolutional Network (2D-DenseNet) framework is presented. Further, a case study with synthetic disturbance events are analyzed. Easy-to-implement formulation, built on the 2D-DenseNet, without hard-to-design parameters, highlight potential aspects for real-life implementation.

42 ENGINEERING↗

Robust Medium-Voltage Distribution System State Estimation using Multi-Source Data

Due to the lack of sufficient online measurements for distribution system observability, pseudo-measurements from short-term load or distributed renewable energy resources (DERs) forecasting are used. However, the accuracy of them is low and thus significantly limits the performance of distribution system state estimation (DSSE). In this paper, a robust DSSE that integrates multi-source measurement data is proposed. Specifically, the historical low-voltage (LV) side smart meters are used to forecast load and DERs injections via the support vector machine (SVM) with optimally tuned parameters. By contrast, the online smart meters at LV side are utilized to derive equivalent power injections at the MV/LV transformers, yielding more accurate pseudo-measurements compared to the forecasted injections. Furthermore, to deal with bad data caused by communication loss, instrumental errors and cyber attacks, robust DSSE that relies on generalized maximum-likelihood (GM)-estimation criterion is developed. The projection statistics are developed to adjust the weights of each measurement, leading to better balance between pseudo- and real-time measurements. Numerical results conducted on modified IEEE 33-bus system with DG integration demonstrate the effectiveness and robustness of the proposed method.

distribution system state estimation↗

iDDS: intelligent distributed dispatch and scheduling for workflow orchestration

The intelligent distributed dispatch and scheduling (iDDS) service is a versatile workflow orchestration system designed for large-scale, distributed scientific computing. iDDS extends traditional workload and data management by integrating data-aware execution, conditional logic, and programmable workflows, enabling automation of complex and dynamic processing pipelines. Originally developed for the ATLAS experiment at the large hadron collider, iDDS has evolved into an experiment-agnostic platform that supports both template-driven workflows and a Function-as-a-Task model for Python-based orchestration. This paper presents the architecture and core components of iDDS, highlighting its scalability, modular message-driven design, and integration with systems such as PanDA and Rucio. We demonstrate its versatility through real-world use cases: fine-grained tape resource optimization for ATLAS, orchestration of large Directed Acyclic Graph (DAG) workflows for the Rubin Observatory, distributed hyperparameter optimization for machine learning applications, active learning for physics analyses, and AI-assisted detector design at the electron–ion collider. By unifying workload scheduling, data movement, and adaptive decision-making, iDDS reduces operational overhead and enables reproducible, high-throughput workflows across heterogeneous infrastructures. We conclude with current challenges and future directions, including interactive, cloud-native, and serverless workflow support.

97 MATHEMATICS AND COMPUTING↗

High-throughput virtual laboratory for drug discovery using massive datasets

Time-to-solution for structure-based screening of massive chemical databases for COVID-19 drug discovery has been decreased by an order of magnitude, and a virtual laboratory has been deployed at scale on up to 27,612 GPUs on the Summit supercomputer, allowing an average molecular docking of 19,028 compounds per second. Over one billion compounds were docked to two SARS-CoV-2 protein structures with full optimization of ligand position and 20 poses per docking, each in under 24 hours. GPU acceleration and high-throughput optimizations of the docking program produced 350× mean speedup over the CPU version (50× speedup per node). GPU acceleration of both feature calculation for machine-learning based scoring and distributed database queries reduced processing of the 2.4 TB output by orders of magnitude. The resulting 50× speedup for the full pipeline reduces an initial 43 day runtime to 21 hours per protein for providing high-scoring compounds to experimental collaborators for validation assays.

97 MATHEMATICS AND COMPUTING↗

Integrated Approach to Ancillary PV Component Reliability Assessment (Final Report)

In this project, we have established a nondestructive, generalized methodology that (1) fuses rich field data with advanced ML for proactive reliability forecasting, (2) dramatically reduces experimental iterations via synthetic dataset generation, and (3) achieves unprecedented regression precision in both anomaly detection and component-level degradation assessment—paving the way for truly predictive maintenance of grid-tied PV inverters under diverse outdoor conditions.

14 SOLAR ENERGY↗

A 1 km soil moisture dataset over eastern CONUS generated by assimilating SMAP data into the Noah-MP land surface model

An improved fine-scale soil moisture (SM) dataset at 1 km grid spacing, covering much of the eastern continental US, was generated by assimilating 9 km Soil Moisture Active Passive (SMAP) SM data into the v4.0.1 Noah-MP land surface model. With 12 ensemble members, the assimilation was carried out using the ensemble Kalman filter algorithm within NASA's Land Information System. The SM analysis for 2016 was fully validated against in situ observations from four different networks and compared with four other existing datasets. Results indicate that this SM analysis surpasses other datasets in top-layer SM distribution, including a machine-learning-based product, despite all SM estimates being less heterogeneous than observed. The analysis of anomalous errors suggests that large similarity in intrinsic errors is likely due to overlapping data sources among the selected SM datasets. More detailed evaluations were performed over two geographic areas. The observations collected by the Atmospheric Radiation Measurement facility in Oklahoma suggest that soil temperature and surface heat fluxes are concurrently simulated with good accuracy. Investigation into the 2016 southeastern US drought response further indicates drier conditions and higher evapotranspiration estimates compared to GLEAMv4.1. Notably, large errors are associated with grids having clay soil textures, underscoring the need for refined model treatments for specific soil types to further improve SM estimates. The dataset is publicly available on Zenodo at https://doi.org/10.5281/zenodo.14370563 (Tai et al., 2024).

Tai, Sheng-Lun [Pacific Northwest National Laborat↗

Quantitative Insight to Fission Gas Pores Distribution in Irradiated Annular U-10Zr Metallic Fuel Using Machine Learning

Metallic fuels, particularly U-10Zr and its performance in reactor irradiation conditions, have been thoroughly investigated and are a promising candidate for next-generation sodium-cooled fast spectrum nuclear reactors. Irradiation in reactors can lead to the formation of fission gas and increased pore formation which can significantly impact fuel performance. Due to the large number of pores and various phases formed in metallic fuel during irradiation, a quantitative description of fission gas pores as a function of irradiation conditions is not yet available, undermining the fidelity of fuel performance modeling to support fuel qualification. It has been difficult to clearly detect pore boundaries and distinguish matrix phases from fission gas pores using optical microscopy by using simple threshold methods working with low magnification images. The pre-trained deep learning model for fission gas pore detection was applied to ~10,260 high magnification scanning electron microscopy images. The model increased the accuracy of fission gas pore segmentation to obtain statistical features, which cannot be processed manually. A pre-trained decision tree model was used to classify pores as isolated or connected pores, providing new insight into the correlation between the movement of lanthanides, solid fission products, and the radial temperature gradient developed in fuel irradiation conditions. This paper emphasizes the potential that artificial intelligence-based machine learning models have to accelerate qualification and support nuclear fuel development.

36 MATERIALS SCIENCE↗

PROTEUS: Machine Learning Driven Resilience for Extreme-scale Systems

The objective of this project is to design, develop, and evaluate scalable software to enhance resilience, data checkpointing, program restart, and analysis. The proposed tasks are to 1) develop scalable machine learning techniques to learn temporal change patterns in a scalable and in-situ manner, and to minimize data movement and maximize learning locally closest to data; 2) design a concise data representation and indexing mechanism to capture the distribution of changes in data that can guarantee point-wise user-defined tolerable errors while reducing the data storage requirements by an order of magnitude or more; 3) develop data reduction techniques as library modules; 4) exploit local SSD for minimizing data movement in storage hierarchy; 5) develop anomaly detection algorithms that can predict corruptions based on learning of emerging patterns; 6) develop software libraries to be incorporated within widely used data formats and APIs; and 7) evaluate the proposed software using DOE scientific applications. The outcomes of the proposed work are to satisfy many synergistic data reduction and resilience requirements for large-scale data intensive applications executed on extreme-scale computing systems. The developed mechanism for error-bound data approximation is directly applicable to existing scientific applications. Through machine learning from historical events and change distribution, this work will enable anomaly detection for DOE computer facility.

97 MATHEMATICS AND COMPUTING↗

Smart sensor for online situational awareness in power grids

Waveforms in power grids typically reveal a certain pattern with specific features and peculiarities driven by the system operating conditions, internal and external uncertainties, etc. This prompts an observation of different types of waveforms at the measurement points (substations). An innovative next-generation smart sensor technology includes a measurement unit embedded with sophisticated analytics for power grid online surveillance and situational awareness. The smart sensor brings additional levels of smartness into the existing phasor measurement units (PMUs) and intelligent electronic devices (IEDs). It unlocks the full potential of advanced signal processing and machine learning for online power grid monitoring in a distributed paradigm. Within the smart sensor are several interconnected units for signal acquisition, feature extraction, machine learning-based event detection, and a suite of multiple measurement algorithms where the best-fit algorithm is selected in real-time based on the detected operating condition. Embedding such analytics within the sensors and closer to where the data is generated, the distributed intelligence mechanism mitigates the potential risks to communication failures and latencies, as well as malicious cyber threats, which would otherwise compromise the trustworthiness of the end-use applications in distant control centers. The smart sensor achieves a promising classification accuracy on multiple classes of prevailing conditions in the power grid and accordingly improves the measurement quality across the power grid.

Dehghanian, Payman↗