Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “distributed machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

i- flow: High-dimensional integration and sampling with normalizing flows

In many fields of science, high-dimensional integration is required. Numerical methods have been developed to evaluate these complex integrals. We introduce the code i-flow, a python package that performs high-dimensional numerical integration utilizing normalizing flows. Normalizing flows are machine-learned, bijective mappings between two distributions. i-flow can also be used to sample random points according to complicated distributions in high dimensions. We compare i-flow to other algorithms for high-dimensional numerical integration and show that i-flow outperforms them for high dimensional correlated integrals. The i-flow code is publicly available on gitlab at https://gitlab.com/i-flow/i-flow.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

VST ATLAS galaxy cluster catalogue I: cluster detection and mass calibration

Taking advantage of ∼4700 deg 2 optical coverage of the Southern sky offered by the VST ATLAS survey, we construct a new catalogue of photometrically selected galaxy groups and clusters using the orca cluster detection algorithm. The catalogue contains ∼22 000 detections with N 200 > 10 and ∼9000 with N 200 > 20. We estimate the photometric redshifts of the clusters using machine learning and find the redshift distribution of the sample to extend to z ∼ 0.7, peaking at z ∼ 0.25. We calibrate the ATLAS cluster mass-richness scaling relation using masses from the MCXC, Planck, ACT DR5, and SDSS redMaPPer cluster samples. We estimate the ATLAS sample to be > 95 per cent complete and > 85 per cent pure at z < 0.35 and in the M 200m >1 x 10 14 h -1 M ⊙ mass range. At z < 0.35, we also find the ATLAS sample to be more complete than redMaPPer, recovering a ~ 40 per cent higher fraction of Abell clusters. This higher sample completeness places the amplitude of the z < 0.35 ATLAS cluster mass function closer to the predictions of a ΛCDM model with parameters based on the Planck CMB analyses, compared to the mass functions of the other cluster samples. However, strong tensions between the observed ATLAS mass functions and models remain. We shall present a detailed cosmological analysis of the ATLAS cluster mass functions in paper II. In the future, optical counterparts to X-ray-detected eROSITA clusters can be identified using the ATLAS sample. The catalogue is also well suited for auxiliary spectroscopic target selection in 4MOST. The ATLAS cluster catalogue is publicly available at http://astro.dur.ac.uk/cosmology/vstatlas/cluster_catalogue/.

79 ASTRONOMY AND ASTROPHYSICS↗

VDiSC: An Open Source Framework for Distributed Smart City Vision and Biometric Surveillance Networks

Recent global growth in the interest of smart cities has led to trillions of dollars of investment toward research and development. These connected cities have the potential to create a symbiosis of technology and society and revolutionize the cost of living, safety, ecological sustainability, and quality of life of societies on a world-wide scale. Some key components of the smart city construct are connected smart grids, self-driving cars, federated learning systems, smart utilities, large-scale public transit, and proactive surveillance systems. While exciting in prospect, these technologies and their subsequent integration cannot be attempted without addressing the potential societal impacts of such a high degree of automation and data sharing. Additionally, the feasibility of coordinating so many disparate tasks will require a fast, extensible, unifying framework. To that end, we propose the Distributed Smart City framework for Vision, or VDiSC. VDiSC serves as a unified biometric API harness that allows for seamless evaluation, deployment, and simple pipeline creation for heterogeneous biometric software. VDiSC additionally provides a fully declarative capability for defining and coordinating custom machine learning and sensor pipelines, allowing the distribution of processes across otherwise incompatible hardware and networks. VDiSC ultimately provides a way to quickly configure, hot-swap, and expand large coordinated or federated systems online without interruptions for maintenance. Because much of the data collected in a smart city contains Personally Identifying Information (PII), VDiSC also provides built-in tools and layers to ensure secure and encrypted streaming, storage, and access of PII data across distributed systems.

Brogan, Joel↗

Deconstruction by C. thermocellum —from microbe mediated to dynamic redistribution of cellulosomes

Clostridium thermocellum is one of the most efficient microorganisms for the deconstruction of cellulosic biomass. To achieve this high level of cellulolytic activity, C. thermocellum uses large multienzyme complexes known as cellulosomes to break down complex polysaccharides, notably cellulose, found in plant cell walls. The attachment of bacterial cells to the nearby substrate via the cellulosome has been hypothesized to be the reason for this high efficiency. The region lying between the cell and the substrate has shown great variation and dynamics that are affected by the growth stage of cells and the substrate used for growth. Here, we used both super-resolution imaging and machine-learning approaches to study the distribution of C. thermocellum cellulosomes at different stages of growth. We show that C. thermocellum initially retains its cellulosomes primarily on the cell surface but then relocates large cellulosome clusters to the interface with biomass, therefore depleting its cell surface of cellulosomes. These results indicate dynamic redistribution of cellulosomes during growth, with a functional shift toward substrate-associated degradation later during growth on biomass.

09 BIOMASS FUELS↗

Modular performance prediction for scientific workflows using Machine Learning

Scientific workflows provide an opportunity for declarative computational experiment design in an intuitive and efficient way. A distributed workflow is typically executed on a variety of resources, and it uses a variety of computational algorithms or tools to achieve the desired outcomes. Such a variety imposes additional complexity in scheduling these workflows on large scale computers. As computation becomes more distributed, insights into expected workload that a workflow presents become critical for effective resource allocation. In this paper, we present a modular framework that leverages Machine Learning for creating precise performance predictions of a workflow. The central idea is to partition a workflow in such a way that makes the task of forecasting each atomic unit manageable and gives us a way to combine the individual predictions efficiently. We recognize a combination of an executable and a specific physical resource as a single module. This gives us a handle to characterize workload and machine power as a single unit of prediction. Overall, our modular technique of creating atomic modules and deployment of longest-path approach to estimate workflow performance, allows the framework to adapt to highly complex nested directed acyclic workflows and scale to new scenarios, since it does not make assumptions of underlying workflow structure. We present performance estimation results of independent workflow modules executed on the XSEDE SDSC Comet cluster using various Machine Learning algorithms. The results provide insights into the behavior and effectiveness of different algorithms in the context of scientific workflow performance prediction.

97 MATHEMATICS AND COMPUTING↗

Machine learning for geophysical characterization of brittleness: Tuscaloosa Marine Shale case study

Brittleness is one of the most important reservoir properties for unconventional reservoir exploration and production. Better knowledge about the brittleness distribution can help to optimize the hydraulic fracturing operation and lower costs. However, there are very few reliable and effective physical models to predict the spatial distribution of brittleness. We have developed a machine learning-based method to predict subsurface brittleness by using multidiscipline data sets, such as seismic attributes, rock physics, and petrophysics information, which allows us to implement the prediction without using a physical model. The method is applied on a data set from Tuscaloosa Marine Shale, and the predicted rock physics template is close to the calculated value from conventional inverted elastic parameters. Therefore, the proposed method helps determine areas of the reservoir that have optimal geomechanical properties for successful hydraulic fracturing.

Geochemistry & Geophysics↗

Out-of-Distribution Detection and Radiological Data Monitoring Using Statistical Process Control

Abstract Machine learning (ML) models often fail with data that deviates from their training distribution. This is a significant concern for ML-enabled devices as data drift may lead to unexpected performance. This work introduces a new framework for out of distribution (OOD) detection and data drift monitoring that combines ML and geometric methods with statistical process control (SPC). We investigated different design choices, including methods for extracting feature representations and drift quantification for OOD detection in individual images and as an approach for input data monitoring. We evaluated the framework for both identifying OOD images and demonstrating the ability to detect shifts in data streams over time. We demonstrated a proof-of-concept via the following tasks: 1) differentiating axial vs. non-axial CT images, 2) differentiating CXR vs. other radiographic imaging modalities, and 3) differentiating adult CXR vs. pediatric CXR. For the identification of individual OOD images, our framework achieved high sensitivity in detecting OOD inputs: 0.980 in CT, 0.984 in CXR, and 0.854 in pediatric CXR. Our framework is also adept at monitoring data streams and identifying the time a drift occurred. In our simulations tracking drift over time, it effectively detected a shift from CXR to non-CXR instantly, a transition from axial to non-axial CT within few days, and a drift from adult to pediatric CXRs within a day—all while maintaining a low false positive rate. Through additional experiments, we demonstrate the framework is modality-agnostic and independent from the underlying model structure, making it highly customizable for specific applications and broadly applicable across different imaging modalities and deployed ML models.

Zamzmi, Ghada↗

Explainable machine learning for incipient anomaly detection in compact molten salt heat exchanger with overlapping feature distributions

High-temperature molten salt-cooled reactors (MSCRs) are a promising next-generation nuclear technology option, offering efficient power conversion and inherent safety features. However, the reliability of these systems depends on the robust operation of heat exchangers (HXs), which are susceptible to failure due to temperature gradients and channel plugging caused by fluid freezing. Conventional monitoring methods, relying on inlet and outlet measurements, lack the spatial resolution needed to detect early-stage faults. We propose a novel design of a compact salt-to-salt matrix-type HX design consisting of interleaved arrays of parallel tubes, with integrated synthetic fiber optic distributed temperature sensing (DTS) to enable localized detection of incipient faults. To evaluate performance of this design, we generate high-fidelity synthetic data using heat transfer computational modeling to simulate channel plugging, and introduce sensor noise for realistic modeling of measurements. The dataset comprises of 97% normal operation and 3% anomaly cases, with each anomaly class representing 1% of the data. These early anomalies result in overlapping temperature profiles between normal and faulty channels, producing a non-separable dataset that challenges traditional classification techniques. We benchmark eight supervised machine learning (ML) models and demonstrate that XGBoost achieves the highest performance. To improve transparency, we develop an explainability framework combining Shapley values and partially ordered sets (POSETs) to quantify and structurally analyze feature importance. This approach identifies both dominant predictors and ambiguous feature relationships, enhancing trust and interpretability. Our results highlight the potential of combining DTS and explainable ML with intelligent feature selection to improve predictive maintenance and ensure operational resilience in advanced nuclear systems.

Prantikos, Konstantinos [Argonne National Laborato↗

Fiber optic computing using distributed feedback

Abstract The widespread adoption of machine learning and other matrix intensive computing algorithms has renewed interest in analog optical computing, which has the potential to perform large-scale matrix multiplications with superior energy scaling and lower latency than digital electronics. However, most optical techniques rely on spatial multiplexing, requiring a large number of modulators and detectors, and are typically restricted to performing a single kernel convolution operation per layer. Here, we introduce a fiber-optic computing architecture based on temporal multiplexing and distributed feedback that performs multiple convolutions on the input data in a single layer. Using Rayleigh backscattering in standard single mode fiber, we show that this technique can efficiently apply a series of random nonlinear projections to the input data, facilitating a variety of computing tasks. The approach enables efficient energy scaling with orders of magnitude lower power consumption than GPUs, while maintaining low latency and high data-throughput.

97 MATHEMATICS AND COMPUTING↗

Roadmap and Benchmarking: Privacy in Federated Load Forecasting

Data-driven techniques for energy demand forecasting continue to emerge with promising impacts on distribution grid planning. However, the development of robust and generalizable machine learning models requires that representative high quality training data are available. Distributed energy resources have begun to embed intelligence, gathering large amounts of data on customer demand, behavior, and household devices that are connected to the grid. Though utilities aggregate meter-level demand data for load shaping, demand response, outage management, reliability planning, and billing applications, there lies an inherent privacy concern in sharing consumption data that may identify individual consumer behavioral patterns. Hence, while sharing the data is crucial, the private sensitive customer data must be safeguarded from being exposed or manipulated. In this study, we propose a roadmap for implementing a based privacy preserving framework to support the advancement of data-driven analytics in data-sensitive distributed energy resources environments. The roadmap incorporates federated learning–a distributed training framework, differential privacy–a statistical framework that provides guarantees to safeguard the leakage of sensitive data, secure multiparty computation and homomorphic encryption– techniques for encrypting model gradients and applying secure aggregation on the server. Moreover, we perform baseline experiments on the federated short-term load forecasting (STLF) task using open-source residential load profile datasets, offering insights into the challenges of integrating differential privacy into federated learning.

Abebe, Waqwoya [Oak Ridge National Laboratory (ORN↗

Extracting the gamma-ray source-count distribution below the Fermi-LAT detection limit with deep learning

We reconstruct the extra-galactic gamma-ray source-count distribution, or dN/dS, of resolved and unresolved sources by adopting machine learning techniques. Specifically, we train a convolutional neural network on synthetic 2-dimensional sky-maps, which are built by varying parameters of underlying source-counts models and incorporate the Fermi-LAT instrumental response functions. The trained neural network is then applied to the Fermi-LAT data, from which we estimate the source count distribution down to flux levels a factor of 50 below the Fermi-LAT threshold. We perform our analysis using 14 years of data collected in the (1,10) GeV energy range. The results we obtain show a source count distribution which, in the resolved regime, is in excellent agreement with the one derived from cataloged sources, and then extends as dN/dS ~ S -2 in the unresolved regime, down to fluxes of 5 · 10 -12 cm -2 s -1 . The neural network architecture and the devised methodology have the flexibility to enable future analyses to study the energy dependence of the source-count distribution.

79 ASTRONOMY AND ASTROPHYSICS↗

Out-of-distribution generalization for learning quantum dynamics

Generalization bounds are a critical tool to assess the training data requirements of Quantum Machine Learning (QML). Recent work has established guarantees for in-distribution generalization of quantum neural networks (QNNs), where training and testing data are drawn from the same data distribution. However, there are currently no results on out-of-distribution generalization in QML, where we require a trained model to perform well even on data drawn from a different distribution to the training distribution. Here, we prove out-of-distribution generalization for the task of learning an unknown unitary. In particular, we show that one can learn the action of a unitary on entangled states having trained only product states. Since product states can be prepared using only single-qubit gates, this advances the prospects of learning quantum dynamics on near term quantum hardware, and further opens up new methods for both the classical and quantum compilation of quantum circuits.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Preliminary Results for Using Uncertainty and Out-of-distribution Detection to Identify Unreliable Predictions.

As machine learning (ML) models are deployed into an ever-diversifying set of application spaces, ranging from self-driving cars to cybersecurity to climate modeling, the need to carefully evaluate model credibility becomes increasingly important. Uncertainty quantification (UQ) provides important information about the ability of a learned model to make sound predictions, often with respect to individual test cases. However, most UQ methods for ML are themselves data-driven and therefore susceptible to the same knowledge gaps as the models themselves. Specifically, UQ helps to identify points near decision boundaries where the models fit the data poorly, yet predictions can score as certain for points that are under-represented by the training data and thus out-of-distribution (OOD). One method for evaluating the quality of both ML models and their associated uncertainty estimates is out-of-distribution detection (OODD). We combine OODD with UQ to provide insights into the reliability of the individual predictions made by an ML model.

97 MATHEMATICS AND COMPUTING↗

Machine Learning-Based PV Reserve Determination Strategy for Frequency Control on the WECC System

This paper proposes a machine learning based strategy, that is suitable for real-time operation, to determine the optimal photovoltaic (PV) power plants reserve for frequency control. The proposed machine learning algorithm is trained and tested on 1,987 offline simulations of a 60% renewable penetration Western Electricity Coordinating Council (WECC) system. On a realistic 1-day operation profile of the WECC system, the ML model demonstrates a savings of more than 40% PV headroom compared to a conservative approach.

14 SOLAR ENERGY↗

Out-of-distribution generalization for learning quantum dynamics

Abstract Generalization bounds are a critical tool to assess the training data requirements of Quantum Machine Learning (QML). Recent work has established guarantees for in-distribution generalization of quantum neural networks (QNNs), where training and testing data are drawn from the same data distribution. However, there are currently no results on out-of-distribution generalization in QML, where we require a trained model to perform well even on data drawn from a different distribution to the training distribution. Here, we prove out-of-distribution generalization for the task of learning an unknown unitary. In particular, we show that one can learn the action of a unitary on entangled states having trained only product states. Since product states can be prepared using only single-qubit gates, this advances the prospects of learning quantum dynamics on near term quantum hardware, and further opens up new methods for both the classical and quantum compilation of quantum circuits.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Anomaly Detection and Mitigation for Wide-Area Damping Control using Machine Learning

In an interconnected multi-area power system, wide-area measurement based damping controllers are used to damp out inter-area oscillations, which jeopardize grid stability and constrain the power flows below to their transmission capacity. The effect of wide-area damping control (WADC) significantly depends on both power and cyber systems. At the cyber system layer, an adversary can inflict the WADC process by compromising either measurement signals, control signals or both. Stealthy and coordinated cyber-attacks may bypass the conventional cybersecurity measures to disrupt the seamless operation of WADC. This paper proposes an anomaly detection (AD) algorithm using supervised Machine Learning and a model-based logic for mitigation. The proposed AD algorithm considers measurement signals (input of WADC) and control signals (output of WADC) as input to evaluate the type of activity such as normal, perturbation (small or large signal faults), attack and perturbation-and-attack. Upon anomaly detection, the mitigation module tunes the WADC signal and sets the control status mode as either wide-area mode or local mode. The proposed anomaly detection and mitigation (ADM) module works inline with the WADC at the control center for attack detection on both measurement and control signals and eliminates the need for ADMs at the geographically distributed actuators. Here, we consider coordinated and primitive data-integrity attack vectors such as pulse, ramp, relay-trip and replay attacks. The performance of the proposed ADM algorithms was evaluated under these attack vector scenarios on a testbed environment for 2-area 4-machine power system. The ADM module shows effective performance with 96:5% accuracy to detect anomalies.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗