Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “decentralized learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

70 records · Page 4

An Intelligent Distributed Ledger Construction Algorithm for IoT

Blockchain is the next generation of secure data management that creates near-immutable decentralized storage. Secure cryptography created a niche for blockchain to provide alternatives to well-known security compromises. However, design bottlenecks with traditional blockchain data structures scale poorly with increased network usage and are extremely computation-intensive. This made the technology difficult to combine with limited devices, like those in Internet of Things networks. In protocols like IOTA, replacement of blockchain's linked-list queue processing with a lightweight dynamic ledger showed remarkable throughput performance increase. However, current stochastic algorithms for ledger construction suffer distinct trade-offs between efficiency and security. This work proposed a machine-learning approach with a multi-arm bandit that resolved these issues and was designed for auditing on limited devices. This algorithm was tested in a reinforcement-learning environment simulating the IOTA ledger's construction with a decision tree. This study showed through regret analysis and experimentation that this approach was secure against impulse manipulation attacks while remaining energy-efficient. Although the IOTA protocol was a pioneer for lightweight distributed ledgers, it is expected that future blockchain protocols will adopt techniques similar to those presented in this work.

multi-arm bandit↗

Enabling Cybersecurity, Situational Awareness and Resilience in Distribution Grids with High Penetration of Photovoltaics (CARE-PV) (Final Report)

Since legacy distribution systems have very limited visibility beyond the substation, high penetration of PV at the grid edge presents some unique operational challenges. One approach to address these challenges is to use information from advanced metering infrastructure (AMI) and µPMUs. However, exploiting this information is impacted by a number of factors, including multi-timescale measurements, volume of data generated, communication network impairments (e.g., information loss and latency) and susceptibility to cyber-attacks. Therefore, one of the critical tasks involved in the management of a distribution grid is to develop complete situational awareness by integrating cyber-security mechanisms with state estimation strategies and leveraging this situational awareness to assure energy services at strategic locations while exploiting AMI/PV inverter/ µPMU data. This CARE-PV project addresses the fundamental challenges in situational awareness and resilience to cyber and physical vectors by exploiting the synergy between innovative modeling, estimation, data analytics, testing and validation using smart PV inverters designed at K-State and facilities at NREL. Specifically, the project involved the development, testing and validation of the following novel enabling technologies: (Thrust 1) Resilience to cyber vectors that impact data integrity was addressed via a two-level defense strategy that combines cyber intrusion detection using self-learning, cooperative smart PV inverters, and a novel moving target defense framework to combat data integrity attacks. (Thrust 2) Resilience to cyber-physical vectors that impact situational awareness by limiting data availability was addressed via novel centralized and decentralized, sparsity-based static and dynamic state estimation approaches that enhance observability even when the underlying system is unobservable. (Thrust 3) Leveraging a unique probabilistic sensitivity analysis approach accompanied by one-of-a-kind dominant influencer set computation, the vulnerability of critical infrastructure at strategic locations was evaluated so that proactive PV-based control strategies can be used to support operations under normal/outage scenarios. These CARE-PV project innovations were demonstrated on both small-scale IEEE and larger utility-scale testbeds (Thrust 4). Feedback from Industry Advisory Board members was used to formulate a commercialization pathway for a subset of CARE-PV technologies. These CARE-PV technologies will ultimately lead to reliable and secure, large-scale integration of renewable energy and mitigate the risk of energy disruption resulting from cyber incidents and other emerging threats within the energy environment.

14 SOLAR ENERGY↗

Efficient Anomaly Detection Driven By Different Machine Learning Architectures And Models

The rapid growth and ubiquitous adoption of the internet and cyber-physical systems (CPS) have fundamentally transformed modern communication, work, and human-system interactions. While networks now form the backbone of critical digital ecosystems, enabling seamless data transmission across diverse, interconnected systems, this increased connectivity also expands the attack surface, making real-time detection of network intrusions and anomalies a pressing challenge. Detecting unusual activities within network infrastructure requires advanced data traffic analysis to differentiate between legitimate and malicious interactions. Traditional approaches to network anomaly detectionâ??such as rule-based and signature-based systemsâ??often depend on predefined patterns to identify known anomalies, limiting their effectiveness against emerging, stealthy, or previously unseen threats. These conventional methods suffer from high false alarm rates and fail to adapt to the ever-evolving nature of network traffic, particularly in large-scale, decentralized environments where data volume, velocity, and variety are constantly increasing. This dissertation presents artificial intelligence (AI)-driven approaches to anomaly detection that leverage graphics processing unit (GPU)-enabled high-performance computing (HPC) platforms for processing massive network traffic data and monitoring the components of cyber-physical systems (CPS) for potentially hazardous conditions. The research advances several key contributions: (1) Designing efficient machine learning techniques for CPS condition monitoring and anomaly detection; (2) enabling federated learning (FL) frameworks that enable distributed detection while preserving data privacy and system resilience; (3) exploring graph-based methodologies combining graph neural networks (GNN) and graph machine learning (ML) approaches for the Internet of Things (IoT) and automotive network security, and (4) performing distributed edge computing optimizations that integrate FL with scalable technologies for reduced communication overhead. Through extensive experiments, these methodologies demonstrate that complex anomaly detection and condition monitoring tasks can be achieved while balancing computational efficiency and detection accuracy through fine-grained network information processing. The frameworks developed in this research establish a robust foundation for network anomaly detection, providing scalable, adaptive, and privacy-preserving solutions for safeguarding CPS and IoT networks in an increasingly interconnected digital landscape. The practical implications of these research findings are significant, as they can inform the development of next-generation network security systems and contribute to the protection of critical infrastructure against sophisticated cyber attacks.

Marfo, William↗

Privacy-Preserving Real-Time Action Detection in Intelligent Vehicles Using Federated Learning-Based Temporal Recurrent Network

This study introduces a privacy-preserving approach for the real-time action detection in intelligent vehicles using a federated learning (FL)-based temporal recurrent network (TRN). This approach enables edge devices to independently train models, enhancing data privacy and scalability by eliminating central data consolidation. Our FL-based TRN effectively captures temporal dependencies, anticipating future actions with high precision. Extensive testing on the Honda HDD and TVSeries datasets demonstrated robust performance in centralized and decentralized settings, with competitive mean average precision (mAP) scores. The experimental results highlighted that our FL-based TRN achieved an mAP of 40.0% in decentralized settings, closely matching the 40.1% in centralized configurations. Notably, the model excelled in detecting complex driving maneuvers, with mAPs of 80.7% for intersection passing and 78.1% for right turns. These outcomes affirm the model’s accuracy in action localization and identification. The system showed significant scalability and adaptability, maintaining robust performance across increased client device counts. The integration of a temporal decoder enabled predictions of future actions up to 2 s ahead, enhancing the responsiveness. Our research advances intelligent vehicle technology, promoting safety and efficiency while maintaining strict privacy standards.

33 ADVANCED PROPULSION SYSTEMS↗

Decentralized Distributed Proximal Policy Optimization (DD-PPO) for High Performance Computing Scheduling on Multi-User Systems

Resource allocation in High Performance Computing (HPC) environments presents a complex and multifaceted challenge for job scheduling algorithms. Beyond the efficient allocation of system resources, schedulers must account for and optimize multiple performance metrics, including job wait time and system throughput. Traditional heuristic-based scheduling algorithms increasingly struggle and lack the efficiency needed to meet the demands and address the complexity and scale of modern HPC systems. Consequently, recent research efforts have focused on leveraging advancements in Artificial Intelligence (AI) and Deep Learning (DL), particularly Reinforcement Learning (RL), to develop more adaptable and intelligent scheduling strategies. Previous RL-based scheduling approaches have explored a range of algorithms, from Deep Q-Networks (DQN) to Proximal Policy Optimization (PPO), and more recently, hybrid methods that integrate Graph Neural Networks (GNNs) with RL techniques. However, a common limitation across these methods is their reliance on relatively small datasets, with few methods being evaluated using large-scale, multi-million-job trace datasets representative of real-world HPC workloads. Moreover, existing RL schedulers face scalability issues due to centralized policy updates, which hinder training efficiency and performance when applied to large datasets. This study introduces a novel RL-based scheduler utilizing Decentralized Distributed Proximal Policy Optimization (DD-PPO) algorithm, which supports large-scale distributed training across multiple workers without requiring parameter synchronization at every step. By eliminating reliance on centralized updates to a shared policy, the DD-PPO scheduler enhances scalability, training efficiency, and sample utilization. Experimental validation using a large real-world dataset containing over 11.5 million job traces collected from petascale HPC systems over six years assesses the influence of dataset scale on training effectiveness and compares DD-PPO performance to traditional and advanced scheduling approaches. The experimental results demonstrate improved scheduling performance in comparison to both heuristic-based schedulers and existing RL-based scheduling algorithms.

AI↗

Multi-scale planning model for robust urban drought response

Increasingly severe droughts are straining municipal water resources and jeopardizing urban water security, but uncertainty in their duration, frequency, and intensity challenges drought planning and response. We develop the Drought Resilient Interscale Portfolio Planning model (DRIPP) to generate optimal planning responses to urban drought. DRIPP is a generalizable multi-scale framework for optimizing dynamic planning strategies of long-term infrastructure deployment and short-term drought response. It integrates climate and hydrological variability with high-fidelity representations of urban water distribution, available technology options, and demand reduction measures to yield robust and cost-effective water supply portfolios that are location-specific. We apply DRIPP in Santa Barbara, California to assess how least cost water supply portfolios vary under different drought scenarios and identify portfolios that are robust across drought scenarios. In Santa Barbara, we find that drought intensity, not duration or frequency, drives cost increases, reliability risk, and regret of overbuilding infrastructure. Under uncertain drought conditions, a diversified technology portfolio that includes both rapidly deployable, decentralized technologies alongside larger centralized technologies minimizes water supply cost while maintaining high robustness to climate uncertainty.

54 ENVIRONMENTAL SCIENCES↗

Secure mmWave Spectrum Sharing with Autonomous Beam Scheduling for 5G and Beyond

Spectrum Sharing (SS) has seen a renewed set of initiatives in 5G with the availability of shared and unlicensed spectrum bands that can be used by multiple cellular service providers and private cellular networks. Beam based transmission, instead of the traditional sector based transmission in conjunction with the spectrum agility of the 5G New Radio (NR) has brought new opportunities to optimized sharing of spectrum. Currently in the U.S., a centralized Spectrum Access Server (SAS) is used to co-ordinate spectrum sharing among networks sharing the same spectrum band. However, SAS becomes a focal point for security attacks and a performance bottleneck. In addition, SAS relies on an Environmental Sensor Network (ESN), separate from the 5G network. Without trusted spectral occupancy information, false reporting of spectrum sensing data can create sub-optimal and unfair spectrum usage. This paper summarizes our recent research findings in using a decentralized scheme for multiple networks to securely share spectrum with autonomous beam scheduling : 1) A new stochastic network framework based on Lyapunov Optimization approach is developed to optimize scheduling at the base stations; 2) Game theoretic (GT) approach is used to formulate the distributed scheduler; 3) Another distributed scheduler with Q-learning is presented that utilizes the Reinforcement Learning (RL) approach; 4) The performance and convergence rate of these distributed solutions to use shared and unlicensed spectrum are compared with existing solutions. Conditions under which the performance of these schedulers approach the theoretical upper bound, which is the performance possible with no interference among the operators sharing the spectrum, are presented; 5) The ability of a base station to use its own user equipment as sensors, for optimal spectrum sharing with base stations in other operator networks, is demonstrated to be an effective approach.

5G↗

Data-driven Optimal Control Strategy for Virtual Synchronous Generator via Deep Reinforcement Learning Approach

This paper aims at developing a data-driven optimal control strategy for virtual synchronous generator (VSG) in the scenario where no expert knowledge or requirement for system model is available. Firstly, the optimal and adaptive control problem for VSG is transformed into a reinforcement learning task. Specifically, the control variables, i.e., virtual inertia and damping factor, are defined as the actions. Meanwhile, the active power output, angular frequency and its derivative are considered as the observations. Moreover, the reward mechanism is designed based on three preset characteristic functions to quantify the control targets: (1) maintaining the deviation of angular frequency within special limits; (2) preserving well-damped oscillations for both the angular frequency and active power output; (3) obtaining slow frequency drop in the transient process. Next, to maximize the cumulative rewards, a decentralized deep policy gradient algorithm, which features model-free and faster convergence, is developed and employed to find the optimal control policy. With this effort, a data-driven adaptive VSG controller can be obtained. By using the proposed controller, the inverter-based distributed generator can adaptively adjust its control variables based on current observations to fulfill the expected targets in model-free fashion. Finally, simulation results validate the feasibility and effectiveness of the proposed approach.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Accelerating Collective Communication in Data Parallel Training across Deep Learning Frameworks

This work develops new techniques within Horovod, a generic communication library supporting data parallel training across deep learning frameworks. In particular, we improve the Horovod control plane by implementing a new coordination scheme that takes advantage of the characteristics of the typical data parallel training paradigm, namely the repeated execution of collectives on the gradients of a fixed set of tensors. Using a caching strategy, we execute Horovod’s existing coordinator-worker logic only once during a typical training run, replacing it with a more efficient decentralized orchestration strategy using the cached data and a global intersection of a bitvector for the remaining training duration. Next, we introduce a feature for end users to explicitly group collective operations, enabling finer grained control over the communication buffer sizes. To evaluate our proposed strategies, we conduct experiments on a world-class supercomputer — Summit. We compare our proposals to Horovod’s original design and observe 2x performance improvement at a scale of 6000 GPUs; we also compare them against tf.distribute and torch.DDP and achieve 12% better and comparable performance, respectively, using up to 1536 GPUs; we compare our solution against BytePS in typical HPC settings and achieve about 20% better performance on a scale of 768 GPUs. Finally, we test our strategies on a scientific application (STEMDL) using up to 27,600 GPUs (the entire Summit) and show that we achieve a near-linear scaling of 0.93 with a sustained performance of 1.54 exaflops (with standard error +- 0.02) in FP16 precision.

Romero, Joshua↗

Transient Optimization of the Cryogenic Moderator System Controller at the Spallation Neutron Source for Improved Performance

The high-energy neutron beam generated at the Spallation Neutron Source (SNS) at Oak Ridge National Laboratory is moderated to use cold (slow) neutrons for scientific discoveries. The Cryogenic Moderator System (CMS) removes heat from the neutron beam using cryogenic hydrogen (H 2 ) moderators connected via heat exchangers to a helium (He) refrigeration loop that dissipates heat using a compressor-brake system. However, the CMS is affected by sporadic losses in beam power, referred to as "beam trips," as these events generate significant disturbances in cooling requirements. To accommodate the heat load transients during beam trips, the CMS uses a decentralized control strategy consisting of four flow valves and one electric heater adjusted by independent proportional-integral (PI) controllers. During the CMS’s initial commissioning, the PI gains were calibrated based only on tracking performance, overlooking their effectiveness in disturbance rejection. A data-driven, control-oriented closed-loop model was developed to recalibrate the PI gains and minimize the transient disturbances caused by beam trips. The model consists of three main components: (1) a physics-based model of the He refrigeration loop, (2) a machine-learning model of the cryogenic H 2 cooling trains, and (3) the control logic used for feedback set-point tracking. Experimental results showed that the recalibrated gains obtained in this study improved the CMS’s transient response during beam trips.

Maldonado Puente, Bryan↗

Super-Resolution for Renewable Energy Resource Data with Wind from Reanalysis Data and Application to Ukraine

With a potentially increasing share of the electricity grid relying on wind to provide generating capacity and energy, there is an expanding global need for historically accurate, spatiotemporally continuous, high-resolution wind data. Conventional downscaling methods for generating these data based on numerical weather prediction have a high computational burden and require extensive tuning for historical accuracy. In this work, we present a novel deep learning-based spatiotemporal downscaling method using generative adversarial networks (GANs) for generating historically accurate high-resolution wind resource data from the European Centre for Medium-Range Weather Forecasting Reanalysis version 5 data (ERA5). In contrast to previous approaches, which used coarsened high-resolution data as low-resolution training data, we use true low-resolution simulation outputs. We show that by training a GAN model with ERA5 as the low-resolution input and Wind Integration National Dataset Toolkit (WTK) data as the high-resolution target, we achieved results comparable in historical accuracy and spatiotemporal variability to conventional dynamical downscaling. This GAN-based downscaling method additionally reduces computational costs over dynamical downscaling by two orders of magnitude. We applied this approach to downscale 30 km, hourly ERA5 data to 2 km, 5 min wind data for January 2000 through December 2023 at multiple hub heights over Ukraine, Moldova, and part of Romania. With WTK coverage limited to North America from 2007–2013, this is a significant spatiotemporal generalization. The geographic extent centered on Ukraine was motivated by stakeholders and energy-planning needs to rebuild the Ukrainian power grid in a decentralized manner. This 24-year data record is the first member of the super-resolution for renewable energy resource data with wind from the reanalysis data dataset (Sup3rWind).

17 WIND ENERGY↗

LC-Opt: Benchmarking Reinforcement Learning and Agentic AI for End-to-End Liquid Cooling Optimization in Data Centers

Liquid cooling is critical for thermal management in high-density data centers with the rising AI workloads. However, machine learning-based controllers are essential to unlock greater energy efficiency and reliability, promoting sustainability. We present LC-Opt, a Sustainable Liquid Cooling (LC) benchmark environment, for reinforcement learning (RL) control strategies in energy-efficient liquid cooling of high-performance computing (HPC) systems. Built on the baseline of a high-fidelity digital twin of Oak Ridge National Lab's Frontier Supercomputer cooling system, LC-Opt provides detailed Modelica-based end-to-end models spanning site-level cooling towers to data center cabinets and server blade groups. RL agents optimize critical thermal controls like liquid supply temperature, flow rate, and granular valve actuation at the IT cabinet level, as well as cooling tower (CT) setpoints through a Gymnasium interface, with dynamic changes in workloads. This environment creates a multi-objective real-time optimization challenge balancing local thermal regulation and global energy efficiency, and also supports additional components like a heat recovery unit (HRU). We benchmark centralized and decentralized multi-agent RL approaches, demonstrate policy distillation into decision and regression trees for interpretable control, and explore LLM-based methods that explain control actions in natural language through an agentic mesh architecture designed to foster user trust and simplify system management. LC-Opt democratizes access to detailed, customizable liquid cooling models, enabling the ML community, operators, and vendors to develop sustainable data center liquid cooling control solutions.

Naug, Avisek [Hewlett Packard Enterprise]↗

Stochastic Gradient-Based Distributed Bayesian Estimation in Cooperative Sensor Networks

Distributed Bayesian inference provides a full quantification of uncertainty offering numerous advantages over point estimates that autonomous sensor networks are able to exploit. However, fully-decentralized Bayesian inference often requires large communication overheads and low network latency, resources that are not typically available in practical applications. In this paper, we propose a decentralized Bayesian inference approach based on stochastic gradient Langevin dynamics, which produces full posterior distributions at each of the nodes with significantly lower communication overhead. We provide analytical results on convergence of the proposed distributed algorithm to the centralized posterior, under typical network constraints. Finally, we also provide extensive simulation results to demonstrate the validity of the proposed approach.

42 ENGINEERING↗

Scaling Building Energy Audits through Machine Learning Methods on Novel Drone Image Data

Building energy audits are time-consuming and labor-intensive. This paper describes a new method using machine learning (ML) techniques on novel data sources (drone images) to improve the identification of building characteristics and retrofit opportunities, and thereby reduce the effort for audits. The new ML method includes: (1) Building footprint extraction using line extraction, polygonization, and polygon-merging, (2) Building envelope extraction using PIX4d modeling software to reconstruct a building 3D model, (3) Visualization tool for viewing images from the 3D model, (4) Window-to-wall ratio (WWR) using state-of-art deep neural network semantic segmentation, (5) Envelope thermal anomaly detection using an unsupervised machine learning clustering algorithm, and (6) Rooftop energy equipment detection based on an object detection algorithm. The testing of this method involved a comparison of additional ML-generated information overlaid on current ‘state-of-practice’ audit and remote assessment baselines using evaluation metrics: labor time and associated cost, marginal benefits of using ML-generated information in workflows for audits and remote assessments, integration potential with existing processes and tools, and replicability/scalability of the method. In two test buildings in California that had comprehensive drawings and meter data available, the ML method effectively generated a building footprint, envelope, rooftop equipment, WWR, and locations of envelope thermal anomalies. Projected target segments of the ML method are sites with minimal drawings and energy data, and underserved sectors such as multistoried housing, disadvantaged communities, and schools for which the ML method can enable identification of building asset characteristics and prioritization of envelope retrofits and decentralized energy equipment retrofits.

Singh, Reshma↗

Deep Multi-Agent Reinforcement Learning for Real-World Signalized Traffic Corridor Control

Signalized traffic control problem has been addressed recently with deep Reinforcement Learning (RL) approaches involving diverse state, action, and reward structures. While significant progress has been noted in the literature, open challenges still remain in the areas of adaptive signal phase timing, coordination in a multi-intersection corridor setting, and consideration of real-world traffic conditions. In the context of deep RL-based problem framing, extensions are needed that enable adaptive signal phase timings in an intersection agent's action space, computationally efficient information sharing among neighboring signalized intersection agents along a corridor, and experimentation in realistic simulation environments. In this paper, we develop a deep Advantage Actor Critic (A2C) multi-agent RL (MARL) approach capturing the research extensions above and apply it within a real-world calibrated Aimsun Next traffic corridor simulation model based on traffic data from the City of Coral Gables, Florida. For a multi-intersection corridor control setting, our numerical simulation experiments with a decentralized A2C MARL algorithm applied at different time periods led to a total average corridor travel delay reduction (expressed in seconds/mile averaged over vehicles) from 4.9% to 19.9% compared to state-of-the-art actuated control.

Shuvo, Salman S. [BATTELLE (PACIFIC NW LAB)]↗

Utility-Scale Operational Consequences for Solar Grid Services

This report delves into the critical aspects of grid services provided by solar inverter-based resources (IBRs), with an emphasis on the evolving landscape of microgrids, virtual power plants (VPPs), aggregators, and distributed energy resource management systems (DERMS). As the energy sector undergoes a transformative shift towards more decentralized and resilient grid architectures, understanding the multifaceted risks associated with these technologies becomes paramount. The report categorizes these risks into organizational, technical, and procedural domains, providing a thorough risk assessment framework that stakeholders can utilize to anticipate and mitigate potential issues. In addressing the increasing complexity of grid interconnections, the report highlights the importance of Cyber-Informed Engineering (CIE). By embedding engineering controls and cybersecurity measures into the early stages of system design, this approach aims to fortify grid infrastructure against emerging cyber threats. The analysis includes an exploration of best practices and strategies for integrating CIE principles to enhance grid security and resilience. To provide practical insights, the report conducts a detailed consequence analysis of various grid services and cyber mitigations that can be applied through the interconnection process. This analysis evaluates the potential impacts of different failure modes and vulnerabilities, offering a clear understanding of the consequences that could arise from disruptions within the energy grid. The findings are further enriched by a series of case studies that illustrate real-world scenarios and lessons learned from past incidents. Through this comprehensive examination of grid services and their criticality, the report aims to prepare industry professionals with the knowledge and tools necessary to navigate the complexities of modern energy systems. By providing a comprehensive approach that includes risk assessment, cybersecurity, and consequence analysis, solar stakeholders can more effectively guarantee the reliability, efficiency, and security of the energy grid.

14 SOLAR ENERGY↗