Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “actor critic”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Graph Partitioning and Sparse Matrix Ordering using Reinforcement Learning and Graph Neural Networks

We present a novel method for graph partitioning, based on reinforcement learning and graph convolutional neural networks. Our approach is to recursively partition coarser representations of a given graph. The neural network is implemented using SAGE graph convolution layers, and trained using an advantage actor critic (A2C) agent. We present two variants, one for finding an edge separator that minimizes the normalized cut or quotient cut, and one that finds a small vertex separator. The vertex separators are then used to construct a nested dissection ordering to permute a sparse matrix so that its triangular factorization will incur less fill-in. The partitioning quality is compared with partitions obtained using METIS and SCOTCH, and the nested dissection ordering is evaluated in the sparse solver SuperLU. Our results show that the proposed method achieves similar partitioning quality as METIS and SCOTCH. Furthermore, the method generalizes across different classes of graphs, and works well on a variety of graphs from the SuiteSparse sparse matrix collection.

97 MATHEMATICS AND COMPUTING↗

Interpreting Primal-Dual Algorithms for Constrained Multiagent Reinforcement Learning

Constrained multiagent reinforcement learning (C-MARL) is gaining importance as MARL algorithms find new applications in real-world systems ranging from energy systems to drone swarms. Most C-MARL algorithms use a primal-dual approach to enforce constraints through a penalty function added to the reward. In this paper, we study the structural effects of this penalty term on the MARL problem. First, we show that the standard practice of using the constraint function as the penalty leads to a weak notion of safety. However, by making simple modifications to the penalty term, we can enforce meaningful probabilistic (chance and conditional value at risk) constraints. Second, we quantify the effect of the penalty term on the value function, uncovering an improved value estimation procedure. We use these insights to propose a constrained multiagent advantage actor critic (C-MAA2C) algorithm. Simulations in a simple constrained multiagent environment affirm that our reinterpretation of the primal-dual method in terms of probabilistic constraints is effective, and that our proposed value estimate accelerates convergence to a safe joint policy.

chance constraints↗

Third-integer Resonant Extraction Regulation System for Mu2e

A third-integer resonant slow extraction system is being developed for Fermilab's Delivery Ring to deliver protons to the upcoming Mu2e experiment. The timescale of the extraction (or spill) duration is 43 milliseconds, which is extremely short and unprecedented. Additionally, the experiment's strict and challenging requirements on the quality of the spill at this time scale has led to the development of a new Spill Regulation System (SRS) design. The SRS primarily consists of three components - slow regulation, fast regulation, and harmonic content suppressor. Contributions to the first two components of the SRS, i.e., Slow Regulation and Fast Regulation subsystems, will be presented in which new adaptive learning algorithm schemes for the slow regulation of the spill -- validated using particle tracking simulations -- shall be described. In addition to these novel methods for the enhancement of the spill regulation system, results of employing Machine Learning in enhancing the performance of the resonant extraction are also presented. At the forefront of applying ML techniques to solve non-linear accelerator control problems, this work includes optimizing the PID gains as well as the replacement of the traditional PID controller using Recurrent Neural Networks and Gated Recurrent Unit (GRU) ML models to achieve efficiencies greater than a PID controller. Cutting-edge on-going Reinforcement Learning efforts, including an actor-critic family of learning algorithms, to regulate the spill rate will be reviewed, as well as present analytical calculations pertaining the transit time of particles in a third-integer resonant extraction. Detailed numerical investigations and validations of such calculations, the model of which could be exported and reliably used in future analytical modeling of any resonant extraction, are discussed.

43 PARTICLE ACCELERATORS↗

Topology-Aware Reinforcement Learning for Voltage Control: Centralized and Decentralized Strategies

Volt-VAR control (VVC) methods based on deep reinforcement learning (DRL) can effectively control distribution grid voltage and minimize power loss by implementing corrective and preventive control measures on the reactive power output of inverter-based distributed energy resources (DERs). However, model-free DRL-based VVC approaches usually cannot capture the important topological feature of the power system since they use a fully-connected network (FCN) to deliver the action. Therefore, this paper proposes a graph convolutional network (GCN)-based DRL approach that can employ the topological information of the network to take better control action for regulating the voltage. Our implementation allows for both centralized and decentralized configurations, utilizing a single agent and multiple agents respectively. Although the centralized GCN-based DRL approach has its advantages of minimizing voltage fluctuation and power loss, it is not suitable for large scale power systems due to its challenges in terms of scalability, computation speed and potential single points of failure. Therefore, these problems can be resolved using the decentralized GCN-based DRL approach. Moreover, to ensure the safe operation of the model, our proposed approach incorporates an exponential barrier function while formulating the reward function for each agent. To validate performance of the proposed approaches, the proposed model is tested on modified IEEE test systems and the performances are measured in terms on voltage fluctuation reduction, minimization of power loss and computational speed. Finally, the results show that the proposed topology-aware approach outperforms the FCN-based DRL approach in terms of reducing voltage fluctuation and minimizing power loss of the network. Moreover, it is shown that the decentralized GCN-based DRL has faster computational speed than other approaches.

42 ENGINEERING↗

Non-Stationary Policy Learning for Multi-Timescale Multi-Agent Reinforcement Learning

In multi-timescale multi-agent reinforcement learning (MARL), agents interact across different timescales. In general, policies for time-dependent behaviors, such as those induced by multiple timescales, are non-stationary. Learning non-stationary policies is challenging and typically requires sophisticated or inefficient algorithms. Motivated by the prevalence of this control problem in real-world complex systems, we introduce a simple framework for learning non-stationary policies for multi-timescale MARL. Our approach uses available information about agent timescales to define and learn periodic multi-agent policies. In detail, we theoretically demonstrate that the effects of non-stationarity introduced by multiple timescales can be learned by a periodic multi-agent policy. To learn such policies, we propose a policy gradient algorithm that parameterizes the actor and critic with phase-functioned neural networks, which provide an inductive bias for periodicity. The framework's ability to effectively learn multi-timescale policies is validated on a gridworld and building energy management environment.

control↗

Optimizing Non-Terrestrial Hybrid RF/FSO Links With Reinforcement Learning: Navigating Through Clouds

In the pursuit of ubiquitous broadband connectivity, there has been a significant shift towards the vertical expansion of communication networks into space, particularly through the exploitation of low Earth orbit (LEO) satellite constellations, which are favored for their relatively low latency. However, this approach faces many challenges that need to be addressed, including atmospheric turbulence, high path loss, and dynamic cloud formations. High-altitude pseudo-satellites (HAPS) have emerged as promising relaying layers between LEO satellites and ground stations, enhancing coverage, latency, and direct terrestrial user connectivity. While radio frequency (RF) bands suffer from congestion and limited bandwidth, free space optical (FSO) communications offer higher data rates, but are susceptible to misalignment and weather-induced signal degradation. To address these challenges, a hybrid RF/FSO approach has been proposed to take advantage of both technologies by dynamic switching between RF and FSO based on propagation channel conditions. This paper introduces a reinforcement learning-based algorithm designed to optimize the trajectory of HAPS, maneuver around cloudy areas, and seamlessly switch between the RF and FSO communication modes to maximize the achievable capacity. The proposed approach aims to maximize system performance by intelligently adapting to environmental conditions and offering a promising solution for next-generation space communication networks.

actor-critic algorithm↗

Two-Stage Deep Reinforcement Learning for Distribution System Voltage Regulation and Peak Load Management

The growing integration of distributed solar photovoltaic (PV) in distribution systems could result in adverse effects during grid operation. This paper develops a two-agent soft actor critic-based deep reinforcement learning (SAC-DRL) solution to simultaneously control PV inverters and battery energy storage systems for voltage regulation and peak demand reduction. The novel two-stage framework, featured with two different control agents, is applied for daytime and nighttime operations to enhance control performance. Comparison results with other control methods on a real feeder in Western Colorado demonstrate that the proposed method can provide advanced voltage regulation with modest active power curtailment and reduce peak load demand from feeder's head.

deep reinforcement learning↗

Two-Stage Deep Reinforcement Learning for Distribution System Voltage Regulation and Peak Load Management: Preprint

The growing integration of distributed solar photovoltaic (PV) in distribution systems could result in adverse effects during grid operation. This paper develops a soft actor critic-based deep reinforcement learning (SAC-DRL) solution to simultaneously control PV inverters and battery energy storage systems for voltage regulation and peak load demand shaving. The novel two-stage framework, featured with two different control agents, is applied for daytime and nighttime operation to enhance the control performance. Comparison results with other control methods on a real feeder in Western Colorado demonstrate that the proposed method can provide advanced voltage regulation with modest active power curtailment for peak demand reduction.

deep reinforcement learning↗

Feature Engineering and Ensemble Methods for Imbalanced ICS Intrusion Detection: Pipeline Audit and Constrained Evaluation

Industries are becoming increasingly connected and are more vulnerable to cyberattacks due to the widened attack surface. Industrial Control Systems (ICS) are among the most critical sectors that malicious actors can target, as such attacks can cause significant operational disruption and physical damage. It is imperative to detect such attacks as early as possible. This paper evaluates constraint-conditioned optimistic performance estimates for traditional ML models in ICS intrusion detection (i.e., estimates obtained under contiguous, non-shuffled temporal evaluation without test-set alteration, but with pre-split feature engineering that may introduce temporal leakage, due to dataset constraints). Our findings are threefold. First, we quantify how iterative feature engineering affects tree-based ensemble performance and examine how pipeline decisions (split strategy, sampling scope, and cleaning policy) can inflate or reduce reported IDS results under constraint-bound evaluation. Second, we compare intrinsic class-imbalance handling across ensemble models. Third, under our current pipeline constraints (including pre-split feature engineering), CatBoost achieves the best performance on Water Storage Tank (accuracy: 0.9831, class-1 F1: 0.9682), while Light- GBM achieves the best performance on Gas Pipeline (accuracy: 0.9618, class-1 F1: 0.9086).

97 MATHEMATICS AND COMPUTING↗

Living-off-the-land Techniques Unlikely to Supplant Energy Sector-Focused OT-Specific Malware

Despite increased reports of energy sector-focused threat actors using living-off-the-land (LOTL) techniques, it is unlikely LOTL techniques will wholly supplant malware in energy sector operational technology (OT)-focused cyber operations. Threat actors leverage LOTL techniques to access energy sector networks, abstracting process information and maintaining persistence. Although threat actors using LOTL techniques have successfully interrupted energy sector industrial control environments, designed features of OT-specific malware likely increase the cyber-physical impact of an attack and delay recovery of critical functions and services. Malicious actors will very likely continue to use LOTL techniques for stealth, while designing malware to bolster final impacts on cyber-physical systems in energy sector OT environments.

99 GENERAL AND MISCELLANEOUS↗

Attack Surface of Wind Energy Technologies in the United States [Slides]

This slide deck presents an overview of the threat landscape for wind energy technologies. It highlights unique cybersecurity considerations for wind, the growing penetration and potential impact of an attack. Standard architectures are shared to highlight where vulnerabilities may exist and attack paths to reach critical infrastructure. We discuss threat actors and attack paths. Several recent events impacting wind assets or wind companies are explained.

17 WIND ENERGY↗

Cooperative and Non-Cooperative UAS Detection

This paper describes a system that enhances airspace situational awareness by detecting and identifying Unmanned Aerial Systems (UAS). This multi-domain solution tracks both cooperative scientific flights as well as non-cooperative intrusions from "bad actors." The system supports a critical push towards safety within NASA’s advanced air mobility mission. After surveying the existing technologies at Langley Research Center, the radar and visual systems were chosen for the primary and secondary detection mechanisms, respectively. These systems were tuned an upgraded to become more sensitive to UAS activity. Additionally, a Remote Identification receiver was procured and integrated into the flight surveillance system.

Lucas Barduson↗

Attack Surface of Renewable Energy Technologies

This slide deck presents an overview of the threat landscape for renewable energy technologies. It highlights he growing penetration of renewable resources and potential impact of an attack. Standard architectures are shared to highlight where vulnerabilities may exist and attack paths to reach critical infrastructure. We discuss threat actors and attack paths. Several recent events impacting renewable energy assets or companies are explained.

14 SOLAR ENERGY↗

A Proposed Evaluation Framework for New and Emerging Low Embodied-Carbon Concrete Technologies

New opportunities for carbon reductions in buildings create a strong need for a common framework and method for those who design, build and influence construction to evaluate lifecycle carbon reductions from design decisions and technology choices. These opportunities include a wide range of low-embodied-carbon concrete materials being rapidly developed and introduced to the market. How to evaluate these newer materials and technologies has become critical for both public- and private-sector actors seeking to decarbonize building constructions by leveraging the Infrastructure Investment and Jobs Act (IIJA) and Inflation Reduction Act (IRA) funds. We propose an evaluation framework to assess the lifecycle carbon reductions from adoption of these technologies, including a subset of key “must have” (1) technical criteria (embodied carbon level, technology development stage); (2) market criteria (market size, scalability); and (3) financial criteria (cost of technology implementation compared to businessas-usual) from a range of options. We discuss how to use the framework and illustrate it using a “heatmap,” rating score and short case study of a promising technology. We also propose a plan to implement this framework that includes (1) standardized measurement and validation methods for verifying emission reductions from these technologies, and (2) avenues to implement real world demonstrations. We conclude with recommendations for next steps on framework refinement and commercialization strategy development.

Singh, Reshma↗

Water Security: Trends, Capabilities, and Research Directions to Secure Water Infrastructure

Water and wastewater sector is target rich and resource poor ~153k water utilities, serve 80% of US population ~16k publicly-owned wastewater systems in the US serve 75% of the population Need scalable solutions to fit small and medium to large systems Federal attention to critical infrastructure continues to grow – particularly in the water sector Increase in water sector incidents and threats for large scale disruption – particularly by nation-state actors and their proxies EPA is the SRMA DHS CISA focuses on critical infrastructure protection across sectors They must work together to secure WWW systems Research capabilities to enable secure water systems Current and future threats Resilience – natural disasters, accidents, cyber-physical attacks

99 - GENERAL AND MISCELLANEOUS↗

Domestic Extremism: Countering the Threat Posed to Critical Assets

Domestic extremism has been a growing concern in the United States in recent months, as illustrated in multiple bulletins from the Department of Homeland Security (DHS) warning law enforcement partners of the heightened threat. As concerns about these actors grows, it is important that facilities in the U.S. and internationally that protect critical assets, such as sensitive information, hazardous materials, or critical infrastructure, have effective methods in place to secure those assets. DE has challenged security systems through the threat of insider attack and violence, creating a new threat to be countered in the Office of Radiological Security’s radiological source security mission. In this effort, we used a literature review and focus group discussions with experts in critical asset security and extremism to understand the nature of the domestic extremist threat, to identify best practices in securing assets, recognize potential gaps in security measures to be corrected, and recommend actions and next steps. Twenty-two subject matter experts participated in a series of five focus group sessions. Questions focused on definitions of domestic extremism, potential changes in the threat, best practices in securing facilities, assets, and personnel, and any perceived gaps. Upon completion of the focus groups, notes were analyzed thematically to identify any recurring patterns in the results. In addition, a review of academic, industry, and government literature was conducted to understand the threat, describe the process of radicalization to extremism, and to identify empirically informed practices in prevention and response. Results of this project demonstrated that further work is needed to define domestic extremism in law, regulation, and policy, to help the U.S. develop a consistent response to the threat within organizations. This is especially important, as SMEs emphasized the need for early intervention in prevention efforts, noting that organizations need clear guidance on when and how to intervene. In addition, the need for social media monitoring was discussed, although challenges remain to do so with appropriate respect for privacy and civil liberties concerns.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗