Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “reinforcement Learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Pareto Optimization of Analog Circuits Using Reinforcement Learning

Analog circuit optimization and design presents a unique set of challenges in the IC design process. Many applications require the designer to optimize for multiple competing objectives, which poses a crucial challenge. Motivated by these practical aspects, we propose a novel method to tackle multi-objective optimization for analog circuit design in continuous action spaces. In particular, we propose to (i) extrapolate current techniques in Multi-Objective Reinforcement Learning to continuous state and action spaces and (ii) provide for a dynamically tunable trained model to query user defined preferences in multi-objective optimization in the analog circuit design context.

99 GENERAL AND MISCELLANEOUS↗

CyRRL (Cyber Resilient Reinforcement Learning for grid voltage control) [SWR-24-115]

This codebase contains a multi-agent, actor-critic reinforcement learning implementation for cyber-resilient grid voltage control. It uses a 123-bus OpenDSS system as the environment, with three-phase power flow translating nodal power injections into solved nodal voltages. The reward function penalizes deviations from nominal voltage as well as reactive power dispatch, while encouraging agents to take actions that result in fast convergence to nominal conditions. The codebase models false data injection attacks and includes functionality for training, testing, hyper-parameter tuning, and visualization.

Murphy, Sinnott [National Renewable Energy Laborat↗

Deep Reinforcement Learning for Automatic Generation Control of Wind Farms

This paper provides a model-free framework for real-time control of wind farms to accurately track a power reference signal. This problem requires tractable dynamical models for capturing the aerodynamic interaction between wind turbines and controllers that can make decisions in realtime given varying atmospheric conditions. In this paper, we propose a deep reinforcement learning framework to provide real-time yaw control of a wind farm. Modifications have been made to FLOw Redirection and Induction in Steady State (FLORIS), a modeling tool that incorporates transient wake behavior. The control problem is formulated to track a synthetic power reference signal based on historical atmospheric (wind speed and direction) information, price signals, and regulation deployment data from U.S. regional transmission operators. Results indicate that a wind farm, with this control paradigm, can achieve good tracking performance when tested with real atmospheric data.

49 EE - Wind and Water Power Program - Wind (EE-4W↗

An Online Reinforcement Learning Controller Design For Mars Ascent Vehicle

This paper presents a neural network (NN) approximator-based online reinforcement learning (ORL) controller design for Mars Ascent Vehicle (MAV) under parametric variation and significant external disturbances. The ORL controller, which does not require any offline training, involves two NNs where an action NN produces optimal short-term control performance while a critic NN evaluates the performance of the action NN using an approximated cost function. The simulation example with comparisons against baseline Proportional-Integral-Derivative (PID) and gain scheduled pole-placement PID (GS-PP-PID) controllers show the proposed controller’s effectiveness and robustness under parametric variation and high external disturbances.

Han Woong Bae↗

Dynamic Role-Based Access Control Policy for Smart Grid Applications: An Offline Deep Reinforcement Learning Approach

Role-based access control (RBAC) is adopted in the information and communication technology domain for authentication purposes. However, due to a very large number of entities within organizational access control (AC) systems, static RBAC management can be inefficient, costly, and can lead to cybersecurity threats. In this paper, a novel hybrid RBAC model is proposed, based on the principles of offline deep reinforcement learning (RL) and Bayesian belief networks. The considered framework utilizes a fully offline RL agent, which models the behavioral history of users as a Bayesian belief-based trust indicator. Thus, the initial static RBAC policy is improved in a dynamic manner through off-policy learning while guaranteeing compliance of the internal users with the security rules of the system. By deploying our implementation within the smart grid domain and specifically within a Distributed Energy Resources (DER) ecosystem, we provide an end-to-end proof of concept of our model. Finally, detailed analysis and evaluation regarding the offline training phase of the RL agent are provided, while the online deployment of the hybrid RL-based RBAC model into the DER ecosystem highlights its key operation features and salient benefits over traditional RBAC models.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Reinforcement Learning for Intentional Islanding in Resilient Power Transmission Systems

Intentional islanding is the process of identifying and deliberately decomposing the transmission network to form self-sustained islands from an endangered network during disruptions to improve resilience and security. Most existing intentional islanding models are offline resilience decision tools and hence do not provide outage responses in a timely manner. In this paper, a reinforcement learning (RL) based model for intentional islanding is developed, which offers real-time switching control, online deployability, and adaptability to varying system conditions. The intentional islanding process is formulated as a Markov decision process, where the optimal transmission switching policy is learned using the RL approach. The control policy is learned over an environment that encompasses a Power System Simulator for Engineering (PSS/E) model of the transmission network, facilitated by an interface to the standard openAI Gym framework. The proposed RL-based methodology aims to form stable and self-sustainable islands by ensuring voltage stability while reducing the power mismatch in the formed islands. A proximal policy optimization algorithm is designed, which is suitable for controlling the on/off status of the switches with multi-layer perceptron as value and actor networks. The effectiveness of the proposed framework in the self-recovery of the grid by island formation is applied on the modified IEEE 39-bus test network and validated by dynamic simulations.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Adaptive Deep Reinforcement Learning Algorithm for Distribution System Cyber Attack Defense With High Penetration of DERs

With grid modernization, smart inverters are increasingly used to execute advanced controls for distribution network reliability. However, this also increases the cyber-attack space. Here this paper focuses on the defense approaches to restore the system to normal operation circumstances in the presence of cyber-attacks. A unique deep reinforcement learning (DRL) method is developed to minimize voltage violations and reduce power losses for impacted feeders. The defense problem is reformulated as a Markov decision-making process to dynamically control DERs while minimizing load shedding. This is achieved via an improved soft actor-critic (SAC)-based DRL algorithm, which can govern DER set points and load-shedding scenarios in discrete and continuous modes via the auto-tune entropy and Gaussian policy features. Numerical comparison results on the modified IEEE 123-node system with other control approaches, such as Volt-VAR (VV), Volt-Watt (VW), and model predictive control (MPC) show that the proposed method can eliminate voltage violations and provide feasible control actions that perform complete mitigation of cyber-threats.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Deep Reinforcement Learning for Distribution System Restoration Using Distributed Energy Resources and Tie-Switches

Distributed energy resources (DERs), such as solar PVs and energy storage, can be used to restore distribution system critical loads after the extreme weather events to increase grid resilience. However, coordinating multiple DERs together with tie-switches for multi-step restoration process under renewable uncertainty is challenging. This paper proposes a deep reinforcement learning to control discrete actions of switching on/off tie switches and DERs for critical load restoration. The restoration problem is first cast into the Markov decision process suitable for DRL. Then, the original soft actor critic (SAC) method for continuous actions has been extended to handle discrete and continuous actions. Numerical comparison results with other stochastic optimization-based approaches on the modified IEEE 33-bus system show that the proposed method can achieve fast critical load restoration in the presence of substation power outage while maintaining system voltage limit throughout the restoration process.

active distribution systems↗

Bi-Level Adaptive Storage Expansion Strategy for Microgrids Using Deep Reinforcement Learning

Battery energy storage (BES) is a versatile resource for the secure and economic operation of microgrids (MGs). Prevailing stochastic optimization-based approaches for BES expansion planning for MGs are computationally complicated. This work proposes a data-driven bi-level multi-period BES expansion planning framework to determine the siting, sizing, and timing of BES installations. The proposed planning framework unifies deep reinforcement learning (DRL) and linear programming, thereby decoupling the determinations for the integer and continuous decision variables in two time scales, respectively. In the upper level, a rainbow DRL agent with quantile regression is trained to provide dynamic planning policies to accommodate stochastic renewable energy resources (RESs), load, and battery price changes efficiently. Further, the lower level computes the optimal operation of MGs with frequency constraints to hedge the islanding contingency. The two levels communicate with one another by exchanging storage configuration and operating expenses in order to accomplish the shared goal of minimizing investment and operation costs. Comparative case studies on an MG are carried out to demonstrate the superiority of the proposed DRL-based solution to the mixed-integer linear programming counterpart on efficiency, scalability, and adaptability.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Visibility-enhanced model-free deep reinforcement learning algorithm for voltage control in realistic distribution systems using smart inverters

Increasing integration of distributed solar photovoltaic (PV) into distribution networks could result in adverse effects on grid operation. Traditional model-based control algorithms require accurate model information that is difficult to acquire and thus are challenging to implement in practice. Here, this paper proposes a surrogate model-enabled grid visibility scheme to empower deep reinforcement learning (DRL) approach for distribution network voltage regulation using PV inverters with minimal system knowledge. In contrast to existing DRL methods, this paper presents and corroborates the adverse impact of missing load information on DRL performance and, based on this finding, proposes a surrogate model methodology to impute load information utilizing observable data. Additionally, a multi-fidelity neural network is utilized to construct the DRL training environment, chosen for its efficient data utilization and enhanced robustness to data uncertainty. The feasibility and effectiveness of the proposed algorithm are assessed by considering DRL testing across varying degrees of observable load information and diverse training environments on a realistic power system.

14 SOLAR ENERGY↗

Risk-Informed Operations and Maintenance Decision Making Using Deep Reinforcement Learning

A challenge for operating nuclear power plants is the significant cost of operations and maintenance, at times consuming up to 66% of annual operating costs. This project aims to build a framework for a risk-informed asset-management tool that integrates inspections, repairs, spare-part inventory, supply chain, and business choices to lower overall O&M costs. Our approach uses a combination of data-driven modeling and deep reinforcement learning to create and implement optimal maintenance policies for the existing nuclear fleet, as well as new advanced reactors. The creation of an asset management tool that uses these advanced methods will give operators new capabilities to help reduce the burden of O&M spending in nuclear power plants.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

A Modular and Transferable Reinforcement Learning Framework for the Fleet Rebalancing Problem

Mobility on demand (MoD) systems show great promise in realizing flexible and efficient urban transportation. However, significant technical challenges arise from operational decision making associated with MoD vehicle dispatch and fleet rebalancing. For this reason, operators tend to employ simplified algorithms that have been demonstrated to work well in a particular setting. To help bridge the gap between novel and existing methods, we propose a modular framework for fleet rebalancing based on model-free reinforcement learning (RL) that can leverage an existing dispatch method to minimize system cost. In particular, by treating dispatch as part of the environment dynamics, a centralized agent can learn to intermittently direct the dispatcher to reposition free vehicles and mitigate against fleet imbalance. We formulate RL state and action spaces as distributions over a grid partitioning of the operating area, making the framework scalable and avoiding the complexities associated with multiagent RL. Numerical experiments, using real-world trip and network data, demonstrate that RL reduces waiting time by 28% to 38% for the same-day evaluation, 17% to 44% for cross-day evaluation, and 22% to 25% for cross-season evaluation compared with no rebalancing scenarios. This approach has several distinct advantages over baseline methods including: improved system cost; high degree of adaptability to the selected dispatch method; and the ability to perform scale-invariant transfer learning between problem instances with similar vehicle and request distributions.

33 ADVANCED PROPULSION SYSTEMS↗

Data-Driven Distribution System Coordinated PV Inverter Control Using Deep Reinforcement Learning

The deployment of distributed solar photovoltaic (PV) systems has increased consistently over the past decades. High penetrations of PVs could cause a series of adverse grid impacts, such as voltage violations. The recent development of smart inverter technologies rises the incentives of developing PV control solutions that regulate the inverter output power and seeking the optimization on system operational objectives. This paper proposes a data-driven control solution based on deep reinforcement learning (DRL) to optimize PV inverters for voltage regulation. The proposed solution can minimize PV real power curtailment while maintaining network voltage at an acceptable range. Comparison results between the proposed DRL control algorithms with deep deterministic policy gradient (DDPG) and volt-var control on a real feeder in west Colorado highlight the advantage of the proposed framework in controlling the system voltage while minimizing the PV real power curtailment.

deep reinforcement learning↗

Chapter Nine - Automated Optimal Control in Energy Systems: The Reinforcement Learning Approach

With the development of smart grid technologies an increasing number of new devices and participants have joined modern energy systems and are inevitably making them more complicated and interdependent than ever. Optimally controlling such a complex energy system and maintaining its operation in a high-efficient, secure, and resilient manner are challenging tasks to the system operators. Fortunately, the revolution in deep learning and artificial intelligence (AI), both from hardware and algorithms perspectives, has provided new ideas and solutions to many previously intractable problems. As a result, this advance in computer science also sparked great research interests in utilizing AI in solving engineering problems related to the modern energy systems. Among many AI techniques, deep reinforcement learning (DRL) has demonstrated great potential for solving sequential optimization problems, which are very common in the engineering domains. Its ability to handle nonlinearity and stochasticity in controlled systems has out-competed many traditional optimal control algorithms. Therefore in this chapter, we focus on the state-of-the-art of DRL concepts and related algorithms, compare their pros and cons with traditional optimal control approaches and discuss the typical workflow for leveraging RL in solving complex problems in modern energy systems.

artificial intelligence↗

Tuning Phase Lock Loop Controller of Grid Following Inverters by Reinforcement Learning to Support Networked Microgrid Operations

The dynamic operation of networked microgrids leads to varying topological configurations and generator commitments and dispatches. These variations correspond to systems with different electrical characteristics. The fixed control gains of high-speed power electronic devices may result in undesirable system performance when the electrical characteristics change significantly. As such, it is necessary to tune the control gains of power electronics devices to adapt to the changing system characteristics. This paper uses observer-based reinforcement learning to automatically tune the proportional-integral (PI) gains of phase lock loop (PLL) controller of grid-following (GFL) inverters to adapt to the changing system strengths, that would be seen in networked microgrid operations. Simulation results using an operational electric distribution system, modeled as networked microgrids, are presented to demonstrate the need and effectiveness of the proposed adaptive controls.

networked microgrids, reinforcement learning, grid↗

Hierarchical Reinforcement Learning of a Short-Range Bond-Order Potential for Silica: Analytic Embedding of Coordination with Classical Efficiency

Reinforcement learning (RL) has recently emerged as a data-efficient strategy to parametrize short-range interatomic potentials. Building on our past RL optimization of pairwise silica models, we extend the framework to a bond-order (Tersoff-type) potential that provides an analytic embedding of local coordination through a three-body term. A hierarchical RL workflow combining continuous-action Monte Carlo Tree Search and property-based rewards efficiently explores the 26-dimensional parameter space, sequentially optimizing lattice parameters, densities, angles, and cohesive energies of 21 silica polymorphs. The resulting models, Q-Tersoff and ML-Tersoff, reproduce the energetic ordering of low-energy phases and capture the angular correlations and amorphous structure factors of silica with improved fidelity over pairwise force fields, while remaining orders of magnitude faster than high-dimensional machine-learned potentials. Both models underperform for elastic constants and high-energy frameworks, delineating the limits of the current analytic form. The approach establishes a general and interpretable route to angle-aware, short-range potentials that bridge physics-based and machine-learned descriptions of silicate materials.

36 MATERIALS SCIENCE↗

Deep Multi-Agent Reinforcement Learning for Real-World Signalized Traffic Corridor Control

Signalized traffic control problem has been addressed recently with deep Reinforcement Learning (RL) approaches involving diverse state, action, and reward structures. While significant progress has been noted in the literature, open challenges still remain in the areas of adaptive signal phase timing, coordination in a multi-intersection corridor setting, and consideration of real-world traffic conditions. In the context of deep RL-based problem framing, extensions are needed that enable adaptive signal phase timings in an intersection agent's action space, computationally efficient information sharing among neighboring signalized intersection agents along a corridor, and experimentation in realistic simulation environments. In this paper, we develop a deep Advantage Actor Critic (A2C) multi-agent RL (MARL) approach capturing the research extensions above and apply it within a real-world calibrated Aimsun Next traffic corridor simulation model based on traffic data from the City of Coral Gables, Florida. For a multi-intersection corridor control setting, our numerical simulation experiments with a decentralized A2C MARL algorithm applied at different time periods led to a total average corridor travel delay reduction (expressed in seconds/mile averaged over vehicles) from 4.9% to 19.9% compared to state-of-the-art actuated control.

Shuvo, Salman S. [BATTELLE (PACIFIC NW LAB)]↗

Reinforcement Learning for Distribution Grid Optimization (PyCIGAR) v0.1

PyCIGAR is a python software package that merges off-the-shelf reinforcement learning libraries (RLLib and Ray) with electric power distribution system simulation tools (OpenDSS and a custom power flow solver built by LBL). PyCIGAR enables the training of neural networks to optimize the behavior of different components in the electric distribution grid, such as control systems in photovoltaic rooftop solar inverters and electric battery storage systems. The software package has been used to train neural networks to update settings in photovoltaic rooftop solar inverter control systems to mitigate cyber attacks on other solar photovoltaic rooftop devices.

Arnold, Daniel↗