Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Deep deterministic policy gradient (DDPG)”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Automated Control of Transactive HVACs in Energy Distribution Systems

Heating, Ventilation, and Air Conditioning (HVAC) systems contribute significantly to a building’s energy consumption. In the recent years, there is an increased interest in developing transactive approaches which could enable automated and flexible scheduling of HVAC systems based on the customer demand and the electricity prices decided by the suppliers. Flexible and automated scheduling of the HVAC systems make it a prime source for participation in residential demand response or transactive energy systems. Therefore, it is of significant interest to identify an optimal strategy to control the HVAC systems. Here, reducing the energy cost while keeping the comfort level acceptable to the users, we argue that such a control strategy should consider both the energy cost and user comfort simultaneously. Accordingly, we develop the control strategy through the solution of an optimization problem that balances between the energy cost and consumer’s dissatisfaction. This optimization enables us to solve a decision-making problem through first price prediction and then choosing HVAC temperature settings throughout the day based on the predicted price, history of the price and HVAC settings, and outside temperature. More specifically, we formulate the control design as a Markov decision process (MDP) using deep neural networks and use Deep Deterministic Policy Gradients (DDPG)-based deep reinforcement learning algorithm to find the optimal control strategy for HVAC systems that balances between electricity cost and user comfort.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Intelligent multi-zone residential HVAC control strategy based on deep reinforcement learning

Residential heating, ventilation, and air conditioning (HVAC) has been considered as an important demand response resource. However, the optimization of residential HVAC control is no trivial task due to the complexity of the thermal dynamic models of buildings and uncertainty associated with both occupant-driven heat loads and weather forecasts. In this paper, we apply a novel model-free deep reinforcement learning (RL) method, known as the deep deterministic policy gradient (DDPG), to generate an optimal control strategy for a multi-zone residential HVAC system with the goal of minimizing energy consumption cost while maintaining the users’ comfort. Here, the applied deep RL-based method learns through continuous interaction with a simulated building environment and without referring to any prior model knowledge. Simulation results show that compared with the state-of-art deep Q network (DQN), the DDPG-based HVAC control strategy can reduce the energy consumption cost by 15% and reduce the comfort violation by 79%; and when compared with a rule-based HVAC control strategy, the comfort violation can be reduced by 98%. In addition, experiments with different building models and retail price models demonstrate that the well-trained DDPG-based HVAC control strategy has high generalization and adaptability to unseen environments, which indicates its practicability for real-world implementation.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Deep reinforcement learning assisted co-optimization of Volt-VAR grid service in distribution networks

With the increasing penetration of distributed energy resources in distribution networks, Volt-VAR control and optimization (VVC/VVO) have become very important to ensure an acceptable quality of service to all customers. System operators can rely on slow-responding utility devices, including capacitor banks and on-load tap changing transformers, along with fast-responding battery and photovoltaic (PV) inverters for the VVC/VVO implementation. Because of variations in response time of these two classes of devices, and different control actions (discrete versus continuous), coordinated and optimal scheduling and operation have become of utmost importance. Here, this paper develops a look-ahead deep reinforcement learning (DRL)-based multi-objective VVO technique to improve the voltage profile of active distribution networks, decrease network and inverter power loss, and save the operational cost of the grid. It proposes a deep deterministic policy gradient (DDPG)-based approach to schedule the optimal reactive and/or active power set-points of fast-responding inverters, and a deep Q-network (DQN)-based DRL agent to schedule the discrete decisions variables of slow-responding assets. The reactive power output of PV and battery smart inverters are scheduled at 30-minute intervals and the capacitors’ commitment status is scheduled with several hour intervals. The proposed framework is validated on the modified IEEE 34-bus and 123-bus test cases with embedded PV and PV-plus-storage. To validate the efficacy of the proposed VVO, it is compared with several scenarios, including the base case without VVO, localized droop control of DERs, DDPG-only, and twin delayed DDPG (TD3) agent-based DRL techniques. The results justify the superior performance of the proposed method to improve the voltage profile, reduce network power loss, and minimize the look-ahead grid operational cost while minimizing the undesirable power losses in inverters as a result of power factor adjustments.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Approximating Nash Equilibrium in Day-ahead Electricity Market Bidding with Multi-agent Deep Reinforcement Learning

In this paper, a day-ahead electricity market bidding problem with multiple strategic generation company (GEN-CO) bidders is studied. The problem is formulated as a Markov game model, where GENCO bidders interact with each other todevelop their optimal day-ahead bidding strategies. Considering unobservable information in the problem, a model-free and data-driven approach, known as multi-agent deep deterministic policy gradient (MADDPG), is applied for approximating the Nash equilibrium (NE) in the above Markov game. The MADDPG algorithm has the advantage of generalization due to the automatic feature extraction ability of the deep neural networks. The algorithm is tested on an IEEE 30-bus system with three competitive GENCO bidders in both an uncongested caseand a congested case. Comparisons with a truthful bidding strategy and state-of-the-art deep reinforcement learning methods including deep Q network and deep deterministic policy gradient (DDPG) demonstrate that the applied MADDPG algorithm can find a superior bidding strategy for all the market participants with increased profit gains. In addition, the comparison with a conventional model-based method shows that the MADDPG algorithm has higher computational efficiency, which is feasible for real-world applications.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Data-Driven Distribution System Coordinated PV Inverter Control Using Deep Reinforcement Learning

The deployment of distributed solar photovoltaic (PV) systems has increased consistently over the past decades. High penetrations of PVs could cause a series of adverse grid impacts, such as voltage violations. The recent development of smart inverter technologies rises the incentives of developing PV control solutions that regulate the inverter output power and seeking the optimization on system operational objectives. This paper proposes a data-driven control solution based on deep reinforcement learning (DRL) to optimize PV inverters for voltage regulation. The proposed solution can minimize PV real power curtailment while maintaining network voltage at an acceptable range. Comparison results between the proposed DRL control algorithms with deep deterministic policy gradient (DDPG) and volt-var control on a real feeder in west Colorado highlight the advantage of the proposed framework in controlling the system voltage while minimizing the PV real power curtailment.

deep reinforcement learning↗

A Deep Reinforcement Learning-based Reserve Optimization in Active Distribution Systems for Tertiary Frequency Regulation

Federal Energy Regulatory Commission (FERC)Orders 841 and 2222 have recommended that distributed energy resources (DERs) should participate in energy and reserve markets; therefore, a mechanism needs to be developed to facilitate DERs’ participation at the distribution level. Although the available reserve from a single distribution system may not be sufficient for tertiary frequency regulation, stacked and coordinated contributions from several distribution systems can enable them participate in tertiary frequency regulation at scale. This paper proposes a deep reinforcement learning (DRL)-based approach for optimization of requested aggregated reserves by system operators among the clusters of DERs. The co-optimization of cost of reserve, distribution network loss, and voltage regulation of the feeders are considered while optimizing the reserves among participating DERs. The proposed framework adopts deep deterministic policy gradient (DDPG), which is an algorithm based on an actor-critic method. The effectiveness of the proposed method for allocating reserves among DERs is demonstrated through case studies on a modified IEEE 34-node distribution system.

deep reinforcement learning, distributed energy re↗

Designing reinforcement learning algorithms for building HVAC control: From experimental observation to simulation comparisons

Advanced supervisory-level control with reinforcement learning (RL) is regarded as a promising solution for HVAC systems to minimize energy consumption while maintaining thermal comfort and indoor air quality. However, most RL applications were conducted in the simulation environment rather than real-world HVAC systems. This paper developed a value-based RL controller termed Deep Q-Network (DQN) for a typical central HVAC system and evaluated its performance in a building test facility. By comparing DQN with a rule-based controller, the study not only demonstrated the cases where DQN could properly maintain indoor comfort but also discussed possible reasons why DQN failed in some other situations. Recognizing the limitations of value-based RL algorithms from the experimental tests, a simulation study was conducted to compare DQN with an alternative RL approach, an actor–critic algorithm termed Deep Deterministic Policy Gradient (DDPG). In scenarios with a relatively large action space, DDPG outperformed DQN by requiring fewer computational resources and achieving better thermal comfort, lower energy consumption, and more stable control actions. The findings suggest that the ability of DDPG to handle continuous control variables more effectively allows for faster convergence in training and more precise control in practice, which enhances the overall efficiency and reliability of the HVAC system.

Guo, Fangzhou↗

Volt-VAR Optimization in Distribution Networks Using Twin Delayed Deep Reinforcement Learning

Modern distribution grids are undergoing new challenges due to the stochastic nature of distributed energy resources (DERs). High penetration of DERs has a significant impact on Volt-VAR profile and system power losses. This work proposes a deep reinforcement learning (DRL)-based Volt-VAR optimization approach for improving voltage profile and reducing system power loss under high penetration of distributed energy resources, such as battery energy storage and solar photovoltaic units in distribution grids. The twin delayed deep deterministic policy gradient (TD3) method-based DRL agent is proposed to configure optimal set-points of reactive power outputs of fast responding smart inverters. The agent schedules the reactive power of inverters according to their physical capabilities, such as minimum allowed power factor, e.g., 0.9 leading/lagging. The reward function of the proposed DRL scheme is designed carefully to ensure a proper voltage profile of the grids with effective scheduling of reactive power outputs from inverters. The performance of the proposed model is verified on modified IEEE 34- and 123-bus systems and compared with base case with no reactive supply by inverters, and local droop Volt-VAR control approach. The results show that the proposed method performs better than the local droop control and deep deterministic policy gradient (DDPG)-based DRL method for reducing voltage fluctuation and minimizing power loss.

Hossain, Rakib↗

Optimizing and Extending the Functionality of EXARL for Scalable Reinforcement Learning [Slides]

The main goal of the Co-Design Summer School 2021 is to provide algorithmic improvements to EXARL framework by improving performance and by adding functionalities. This presentation includes an introduction to reinforcement learning and to EXARL. The researchers expanded the capability of EXARL by including additional agents like (Asynchronized) Advantage Actor Critic (A2C/A3C) and Twin Delayed Deep Deterministic Policy Gradient (TD3). They also explored algorithmic improvements such as v-trace and Prioritized Experience Replay. They found that A2C/A3C performed best with v-trace and outperformed Deep Q-Network (DQN) on both the CartPole game and the ExaBooster scientific environment. Additionally, they found that TD3 performed as good as the existing Deep Deterministic Policy Gradient (DDPG) agent and that adding Prioritized Experience Replay to DDPG accelerated convergence.

97 MATHEMATICS AND COMPUTING↗

Hierarchical Flexibility Offering Strategy for Integrated Hybrid Resources in Real-time Energy Markets

This paper proposes a hierarchical model for determining the energy flexibility offering strategy of integrated hybrid resources (IHRs) in power distribution systems to participate in real-time energy markets. The proposed model utilizes the scalability, fast response time, and uncertainty observation of deep reinforcement learning (DRL) to overcome the scalability issue of operating numerous flexible resources and deliverability of energy flexibility to the real-time markets in the presence of the network constraints. To that end, the power distribution system is divided into multiple IHRs, where different types of flexible loads, energy storage systems, and solar plants with controllable inverters are operated through local IHR controllers, trained by deep deterministic policy gradient (DDPG) algorithm. Active power request and reactive power capacity of IHRs are then transmitted to a central flexibility controller, where a quadratic optimization model ensures the deliverability of the energy flexibility to the real-time energy market by satisfying the distribution network constraints. The proposed model is implemented on the 123-bus test power distribution system, demonstrating the capability of DRL-based hierarchical model for scalable operation of IHRs in order to offer deliverable energy flexibility to the real-time energy market.

Majidi, Majid↗

Multi-task deep reinforcement learning for intelligent multi-zone residential HVAC control

In this short communication, a data-driven deep reinforcement learning (deep RL) method is applied to minimize HVAC users’ energy consumption costs while maintaining users’ comfort. The applied deep RL method's efficiency is enhanced by conducting multi-task learning that can achieve an economic control strategy for a multi-zone residential HVAC system in both cooling and heating scenarios. The applied multi-task deep RL method is compared with a rule-based benchmark case and a single-task deep deterministic policy gradient algorithm to verify its effective and generalized application in optimizing HVAC operation.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗