Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Policy gradient”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Comprehensive assessment of deep reinforcement learning approaches for economic dispatch in nuclear-driven microgrids

As the electrical grid integrates more variable renewable energy sources such as wind and solar, the demand for distributed and flexible systems to address this increased variability becomes critical. Nuclear-driven microgrids provide a promising solution by offering stable generation to complement intermittent renewables, ensuring grid reliability and operating efficiency. This paper proposes a recurrent deep reinforcement learning framework for optimal economic dispatch in a nuclear-powered microgrid integrating renewable energy sources, small modular reactors, battery storage systems, and balance-of-plant dynamics. A three-agent control architecture is developed, where demand and renewable energy agents act as forecasters, and a reinforcement learning-based dispatch agent performs real-time energy allocation. A nonlinear programming formulation is first used to generate an optimal baseline for benchmarking. The proposed dispatch controller, based on Proximal Policy Optimization enhanced with Long Short-Term Memory networks, exploits temporal correlations in system dynamics by taking advantage of the time series used as inputs to improve policy robustness under uncertainty. Comparative analysis against established deep reinforcement learning methods, including Proximal Policy Optimization with a feedforward architecture, Soft Actor-Critic, and Twin Delayed Deep Deterministic Policy Gradient, demonstrates superior performance. Numerical results indicate that the proposed controller achieves a 0.39% cost reduction relative to the nonlinear programming benchmark and outperforms other learning-based methods by generating additional revenue of up to 0.35%. All reinforcement learning controllers compute dispatch actions in less than 0.3 s, resulting in a computational speedup of more than three orders of magnitude over the nonlinear programming baseline. The findings of this paper highlight their applicability for real-time operation and control in nuclear-integrated microgrids under volatile operating conditions.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Hamiltonian switching control of noisy bipartite qubit systems

Abstract We develop a Hamiltonian switching ansatz for bipartite control that is inspired by the quantum approximate optimization algorithm, to mitigate environmental noise on qubits. We demonstrate the control for a central spin coupled to bath spins via isotropic Heisenberg interactions, and then make physical applications to the protection of quantum gates performed on superconducting transmon qubits coupling to environmental two-level-systems (TLSs) through dipole-dipole interactions, as well as on such qubits coupled to both TLSs and a Lindblad bath. The control field is classical and acts only on the system qubits. We use reinforcement learning with policy gradient to optimize the Hamiltonian switching control protocols, using a fidelity objective for specific target quantum gates. We use this approach to demonstrate effective suppression of both coherent and dissipative noise, with numerical studies achieving target gate implementations with fidelities over 0.9999 (four nines) in the majority of our test cases and showing improvement beyond this to values of 0.999 999 999 (nine nines) upon a subsequent optimization by GRadient Ascent Pulse Engineering (GRAPE). We analyze how the control depth, total evolution time, number of environmental TLS, and choice of optimization method affect the fidelity achieved by the optimal protocols and reveal some critical behaviors of bipartite control of quantum gates.

Physics↗

Federated Deep Reinforcement Learning for Decentralized VVO of BTM DERs

The future of grid control requires a hybrid approach combining centralized and decentralized methods to fully utilize the potential of smart edge devices with artificial intelligence (AI) capabilities. This paper aims to develop and evaluate a federated deep reinforcement learning (FDRL) framework for decentralized adaptive volt-var optimization (VVO) of behind-the-meter (BTM) distributed energy resources (DERs). First, this paper models a single deep reinforcement learning (DRL) agent using the Markov Decision Process (MDP) framework for decentralized adaptive VVO of BTM DERs. Two DRL algorithms, soft actor-critic (SAC) and twin-delayed deep deterministic policy gradient (TD3), are compared for their effectiveness in optimizing VVO. Results show that TD3 outperforms SAC, achieving a 71.3% improvement in mean reward. Finally, the DRL agent is deployed within the FDRL framework, using the Flower platform, to enhance learning, provide adaptive control, and ensure data privacy for BTM DERs.

Ravi, Abhijith↗

Post-Disaster Microgrid Formation for Enhanced Distribution System Resilience

This paper proposes a deep reinforcement learning (DRL) based approach for post-disaster critical load restoration in active distribution systems to form microgrids through network reconfiguration to minimize critical load curtailments. Distribution networks are represented as graph networks, and optimal network configurations with microgrids are obtained by searching for the optimal spanning forest. The constraints to the research question being explored are the radial topology and power balance. Unlike existing analytical and population-based approaches, which necessitate the repetition of entire analyses and computation for each outage scenario to find the optimal spanning forest, the proposed approach, once properly trained, can quickly determine the optimal, or near-optimal, spanning forest even when outage scenarios change. When multiple lines fail in the system, the proposed approach forms microgrids with distributed energy resources in active distribution systems to reduce critical load curtailment. The proposed DRL-based model learns the action-value function using the REINFORCE algorithm, which is a model-free reinforcement learning technique based on stochastic policy gradients. A case study was conducted on a 33-node distribution test system, demonstrating the effectiveness of the proposed approach for post-disaster critical load restoration.

active distribution systems↗

Optimal Control of SOEC-Based Hydrogen Production Systems for Demand Response Using Deep Reinforcement Learning in Smart Grids

Solid oxide electrolysis cell (SOEC) hydrogen production technology can range in size from small, appliance-size equipment to large-scale, central production facilities that can be tied directly to renewable or non-greenhouse-gas-emitting forms of electricity production, making it an ideal resource for demand response (DR). The SOEC hydrogen production system is a complex integrated system that encompasses fluid dynamics, electrical dynamics, and electrochemical and thermal dynamics, all of which involve non-linearity and non-convexity. Proper control of the SOEC hydrogen production system is crucial to enable its participation in the DR program. Here, to overcome the difficulty of designing an explicit control law for such nonlinear systems with nonconvex optimization features in DR applications, deep reinforcement learning (DRL) is explored to achieve the optimal control of the SOEC system for DR participation. Specifically, a twin delayed deterministic policy gradient (TD3) control framework is applied to achieve optimal response performance during DR events by considering power tracking error and hydrogen production efficiency with a suitable reward function. Two case studies with grid connections for tracking different DR commands were investigated. The first case study involved operating conditions reaching the boundaries, while the second involved operating conditions within the boundaries. The results showed that the proposed DRL-based control for SOEC can track the DR signal in a timely manner while maintaining high energy efficiency.

08 HYDROGEN↗

From Sim to Real: A Pipeline for Training and Deploying Traffic Smoothing Cruise Controllers

Designing and validating controllers for connected and automated vehicles to enhance traffic flow presents significant challenges, from the complexity of replicating real-world stop-and-go traffic dynamics in simulation, to the intricacies involved in transitioning from simulation to actual deployment. In this work, we present a full pipeline from data collection to controller deployment. Specifically, we collect 772 km of driving data from the I-24 in Tennessee, and use it to build a one-lane simulator, placing simulated vehicles behind real-world trajectories. Using policy-gradient methods with an asymmetric critic, we improve fuel efficiency by over 10% when simulating congested scenarios. Our comprehensive approach includes reinforcement learning for controller training, software verification, hardware validation and setup, and navigating various sim-to-real challenges. Furthermore, we analyze the controller's behavior and wave-smoothing properties, and deploy it on four Toyota Rav4’s in a real-world validation experiment on the I-24. Lastly, we release the driving dataset, the simulator and the trained controller, to enable future benchmarking and controller design.

42 ENGINEERING↗

GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning

Large language models (LLMs) are increasingly adapted to downstream tasks via reinforcement learning (RL) methods like Group Relative Policy Optimization (GRPO), which often require thousands of rollouts to learn new tasks. We argue that the interpretable nature of language often provides a much richer learning medium for LLMs, compared to policy gradients derived from sparse, scalar rewards. To test this, we introduce GEPA (Genetic-Pareto), a prompt optimizer that thoroughly incorporates natural language reflection to learn high-level rules from trial and error. Given any AI system containing one or more LLM prompts, GEPA samples trajectories (e.g., reasoning, tool calls, and tool outputs) and reflects on them in natural language to diagnose problems, propose and test prompt updates, and combine complementary lessons from the Pareto frontier of its own attempts. As a result of GEPA's design, it can often turn even just a few rollouts into a large quality gain. Across six tasks, GEPA outperforms GRPO by 6% on average and by up to 20%, while using up to 35x fewer rollouts. GEPA also outperforms the leading prompt optimizer, MIPROv2, by over 10% (e.g., +12% accuracy on AIME-2025), and demonstrates promising results as an inference-time search strategy for code optimization. We release our code at https://github.com/gepa-ai/gepa.

97 MATHEMATICS AND COMPUTING↗

A Novel Multi-Agent Deep Reinforcement Learning-enabled Distributed Power Allocation Scheme for mmWave Cellular Networks

We consider the power allocation problem over shared spectrum for millimeter-Wave (mmWave) cellular downlink. Existing approaches usually find sub-optimal solutions by solving a non-convex optimization which leads to scalability issues due to centralized control. Therefore, distributed and adaptive approaches are desirable. Recently, model-free Deep Reinforcement Learning (DRL) has achieved success in such wireless resource management tasks. By modeling the radio environment as a Markov Decision Process (MDP) with the base stations (BSs) being the agents, power allocation can be automated at the agent level with comparable throughput performance to conventional centralized schemes. The multi-agent setting presents new challenges as the radio environment is impacted by the joint actions of the agents and is no longer stationary from any individual agent’s perspective. Existing literature bypasses this non-stationarity violation by ignoring it which may cause performance degradation. To tackle this issue, we propose a distributed continuous power allocation scheme based on a modified version of multi-agent Deep Deterministic Policy Gradient (MADDPG) that is tailored for the distributed multiple-agent setting. The proposed scheme employs a centralized-training distributed-execution framework where Q-functions are trained over subsets of BSs while each BS determines its transmit power based only on its own local observation. It admits constant per-BS communication and computation complexity and is thus scalable to large networks. Numerical evaluation shows that the proposed scheme adapts well to a wide range of interference conditions and can achieve comparable or better performance than several state-of-the-art non-learning approaches.

99 GENERAL AND MISCELLANEOUS↗

Hierarchical Flexibility Offering Strategy for Integrated Hybrid Resources in Real-time Energy Markets

This paper proposes a hierarchical model for determining the energy flexibility offering strategy of integrated hybrid resources (IHRs) in power distribution systems to participate in real-time energy markets. The proposed model utilizes the scalability, fast response time, and uncertainty observation of deep reinforcement learning (DRL) to overcome the scalability issue of operating numerous flexible resources and deliverability of energy flexibility to the real-time markets in the presence of the network constraints. To that end, the power distribution system is divided into multiple IHRs, where different types of flexible loads, energy storage systems, and solar plants with controllable inverters are operated through local IHR controllers, trained by deep deterministic policy gradient (DDPG) algorithm. Active power request and reactive power capacity of IHRs are then transmitted to a central flexibility controller, where a quadratic optimization model ensures the deliverability of the energy flexibility to the real-time energy market by satisfying the distribution network constraints. The proposed model is implemented on the 123-bus test power distribution system, demonstrating the capability of DRL-based hierarchical model for scalable operation of IHRs in order to offer deliverable energy flexibility to the real-time energy market.

Majidi, Majid↗

Optimal Coordination of Electric Vehicles for Grid Services using Deep Reinforcement Learning

Recent research has shown the effectiveness of reinforcement learning (RL) in coordinating electric vehicles (EVs) with vehicle-to-grid capabilities for grid services. However, many of these studies rely on lookup table and deep Q-network techniques, which can be impractical when dealing with continuous states and actions. In addition, existing RL designs inadequately account for battery aging effects, EV user satisfaction, uncertain departure and arrival time, and trip distance, which may compromise effective coordination. This paper aims to bridge these gaps by developing an innovative deep deterministic policy gradient-based RL framework for optimal coordination of EVs. Case studies were carried out using a test system with 100 EVs, and numerical analysis results showed that the proposed RL framework can effectively coordinate EVs to maximize economic benefits and user satisfaction while ensuring the expected battery lifespan.

Das, Avijit↗

Variational actor-critic algorithms,

We introduce a class of variational actor-critic algorithms based on a variational formulation over both the value function and the policy. The objective function of the variational formulation consists of two parts: one for maximizing the value function and the other for minimizing the Bellman residual. Besides the vanilla gradient descent with both the value function and the policy updates, we propose two variants, the clipping method and the flipping method, in order to speed up the convergence. We also prove that, when the prefactor of the Bellman residual is sufficiently large, the fixed point of the algorithm is close to the optimal policy.

97 MATHEMATICS AND COMPUTING↗

Binary Quantum Control Optimization with Uncertain Hamiltonians

Optimizing the controls of quantum systems plays a crucial role in advancing quantum technologies. The time-varying noises in quantum systems and the widespread use of inhomogeneous quantum ensembles raise the need for high-quality quantum controls under uncertainties. In this paper, we consider a stochastic discrete optimization formulation of a discretized binary optimal quantum control problem involving Hamiltonians with predictable uncertainties. We propose a sample-based reformulation that optimizes both risk-neutral and risk-averse measurements of control policies, and solve these with two gradient-based algorithms using sum-up-rounding approaches. Furthermore, we discuss the differentiability of the objective function and prove upper bounds of the gaps between the optimal solutions to binary control problems and their continuous relaxations. We conduct numerical simulations on various sized problem instances based on two applications of quantum pulse optimization; we evaluate different strategies to mitigate the impact of uncertainties in quantum systems. In conclusion, we demonstrate that the controls of our stochastic optimization model achieve significantly higher quality and robustness compared with the controls of a deterministic model.

conditional value-at-risk (CVaR)↗

Climate Warming Alters Nutrient Storage in Seasonally Dry Forests: Insights From a 2,300 m Elevation Gradient

Understanding potential response of forest carbon (C) and nutrient storage to warming is important for climate mitigation policies. Unfortunately, those responses are difficult to predict in seasonally dry forests, in part, because ecosystem processes are highly sensitive to both changes in temperature and precipitation. We investigated how warming might alter stocks of C, nitrogen (N), and phosphorus (P) in vegetation and the entire regolith (soil + weathered bedrock or “saprock”) using a space-for-time substitution along a bioclimatic gradient in the Sierra Nevada, California. The pine-oak and mixed-conifer forests between 1,160–2,015 m elevation have more optimal climates (not too dry or hot) for ecosystem productivity, soil weathering, and cycling of essential elements than the oak savannah (405 m) and subalpine forest (2,700 m). We found decreases in overstory vegetation nutrient stocks with decreasing elevation because of enhanced water limitation and greater occurrence of disturbances. Stocks of C, N, and P in the entire regolith peaked at the pine-oak and mixed-conifer forests across the bioclimatic gradient, driven by thicker regolith profiles and greater nutrient input rates. These observations suggest long-term warming will decrease ecosystem nutrient storage at the warmer, transitional pine-oak zone, but will increase nutrient storage at the colder, subalpine zone. Assuming steady-state conditions, we found the mean residence time of ecosystem C decreased with projected rising air temperatures and increased following a major drought event across the bioclimatic gradient. Our study emphasizes potentially elevation-dependent changes in nutrient storage and C persistence with warming in seasonally dry forests.

54 ENVIRONMENTAL SCIENCES↗

From Optimization to Sampling Through Gradient Flows

Optimization and sampling algorithms play a central role in science and engineering as they enable finding optimal predictions, policies, and recommendations, as well as expected and equilibrium states of complex systems. The notion of “optimality” is formalized by the choice of an objective function, while the notion of an “expected” state is specified by a probabilistic model for the distribution of states. Optimizing rugged objective functions and sampling multimodal distributions is computationally challenging, especially in high-dimensional problems. Here, for this reason, many optimization and sampling methods have been developed by researchers working in disparate fields such as Bayesian statistics, molecular dynamics, genetics, quantum chemistry, machine learning, weather forecasting, econometrics, and medical imaging.

Trillos, N. García↗

A Two-Stage Quantum Reinforcement Learning Method for Multi-Objective Transmission Switching

Multi-objective transmission switching (MO-TS) problems involve the strategic reconfiguration of network topology to simultaneously optimize multiple objectives. As the system scale increases, finding feasible solutions becomes increasingly challenging due to the problem's nonlinearity and high computational complexity. To address these challenges, this paper proposes a two-stage quantum reinforcement learning method that leverages potential quantum advantages for MO-TS. In the first stage, candidate switching lines are identified using a graph-theoretical approach to reduce the problem's dimensionality. The second stage introduces a quantum-classical reinforcement learning framework, where a learnable measurement-based CNN-ResVQC architecture is developed to effectively reduce the input dimension for quantum processing, mitigate vanishing gradients, and enhance trainability while improving the quantum circuit's flexibility in modeling complex decision policies for MO-TS. Numerical studies on IEEE 14-bus, 57-bus, and 118-bus systems demonstrate that the proposed algorithm achieves superior training stability and faster convergence with approximately 1% of the network parameters required by classical algorithms, highlighting its effectiveness, efficiency, and scalability. Furthermore, the practicality is validated through its stable convergence under three common quantum noise channels.

99 GENERAL AND MISCELLANEOUS↗

Method and apparatus for constructing informative outcomes to guide multi-policy decision making

In Multi-Policy Decision-Making (MPDM), many computationally-expensive forward simulations are performed in order to predict the performance of a set of candidate policies. In risk-aware formulations of MPDM, only the worst outcomes affect the decision making process, and efficiently finding these influential outcomes becomes the core challenge. Recently, stochastic gradient optimization algorithms, using a heuristic function, were shown to be significantly superior to random sampling. In this disclosure, it was shown that accurate gradients can be computed-even through a complex forward simulation—using approaches similar to those in dep networks. The proposed approach finds influential outcomes more reliably, and is faster than earlier methods, allowing one to evaluate more policies while simultaneously eliminating the need to design an easily-differentiable heuristic function.

Olson, Edwin↗

Improved assessment of mangrove forests in Sundarbans East Wildlife Sanctuary using WorldView 2 and TanDEM-X high resolution imagery

Recent developments of remote sensing techniques which can capture both the structure and function of the ecosystem provide a more representative view of the landscape. These unique Earth observations were used to help improve traditional forestry surveys by providing species-specific land cover classes for mangrove forests in the Sundarbans East Wildlife Sanctuary. By combining optical data from WorldView2 (WV2; 2 m pixel) and a canopy height model derived using radar data from TanDEM-X (TDX; 12 m pixel), we identified nine mangrove and five non-mangrove classes by following an Iterative Self-Organizing Data Analysis Algorithm. Three dominant mangrove species accounted for nearly 50% of the sanctuary. Heritieria fomes disproportionately covered the largest area at 43%, overturning previous field-based estimates of Excoecaria agallocha dominance. E. agallocha and Sonneratia apetala, covered 3% and 1.47% of the sanctuary, respectively. Four mixed species classes were also identified with clear vegetation zonation patterns that trended toward species homogeneity with increasing distance from shore. The overall land cover accuracy (WV2: 89.33%; WV2-TDX: 89.89%), the Kappa Coefficient (WV2:0.88; WV2-TDX: 0.89) and change statistics between WV2 and WV2-TDX landcover classifications indicate that the WV2 imagery can separate mangrove community types without structural data. The combination of the land cover classifications and the canopy height model indicated that H. fomes were not only the most dominant forest but also, on average, the tallest (12.3 m) among the other eight mangrove types. Our large-scale mapping with high resolution optical and radar platforms can capture subtle changes in mangrove vegetation and canopy structural gradients more accurately and be used to monitor biodiversity changes and Aichi Biodiversity Targets and Indicators, which would contribute to biodiversity policy updating.

Md Mizanur Rahman↗

Hyperspectral leaf reflectance of grasses varies with evolutionary lineage more than with site

Abstract To predict ecological responses at broad environmental scales, grass species are commonly grouped into two broad functional types based on photosynthetic pathway. However, closely related species may have distinctive anatomical and physiological attributes that influence ecological responses, beyond those related to photosynthetic pathway alone. Hyperspectral leaf reflectance can provide an integrated measure of covarying leaf traits that may result from phylogenetic trait conservatism and/or environmental conditions. Understanding whether spectra‐trait relationships are lineage specific or reflect environmental variation across sites is necessary for using hyperspectral reflectance to predict plant responses to environmental changes across spatial scales. We measured hyperspectral leaf reflectance (400–2400 nm) and 12 structural, biochemical, and physiological leaf traits from five grass‐dominated sites spanning the Great Plains of North America. We assessed if variation in leaf reflectance spectra among grass species is explained more by evolutionary lineage (as captured by tribes or subfamilies), photosynthetic pathway (C 3 or C 4 ), or site differences. We then determined whether leaf spectra can be used to predict leaf traits within and across lineages. Our results using redundancy analysis ordination (RDA) show that grass tribe identity explained more variation in leaf spectra (adjusted R 2 = 0.12) than photosynthetic pathway, which explained little variation in leaf spectra (adjusted R 2 = 0.00). Furthermore, leaf reflectance from the same tribe across multiple sites was more similar than leaf reflectance from the same site across tribes (adjusted R 2 = 0.12 and 0.08, respectively). Across all sites and species, trait predictions based on spectra ranged considerably in predictive accuracies ( R 2 = 0.65 to <0.01), but R 2 was >0.80 for certain lineages and sites. The relationship between Vc max , a measure of photosynthetic capacity, and spectra was particularly promising. Chloridoideae, a lineage more common at drier sites, appears to have distinct spectra‐trait relationships compared with other lineages. Overall, our results show that evolutionary relatedness explains more variation in grass leaf spectra than photosynthetic pathway or site, but consideration of lineage‐ and site‐specific trait relationships is needed to interpret spectral variation across large environmental gradients.

Pau, Stephanie [Department of Geography University↗