Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “policy optimization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Modal cost analysis for linear matrix-second-order systems

Reduced models and reduced controllers for systems governed by matrix-second-order differential equations are obtained by retaining those modes which make the largest contributions to quadratic control objectives. Such contributions, expressed in terms of modal data, used as mode truncation criteria, allow the statement of the specific control objectives to influence the early model reduction from very high order models which are available, for example, from finite element methods. The relative importance of damping, frequency, and eigenvector in the mode truncation decisions are made explicit for each of these control objectives: attitude control, vibration suppression and figure control. The paper also shows that using modal cost analysis (MCA) on the closed loop modes of the optimally controlled system allows the construction of reduced control policies which feedback only those closed loop modal coordinates which are most critical to the quadratic control performance criterion. In this way, the modes which should be controlled (and hence the modes which must be observable by choice of measurements), are deduced from truncations of the optimal controller.

Skelton, R. E.↗

Federated Deep Reinforcement Learning for Decentralized VVO of BTM DERs

The future of grid control requires a hybrid approach combining centralized and decentralized methods to fully utilize the potential of smart edge devices with artificial intelligence (AI) capabilities. This paper aims to develop and evaluate a federated deep reinforcement learning (FDRL) framework for decentralized adaptive volt-var optimization (VVO) of behind-the-meter (BTM) distributed energy resources (DERs). First, this paper models a single deep reinforcement learning (DRL) agent using the Markov Decision Process (MDP) framework for decentralized adaptive VVO of BTM DERs. Two DRL algorithms, soft actor-critic (SAC) and twin-delayed deep deterministic policy gradient (TD3), are compared for their effectiveness in optimizing VVO. Results show that TD3 outperforms SAC, achieving a 71.3% improvement in mean reward. Finally, the DRL agent is deployed within the FDRL framework, using the Flower platform, to enhance learning, provide adaptive control, and ensure data privacy for BTM DERs.

Ravi, Abhijith↗

Probabilistic Deliverability Assessment of Distributed Energy Resources via Scenario-Based AC Optimal Power Flow

As electric grids decarbonize and distributed energy resources (DERs) become increasingly prevalent, interconnection assessments must evolve to reflect operational variability and control flexibility. This paper highlights key modeling limitations observed in practice and reviews approaches for modeling uncertainty. It then introduces a Probabilistic Deliverability Assessment (PDA) framework designed to complement and extend existing procedures. The framework integrates scenario-based AC optimal power flow (AC OPF), corrective dispatch, and optional multi-temporal constraints. Together, these form a structured methodology for quantifying DER utilization, deliverability, and reliability under uncertainty in load, generation, and topology. Outputs include interpretable metrics with confidence intervals that inform siting decisions and evaluate compliance with reliability thresholds across sampled operating conditions. A case study on Puerto Rico’s publicly available bulk power system model demonstrates the framework’s application using minimal input data, consistent with current interconnection practice. Across staged fossil generation retirements, the PDA identifies high-value DER sites and regions requiring additional reactive power support. Results are presented through mean dispatch signals, reliability metrics, and geospatial visualizations, demonstrating how the framework provides transparent, data-driven siting recommendations. The framework’s modular design supports incremental adoption within existing workflows, encouraging broader use of AC OPF in interconnection and planning contexts.

14 SOLAR ENERGY↗

Quantifying Energy Justice Goals in the Power Sector: Developing and Using Metrics

New policy goals are explicitly guiding the future grid toward greater energy equity. At the same time, policy goals also guide the grid toward decarbonization and resilience, while maintaining cost, security, and reliability cornerstone requirements. In conclusion, these pressures create a dilemma: moving urgently to address climate change and respond to energy disruptions, but also slowly to engage communities on the climate frontlines in earnest.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Development of A Composite Drought Indicator for Operational Drought Monitoring in the MENA Region

This paper presents the composite drought indicator (CDI) that Jordanian, Lebanese, Moroccan, and Tunisian government agencies now produce monthly to support operational drought management decision making, and it describes their iterative co-development processes. The CDI is primarily intended to monitor agricultural and ecological drought on a seasonal time scale. It uses remote sensing and modelled data inputs, and it reflects anomalies in precipitation, vegetation, soil moisture, and evapotranspiration. Following quantitative and qualitative validation assessments, engagements with policymakers, and consideration of agencies’ technical and institutional capabilities and constraints, we made changes to CDI input data, modelling procedures, and integration to tailor the system for each national context. We summarize validation results, drought modelling challenges and how we overcame them through CDI improvements, and we describe the monthly CDI production process and outputs. Finally, we synthesize procedural and technical aspects of CDI development and reflect on the constraints we faced as well as trade-offs made to optimize the CDI for operational monitoring to support policy decision-making—including aspects of salience, credibility, and legitimacy—within each national context.

Karim Bergaoui↗

Cost and Benefit Analysis of Mitigating, Tracking, and Remediating Orbital Debris

Orbital debris may collide with crewed and robotic spacecraft, placing them at risk. The wide range of debris, from 9,000-kilogram rocket bodies to millions of millimeter-size debris, has led to a similarly wide range of proposed actions for addressing the risks posed by debris. However, the costs and benefits of these actions have historically been unknown. This is a challenge for decision makers who are choosing which actions to support through technology development or policy changes. NASA’s Office of Technology, Policy, and Strategy is addressing these technical and economic uncertainties by building a capability to (1) complete rigorous calculations of the net present value of each action, (2) identify an optimal portfolio of actions to reduce risk, and (3) quantitatively analyze policies related to space sustainability. This report describes our progress toward that capability and to solicit feedback from the space and economic communities. Our previous work, referred to here as Phase 1, assessed the costs and benefits of performing debris remediation on operationally relevant timescales. The current analysis contains major updates to the risk model used in Phase 1 and expands the breadth of actions considered to include mitigating the creation of debris, improving the ability to track debris, and more methods for cleaning up existing debris. We demonstrate that our approach of measuring risks in dollars allows for the effectiveness of seemingly incommensurate actions to be compared and generates insights that other approaches to measuring risk have missed.

Jericho Locke↗

Cost and Benefit Analysis of Mitigating, Tracking, and Remediating Orbital Debris

Presentation accompanying the technical report with the same name. The abstract for the report and the full study is repeated below. Orbital debris may collide with crewed and robotic spacecraft, placing them at risk. The wide range of debris, from 9,000-kilogram rocket bodies to millions of millimeter-size debris, has led to a similarly wide range of proposed actions for addressing the risks posed by debris. However, the costs and benefits of these actions have historically been unknown. This is a challenge for decision makers who are choosing which actions to support through technology development or policy changes. NASA’s Office of Technology, Policy, and Strategy is addressing these technical and economic uncertainties by building a capability to (1) complete rigorous calculations of the net present value of each action, (2) identify an optimal portfolio of actions to reduce risk, and (3) quantitatively analyze policies related to space sustainability. This report describes our progress toward that capability and to solicit feedback from the space and economic communities. Our previous work, referred to here as Phase 1, assessed the costs and benefits of performing debris remediation on operationally relevant timescales. The current analysis contains major updates to the risk model used in Phase 1 and expands the breadth of actions considered to include mitigating the creation of debris, improving the ability to track debris, and more methods for cleaning up existing debris. We demonstrate that our approach of measuring risks in dollars allows for the effectiveness of seemingly incommensurate actions to be compared and generates insights that other approaches to measuring risk have missed.

Jericho Locke↗

A Cost and Benefit Analysis of Orbital Debris Remediation, Mitigation,Tracking, and Characterization

Orbital debris may collide with crewed and robotic spacecraft, placing them at risk. The wide range of debris, from 9,000-kilogram rocket bodies to millions of millimeter-size debris, has led to a similarly wide range of proposed actions for addressing the risks posed by debris. However, the costs and benefits of these actions have historically been unknown. This is a challenge for decision makers who are choosing which actions to support through technology development or policy changes. NASA’s Office of Technology, Policy, and Strategy is addressing these technical and economic uncertainties by building a capability to (1) complete rigorous calculations of the net present value of each action, (2) identify an optimal portfolio of actions to reduce risk, and (3) quantitatively analyze policies related to space sustainability. This report describes our progress toward that capability and to solicit feedback from the space and economic communities. Our previous work(Colvin, Karcz, and Wusk. 2023), referred to here as Phase 1, assessed the costs and benefits of performing debris remediation on operationally relevant timescales. The current paper summarizes the work of Locke and Colvin (2024), which contains major updates to the risk model used in Phase 1 and expands the breadth of actions considered to include mitigating the creation of debris, improving the ability to track debris, and more methods for cleaning up existing debris. We demonstrate that our approach of measuring risks in dollars allows for the effectiveness of seemingly incommensurate actions to be compared and generates insights that other approaches to measuring risk have missed.

Orbital Debris↗

Scale-Bridging Optimization Framework for Desalination Integrated Produced Water Networks

In this work, we develop a Pyomo-based non-linear optimization strategy that includes rigorous MVR models. The detailed desalination unit is integrated into the multiperiod produced water network problem using the trust region filter (TRF) method. TRF decomposes the integrated problem into a master problem consisting of the network variables and a simplified surrogate model for the detailed desalination unit. The surrogate is updated using zero and first-order corrections from the optimal solution of the detailed models at every iteration. This framework allows us to co-optimize the design of the desalination units and operating policy for the multiperiod network. A common design is ensured across all periods using global capacity constraints. We validate the solution obtained using the TRF method by solving the full integrated problem for small network instances and show our results on real case studies on produced water networks from the Permian and Appalachian basins. In this work, we describe our TRF formulation, give details on our implementation in Pyomo, and analyze the results obtained by solving the optimization problem using IPOPT. We also present a discussion on the computational efficiency and scaling using the TRF approach against a full-scale integration of the rigorous models within the water network.

Naik, Sakshi↗

Achieving the Proper Balance Between Crew and Public Safety

A paramount objective of all human-rated launch and reentry vehicle developers is to ensure that the risks to both the crew onboard and the public are minimized within reasonable cost, schedule, and technical constraints. Past experience has shown that proper attention to range safety requirements necessary to ensure public safety must be given early in the design phase to avoid additional operational complexities or threats to the safety of people onboard, and the design engineers must give these requirements the same consideration as crew safety requirements. For human spaceflight, the primary purpose and operational concept for any flight safety system is to protect the public while maximizing the likelihood of crew survival. This paper will outline the policy considerations, technical issues, and operational impacts regarding launch and reentry vehicle failure scenarios where crew and public safety are intertwined and thus addressed optimally in an integrated manner. An overview of existing range and crew safety policy requirements will be presented. Application of these requirements and lessons learned from both the Space Shuttle and Constellation Programs will also be discussed. Using these past programs as examples, the paper will detail operational, design, and analysis approaches to mitigate and balance the risks to people onboard and in the public. Manned vehicle perspectives from the Federal Aviation Administration (FAA) and Air Force organizations that oversee public safety will be summarized as well. Finally, the paper will emphasize the need to factor policy, operational, and analysis considerations into the early design trades of new vehicles to help ensure that both crew and public safety are maximized to the greatest extent possible.

Gowan, John↗

Aircraft Trajectory Optimization and Contrails Avoidance in the Presence of Winds

There are indications that persistent contrails can lead to adverse climate change, although the complete effect on climate forcing is still uncertain. A flight trajectory optimization algorithm with fuel and contrails models, which develops alternative flight paths, provides policy makers the necessary data to make tradeoffs between persistent contrails mitigation and aircraft fuel consumption. This study develops an algorithm that calculates wind-optimal trajectories for cruising aircraft while avoiding the regions of airspace prone to persistent contrails formation. The optimal trajectories are developed by solving a non-linear optimal control problem with path constraints. The regions of airspace favorable to persistent contrails formation are modeled as penalty areas that aircraft should avoid and are adjustable. The tradeoff between persistent contrails formation and additional fuel consumption is investigated, with and without altitude optimization, for 12 city-pairs in the continental United States. Without altitude optimization, the reduction in contrail travel times is gradual with increase in total fuel consumption. When altitude is optimized, a two percent increase in total fuel consumption can reduce the total travel times through contrail regions by more than six times. Allowing further increase in fuel consumption does not seem to result in proportionate decrease in contrail travel times.

Ng, Hok K.↗

Concepts and Challenges for Environmentally Friendly En Route Operations

A flight trajectory optimization algorithm with fuel and contrails models, which develops alternative flight paths, provides policy makers the necessary data to make tradeoffs between persistent contrails mitigation and aircraft fuel consumption. This study develops an algorithm that calculates wind-optimal trajectories for cruising aircraft while avoiding the regions of airspace prone to persistent contrails formation. The optimal trajectory is derived using Singular Perturbation Method. The regions of airspace favorable to persistent contrails formation are modeled as high-risk areas that aircraft should avoid and are adjustable. The tradeoffs between persistent contrails formation and additional travel time are investigated for wind-optimal trajectories and various contrails-avoidance trajectories at 10 different cruising altitudes for flights departing from Chicago and San Diego to New York. The additional travel times required for avoiding 100% persistent contrails formation at various flight altitudes ranged from approximately 0% to 4.3% for flights from Chicago to New York. For flights between San Diego and New York, additional traveling times vary between 1.3% and 5% depending on the cruise altitude and the percentage of contrail avoidance. Talk will present the results of aircraft fuel consumptions that are proportional to the travel time for cruising aircraft.

Sridhar, Banavar↗

Decision-Aiding and Optimization for Vertical Navigation of Long-Haul Aircraft

Most decisions made in the cockpit are related to safety, and have therefore been proceduralized in order to reduce risk. There are very few which are made on the basis of a value metric such as economic cost. One which can be shown to be value based, however, is the selection of a flight profile. Fuel consumption and flight time both have a substantial effect on aircraft operating cost, but they cannot be minimized simultaneously. In addition, winds, turbulence, and performance vary widely with altitude and time. These factors make it important and difficult for pilots to (a) evaluate the outcomes associated with a particular trajectory before it is flown and (b) decide among possible trajectories. The two elements of this problem considered here are: (1) determining what constitutes optimality, and (2) finding optimal trajectories. Pilots and dispatchers from major u.s. airlines were surveyed to determine which attributes of the outcome of a flight they considered the most important. Avoiding turbulence-for passenger comfort-topped the list of items which were not safety related. Pilots' decision making about the selection of flight profile on the basis of flight time, fuel burn, and exposure to turbulence was then observed. Of the several behavioral and prescriptive decision models invoked to explain the pilots' choices, utility maximization is shown to best reproduce the pilots' decisions. After considering more traditional methods for optimizing trajectories, a novel method is developed using a genetic algorithm (GA) operating on a discrete representation of the trajectory search space. The representation is a sequence of command altitudes, and was chosen to be compatible with the constraints imposed by Air Traffic Control, and with the training given to pilots. Since trajectory evaluation for the GA is performed holistically, a wide class of objective functions can be optimized easily. Also, using the GA it is possible to compare the costs associated with different airspace design and air traffic management policies. A decision aid is proposed which would combine the pilot's notion of optimality with the GA-based optimization, provide the pilot with a number of alternative pareto-optimal trajectories, and allow him to consider unmodelled attributes and constraints in choosing among them. A solution to the problem of displaying alternatives in a multi-attribute decision space is also presented.

Patrick, Nicholas J. M.↗

A robust model predictive control algorithm for uncertain nonlinear systems that guarantees resolvability

A robustly stabilizing MPC (model predictive control) algorithm for uncertain nonlinear systems is developed that guarantees resolvability. With resolvability, initial feasibility of the finite-horizon optimal control problem implies future feasibility in a receding-horizon framework. The control consists of two components; (i) feed-forward, and (ii) feedback part. Feed-forward control is obtained by online solution of a finite-horizon optimal control problem for the nominal system dynamics. The feedback control policy is designed off-line based on a bound on the uncertainty in the system model. The entire controller is shown to be robustly stabilizing with a region of attraction composed of initial states for which the finite-horizon optimal control problem is feasible. The controller design for this algorithm is demonstrated on a class of systems with uncertain nonlinear terms that have norm-bounded derivatives and derivatives in polytopes. An illustrative numerical example is also provided.

feed - forward↗

Small Body GN&C Research Report: A Robust Model Predictive Control Algorithm with Guaranteed Resolvability

A robustly stabilizing MPC (model predictive control) algorithm for uncertain nonlinear systems is developed that guarantees the resolvability of the associated finite-horizon optimal control problem in a receding-horizon implementation. The control consists of two components; (i) feedforward, and (ii) feedback part. Feed-forward control is obtained by online solution of a finite-horizon optimal control problem for the nominal system dynamics. The feedback control policy is designed off-line based on a bound on the uncertainty in the system model. The entire controller is shown to be robustly stabilizing with a region of attraction composed of initial states for which the finite-horizon optimal control problem is feasible. The controller design for this algorithm is demonstrated on a class of systems with uncertain nonlinear terms that have norm-bounded derivatives, and derivatives in polytopes. An illustrative numerical example is also provided.

model predictive control algorithm (MPC)↗

Feedback Optimization of Incentives for Distribution Grid Services

Energy prices and net power injection limitations regulate the operations in distribution grids and typically ensure that operational constraints are met. Nevertheless, unexpected or prolonged abnormal events could undermine the grid's functioning. During contingencies, customers could contribute effectively to sustaining the network by providing services. Herein this paper proposes an incentive mechanism that promotes users' active participation by essentially altering the energy pricing rule. The incentives are modeled via a linear function whose parameters can be computed by the system operator (SO) by solving an optimization problem. Feedback-based optimization algorithms are then proposed to seek optimal incentives by leveraging measurements from the grid, even in the case when the SO does not have a full grid and customer information. Numerical simulations on a standard testbed validate the proposed approach.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Binary Quantum Control Optimization with Uncertain Hamiltonians

Optimizing the controls of quantum systems plays a crucial role in advancing quantum technologies. The time-varying noises in quantum systems and the widespread use of inhomogeneous quantum ensembles raise the need for high-quality quantum controls under uncertainties. In this paper, we consider a stochastic discrete optimization formulation of a discretized binary optimal quantum control problem involving Hamiltonians with predictable uncertainties. We propose a sample-based reformulation that optimizes both risk-neutral and risk-averse measurements of control policies, and solve these with two gradient-based algorithms using sum-up-rounding approaches. Furthermore, we discuss the differentiability of the objective function and prove upper bounds of the gaps between the optimal solutions to binary control problems and their continuous relaxations. We conduct numerical simulations on various sized problem instances based on two applications of quantum pulse optimization; we evaluate different strategies to mitigate the impact of uncertainties in quantum systems. In conclusion, we demonstrate that the controls of our stochastic optimization model achieve significantly higher quality and robustness compared with the controls of a deterministic model.

conditional value-at-risk (CVaR)↗

SatNet: A Benchmark for Satellite Scheduling Optimization

Satellites provide essential services such as networking and weather tracking, and the number of near-earth and deep space satellites are expected to grow rapidly in the coming years. Communications with terrestrial ground stations is one of the critical functionalities of any space mission. Satellite scheduling is a problem that has been scientifically investigated since the 1970s. A central aspect of this problem is the need to consider resource contention and satellite visibility constraints as they require line of sight. Due to the combinatorial nature of the problem, prior solutions such as linear programs and evolutionary algorithms require extensive compute capabilities to output a feasible schedule for each scenario. Machine learning based scheduling can provide an alternative solution by training a model with historical data and generating a schedule quickly with model inference. We present SatNet, a benchmark for satellite scheduling optimization based on historical data from the NASA Deep Space Network. We propose formulation of the satellite scheduling problem as a Markov Decision Process and use reinforcement learning (RL) policies to generate schedules. The nature of constraints imposed by SatNet differ from other combinatorial optimization problems such as vehicle routing studied in prior literature. Our initial results indicate that RL is an alternative optimization approach that can generate candidate solutions of comparable quality to existing state-of-the-practice results. However, we also find that RL policies overfit to the training dataset and do not generalize well to new data, thereby necessitating continued research on reusable and generalizable agents.

Wilson, Brian↗