Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Inverse Reinforcement Learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Use of Inverse Reinforcement Learning for Identity Prediction

We adopt Markov Decision Processes (MDP) to model sequential decision problems, which have the characteristic that the current decision made by a human decision maker has an uncertain impact on future opportunity. We hypothesize that the individuality of decision makers can be modeled as differences in the reward function under a common MDP model. A machine learning technique, Inverse Reinforcement Learning (IRL), was used to learn an individual's reward function based on limited observation of his or her decision choices. This work serves as an initial investigation for using IRL to analyze decision making, conducted through a human experiment in a cyber shopping environment. Specifically, the ability to determine the demographic identity of users is conducted through prediction analysis and supervised learning. The results show that IRL can be used to correctly identify participants, at a rate of 68% for gender and 66% for one of three college major categories.

Hayes, Roy↗

ARM-IRL: Adaptive Resilience Metric Quantification Using Inverse Reinforcement Learning

The resilience of safety-critical systems is gaining importance due to the rise in cyber and physical threats, especially within critical infrastructure. Traditional static resilience metrics may not capture dynamic system states, leading to inaccurate assessments and ineffective responses to cyber threats. This work aims to develop a data-driven, adaptive method for resilience metric learning. We propose a data-driven approach using inverse reinforcement learning (IRL) to learn a single, adaptive resilience metric. The method infers a reward function from expert control actions. Unlike previous approaches using static weights or fuzzy logic, this work applies adversarial inverse reinforcement learning (AIRL), training a generator and discriminator in parallel to learn the reward structure and derive an optimal policy. The proposed approach is evaluated on multiple scenarios: optimal communication network rerouting, power distribution network reconfiguration, and cyber–physical restoration of critical loads using the IEEE 123-bus system. The adaptive, learned resilience metric enables faster critical load restoration in comparison to conventional RL approaches.

97 MATHEMATICS AND COMPUTING↗

Ground Delay Program Analytics with Behavioral Cloning and Inverse Reinforcement Learning

We used historical data to build two types of model that predict Ground Delay Program implementation decisions and also produce insights into how and why those decisions are made. More specifically, we built behavioral cloning and inverse reinforcement learning models that predict hourly Ground Delay Program implementation at Newark Liberty International and San Francisco International airports. Data available to the models include actual and scheduled air traffic metrics and observed and forecasted weather conditions. We found that the random forest behavioral cloning models we developed are substantially better at predicting hourly Ground Delay Program implementation for these airports than the inverse reinforcement learning models we developed. However, all of the models struggle to predict the initialization and cancellation of Ground Delay Programs. We also investigated the structure of the models in order to gain insights into Ground Delay Program implementation decision making. Notably, characteristics of both types of model suggest that GDP implementation decisions are more tactical than strategic: they are made primarily based on conditions now or conditions anticipated in only the next couple of hours.

Bloem, Michael↗

MSD CoP Webinar: Modeling the Operations of Reservoir Systems with LLMs and Inverse Reinforcement Learning

Context: This panel featured three presentations centered on the common theme of applying LLMs and inverse reinforcement learning (IRL) to capture the complex human-environment interactions that are central to the operation of reservoir systems. Dr. Wyatt Arnold will kick off the webinar with a talk on how analyzing LLM chain-of-thought reasoning reveals sophisticated quantitative justification and risk awareness, showing promise as a bridge between quantitative models and value-driven water management decisions. Next, Dr. Matteo Giuliani will build on this with a discussion demonstrating that AI- and IRL-driven approaches can infer the trade-offs between flood control and water supply using historical observations. Finally, Dr. Rohan Singh Wilkho will close the webinar with a talk establishing IRL as a generalizable diagnostic tool for decoding decision-making in managed hydrologic and human-infrastructure systems. Across the three presentations, the application of LLMs and IRL opens new possibilities for the development of adaptive, transparent, and human-aware models supporting water management in an increasingly uncertain future. Presenters: Wyatt Arnold (Politecnico di Milano); Matteo Giuliani (Politecnico di Milano); Rohan Singh Wilkho (Cornell University) Moderator: Patrick M. Reed (MSD CoP Facilitation Team); Stefano Galelli (MSD CoP AI Working Group Co-Chair); David Gold (MSD CoP AI Working Group Co-Chair) This webinar was held on: June 23rd, 2026 from 12-1 PM EST.

Arnold, Wyatt [Politecnico di Milano]↗

Inverse reinforcement learning control for building energy management

Reinforcement learning (RL) based control is widely considered a promising approach in building automation and control as it has demonstrated the potential to deal with complex objectives in adjacent domains like robotics, autonomous vehicles, gaming applications, and advertisement recommendations. When applied to any environment, model-free RL learns to improve its control performance over time without requiring a control model, by receiving and then analyzing feedback from the building environment after each control action. Operational objectives are becoming increasingly complex through the simultaneous consideration of thermal comfort, carbon emissions, grid services, and indoor air quality. In this context, conventional rule-based control approaches are proving sub-optimal, mostly heuristic, and inadequate. The model-free and self learning nature of RL appears promising and attractive as it may address the scalability issues associated with advanced control approaches. However, it suffers from long training times and unstable control behavior during the early stages of its learning process, which makes it unsuitable to be applied directly to buildings. This paper addresses these issues using an inverse reinforcement learning approach (IRL), a technique utilized to learn the objective of a controller agent which is considered an expert in its respective domain. Here, we consider a rule-based control as the expert demonstrator. IRL is different from a direct imitation (i.e., direct mapping of states to actions) of control actions as it tries to find the underlying intent of an expert's policy, providing the controller with a better-generalized policy for unseen states or environments with slightly different dynamics. This approach propels the RL controller's policy to levels similar to or better than that of a rule-based policy before it starts learning by interacting with the building. This makes RL for building energy management applications more practical as it prevents the erratic and exploratory behavior in the initial training period, simultaneously speeding up the learning process when compared to applying an untrained RL agent directly to a building environment.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

A Bayesian Approach for Quantifying Data Scarcity when Modeling Human Behavior via Inverse Reinforcement Learning

Computational models that formalize complex human behaviors enable study and understanding of such behaviors. However, collecting behavior data required to estimate the parameters of such models is often tedious and resource intensive. Thus, estimating dataset size as part of data collection planning (also known as Sample Size Determination) is important to reduce the time and effort of behavior data collection while maintaining an accurate estimate of model parameters. In this paper, we present a sample size determination method based on Uncertainty Quantification (UQ) for a specific Inverse Reinforcement Learning (IRL) model of human behavior, in two cases: 1) pre-hoc experiment design—conducted in the planning stage before any data is collected, to guide the estimation of how many samples to collect; and 2) post-hoc dataset analysis—performed after data is collected, to decide if the existing dataset has sufficient samples and whether more data is needed. Here, we validate our approach in experiments with a realistic model of behaviors of people with Multiple Sclerosis (MS) and illustrate how to pick a reasonable sample size target. Our work enables model designers to perform a deeper, principled investigation of effects of dataset size on IRL.

97 MATHEMATICS AND COMPUTING↗

A Modified Maximum Entropy Inverse Reinforcement Learning Approach for Microgrid Energy Scheduling

Increasing popularity of integrating distributed energy resources (DERs) into the power system brings a challenge to optimize the microgrid dispatch policy. The reinforcement learning methods suffer from a long-time problem with the theoretical assumption of the objective/reward function for the microgrid system. Although the traditional inverse reinforcement learning (IRL) approaches can solve this problem to some extent, they encounter a limitation of complex computations for state visitation frequency in the large and continuous state space. To alleviate this limitation, we propose a modified maximum entropy IRL (MMIRL) method to extract the reward function from the expert demonstrations for solving the microgrid energy scheduling problem. The proposed MMIRL algorithm is promising in recovering the reward function and learning the dispatch policy compared to conventional approaches. Case studies are performed in an energy arbitrage problem and a microgrid system with DERs. Results substantiate that the proposed MMIRL approach can learn the dispatch policy with more than 99% efficiency and outperforms other comparative methods.

artificial intelligence, reinforcement learning, m↗

Inverse Reinforcement Learning based Bayesian Goal Inference Method for Early Nuclear Proliferation Detection

Traditional methods for detection of nuclear proliferation indicators are usually applied after nuclear proliferation has already occurred. There is a need to advance these methods to perform early detection of nuclear proliferation indicators. In this project, we formulated an early detection problem as a sequential, decision-making, goal inference problem based on research publications of authors, to determine whether it is possible to infer whether an author will publish on a research activity before it has occurred. To develop and test our approach, we selected a civil nuclear activity for our case study. We constructed a state-action-state transition graph from publications of authors associated with the activity and the co-authors of their publications, using titles, abstracts, and author publication sequences. We then used inverse reinforcement learning to model the goal-directed behavior of authors in trajectories that terminate at selected goal states. Using a Bayesian formulation, we computed the probability that authors would reach each selected state from partially observed trajectories of their state transitions in their research topic space. The state with the highest probability was selected as the most probable goal state. Based on our results, we found that 60% of the times we can infer the correct goal state early; sometimes the inference is either delayed, or multiple states could be inferred as goal states. Overall, our results show that it is possible to perform early detection of research activities of authors in a nuclear technology area. Further research is necessary to establish a more accurate understanding of how topic modeling, topic space grid discretization, and the extent of overlap among trajectories of different goal states, affect the goal inference results. The methods developed in this work may be used to enhance data-driven methods for early detection of nuclear proliferation indicators.

97 MATHEMATICS AND COMPUTING↗

Toward Justifiable Trust in Autonomous Systems Incorporating Human Knowledge in Autonomous Systems through Machine Learning

Trust in Autonomous Systems is largely about humans trusting the decisions made by autonomous systems. This trust can be increased through learning from domain experts. In particular, autonomous systems can learn offline from past mission operations before conducting any operations of its own. Additionally, autonomous systems can learn online by obtaining human feedback during operations. We will discuss several classes of machine learning methods and our application of them to autonomous systems. The first class of methods is anomaly detection, which uses operations data to identify examples of anomalous operations. The second class of methods is inverse reinforcement learning, also known as apprenticeship learning, that takes past operations data as input and yields a controller that is able to duplicate the operations described by the data. The third class is active learning, which identifies examples on which the model is most uncertain and requests domain expert feedback.

Oza, Nikunj C.↗

Early Inference of Nuclear Technology-Directed Research Activities of Authors from Scientific Publications

Nuclear research articles can provide information about early nuclear proliferation indicators such as influential research entities and technology capability levels of a country, but detection of nuclear activities typically occurs after they have started. We investigate the extent to which nuclear research articles can be used to infer whether a research entity will acquire or develop a nuclear technology before it happens. Early detection of nuclear proliferation or technology development indicators from data is challenging due to partial observability, sparse and unlabeled information, and confounding signals from multiple concurrent activities. This paper presents the early detection problem as a sequential decision-making, goal inference problem, where the objective is to characterize and predict an individual’s, organization’s, or a country’s intent (unobserved goal-directed behavior) towards developing a nuclear capability from partially observed sequences of their research publications, using inverse reinforcement learning and Bayesian goal inference methods. A computational framework is presented, and its application demonstrated using 29,196 Scopus records for a case study related to a civil nuclear capability. The case study results serve as a proof-of-concept demonstration for inference of technology-directed research activity of authors who publish in the nuclear domain. The inference method, combined with advanced computing, may be used to assess and monitor activities pertaining to early developmental stages of a nuclear technology or capability, which in turn can help to identify and prioritize activities with nuclear proliferation potential for further investigation.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

RLMolLM: Reinforcement Learning-Enhanced Language Model Framework for Inverse Molecular Design

Inverse molecular design faces significant challenges due to vast chemical space and complex property requirements. While language models show promise for molecular generation, they struggle with validity, multi-property optimization, and structural constraints. This work presents RLMolLM, a reinforcement learning framework combining Proximal Policy Optimization (PPO) with genetic algorithms to address these limitations. Our approach optimizes multiple user-specified properties including quantitative estimates of drug-likeness (QED), synthetic accessibility (SA), and ADMET (absorption, distribution, metabolism, excretion, and toxicity) endpoints without requiring complete model retraining, while maintaining capability for scaffold-constrained generation where specific substructures must be preserved. We outperform state-of-the-art methods for molecular optimization, achieving best QED scores across GDB13, Moses, and Zinc datasets with up to 31% improvement over previous methods while maintaining excellent validity, uniqueness, and novelty metrics. For simultaneous multi-property optimization, our framework achieves substantial improvements in ADMET properties including 4.5-fold reduction in hERG toxicity and enhanced Caco-2 permeability compared to Moses dataset. Under structural constraints, the framework significantly improves molecular validity while preserving scaffolds and effectively optimizing properties. In conclusion, this versatile solution advances pharmaceutical and materials molecular design through effective integration of reinforcement learning and genetic algorithms with multi-property optimization and scaffold preservation.

Genetic algorithms↗

Machine learning enabled discovery of superhard and ultrahard carbon polymorphs

The demand for multifunctional materials has motivated the move from near-equilibrium materials to metastable i.e. out-of-equilibrium phases that can meet several desired target properties. The search for such metastable phases with exotic properties is non-trivial and often serendipitous. Inverse design approaches based on evolutionary search have been powerful tools, but such traditional searches have focused on identifying primarily stable and metastable materials with the lowest enthalpy. The inverse design of materials, with a focus on a desired property such as, for example, hardness is a challenging task because of the expensive computational cost involved in sampling multiple structures. The recent advances in machine learning have brought new powerful AI techniques to the forefront which can potentially revolutionize the inverse design and discovery of materials, especially metastable phases capable of meeting multifunctionality. Here, in this work, we develop and apply an automated reinforcement learning workflow for inverse design that integrates first principles physics and atomistic simulations with machine learning (ML), and high-performance computing to allow rapid exploration of the superhard and ultrahard metastable phases of Carbon. We demonstrate an automatic machine learning based inverse design workflow to map new undiscovered metastable states ranging from near equilibrium to those far-from-equilibrium that satisfy multiple property objectives, specifically bulk moduli, shear moduli and hardness. We create a comprehensive library of carbon stable and metastable phases with varying hardness and subsequently shortlist 10 top performing candidate carbon structures, including two newly reported phases, based on their hardness and characterize their temperature dependent mechanical properties. A neural network model is built using featurization of allotropes of carbon to predict the quasi-harmonic Gibbs free energies. The Gibbs free energies of the top performing phases are analyzed to get an estimate of the experimental synthesizability of these superhard and ultrahard carbon phases. In general, we show using machine learning based inverse design approaches how hitherto inaccessible metastable states can be identified and potentially synthesized to meet the demand for multifunctional materials.

Balasubramanian, Karthik [Univ. of Illinois, Chica↗

A Continuous Action Space Tree search for INverse desiGn (CASTING) framework for materials discovery

Abstract Material properties share an intrinsic relationship with their structural attributes, making inverse design approaches crucial for discovering new materials with desired functionalities. Reinforcement Learning (RL) approaches are emerging as powerful inverse design tools, often functioning in discrete action spaces. This constrains their application in materials design problems, which involve continuous search spaces. Here, we introduce an RL-based framework CASTING (Continuous Action Space Tree Search for inverse design), that employs a decision tree-based Monte Carlo Tree Search (MCTS) algorithm with continuous space adaptation through modified policies and sampling. Using representative examples like Silver (Ag) for metals, Carbon (C) for covalent systems, and multicomponent systems such as graphane, boron nitride, and complex correlated oxides, we showcase its accuracy, convergence speed, and scalability in materials discovery and design. Furthermore, with the inverse design of super-hard Carbon phases, we demonstrate CASTING’s utility in discovering metastable phases tailored to user-defined target properties and preferences.

36 MATERIALS SCIENCE↗

Stochastic Learning Approach for Binary Optimization: Application to Bayesian Optimal Design of Experiments

Here, we present a novel stochastic approach to binary optimization suited for optimal experimental design (OED) for Bayesian inverse problems governed by mathematical models such as partial differential equations. The OED utility function, namely, the regularized optimality criterion, is cast into a stochastic objective function in the form of an expectation over a multivariate Bernoulli distribution. The probabilistic objective is then solved by using a stochastic optimization routine to find an optimal observational policy. This formulation (a) is generally applicable to binary optimization problems with soft constraints and is ideal for OED and sensor placement problems; (b) does not require differentiability of the original objective function (e.g., a utility function in OED applications) with respect to the design variable, and thus it enables direct employment of sparsity-enforcing penalty functions such as $\ell_0$, without needing to utilize a continuation procedure or apply a rounding technique; (c) exhibits much lower computational cost than traditional gradient-based relaxation approaches; and (d) can be applied to both linear and nonlinear OED problems with proper choice of the utility function. The proposed approach is analyzed from an optimization perspective with detailed convergence analysis of the optimization approach and is also analyzed from a machine learning perspective with correspondence to policy gradient reinforcement learning. The approach is demonstrated numerically by using an idealized two-dimensional Bayesian linear inverse problem and validated by extensive numerical experiments carried out for sensor placement in a parameter identification setup.

97 MATHEMATICS AND COMPUTING↗

Towards Intelligent Control for Next Generation Aircraft

NASA Aeronautics Subsonic Fixed Wing Project is focused on mitigating the environmental and operation impacts expected as aviation operations triple by 2025. The approach is to extend technological capabilities and explore novel civil transport configurations that reduce noise, emissions, fuel consumption and field length. Two Next Generation (NextGen) aircraft have been identified to meet the Subsonic Fixed Wing Project goals - these are the Hybrid Wing-Body (HWB) and Cruise Efficient Short Take-Off and Landing (CESTOL) aircraft. The technologies and concepts developed for these aircraft complicate the vehicle s design and operation. In this paper, flight control challenges for NextGen aircraft are described. The objective of this paper is to examine the potential of state-of-the-art control architectures and algorithms to meet the challenges and needed performance metrics for NextGen flight control. A broad range of conventional and intelligent control approaches are considered, including dynamic inversion control, integrated flight-propulsion control, control allocation, adaptive dynamic inversion control, data-based predictive control and reinforcement learning control.

Acosta, Diana Michelle↗

Autonomous Output‐Oriented Aerosol Jet Printing Enabled by Hybrid Machine Learning

Additive manufacturing (AM) is rapidly revolutionizing modern manufacturing with recent progress in advanced printing methods and improved properties of printed materials. However, traditional AM methods are limited by their input‐oriented nature, which demands tedious trial‐and‐error tuning of printing parameters to achieve desired output properties. Here, in this work, an output‐oriented artificial intelligence‐integrated AM (AIAM) method is reported that enables an user to specify desired output properties while the printer autonomously discovers the optimal input printing parameters by integrating hybrid machine learning models and in situ measurements. Based on a predictive mapping between the input printing parameters and the output properties of interests established with <20 experiments designed by active learning, inverse design tasks are performed to intelligently generate the printing parameter settings that lead to desired outcomes using reinforcement learning. This method is demonstrated by autonomous aerosol jet printing (AJP) of conductive polymer films and achieving user‐defined electrical resistances with an ultralow error of 3.7%. The AIAM method, with its output‐oriented nature, holds the potential to significantly improve the autonomy, predictability, efficiency, and accessibility of the AM processes, which will unlock new possibilities in the autonomous and intelligent printing of a broad range of functional materials and devices.

36 MATERIALS SCIENCE↗

Artificial intelligence-driven approaches for materials design and discovery

Materials design is an important component of modern science and technology, yet traditional approaches rely heavily on trial and error and can be inefficient. Computational techniques, enhanced by modern artificial intelligence, have reshaped the landscape of designing new materials. Among these approaches, inverse design has shown great promise in designing materials that meet specific property requirements. Here, in this Review, we present key computational advances in materials design over the past few decades. We follow the evolution of relevant materials design techniques, from high-throughput forward machine learning methods and evolutionary algorithms, to advanced artificial intelligence strategies such as reinforcement learning and deep generative models. We highlight the paradigm shift from conventional screening approaches to inverse generation driven by deep generative models. Finally, we discuss current challenges and future perspectives of materials inverse design. This Review may serve as a brief guide to the approaches, progress and outlook of designing future functional materials with technological relevance.

computational methods↗

End-to-end optimization for battery materials and molecules by combining graph neural networks and reinforcement learning

The National Renewable Energy Laboratory (NREL), together with the Colorado School of Mines (CSM) and Colorado State University (CSU), has developed a machine learning-enhanced approach to design new battery materials. Currently, such materials are designed in part via numerous expensive high-fidelity computational simulations that predict the performance of a given composition. Even with computational screening tools, the vast landscape of possible molecular or crystal structures exceeds current and future computational capacity. Improving the efficiency by which new materials can be optimized will therefore disrupt the cost, risk, and time required to bring new energy solutions to the marketplace. Predicting the properties of an organic molecule or periodic crystalline material given its structure has grown increasingly common. These approaches leverage large-scale computational and experimental databases and ML approaches such as graph neural networks. The inverse design problem of finding a material that possesses desired properties is substantially more challenging, since enumerating all valid material structures is not feasible. In this project, we leveraged recent success in reinforcement learning to efficiently navigate this high-dimensional search space. Just as algorithms can find the optimal chess moves from nearly limitless options, we train an approach to evolve a simple starting structure into a complex structure that possess the desired properties. Our solution has been demonstrated by applying it to two related design application tasks for short- and long-term energy storage, respectively: (1) the design of solid-state ion conductors and (2) the design of organic redox-active materials. The project has resulted an open-source software library for material design, documented examples of applying the library to both organic and inorganic material optimization, and peer-reviewed publications detailing the data, computational models, and resulting candidate materials.

25 ENERGY STORAGE↗