Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Inverse Reinforcement Learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

ARM-IRL: Adaptive Resilience Metric Quantification Using Inverse Reinforcement Learning

The resilience of safety-critical systems is gaining importance due to the rise in cyber and physical threats, especially within critical infrastructure. Traditional static resilience metrics may not capture dynamic system states, leading to inaccurate assessments and ineffective responses to cyber threats. This work aims to develop a data-driven, adaptive method for resilience metric learning. We propose a data-driven approach using inverse reinforcement learning (IRL) to learn a single, adaptive resilience metric. The method infers a reward function from expert control actions. Unlike previous approaches using static weights or fuzzy logic, this work applies adversarial inverse reinforcement learning (AIRL), training a generator and discriminator in parallel to learn the reward structure and derive an optimal policy. The proposed approach is evaluated on multiple scenarios: optimal communication network rerouting, power distribution network reconfiguration, and cyber–physical restoration of critical loads using the IEEE 123-bus system. The adaptive, learned resilience metric enables faster critical load restoration in comparison to conventional RL approaches.

97 MATHEMATICS AND COMPUTING↗

MSD CoP Webinar: Modeling the Operations of Reservoir Systems with LLMs and Inverse Reinforcement Learning

Context: This panel featured three presentations centered on the common theme of applying LLMs and inverse reinforcement learning (IRL) to capture the complex human-environment interactions that are central to the operation of reservoir systems. Dr. Wyatt Arnold will kick off the webinar with a talk on how analyzing LLM chain-of-thought reasoning reveals sophisticated quantitative justification and risk awareness, showing promise as a bridge between quantitative models and value-driven water management decisions. Next, Dr. Matteo Giuliani will build on this with a discussion demonstrating that AI- and IRL-driven approaches can infer the trade-offs between flood control and water supply using historical observations. Finally, Dr. Rohan Singh Wilkho will close the webinar with a talk establishing IRL as a generalizable diagnostic tool for decoding decision-making in managed hydrologic and human-infrastructure systems. Across the three presentations, the application of LLMs and IRL opens new possibilities for the development of adaptive, transparent, and human-aware models supporting water management in an increasingly uncertain future. Presenters: Wyatt Arnold (Politecnico di Milano); Matteo Giuliani (Politecnico di Milano); Rohan Singh Wilkho (Cornell University) Moderator: Patrick M. Reed (MSD CoP Facilitation Team); Stefano Galelli (MSD CoP AI Working Group Co-Chair); David Gold (MSD CoP AI Working Group Co-Chair) This webinar was held on: June 23rd, 2026 from 12-1 PM EST.

Arnold, Wyatt [Politecnico di Milano]↗

Inverse reinforcement learning control for building energy management

Reinforcement learning (RL) based control is widely considered a promising approach in building automation and control as it has demonstrated the potential to deal with complex objectives in adjacent domains like robotics, autonomous vehicles, gaming applications, and advertisement recommendations. When applied to any environment, model-free RL learns to improve its control performance over time without requiring a control model, by receiving and then analyzing feedback from the building environment after each control action. Operational objectives are becoming increasingly complex through the simultaneous consideration of thermal comfort, carbon emissions, grid services, and indoor air quality. In this context, conventional rule-based control approaches are proving sub-optimal, mostly heuristic, and inadequate. The model-free and self learning nature of RL appears promising and attractive as it may address the scalability issues associated with advanced control approaches. However, it suffers from long training times and unstable control behavior during the early stages of its learning process, which makes it unsuitable to be applied directly to buildings. This paper addresses these issues using an inverse reinforcement learning approach (IRL), a technique utilized to learn the objective of a controller agent which is considered an expert in its respective domain. Here, we consider a rule-based control as the expert demonstrator. IRL is different from a direct imitation (i.e., direct mapping of states to actions) of control actions as it tries to find the underlying intent of an expert's policy, providing the controller with a better-generalized policy for unseen states or environments with slightly different dynamics. This approach propels the RL controller's policy to levels similar to or better than that of a rule-based policy before it starts learning by interacting with the building. This makes RL for building energy management applications more practical as it prevents the erratic and exploratory behavior in the initial training period, simultaneously speeding up the learning process when compared to applying an untrained RL agent directly to a building environment.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

A Bayesian Approach for Quantifying Data Scarcity when Modeling Human Behavior via Inverse Reinforcement Learning

Computational models that formalize complex human behaviors enable study and understanding of such behaviors. However, collecting behavior data required to estimate the parameters of such models is often tedious and resource intensive. Thus, estimating dataset size as part of data collection planning (also known as Sample Size Determination) is important to reduce the time and effort of behavior data collection while maintaining an accurate estimate of model parameters. In this paper, we present a sample size determination method based on Uncertainty Quantification (UQ) for a specific Inverse Reinforcement Learning (IRL) model of human behavior, in two cases: 1) pre-hoc experiment design—conducted in the planning stage before any data is collected, to guide the estimation of how many samples to collect; and 2) post-hoc dataset analysis—performed after data is collected, to decide if the existing dataset has sufficient samples and whether more data is needed. Here, we validate our approach in experiments with a realistic model of behaviors of people with Multiple Sclerosis (MS) and illustrate how to pick a reasonable sample size target. Our work enables model designers to perform a deeper, principled investigation of effects of dataset size on IRL.

97 MATHEMATICS AND COMPUTING↗

A Modified Maximum Entropy Inverse Reinforcement Learning Approach for Microgrid Energy Scheduling

Increasing popularity of integrating distributed energy resources (DERs) into the power system brings a challenge to optimize the microgrid dispatch policy. The reinforcement learning methods suffer from a long-time problem with the theoretical assumption of the objective/reward function for the microgrid system. Although the traditional inverse reinforcement learning (IRL) approaches can solve this problem to some extent, they encounter a limitation of complex computations for state visitation frequency in the large and continuous state space. To alleviate this limitation, we propose a modified maximum entropy IRL (MMIRL) method to extract the reward function from the expert demonstrations for solving the microgrid energy scheduling problem. The proposed MMIRL algorithm is promising in recovering the reward function and learning the dispatch policy compared to conventional approaches. Case studies are performed in an energy arbitrage problem and a microgrid system with DERs. Results substantiate that the proposed MMIRL approach can learn the dispatch policy with more than 99% efficiency and outperforms other comparative methods.

artificial intelligence, reinforcement learning, m↗

Inverse Reinforcement Learning based Bayesian Goal Inference Method for Early Nuclear Proliferation Detection

Traditional methods for detection of nuclear proliferation indicators are usually applied after nuclear proliferation has already occurred. There is a need to advance these methods to perform early detection of nuclear proliferation indicators. In this project, we formulated an early detection problem as a sequential, decision-making, goal inference problem based on research publications of authors, to determine whether it is possible to infer whether an author will publish on a research activity before it has occurred. To develop and test our approach, we selected a civil nuclear activity for our case study. We constructed a state-action-state transition graph from publications of authors associated with the activity and the co-authors of their publications, using titles, abstracts, and author publication sequences. We then used inverse reinforcement learning to model the goal-directed behavior of authors in trajectories that terminate at selected goal states. Using a Bayesian formulation, we computed the probability that authors would reach each selected state from partially observed trajectories of their state transitions in their research topic space. The state with the highest probability was selected as the most probable goal state. Based on our results, we found that 60% of the times we can infer the correct goal state early; sometimes the inference is either delayed, or multiple states could be inferred as goal states. Overall, our results show that it is possible to perform early detection of research activities of authors in a nuclear technology area. Further research is necessary to establish a more accurate understanding of how topic modeling, topic space grid discretization, and the extent of overlap among trajectories of different goal states, affect the goal inference results. The methods developed in this work may be used to enhance data-driven methods for early detection of nuclear proliferation indicators.

97 MATHEMATICS AND COMPUTING↗

Early Inference of Nuclear Technology-Directed Research Activities of Authors from Scientific Publications

Nuclear research articles can provide information about early nuclear proliferation indicators such as influential research entities and technology capability levels of a country, but detection of nuclear activities typically occurs after they have started. We investigate the extent to which nuclear research articles can be used to infer whether a research entity will acquire or develop a nuclear technology before it happens. Early detection of nuclear proliferation or technology development indicators from data is challenging due to partial observability, sparse and unlabeled information, and confounding signals from multiple concurrent activities. This paper presents the early detection problem as a sequential decision-making, goal inference problem, where the objective is to characterize and predict an individual’s, organization’s, or a country’s intent (unobserved goal-directed behavior) towards developing a nuclear capability from partially observed sequences of their research publications, using inverse reinforcement learning and Bayesian goal inference methods. A computational framework is presented, and its application demonstrated using 29,196 Scopus records for a case study related to a civil nuclear capability. The case study results serve as a proof-of-concept demonstration for inference of technology-directed research activity of authors who publish in the nuclear domain. The inference method, combined with advanced computing, may be used to assess and monitor activities pertaining to early developmental stages of a nuclear technology or capability, which in turn can help to identify and prioritize activities with nuclear proliferation potential for further investigation.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

RLMolLM: Reinforcement Learning-Enhanced Language Model Framework for Inverse Molecular Design

Inverse molecular design faces significant challenges due to vast chemical space and complex property requirements. While language models show promise for molecular generation, they struggle with validity, multi-property optimization, and structural constraints. This work presents RLMolLM, a reinforcement learning framework combining Proximal Policy Optimization (PPO) with genetic algorithms to address these limitations. Our approach optimizes multiple user-specified properties including quantitative estimates of drug-likeness (QED), synthetic accessibility (SA), and ADMET (absorption, distribution, metabolism, excretion, and toxicity) endpoints without requiring complete model retraining, while maintaining capability for scaffold-constrained generation where specific substructures must be preserved. We outperform state-of-the-art methods for molecular optimization, achieving best QED scores across GDB13, Moses, and Zinc datasets with up to 31% improvement over previous methods while maintaining excellent validity, uniqueness, and novelty metrics. For simultaneous multi-property optimization, our framework achieves substantial improvements in ADMET properties including 4.5-fold reduction in hERG toxicity and enhanced Caco-2 permeability compared to Moses dataset. Under structural constraints, the framework significantly improves molecular validity while preserving scaffolds and effectively optimizing properties. In conclusion, this versatile solution advances pharmaceutical and materials molecular design through effective integration of reinforcement learning and genetic algorithms with multi-property optimization and scaffold preservation.

Genetic algorithms↗

Machine learning enabled discovery of superhard and ultrahard carbon polymorphs

The demand for multifunctional materials has motivated the move from near-equilibrium materials to metastable i.e. out-of-equilibrium phases that can meet several desired target properties. The search for such metastable phases with exotic properties is non-trivial and often serendipitous. Inverse design approaches based on evolutionary search have been powerful tools, but such traditional searches have focused on identifying primarily stable and metastable materials with the lowest enthalpy. The inverse design of materials, with a focus on a desired property such as, for example, hardness is a challenging task because of the expensive computational cost involved in sampling multiple structures. The recent advances in machine learning have brought new powerful AI techniques to the forefront which can potentially revolutionize the inverse design and discovery of materials, especially metastable phases capable of meeting multifunctionality. Here, in this work, we develop and apply an automated reinforcement learning workflow for inverse design that integrates first principles physics and atomistic simulations with machine learning (ML), and high-performance computing to allow rapid exploration of the superhard and ultrahard metastable phases of Carbon. We demonstrate an automatic machine learning based inverse design workflow to map new undiscovered metastable states ranging from near equilibrium to those far-from-equilibrium that satisfy multiple property objectives, specifically bulk moduli, shear moduli and hardness. We create a comprehensive library of carbon stable and metastable phases with varying hardness and subsequently shortlist 10 top performing candidate carbon structures, including two newly reported phases, based on their hardness and characterize their temperature dependent mechanical properties. A neural network model is built using featurization of allotropes of carbon to predict the quasi-harmonic Gibbs free energies. The Gibbs free energies of the top performing phases are analyzed to get an estimate of the experimental synthesizability of these superhard and ultrahard carbon phases. In general, we show using machine learning based inverse design approaches how hitherto inaccessible metastable states can be identified and potentially synthesized to meet the demand for multifunctional materials.

Balasubramanian, Karthik [Univ. of Illinois, Chica↗

A Continuous Action Space Tree search for INverse desiGn (CASTING) framework for materials discovery

Abstract Material properties share an intrinsic relationship with their structural attributes, making inverse design approaches crucial for discovering new materials with desired functionalities. Reinforcement Learning (RL) approaches are emerging as powerful inverse design tools, often functioning in discrete action spaces. This constrains their application in materials design problems, which involve continuous search spaces. Here, we introduce an RL-based framework CASTING (Continuous Action Space Tree Search for inverse design), that employs a decision tree-based Monte Carlo Tree Search (MCTS) algorithm with continuous space adaptation through modified policies and sampling. Using representative examples like Silver (Ag) for metals, Carbon (C) for covalent systems, and multicomponent systems such as graphane, boron nitride, and complex correlated oxides, we showcase its accuracy, convergence speed, and scalability in materials discovery and design. Furthermore, with the inverse design of super-hard Carbon phases, we demonstrate CASTING’s utility in discovering metastable phases tailored to user-defined target properties and preferences.

36 MATERIALS SCIENCE↗

Stochastic Learning Approach for Binary Optimization: Application to Bayesian Optimal Design of Experiments

Here, we present a novel stochastic approach to binary optimization suited for optimal experimental design (OED) for Bayesian inverse problems governed by mathematical models such as partial differential equations. The OED utility function, namely, the regularized optimality criterion, is cast into a stochastic objective function in the form of an expectation over a multivariate Bernoulli distribution. The probabilistic objective is then solved by using a stochastic optimization routine to find an optimal observational policy. This formulation (a) is generally applicable to binary optimization problems with soft constraints and is ideal for OED and sensor placement problems; (b) does not require differentiability of the original objective function (e.g., a utility function in OED applications) with respect to the design variable, and thus it enables direct employment of sparsity-enforcing penalty functions such as $\ell_0$, without needing to utilize a continuation procedure or apply a rounding technique; (c) exhibits much lower computational cost than traditional gradient-based relaxation approaches; and (d) can be applied to both linear and nonlinear OED problems with proper choice of the utility function. The proposed approach is analyzed from an optimization perspective with detailed convergence analysis of the optimization approach and is also analyzed from a machine learning perspective with correspondence to policy gradient reinforcement learning. The approach is demonstrated numerically by using an idealized two-dimensional Bayesian linear inverse problem and validated by extensive numerical experiments carried out for sensor placement in a parameter identification setup.

97 MATHEMATICS AND COMPUTING↗

Autonomous Output‐Oriented Aerosol Jet Printing Enabled by Hybrid Machine Learning

Additive manufacturing (AM) is rapidly revolutionizing modern manufacturing with recent progress in advanced printing methods and improved properties of printed materials. However, traditional AM methods are limited by their input‐oriented nature, which demands tedious trial‐and‐error tuning of printing parameters to achieve desired output properties. Here, in this work, an output‐oriented artificial intelligence‐integrated AM (AIAM) method is reported that enables an user to specify desired output properties while the printer autonomously discovers the optimal input printing parameters by integrating hybrid machine learning models and in situ measurements. Based on a predictive mapping between the input printing parameters and the output properties of interests established with <20 experiments designed by active learning, inverse design tasks are performed to intelligently generate the printing parameter settings that lead to desired outcomes using reinforcement learning. This method is demonstrated by autonomous aerosol jet printing (AJP) of conductive polymer films and achieving user‐defined electrical resistances with an ultralow error of 3.7%. The AIAM method, with its output‐oriented nature, holds the potential to significantly improve the autonomy, predictability, efficiency, and accessibility of the AM processes, which will unlock new possibilities in the autonomous and intelligent printing of a broad range of functional materials and devices.

36 MATERIALS SCIENCE↗

Artificial intelligence-driven approaches for materials design and discovery

Materials design is an important component of modern science and technology, yet traditional approaches rely heavily on trial and error and can be inefficient. Computational techniques, enhanced by modern artificial intelligence, have reshaped the landscape of designing new materials. Among these approaches, inverse design has shown great promise in designing materials that meet specific property requirements. Here, in this Review, we present key computational advances in materials design over the past few decades. We follow the evolution of relevant materials design techniques, from high-throughput forward machine learning methods and evolutionary algorithms, to advanced artificial intelligence strategies such as reinforcement learning and deep generative models. We highlight the paradigm shift from conventional screening approaches to inverse generation driven by deep generative models. Finally, we discuss current challenges and future perspectives of materials inverse design. This Review may serve as a brief guide to the approaches, progress and outlook of designing future functional materials with technological relevance.

computational methods↗

End-to-end optimization for battery materials and molecules by combining graph neural networks and reinforcement learning

The National Renewable Energy Laboratory (NREL), together with the Colorado School of Mines (CSM) and Colorado State University (CSU), has developed a machine learning-enhanced approach to design new battery materials. Currently, such materials are designed in part via numerous expensive high-fidelity computational simulations that predict the performance of a given composition. Even with computational screening tools, the vast landscape of possible molecular or crystal structures exceeds current and future computational capacity. Improving the efficiency by which new materials can be optimized will therefore disrupt the cost, risk, and time required to bring new energy solutions to the marketplace. Predicting the properties of an organic molecule or periodic crystalline material given its structure has grown increasingly common. These approaches leverage large-scale computational and experimental databases and ML approaches such as graph neural networks. The inverse design problem of finding a material that possesses desired properties is substantially more challenging, since enumerating all valid material structures is not feasible. In this project, we leveraged recent success in reinforcement learning to efficiently navigate this high-dimensional search space. Just as algorithms can find the optimal chess moves from nearly limitless options, we train an approach to evolve a simple starting structure into a complex structure that possess the desired properties. Our solution has been demonstrated by applying it to two related design application tasks for short- and long-term energy storage, respectively: (1) the design of solid-state ion conductors and (2) the design of organic redox-active materials. The project has resulted an open-source software library for material design, documented examples of applying the library to both organic and inorganic material optimization, and peer-reviewed publications detailing the data, computational models, and resulting candidate materials.

25 ENERGY STORAGE↗

End-to-End Optimization for Battery Materials and Molecules by Combining Graph Neural Networks and Reinforcement Learning

The National Renewable Energy Laboratory (NREL), together with the Colorado School of Mines (CSM) and Colorado State University (CSU), has developed a machine learning-enhanced approach to the design of new battery materials. Currently, such materials are designed in part via numerous expensive high-fidelity computational simulations that predict the performance of a given composition. Even with computational screening tools, the vast landscape of possible molecular or crystal structures exceeds current and future computational capacity. Improving the efficiency by which new materials can be optimized will therefore disrupt the cost, risk, and time required to bring new energy solutions to the marketplace. Predicting the properties of an organic molecule or periodic crystalline material given its structure has grown increasingly common. These approaches leverage large-scale computational and experimental databases and ML approaches such as graph neural networks. The inverse design problem of finding a material that possesses desired properties is substantially more challenging, since enumerating all valid material structures is not feasible. In this project, we leveraged recent success in reinforcement learning to efficiently navigate this high-dimensional search space. Just as algorithms can find the optimal chess moves from nearly limitless options, we train an approach to evolve a simple starting structure into a complex structure that possess the desired properties. Our solution has been demonstrated by applying it to two related design application tasks for short- and long-term energy storage, respectively: (1) the design of solid-state ion conductors and (2) the design of organic redox-active materials. The project has resulted an open-source software library for material design, documented examples of applying the library to both organic and inorganic material optimization, and peer-reviewed publications detailing the data, computational models, and resulting candidate materials.

25 ENERGY STORAGE↗

A machine learning approach to determine the elastic properties of printed fiber-reinforced polymers

This work focuses on the simultaneous determination of the elastic constants and the fiber orientation state for a short fiber-reinforced polymer composite by performing a minimum of experimental tests. Here we introduce a methodology that enables the inverse determination of fiber orientation state and the in-situ polymer properties by performing tensile tests at the composite coupon level. We demonstrate the approach for the extrusion deposition additive manufacturing (EDAM) process to illustrate one application of the methodology, but the development is such that it can be applied to short fiber-reinforced polymer (SFRP) systems processed via other methods. Currently, developing composites additive manufacturing digital twins require extensive material characterization. In particular, the mechanical characterization of the orthotropic elastic properties of a composite involves extensive sample preparation and testing, therefore the elasticity tensor is generally populated using a micromechanics model. This, however, requires measuring the fiber orientation state in addition to knowing the constituent material properties. Experimentally measuring the fiber orientation state can be tedious and time consuming. Further, optical methods are limited to resolving the orientation of cylindrical fibers or cluster of non-cylindrical fibers, and computed tomography (CT) methods scan regions of volume that are much smaller than a full printed bead. Therefore, we propose a methodology, accelerated by machine learning, to identify the anisotropic mechanical properties and fiber orientation state at the same time. Early results show that inference of the fiber orientation and composite properties is possible with as few as three tensile tests. Our results show that a combination of the choice of the micromechanics model and reliable set of experiments can yield the nine elastic constants, as well as, the fiber orientation state.

36 MATERIALS SCIENCE↗

Bayesian sequential optimal experimental design for nonlinear models using policy gradient reinforcement learning

We present a mathematical framework and computational methods for optimally designing a finite sequence of experiments. This sequential optimal experimental design (sOED) problem is formulated as a finite-horizon partially observable Markov decision process (POMDP) under a Bayesian setting and with information-theoretic utilities. The formulation is general and may accommodate continuous random variables, non-Gaussian posteriors, and nonlinear forward models. The sOED design policy incorporates elements of feedback and lookahead simultaneously, and we show it to generalize the commonly-used batch and greedy design strategies. We solve for the sOED policy using the policy gradient (PG) method from reinforcement learning, and provide a derivation for the PG expression in the sOED context. Adopting an actor-critic approach, the policy and value functions are parameterized using deep neural networks and improved via PG estimates produced from simulated episodes of designs and observations. The new PG-sOED algorithm is first validated on a linear-Gaussian benchmark, and then compared against other design baselines on a sensor movement problem for contaminant source inversion in a convection-diffusion field. As a result, we provide explanation for the policy behaviors using knowledge of the underlying physical process.

97 MATHEMATICS AND COMPUTING↗

Machine learning-driven design and self-sensing capabilities of automotive bumper lattices for adaptive impact response

We present a novel approach to design an automotive bumper energy absorber using carbon fiber reinforced polymer composites, optimized to meet conflicting performance requirements for two distinct impact scenarios. The design must satisfy both a low-speed (2.5 mph) pendulum intrusion test, simulating vehicle-to-vehicle collisions, and a high-speed (25 mph) leg flexion test, replicating pedestrian impacts. These tests demand opposing deformation characteristics: high flexibility (deformation < 85 mm) for the former and high stiffness (deformation < 22 mm) for the latter. To address these contradictory requirements, we developed a machine learning (ML) framework for inverse optimization of lattice designs and material selection. Unlike traditional iterative design processes, our ML model directly outputs optimal design parameters and material choices based on target performance inputs. The energy absorber was fabricated using advanced additive manufacturing techniques, including extrusion deposition and digital light processing. The integration of carbon fibers provides multifunctionality to the bumper structure, enabling self-sensing capabilities through changes in electrical resistivity under compression. This electrical response demonstrates high repeatability under multiple cycles at 2% compression and exhibits distinct signatures during crack formation under high deformation. This research offers adaptive performance through innovative design methodologies and smart material integration. The approach has potential applications in various fields requiring adaptive energy absorption and real-time structural health monitoring.

Chawla, Komal [ORNL] (ORCID:0000000190327565)↗