Engineering PapersSearch

SEARCH · Engineering Papers

Results for “REINFORCEMENT”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Additively Reinforced Thermoformable PETG Composite Sheets for Improved Structural Efficiency

Thermoforming of short-fiber reinforced thermoplastic sheets offers a viable pathway for producing lightweight composite components; however, inherent anisotropy in fiber-reinforced sheets can limit structural performance under multidirectional loading. In this work, short carbon fiber, glass fiber, and hybrid fiber–reinforced PETG sheets were evaluated as candidate feedstock materials for thermoforming, with flexural and tensile testing performed both along the primary fiber direction and in the off-axis orientation to establish baseline stiffness, strength, and anisotropy. As expected, short carbon fiber PETG exhibited the highest stiffness and strength in the primary fiber direction, while all systems showed reduced performance in the off-axis direction. This off-axis performance reduction provides clear justification for the use of additive reinforcement when such thermoformed sheets are intended for structural applications. The intended manufacturing sequence involves thermoforming the reinforced sheet first, followed by the application of additively manufactured lattice reinforcement; therefore, the reinforcement strategy does not impose limitations on sheet formability during thermoforming. Post-forming lattice reinforcement significantly reduced load-normalized displacement by approximately 95–99% relative to non-reinforced sheets and improved weight-normalized stiffness by ~70%. These findings demonstrate that geometry-driven additive reinforcement can effectively compensate for off-axis property reductions in thermoformed PETG composites, enabling enhanced multidirectional structural performance without compromising manufacturability.

Talabi, Isaac [ORNL] (ORCID:0000000340215594)

Additively Manufactured Carbon Fiber-Reinforced Siliconized Silicon Carbide Composites Using Carbon Fiber-Reinforced Poly-Ether-Ether-Ketone (PEEK) as a Precursor

Herein, we report a method to additively manufacture carbon fiber-reinforced siliconized silicon carbide composites. The process involves the pyrolysis of a 3D-printed carbon fiber-reinforced poly-ether-ether-ketone (PEEK) composite to produce a porous carbon fiber-reinforced carbon matrix composite preform, which is subsequently infiltrated with molten silicon to obtain a carbon fiber-reinforced siliconized silicon carbide composite. A key aspect of the method is limiting polymer melt flow during pyrolysis of PEEK, which is achieved by thermally annealing the 3D-printed carbon fiber-reinforced PEEK preform in air at a temperature below PEEK’s melting temperature. Rheological and differential scanning calorimetry (DSC) measurements demonstrate that the thermal annealing treatment altered the melting behavior of PEEK, while NMR and FTIR measurements provided a mechanistic explanation for the structural changes responsible for the behavior. It was also found that dimensional changes during pyrolysis were anisotropic with greater shrinkage in the stacking direction of the material.

Yoon, Bola [ORNL] (ORCID:0000000260875373)

Barrier reinforcement for enhanced perovskite solar cell stability under reverse bias

Stability of perovskite solar cells (PSCs) under light, heat, humidity and their combinations have been notably improved recently. However, PSCs have poor reverse-bias stability that limits their real-world application. Here we report a systematic study on the degradation mechanisms of p-i-n structure PSCs under reverse bias. The oxidation of iodide by injected holes at the cathode side initialize the reverse-bias-induced degradation, then the generated neutral iodine oxidizes metal electrode such as copper, followed by drift of Cu+ into perovskites and its reduction by injected electrons, resulting in localized metallic filaments and thus device breakdown. A reinforced barrier with combined lithium fluoride, tin oxide and indium tin oxide at the cathode side reduces device dark current and avoids the corrosion of Cu0. It dramatically increases breakdown voltage to above -20 V and improved the T90 lifetime of PSCs to ~1,000 h under -1.6 V. The modified minimodule also maintained over 90% of its initial performance after 720 h of shadow tests.

degradation mechanisms

Reinforcement Learning-Based Oscillation Dampening: Scaling Up Single-Agent Reinforcement Learning Algorithms to a 100-Autonomous-Vehicle Highway Field Operational Test

In this article, we explore the technical details of the reinforcement learning (RL) algorithms that were deployed in the largest field test of automated vehicles designed to smooth traffic flow in history as of 2023, uncovering the challenges and breakthroughs that come with developing RL controllers for automated vehicles. We delve into the fundamental concepts behind RL algorithms and their application in the context of self-driving cars, discussing the developmental process from simulation to deployment in detail, from designing simulators to reward function shaping. We present the results in both simulation and deployment, discussing the flow-smoothing benefits of the RL controller. From understanding the basics of Markov decision processes to exploring advanced techniques such as deep RL, our article offers a comprehensive overview and deep dive of the theoretical foundations and practical implementations driving this rapidly evolving field. We also showcase real-world case studies and alternative research projects that highlight the impact of RL controllers in revolutionizing autonomous driving. From tackling complex urban environments to dealing with unpredictable traffic scenarios, these intelligent controllers are pushing the boundaries of what automated vehicles can achieve. Furthermore, we examine the safety considerations and hardware-focused technical details surrounding deployment of RL controllers into automated vehicles. As these algorithms learn and evolve through interactions with the environment, ensuring their behavior aligns with safety standards becomes crucial. Here, we explore the methodologies and frameworks being developed to address these challenges, emphasizing the importance of building reliable control systems for automated vehicles.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Reinforcement expectation in the honeybee ( Apis mellifera ): Can downshifts in reinforcement show conditioned inhibition?

When animals learn the association of a conditioned stimulus (CS) with an unconditioned stimulus (US), later presentation of the CS invokes a representation of the US. When the expected US fails to occur, theoretical accounts predict that conditioned inhibition can accrue to any other stimuli that are associated with this change in the US. Empirical work with mammals has confirmed the existence of conditioned inhibition. But the way it is manifested, the conditions that produce it, and determining whether it is the opposite of excitatory conditioning are important considerations. Invertebrates can make valuable contributions to this literature because of the well-established conditioning protocols and access to the central nervous system (CNS) for studying neural underpinnings of behavior. Nevertheless, although conditioned inhibition has been reported, it has yet to be thoroughly investigated in invertebrates. Here, we evaluate the role of the US in producing conditioned inhibition by using proboscis extension response conditioning of the honeybee (Apis mellifera). Specifically, using variations of a “feature-negative” experimental design, we use downshifts in US intensity relative to US intensity used during initial excitatory conditioning to show that an odorant in an odor–odor mixture can become a conditioned inhibitor. We argue that some alternative interpretations to conditioned inhibition are unlikely. However, we show variation across individuals in how strongly they show conditioned inhibition, with some individuals possibly revealing a different means of learning about changes in reinforcement. We discuss how the resolution of these differences is needed to fully understand whether and how conditioned inhibition is manifested in the honeybee, and whether it can be extended to investigate how it is encoded in the CNS. It is also important for extension to other insect models. In particular, work like this will be important as more is revealed of the complexity of the insect brain from connectome projects.

60 APPLIED LIFE SCIENCES

Graph reinforcement learning for exploring model spaces beyond the standard model

We present a methodology for performing scans of beyond the standard model (BSM) parameter spaces with reinforcement learning. We identify a novel procedure using graph neural networks that is capable of exploring spaces of models without the user specifying a fixed particle content, allowing broad classes of BSM models to be explored—in theory, the technique is applicable to nearly any model space with a prespecified gauge group. We provide a generic procedure by which a suitable graph grammar can be developed for any BSM model that features user-specified symmetry groups and a finite number of different possible particle species, the use of which is applicable to a variety of machine learning tasks over the actions of BSM theories beyond our particular reinforcement learning use case. As a proof of concept, we construct the graph grammar for theories with vectorlike leptons that may or may not be charged under a dark U ( 1 ) group, inspired by portal matter extensions of the sub-GeV vector portal/kinetic mixing simplified dark matter models. We then use this graph grammar to create a reinforcement learning environment tasked with creating models with these vectorlike leptons that are consistent with a list of a variety of precision observables. The reinforcement learning agent succeeds in developing models that can address the observed muon anomalous magnetic moment discrepancy while remaining consistent with flavor violation and electroweak precision observables, including both constructions that have previously been studied as well as new models that have not, to our knowledge, previously been identified. By inspecting the resulting ensembles of models that the agent produces and experimenting with different configurations for our reinforcement learning environment and graph grammar, we also infer various lessons about the development of these environments that can be transferable to reinforcement learning scans of more complicated model spaces and comment on future directions for the development of this technique into a more mature tool. Published by the American Physical Society 2025

Wojcik, George N.

Life cycle assessment of coir fiber-reinforced composites for automotive applications

Past decades have seen an increasing prevalence of natural fiber-reinforced composites (NFRCs) due to growing conscientiousness around sustainability and a push towards vehicle lightweighting. The environmentally friendly and sustainable claims of NFRCs need to be validated due to their large variability and variety, particularly where material substitutions are concerned, such as in substituting glass fiber with natural fiber. Additionally, the objective of this work is to determine the cumulative energy demand (CED) and greenhouse gas emissions (GHG) associated with an automotive part (of volume 0.001 m3) made from 40 wt% coir fiber-reinforced polypropylene (PP) and compared with a similar part made from 40 wt% glass fiber reinforced PP. SimaPro v. 9.0.0.49 was used for the analysis, whereas inventory data were collected from databases, such as Ecoinvent 3, Transportation Energy Databook, Greet model 2022, and published papers. The results showed that CED and GHG associated with the coir fiber-reinforced composite part were lower than the glass fiber-reinforced composite part for both cradle-to-gate (~34–40%) and cradle-to-grave (excluding end-of-life) (~24%) analysis.

36 MATERIALS SCIENCE

A Generation-Storage Coordination Dispatch Strategy for Power System Based on Causal Reinforcement Learning

In the backdrop of global energy transformation, power systems integrating high proportions of renewable energy sources are facing unprecedented challenges in operational stability and dispatch efficiency. To address these challenges, this study introduces a generation-storage coordination real-time dispatch strategy based on Causal Power System Dynamic Reinforcement Learning (CPSDRL). Diverging from traditional reinforcement learning approaches, CPSDRL innovatively incorporates causal inference within the state prediction model - the crux of model-based reinforcement learning - thereby establishing the Power Causal Dynamic Model (PCDM). Assisted by the prior knowledge of power systems, the model significantly enhances prediction accuracy and reliability through a two-stage training process. Utilizing PCDM, this study further applies a direct policy search algorithm to optimize the real-time dispatch strategy. Experimental results indicate that the proposed method improves the stability of generation-storage coordination real-time dispatch and exhibits competitive advantages in sample efficiency and computational speed, compared to traditional model-based and model-free reinforcement learning algorithms. This method is expected to enhance the practicality and adaptability of causal reinforcement learning techniques in power system scheduling and control.

causal reinforcement learning

Polymer-fiber-reinforced polymers with enhanced interfacial bonding between polypropylene fiber and polyethylene matrix

Self-reinforced composites (SRCs) consist of reinforcing fibers and a base matrix made of the same thermoplastic polymer, offering lightweight, recyclability, and sustainability benefits. However, limited research exists on composites where the reinforcing thermoplastic polymer fibers differ from the base thermoplastic matrix. Here, this study focuses on investigating the mechanical behavior of such composites and exploring different surface modification methods to enhance the fiber/matrix interfacial bonding using polypropylene fibers and a polyethylene matrix as an example. It is shown that surface treatment with a commercial adhesion promoter containing n-butyl acetate significantly improves the interfacial shear strength between polypropylene fibers and the polyethylene matrix, increasing it by 145% compared to other methods investigated. Additionally, increasing the length of the embedded polymer fiber in the matrix leads to a notable increase in specific interfacial energy. Consequently, the thermoplastic polymer-fiber-reinforced polymers (PFRPs) using surface-treated woven polypropylene fabrics and a polyethylene matrix exhibit a 20% higher tensile strength and a 65% higher toughness compared to non-treated PFRPs. This study also shows that specific mechanical properties (normalized by the composite density) of the investigated woven PFRPs are similar to those of non-treated SRCs under uni-axial tension. Particularly, their ductility outperforms carbon-/glass-/aramid-fiber-reinforced polymers by at least 6 times at a same fiber volume fraction. The investigation of such composites and the exploration of surface modification methods present important progress in the field of thermoplastic PFRPs, which serve as a solution for addressing concerns related to recyclability and sustainability.

Fiber pull-out

Graphene reinforced UHMWPE fibers

Thermoplastic polymers are increasingly used in electric vehicles, hydrogen fuel cell vehicles, and other decarbonization applications due to their lightweight and formability. Higher-strength polymers are needed to supplant metals in the vehicle structure, thereby reducing mass and improving efficiency. Ultra-high molecular weight polyethylene (UHMWPE) fibers possess one of the highest strength-to-weight ratios of technical polymers, and further improvement via reinforcement by nanofillers, such as graphene, will expand their performance envelope. Here, in this work, UHMWPE/graphene nanocomposite fibers were gel spun and characterized for their morphological, microstructural, thermal, and mechanical properties. The addition of a low fraction of graphene improved the tensile strength of the fibers by 25% and tensile modulus by 32%. Differential scanning calorimetry showed an increase in melting temperature and degree of crystallinity, which indicates improved coordination of the molecular chains induced by the addition of graphene. The reinforcement also affected the cross-sectional shape of the fibers; the aspect ratio of the fibers’ elliptical shape declined with increasing graphene content showing the skeletal effect of the graphene nanofillers in the polymer matrix. The reinforcing effect of graphene declined above a threshold concentration, and theoretical modeling was applied to demonstrate that increased agglomerates led to reduced properties. This work demonstrates a simple, effective method to produce graphene-reinforced UHMWPE fibers and lays a foundation for understanding the potential for leveraging graphene to form ultra-high-performance nanocomposite fibers for myriad engineering applications.

36 MATERIALS SCIENCE

Comprehensive assessment of deep reinforcement learning approaches for economic dispatch in nuclear-driven microgrids

As the electrical grid integrates more variable renewable energy sources such as wind and solar, the demand for distributed and flexible systems to address this increased variability becomes critical. Nuclear-driven microgrids provide a promising solution by offering stable generation to complement intermittent renewables, ensuring grid reliability and operating efficiency. This paper proposes a recurrent deep reinforcement learning framework for optimal economic dispatch in a nuclear-powered microgrid integrating renewable energy sources, small modular reactors, battery storage systems, and balance-of-plant dynamics. A three-agent control architecture is developed, where demand and renewable energy agents act as forecasters, and a reinforcement learning-based dispatch agent performs real-time energy allocation. A nonlinear programming formulation is first used to generate an optimal baseline for benchmarking. The proposed dispatch controller, based on Proximal Policy Optimization enhanced with Long Short-Term Memory networks, exploits temporal correlations in system dynamics by taking advantage of the time series used as inputs to improve policy robustness under uncertainty. Comparative analysis against established deep reinforcement learning methods, including Proximal Policy Optimization with a feedforward architecture, Soft Actor-Critic, and Twin Delayed Deep Deterministic Policy Gradient, demonstrates superior performance. Numerical results indicate that the proposed controller achieves a 0.39% cost reduction relative to the nonlinear programming benchmark and outperforms other learning-based methods by generating additional revenue of up to 0.35%. All reinforcement learning controllers compute dispatch actions in less than 0.3 s, resulting in a computational speedup of more than three orders of magnitude over the nonlinear programming baseline. The findings of this paper highlight their applicability for real-time operation and control in nuclear-integrated microgrids under volatile operating conditions.

24 POWER TRANSMISSION AND DISTRIBUTION

Reinforcement Learning for Anomaly Detection in Nuclear Power Plant Operation and Maintenance

In nuclear power plants (NPPs), timely identification of sensor and human errors is critical to ensure safe and efficient plant operations. Anomaly detection models can be employed for this task. However, traditional anomaly detection approaches may have high dependency on labeled datasets and struggle with adaptability in complex, dynamic environments. Reinforcement learning (RL) has demonstrated significant potential in fault diagnosis and anomaly detection; however, its application to anomaly detection in NPPs remains a relatively underexplored research direction. Hence, to address this gap, in this study, we present a novel physics-informed reinforcement learning model, PIRL-AD: Physics-Informed Reinforcement Learning for Anomaly Detection, that integrates domain knowledge from calorimetric equations into the RL framework for enhanced sensor and human error anomaly detection. We evaluate the performance of PIRL-AD against a non-physics informed RL benchmark and a support vector machine (SVM) on data collected from a forced flow loop testbed. Experimental results suggest that PIRL-AD outperforms other baselines on a range of anomalous datasets that include both sensor and human-induced anomalies across key performance metrics, statistically outperforming the RL and SVM benchmarks with respect to geometric mean (respectively, 92.96% vs. 91.06% vs. 83.01%) and F1-score (respectively, 89.23% vs. 86.98% vs. 77.01%). Furthermore, the findings suggest the potential of physics-integrated reinforcement learning models for enhanced anomaly detection performance in NPPs.

Reinforcement learning

Evaluation of desiccation cracking characteristics of inorganic micro-fiber-reinforced engineered barrier material (IMEBM) for geological repository

Abstract Buffer material is crucial for the engineering barrier system to dispose of high-level radioactive waste in a geological repository. A reliable buffer material should be able to maintain good sealing characteristics and minimize desiccation cracking. In this study, the effectiveness of inorganic fiber-reinforced engineering barrier material in reducing desiccation cracks in bentonite was evaluated via desiccation tests, image analysis, and air permeability tests. The effects of fiber type (E-glass fiber and basalt fiber) and fiber content (0.0%, 0.5%, 1.0%, 1.5%, 3.0%, and 5.0% of dry weight of the bentonite) on the development of desiccation cracks in the fiber–bentonite mixtures with the same given initial moisture content were evaluated. The results indicated that the addition of fibers could significantly reduce the crack size and area in bentonite during the drying. Basalt fibers showed a slightly better reinforcement effect than E-glass fibers when the fiber content was lower than 3.0%. The addition of fibers prevented the development of penetrating cracks and significantly reduced the permeability of the bentonite after drying. The permeabilities of basalt fiber- and E-glass fiber-reinforced bentonite composites with 3% reinforcement were 5.81 × 10 –11 m 2 and 7.24 × 10 –11 m 2 , respectively, which were 64 and 51 times smaller than that of pure bentonite. X-ray–CT observation of the internal structure of the samples after drying showed that the addition of fibers significantly changed the crack morphology and potentially increased the tortuosity.

Feng, Yuan

Additive manufacturing of carbon fiber-reinforced thermoset composites via in-situ thermal curing

Fiber-reinforced polymer composites are lightweight structural materials widely used in the transportation and energy industries. Current approaches for the manufacture of composites require expensive tooling and long, energy-intensive processing, resulting in a high cost of manufacturing, limited design complexity, and low fabrication rates. Here, we report rapid, scalable, and energy-efficient additive manufacturing of fiber-reinforced thermoset composites, while eliminating the need for tooling or molds. Use of a thermoresponsive thermoset resin as the matrix of composites and localized, remote heating of carbon fiber reinforcements via photothermal conversion enables rapid, in-situ curing of composites without further post-processing. Rapid curing and phase transformation of the matrix thermoset, from a liquid or viscous resin to a rigid polymer, immediately upon deposition by a robotic platform, allows for the high-fidelity, freeform manufacturing of discontinuous and continuous fiber-reinforced composites without using sacrificial support materials. This method is applicable to a variety of industries and will enable rapid and scalable manufacture of composite parts and tooling as well as on-demand repair of composite structures.

36 MATERIALS SCIENCE

Design Trade-Offs in Composite Fuel Cell Membranes: Effects of Reinforcement and Chemical Additives

Perfluorosulfonic acid (PFSA) membranes are critical components in proton exchange membrane fuel cells, where performance depends on balancing ionic conductivity, mechanical durability, and chemical stability. This study characterizes a composite membrane (NC700) featuring PFSA-impregnated expanded polytetrafluoroethylene (ePTFE) reinforcement and cerium-based radical scavengers, benchmarked against unreinforced NR211. Complementary techniques, including electron microscopy, X-ray scattering, infrared spectroscopy, thermogravimetric analysis, and dynamic mechanical analysis, identify the structural and compositional strategies employed in NC700. Water sorption isotherms reveal lower water uptake for NC700 across all conditions, attributed to reinforcement and cerium incorporation. Reinforcement reduces in-plane swelling from 11% to 2.1% at 90% RH, confirming strong swelling anisotropy, while maintaining mechanical properties at elevated temperatures. While the ionic conductivity of NC700 is approximately 10% lower than that of NR211, the reduced thickness yields a 40% decrease in calculated area-specific resistance, suggesting the composite architecture can favorably shift the conductivity-stability trade-off. The composite structure also reduces gas permeability, indicating potential for improved separator function alongside favorable transport properties. Systematic deconvolution of reinforcement and additive contributions shows that conductivity losses from cerium incorporation are largely offset by gains from the lower equivalent-weight polymer, providing quantitative relationships that may guide composite membrane design for fuel cells and other electrochemical applications.

25 ENERGY STORAGE

Virtual to Physical: Reinforcement Learning to Optimize SNS Particle Accelerator Controls

Complex accelerators must have control systems that can handle dynamic nonlinear environments. This makes traditional control methods unsuitable as they can struggle to adapt to these uncertainties. This provides an ideal environment for reinforcement learning algorithms as they are adaptable and generalizable. We present a reinforcement learning pipeline that can effectively handle the dynamics of a complex accelerator. We test and prove our pipelines capabilities on multiple environments including the Spallation Neutron Source (SNS) and the Beam Test Facility (BTF) at Oakridge National Lab (ORNL). Due to the limited time available to train an online algorithm like reinforcement learning on a real accelerator, we utilize a virtual twin accelerator (VIRAC) developed by ORNL to pretrain the policy and show its ability to converge in the virtual environment. We then test the adaptability of the pretrained RL model by applying it on the real accelerator and comparing the results. Utilizing our Scientific Optimization and Controls Toolkit (SOCT) and open-source standards such as Gymnasium we create and solve for a MEBT orbit correction problem in the SNS and an emittance maximization problem in the BTF. We show how Twin Delayed Deep Deterministic Policy Gradient (TD3) can solve this optimization environment in the virtual accelerator and transfer this policy onto the real accelerator for inference and model retraining. We show how reinforcement learning can be utilized as a control system for complex accelerators and provide a model pipeline for how an implementation performs and can be adapted to new accelerator control problems.

Kasparian, Armen [Thomas Jefferson National Accele

Model-based Hierarchical Reinforcement Learning for Improved Physical Security Design: A Prototype

Prior work in FY24 developed an adversarial AI agent aid in path analysis of physical protection systems. This agent, trained using a model-based reinforcement learning algorithm, was able to successfully learn the most vulnerable path in facilities. It was able to extend the current state of practice for physical protection design by exhibiting dynamic behavior based on current environmental conditions. Whereas PathTrace largely performs a static, graph-based analysis, the AI agent was able to make decisions based on relative position in the facility, current conditions (was the adversarial agnet discovered?), and proximity to secondary targets. The agent demonstrated some novel capabilities, but had limitations that need to be resolved before it can be used for production purposes. For example, the adversarial agent generalizes poorly and takes a relatively long time to train. Nonetheless, there is still considerable promise for developing the adversarial agent further in order to explore even richer, more dynamic behaviors (e.g., adversary motivations, environmental debris, and more). This work considers a complementary idea; development of a planning agent. The planning agent is envisioned as an auto-complete-like tool that can help accelerate security system design by human experts. The agent would respect existing barriers and sensors placed by a human expert while offering cost-effective suggestions (i.e., implicitly balancing effectiveness with cost) to improve the design. The goal is for this agent to be part of an expert’s toolbox, not to totally upend the current state-of-practice, or to displace human experts. The ultimate goal would be concurrent training of both the adversarial and planning agent together, to learn entirely through self-play. This would represent an entirely new way of performing system deign. We selected a hierarchical, model-based reinforcement learning algorithm to serve as the planning agent. This is an extension of concepts used in the prior FY24 adversarial agent work. There, we had a single agent acting an environment. Here, we have two different sub-agents (policies), working together, to form a complete agent. There is a manager policy, which can select abstract goals on slower time scales, and a worker, which performs primitive actions to reach goals selected by the manager. It is worth noting that this class of algorithm is challenging to work with. From our understanding, our work is one of the first successful uses of model-based reinforcement learning (MBRL) in nuclear energy1 , and likely the first hierarchical model-based reinforcement learning application in nuclear energy. Further, this work is one of the first known attempts to apply AI to perform a design tasks in nuclear energy. Consequently, there were significant implementation challenges and the bulk of the work was focused on successful implementation and algorithm design. The results presented here are very low technology readiness level as a consequence of the lack of related literature, but still represent a significant step forward in the pursuit of applied AI for design.

42 ENGINEERING

Explainable and Differentiable Reinforcement Learning for Multi-objective Optimization in Particle Accelerators

Operating particle accelerators involves optimizing multiple goals simultaneously, which can be challenging due to trade-offs among objectives. While evolutionary algorithms like the genetic algorithm (GA) have been used for various Multi-Objective Optimization (MOO) tasks, they are not inherently suited for complex control problems. This talk highlights two variations of Reinforcement Learning (RL) for concurrently optimizing heat load and trip rates at the Continuous Electron Beam Accelerator Facility (CEBAF). The problem involves strict constraints on individual states, actions, and overall energy requirements of the beam. First, this talk highlights how differentiability can be harnessed through a Deep Differentiable Reinforcement Learning (DDRL) approach to address MOO issues within particle accelerators. We examine the DDRL method alongside Model Free Reinforcement Learning (MFRL), GA, and Bayesian Optimization (BO). The performance of these methods is assessed by generating a Pareto-front for two objectives. Our findings indicate that DDRL excels in handling high-dimensional problems more effectively than MFRL, BO, and GA. Next, we will show integration of explainable physics-based constraints into RL algorithms to enhance trans- parency and trust in decision-making processes by enabling users to verify that agents adhere to established physical principles. This surrogate function can be modeled using neural networks or sparse dictionary mod- els. By examining the mathematical form of the learned constraint function, we are able to confirm the agent has learned to use the established physics of each environment provided but the surrogate model. In addi- tion, we find that the introduction of a mathematical functional dictionary based surrogate model enables our reinforcement learning algorithms to reliably converge for difficult high-dimensional accelerator controls environments.

Rajput, Kishansingh [Thomas Jefferson National Acc