Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “REINFORCEMENT”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Post-Deployment Characterization of Glass Fiber-Reinforced Thermoset and Thermoplastic Composite Tidal Turbine Blades

In 2021, the National Renewable Energy Laboratory (NREL) supported Verdant Power with the most successful tidal energy deployment in U.S. history. Three of their Gen5d 5 m turbines were deployed as part of the Roosevelt Island Tidal Energy project. Initially, the three rotors initially deployed were manufactured from glass fiber-reinforced epoxy composites. Midway through the deployment, one rotor was replaced with one manufactured at NREL. The new rotor utilized a novel infusible thermoplastic resin system. Since the deployment, one epoxy rotor and one thermoplastic rotor were returned to NREL for continued materials and manufacturing research. The two rotors underwent full-scale structural testing before being sectioned and cut into specimens for a variety of manufacturing quality tests, thermomechanical characterization, and evaluation of material performance in marine environments to understand the key differences between the fiberglass-reinforced epoxy and Elium composites used for the respective rotors. Matrix burn-off tests showed that the Elium blades had a considerably higher fiber volume fraction compared to the epoxy blades (61% vs. 49%). Environmental aging of the specimens showed that the epoxy laminates absorbed more water over the conditioning period; however, it was determined that the Elium laminates had higher diffusion coefficients, so they initially absorbed water faster. Finally, one full epoxy blade and one full Elium blade were conditioned at ambient temperatures for up to 11 months, while periodic mass measurements were taken. The datasets were extrapolated to assume a full 20-year operational life span, and it was determined that the blades would not reach full saturation during that time span.

composite manufacturing↗

Post-Deployment Characterization of Glass Fiber-Reinforced Thermoset and Thermoplastic Composite Tidal Turbine Blades

In 2021, the National Renewable Energy Laboratory (NREL) supported Verdant Power with the most successful tidal energy deployment in U.S. history. Three of their Gen5d 5 m turbines were deployed as part of the Roosevelt Island Tidal Energy project. Initially, the three deployed rotors were manufactured from glass fiber-reinforced epoxy composites. Midway through the deployment, one rotor was replaced with one manufactured at NREL. The new rotor utilized a novel infusible thermoplastic resin system (Elium from Arkema). Since the deployment, one epoxy rotor and one thermoplastic rotor were returned to NREL for continued materials and manufacturing research. The two rotors underwent full-scale structural testing before being sectioned and cut into specimens for a variety of manufacturing quality tests, thermomechanical characterization, and evaluation of material performance in marine environments to understand the key differences between the fiberglass-reinforced epoxy and Elium composites used for the respective rotors. Matrix burn-off tests showed that the Elium blades had a considerably higher fiber volume fraction compared to the epoxy blades (61% vs. 49%). Environmental aging of the specimens showed that the epoxy laminates absorbed more water over the conditioning period; however, it was determined that the Elium laminates had higher diffusion coefficients, so they initially absorbed water faster. Finally, one full epoxy blade and one full Elium blade were conditioned at ambient temperatures for up to 11 months, while periodic mass measurements were taken. The datasets were extrapolated to assume a full 20-year operational life span, and it was determined that the blades would not reach full saturation during that time span.

composite manufacturing↗

CDRL: Certification-Driven Reinforcement Learning for Neutrino Flavor Model Discovery

Many scientific discovery problems require searching combinatorial hypothesis spaces under complex domain constraints. Reinforcement learning (RL) offers a promising approach, but existing methods rely on scalar rewards that provide limited information about why candidate solutions fail, leading agents to repeatedly explore invalid regions. We introduce Certification-Driven Reinforcement Learning (CDRL), a framework that leverages structured feedback from symbolic reasoning tools. When a candidate violates domain constraints, these tools produce certificates identifying the actions responsible for failure. CDRL converts these certificates into reusable constraints that eliminate classes of invalid solutions and guide exploration toward valid regions. We evaluate CDRL on neutrino flavor model discovery in theoretical particle physics, where the hypothesis space exceeds $10^{26}$ possible models, and compare it with the state-of-the-art RL approach previously used for this task. Across three theory spaces, CDRL achieves up to 1.95$\times$ higher valid model rates and up to 6.33$\times$ higher neutrino model rates while evaluating up to 4$\times$ fewer candidates. We further extract 40 interpretable rules from search trajectories using a post-hoc decision-tree framework and show that reusing them as soft constraints yields gains of up to 2$\times$ in valid model rates and 3$\times$ in neutrino model discovery across all three theory spaces. These results suggest that CDRL uncovers reusable structure in combinatorial search spaces and provides a general framework for scientific model discovery.

Jha, Piyush [Georgia Tech., Atlanta; Georgia Tech]↗

A Physics-Informed Reinforcement Learning Framework for Economic-Thermal Co-Optimization of Crypto Mining Data Centers: Preprint

The rapid expansion of cryptocurrency mining has created a new class of high-density data centers characterized by extreme thermal flux and high sensitivity to volatile economic markets. Traditional thermal management strategies, typically reliant on rule-based control, maintain static setpoints that fail to account for fluctuating electricity prices and cryptocurrency values - factors critical to mining profitability. To address this, we present a physics-informed reinforcement learning (PIRL) framework for economic-thermal co-optimization in crypto mining data centers. This framework consists of a proximal policy optimization (PPO) agent, a virtual testbed powered by high-fidelity physics-based models, and an interactive frontend dashboard. The PPO agent is trained using the virtual testbed and strict hardware safety limits. This physics-informed approach allows the agent to learn a stochastic policy that dynamically balances mining revenue against operational costs by co-optimizing HVAC cooling setpoints and IT computational hashrate. The simulation results demonstrate that the integrated framework achieved an 8.62% increase in net operational profit compared to traditional baseline strategies while strictly adhering to safety-critical temperature constraints (coolant supply temperature < 32 degrees C). This work provides a scalable template for the deployment of reinforcement learning in mission critical facilities where economic volatility and physical safety must be managed simultaneously.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Nuclear microreactor transient and load-following control with deep reinforcement learning

The economic feasibility of nuclear microreactors will depend on minimizing operating costs through advancements in autonomous control, especially when these microreactors are operating alongside other types of energy systems (e.g., renewable energy). This study explores the application of deep reinforcement learning (RL) for real-time drum control in microreactors, exploring performance in regard to load-following scenarios. By leveraging a point kinetics model with thermal and xenon feedback, we first establish a baseline using a single-output RL agent, then compare it against a traditional proportional–integral–derivative (PID) controller. This study demonstrates that RL controllers, including both single- and multi-agent RL (MARL) frameworks, can achieve similar or even superior load-following performance as traditional PID control across a range of load-following scenarios. In short transients, the RL agent was able to reduce the tracking error rate in comparison to PID by one half to one third. Over extended 300-minute load-following scenarios in which xenon feedback becomes a dominant factor, PID maintained better accuracy, but RL still remained within a 1% error margin despite being trained only on short-duration scenarios. This highlights RL’s strong ability to generalize and extrapolate to longer, more complex transients, affording substantial reductions in training costs and reduced overfitting. Furthermore, when control was extended to multiple drums, MARL enabled independent drum control as well as maintained reactor symmetry constraints without sacrificing performance---an objective that standard single-agent RL could not learn. We also found that, as increasing levels of Gaussian noise were added to the power measurements, the RL controllers were able to maintain lower error rates than PID, and to do so with at least 10% and upwards of 150% less control effort. These findings illustrate RL's potential for autonomous nuclear reactor control, laying the groundwork for future integration into high-fidelity simulations and experimental validation efforts.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Optimization of the FRIB beam dump: a hybrid genetic algorithm and reinforcement learning approach

The operational envelope of high-power-density systems, such as particle accelerators and advanced nuclear energy systems, is critically constrained by the need to manage extreme thermal loads. To address this, we present a novel hybrid optimization framework combining a genetic algorithm (GA) with a soft actor-critic (SAC) deep reinforcement learning agent. This framework was applied to a practical high-heat-flux problem: redesigning the beam dump at the Facility for Rare Isotope Beams (FRIB) for a power upgrade from 20 kW to 50 kW. The resulting design, validated by three-dimensional conjugate heat transfer simulations, suppresses hazardous hot spots and yields a markedly more uniform temperature distribution. This provides a robust operating margin, increasing the average power-handling capability by 72% relative to the current design, demonstrating the framework’s potential to solve complex thermal management challenges in both accelerator technology and advanced nuclear systems.

Accelerator↗

Visibility-enhanced model-free deep reinforcement learning algorithm for voltage control in realistic distribution systems using smart inverters

Increasing integration of distributed solar photovoltaic (PV) into distribution networks could result in adverse effects on grid operation. Traditional model-based control algorithms require accurate model information that is difficult to acquire and thus are challenging to implement in practice. Here, this paper proposes a surrogate model-enabled grid visibility scheme to empower deep reinforcement learning (DRL) approach for distribution network voltage regulation using PV inverters with minimal system knowledge. In contrast to existing DRL methods, this paper presents and corroborates the adverse impact of missing load information on DRL performance and, based on this finding, proposes a surrogate model methodology to impute load information utilizing observable data. Additionally, a multi-fidelity neural network is utilized to construct the DRL training environment, chosen for its efficient data utilization and enhanced robustness to data uncertainty. The feasibility and effectiveness of the proposed algorithm are assessed by considering DRL testing across varying degrees of observable load information and diverse training environments on a realistic power system.

14 SOLAR ENERGY↗

Development of algorithms for augmenting and replacing conventional process control using reinforcement learning

Here, this work seeks to allow for the online operation and training of model-free reinforcement learning (RL) agents but limit the risk to system equipment and personnel. The parallel implementation of RL alongside more conventional process control (CPC) allows for the RL algorithm to learn from CPC. The past performance of both methods are assessed on a continuous basis allowing for a transition from CPC to RL and, if needed, transitioning back to CPC from RL. This allows for the RL algorithm to slowly and safely assume control of the process without significant degradation in control performance. It is shown that the RL can derive a near optimal policy even when coupled with a suboptimal CPC. It is also demonstrated that the coupled RL-CPC algorithm learns at a faster rate than traditional RL methods of exploration while the algorithm’s performance does not deteriorate below CPC, even when exposed to an unknown operating condition.

30 DIRECT ENERGY CONVERSION↗

Optimizing on-ramp merging for connected and automated vehicles: A hierarchical approach using deep reinforcement learning and optimal control

On-ramp merging for Connected and Automated Vehicles (CAVs) presents significant challenges in dynamic traffic environments. Traditional methods and recent learning-based approaches often fail to simultaneously address decision-making complexity and execution precision under fluctuating conditions. This study introduces a novel hierarchical framework that combines: (1) a high-level Deep Reinforcement Learning (DRL) module that coordinates merging sequences through Virtual Traffic Signals (VTS) with Yield/Green phases and (2) a low-level optimal controller generating collision-free speed trajectories via pseudospectral convex optimization. A convolutional autoencoder compresses high-dimensional traffic states to enhance responsiveness. Extensive simulations demonstrate a 12.5% improvement in mainline throughput a 28% reduction in emergency braking events, and 31.66% lower fuel consumption compared to baseline methods. Furthermore, the framework’s effectiveness in coordinating CAV merges highlights its potential for real-world deployment. Future work will extend validation to multi-lane scenarios with mixed traffic and large-scale multiple merging points.

Connected and automated vehicles↗

Energy performance evaluation of the ASHRAE Guideline 36 control and reinforcement learning–based control using field measurements

This study evaluates the energy performance of ASHRAE Guideline 36–compliant control (ASHRAE 36 control) and reinforcement learning (RL)–based control through experimental field tests and a simulation study. Three field tests were conducted at Oak Ridge National Laboratory’s commercial building test facility in Oak Ridge, Tennessee: a baseline with a baseline conventional control, a test with ASHRAE 36 control, and a test with RL-based control. The selected ASHRAE 36 controls were trim and respond control, as well as variable air volume (VAV) box control. We compared the measured supply air temperature of the rooftop unit, VAV box supply air temperature, and VAV box supply airflow rate across the three test cases. The field data indicated that ASHRAE 36 controls operated as specified by ASHRAE Guideline 36. Based on these data, ASHRAE 36 control achieved a 45 % reduction in hourly averaged HVAC energy consumption compared with the baseline, and RL-based control achieved a 66 % reduction. These potential annual energy savings were confirmed using a calibrated whole-building energy model. Compared with the baseline, ASHRAE 36 control reduced HVAC energy consumption by 42 %, and RL-based control achieved a 54 % reduction. Furthermore, RL-based control reduced total HVAC energy consumption by 21 % more than ASHRAE 36 control.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Topology-driven compressive behavior of Inconel 718 lattice structures with Z-strut reinforcement fabricated by laser powder bed fusion

This study investigates the compressive deformation behavior and mechanical performance of Inconel 718 lattice structures fabricated by laser powder bed fusion (LPBF). Four unit-cell topologies—BCC, FCC, BCCZ, and FCCZ—were designed with a unit-cell size of 3 mm and fabricated under identical process conditions to isolate the effect of topology. Measured relative densities ranged from 14.65% to 17.72%. Compressive testing showed that Z-strut-reinforced topologies (BCCZ: 54.6 MPa; FCCZ: 80.2 MPa) exhibited higher strength than their unreinforced counterparts, which may be associated with mixed-mode deformation behavior enabled by the vertically aligned Z-struts. Finite element simulations and Digital Image Correlation (DIC) analysis support the observation of a transition from node-dominated deformation in BCC/FCC to mixed-mode deformation in BCCZ/FCCZ. These findings suggest that unit-cell topology is a key design variable for tailoring deformation mechanisms in LPBF lattice structures.

36 MATERIALS SCIENCE↗

A safe reinforcement learning algorithm for supervisory control of power plants

Traditional control theory-based methods require tailored engineering for each system and constant fine-tuning. In power plant control, one often needs to obtain a precise representation of the system dynamics and carefully design the control scheme accordingly. Model-free Reinforcement learning (RL) has emerged as a promising solution for control tasks due to its ability to learn from trial-and-error interactions with the environment. It eliminates the need for explicitly modeling the environment’s dynamics, which is potentially inaccurate. However, the direct imposition of state constraints in power plant control raises challenges for standard RL methods. To address this, we propose a chance-constrained RL algorithm based on Proximal Policy Optimization for supervisory control. Our method employs Lagrangian relaxation to convert the constrained optimization problem into an unconstrained objective, where trainable Lagrange multipliers enforce the state constraints. In conclusion, our approach achieves the smallest distance of violation and violation rate in a load-follow maneuver for an advanced Nuclear Power Plant design.

constrained optimization↗

Deep reinforcement learning for optimal control of induction welding process

Optimizing induction welding (IW) process parameters for the application of joining thermoplastic composites is challenging as it requires achieving complex spatiotemporal thermal characteristics along the weld-line to obtain desired weld quality. We formulate an optimal control problem which captures these requirements and seeks to optimize the IW coil speed using a fast-acting dynamic IW process model. We develop a novel Deep Reinforcement Learning (DRL) framework to solve this computationally challenging control problem and demonstrate via simulation study that the learned DRL feedback control policy results in better spatiotemporal thermal characteristics as compared to the current state-of-the-art.

36 MATERIALS SCIENCE↗

Development and assessment of hierarchical multi-reward reinforcement learning based potential for silicene with state-of-the-art models

We develop a new interatomic force field for Silicene, a 2D material with a buckled hexagonal lattice structure with high polymorphism. We introduce new parameterizations of a Tersoff model using a hierarchical multi-reward reinforcement learning (RL) methodology coupled with a continuous Monte Carlo Tree Search optimization. Our model significantly outperforms existing methods by enhancing the accuracy of predictions for the structural and thermodynamic properties of seven silicene polymorphs-including structure, energy, equation of state, elasticity, and phonon dispersion-when compared to established models. We further make a comprehensive comparison of the various models in predicting the mechanical and thermal properties of silicene. We trace the origin of the improved performance to the description of the angular dependence in the bond-order term, suggesting that modifying the angular terms in short-range models is essential to capture the structural diversity in low dimensional systems.

2D materials↗

Optimizing processing conditions for additively reinforced thermoforming (ART) in convergent manufacturing

This study utilized additively reinforced thermoforming (ART) to enhance the thermomechanical properties of polyethylene terephthalate glycol (PETG) sheet. ART materials were produced by overprinting PETG/carbon fiber filament (PETG/CF) on neat PETG sheets at varying conditions. The mechanical properties of the PETG sheet, PETG/CF, and ART materials were assessed, showing that ART exhibited superior tensile strength and modulus of elasticity. The tensile strength and modulus in the x-direction for ART at 265°C were 57.32 ± 2.9 MPa and 3.41 ± 0.4 GPa, respectively, compared to 49.1 ± 0.5 MPa and 1.92 ± 0.09 GPa for neat PETG. Microstructural analysis revealed strong interfacial adhesion between layers, while thermogravimetric analysis (TGA), differential scanning calorimetry (DSC), and heat deflection temperature analysis provided insights into the ART material's thermoforming behavior, aiding design optimization for enhanced stiffness, reduced necking, and improved customization. In conclusion, this information can be used to design for the thermoforming operation.

Additive reinforcement↗

An analysis of physics limited dispatch of nuclear renewable integrated energy systems using deep reinforcement learning and dynamic modeling

Previous approaches to dispatching nuclear integrated energy systems (NIES) have focused on the profitability and flexibility of these systems to operate on energy grids with highly variable pricing. However, due to the complexity involved in modeling and designing these systems, there has been less emphasis on ensuring that these dispatch strategies are physically achievable. It is imperative to develop methods that allow the system to remain within the desired NIES operating conditions and perform this based on realistic limited forecasted information. This research employs next generation artificial intelligence, namely deep reinforcement learning (DRL), and a dynamic system model written in Modelica to find a safe and profitable dispatch strategy for a solar nuclear hybrid design. The DRL agent is shown to find a novel dispatch strategy that manages both power ramping and power levels while respecting operational limits. This DRL-based dispatch is compared to other dispatching strategies including an optimal design solution from mixed integer linear programming (MILP). It is found that incorporating the physics of such a tightly coupled NIES limits the profitability of the MILP-based dispatch strategy. As a result, the MILP solution overestimates the design’s generated revenue. In contrast, DRL significantly reduces the number of breaches of safe operational conditions during energy arbitrage while maintaining profitability. Furthermore, this work paves the way for a more detailed assessment of NIES profitability and could be used to aid operator decisions on future NIES projects.

14 - SOLAR ENERGY↗

Acetolysis for Epoxy-Amine Carbon Fibre-Reinforced Polymer Recycling

Carbon fibre-reinforced polymers (CFRPs) are used in many applications in the global energy transition, including for lightweighting aircraft and vehicles and in wind turbine blades, shipping containers and gas storage vessels1,2,3,4. Given the high cost and energy-intensive manufacture of CFRPs5,6,7, recycling strategies are needed that recover intact carbon fibres and the epoxy-amine resin components. Here we show that acetic acid efficiently depolymerizes both aliphatic and aromatic epoxy-amine thermosets used in CFRPs to recoverable monomers, yielding pristine carbon fibres. Deconstruction of materials from multiple sectors demonstrates the broad applicability of this approach, providing clean fibres from 2 h reactions. The optimal conditions were scaled to 80.0 g of post-consumer CFRPs, and demonstrative composites were fabricated from the recycled carbon fibres, which were recycled two more times, maintaining their strength throughout. Process modelling and techno-economic analysis, with feedstock cost informed by wind turbine blade waste generation8, indicates this method is cost effective, with a minimum selling price of US$1.50 per kg for recycled carbon fibres whereas life cycle assessment shows process greenhouse gas emissions around 99% lower than virgin carbon fibre production. Overall, this approach could enable recycling of industrial CFRPs as it provides clean, mechanically viable recycled carbon fibres and recoverable resin monomers from the thermoset.

09 BIOMASS FUELS↗

Entanglement engineering of optomechanical systems by reinforcement learning

Entanglement is fundamental to quantum information science and technology, yet controlling and manipulating entanglement—so-called entanglement engineering—for arbitrary quantum systems remains a formidable challenge. There are two difficulties: the fragility of quantum entanglement and its experimental characterization. We develop a model-free deep reinforcement-learning (RL) approach to entanglement engineering, in which feedback control together with weak continuous measurement and partial state observation is exploited to generate and maintain desired entanglement. We employ quantum optomechanical systems with linear or nonlinear photon–phonon interactions to demonstrate the workings of our machine-learning-based entanglement engineering protocol. In particular, the RL agent sequentially interacts with one or multiple parallel quantum optomechanical environments, collects trajectories, and updates the policy to maximize the accumulated reward to create and stabilize quantum entanglement over an arbitrary amount of time. The machine-learning-based model-free control principle is applicable to the entanglement engineering of experimental quantum systems in general.

97 MATHEMATICS AND COMPUTING↗