Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “RL”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

81 records · Page 5

Planetary Boundary-Layer Height (PBLHT) Value-Added Product: Remote-Sensing Retrievals

The planetary boundary layer (PBL) is fundamental to numerous atmospheric processes, including aerosol mixing and transport, cloud evolution, and precipitation formation. A critical parameter in these studies is the PBL height (PBLHT). This vertical depth is essential for characterizing PBL structures in numerical simulations and serves as a primary metric for estimating flux exchanges between the Earth’s surface and the atmosphere. Radiosonde (SONDE) observations provide high-vertical-resolution measurements of temperature and moisture profiles and are widely used to estimate PBLHT (Liu and Liang 2010, Seidel et al. 2010). The U.S. Department of Energy Atmospheric Radiation Measurement (ARM) User Facility’s PBLHT value-added product (VAP) for radiosonde measurements, known as PBLHTSONDE, applies three commonly used methods—the Heffter (1980) method, the Liu and Liang (2010) method, and the bulk Richardson number approach (Seibert et al. 2000)—to derive PBLHT. The PBLHTSONDE VAP operates routinely at ARM observatories and mobile facilities, with data available from the ARM Data Center shortly after sounding observations are collected (Sivaraman et al. 2013). However, radiosonde observations are limited by their low temporal resolution. Most stations launch soundings only twice daily, which constrains the ability to investigate and characterize the temporal evolution of the PBL using radiosonde data alone. The use of continuous remote-sensing observations provides high temporal resolution of PBLHT estimates. These observations include aerosol lidars (Dang et al. 2019, Su et al. 2020), Doppler lidar (DL; Tucker et al. 2009, Krishnamurthy et al. 2021), and water vapor and/or temperature lidars and radiometers (Turner et al. 2014). These observations provide valuable data on the PBL’s thermodynamic properties (e.g., water vapor and/or temperature lidars and radiometers), dynamic properties (e.g., DL), and distribution of tracer substances (e.g., aerosol lidars), all of which can be used to estimate PBLHT. ARM developed PBLHT estimates from the micropulse lidar (MPL; PBLHTMPL), Doppler lidar (PBLHTDL), and combined Raman lidar (RL)/atmospheric emitted radiance interferometer (AERI) thermodynamic profiles (PBLHTTHERMO). Each estimate captures different physical characteristics of the boundary layer—aerosol tracers, vertical velocity turbulence, and thermodynamic structure—and exhibits distinct strengths and limitations depending on the PBL regime and time of day. In addition, the ARM ceilometer (CEIL) provides three potential PBLHT candidates derived from the vendor's built-in algorithm. Building on these individual retrievals, ARM developed the PBLHTBEML VAP, which combines the four remote-sensing-based estimates with ancillary meteorological variables using the machine learning approach of Zhang et al. (2025) to produce a best-estimate PBLHT at 10-minute resolution.

54 ENVIRONMENTAL SCIENCES↗

Reinforcement Learning Control for Enhancing Marine Hydrokinetic Turbine Energy Generation

This paper proposes a reinforcement learning-based method to maximize power generation for a direct-drive marine hydrokinetic turbine. A high levelized cost of energy (LCOE) is preventative in the widespread adoption of many marine energy conversion technologies. A straightforward way to reduce LCOE is to increase conversion efficiency and ensure maximum energy generation. The proposed method utilizes a damping control methodology, varying applied generator torque via a linear relationship between the applied damping coefficient and rotor speed. A state-action-reward-state-action (SARSA) algorithm has been used to learn the optimal control action for a given flow velocity. The proposed SARSA methodology uses Gaussian radial basis functions to create a three-dimensional surface to estimate the relationship between damping coefficient, incoming flow velocity, and coefficient of power (C p ). Here, the SARSA algorithm was compared against a baseline optimal tip speed ratio controller over a year-long flow velocity case profile while considering the effects of biofouling on the turbine system, where the proposed RL method generated 0.92% more energy than the baseline.

Damp↗

Unconventional compute methods and future challenges for superconducting digital computing

Superconducting digital computing (SDC) based on Josephson junctions (JJs) offers significant potential for enhancing compute throughput and reducing energy consumption compared to conventional room-temperature CMOS-based approaches. Current superconducting logic families exhibit diverse characteristics in clocking strategies, power management, and information encoding techniques. This paper reviews recent advancements in unconventional computing methods specifically designed for superconducting digital circuits, emphasizing temporal computing and pulse-train representations. Notable techniques include race logic (RL), temporal pulse train computing (U-SFQ), and temporal multipliers, each offering unique performance and area advantages suited to superconducting implementations. Additionally, this paper reviews innovations in superconducting coarse-grain reconfigurable architectures (CGRA), superconducting-specific on-chip communication architectures, cryogenic sensor interfaces, and quantum computing control electronics. Finally, we highlight research challenges that should be addressed to facilitate the widespread adoption of superconducting digital computing.

EDA tools↗

ARM-IRL: Adaptive Resilience Metric Quantification Using Inverse Reinforcement Learning

The resilience of safety-critical systems is gaining importance due to the rise in cyber and physical threats, especially within critical infrastructure. Traditional static resilience metrics may not capture dynamic system states, leading to inaccurate assessments and ineffective responses to cyber threats. This work aims to develop a data-driven, adaptive method for resilience metric learning. We propose a data-driven approach using inverse reinforcement learning (IRL) to learn a single, adaptive resilience metric. The method infers a reward function from expert control actions. Unlike previous approaches using static weights or fuzzy logic, this work applies adversarial inverse reinforcement learning (AIRL), training a generator and discriminator in parallel to learn the reward structure and derive an optimal policy. The proposed approach is evaluated on multiple scenarios: optimal communication network rerouting, power distribution network reconfiguration, and cyber–physical restoration of critical loads using the IEEE 123-bus system. The adaptive, learned resilience metric enables faster critical load restoration in comparison to conventional RL approaches.

97 MATHEMATICS AND COMPUTING↗

GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning

Large language models (LLMs) are increasingly adapted to downstream tasks via reinforcement learning (RL) methods like Group Relative Policy Optimization (GRPO), which often require thousands of rollouts to learn new tasks. We argue that the interpretable nature of language often provides a much richer learning medium for LLMs, compared to policy gradients derived from sparse, scalar rewards. To test this, we introduce GEPA (Genetic-Pareto), a prompt optimizer that thoroughly incorporates natural language reflection to learn high-level rules from trial and error. Given any AI system containing one or more LLM prompts, GEPA samples trajectories (e.g., reasoning, tool calls, and tool outputs) and reflects on them in natural language to diagnose problems, propose and test prompt updates, and combine complementary lessons from the Pareto frontier of its own attempts. As a result of GEPA's design, it can often turn even just a few rollouts into a large quality gain. Across six tasks, GEPA outperforms GRPO by 6% on average and by up to 20%, while using up to 35x fewer rollouts. GEPA also outperforms the leading prompt optimizer, MIPROv2, by over 10% (e.g., +12% accuracy on AIME-2025), and demonstrates promising results as an inference-time search strategy for code optimization. We release our code at https://github.com/gepa-ai/gepa.

97 MATHEMATICS AND COMPUTING↗

Best estimate of the planetary boundary layer height from multiple remote sensing measurements

Remote sensing measurements have been widely used to estimate the planetary boundary layer height (PBLHT). Each remote sensing approach offers unique strengths and faces different limitations. In this study, we use machine learning (ML) methods to produce a best-estimate PBLHT (PBLHT-BE-ML) by integrating four PBLHT estimates derived from remote sensing measurements at the Department of Energy (DOE) Atmospheric Radiation Measurement (ARM) Southern Great Plains (SGP) observatory. Three ML models – random forest (RF) classifier, RF regressor, and light gradient-boosting machine (LightGBM) – were trained on a dataset from 2017 to 2023 that included radiosonde, various remote sensing PBLHT estimates, and atmospheric meteorological conditions. Evaluations indicated that PBLHT-BE-ML from all three models improved alignment with the PBLHT derived from radiosonde data (PBLHT-SONDE), with LightGBM demonstrating the highest accuracy under both stable and unstable boundary layer conditions. Feature analysis revealed that the most influential input features at the SGP site were the PBLHT estimates derived from (a) potential temperature profiles retrieved using Raman lidar (RL) and atmospheric emitted radiance interferometer (AERI) measurements (PBLHT-THERMO), (b) vertical velocity variance profiles from Doppler lidar (PBLHT-DL), and (c) aerosol backscatter profiles from micropulse lidar (PBLHT-MPL). The trained models were then used to predict PBLHT-BE-ML at a temporal resolution of 10 min, effectively capturing the diurnal evolution of PBLHT and its significant seasonal variations, with the largest diurnal variation observed over summer at the SGP site. We applied these trained models to data from the ARM Eastern Pacific Cloud Aerosol Precipitation Experiment (EPCAPE) field campaign (EPC), where the PBLHT-BE-ML, particularly with the LightGBM model, demonstrated improved accuracy against PBLHT-SONDE. Analyses of model performance at both the SGP and EPC sites suggest that expanding the training dataset to include various surface types, such as ocean and ice-covered areas, could further enhance ML model performance for PBLHT estimation across varied geographic regions.

Zhang, Damao [Pacific Northwest National Laborator↗

Learning to Trigger: Reinforcement Learning at the Large Hadron Collider

High-throughput scientific facilities such as the Large Hadron Collider depend on real-time event filtering (\textit{triggering}) under tight constraints on bandwidth, latency, and storage. In practice, trigger menus are largely static and hand-tuned and can become suboptimal as detector conditions, pileup, and background composition drift over time. We cast online threshold tuning as a sequential decision-making problem: a reinforcement learning agent ingests streaming summaries of recent rates and signal-sensitive features and updates trigger thresholds to maximize signal efficiency while tracking a target background rate within a tolerance band. We adapt Group-Filtered Policy Optimization (GFPO) to streaming control and introduce two variants (GFPO-F, GFPO-FR) that enforce background rate feasibility during training. On a benchmark that emulates realistic collider operation, we study two representative triggers: a total transverse energy ($H_{T}$) trigger sensitive to pileup variation, and an anomaly-detection (AD) trigger based on reconstruction loss for rare or non-standard signatures. On Monte Carlo streams, our agent increases the fraction of in-tolerance time intervals by 48% ($H_T$) and 28% (AD), with a cumulative gain of up to 2% in signal efficiency on those in-tolerance intervals. Transferring from simulation to \emph{real} collision data (CMS Run 283408), the same agent, without fine-tuning, achieves a 56% ($H_T$) and 28% (AD) in-tolerance improvement over baselines, with further signal-efficiency gain on both triggers. To our knowledge, this is the \emph{first} demonstration of RL-based trigger control on real Large Hadron Collider collision data. Code is available at https://github.com/Zixind/GFPO_LHC (see repo for details).

Ding, Zixin [Chicago U.]↗

Intern Poster Session 08/13: Autonomous Nuclear Robotics: Applications in nuclear waste inspection and hot cell experiments

The nuclear industry is experiencing renewed interest in autonomous robotics, yet most deployed systems remain teleoperated with limited autonomy. This work presents two contributions toward fully autonomous nuclear robotic systems: autonomous waste inspection at the Hanford Site and an autonomous hot cell laboratory framework. Inspections of Hanford's underground waste storage tanks are performed manually at significant cost and personnel exposure. We developed a reinforcement-learning (RL) training pipeline for a custom-built inspection arm. In parallel, we are designing an autonomous laboratory framework for post-irradiation examination in hot cells at the Specimen Preparation Laboratory (SPL) that integrates computer vision, task and motion planning, hardware execution, and operator-in-the-loop control. These systems demonstrate a path toward safer, more efficient nuclear operations by reducing human exposure while maintaining rigorous human oversight at critical decision points.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

From Sim to Real: A Pipeline for Training and Deploying Traffic Smoothing Cruise Controllers

Designing and validating controllers for connected and automated vehicles to enhance traffic flow presents significant challenges, from the complexity of replicating real-world stop-and-go traffic dynamics in simulation, to the intricacies involved in transitioning from simulation to actual deployment. In this work, we present a full pipeline from data collection to controller deployment. Specifically, we collect 772 km of driving data from the I-24 in Tennessee, and use it to build a one-lane simulator, placing simulated vehicles behind real-world trajectories. Using policy-gradient methods with an asymmetric critic, we improve fuel efficiency by over 10% when simulating congested scenarios. Our comprehensive approach includes reinforcement learning for controller training, software verification, hardware validation and setup, and navigating various sim-to-real challenges. Furthermore, we analyze the controller's behavior and wave-smoothing properties, and deploy it on four Toyota Rav4’s in a real-world validation experiment on the I-24. Lastly, we release the driving dataset, the simulator and the trained controller, to enable future benchmarking and controller design.

42 ENGINEERING↗