Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “robust execution”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Scalable Federated Learning for Scientific Foundation Models on Leadership-Class Systems

Federated learning (FL) at leadership-class HPC systems remains largely unexplored, despite growing interest in deploying federated workflows on modern HPC systems. This paper provides the first system-level empirical characterization of federated fine-tuning of pretrained foundation models on an exascale supercomputer under a multi-node deployment. Using up to 96 concurrent FL clients deployed across Frontier nodes, we study the impact of client scale, model size, data heterogeneity, partial participation, and differential privacy on runtime, communication overhead, and convergence stability. Our results show that pretrained transformer models remain robust to heterogeneity, client dropout, and privacy noise, while system efficiency degrades rapidly with scale as synchronizat and orchestration dominate runtime. We further demonstrate that system-aware execution strategies, including intra-node aggregation and early aggregation, significantly reduce wall-clock time without degrading model quality. These findings establish a practical performance baseline and inform the design of communication-efficient FL systems on leadership-class HPC platforms.

Kotevska, Olivera [ORNL] (ORCID:0000000316772243)↗

EnergyPlus Model Context Protocol Server (EnergyPlus-MCP) v0.1

EnergyPlus-MCP is the first open-source Model Context Protocol server specifically designed for EnergyPlus building energy simulation. This innovative software enables AI assistants and other applications to interact programmatically with EnergyPlus through a standardized, secure interface, eliminating traditional technical barriers in building energy modeling. The software provides specialized tools across five functional domains: server management, model configuration and loading, comprehensive building component inspection, systematic model modification, and simulation execution with results visualization. Key features include automated HVAC system discovery and topology mapping, advanced schedule analysis, intelligent model validation, and interactive visualization capabilities. EnergyPlus-MCP's layered architecture ensures robust separation between protocol communication and domain expertise, enabling scalable deployment across organizations, educational institutions, and research teams. Unlike direct LLM approaches that suffer from inconsistent results and security gaps, EnergyPlus-MCP provides validated, reliable interactions while maintaining scientific rigor. This democratizes sophisticated building energy analysis, making EnergyPlus accessible to broader audiences through conversational interfaces and streamlined workflows.

Li, Han [Lawrence Berkeley National Laboratory (LB↗

Current Best Practices on Wildfire Risk Reduction for Electric Transmission and Distribution Systems

This report provides a set of best practices in response to Section 4(d) of Executive Order 14308, Empowering Commonsense Wildfire Prevention and Response. The information contained herein is expected to be used in conjunction with other materials at the Federal Energy Regulatory Commission Wildfire Risk Mitigation Technical Conference (Docket No. AD25-16-000), with possible direction to the North American Electric Reliability Corporation to take action. The objective of this report is to provide an overview of existing and emerging best practices currently employed or planned by utilities for wildfire mitigation, demonstrating how these efforts align with the executive order’s emphasis on reducing electric utility–caused wildfires while also balancing cost-effectiveness. Additionally, while most practices in utility-developed wildfire mitigation plans focus on risk reduction through robustness and operational reliability, this report also discusses best practices for resilience. The best practices are adopted from publicly available utility wildfire mitigation plans from the United States and Canada, recent findings from wildfire risk reduction research, and industry engagement.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Trajectory Shaper: A Solution for Disrupted Cooperative Adaptive Cruise Control

Cooperative adaptive cruise control (CACC) can effectively reduce energy consumption, alleviate traffic congestion, and enhance safety. However, communication-related constraints and uncooperative vehicle users can disrupt CACC during real-world operations, significantly undermining the putative benefits of CACC. To alleviate the negative impacts of disrupted CACC, this study develops the trajectory shaper (TS) methods as backup solutions for two scenarios: (i) communication between vehicles is infeasible, and vehicles execute adaptive cruise control (ACC) using local sensor measurements; (ii) follower vehicles reject forming a cooperative platoon and execute their local distributed controllers using the information attained via communication. When communication is infeasible, a distributed TS is devised on each vehicle to modify the sensor measurements, enabling safe and efficient ACC operations. When communication is available but uncooperative agents are involved, the lead vehicle of the platoon executes a centralized TS to modify the information shared with uncooperative agents, achieving optimal platoon-level performance. The centralized and distributed TSs are implemented based on the model predictive control algorithms to yield optimal modifications on input information. Robustness is also factored to tackle model uncertainties during TS operations to ensure safety and efficiency. Numerical experiments validate the control performance of the proposed TSs.

Zhou, Anye [ORNL] (ORCID:0000000301455579)↗

Repositioning Quantum Cellular Automata for Dependable Quantum-Classical Systems

Quantum Cellular Automata (QCA) provides a structured model of distributed quantum computation with inherent locality and regularity properties that are suited to dependable execution. However, QCA remain largely absent from discussions on reproducibility, fault management, and orchestration in heterogeneous quantum-classical systems. We propose a dual-axis framework that situates QCA within both computation and physical realizability, revealing regions where robust, scalable, and hardware-constrained quantum dynamics may reside. By revisiting prior results through the lens of reproducibility and architecture resilience, we suggest that QCA offers a potential substrate for benchmarking and system-level co-design.

Stapleton, Nicholas [ORNL] (ORCID:0000000335305325↗

An end-to-end workflow for executing a classically bootstrapped variational quantum algorithm on an academic quantum computer

Academic quantum computing platforms often face unique challenges in executing quantum workloads due to fragmented software environments and limited engineering support. Unlike commercial ecosystems, academic devices typically evolve without full-stack integration in mind, making it difficult to run complex applications—such as variational quantum algorithms (VQA)—reliably and efficiently. Issues such as incompatible software layers and lack of automated job management significantly increase the overhead of theory-experiment collaboration. To address these challenges, we develop a modular, end-to-end workflow that decouples application-layer code from low-level hardware control, automates circuit submission and result collection, and supports fine-grained circuit-level job scheduling and recovery. The architecture employs a dual-end application programming interface (API) design, enabling robust operation across unstable or resource-constrained hardware backends. For practical use, the framework is lightweight and user-friendly, allowing rapid prototyping of full-stack workflows using basic Python tools. We validate this workflow on a high-fidelity trapped-ion quantum computer by demonstrating a variational quantum eigensolver (VQE) experiment with a classically bootstrapped ansatz initialization technique. The system successfully executed over 60,000 circuits across multiple molecular test cases with minimal human intervention, highlighting the framework’s effectiveness in enabling reproducible, resilient quantum experimentation in academic settings.

Clifford↗

MVP: a modular viromics pipeline to identify, filter, cluster, annotate, and bin viruses from metagenomes

While numerous computational frameworks and workflows are available for recovering prokaryote and eukaryote genomes from metagenome data, only a limited number of pipelines are designed specifically for viromics analysis. With many viromics tools developed in the last few years alone, it can be challenging for scientists with limited bioinformatics experience to easily recover, evaluate quality, annotate genes, dereplicate, assign taxonomy, and calculate relative abundance and coverage of viral genomes using state-of-the-art methods and standards. Here, we describe Modular Viromics Pipeline (MVP) v.1.0, a user-friendly pipeline written in Python and providing a simple framework to perform standard viromics analyses. MVP combines multiple tools to enable viral genome identification, characterization of genome quality, filtering, clustering, taxonomic and functional annotation, genome binning, and comprehensive summaries of results that can be used for downstream ecological analyses. Overall, MVP provides a standardized and reproducible pipeline for both extensive and robust characterization of viruses from large-scale sequencing data including metagenomes, metatranscriptomes, viromes, and isolate genomes. As a typical use case, we show how the entire MVP pipeline can be applied to a set of 20 metagenomes from wetland sediments using only 10 modules executed via command lines, leading to the identification of 11,656 viral contigs and 8,145 viral operational taxonomic units (vOTUs) displaying a clear beta-diversity pattern. Further, acting as a dynamic wrapper, MVP is designed to continuously incorporate updates and integrate new tools, ensuring its ongoing relevance in the rapidly evolving field of viromics. MVP is available at https://gitlab.com/ccoclet/mvp and as versioned packages in PyPi and Conda.

59 BASIC BIOLOGICAL SCIENCES↗

ARCS: Agentic Retrieval-Augmented Code Synthesis with Iterative Refinement

Agentic Retrieval-Augmented Code Synthesis with Iterative RefinementIn supercomputing, efficient and optimized code generation is essential to leverage high-performance systems effectively. We have developed Agentic Retrieval-Augmented Code Synthesis (ARCS), an advanced framework for accurate, robust, and efficient code generation, completion, and translation. ARCS integrates Retrieval-Augmented Generation (RAG) with Chain-of-Thought (CoT) reasoning to systematically break down and iteratively refine complex programming tasks. An agent-based RAG mechanism retrieves relevant code snippets, while real-time execution feedback drives the synthesis of candidate solutions. This process is formalized as a state-action search tree optimization, balancing code correctness with editing efficiency. Evaluations on the Geeks4Geeks and HumanEval benchmarks demonstrate that ARCS significantly outperforms traditional prompting methods in translation and generation quality. By enabling scalable and precise code synthesis, ARCS offers transformative potential for automating and optimizing code development in supercomputing applications, enhancing computational resource utilization

Bhattarai, Manish [Los Alamos National Labs]↗

Docker Containers for MCNP ® Development

Containers are a revolutionary technology in software development and deployment that provides a lightweight, portable environment for ensuring consistency across multiple computing environments. In anticipation of the MCNP 6.3.1 release, two Docker container images have been released on DockerHub for general use. The MCNP source code is not included in the images, and users are still required to obtain it through RSICC. The images produced by Docker are compliant with the OCI (Open Container Initiative) standards, ensuring compatibility with other container engines such as Podman or Kubernetes’ CRI-O. Initially, the images are stored under the author’s personal space on DockerHub (docker.io/azukaitis), but they will be relocated to a dedicated MCNP group space once approved. In the future, they will also be available through the registry feature of the https://github.com/lanl/mcnp-containers project. The use of Docker provides a pre-configured environment for building and running MCNP, ensuring reproducibility of results across various host architectures. This significantly improves consistency when running MCNP on different systems. Notably, executables and installers from the Docker images have successfully passed the MCNP development branch testing suite on x86-64 architectures, including Windows, macOS, and Linux operating systems. Furthermore, testing has demonstrated compatibility with macOS Docker in emulation mode on the latest Apple Mac M2 Ultra hardware, ensuring robust support even on the latest platforms. In this document, we will provide a step-by-step guide to using the Docker images across multiple platforms. Additionally, we will present performance numbers for building and running the MCNP test suite.

97 MATHEMATICS AND COMPUTING↗

Workflows for Science: A comprehensive guide for ensemble workflow tools usage with applications on OLCF systems

The growing demand for robust computational and workflow environments for scientific applications and user communities at the Oak Ridge Leadership Computing Facility (OLCF) has prompted collaboration with ensemble tools development teams and facility users to produce this technical paper. We connect science applications to the RADICAL-Pilot (RP) workflow tool to execute ensemble instantiations using the Frontier supercomputer. The documented installation, usage, and execution demonstrates how RP streamlines scientific workflows at OLCF. We outline the specific steps OLCF users can follow to integrate this tool with their applications and advance their research. This document stands as a comprehensive guide to OLCF users of ensemble workflow tools with examples on real applications using the Frontier supercomputer.

97 MATHEMATICS AND COMPUTING↗

DEReliction: A Cybersecurity Vulnerability Assessment Methodology for Distributed Energy Resources

With the increasing integration of Distributed Energy Resources (DER) into the electric grid, maintaining grid reliability and resilience requires that these devices remain secure. This paper discusses a cybersecurity vulnerability assessment methodology that incorporates best practices from Sandia National Laboratories, SANS Institute, OWASP Foundation, and other web and Internet of Things (IoT) penetration testing (“pen testing”) programs, courses, and frameworks for assessing the security posture of devices. The methodology involves five sequential steps: (1) Collect Public Information, (2) Extract Hardware Details, (3) Inventory Software Components, (4) Identify Vulnerabilities, and (5) Test Vulnerabilities. Each step uncovers potential weaknesses in both hardware and software components of DER devices, considering adversary tactics, techniques, and procedures (TTPs), and potential attack vectors along the way. The results from the execution of this method on multiple residential- and small commercial-scale photovoltaic (PV) inverters reveled hardware and software vulnerabilities, which highlight the benefit of taking a methodical approach to discover vulnerabilities. While the specific vulnerability details are not shared here, a generalized overview of findings underscore the importance of robust security assessments for DER devices. Adoption of an assessment framework of this kind will identify and mitigate cybersecurity threats and bolster the resilience of DER-integrated electric grids.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Advanced Fuels Campaign Execution Plan

The Advanced Fuels Campaign (AFC) Execution Plan details the strategy, mission, scope, and goals—both near-term and long-term—along with the structure and organization of nuclear fuels and materials research, development, and demonstration (RD&D) activities within the Fuel Cycle Technologies (FCT) program. The FCT program, tasked by the U.S. Department of Energy (DOE), employs a science-based approach to advance fuel technologies. This approach integrates theory, experiments, and multi-scale modeling and simulation (M&S) to develop a predictive understanding of fuel fabrication processes and fuel/cladding performance under irradiation, moving beyond traditional empirical methods. The long-term goals of the AFC are guided by the AFC Strategic Plan and align with the DOE Office of Nuclear Energy (NE) Roadmap [1], which outlines a multi-decade vision for demonstrating and qualifying advanced fuel forms to support diverse fuel cycle options. Near-term goals focus on enhancing accident tolerant fuels (ATF) for Light Water Reactors (LWR), a significant challenge that demands balancing immediate objectives with ongoing progress toward advanced reactor missions. Accelerating the traditional fuel qualification process to meet ATF objectives is another critical challenge. A detailed set of 5-year goals, summarized below, has been developed in line with the overarching science-based fuel development approach: • Advanced LWR Fuel Technologies: By 2027, support the development of advanced LWR fuel technologies with improved performance and enhanced accident tolerance. This includes high burnup (HBu), low enriched uranium (LEU)+, coated cladding, and doped fuel, aimed at complementing industry-led significant LWR uprates and plant refurbishments. • Tristructural Isotropic (TRISO) Fuel: Achieve qualification by 2028 and develop improved designs for emerging markets. • Metal Fuel: Achieve qualification by 2028 and develop improved designs for emerging markets. • Molten Salt Fuel: By 2027, deploy a robust program that enables fuel salt qualification technologies needed to support fuel salt research and development (R&D), focusing on emergent needs to derisk fuel salt production and utilization in advanced reactors. • Long-Term ATF: Develop fuel technologies that enable significant power uprates (~50%) in refurbished or new LWRs while optimizing fissile material utilization and waste disposal. The 5-year milestones in the AFC Execution Plan are contingent on an assumed budget. This Execution Plan will be updated annually to reflect actual funding profiles as budget guidance becomes available, ensuring milestones are adjusted accordingly. In summary, the AFC Execution Plan presents a comprehensive strategy to advance nuclear fuel technologies through a science-based approach, addressing both near-term and long-term goals while adapting to funding realities.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

ACDC (Automated Campbell Diagram Code) [SWR-26-042]

This application provides a web-based graphical user interface to generating Campbell Diagrams and visualizing mode shapes for OpenFAST turbine models. Determining the aeroelastic stability and dynamic characteristics of wind turbines is a critical step in turbine design and analysis. Historically, extracting natural frequencies and mode shapes from OpenFAST—the industry-standard whole-turbine simulation code—has been a fragmented and tedious process. It required manual model configuration, command-line linearization execution, and complex post-processing via proprietary scripts to handle rotating-frame dynamics. To address these workflow bottlenecks, we present the Automated Campbell Diagram Code (ACDC), an open-source graphical software tool developed by the National Laboratory of the Rockies (NLR) under the DOE-funded Distributed Wind Aeroelastic Modeling (dWAM) project. ACDC streamlines the end-to-end linearization and stability analysis workflow into a single, intuitive cross-platform application. The software guides users through OpenFAST model configuration, definition of operating points, and the automated execution of steady-state trim and linearization simulations. Under the hood, ACDC automates the complex mathematical post-processing steps required for rotating systems, including Multi-Blade Coordinate (MBC) transformations, eigenanalysis, and advanced modal tracking utilizing the Modal Assurance Criterion (MAC) and spectral clustering. Finally, ACDC processes these results to automatically generate Campbell diagrams and features a robust 3D visualization engine to animate full-system mode shapes. By eliminating the reliance on external post-processing environments and manual data manipulation, ACDC significantly accelerates dynamic analysis and lowers the barrier to entry for wind energy researchers and engineers.

Summerville, Brent [National Laboratory of the Roc↗

Bridging paradigms: Designing for HPC-Quantum convergence

Here, this paper presents a comprehensive software stack architecture for integrating quantum computing (QC) capabilities with High-Performance Computing (HPC) environments. While quantum computers show promise as specialized accelerators for scientific computing, their effective integration with classical HPC systems presents significant technical challenges. We propose a hardware-agnostic software framework that supports both current noisy intermediate-scale quantum devices and future fault-tolerant quantum computers, while maintaining compatibility with existing HPC workflows. The architecture includes a quantum gateway interface, standardized APIs for resource management, and robust scheduling mechanisms to handle both simultaneous and interleaved quantum–classical workloads. Key innovations include: (1) a unified resource management system that efficiently coordinates quantum and classical resources, (2) a flexible quantum programming interface that abstracts hardware-specific details, (3) A Quantum Platform Manager API that simplifies the integration of various quantum hardware systems, and (4) a comprehensive tool chain for quantum circuit optimization and execution. We demonstrate our architecture through implementation of quantum–classical algorithms, including the variational quantum linear solver, showcasing the framework’s ability to handle complex hybrid workflows while maximizing resource utilization. This work provides a foundational blueprint for integrating QC capabilities into existing HPC infrastructures, addressing critical challenges in resource management, job scheduling, and efficient data movement between classical and quantum resources.

97 MATHEMATICS AND COMPUTING↗

Lessons Learned from Ecosystem-Scale Experimental Field Studies (Workshop Report)

Efforts to understand and predict ecosystem responses to environmental change require long-term, large-scale, spatially representative experiments and observations that capture natural variability, test predictive models, and generate transferable knowledge. Such studies are indispensable for unraveling the complexities of terrestrial ecosystems and their responses to disturbances and evolving environmental conditions, while generating the data necessary for developing mechanistic models and predictive tools that inform decision-making processes. Having a rich history of designing and executing large-scale ecosystem experiments, the U.S. Department of Energy’s Environmental System Science program convened a workshop in January 2025 that brought together leaders in the field to distill critical lessons from decades of experience in large-scale experiments. The workshop aimed to (1) provide an ecosystem experiment primer for best practices, thus ensuring a high scientific return on investment for funding agencies, and (2) offer a robust framework for the design and management of future research initiatives. This report synthesizes insights and experiences from workshop participants and is structured to capture the entire research life cycle, from goal setting and design to operations, adaptive management, team dynamics, collaborations, and the often overlooked aspect of decommissioning. By synthesizing decision-making and lessons learned across diverse research approaches, the report aims to provide a template of essential factors to consider when designing successful long-term, large-scale ecosystem experiments.

54 ENVIRONMENTAL SCIENCES↗

Cyber Conservative Operations

In an era of increasingly sophisticated and pervasive cyber threats, robust cyber resilience strategies are more critical than ever. This paper introduces the concept of Cyber Conservative Operations, a proactive approach designed to assist critical infrastructure owners and operators in managing risk and maintaining resilience in the face of imminent, yet not occurring, cyber events. By leveraging the principles of Cyber-Informed Engineering (CIE) and the energy sector’s practice of conservative operations, Cyber Conservative Operations offer a framework for planning and executing coordinated active defense and resilience-oriented actions ahead of cyber events. This approach aims to minimize the consequences of digitally-enabled hazards and ensure swift recovery from anticipated cyber threats. Cyber Conservative Operations enhance the ability of owners and operators to execute pre-planned defense and resilience actions, thereby reducing the impact of digitally-enabled hazards. These operations are crucial for addressing impacts that cannot be precisely quantified or forecasted and for preparing organizations for rapid recovery from imminent digital threats. The paper begins with a discussion of conservative operations as applied to the bulk power system (BPS), a current mechanism allowing BPS entities to enact defensive operating plans to mitigate impending grid unreliability. Building on this model, we present a concept for implementing cyber conservative operations at asset owner facilities. The paper concludes with two case studies of cyber conservative operations drawn from high-profile cyber events, illustrating the practical application and benefits of this proactive approach.

42 - ENGINEERING↗

Bridging Cloud and Edge Computing at NREL Using CONNECT: Cloud Optimized Networking for Next-Gen Edge Computing Technologies [Slides]

CONNECT is an innovative on-premise hardware and software solution that integrates edge and cloud computing infrastructure at NREL. Built on the AWS Greengrass middleware and leveraging the MQTT protocol, CONNECT enables real-time data streaming from IoT devices and gateways to both cloud and local services, empowering researchers to rapidly capture, analyze, and act upon edge-generated data while leveraging cloud capabilities. The platform addresses research infrastructure challenges by providing a pre-approved platform which is already configured with the correct networking and cybersecurity baselines thus eliminating procurement delays and enabling on-demand availability. CONNECT's hybrid architecture efficiently manages burstable workloads, allowing research teams to dynamically scale computational capacity, handle peak data loads, and reduce operational bottlenecks. Advanced capabilities include built-in GPU support for executing machine learning models which enables low-latency inference at the edge from models trained in the cloud. This architecture supports real-time analytics and filtering, providing a mechanism to allow only transmitting and processing high-value data. Cloud-based configuration management permits engineers to manage on-premise systems remotely, optimizing operational efficiency. By bridging edge and cloud computing, CONNECT provides NREL researchers with a flexible, scalable platform that accelerates scientific discovery while maintaining robust security and performance standards.

97 MATHEMATICS AND COMPUTING↗

End-to-end protocol for high-quality quantum approximate optimization algorithm parameters with few shots

The quantum approximate optimization algorithm (QAOA) is a quantum heuristic for combinatorial optimization that has been demonstrated to scale better than state-of-the-art classical solvers for some problems. For a given problem instance, QAOA performance depends crucially on the choice of the parameters. While average-case optimal parameters are available in many cases, meaningful performance gains can be obtained by fine-tuning these parameters for a given instance. This task is especially challenging, however, when the number of circuit executions (shots) is limited. In this work, we develop an end-to-end protocol that combines multiple parameter settings and fine-tuning techniques. We use large-scale numerical experiments to optimize the protocol for the shot-limited setting and observe that optimizers with the simplest internal model (linear) perform best. We implement the optimized pipeline on a trapped-ion processor using up to 32 qubits and 5 QAOA layers, and we demonstrate that the pipeline is robust to small amounts of hardware noise. To the best of our knowledge, these are the largest demonstrations of QAOA parameter fine-tuning on a trapped-ion processor in terms of two-qubit gate count.

quantum algorithms & computation↗