Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Workflow Management Systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

User-Focused Tools to Enhance IT/OT Cyber Resilience within the Power Grid

The power grid is undergoing several changes that are increasing its complexity as nexuses between electric-gas, transmission-distribution, and energy-communications continue to become increasingly critical. This system is heavily dependent on communication infrastructure and controls, and it relies on humans in operational technology (OT) and information technology (IT) roles to manage the increasing breadth, depth, and speed of data. Many technical challenges have presented themselves and will need to be addressed to provide reliable grid operations. With increased reliance on distributed controls and communication infrastructure, cybersecurity becomes an inherent requirement. When considering current cyber-physical security solutions for the power grid, one can notice a clear divide between information technology and operation technology networks. However, in real-life applications, these networks are interdependent. This work presents results of interviews with key utility cybersecurity personnel, analyzes the results, and makes recommendations towards solution of existing technical and operational challenges realized. Existing workflows are presented, wireframe interviews are discussed, and tool requirements are described. The existence of easy-to-implement solutions, based on existing energy management systems, highlight the potential for real-life applications.

cybersecurity, resilience, user-centered design, p↗

Frontier Job-Centric Telemetry Dataset

Comprehensive analysis of high-performance computing (HPC) systems requires linking workload execution to system behavior. This kind of analysis is vital for diagnosing performance issues, managing capacity, detecting anomalous workloads, and understanding how applications interact with system hardware. This job-centric telemetry dataset unifies scheduler job records with node-level measurements, enabling direct association between workloads and their corresponding power, thermal, and performance characteristics. It contains sanitized, scheduler related metadata for 152,400 individual jobs that ran on the Frontier supercomputer and ended on selected days throughout 2024 and 2025, a subpopulation of ~6.8% of the total number of allocated jobs with non-zero run time on the system over that same period. Each is linked with files that contain telemetry time series records of the power utilization and temperature behavior of its allocated nodes and their processors during the run time of the job. Where available, a portion of the job files also contain network performance time series. Jobs are sampled from select days that reflect normal levels of user activity and possess job size distributions with large numbers of leadership class jobs (>20% of Frontier nodes). Jobs in this dataset attempt to best represent successful user workflows.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

AI-Ready Control System for the Fermilab Accelerator Complex

Reliable, high-intensity operation of the Fermilab Accelerator Complex is critical to the success of the Long-Baseline Neutrino Facility and Deep Underground Neutrino Experiment. We describe the requirements and infrastructure necessary to support routine use of artificial intelligence and machine learning (AI/ML) in the accelerator control system. Three capabilities are identified: a machine learning operations (MLOps) framework standardizing the lifecycle of AI/ML automation from data management through deployment and monitoring; a data quality framework defining and enforcing standards required to build trustworthy AI/ML applications; and workflow integration with large language models to assist physicists, engineers, and operators with information retrieval, code development, and routine analysis. Use cases spanning beam diagnostics, beam control, and support system automation illustrate the technical requirements across the complex.

43 PARTICLE ACCELERATORS↗

A galactic approach to neutron scattering science

Neutron scattering science is leading to significant advances in our understanding of materials and will be key to solving many of the challenges that society is facing today. Improvements in scientific instruments are actually making it more difficult to analyze and interpret the results of experiments due to the vast increases in the volume and complexity of data being produced and the associated computational requirements for processing that data. New approaches to enable scientists to leverage computational resources are required, and Oak Ridge National Laboratory (ORNL) has been at the forefront of developing these technologies. We recently completed the design and initial implementation of a neutrons data interpretation platform that allows seamless access to the computational resources provided by ORNL. For the first time, we have demonstrated that this platform can be used for advanced data analysis of correlated quantum materials by utilizing the world's most powerful computer system, Frontier. In particular, we have shown the end-to-end execution of the DCA++ code to determine the dynamic magnetic spin susceptibility χ(q, ω) for a single-band Hubbard model with Coulomb repulsion U/t = 8 in units of the nearest-neighbor hopping amplitude t and an electron density of n = 0.65. The following work describes the architecture, design, and implementation of the platform and how we constructed a correlated quantum materials analysis workflow to demonstrate the viability of this system to produce scientific results.

97 MATHEMATICS AND COMPUTING↗

HDF5 in the exascale era: Delivering efficient and scalable parallel I/O for exascale applications

Accurately modeling real-world systems requires scientific applications at exascale to generate massive amounts of data and manage data storage efficiently. However, parallel input and output (I/O) faces challenges due to new application workflows and the state-of-the-art memory, interconnect, and storage architectures considered in exascale designs. The storage hierarchy has expanded with node-local persistent memory, solid-state storage, and traditional disk and tape-based storage, thus requiring efficiency at each layer and much more efficient data movement among these layers. This paper discusses how the ExaHDF5 project improved the I/O performance and data management for exascale architectures by enhancing HDF5, a widely used parallel I/O library. The team developed an Asynchronous I/O Virtual Object Layer (VOL) connector that allowed overlapping I/O with computation. They also created a Cache VOL to complement asynchronous I/O by incorporating fast storage layers, such as burst buffer and node-local storage, into the parallel I/O workflow through caching and staging data. Additionally, the team enabled data aggregation and I/O at the node level by using a Subfiling Virtual File Driver (VFD). To demonstrate superior I/O performance with HDF5 at exascale, the ExaHDF5 team collaborated with several exascale applications. In this paper, we show I/O performance improvements for three applications: Cabana (a particle-based simulation library), EQSIM (a regional earthquake simulation software), and E3SM (a climate system modeling library).

Asynchronous I/Ol↗

Waste Compliance and Tracking System (WCATS) Version 3 Requirements Document

This document describes the end-user requirements for the Waste Compliance and Tracking System (WCATS) project in accordance with the WCATS Software Quality Management Plan, EPC-WMP-WCATSPLAN-001. The WCATS application shall support the generation, characterization, processing, and shipment of LANL radioactive, hazardous, and industrial waste. Regulatory drivers include RCRA hazardous waste, DOT shipping, NNSA nuclear material control and accountability, DOE nuclear safety, TSDF permit, and transuranic waste certification requirements. The system will utilize a task-based architecture that supports the spectrum of treatment, storage, disposal, administrative, and characterization based unit operations necessary to manage waste from cradle to grave. The application design shall readily accommodate new facilities, processes, workflow, signature requirements, and so forth, via end-user established metadata. WCATS will provide support for representing waste storage and disposal facilities, buildings, rooms, and grid layouts (x, y, z) to support waste and radioactive material inventory management. Nuclear material at risk (MAR), DOE hazard rating (e.g., Category II facility) compliance per DOE-STD-1027, and permit inventory requirements will be configurable for any storage or disposal facility, or waste operation, and the system will automatically evaluate and enforce those requirements. In addition, the application will support the characterization and management of the entire range of hazardous and radioactive wastes (TRU, MTRU, LLW, MLLW, hazardous waste, etc.) that might be colocated or processed at a permitted facility. Some capabilities not found in traditional systems include user-defined tank systems for liquid waste, user-defined work paths (i.e., sequence of operations), and an equipment subsystem for tracking the calibration, maintenance, and inspection of tools used to process waste, such as torque wrenches, scales, pH probes, etc. The application incorporates a desktop and mobile user interface as shown in Figure 1. The mobile interface supports field operations, such as waste item characterization, intra-facility transfers, internal and external audits, and shipment preparation and receipt.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

GDSA framework, a computational framework for complex modeling problems in radioactive waste management

This paper details a computational framework to produce automated, graphical workflows, and how this framework can be deployed to support complex modeling problems like those in nuclear engineering. Key benefits of the framework include: automating previously manual workflows; intuitive construction and communication of workflows through a graphical interface; and automated file transfer and handling for workflows deployed across heterogeneous computing resources. This paper demonstrates the framework's application to probabilistic post-closure performance assessment of systems for deep geologic disposal of nuclear waste. However, the framework is a general capability that can help users running a variety of computational studies.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Open Reproducible Electron Microscopy Data Analysis

Electron microscopy (EM) is a cornerstone technique in the materials and biological sciences capable of imaging structures at nano- to atomic-scale resolution. Advances in technologies mean that one acquires datasets at increasing data rates and sizes. These advancements present enormous opportunities for researchers to understand complex systems. However, processing the resulting large-scale, complex data in a reproducible and shareable way is a real challenge for researchers. The building, managing, and maintaining complex workflows in a reproducible manner requires extensive knowledge in several areas outside the researchers’ core skill sets, such as software engineering, data science, and high-performance computing (HPC). Our work demonstrates an innovative approach to solving these problems, enabling reproducible EM data analysis through container encapsulated pipelines. Using modern container technologies, we encapsulate processing elements and connect them using shared memory. We expose user-friendly, advanced algorithms and tools to allow end users to utilize without expert programming skills. The platform enables reproducible, scalable, shareable pipelines for the analysis and visualization of EM data. Focusing on interoperability, we leverage the DOE and other agencies’ existing investments to provide a powerful software platform for EM data analysis.

Harris, Christopher↗

FY 2025 Multidimensional Data Correlation Platform: Unified Software Architecture for Advanced Materials and Manufacturing Technologies Data Management and Processing

The Advanced Materials and Manufacturing Technologies (AMMT) program continues to advance a data-driven approach to demonstrate the utility of additive manufacturing for fabricating components for nuclear applications. A key scientific goal is to leverage data to better understand manufacturing outcomes and thereby improve the performance, reliability, and lifespan of nuclear components. Ultimately, this effort supports the development of standards for certification and qualification of additively manufactured components, enabling broader industry adoption. In support of this objective, the AMMT program is building and deploying a data management platform to record, index, analyze, and make available the manufacturing data generated across the AMMT program. In FY 2023, the team conceptualized the architecture of the platform and, in FY 2024, deployed the first functional version at the Oak Ridge National Laboratory (ORNL) Manufacturing Demonstration Facility (MDF). In FY 2025, the platform was officially opened to all AMMT members. To enable this expansion, core modifications and enhancements were developed, including improvements to the user interface and workflows for data entry and retrieval. Most notably, robust security and access control mechanisms were implemented to protect data and manage information sharing. This effort featured a logging system, protected views, and controlled access mechanisms. This report documents these enhancements and the transition of the platform into program-wide use.

36 MATERIALS SCIENCE↗

MSD CoP Webinar: "Advances in MSD-LIVE to Support the MSD Community of Practice"

Context: This webinar was hosted by the MultiSector Dynamics Community of Practice (MSD CoP; https://multisectordynamics.org). Advances in MSD-LIVE to Support the MSD Community of Practice Presenters: Casey Burleyson and Zoe Guillen (Pacific Northwest National Laboratory) Abstract: The MultiSector Dynamics Living, Intuitive, Value-adding, Environment (MSD-LIVE; msdlive.org) is a cloud-based data management system and advanced computing platform that enables MSD researchers to document and archive their data, run their models and analysis tools, and share their data, software, and workflows within the MSD Community of Practice. Recently, several high-profile datasets have attracted many new users to MSD-LIVE. This webinar has two goals: 1) To refamiliarize the MSD community and new users with the components of the platform (e.g., the data repository, model training notebooks, and data dashboards) and to highlight examples of how these components are advancing MSD science and 2) To demonstrate new features in v3 of the platform, released in late 2025. The main new feature in v3 is the ability to interactively explore data in MSD-LIVE without downloading it. MSD-LIVE users can now click a button in our data repository and launch a blank Jupyter notebook with access to the underlying data on AWS. Users can use the notebook to write analysis, visualization, or subsetting routines that process the data directly on the AWS cloud. We also added a GitHub integration feature that allows users to share analysis or visualization code they develop with the community of MSD-LIVE users. The webinar will wrap up with a look at what's coming next for MSD-LIVE in 2026. Moderator: Patrick M. Reed (MSD CoP Facilitation Team) This webinar was held on: May 12th, 2026 from 1-2 PM EST.

Open Science↗

INTEGRATION OF DATA ANALYTICS WITH SYSTEM HEALTH PROGRAMS

Industry equipment reliability and asset management programs are essential elements that help ensure the safe and economical operation of nuclear power plants. The effectiveness of these programs is addressed in several industry developed and regulatory programs. However, these programs have proven to be labor intensive and expensive. There is an opportunity to significantly enhance the collection, analysis, and use of this information to provide more cost-effective plant operation. Additionally, there is an acute industry need to leverage advanced technology to reduce costs and improve operational effectiveness. The goal of this paper is to provide effective and efficient analytical methods and tools to support risk-informed decisions for the equipment reliability and asset management programs at nuclear power plants. This is accomplished by creating a direct bridge between component health/lifecycle data and decision making (e.g., maintenance scheduling and project prioritization). Here we are supporting typical system engineer decisions regarding maintenance activity scheduling and component ageing management. This is performed in a risk-informed context where herein the term “risk” is broadly constructed to include both plant reliability and economics. This framework combines data analytics tools to analyze equipment reliability data with risk-informed methods designed to support system engineer decisions (e.g., maintenance and replacement schedules, optimal maintenance posture) in a customizable workflow. A challenge is that the structure of this workflow strongly depends on the decision that needs to be made, the type of data available, and the constraints that need to be considered. Current methods are designed to provide specific answers to specific problems; however, these methods might prove to be inadequate even when problem settings slightly change (e.g., different types of requirements, additional dependencies between system reliability and economics). We tackled this challenge by designing framework in a flexible and modular fashion such that the user can assemble and customize his/her own workflow that integrates SSC economic lifecycle models (e.g., maintenance and replacement costs), system reliability models, and optimization methods.

97 - MATHEMATICS AND COMPUTING↗

Development of a nearly autonomous management and control system for advanced reactors

This presentation discusses the development of a Nearly Autonomous Management and Control (NAMAC) system for advanced reactors. This presentation first introduces the NAMAC-enabled plant control. Next, this presentation introduces major technologies of implementing NAMAC, including operational workflow, digital twin, and advanced machine learning algorithms. To demonstrate the capability of the NAMAC system, NAMAC is connected with a simulator of the Experimental Breeder Reactor II, where NAMAC is activated to provide recommendations during a single loss of flow accident.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Active learning path-dependent properties using a cloud-based materials acceleration platform

Solid state materials are central to many modern technologies in which a given material may be exposed to a variety of environments. The material properties often vary with the sequence of environments in an irreversible manner, resulting in a quintessential path-dependency in experimental observables. While sequential learning techniques have been effectively deployed for accelerating learning of state properties of materials, they often use a consistent environment path in all experiments. To elevate such techniques for making optimal decisions in experimental investigations of path-dependent properties, we introduce an iterated expected information gain acquisition function that optimizes over entire experimental trajectories. This approach is implemented within a cloud-based Materials Acceleration Platform architecture utilizing an event-driven stateful broker coupled with remote HELAO (Hierarchical Experimental Laboratory Automation and Orchestration) instances and an AI science manager. The platform's efficacy was demonstrated through a case study optimizing multi-step spectro-electrochemical experiments to identify optically stable potential windows in (Co–Ni–Sb)O z metal oxides. The system successfully integrated AI-driven experiment design, remote laboratory automation, and cloud-based data infrastructure, validating the platform's capability for managing complex, adaptive, path-dependent workflows in materials discovery.

Guevarra, Dan [California Institute of Technology ↗

Parallel Battery: The Framework and Process for an Intelligent and Ecological Battery System and Related Services

The concept, framework, process methodology and applications of parallel battery were proposed from both virtual and real aspects.The parallel battery was an application of ACP-based parallel intelligence in battery and related energy system areas.The real battery system was running with its equivalent, and the artificial battery system was in a virtual space, in a parallel and interactive manner.The artificial battery system contained the descriptive, predictive, and prescriptive functions on the real battery and related energy systems.There was a closed-loop workflow between the real battery system and the artificial battery system, which iteratively optimizes the battery and related energy systems, leading to a new paradigm of intelligent and ecological parallel battery system management.

ACP approach↗

Darshan for HEP applications

Modern HEP workflows must manage increasingly large and complex data collections. HPC facilities may be employed to help meet these workflows’ growing data processing needs. However, a better understanding of the I/O patterns and underlying bottlenecks of these workflows is necessary to meet the performance expectations of HPC systems.Darshan is a lightweight I/O characterization tool that captures concise views of HPC application I/O behavior. It intercepts application I/O calls at runtime, records file access statistics for each process, and generates log files detailing application I/O access patterns.Typical HEP workflows include event generation, detector simulation, event reconstruction, and subsequent analysis stages. A study of the I/O behavior of the ATLAS simulation and filtering stage, and the CMS simulation workflow using Darshan is presented, including insights into the I/O operations and data access size.

Wang, Rui↗

Distributed Computing for the Project 8 Experiment

The Project 8 collaboration aims to measure the absolute neutrino mass or improve on the current limit by measuring the tritium beta decay electron spectrum. We present the current distributed computing model for the Project 8 experiment. Project 8 is in its second phase of data taking with a near continuous data rate of 1Gbps. The current computing model uses DIRAC (Distributed Infrastructure with Remote Agent Control) for its workflow and data management. A detailed meta-data assignment using the DIRAC File Catalog is used to automate raw data transfers and subsequent stages of data processing. The DIRAC system is deployed on containers managed using a Kubernetes cluster to provide a scalable infrastructure. A modified DIRAC Site Director provides the ability to submit jobs using Singularity on opportunistic High-Performance Computing (HPC) sites.

Distributed Computing, Kubernetes, DIRAC, Project ↗

Interpretable Models for Workflow Differentiation in High-Performance Scientific Networks

Scientific workflows in high-performance networks spawn hundreds of interdependent flows that must be managed collectively—yet existing network classifiers treat each flow in isolation, leading to fragmented QoS decisions and missed interflow patterns. We present a novel traffic classification solution that operates at the workflow level, distinguishing entire filetransfer operations from streaming analytics by capturing how concurrent flows interact and burst together. We introduce a workflow identification window (WIW) that ingests raw packet headers from parallel flows into unified tensors, preserving the spatial-temporal patterns that differentiate scientific workflows. This approach achieves 98.7% accuracy using CNN, LSTM, and hybrid architectures, while maintaining 84% accuracy on production traffic collected a week later—demonstrating robustness to temporal drift. By integrating SHAP and GradCAM explainability, we reveal that early-packet timing patterns and cross-flow correlations drive classification decisions, providing operators with interpretable insights. Our system enables coherent workflow-level QoS enforcement and dynamic bandwidth allocation in scientific networks, eliminating manual per-flow configuration while maintaining classification latency at millisecond level.

Giannakou, Anna [LBL, Berkeley]↗

VA EDH Advanced Software Pipeline Framework Report: Enhancing Automation and Scalability

The VA Environmental Determinants of Health (EDH) Advanced Software Pipeline Framework is designed to enhance the efficiency, scalability, and security of geospatial data processing workflows. This framework integrates modern data orchestration and containerization technologies, including Prefect for workflow automation, Docker for containerization, and PostgreSQL/PostGIS for geospatial data storage and analysis. It ensures standardized, reproducible, and automated data processing, supporting VA objectives related to substance use risk assessment and recovery research. The pipeline addresses key scalability and performance challenges through horizontal and vertical scaling, high-performance computing (HPC) integration, parallel processing, task caching, and dynamic resource allocation. These optimizations improve throughput and reduce latency, allowing the system to efficiently manage large and complex datasets. Additionally, security and compliance measures—such as data encryption (SSL), Role-Based Access Control (RBAC), and adherence to GDPR and HIPAA standards—safeguard sensitive information throughout data transmission and storage. A key implementation of this framework includes the automation of shelter list geolocation workflows, ensuring that up-to-date data is readily available for VA decision-making. Lessons learned from this project include the transition from in-memory processing to incremental storage writes, improving resource management and reliability. Future enhancements aim to expand automation, integrate AI-driven anomaly detection, and incorporate high-performance computing resources. This framework provides a scalable, secure, and adaptable solution for managing geospatial datasets, reinforcing the VA’s ability to support clinical and strategic initiatives through data-driven decision-making.

97 MATHEMATICS AND COMPUTING↗