Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Workflow Management Systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Integrated End-to-end Performance Prediction and Diagnosis for Extreme Scientific Workflows

This report details recent progress for the ASCR funded project “Integrated End-to-end Performance Prediction and Diagnosis for Extreme Scientific Workflows”. We refer to the project as IPPD/2, reflecting the 2017 renewal under expanded scope and partners In IPPD/2, we increased our research scope to include data motion. We are focusing on three major aspects: a) observe how data is generated, distributed, and used; b) analyze how data is (repeatedly) consumed with a focus both on repeated patterns and anomalies; and c) explore how to optimize data motion. This new work on data motion will augment and complement IPPD/2’s research that focused on the computational aspects of tasks. We leverage and extend our existing tools and demonstrate our work on the Belle II workflow suite as well as on workflows from NSLS-II. The highlights of our work are as follows: Provenance for Workflows: Provenance is used to provide information enabling quality control, re-run computational workflows, and reproduce results. IPPD/2 has been building a scalable provenance management system that enables the capture of provenance from the high-level workflow through all relevant system levels in one integrated environment. Leveraging this work, our recent efforts have included using provenance as an enabling technique. Workload characterization: Leveraging provenance and analysis, we characterize data movement within network, storage, and memory over a variety of workloads. This characterization enables an understanding by performance analysts and application developers of the range of behaviors that could be expected. Performance Prediction for Workflows: The goal of modeling distributed workflows is to understand performance bottlenecks and enable more intelligent task scheduling to optimize selected metrics of interest (e.g., task throughput or output data rate). IPPD/2 has utilized both analytical and AI/ML modeling methodologies for performance modeling. Advanced Scheduling and Fault Modeling for Workflows: Scheduling of large-scale scientific workflows on geographically distributed resources is a challenging problem. To improve workflow throughput, we combined novel scheduling algorithms with task predictions from performance modeling and fault modeling. Dynamically Alleviating Bottlenecks in Workflows: Exploiting our provenance, analysis, and modeling efforts, we have explored and developed several techniques for dynamically detecting and alleviating bottlenecks in data movement. In particular, we have spent considerable effort demonstrating our techniques on production-like workflow configurations.

97 MATHEMATICS AND COMPUTING↗

RISE: Reducing I/O Contention in Staging-based Extreme-Scale In-situ Workflows

While in-situ workflow formulations have addressed some of the data-related challenges associated with extreme-scale scientific workflows, these workflows involve complex interactions and different modes of data exchange. In the context of increasing system complexity, such workflows present significant resource management challenges, requiring complex cost-performance tradeoffs. This paper presents RISE, an intelligent staging-based data management middleware, which builds on the DataSpaces framework and performs intelligent scheduling of data management operations to reduce I/O contention. In RISE, data are always written immediately to local buffers to reduce the effect of the transfer impact upon application performance. RISE identifies applications’ data access patterns and moves data towards data consumers only when the network is expected to be idle, reducing the impact of asynchronous background data movement upon critical data read/write requests. Here, we experimentally demonstrate that RISE can take advantage of staging nodes to offload data during writes without degrading application data movement performance.

97 MATHEMATICS AND COMPUTING↗

SSC Engineering Analysis

A package for the automation of the Engineering Analysis (EA) process at the Stennis Space Center has been customized. It provides the ability to assign and track analysis tasks electronically, and electronically route a task for approval. It now provides a mechanism to keep these analyses under configuration management. It also allows the analysis to be stored and linked to the engineering data that is needed to perform the analysis (drawings, etc.). PTC s (Parametric Technology Corp o ration) Windchill product was customized to allow the EA to be created, routed, and maintained under configuration management. Using Infoengine Tasks, JSP (JavaServer Pages), Javascript, a user interface was created within the Windchill product that allows users to create EAs. Not only does this interface allow users to create and track EAs, but it plugs directly into the out-ofthe- box ability to associate these analyses with other relevant engineering data such as drawings. Also, using the Windchill workflow tool, the Design and Data Management System (DDMS) team created an electronic routing process based on the manual/informal approval process. The team also added the ability for users to notify and track notifications to individuals about the EA. Prior to the Engineering Analysis creation, there was no electronic way of creating and tracking these analyses. There was also a feature that was added that would allow users to track/log e-mail notifications of the EA.

Ryan, Harry↗

MADA: Multi-Agent Design Assistant

MADA (Multi-Agent Design Assistant) is a Large Language Model (LLM) powered multi-agent framework that coordinates specialized agents for complex design workflows. The system was designed for HPC workflows with the following agents in mind: 1) A Job Management Agent (JMA) launches and manages ensemble simulations on HPC systems, 2) a Geometry Agent (GA) generates meshes, and 3) an Inverse Design Agent (IDA) proposes new designs informed by simulation outcomes. Our framework reduces cumbersome manual workflow setup, and enables automated design exploration at scale. However, the software also enables users to rapidly create new multi-agent systems. Simply define new agents in a configuration file, giving each their own set of tools (via MCP), and then chat and prompt your new multi-agent system. Is

Gunnarson, BrianS [Lawrence Livermore National Lab↗

First-generation college student forges ahead, now key to Lab's mission

Fatima Woody was just 17 and a student at Pojoaque Valley High School when she first started her career as a Los Alamos Neutron Science Center receptionist. Now she's in a crucial role that keeps plutonium pit production and other mission processes operating with as little interruption as possible. Over nearly four decades, Fatima has gradually advanced from her initial positions as a receptionist and administrative secretary to become a computer technician, then a computer system professional who specializes in project management. Today, she coordinates the workflow for nearly two dozen deployed information technology technicians who keep the computer systems and networks operating at the high-tech complex that houses the Lab's Plutonium Facility.

99 GENERAL AND MISCELLANEOUS↗

Software and computing for Run 3 of the ATLAS experiment at the LHC

The ATLAS experiment has developed extensive software and distributed computing systems for Run 3 of the LHC. These systems are described in detail, including software infrastructure and workflows, distributed data and workload management, database infrastructure, and validation. The use of these systems to prepare the data for physics analysis and assess its quality are described, along with the software tools used for data analysis itself. An outlook for the development of these projects towards Run 4 is also provided.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Practical Battery Thermal Modeling Techniques

Lithium-ion batteries are thermo-electrochemical devices, whereby nearly every facet of their functionality and performance are thermally driven. As a result, it is important to have thermal modeling techniques that effectively capture the intricacies of both the electrochemical nature of the battery and also the complex thermal network that typically results from the design of the battery thermal management system. Here we present a thermal modeling workflow and a set of general assumptions for how to construct a thermal model of a Li-ion battery pack. We use a 14-cell bank of 18650-format Li-ion cells, loosely based on a proposed alternative battery design for Orion, as the example. Although the workflow is performed with Thermal Desktop and related utilities, the focus of this presentation is less about software specific techniques, but rather is focused on the assumptions and conditions that should be used in a model (regardless of the tool used to build the model). Example cases and results will be presented for charge, discharge, and thermal runaway.

lithium-ion battery↗

$\mathrm{RADICAL}$-Pilot and $\mathrm{PMIx}$/$\mathrm{PRRTE}$: Executing Heterogeneous Workloads at Large Scale on Partitioned $\mathrm{HPC}$ Resources

Execution of heterogeneous workflows on high-performance computing (HPC) platforms present unprecedented resource management and execution coordination challenges for runtime systems. Task heterogeneity increases the complexity of resource and execution management, limiting the scalability and efficiency of workflow execution. Re-source partitioning and distribution of tasks execution over portioned re-sources promises to address those problems but we lack an experimental evaluation of its performance at scale. Here this paper provides a performance evaluation of the Process Management Interface for Exascale (PMIx) and its reference implementation PRRTE on the leadership-class HPC plat-form Summit, when integrated into a pilot-based runtime system called RADICAL-Pilot. We partition resources across multiple PRRTE Distributed Virtual Machine (DVM) environments, responsible for launching tasks via the PMIx interface. We experimentally measure the work-load execution performance in terms of task scheduling/launching rate and distribution of DVM task placement times, DVM startup and termination overheads on the Summit leadership-class HPC platform. Integrated solution with PMIx/PRRTE enables using an abstracted, standardized set of interfaces for orchestrating the launch process, dynamic process management and monitoring capabilities. It extends scaling capabilities allowing to overcome a limitation of other launching mechanisms (e.g., JSM/LSF). Explored different DVM setup configurations provide insights on DVM performance and a layout to leverage it. Our experimental results show that heterogeneous workload of 65,500 tasks on 2048 nodes, and partitioned across 32 DVMs, runs steady with resource utilization not lower than 52%. While having less concurrently executed tasks resource utilization is able to reach up to 85%, based on results of heterogeneous workload of 8200 tasks on 256 nodes and 2 DVMs.

97 MATHEMATICS AND COMPUTING↗

Empowering Scientific Discovery Through Computing at the Advanced Photon Source

This paper explores the challenges and solutions for managing and processing the vast amount of data generated by the Advanced Photon Source (APS), a synchrotron light source facility producing ultra-bright x-rays for diverse scientific domains. With 68 experimental beamlines covering materials research, biology, and more, the APS serves a wide user base across academia, government, and industry. The ongoing upgrade of the APS storage ring and installation of new instruments will amplify data generation and processing demands. This paper discusses the approach to address these demands through automated data processing using standardized workflows that produce faster scientific insights. The APS Data Management System coordinates various data related tasks to manage storage, data transfer, metadata cataloging, data processing, and interfaces with tools provided by Globus. Through integration with the Argonne Leadership Computing Facility (ALCF), APS users can efficiently access high-performance computing resources. Standardized workflows have led to reduced computational burdens on scientists and greater accessibility of high performance computing resources. We demonstrate how standardization and collaboration enable scientists to rapidly convert raw data into meaningful scientific results, establishing a streamlined path from data collection to analysis and ultimately to publication.

Parraga, Hannah↗

Enabling Low-Overhead HT-HPC Workflows at Extreme Scale using GNU Parallel

GNU Parallel is a versatile and powerful tool for process parallelization widely used in scientific computing. This paper demonstrates its effective application in high-performance computing (HPC) environments, particularly focusing on its scalability and efficiency in executing large-scale high-throughput high-performance computing (HT-HPC) workflows. Through real-world examples, we highlight GNU Parallel’s performance across various HPC workloads, including GPU computing, container-based workloads, and node-local NVMe storage. Our results on two leading supercomputers, OLCF’s Frontier and NERSC’s Perlmutter, showcase GNU Parallel’s rapid process dispatching ability and its capacity to maintain low overhead even at extreme scales. We explore GNU Parallel’s application in massive parallel file transfers using a scheduled Data Transfer Node (DTN) cluster, emphasizing its broad utility in diverse scientific workflows. Beyond its direct application as a viable workflow manager, GNU Parallel can be employed in conjunction with other workflow systems as a "last-mile" parallelizing driver and as a quick prototyping tool to design and extract parallel profiles from application executions. We then argue that the potential for GNU Parallel to transform workflow management at extreme scales is substantial, paving the way for more efficient and effective scientific discoveries.

Maheshwari, Ketan↗

PIPES (Pipeline for Integrated Projects in Energy Systems) [SWR-24-89]

The Pipeline for Integrated Projects in Energy Systems (PIPES) is a comprehensive project, data, and workflow management tool designed for integrated modeling teams. PIPES facilitates the management of data requirements, tasks, and progress tracking, serving as a higher-level integration layer that works across various data and modeling software. This tool integrates models, data, and tools to perform large-scale, integrated analysis work at scale. PIPES is designed to streamline integrated modeling projects, enhance collaboration, and ensure the quality and efficiency of data management and workflow processes. https://github.com/nrel-pipes/pipes-api https://github.com/nrel-pipes/pipes-web https://github.com/nrel-pipes/nrel-pipes

Gu, Jianli↗

Consulting report on the NASA technology utilization network system

The purposes of this consulting effort are: (1) to evaluate the existing management and production procedures and workflow as they each relate to the successful development, utilization, and implementation of the NASA Technology Utilization Network System (TUNS) database; (2) to identify, as requested by the NASA Project Monitor, the strengths, weaknesses, areas of bottlenecking, and previously unaddressed problem areas affecting TUNS; (3) to recommend changes or modifications of existing procedures as necessary in order to effect corrections for the overall benefit of NASA TUNS database production, implementation, and utilization; and (4) to recommend the addition of alternative procedures, routines, and activities that will consolidate and facilitate the production, implementation, and utilization of the NASA TUNS database.

Hlava, Marjorie M. K.↗

A low-cost centralized HVAC control system solution for energy savings, load shedding, and improved maintenance

University campuses rely on centralized controls for managing and optimizing complex HVAC systems in larger buildings. However, most campuses also have many smaller buildings with packaged HVAC systems controlled by a stand-alone thermostat. Even when these distributed and often overlooked systems have modern programmable thermostats, they cannot be centrally monitored or controlled, and they are typically not programmed adequately. This paper describes the implementation of a low-cost centralized control solution for these systems serving smaller campus buildings, mostly under 5,000 sf and representative of light commercial spaces. Thanks to advances in technology spurred by residential and commercial IoT developments, simple networked thermostat solutions exist that can easily replace original thermostats, and, connect these systems to a web-based portal for monitoring and control. We show that, with small customizations, these platforms can be integrated into facility management workflows. Beyond the energy savings potential from improved scheduling and closer management of these systems, there are significant advantages for maintenance crews since these systems can now be monitored on smart phones or tablets. A grid-responsive load-shedding program has also been implemented for additional cost savings. The networked thermostats can also be connected to additional systems such as economizer controls for improved ventilation management and energy savings. With data from these systems integrated centrally, it can also be used for improved analytics and fault detection. A toolkit has been developed to share the program with other campuses, whether for energy savings, improved management of ventilation, or a more proactive maintenance approach.

Fauchier-Magnan, Nicolas↗

iDDS: intelligent distributed dispatch and scheduling for workflow orchestration

The intelligent distributed dispatch and scheduling (iDDS) service is a versatile workflow orchestration system designed for large-scale, distributed scientific computing. iDDS extends traditional workload and data management by integrating data-aware execution, conditional logic, and programmable workflows, enabling automation of complex and dynamic processing pipelines. Originally developed for the ATLAS experiment at the large hadron collider, iDDS has evolved into an experiment-agnostic platform that supports both template-driven workflows and a Function-as-a-Task model for Python-based orchestration. This paper presents the architecture and core components of iDDS, highlighting its scalability, modular message-driven design, and integration with systems such as PanDA and Rucio. We demonstrate its versatility through real-world use cases: fine-grained tape resource optimization for ATLAS, orchestration of large Directed Acyclic Graph (DAG) workflows for the Rubin Observatory, distributed hyperparameter optimization for machine learning applications, active learning for physics analyses, and AI-assisted detector design at the electron–ion collider. By unifying workload scheduling, data movement, and adaptive decision-making, iDDS reduces operational overhead and enables reproducible, high-throughput workflows across heterogeneous infrastructures. We conclude with current challenges and future directions, including interactive, cloud-native, and serverless workflow support.

97 MATHEMATICS AND COMPUTING↗

EARTHDATA PUB: A Data Publication Workflow Solution for NASA’s EOSDIS

Each NASA Distributed Active Archive Center (DAAC) faces the challenge of dealing with an increasingly diverse number of publishable data products from diverse data producers. Data producers, on the other hand, may experience pain points when interacting with the EOSDIS for the first time or when publishing different data at different DAACs. As a result, there has been a growing need to develop a common software framework that serves as a common interface for data producers, rigorously defines the data publication procedure for DAAC staff, facilitates the management of various data publication processes, and tracks the progress of data publication. This software should also account for the different configurations at different DAACs. Currently, two primary data publication workflow and tracking tools exist in operation at the EOSDIS: Semi-Automated ingest System (SAuS) and Data Publication workflow Portal (DAPPeR). However, neither tool is cloud-ready. Automated data processing could be managed by Cumulus, an EOSDIS cloud-based data ingest, archive and management system. However, Cumulus does not support manual tasks or on-premise implementations. We propose to develop the Earthdata Publication Minimum Viable Product (Earthdata Pub MVP) -- a cloud-hosted solution that works with both cloud and on-premise systems and implements the communications and exchange requirements generated by the Earthdata Pub information architecture team.

Rice, Justin L.↗

A2SD: Accelerating Scientific Innovation Through Autonomous Discovery Systems

The 2025 Advancing Autonomous Scientific Discovery (A2SD) workshop convened researchers from academia, national laboratories, and industry to explore the transformative role of autonomy in scientific discovery. The workshop highlighted a convergence of artificial intelligence, robotics, and computational workflows into autonomous systems capable of accelerating the scientific process. Presentations and discussions spanned autonomous experimentation, intelligent workflow orchestration, digital twins, and agent-based systems for managing complex research ecosystems. Key challenges discussed included interoperability across heterogeneous infrastructures, near real-time data management under FAIR principles, reproducibility, and the integration of human oversight. The workshop also emphasized the need for modular software interfaces, federated learning models, and education initiatives to support a next-generation scientific workforce.

Taufer, Michela [University of Tennessee, Knoxvill↗

kessel

Kessel is a tool to create and drive continuous integration (CI) and developer workflows through a unified interface across multiple code projects and environments. It serves as a driver and integration layer for build systems and package managers, providing a flexible library of reusable components to build and execute complex workflows consistently.

Berger, Richard [@lanl]↗

Exploration Technologies for Operations

Although the International Space Station (ISS) assembly has been completed, the Operations support teams continue to seek more efficient and effective ways to prepare for and conduct the ISS operations and future exploration missions beyond low earth orbit. This search for improvement has led to a significant collaboration between the NASA research and advanced software development community at NASA Ames Research Center and the Mission Operations community at NASA Johnson Space Center. Since 2001, NASA Ames Research Center has been developing and applying its advanced intelligent systems and human systems integration research to mission operations tools for several of the unmanned Mars missions operations. Since 2006, NASA Ames Research Center has also been developing and applying its advanced intelligent systems and human systems integration research to mission operations tools for manned operations support with the Mission Operations Directorate at NASA Johnson Space Center. This paper discusses the completion of the development and deployment of a variety of intelligent and human systems technologies adopted for manned mission operations. The technologies associated with the projects include advanced software systems for operations and human-centered computing. Human-centered computing looks to the processes and procedures that people do to perform any given job, then attempts to identify opportunities to improve these processes and procedures. In particular, for mission operations, improvements are quantified by specifically identifying how a tool can increase a persons efficiency, enhance a persons functional capability, andor improve the assurance of a persons decisions. The Ames development team has collaborated with the Mission Operations team to identify areas of efficiencies through technology infusion applications in support of the Plan, Train, and Fly activities of human-spaceflight mission operations. The specific applications discussed in this paper are in the areas of mission planning systems, mission operations design modeling and workflow automation, advanced systems monitoring, mission control technologies, search tools, training management tools, spacecraft solar array management, spacecraft power management, and spacecraft attitude planning. We discuss these specific projects between the Ames Research Center and the Johnson Space Centers Mission Operations Directorate, and how these technologies and projects are enhancing the mission operations support for the International Space Station. We also discuss the challenges, problems, and successes associated with long-distance and multi-year development projects between the research team at Ames and the Mission Operations customers at Johnson Space center. Finally, we discuss how these technology infusion applications and underlying technologies might be used in the future to support on-board operations of the crew and spacecraft systems as human exploration expands beyond low earth orbit to destinations in the solar system where communications delays will require more on-board autonomy and planning by the crew. Longer communications delays will require that the ground mission operations support will be primarily strategic in nature, while the tactical level of planning, systems monitoring and control, and failure analysisisolationrecovery will be the responsibility of both the spacecraft autonomous systems and the crew. Our expectation is that the technologies

mission operations↗