Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Workflow Management Systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

kessel

Kessel is a tool to create and drive continuous integration (CI) and developer workflows through a unified interface across multiple code projects and environments. It serves as a driver and integration layer for build systems and package managers, providing a flexible library of reusable components to build and execute complex workflows consistently.

Berger, Richard [@lanl]↗

Workflows Community Summit 2022: A Roadmap Revolution

Scientific workflows have become integral tools in broad scientific computing use cases. Science discovery is increasingly dependent on workflows to orchestrate large and complex scientific experiments that range from the execution of a cloud-based data preprocessing pipeline to multi-facility instrument-to-edge-to-HPC computational workflows. Given the changing landscape of scientific computing (often referred to as a computing continuum) and the evolving needs of emerging scientific applications, it is paramount that the development of novel scientific workflows and system functionalities seek to increase the efficiency, resilience, and pervasiveness of existing systems and applications. Specifically, the proliferation of machine learning/artificial intelligence (ML/AI) workflows, need for processing large-scale datasets produced by instruments at the edge, intensification of near real-time data processing, support for long-term experiment campaigns, and emergence of quantum computing as an adjunct to HPC, have significantly changed the functional and operational requirements of workflow systems. Workflow systems now need to, for example, support data streams from the edge-to-cloud-to-HPC, enable the management of many small-sized files, allow data reduction while ensuring high accuracy, orchestrate distributed services (workflows, instruments, data movement, provenance, publication, etc.) across computing and user facilities, among others. Further, to accelerate science, it is also necessary that these systems implement specifications/standards and APIs for seamless (horizontal and vertical) integration between systems and applications, as well as enable the publication of workflows and their associated products according to the FAIR principles.

97 MATHEMATICS AND COMPUTING↗

Exploration Technologies for Operations

Although the International Space Station (ISS) assembly has been completed, the Operations support teams continue to seek more efficient and effective ways to prepare for and conduct the ISS operations and future exploration missions beyond low earth orbit. This search for improvement has led to a significant collaboration between the NASA research and advanced software development community at NASA Ames Research Center and the Mission Operations community at NASA Johnson Space Center. Since 2001, NASA Ames Research Center has been developing and applying its advanced intelligent systems and human systems integration research to mission operations tools for several of the unmanned Mars missions operations. Since 2006, NASA Ames Research Center has also been developing and applying its advanced intelligent systems and human systems integration research to mission operations tools for manned operations support with the Mission Operations Directorate at NASA Johnson Space Center. This paper discusses the completion of the development and deployment of a variety of intelligent and human systems technologies adopted for manned mission operations. The technologies associated with the projects include advanced software systems for operations and human-centered computing. Human-centered computing looks to the processes and procedures that people do to perform any given job, then attempts to identify opportunities to improve these processes and procedures. In particular, for mission operations, improvements are quantified by specifically identifying how a tool can increase a persons efficiency, enhance a persons functional capability, andor improve the assurance of a persons decisions. The Ames development team has collaborated with the Mission Operations team to identify areas of efficiencies through technology infusion applications in support of the Plan, Train, and Fly activities of human-spaceflight mission operations. The specific applications discussed in this paper are in the areas of mission planning systems, mission operations design modeling and workflow automation, advanced systems monitoring, mission control technologies, search tools, training management tools, spacecraft solar array management, spacecraft power management, and spacecraft attitude planning. We discuss these specific projects between the Ames Research Center and the Johnson Space Centers Mission Operations Directorate, and how these technologies and projects are enhancing the mission operations support for the International Space Station. We also discuss the challenges, problems, and successes associated with long-distance and multi-year development projects between the research team at Ames and the Mission Operations customers at Johnson Space center. Finally, we discuss how these technology infusion applications and underlying technologies might be used in the future to support on-board operations of the crew and spacecraft systems as human exploration expands beyond low earth orbit to destinations in the solar system where communications delays will require more on-board autonomy and planning by the crew. Longer communications delays will require that the ground mission operations support will be primarily strategic in nature, while the tactical level of planning, systems monitoring and control, and failure analysisisolationrecovery will be the responsibility of both the spacecraft autonomous systems and the crew. Our expectation is that the technologies

mission operations↗

EnergyPlus-MCP: A model-context-protocol server for ai-driven building energy modeling

Traditional building energy modeling with the EnergyPlus building performance simulation engine requires domain expertise, programming skills, and intensive manual efforts limiting its effective adoption. This paper introduces EnergyPlus-MCP, the first open-source Model Context Protocol (MCP) server specifically designed for EnergyPlus simulation workflows, establishing a new foundational infrastructure for AI-driven building energy modeling. The MCP server implements a layered architecture with 35 specialized tools spanning model management, editing and analysis, HVAC and other systems configuration inspection, and simulation execution, enabling Large Language Models to interact with EnergyPlus through conversational interfaces. The server addresses critical workflow barriers by automating model validation, streamlining energy efficiency measures modification, and providing intelligent output management with interactive visualization. Through practical demonstrations using a multi-zone building retrofit analysis, we show how the EnergyPlus-MCP server significantly reduces manual efforts while maintaining full simulation rigor. By providing accessible natural language interfaces to sophisticated building energy analysis, this approach enables scalable deployment of simulation expertise across public and private organizations, educational institutions, and research teams, fundamentally transforming traditional building energy modeling practices.

AI↗

STNS01-44 BEE – FY22-1: Enhanced BEE Client [Slide]

BEE provides a portable, modular, HPC-focused workflow engine capable of managing containerized applications at scale. In FY22 BEE is completing enhancements and refinements that will complete the major development work of the workflow system. The first major milestone is the development of a graphical client. The second milestone will be the ability for BEE to automatically restart checkpointed tasks. The final milestone for FY22 will be the ability to launch and manage multiple simultaneous workflows.

97 MATHEMATICS AND COMPUTING↗

Osprey Framework v0.2.2

The Alpha Berkeley Framework is a software architecture for building agentic AI systems that coordinate multi-step workflows in scientific and industrial environments. It is based on a plan-first orchestration model, where natural language requests are translated into execution plans with explicit dependencies and optional human approval. The framework includes capability classification, which selects relevant tools on a per-task basis to keep orchestration efficient as the number of available tools grows. It incorporates task extraction methods that compress conversational context and integrate external resources such as databases, APIs, and knowledge bases into structured, machine-readable tasks. Execution is supported by modular services with checkpointing, artifact management, and error handling, allowing workflows to be paused, inspected, and resumed. The system is designed for deployment in production environments, supporting both local and containerized execution as well as integration with HPC clusters. Interfaces include command-line tools, browser-based workflows, and containerized services. The framework has been demonstrated in tutorial examples and deployed at the Advanced Light Source, where it coordinates accelerator control and analysis workflows.

Hellert, Thorsten [Lawrence Berkeley National Labo↗

Creating Apptainer Workflows with Docker-Compose-like Utilities

Creating Apptainer Workflows with Docker-Compose-like Utilities In this presentation, I will explore the utilization of a tool called process-compose, inspired by docker-compose, to create Apptainer-based services. This approach allows for easy deployment and management of fully containerized applications on High Performance Computing (HPC) systems without requiring elevated privileges. Benefits to the Ecosystem: By incorporating process-compose and Apptainer, I aim to address several key challenges in the HPC ecosystem: Simplified Workflow Management: Process-compose provides a user-friendly interface for defining and managing complex containerized application services, reducing the setup time and lowering the barrier to entry for new users. Enhanced Portability: Apptainer ensures that containerized applications can run consistently across different HPC environments, promoting greater portability and reducing compatibility issues. Process-compose is also a single binary that does not need to be installed by admin level users. Community Driven Solutions: This approach aligns with the goals of the High Performance Software Foundation (HPSF) to advance community-driven solutions. By sharing our experiences and insights, I hope to foster collaboration and innovation within the HPC community. Increased Productivity: The combination of process-compose and Apptainer streamlines the serve deployment process, allowing researchers and developers to focus more on their scientific work rather than the intricacies of system or service administration. Through this presentation, attendees will gain valuable insights into the practical implementation of containerized workflows on HPC systems, learn about the benefits of using process-compose and Apptainer, and understand how these tools can contribute to a more efficient HPC ecosystem.

97 - MATHEMATICS AND COMPUTING↗

Centralized Interactive Phenomics Resource: an integrated online phenomics knowledgebase for health data users

Development of clinical phenotypes from electronic health records (EHRs) can be resource intensive. Several phenotype libraries have been created to facilitate reuse of definitions. However, these platforms vary in target audience and utility. Here, we describe the development of the Centralized Interactive Phenomics Resource (CIPHER) knowledgebase, a comprehensive public-facing phenotype library, which aims to facilitate clinical and health services research. The platform was designed to collect and catalog EHR-based computable phenotype algorithms from any healthcare system, scale metadata management, facilitate phenotype discovery, and allow for integration of tools and user workflows. Phenomics experts were engaged in the development and testing of the site. The knowledgebase stores phenotype metadata using the CIPHER standard, and definitions are accessible through complex searching. Phenotypes are contributed to the knowledgebase via webform, allowing metadata validation. Data visualization tools linking to the knowledgebase enhance user interaction with content and accelerate phenotype development. The CIPHER knowledgebase was developed in the largest healthcare system in the United States and piloted with external partners. The design of the CIPHER website supports a variety of front-end tools and features to facilitate phenotype development and reuse. Health data users are encouraged to contribute their algorithms to the knowledgebase for wider dissemination to the research community, and to use the platform as a springboard for phenotyping. CIPHER is a public resource for all health data users available at https://phenomics.va.ornl.gov/ which facilitates phenotype reuse, development, and dissemination of phenotyping knowledge.

60 APPLIED LIFE SCIENCES↗

Extreme-scale workflows: A perspective from the JLESC international community

The Joint Laboratory for Extreme-Scale Computing (JLESC) focuses on software challenges in high-performance computing systems to meet the needs of today’s science campaigns, which often require large resources, consist of multiple tasks, and generate vast amounts of data. In this context, extreme-scale workflows have been the key factor in enabling scientific discoveries by helping scientists automate the dependencies and data exchanges between workflow tasks, instead of managing those manually. Here, in this paper, we present representative extreme-scale workflows and feature workflow systems developed by JLESC participating institutions. We present lessons learned while developing these tools, alongside with the open challenges and future research directions in the field of extreme-scale workflows.

97 MATHEMATICS AND COMPUTING↗

Bridging paradigms: Designing for HPC-Quantum convergence

Here, this paper presents a comprehensive software stack architecture for integrating quantum computing (QC) capabilities with High-Performance Computing (HPC) environments. While quantum computers show promise as specialized accelerators for scientific computing, their effective integration with classical HPC systems presents significant technical challenges. We propose a hardware-agnostic software framework that supports both current noisy intermediate-scale quantum devices and future fault-tolerant quantum computers, while maintaining compatibility with existing HPC workflows. The architecture includes a quantum gateway interface, standardized APIs for resource management, and robust scheduling mechanisms to handle both simultaneous and interleaved quantum–classical workloads. Key innovations include: (1) a unified resource management system that efficiently coordinates quantum and classical resources, (2) a flexible quantum programming interface that abstracts hardware-specific details, (3) A Quantum Platform Manager API that simplifies the integration of various quantum hardware systems, and (4) a comprehensive tool chain for quantum circuit optimization and execution. We demonstrate our architecture through implementation of quantum–classical algorithms, including the variational quantum linear solver, showcasing the framework’s ability to handle complex hybrid workflows while maximizing resource utilization. This work provides a foundational blueprint for integrating QC capabilities into existing HPC infrastructures, addressing critical challenges in resource management, job scheduling, and efficient data movement between classical and quantum resources.

97 MATHEMATICS AND COMPUTING↗

TPSAS-NF1676L-33992-DND

The CERES Science Team integrates and fuses observations from 6 CERES instruments aboard the Terra, Aqua, S-NPP, and NOAA-20 missions with data from more than 20 other unique data sources. Following the November 2017 launch of CERES Flight Model 6 (FM6) onboard NOAA20, CERES has now amassed over 80 instrument-years of valuable Earth radiation budget data. The rapidly growing volume of CERES data coupled with the introduction of new data products alongside improvements to existing science algorithms fosters the requirement for faster, more flexible, and scalable data production and orchestration. New virtualized, cloud-centric compute hardware hosted by the NASA Langley Research Center’s (LaRC) Atmospheric Sciences Data Center (ASDC) provides an ideal environment for these ever-increasing data production demands for CERES. This poster discusses updates to the implementation of the CERES Data Management Team’s (DMT) CERES AuTomAted job Loading sYSTem (CATALYST), a custom data processing workflow engine for CERES, to use on-demand computing resources to perform automated CERES data production processing in a Linux-based container environment. Linux containers provide CERES the flexibility to build multiple production environments in containers tailored for specific workloads and allow effortless provisioning of resources based on the CERES Science Team’s data production requirements.

Thomas N. Hillyer↗

A Task Based Approach for Co-Scheduling Ensemble Workloads on Heterogeneous Nodes

Scientific workflows consist of multiple, connected applications, with data and results flowing from one to another in a pipeline. Traditionally, such workflows are executed in sequential order, storing intermediate data in storage disks. Co-scheduling application workflows concurrently on the same compute nodes would greatly reduce the cost of moving data to/from storage and allow real-time analysis of intermediate results. Nevertheless, most parallel programming runtimes do not allow seamless integration of various applications in a scientific workflow, in part due to the complexity of managing data and resources. The situation is even more complicated for heterogeneous systems. In this work we extend the Minos Computing Library (MCL) runtime to accelerate pipe-lined and parallel workloads where multiple applications are running in the same system. MCL’s asynchronous task library and runtime dynamically manages resources to allow co-scheduling of multiple processes sharing heterogeneous resources. In addition, we design a custom ex- tension of the Open Compute Language (OpenCL) to enable multiple processes to share device memory. We enable MCL to coordinate these shared buffers to allow for easy, fast data sharing between applications. Using malleable micro-benchmarks and two application workflows that combine scientific simulation and AI-based analysis, we show that our method outperforms traditional approaches.

Index Terms—Parallel systems, Scheduling and Task ↗

Towards a Standard Process Management Infrastructure for Workflows Using Python

Orchestrating the execution of ensembles of processes lies at the core of scientific workflow engines on large scale parallel platforms. This is usually handled using platform-specific command line tools, with limited process management control and potential strain on system resources. The PMIx standard provides a uniform interface to system resources. The low level C implementation of PMIx has hampered its use in workflow engines, leading to the development of Python binding that has yet to gain traction. In this paper, we present our work to harden the PMIx Python client, demonstrating its usability using a prototype Python driver to orchestrate the execution of an ensemble of processes. We present experimental results using the prototype on the Summit supercomputer at Oak Ridge National Laboratory. This work lays the foundation for wider adoption of PMIx for workflow engines, and encourages wider support of more PMIx functionality in vendor provided system software stacks.

Elwasif, Wael↗

The future of self-driving laboratories: from human in the loop interactive AI to gamification

Recent developments in artificial intelligence (AI) and machine learning (ML), implemented through self-driving laboratories (SDLs), are rapidly creating unprecedented opportunities for the accelerated discovery and optimization of materials. This paper provides a joint analysis of SDLs from both academic and industry perspectives, highlighting the importance of integrating human intelligence in these systems. It discusses the necessity of careful planning in SDL design across physical, data, and workflow dimensions, including instrumental setup, experimental workflow, data management, and human–SDL interaction. The significance of integrating human input within SDLs, especially as the focus shifts from individual tools and tasks to the creation and management of complex workflows, is emphasized. The paper stresses on the crucial role of reward function design in developing forward-looking workflows and examines the interplay between hardware evolution, ML application across chemical processes, and the influence of reward systems in research. Ultimately, the article advocates for a future where SDLs blend human intuition in hypothesis formulation with AI's precision, speed, and data-handling capabilities.

97 MATHEMATICS AND COMPUTING↗

AI-assisted detector design for the EIC (AID(2)E)

Artificial Intelligence is poised to transform the design of complex, large-scale detectors like ePIC at the future Electron Ion Collider. Featuring a central detector with additional detecting systems in the far forward and far backward regions, the ePIC experiment incorporates numerous design parameters and objectives, including performance, physics reach, and cost, constrained by mechanical and geometric limits. This project aims to develop a scalable, distributed AI-assisted detector design for the EIC (AID(2)E), employing state-of-the-art multiobjective optimization to tackle complex designs. Supported by the ePIC software stack and using G EANT 4 simulations, our approach benefits from transparent parameterization and advanced AI features. The workflow leverages the PanDA and iDDS systems, used in major experiments such as ATLAS at CERN LHC, the Rubin Observatory, and sPHENIX at RHIC, to manage the compute intensive demands of ePIC detector simulations. Tailored enhancements to the PanDA system focus on usability, scalability, automation, and monitoring. Ultimately, this project aims to establish a robust design capability, apply a distributed AI-assisted workflow to the ePIC detector, and extend its applications to the design of the second detector (Detector-2) in the EIC, as well as to calibration and alignment tasks. Additionally, we are developing advanced data science tools to efficiently navigate the complex, multidimensional trade-offs identified through this optimization process.

97 MATHEMATICS AND COMPUTING↗

Accelerating Resilience of the Community through Holistic Engagement and use of Renewables (ARCHER) Planning Framework

The primary objective of the Accelerating Resilience of the Community through Holistic Engagement and Use of Renewables (ARCHER) initiative was to identify and incorporate the unique variations in energy burden, social vulnerability, living conditions, and access to essential services that differ across communities. By accounting for these localized factors—down to the neighborhood level—the project supports more targeted and effective investments in community resilience. The framework seeks to establish practical planning guidance, methods, and performance measures for community energy resilience, integrate community-level and electric utility system resilience planning, and assess its effectiveness through comparison with conventional and operational planning approaches. A key component of the project was its data exchange platform, which is used to evaluate and demonstrate the tools, methodologies, and planning approaches developed through ARCHER. This open-source platform enables developers and vendors of distribution and outage management systems to build upon the research by incorporating its concepts into their own tools and workflows. This capability is enabled by the transparent availability of data, functional requirements, and the underlying information model. The project yielded several important insights. First, meaningful engagement with communities is essential to achieving comprehensive resilience outcomes. Second, resilience planning is most effective when electric grid considerations and broader community needs are addressed in a coordinated manner. Third, the use of platforms that allow for real-time input from communities can enhance utility responsiveness during restoration activities. Fourth, a structured and systematic planning approach can successfully translate ARCHER concepts into practice. Fifth, the development of an integrated metric that reflects both grid performance and community impacts provides a more holistic basis for evaluating resilience. Finally, incorporating community engagement and equity considerations into grid operations is critical, particularly during severe weather events that result in extended outages.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Data-driven modeling to enhance municipal water demand estimates in response to dynamic climate conditions

Altered precipitation and temperature patterns from a changing climate will affect supply, demand, and overall municipal water system operations throughout the arid western U.S. While supply forecasts leverage hydrological models to connect climate influences with surface water availability, demand forecasts typically estimate water use independent of climate and other externalities. Stemming from an increased focus on seasonal water demand management, we use the Salt Lake City, Utah municipal water system as a test bed to assess model accuracy versus complexity trade-offs between simple climate-independent econometric-based models and complex climate-sensitive data-driven models to average to extreme wet and dry climate conditions—representative of a new climate normal. Here, the climate-independent model displayed low performance during extreme dry conditions with predictions exceeding 90% and 40% of the observed monthly and seasonal volumetric demands, respectively, which we attribute to insufficient model complexity. The climate-sensitive models displayed greater accuracy in all conditions, with an ordinary least squares model demonstrating a measurable reduction in prediction bias (3.4% vs. -27.3%) and RMSE (74.0 lpcd vs. 294 lpcd) compared to the climate-independent model. The climate-sensitive workflow increased model accuracy and characterized climate-demand interactions, demonstrating a novel tool to enhance water system management.

54 ENVIRONMENTAL SCIENCES↗

Web Time-Management Tool

Oak Grove Reactor, developed by Oak Grove Systems, is a new software program that allows users to integrate workflow processes. It can be used with portable communication devices. The software can join e-mail, calendar/scheduling and legacy applications into one interactive system via the web. Priority tasks and due dates are organized and highlighted to keep the user up to date with developments. Reactor works with existing software and few new skills are needed to use it. Using a web browser, a user can can work on something while other users can work on the same procedure or view its status while it is being worked on at another site. The software was developed by the Jet Propulsion Lab and originally put to use at Johnson Space Center.

Source record↗