Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “scientific workflow”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Enhancing Monte Carlo Workflows for Nuclear Reactor Analysis with Metamodel-Driven Modeling

Monte Carlo codes are essential components of many reactor physics simulation workflows as high-fidelity continuous-energy neutron transport solvers. Among Monte Carlo radiation transport codes, MCNP is particularly notable due to its diverse simulation capabilities, large user base, and long validation history. Despite being a powerful simulation tool, MCNP provides limited capabilities to allow automated execution, model transformation, or support for user-defined logic and abstractions that limit its compatibility with modern workflows. Here, to better integrate MCNP into a modern scientific workflow, we have developed an intuitive yet full-featured MCNP Application Program Interface (API) in Python, named MCNPy, which provides a specialized set of classes for MCNP input development. Moreover, to guarantee that our reading, writing, and modeling capabilities remain self-consistent (and to render the huge scope of the MCNP API manageable), we have adopted a strategy of model-driven software development in which a generalized model of the MCNP input format has been created. From this generalized model, or “metamodel,” problem-specific implementations such as an engine for input validation or a codebase for programmatic operations may be automatically generated. Since MCNPy primarily acts as a Python front-end to the underlying Java API that directly interfaces with the metamodel, it is intrinsically linked to the metamodel and thus remains maintainable. With MCNPy, users can programmatically read, write, and modify any syntactically valid MCNP input file regardless of its origin. These capabilities allow users to automate complicated tasks like design optimization and model translation for nuclear systems. As examples, this work demonstrates the use of MCNPy to find the critical radius of a plutonium sphere and to translate a 9000+ line MCNP input file into a corresponding OpenMC model.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

LaRC SmartLab Apps For Instrument Control And Data Processing: Optical Micrometer Data Visualizer

The LaRC Smart Lab applications are a series of software tools to greatly enhance researcher efficiency by streamlining and automating workflows. Python scripts and applications are increasingly being used in scientific workflows, including for instrument control and data processing. Interactive Python scripting environments such as Jupyter Lab provide powerful tools for using Python. In some use cases, the development of standalone applications with dedicated graphical user interfaces (GUIs) can enhance the utility of the code and open it up to more users, including non-programmers. Here, we describe a GUI based optical micrometer data visualization application developed as part of the LaRC SmartLab project. We highlight its use in visualizing experimental data and briefly discuss its implementation to give pointers to programmers who wish develop work based on this application's or similar co de.

LaRC SmartLab↗

LaRC SmartLab Apps For Instrument Control and Data Processing: Laboratory Environment Monitor

The LaRC SmartLab applications are a series of software tools to greatly enhance researcher efficiency by streamlining and automating workflows. Python scripts and applications are increasingly being used in scientific workflows, including for instrument control and data processing. Interactive Python scripting environments such as JupyterLab provide powerful tools for using Python. In some use cases, the development of standalone applications with dedicated graphical user interfaces can enhance the utility of the code and open it up to more users, including non-programmers. Here, we describe a Python based application for communicating with, and displaying data from, iTHX Temperature, Humidity, and Dew Point probes. We discuss the set up and use of the application as well as its implementation. We also highlight the use of Simulated probes to enable users and developers to familiarize with or debug the application, even when they do not have access to the physical hardware in the laboratory.

LaRC SmartLab↗

Towards a Standard Process Management Infrastructure for Workflows Using Python

Orchestrating the execution of ensembles of processes lies at the core of scientific workflow engines on large scale parallel platforms. This is usually handled using platform-specific command line tools, with limited process management control and potential strain on system resources. The PMIx standard provides a uniform interface to system resources. The low level C implementation of PMIx has hampered its use in workflow engines, leading to the development of Python binding that has yet to gain traction. In this paper, we present our work to harden the PMIx Python client, demonstrating its usability using a prototype Python driver to orchestrate the execution of an ensemble of processes. We present experimental results using the prototype on the Summit supercomputer at Oak Ridge National Laboratory. This work lays the foundation for wider adoption of PMIx for workflow engines, and encourages wider support of more PMIx functionality in vendor provided system software stacks.

Elwasif, Wael↗

Toward Real-Time Analysis of Synchrotron Micro-Tomography Data: Accelerating Experimental Workflows with AI and HPC

ynchrotron light sources are routinely used to perform imaging experiments. In this paper, we review the relevant computational stages, identify bottlenecks, and highlight future opportunities to streamline data acquisition for experimental microscopy workflows. We demonstrate our preliminary exploration with an end-to-end scientific workflow on Summit based on micro-computed tomography data. Computational elements include: 1) reconstruction of volumetric image data; 2) denoising with deep neural networks; and 3) non-local means based segmentation and quantitative analysis.

Mcclure, James↗

NREL Stratus - Enabling Workflows to Fuse Data Streams, Modeling, Simulation, and Machine Learning

Integrating cloud services into advanced computing facilities provides significant new capabilities over focusing solely on traditional high performance computing (HPC) workloads. This brings complementary capabilities as well as enabling new focused roles for HPC. They are especially potent for workflows that fuse data streams, modeling and simulation ('modsim') and machine learning. A key challenge to adopting a hybrid edge-cloud-HPC model is to align optimal capability, data, and user intent on the right resources for each step in a workflow.?The NREL Stratus service provides a basis for this: Stratus layers capabilities needed to make?cloud services accessible to a lab-based scientific community on commercial offerings, and; currently supports upwards of 200 projects ranging from IOT integration to traditional modeling and simulation. This provides a real-world inventory of scientific workflow elements. A growing knowledge base enables placing these elements appropriately between the edge, cloud, and traditional HPC. This paper outlines a vision via reference architecture and the application of that architecture in a typical workflow highlighting multiple components: sensor data intake, cleaning and transforming (edge/cloud suitable); generation of synthetic data through modsim, computationally heavy ML training and hyperparameter optimization (HPC suitable), and; inference and deployment (cloud ideal). Every step in such a workflow involves a cost-benefit analysis regarding the data movement, computational efficiency, availability, latency, and resource capabilities. The reference architecture and examples outlined allow for understanding new opportunities in the context of emerging workflows that combine IOT, cloud, and HPC to bolster scientific productivity.

AI↗

28 NREL Stratus - Enabling Workflows to Fuse Data Streams, Modeling, Simulation, and Machine Learning: Preprint

Integrating cloud services into advanced computing facilities provides significant new capabilities over focusing solely on traditional high performance computing (HPC) workloads. This brings complementary capabilities as well as enabling new focused roles for HPC. They are especially potent for workflows that fuse data streams, modeling and simulation ('modsim') and machine learning. A key challenge to adopting a hybrid edge-cloud-HPC model is to align optimal capability, data, and user intent on the right resources for each step in a workflow.?The NREL Stratus service provides a basis for this: Stratus layers capabilities needed to make?cloud services accessible to a lab-based scientific community on commercial offerings, and; currently supports upwards of 200 projects ranging from IOT integration to traditional modeling and simulation. This provides a real-world inventory of scientific workflow elements. A growing knowledge base enables placing these elements appropriately between the edge, cloud, and traditional HPC. This paper outlines a vision via reference architecture and the application of that architecture in a typical workflow highlighting multiple components: sensor data intake, cleaning and transforming (edge/cloud suitable); generation of synthetic data through modsim, computationally heavy ML training and hyperparameter optimization (HPC suitable), and; inference and deployment (cloud ideal). Every step in such a workflow involves a cost-benefit analysis regarding the data movement, computational efficiency, availability, latency, and resource capabilities. The reference architecture and examples outlined allow for understanding new opportunities in the context of emerging workflows that combine IOT, cloud, and HPC to bolster scientific productivity.

AI↗

Composable optimization and control toolkit for scientific applications

Applications of Artificial Intelligence (AI) and Machine Learning (ML) can improve the computational efficiency and scientific research output. In order to improve interoperability and reuse of AI/ML software, a composable approach is required. This talk presents a composable approach for scientific workflow development that allows seamless integration of various modules developed by independent researchers. These practices will reduce redundant software development by allowing re-use of workflow modules across projects, teams, departments and facilities. We will present three use cases that follow the composable approach namely, Scientific Optimization and Control Toolkit (SOCT), SciDAC QuantOm workflow, and JLab Nuclear Physics experimental workflows. This talk will dive deeper into SOCT and present the details of the composable code development for optimization and control algorithms using reinforcement learning.

Rajput, Kishansingh↗

Deploying Machine Learning Workflows into HPC environment

Outline: Workflows Overview; Common Workflow Language (CWL), Example of a CWL; BEE Overview; Machine Learning Components; Machine Learning Scientific Workflow using CWL → A new test case for BEE; Discussion: Benefits and Caveats of current ML workflow; Conclusion.

97 MATHEMATICS AND COMPUTING↗

Real-World Experiences Adopting Workflows at Exascale on the ExaAM Project

The purpose of this study is to discuss the experiential lessons associated with adopting scientific workflows in the Exascale Additive Manufacturing project (ExaAM) through the lens of Perceived Characteristic of Innovation (PCI). Besides the implementation, the factors we considered critical to the adoption of the workflow are provenance, sustainable automation, implementation challenges, and integration/compatibility challenges. Through conversations and interviews among the program managers, project leads, and software engineers, we have developed critical insight and strategies to overcome the obstacles and augment the successful adoption and long-term use of these workflows in ExaAM and beyond. We hope our work will pave the way for others in the research community to develop and use workflows in their respective science domains.

Malviya, Addi Thakur↗

RISE: Reducing I/O Contention in Staging-based Extreme-Scale In-situ Workflows

While in-situ workflow formulations have addressed some of the data-related challenges associated with extreme-scale scientific workflows, these workflows involve complex interactions and different modes of data exchange. In the context of increasing system complexity, such workflows present significant resource management challenges, requiring complex cost-performance tradeoffs. This paper presents RISE, an intelligent staging-based data management middleware, which builds on the DataSpaces framework and performs intelligent scheduling of data management operations to reduce I/O contention. In RISE, data are always written immediately to local buffers to reduce the effect of the transfer impact upon application performance. RISE identifies applications’ data access patterns and moves data towards data consumers only when the network is expected to be idle, reducing the impact of asynchronous background data movement upon critical data read/write requests. Here, we experimentally demonstrate that RISE can take advantage of staging nodes to offload data during writes without degrading application data movement performance.

97 MATHEMATICS AND COMPUTING↗

From Edge to HPC: Investigating Cross-Facility Data Streaming Architectures

In this paper, we investigate three cross-facility data streaming architectures, Direct Streaming (DTS), Proxied Streaming (PRS), and Managed Service Streaming (MSS). We examine their architectural variations in data flow paths and deployment feasibility, and detail their implementation using the Data Streaming to HPC (DS2HPC) architectural framework and the SciStream memory-to-memory streaming toolkit on the production-grade Advanced Computing Ecosystem (ACE) infrastructure at Oak Ridge Leadership Computing Facility (OLCF). We present a workflow-specific evaluation of these architectures using three synthetic workloads derived from the streaming characteristics of scientific workflows. Through simulated experiments, we measure streaming throughput, round-trip time, and overhead under work sharing, work sharing with feedback, and broadcast and gather messaging patterns commonly found in AI-HPC communication motifs. Our study shows that DTS offers a minimal-hop path, resulting in higher throughput and lower latency, whereas MSS provides greater deployment feasibility and scalability across multiple users but incurs significant overhead. PRS lies in between, offering a scalable architecture whose performance matches DTS in most cases.

George, Anjus [ORNL] (ORCID:0000000179737061)↗

Management and Storage of Scientific Data

Scientific discoveries rely heavily on efficient access, search, and management of massive data sets. Data management technologies have, for decades, provided foundational capabilities for scientific computing. Just as storage, input/output (I/O), and data management have been fundamental to simulation-based science for many years, so too are capable data-management technologies key to the success of today’s scientific workflows utilizing data intensive and machine learning (ML) techniques. The Department of Energy, Office of Science, Advanced Scientific Computing Research (ASCR) program has invested broadly in data-management research focused on high-performance computing (HPC) systems, from parallel file systems that store data to application software that makes these systems more productive. Still, advances in technology combined with growing diversity of supported science strongly motivate continued investment in this area. In January 2022, ASCR convened a workshop to identify priority research directions in the area of data management for high-performance and scientific computing. Attendees were challenged to identify promising approaches that would support the breadth of the DOE mission, including the explosion of artificial intelligence (AI) uses and the growing needs of experimental and observational science. Technological and science drivers were identified and considered as they relate to key aspects of data management such as interfaces, architectural design, and FAIR principles (Findable, Accessible, Interoperable, and Reusable). The thoughts of the workshop participants were distilled into a set of four priority research directions with the potential for high impact on DOE science. These research directions are summarized in the following pages.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Designing workflows for materials characterization

Experimental science is enabled by the combination of synthesis, imaging, and functional characterization organized into evolving discovery loop. Synthesis of new material is typically followed by a set of characterization steps aiming to provide feedback for optimization or discover fundamental mechanisms. However, the sequence of synthesis and characterization methods and their interpretation, or research workflow, has traditionally been driven by human intuition and is highly domain specific. Here, we explore concepts of scientific workflows that emerge at the interface between theory, characterization, and imaging. In this study, we discuss the criteria by which these workflows can be constructed for special cases of multiresolution structural imaging and functional characterization, as a part of more general material synthesis workflows. Some considerations for theory–experiment workflows are provided. We further pose that the emergence of user facilities and cloud labs disrupts the classical progression from ideation, orchestration, and execution stages of workflow development. To accelerate this transition, we propose the framework for workflow design, including universal hyperlanguages describing laboratory operation, ontological domain matching, reward functions and their integration between domains, and policy development for workflow optimization. These tools will enable knowledge-based workflow optimization; enable lateral instrumental networks, sequential and parallel orchestration of characterization between dissimilar facilities; and empower distributed research.

36 MATERIALS SCIENCE↗

Joint Genome Institute Analysis Workflow Service (JAWS) v2.7

The U.S. Department of Energy Joint Genome Institute (JGI) has developed the JGI Analysis Workflow Service (JAWS) as a distributed framework to run computational workflows across diverse high-performance computing (HPC) and cloud environments. JAWS enhances the reusability, scalability, and robustness of scientific workflows by orchestrating data movement, code execution, and results retrieval across multiple DOE facilities. At its core, JAWS integrates the Cromwell workflow engine to run workflows expressed in the Workflow Description Language (WDL), ensuring portability and interoperability. To provide consistent runtime environments, JAWS employs container technologies such as Shifter, Apptainer, and Docker. Workflow tasks are managed via HTCondor on HPC backends, while Globus ensures secure and efficient data transfer between sites. JAWS is deployed as a multi-site workflow manager across national laboratory computing facilities, with dedicated instances supporting community projects such as the National Microbiome Data Collaborative (NMDC) and KBase. This distributed, service-oriented architecture enables users to "write once, run anywhere," providing scalable, production-quality workflow execution.

Kirton, Edward↗

Position Papers for the ASCR Workshop on the Management and Storage of Scientific Data

The purpose of this workshop is to identify priority research directions in the area of data management for high-performance and scientific computing above and beyond HPC’s traditional "the parallel file system is the data-management system" model. Supporting the breadth of the DOE mission, including the explosion of AI uses and the growing needs of experimental and observational science, motivates revisiting our assumptions about data management. There are many facets of this topic to explore including: (1) Interfaces for accessing data that resides on traditional persistent storage as well as memory devices; (2) Storage-system architecture design that supports scientific workflows on varied hierarchical storage and networking devices; (3) Devising metadata management infrastructure to support FAIR principles (Findability, Accessibility, Interoperability, and Reusability); (4) Capturing provenance information about scientific data; (5) Utilizing AI to learn I/O patterns of emerging workloads for efficient data management; (6) Providing data management support for AI and complex workflows; and (7) Understanding the overlap between traditional storage systems and I/O (SSIO) efforts and data management. While the program committee has identified these topics as important areas for discussion, we welcome position papers from the community that propose additional topics of interest for discussion at the workshop. The workshop agenda will include breakout sessions for discussing these and selected topic areas to inform priority research directions for data management for high-performance and scientific computing.

97 MATHEMATICS AND COMPUTING↗