Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Workflow”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Orchestration of materials science workflows for heterogeneous resources at large scale

In the era of big data, materials science workflows need to handle large-scale data distribution, storage, and computation. Any of these areas can become a performance bottleneck. We present a framework for analyzing internal material structures (e.g., cracks) to mitigate these bottlenecks. We demonstrate the effectiveness of our framework for a workflow performing synchrotron X-ray computed tomography reconstruction and segmentation of a silica-based structure. Our framework provides a cloud-based, cutting-edge solution to challenges such as growing intermediate and output data and heavy resource demands during image reconstruction and segmentation. Specifically, our framework efficiently manages data storage, scaling up compute resources on the cloud. The multi-layer software structure of our framework includes three layers. A top layer uses Jupyter notebooks and serves as the user interface. A middle layer uses Ansible for resource deployment and managing the execution environment. A low layer is dedicated to resource management and provides resource management and job scheduling on heterogeneous nodes (i.e., GPU and CPU). At the core of this layer, Kubernetes supports resource management, and Dask enables large-scale job scheduling for heterogeneous resources. The broader impact of our work is four-fold: through our framework, we hide the complexity of the cloud’s software stack to the user who otherwise is required to have expertise in cloud technologies; we manage job scheduling efficiently and in a scalable manner; we enable resource elasticity and workflow orchestration at a large scale; and we facilitate moving the study of nonporous structures, which has wide applications in engineering and scientific fields, to the cloud. While we demonstrate the capability of our framework for a specific materials science application, it can be adapted for other applications and domains because of its modular, multi-layer architecture.

97 MATHEMATICS AND COMPUTING↗

SWARM: Reimagining scientific workflow management systems in a distributed world

Modern scientific workflows process massive amounts of data from diverse instruments and sensors, leveraging geographically distributed, heterogeneous compute and storage resources—from leadership-class systems to edge devices—connected by high-performance networks. The diversity of resources introduces challenges in harnessing their full potential, with resilience issues arising across applications, system software, networks, storage, and hardware. Today, workflow management systems (WMS) coordinate the execution of computation and data management tasks across target resources. However, WMS’s centralized nature makes them vulnerable to faults and scalability issues that may result in failures of entire computational campaigns. In conclusion, this paper introduces a novel agentic framework for workflow management, fully distributing and decentralizing the WMS functions and modeling them as swarm intelligence agents infused with advanced artificial intelligence solutions and traditional distributed computing algorithms that can make coordinated decisions in the presence of failures of the underlying cyberinfrastructure.

Swarm intelligence↗

NNSA/CEA Workflow (Workshop Report)

This report, prepared by the NNSA/CEA Workflows Working Group, briefly summarizes the presentations in the areas of domain specific workflows, end user environments, data management, job and resource management, and infrastructure, and then identifies six broad areas for potential collaboration. A key finding is that users could benefit from greater interoperability, compatibility, and composability of the workflow technologies under development and that point-to-point collaboration opportunities should be identified to explore these aspects.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Image processing workflow yielding high contrast synchrotron nanoscale computed tomography data from Ni-YSZ electrodes

The operating lifetime of Ni-YSZ fuel electrodes used in solid oxide electrolysis cells and fuel cells (SOECs and SOFCs) is limited by Ni redistribution, one of the primary degradation mechanisms that must be overcome to extend the longevity and maximize the performance of SOECs and SOFCs. To achieve this, 3D microstructural data is needed to relate both initial performance and performance loss over time to microstructural properties and their evolution throughout operation under various conditions. However, 3D microstructure data remains relatively scarce within the literature due to multiple challenges in acquiring and analyzing such data reliably. This work presents a workflow for acquiring and processing synchrotron X-ray nanoscale computed tomography (nano-CT) data from Ni-YSZ electrodes. Parameters for each step in the nano-CT workflow are described up to the final result (a 3D reconstruction), with particular emphasis on image alignment using freely available software. Following the results of a parametric sweep of the image alignment step, high contrast, low signal-to-noise 3D nano-CT data is obtained with relatively short compute times. While the exact methods best suited to samples with different microstructural qualities, or similar Ni-YSZ nano-CT data obtained from other sources may deviate from the solution found herein, this work also generalizes the decision points and evaluation of each step to provide a starting point to adapt this workflow to other datasets.

08 HYDROGEN↗

3-Dimensional Advanced Workflow

Have you ever been frustrated with the 2-dimensional nature of implementing complex Advanced Workflows (AWF) in your RSA Archer environment? Does the lack of concurrent, parallel workflows make you scratch your head and ask, ‘Why?’ Has your AWF gotten so complicated and convoluted that navigating through a maze-like plate of spaghetti looks simple and straightforward by comparison? If so, come listen to how NASA leverages Questionnaires to add an extra dimension and welcome simplification to complex Advanced Workflow configurations!

RSA Archer↗

Open Science Approach to Analyze Climate-Crop Relationships in the US Leveraging GES DISC and Galaxy Workflows

Understanding the intricate relationship between climate variability and agricultural production is crucial for ensuring food security. This study investigates the impact of climate parameters, such as temperature, precipitation, and soil moisture, on major US crop yields. Adopting an open science approach, the study analyzes the impact of climate on agricultural production in the United States. The Galaxy workflow engine serves as the primary tool for integrating climate data from the Goddard Earth Sciences Data and Information Services Center (GES DISC), retrieved via the Giovanni system, with yield statistics from the United States Department of Agriculture’s National Agricultural Statistics Service (USDA NASS). Extensions for reading, preprocessing, and analyzing external data have been developed, enabling the creation of workflows within the Galaxy platform. The development of a reproducible workflow allows for the calculation of seasonal climate averages, which are then assessed for their correlation with crop yields. This methodology ensures the replicability of the research, promoting transparency and collaboration in the scientific community. Correlational and regression analyses have been applied to different sub-zones and crops. The findings from this research offer valuable insights into the relationship between climate parameters and crop yields. These insights contribute to a deeper understanding of climate-crop relationships, providing a solid foundation for informed decision-making in the agricultural sector. The high correlation values indicate a significant relationship between climate parameters and crop yields, underscoring the importance of considering climate factors in agricultural planning and policymaking. This research also exemplifies the power of open science in advancing our understanding of complex environmental and agricultural phenomena. By leveraging open data and services, it provides a robust and replicable framework for future studies in this critical field.

Open science↗

Creating Apptainer Workflows with Docker-Compose-like Utilities

Creating Apptainer Workflows with Docker-Compose-like Utilities In this presentation, I will explore the utilization of a tool called process-compose, inspired by docker-compose, to create Apptainer-based services. This approach allows for easy deployment and management of fully containerized applications on High Performance Computing (HPC) systems without requiring elevated privileges. Benefits to the Ecosystem: By incorporating process-compose and Apptainer, I aim to address several key challenges in the HPC ecosystem: Simplified Workflow Management: Process-compose provides a user-friendly interface for defining and managing complex containerized application services, reducing the setup time and lowering the barrier to entry for new users. Enhanced Portability: Apptainer ensures that containerized applications can run consistently across different HPC environments, promoting greater portability and reducing compatibility issues. Process-compose is also a single binary that does not need to be installed by admin level users. Community Driven Solutions: This approach aligns with the goals of the High Performance Software Foundation (HPSF) to advance community-driven solutions. By sharing our experiences and insights, I hope to foster collaboration and innovation within the HPC community. Increased Productivity: The combination of process-compose and Apptainer streamlines the serve deployment process, allowing researchers and developers to focus more on their scientific work rather than the intricacies of system or service administration. Through this presentation, attendees will gain valuable insights into the practical implementation of containerized workflows on HPC systems, learn about the benefits of using process-compose and Apptainer, and understand how these tools can contribute to a more efficient HPC ecosystem.

97 - MATHEMATICS AND COMPUTING↗

Performance Improvements on SNS and HFIR Instrument Data Reduction Workflows Using Mantid

Performance of data reduction workflows at the High Flux Isotope Reactor (HFIR) and the Spallation Neutron Source (SNS) at Oak Ridge National Laboratory (ORNL) is mainly determined by the time spent loading raw measurement events stored in large and sparse datasets. This paper describes: (1) our long-term view to leverage SNS and HFIR data management needs with our experience at ORNL’s world-class high performance computing (HPC) facilities, and (2) our short-term efforts to speed up current workflows using Mantid, a data analysis and reduction community framework used across several neutron scattering facilities. We show that minimally invasive short-term improvements in metadata management have a moderate impact in speeding up current production workflows. We propose a more disruptive domain-specific solution: the No Cost Input Output (NCIO) framework, we provide an overview, the risks and challenges in NCIO’s adoption by HFIR and SNS stakeholders.

Godoy, William↗

Transitioning from File-Based HPC Workflows to Streaming Data Pipelines with openPMD and ADIOS2

This paper aims to create a transition path from file-based IO to streaming-based workflows for scientific applications in an HPC environment. By using the openPMP-api, traditional workflows limited by filesystem bottlenecks can be overcome and flexibly extended for in situ analysis. The openPMD-api is a library for the description of scientific data according to the Open Standard for Particle-Mesh Data (openPMD). Its approach towards recent challenges posed by hardware heterogeneity lies in the decoupling of data description in domain sciences, such as plasma physics simulations, from concrete implementations in hardware and IO. The streaming backend is provided by the ADIOS2 framework, developed at Oak Ridge National Laboratory. This paper surveys two openPMD-based loosely-coupled setups to demonstrate flexible applicability and to evaluate performance. In loose coupling, as opposed to tight coupling, two (or more) applications are executed separately, e.g. in individual MPI contexts, yet cooperate by exchanging data. This way, a streaming-based workflow allows for standalone codes instead of tightly-coupled plugins, using a unified streaming-aware API and leveraging high-speed communication infrastructure available in modern compute clusters for massive data exchange. We determine new challenges in resource allocation and in the need of strategies for a flexible data distribution, demonstrating their influence on efficiency and scaling on the Summit compute system. The presented setups show the potential for a more flexible use of compute resources brought by streaming IO as well as the ability to increase throughput by avoiding filesystem bottlenecks.

Poeschel, Franz↗

Toward an Autonomous Workflow for Single Crystal Neutron Diffraction

The operation of the neutron facility relies heavily on beamline scientists. Some experiments can take one or two days with experts making decisions along the way. Leveraging the computing power of HPC platforms and AI advances in image analyses, here we demonstrate an autonomous workflow for the single-crystal neutron diffraction experiments. The workflow consists of three components: an inference service that provides real-time AI segmentation on the image stream from the experiments conducted at the neutron facility, a continuous integration service that launches distributed training jobs on Summit to update the AI model on newly collected images, and a frontend web service to display the AI tagged images to the expert. Ultimately, the feedback can be directly fed to the equipment at the edge in deciding the next-step experiment without requiring an expert in the loop. With the analyses of the requirements and benchmarks of the performance for each component, this effort serves as the first step toward an autonomous workflow for real-time experiment steering at ORNL neutron facilities.

Yin, Junqi↗

Towards a Standard Process Management Infrastructure for Workflows Using Python

Orchestrating the execution of ensembles of processes lies at the core of scientific workflow engines on large scale parallel platforms. This is usually handled using platform-specific command line tools, with limited process management control and potential strain on system resources. The PMIx standard provides a uniform interface to system resources. The low level C implementation of PMIx has hampered its use in workflow engines, leading to the development of Python binding that has yet to gain traction. In this paper, we present our work to harden the PMIx Python client, demonstrating its usability using a prototype Python driver to orchestrate the execution of an ensemble of processes. We present experimental results using the prototype on the Summit supercomputer at Oak Ridge National Laboratory. This work lays the foundation for wider adoption of PMIx for workflow engines, and encourages wider support of more PMIx functionality in vendor provided system software stacks.

Elwasif, Wael↗

Optimizing Data Movement for GPU-Based In-Situ Workflow Using GPUDirect RDMA

The extreme-scale computing landscape is increasingly dominated by GPU-accelerated systems. At the same time, in-situ workflows that employ memory-to-memory inter-application data exchanges have emerged as an effective approach for leveraging these extreme-scale systems. In the case of GPUs, GPUDirect RDMA enables third-party devices, such as network interface cards, to access GPU memory directly and has been adopted for intra-application communications across GPUs. In this paper, we present an interoperable framework for GPU-based in-situ workflows that optimizes data movement using GPUDirect RDMA. Specifically, we analyze the characteristics of the possible data movement pathways between GPUs from an in-situ workflow perspective, and design a strategy that maximizes throughput. Furthermore, we implement this approach as an extension of the DataSpaces data staging service, and experimentally evaluate its performance and scalability on a current leadership GPU cluster. The performance results show that the proposed design reduces data-movement time by up to 53% and 40% for the sender and receiver, respectively, and maintains excellent scalability for up to 256 GPUs.

Zhang, Bo↗

Electron-phonon coupling from GW perturbation theory: Practical workflow combining BerkeleyGW, ABINIT, and EPW

Here we present a workflow of practical calculations of electron-phonon (e-ph) coupling with many-electron correlation effects included using the GW perturbation theory (GWPT). This workflow combines BerkeleyGW, ABINIT, and EPW software packages to enable accurate e-ph calculations at the GW self-energy level, going beyond standard calculations based on density functional theory (DFT) and density-functional perturbation theory (DFPT). This workflow begins with DFT and DFPT calculations (ABINIT) as starting point, followed by GW and GWPT calculations (BerkeleyGW) for the quasiparticle band structures and e-ph matrix elements on coarse electron k- and phonon q-grids, which are then interpolated to finer grids through Wannier interpolation (EPW) for computations of various e-ph coupling determined physical quantities such as the electron self-energies or solutions of anisotropic Eliashberg equations, among others. A gauge-recovering symmetry unfolding technique is developed to reduce the computational cost of GWPT (as well as DFPT) while fulfilling the gauge consistency requirement for Wannier interpolation.

74 ATOMIC AND MOLECULAR PHYSICS↗

Automated workflow for non-empirical Wannier-localized optimal tuning of range-separated hybrid functionals

Here, we introduce an automated workflow for generating non-empirical Wannier-localized optimally-tuned screened range-separated hybrid (WOT-SRSH) functionals. WOT-SRSH functionals have been shown to yield highly accurate fundamental band gaps, band structures, and optical spectra for bulk and 2D semiconductors and insulators. Our workflow automatically and efficiently determines the WOT-SRSH functional parameters for a given crystal structure and composition, approximately enforcing the correct screened long-range Coulomb interaction and an ionization potential ansatz. In contrast to previous manual tuning approaches, our tuning procedure relies on a new search algorithm that only requires a few hybrid functional calculations with minimal user input. We demonstrate our workflow on 23 previously studied semiconductors and insulators, reporting the same high level of accuracy. By automating the tuning process and improving its computational efficiency, the approach outlined here enables applications of the WOT-SRSH functional to compute spectroscopic and optoelectronic properties for a wide range of materials.

Gant, Stephen E. [University of California, Berkel↗

DTLMod: A simulation framework for in situ workflow optimization

In situ processing workflows have become essential for coping with the explosion in data volume and velocity in large-scale scientific computing, providing domain scientists with early insights at runtime. Multiple frameworks implement this paradigm through a data transport layer (DTL), offering different data access modes and deployment schemes, but researchers currently lack the appropriate tools to assess design and deployment options before committing to costly real experiments. We introduce DTLMod, an open-source simulated DTL that enables performance evaluation of in situ workflow configurations at scale. Built on SimGrid, it links into any SimGrid-based simulator and is available in C++ and Python. We evaluate DTLMod along four axes: scalability (tens of thousands of simulated processes across interconnected clusters in seconds, with linear memory scaling), versatility (three implementation variants trading fidelity for speed), accuracy (simulated times faithfully reflecting real behavior), and practical utility (two use cases demonstrating evidence-based workflow design decisions).

Suter, Fred [ORNL] (ORCID:0000000319021955)↗

A Unified Workflow for Sensitivity-Based Kinetic Analysis in Microkinetic Models

Degrees of rate control (DRC), apparent activation energies, and apparent reaction orders are established local sensitivity diagnostics for interpreting microkinetic models, but applying them routinely to large mechanisms often requires substantial reaction-specific bookkeeping, perturbation design, and postprocessing. Here, in this study, we present a unified derivative-based workflow that evaluates these quantities from a single compiled reaction-network model and target-rate definition. For any user-provided microkinetic model, the workflow compiles the mechanism into stoichiometrically consistent mass-action rate equations, solves the surface dynamics, and uses automatic differentiation to compute sensitivities with respect to rate constants, temperature, and gas partial pressures. By combining their calculations in the same framework, the workflow clearly demonstrates the relationships between different DRCs and the apparent activation energy. Using existing examples of propylene partial oxidation and methane oxidation on Pd(100), we verify expected transient redistribution of rate control, distinguish net Campbell DRCs from one-sided directional sensitivities, and show how apparent activation energy can be reconstructed either from one-sided DRCs or from state-based DRCs while critical mechanistic insights are obtained consistently. In the methane oxidation case, a pathway-subset test further illustrates how a simplified mechanism preserves key kinetic signatures of a full model, showing the potential of our user-friendly tool for model construction beyond kinetic analysis.

36 MATERIALS SCIENCE↗

Automated Adsorption Workflow for Semiconductor Surfaces and the Application to Zinc Telluride

Surface adsorption is a crucial step in numerous processes, including heterogeneous catalysis, where the adsorption of key species is often used as a descriptor of efficiency. We present here an automated adsorption workflow for semiconductors which employs density functional theory calculations to generate adsorption data in a high-throughput manner. Starting from a bulk structure, the workflow performs an exhaustive surface search, followed by an adsorption structure construction step, which generates a minimal energy landscape to determine the optimal adsorbate-surface distance. An extensive set of energy-based, charge-based, geometric, and electronic descriptors tailored toward catalysis research are computed and saved to a personal user database. Finally, the application of the workflow to zinc telluride, a promising CO 2 reduction photocatalyst, is presented as a case study to illustrate the capabilities of this method and its potential as a material discovery tool.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Davis Computational Spectroscopy Workflow—From Structure to Spectra

Here, we describe an automated workflow that connects a series of atomic simulation tools to investigate the relationship between atomic structure, lattice dynamics, materials properties, and inelastic neutron scattering (INS) spectra. Starting from the atomic simulation environment (ASE) as an interface, we demonstrate the use of a selection of calculators, including density functional theory (DFT) and density functional tight binding (DFTB), to optimize the structures and calculate interatomic force constants. We present the use of our workflow to compute the phonon frequencies and eigenvectors, which are required to accurately simulate the INS spectra in crystalline solids like diamond and graphite as well as molecular solids like rubrene. We have also implemented a machine-learning force field based on Chebyshev polynomials called the Chebyshev interaction model for efficient simulation (ChIMES) to improve the accuracy of the DFTB simulations. We then explore the transferability of our DFTB/ChIMES models by comparing simulations derived from different training sets. We show that DFTB/ChIMES demonstrates ~100× reduction in computational expense while retaining most of the accuracy of DFT as well as yielding high accuracy for different materials outside of our training sets. The DFTB/ChIMES method within the workflow expands the possibilities to use simulations to accurately predict materials properties of increasingly complex structures that would be unfeasible with ab initio methods.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗