Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “compute workflow”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Enabling HPC Scientific Workflows for Serverless

The convergence of edge computing, big data analytics, and AI with traditional scientific calculations is increasingly being adopted in HPC workflows. Workflow management systems are crucial for managing and orchestrating these complex computational tasks. However, it is difficult to identify patterns within the growing population of HPC workflows. Serverless has emerged as a novel computing paradigm, offering dynamic resource allocation, quick response time, fine-grained resource management and auto-scaling. In this paper, we propose a framework to enable HPC scientific workflows on serverless. Our approach integrates a widely used traditional HPC workflow generator with an HPC serverless workflow management system to create benchmark suites of scientific workflows with diverse characteristics. These workflows can be executed on different serverless platforms. We comprehensively compare executing workflows on traditional local containers and serverless computing platforms. Our results show that serverless can reduce CPU and memory usage respectively by 78.11% and 73.92% without compromising performance.

Andrei da silva, Anderson↗

A dynamic solvent chamber propagation estimation framework using RNN for warm solvent injection in heterogeneous reservoirs

Warm solvent injection (WSI), injecting low-temperature solvent into formations to reduce the viscosity of heavy oil, is a clean technology for heavy oil production through reducing greenhouse gas emissions and water usage. The success of WSI operation depends on the uniform development and propagation of solvent chambers in reservoirs. However, reservoir heterogeneity stemming from shale barriers plays a detrimental role in the conformance of solvent chamber development and oil production rate. In this work, we developed a novel recurrent neural network (RNN)-based framework with the capability of efficiently tracking and estimating the solvent chamber positions in heterogeneous reservoirs based on only production time-series data. The developed estimation model utilizes the “sequence-to-sequence" mapping methodology to correlate observed production time-series sequence and solvent chamber edge sequence via a long short-term memory (LSTM) algorithm. The trained RNN models exhibit high accuracy, evidenced by the predicted dynamic solvent chamber locations match the corresponding true locations from numerical simulation, with a high coefficient of determination (R 2 ) and a low mean squared error. Specifically, the achieved R 2 values exceed 0.98 on both the training and testing data. The developed RNN-based workflow was tested via several cases from both regularly- and irregularly-shaped shale barriers, and the results were promising. The predicted solvent chambers showed strong agreement with those obtained from numerical simulations. The major benefits of this workflow include reducing computational time and saving overall monitoring and tracking costs for conventional techniques. In conclusion, the present work would provide a good demonstration of the capability of practical integration of machine learning methods in solving engineering problems.

58 GEOSCIENCES↗

Data Science Meets Physical Organic Chemistry

At the heart of synthetic chemistry is the holy grail of predictable catalyst design. In particular, researchers involved in reaction development in asymmetric catalysis have pursued a variety of strategies toward this goal. This is driven by both the pragmatic need to achieve high selectivities and the inability to readily identify why a certain catalyst is effective for a given reaction. While empiricism and intuition have dominated the field of asymmetric catalysis since its inception, enantioselectivity offers a mechanistically rich platform to interrogate catalyst-structure response patterns that explain the performance of a particular catalyst or substrate. In the early stages of an asymmetric reaction development campaign, the overarching mechanism of the reaction, catalyst speciation, the turnover limiting step, and many other details are unknown or posited based on related reactions. Considering the unclear details leading to a successful reaction, initial enantioselectivity data are often used to intuitively guide the ultimate direction of optimization. However, if the conditions of the Curtin-Hammett principle are satisfied, then measured enantioselectivity can be directly connected to the ensemble of diastereomeric transition states (TSs) that lead to the enantiomeric products, and the associated free energy difference between competing TSs (ΔΔ G ‡ = - RT ln[( S )/( R )], where ( S ) and ( R ) represent the concentrations of the enantiomeric products). We, and others, speculated that this important piece of information can be leveraged to guide reaction optimization in a quantitative way. Although traditional linear free energy relationships (LFERs), such as Hammett plots, have been used to illuminate important mechanistic features, we sought to develop data science derived tools to expand the power of LFERs in order to describe complex reactions frequently encountered in modern asymmetric catalysis. Specifically, we investigated whether enantioselectivity data from a reaction can be quantitatively connected to the attributes of reaction components, such as catalyst and substrate structural features, to harness data for asymmetric catalyst design. In this context, we developed a workflow to relate computationally derived features of reaction components to enantioselectivity using data science tools. The mathematical representation of molecules can incorporate many aspects of a transformation, such as molecular features from substrate, product, catalyst, and proposed transition states. Statistical models relating these features to reaction outputs can be used for various tasks, such as performance prediction of untested molecules. Perhaps most importantly, statistical models can guide the generation of mechanistic hypotheses that are embedded within complex patterns of reaction responses. Overall, merging traditional physical organic experiments with statistical modeling techniques creates a feedback loop that enables both evaluation of multiple mechanistic hypotheses and future catalyst design. In this Account, we highlight the evolution and application of this approach in the context of a collaborative program based on chiral phosphoric acid catalysts (CPAs) in asymmetric catalysis.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Machine-Learning-Driven Discovery of Water Splitting BaFe 2 O 4 and Human-in-the-Loop Improvement via Al-Substitution for Increased Thermal Stability

Thermochemical hydrogen (TCH) production offers a promising method for converting thermal energy into hydrogen fuel through heat-driven redox cycles of metal oxides. Here, in this work a defect graph neural network (dGNN) was used to predict oxygen vacancy formation energies ΔH V O combined with Materials Project predictions of oxygen chemical potential stability to screen candidate oxides via high-throughput database analysis. BaFe 2 O 4 was identified as a promising material for experimental validation based on its predicted ΔH V O , oxygen chemical potential stability range, and potential for tunable substitutions to improve thermal properties. Experimental validation using thermogravimetric analysis (TGA), stagnation flow reactor (SFR), X-ray diffraction (XRD), and electron microscopy confirmed positive water-splitting behavior but also revealed limitations in thermal stability under aggressive reduction conditions. To address this, a human-in-the-loop modification strategy was employed introducing Al substitution in BaFe 2–x Al x O 4 ; this modification improves thermal stability, alters the crystal structure and enhances overall performance. These results demonstrate a combined computational and experimental workflow in which machine learning accelerates identification of promising candidates, while targeted experimental design enables optimization of functional performance. This approach advances the development of robust, cost-effective TCH materials and highlights the importance of integrating data-driven discovery with human-guided materials design in paving the way for scalable hydrogen production technologies.

organic↗

MVP: a modular viromics pipeline to identify, filter, cluster, annotate, and bin viruses from metagenomes

While numerous computational frameworks and workflows are available for recovering prokaryote and eukaryote genomes from metagenome data, only a limited number of pipelines are designed specifically for viromics analysis. With many viromics tools developed in the last few years alone, it can be challenging for scientists with limited bioinformatics experience to easily recover, evaluate quality, annotate genes, dereplicate, assign taxonomy, and calculate relative abundance and coverage of viral genomes using state-of-the-art methods and standards. Here, we describe Modular Viromics Pipeline (MVP) v.1.0, a user-friendly pipeline written in Python and providing a simple framework to perform standard viromics analyses. MVP combines multiple tools to enable viral genome identification, characterization of genome quality, filtering, clustering, taxonomic and functional annotation, genome binning, and comprehensive summaries of results that can be used for downstream ecological analyses. Overall, MVP provides a standardized and reproducible pipeline for both extensive and robust characterization of viruses from large-scale sequencing data including metagenomes, metatranscriptomes, viromes, and isolate genomes. As a typical use case, we show how the entire MVP pipeline can be applied to a set of 20 metagenomes from wetland sediments using only 10 modules executed via command lines, leading to the identification of 11,656 viral contigs and 8,145 viral operational taxonomic units (vOTUs) displaying a clear beta-diversity pattern. Further, acting as a dynamic wrapper, MVP is designed to continuously incorporate updates and integrate new tools, ensuring its ongoing relevance in the rapidly evolving field of viromics. MVP is available at https://gitlab.com/ccoclet/mvp and as versioned packages in PyPi and Conda.

59 BASIC BIOLOGICAL SCIENCES↗

Moment Tensor Inversion Toolkit

The MTINV toolkit (2002-present) is a collection of computer codes and applications written to invert for the moment tensor of a seismic source given the three components of ground motion recorded at regional seismic stations (e.g., Ichinose et al., 2003). The computer codes and workflow are organized to generate moment tensor solutions for a range of source depths and origin times because of the trade-off between these two quantities. The metric used is the variance reduction and variance reduction modulated by the percent double-couple to determine the best-fit moment-tensor solution. We can solve for a deviatoric moment tensor with a constraint added for no isotropic component although this constraint can be lifted for estimating the full moment tensor like mining collapses or explosion sources.

Ichinose, GeneA↗

Review of Tools for Modeling Core Radial Expansion in Liquid Metal-Cooled Fast Reactors

Core radial expansion in liquid-metal cooled fast reactor systems is a well-known phenomenon that produces strong reactivity feedback effects. It is crucial from a safety standpoint to design a reactor with the proper core restraint system to guide expansion of the fuel assemblies into a formation that produces negative reactivity feedback in accident conditions. However, the modeling of core radial expansion and its associated reactivity feedback is extremely challenging. Liquid metal-cooled fast reactor designers (both industry and commercial) have pointed to improved modeling of core radial expansion as a key modeling challenge for these types of reactors. The Department of Energy Nuclear Energy Advanced Modeling and Simulation (DOE-NEAMS) program has an important role to play in helping to develop a robust, predictive tool for this phenomenon. A literature review has been performed to identify current simulation capabilities applicable to the prediction of core radial expansion in liquid metal-cooled fast reactors and the resultant reactivity feedback effects. An overview as well as a detailed description of the multi-physics phenomena underlying core radial expansion is given in the first two chapters. The objective of this report is to identify existing software that has been or could be applied to the core radial expansion problem, and then to identify possible paths to develop a robust, coupled Multiphysics tool for advanced analysis, in which all three physics (radiation transport, structural mechanics, thermal fluids) are taken into account considering feedback effects from each physics to the other. Several computer codes and workflows have been developed in the United States and abroad to perform individual aspects of radial expansion calculations or to perform loosely coupled analysis. No code system currently exists which tightly and robustly couples the neutronics, thermal hydraulics, and thermal mechanical physics with enough detail to fully resolve the complex core radial expansion reactivity feedback effects. The future availability of such a code system is of vital importance to fully understanding the reactivity feedback effects that occur due to radial expansion, and consequently to optimizing the design of the core restraint system. A potential code development path forward using MOOSE-based tools is suggested in order to leverage the natural tight coupling and robustness of MOOSE-based applications. Additionally, potential gaps for modeling and common assumptions are listed which require further investigation to ensure accuracy for prediction of deformations.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Enabling Low-Overhead HT-HPC Workflows at Extreme Scale using GNU Parallel

GNU Parallel is a versatile and powerful tool for process parallelization widely used in scientific computing. This paper demonstrates its effective application in high-performance computing (HPC) environments, particularly focusing on its scalability and efficiency in executing large-scale high-throughput high-performance computing (HT-HPC) workflows. Through real-world examples, we highlight GNU Parallel’s performance across various HPC workloads, including GPU computing, container-based workloads, and node-local NVMe storage. Our results on two leading supercomputers, OLCF’s Frontier and NERSC’s Perlmutter, showcase GNU Parallel’s rapid process dispatching ability and its capacity to maintain low overhead even at extreme scales. We explore GNU Parallel’s application in massive parallel file transfers using a scheduled Data Transfer Node (DTN) cluster, emphasizing its broad utility in diverse scientific workflows. Beyond its direct application as a viable workflow manager, GNU Parallel can be employed in conjunction with other workflow systems as a "last-mile" parallelizing driver and as a quick prototyping tool to design and extract parallel profiles from application executions. We then argue that the potential for GNU Parallel to transform workflow management at extreme scales is substantial, paving the way for more efficient and effective scientific discoveries.

Maheshwari, Ketan↗

Accelerating Scientific Workflows on HPC Platforms with In Situ Processing

Scientific workflows drive most modern large-scale science breakthroughs by allowing scientists to define their computations as a set of jobs executed in a given order based on their data dependencies. Workflow management systems (WMSs) have become key to automating scientific workflows-executing computational jobs and orchestrating data transfers between those jobs running on complex high-performance computing (HPC) platforms. Traditionally, WMSs use files to communicate between jobs: a job writes out files that are read by other jobs. However, HPC machines face a growing gap between their storage and compute capabilities. To address that concern, the scientific community has adopted a new approach called in situ, which bypasses costly parallel filesystem I/O operations with faster in-memory or in-network communications. When using in situ approaches, communication and computations can be interleaved. In this work, we leverage the Decaf in situ dataflow framework to accelerate task-based scientific workflows managed by the Pegasus WMS, by replacing file communications with faster MPI messaging. We propose a new execution engine that uses Decaf to manage communications within a sub-workflow (i.e., set of jobs) to optimize inter-job communications. We consider two workflows in this study: (i) a synthetic workflow that benchmarks and compares file- and MPI-based communication; and (ii) a realistic bioinformatics workflow that computes mu-tational overlaps in the human genome. Experiments show that in situ communication can improve the bioinformatics workflow execution time by 22% to 30% compared with file communication. Our results motivate further opportunities and challenges for bridging traditional WMSs with in situ frameworks.

Decaf↗

Cross-Facility Orchestration of Electrochemistry Experiments and Computations

Instrument-computing ecosystems supporting automated electrochemical workflows typically require the integration of disparate instruments such as syringe pump, fraction collector, and potentiostat, all connected to an electrochemical cell. These specialized instruments with custom software and interfaces are not typically designed for network integration and remote automation. We developed a networked ecosystem of these instruments and computing platforms, which includes software to enable automated workflow orchestration from remote computers. Specifically, we developed Python wrappers of APIs and custom Pyro client-server modules to support remote operation of these instruments over the ecosystem network. Herein, we describe a specific workflow for generating and validating voltammogram (I-V) measurements of an electrolyte solution pumped into the electrochemical cell. We demonstrate the orchestration of this workflow which is composed using a Jupyter notebook and executed on a remote computer.

Al Najjar, Anees↗

Deep learning-accelerated 3D carbon storage reservoir pressure forecasting based on data assimilation using surface displacement from InSAR

Fast forecasting of the reservoir pressure distribution during geologic carbon storage (GCS) by assimilating monitoring data is a challenging problem. Due to high drilling cost, GCS projects usually have spatially sparse measurements from few wells, leading to high uncertainties in reservoir pressure prediction. To address this challenge, we use low-cost Interferometric Synthetic-Aperture Radar (InSAR) data as monitoring data to infer reservoir pressure build up. We develop a deep learning-accelerated workflow to assimilate surface displacement maps interpreted from InSAR and to forecast dynamic reservoir pressure. Employing an Ensemble Smoother Multiple Data Assimilation (ES-MDA) framework, the workflow updates three-dimensional (3D) geologic properties and predicts reservoir pressure with quantified uncertainties. We use a synthetic commercial-scale GCS model with bimodally distributed permeability and porosity to demonstrate the efficacy of the workflow. A two-step CNN-PCA approach is employed to parameterize the bimodal fields. The computational efficiency of the workflow is boosted by two residual U-Net based surrogate models for surface displacement and reservoir pressure predictions, respectively. The workflow can complete data assimilation and reservoir pressure forecasting in half an hour on a personal computer.

25 ENERGY STORAGE↗

Fast Reactor Physics Model Verification Studies using ARC and PyARC Workflows

PyARC was recently developed at Argonne National Laboratory to automate many of the tasks required in the ARC (Argonne Reactor Computation) fast reactor simulation workflow, from input file generation, code execution, data transfer between ARC codes, and output postprocessing. PyARC will likely be the path forward to train new users of the ARC codes with the goal of wide adoption by the national laboratories, academia, and industry. In particular, for the ANL-JAEA collaboration under the Civil Nuclear Working Group (CNWG) project agreement NE-01, PyARC will be used to model the Joyo and EBR-II reactors for comparisons with measured data and calculated results from JAEA (Task 3: Fast Reactor Fuel and Core). As an additional avenue for verification and validation, this report investigates the use of PyARC towards a variety of existing ARC-based reactor models, in order to understand its efficacy in replicating the behavior of base ARC codes and better understand any limitations within modeling realistic fast reactor problems. To this end, PyARC was used to model the Joyo MKI, RBEC Benchmark-M, PRISM Mod-B, and EBR-II Run 138B cores, and its results were compared to those from existing ARC-based models. It was found that for hexagonal-based geometries PyARC was able to replicate the behavior of ARC codes to within 10 pcm for small reactor cores, and ~150pcm difference in eigenvalue for larger cores. These discrepancies are attributed primarily to differences in local mesh refinement options between ARC and PyARC, which currently cannot be resolved with PyARC’s latest version (1.6.0). In some of these cases, PyARC was used to model steady-state problems with initial core compositions originating from a prior REBUS depletion calculation. While PyARC was not designed to support such steady-state calculations, workarounds were applied to replicate the behavior of ARC-based calculations as closely as possible. Thus, these results demonstrate the wide extent to which they can be applied to fast reactor problems while still providing immense benefit to the user in terms of automating and standardizing common routines within the fast reactor analysis workflow. This study concluded that PyARC will be suitable for modeling the steady-state conditions of the EBR-II and Joyo fast reactors as part of the CNWG project agreement.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Enabling Seamless Transitions from Experimental to Production HPC for Interactive Workflows

The evolving landscape of scientific computing requires seamless transitions from experimental to production HPC environments for interactive workflows. This paper presents a structured transition pathway developed at OLCF that bridges the gap between development testbeds and production systems. We address both technological and policy challenges, introducing frameworks for data streaming architectures, secure service interfaces, and adaptive resource scheduling for time-sensitive workloads and improved HPC interactivity. Our approach transforms traditional batch-oriented HPC into a more dynamic ecosystem capable of supporting modern scientific workflows that require near real-time data analysis, experimental steering, and cross-facility integration.

Etz, Brian [ORNL] (ORCID:0000000208554863)↗

Applications of Deep Learning to physics workflows

Modern large-scale physics experiments create datasets with sizes and streaming rates that can exceed those from industry leaders such as Google Cloud and Netflix. Fully processing these datasets requires both sufficient compute power and efficient workflows. Recent advances in Machine Learning (ML) and Artificial Intelligence (AI) can either improve or replace existing domain-specific algorithms to increase workflow efficiency. Not only can these algorithms improve the physics performance of current algorithms, but they can often be executed more quickly, especially when run on coprocessors such as GPUs or FPGAs. In the winter of 2023, MIT hosted the Accelerating Physics with ML at MIT workshop, which brought together researchers from gravitational-wave physics, multi-messenger astrophysics, and particle physics to discuss and share current efforts to integrate ML tools into their workflows. The following white paper highlights examples of algorithms and computing frameworks discussed during this workshop and summarizes the expected computing needs for the immediate future of the involved fields.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Adaptive elasticity policies for staging-based in situ visualization

In situ processing aims to alleviate the growing gap between computation and I/O capabilities by performing data processing close to the data source. In situ processing is widely used to process data generated by multiple data sources, including observation data from edge devices or scientific observational facilities and the simulation data generated by scientific computation on a high-performance computing (HPC) platform. For a scientific workflow that is run on an HPC platform and composed of a simulation program and an in situ data analytics or visualization (abbreviated as ana/vis) task, there is an implicit assumption that the computing resources assigned to the workflow keep static during the workflow execution. However, with the converging trend between the HPC and cloud computing platform, running the in situ ana/vis task in an elastic way is promising to decrease its overhead and improve its resource utilization rate. Resource elasticity represents the ability to change resource configurations such as the number of computing nodes/processes during workflow execution. An elastic job may dynamically adjust resource configurations; it may use a few resources at the beginning and more resources toward the end of the job when interesting data appear. However, it is hard to predict a priori how many computing nodes/processes need to be added/removed during the workflow execution to adapt to changing workflow needs. How to efficiently guide elasticity operations, such as growing or shrinking the number of processes used for in situ analysis during workflow execution, is an open-ended research question. In this article, we present adaptive elasticity policies that adopt workflow runtime information collected during workflow execution to predict how to trigger the addition/removal of processes in order to minimize in situ processing overhead. Taking in situ visualization tasks as an example, we integrate the presented elasticity policies into a staging-based elastic workflow and evaluate its efficiency in multiple elasticity scenarios. Compared with the situation without elasticity or with a static elasticity policy that uses a fixed number of processes for each rescaling operation, the adaptive elasticity policy can save overhead in finding a proper resource configuration and improve resource utilization efficiency. Furthermore, one experiment illustrates that the adaptive elasticity policy saves 41% of core-hours compared with the situation without the resource elasticity.

97 MATHEMATICS AND COMPUTING↗

ESnet-JLab FPGA Accelerated Transport (data plane) [EJFAT (udplb)] v1.0

The ESnet-JLab FPGA Accelerated Transport system is a solution for streaming high-speed scientific measurement data from Data Acquisition Systems (DAQs) to high-performance computing facilties. It is generally compatible with many science workflows, and makes no assumptions about the specifics of any particular experiment. This program (udplb) implements the data plane portion of the EJFAT system. It is an FPGA design that rewrites and forwards data packets from a UDP-based scientific workflow to high-performance compute nodes. It depends on another program (udplbd, disclosed separately) to implement the control system.

Bengough, Peter [Malleable Networks, Inc.]↗

Exascale workflow applications and middleware: An ExaWorks retrospective

Exascale computers offer transformative capabilities to combine data-driven and learning-based approaches with traditional simulation applications to accelerate scientific discovery and insight. However, these software combinations and integrations are difficult to achieve due to the challenges of coordinating and deploying heterogeneous software components on diverse and massive platforms. Here, we present the ExaWorks project, which addresses many of these challenges. We developed a workflow Software Development Toolkit (SDK), a curated collection of workflow technologies that can be composed and interoperated through a common interface, engineered following current best practices, and specifically designed to work on HPC platforms. ExaWorks also developed PSI/J, a job management abstraction API, to simplify the construction of portable software components and applications that can be used over various HPC schedulers. The PSI/J API is a minimal interface for submitting and monitoring jobs and their execution state across multiple and commonly used HPC schedulers. We also describe several leading and innovative workflow examples of ExaWorks tools used on DOE leadership platforms. Furthermore, we discuss how our project is working with the workflow community, large computing facilities, and HPC platform vendors to address the requirements of workflows sustainably at the exascale.

97 MATHEMATICS AND COMPUTING↗