Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Workflow”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

STNS01-21 BEE - FY21 P6-2: Archive, clone, and re-run workflows [Slide]

BEE will give ECP a tool that great simplifies the deployment of containerized workflows on the next generation of pre-exascale and exascale systems, as well as public and private clouds. BEE allows scientists to describe their workflow using the Common Workflow Language and then deploy that workflow across the entire spectrum of systems without having to learn the specifics of each container runtime, HPC resource manager, or cloud API. BEE also streamlines the curation and sharing of common workflows among the scientific community.

97 MATHEMATICS AND COMPUTING↗

Integrated End-to-end Performance Prediction and Diagnosis for Extreme Scientific Workflows (IPPD) (Final Report)

This report details the accomplishments from the ASCR funded project “Integrated End-to-end Performance Prediction and Diagnosis for Extreme Scientific Workflows” under the award numbers FWP-66406 and DE-SC0012630, with a focus on the UC San Diego (Award No. DE-SC0012630) part of the accomplishments. We refer to the project as IPPD. The main activities of IPPD were centered on the development and integration of provenance information to capture empirically workflow information and identify the sources of bottlenecks as well as variability, and modeling and simulation to provide insights and predictively explore multiple scenarios for development of advanced techniques for optimization of resources and workflow execution. After a Phase I of the project, a Phase II effort focused on three major aspects: a) observe how data is generated, distributed, and used; b) analyze how data is (repeatedly) consumed with a focus both on repeated patterns and anomalies; and c) explore how to optimize data motion. The project leveraged and extended our existing tools with new research and demonstrated our work on the Belle II workflow suite as well as on workflows from NSLS-II.

97 MATHEMATICS AND COMPUTING↗

Predictive Capability Maturity Model Demonstration for Cylindrical Cavity Coupling Using Gemma in the Next Generation Workflow

The predictive capability maturity model (PCMM) uses the expert elicitation process to generate credibility evidence for a particular analysis. To ensure Gemma has the capability to efficiently produce this credibility evidence, next generation workflows (NGW) are created for the solution verification, calibration/validation, and input uncertainty quantification portions of the PCMM assessment. These workflows are then used on the Higgins cylinder problem, which is representative of applications involving external-to-internal electromagnetic field coupling through a slot. The uncertainties calculated using these workflows are then used to calculate the validation comparison error and the validation uncertainty for the model following the American Society of Mechanical Engineers (ASME) verification and validation (V&V) 20 standard. These workflows will enable analysts to iterate each element of PCMM more efficiently than if completed without using a NGW workflow. An example of this iterative process is shown in Section 7.2.

42 ENGINEERING↗

Lambda-PFLOTRAN 1.0: a workflow for incorporating organic matter chemistry informed by ultra high resolution mass spectrometry into biogeochemical modeling

Abstract. Organic matter (OM) composition plays a central role in microbial respiration of dissolved organic matter and subsequent biogeochemical reactions. Here, a direct connection of organic matter chemistry and thermodynamics to reactive transport simulators has been achieved through the newly developed Lambda-PFLOTRAN workflow tool that succinctly incorporates carbon chemistry data generated from Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) into reaction networks to simulate organic matter degradation and the resulting biogeochemistry. Lambda-PFLOTRAN is a Python-based workflow, executed through a Jupyter notebook interface, that digests raw FTICR-MS data, develops a representative reaction network based on substrate-explicit thermodynamic modeling (also termed lambda modeling due to its key thermodynamic parameter λ used therein), and completes a biogeochemical simulation with the open source, reactive flow and transport code PFLOTRAN. The workflow consists of the following five steps: configuration, thermodynamic (lambda) analysis, sensitivity analysis, parameter estimation, and simulation output and visualization. Two test cases are provided to demonstrate the functionality of the Lambda-PFLOTRAN workflow. The first test case uses laboratory incubation data of temporal oxygen depletion to fit lambda parameters (i.e., maximum utilization rate and microbial carrying capacity). A slightly more complex second test case fits multiple lambda formulation and soil organic matter release parameters to temporal greenhouse gas generation measured during a soil incubation. Overall, the Lambda-PFLOTRAN workflow facilitates upscaling by using molecular-scale characterization to inform biogeochemical processes occurring at larger scales.

58 GEOSCIENCES↗

Watershed Workflow: A toolset for parameterizing data-intensive, integrated hydrologic models

Integrated, distributed hydrologic models leverage advances in computational power and data accessibility to improve predictive understanding of the water cycle. While impressive advances in this area of environmental modeling have been accomplished, such models are still rarely used, partially because of difficulty integrating model and data. This research describes the release of Watershed Workflow version 1.2, a new library aiming to automate and enable complex workflows defining inputs to high resolution, integrated, distributed hydrologic models. Watershed Workflow provides tools enabling the discovery, acquisition, mapping, and coordination of watershed geometry, land cover, soil properties, and meteorological data. It enables the construction of unstructured meshes that incorporate this data, and provides tools for automating a “first” simulation on any watershed in the United States. We present the design of the workflow tool, and describe best practices for its usage, culminating in a final example from watershed specification to simulation at the Coweeta Hydrologic Laboratory.

Integrated hydrologic modeling↗

A workflow to assess the efficacy of brine extraction for managing injection-induced seismicity potential using data from a CO 2 injection site near Decatur, Illinois

Injection of CO 2 for storage in a geologic formation increases pore pressure and alters in situ stresses. Depending on the orientation of any existing fault and fracture planes, such as critically stressed planes, this stress alteration will modify normal stresses acting on planes and could result in frictional sliding and release stored energy in the form of seismicity. Brine extraction (BE) is a technique that can be applied prior to, or during, CO 2 injection to reduce pore pressure for increasing storage capacity and, potentially, for reducing the likelihood of frictional sliding. Here a workflow is described to assess the efficacy of BE for mitigating frictional sliding (i.e., seismicity) during injection and entails: site characterization, stress calculations and failure assessment, static and dynamic modeling, and BE operational planning. Site characterization describes the stress field used to calculate the Coulomb Failure Function (CFF) that constrains allowable pore pressure changes and injection rates in the numerical simulation of CO 2 injection scenarios. The inclusion of BE in the workflow allows for determination of the potential need for pressure reduction, and evaluation of the effectiveness of this operation. Example application of the workflow using an injection field dataset near Decatur, IL, provides insight on fracture planes and stresses at the site, formation properties and the impact of variable CO 2 injection-rate targets on whether BE plans are required. The study workflow indicates that BE could enhance CO 2 injection rate by 39% and correspondingly reduce the potential for injection-induced seismicity as indicated by a reduction in CFF.

58 GEOSCIENCES↗

Wilkins: HPC in situ workflows made easy

In situ approaches can accelerate the pace of scientific discoveries by allowing scientists to perform data analysis at simulation time. Current in situ workflow systems, however, face challenges in handling the growing complexity and diverse computational requirements of scientific tasks. In this work, we present Wilkins, an in situ workflow system that is designed for ease-of-use while providing scalable and efficient execution of workflow tasks. Wilkins provides a flexible workflow description interface, employs a high-performance data transport layer based on HDF5, and supports tasks with disparate data rates by providing a flow control mechanism. Wilkins seamlessly couples scientific tasks that already use HDF5, without requiring task code modifications. We demonstrate the above features using both synthetic benchmarks and two science use cases in materials science and cosmology.

HPC↗

Reward Driven Workflows for Unsupervised Explainable Analysis of Phases and Ferroic Variants From Atomically Resolved Imaging Data

Rapid progress in aberration corrected electron microscopy necessitates development of robust methods for the identification of phases, ferroic variants, and other pertinent aspects of materials structure from imaging data. While unsupervised methods for clustering and classification are widely used for these tasks, their performance can be sensitive to hyperparameter selection in the analysis workflow. In this study, the effects of descriptors and hyperparameters are explored on the capability of unsupervised ML methods to distill local structural information, exemplified by the discovery of polarization and lattice distortion in Sm − dopped BiFeO 3 (BFO) thin films. It is demonstrated that a reward-driven approach can be used to optimize these key hyperparameters across the full workflow, where rewards are designed to reflect domain wall continuity and straightness, ensuring that the analysis aligns with the material's physical behavior. This approach allows the discovery of local descriptors that are best aligned with the specific physical behavior, providing insight into the fundamental physics of materials. The reward driven workflow is further extended to disentangle structural factors of variation via an optimized variational autoencoder (VAE). Lastly, the importance of well-defined rewards is explored as a quantifiable measure of the success of the workflow.

Barakati, Kamyar [University of Tennessee, Knoxvil↗

Asynchronous Execution of Heterogeneous Tasks in ML-Driven HPC Workflows

Heterogeneous scientific workflows consist of numerous types of tasks that require execution on heterogeneous resources. Asynchronous execution of those tasks is crucial to improve resource utilization, task throughput and reduce workflows' makespan. Therefore, middleware capable of scheduling and executing different task types across heterogeneous resources must enable asynchronous execution of tasks. In this paper, we investigate the requirements and properties of the asynchronous task execution of machine learning (ML)-driven high-performance computing (HPC) workflows. We model the degree of asynchronicity permitted for arbitrary workflows and propose key metrics that can be used to determine qualitative benefits when employing asynchronous execution. Our experiments represent relevant scientific drivers, we perform them at scale on Summit, and we show that the performance enhancements due to asynchronous execution are consistent with our model.

97 MATHEMATICS AND COMPUTING↗

Portable Acceleration of CMS Computing Workflows with Coprocessors as a Service

Computing demands for large scientific experiments, such as the CMS experiment at the CERN LHC, will increase dramatically in the next decades. To complement the future performance increases of software running on central processing units (CPUs), explorations of coprocessor usage in data processing hold great potential and interest. Coprocessors are a class of computer processors that supplement CPUs, often improving the execution of certain functions due to architectural design choices. We explore the approach of Services for Optimized Network Inference on Coprocessors (SONIC) and study the deployment of this as-a-service approach in large-scale data processing. In the studies, we take a data processing workflow of the CMS experiment and run the main workflow on CPUs, while offloading several machine learning (ML) inference tasks onto either remote or local coprocessors, specifically graphics processing units (GPUs). With experiments performed at Google Cloud, the Purdue Tier-2 computing center, and combinations of the two, we demonstrate the acceleration of these ML algorithms individually on coprocessors and the corresponding throughput improvement for the entire workflow. This approach can be easily generalized to different types of coprocessors and deployed on local CPUs without decreasing the throughput performance. We emphasize that the SONIC approach enables high coprocessor usage and enables the portability to run workflows on different types of coprocessors.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

A robust deep learning workflow to predict multiphase flow behavior during geological C O 2 sequestration injection and Post-Injection periods

Simulation of multiphase flow in porous media is essential to manage the geologic CO 2 sequestration (GCS) process, and physics-based simulation approaches usually take prohibitively high computational cost due to the nonlinearity of the coupled physics. This paper contributes to the development and evaluation of a deep learning workflow that accurately and efficiently predicts the temporal-spatial evolution of pressure and CO 2 plumes during injection and post-injection periods of GCS operations. Based on a Fourier Neural Operator, the deep learning workflow takes input variables or features including rock properties, well operational controls and time steps, and predicts the state variables of pressure and CO 2 saturation. To further improve the predictive fidelity, separate deep learning models are trained for CO 2 injection and post-injection periods due to the difference in primary driving force of fluid flow and transport during these two phases. We also explore different combinations of features to predict the state variables. We use a realistic example of CO 2 injection and storage in a 3D heterogeneous saline aquifer, and apply the deep learning workflow that is trained from physics-based simulation data and emulate the physics process. Through this numerical experiment, we demonstrate that using two separate deep learning models to distinguish post-injection from injection period generates the most accurate prediction of pressure, and a single deep learning model of the whole GCS process including the cumulative injection volume of CO 2 as a deep learning feature, leads to the most accurate prediction of CO 2 saturation. For the post-injection period, it is key to use cumulative CO 2 injection volume to inform the deep learning models about the total carbon storage when predicting either pressure or saturation. The deep learning workflow not only provides high predictive fidelity across temporal and spatial scales, but also offers a speedup of 250 times compared to full physics reservoir simulation, and thus will be a significant predictive tool for engineers to manage the long-term process of GCS.

58 GEOSCIENCES↗

Critical Assessment of MetaProteome Investigation (CAMPI): a multi-laboratory comparison of established workflows

Metaproteomics has matured into a powerful tool to assess functional interactions in microbial communities. While many metaproteomic workflows are available, the impact of method choice on results remains unclear. Here, we carry out a community-driven, multi-laboratory comparison in metaproteomics: the critical assessment of metaproteome investigation study (CAMPI). Based on well-established workflows, we evaluate the effect of sample preparation, mass spectrometry, and bioinformatic analysis using two samples: a simplified, laboratory-assembled human intestinal model and a human fecal sample. We observe that variability at the peptide level is predominantly due to sample processing workflows, with a smaller contribution of bioinformatic pipelines. These peptide-level differences largely disappear at the protein group level. While differences are observed for predicted community composition, similar functional profiles are obtained across workflows. CAMPI demonstrates the robustness of present-day metaproteomics research, serves as a template for multi-laboratory studies in metaproteomics, and provides publicly available data sets for benchmarking future developments.

59 BASIC BIOLOGICAL SCIENCES↗

A rule-free workflow for the automated generation of databases from scientific literature

Abstract In recent times, transformer networks have achieved state-of-the-art performance in a wide range of natural language processing tasks. Here we present a workflow based on the fine-tuning of BERT models for different downstream tasks, which results in the automated extraction of structured information from unstructured natural language in scientific literature. Contrary to existing methods for the automated extraction of structured compound-property relations from similar sources, our workflow does not rely on the definition of intricate grammar rules. Hence, it can be adapted to a new task without requiring extensive implementation efforts and knowledge. We test our data-extraction workflow by automatically generating a database for Curie temperatures and one for band gaps. These are then compared with manually curated datasets and with those obtained with a state-of-the-art rule-based method. Furthermore, in order to showcase the practical utility of the automatically extracted data in a material-design workflow, we employ them to construct machine-learning models to predict Curie temperatures and band gaps. In general, we find that, although more noisy, automatically extracted datasets can grow fast in volume and that such volume partially compensates for the inaccuracy in downstream tasks.

36 MATERIALS SCIENCE↗

A simulation framework for evaluating electronic order workflows in integrated health records

Electronic health record (EHR) systems are critical to modern healthcare delivery, yet the dynamic workflows that govern electronic order processing remain underexplored. Inefficiencies in these digital pathways can cause delays in care, repetitive workloads, and even patient harm. This study presents a discrete-event simulation framework used to reconstruct and evaluate EHR-based order workflows in a large integrated healthcare system. Using real-world data extracted from the Veterans Health Administration’s Corporate Data Warehouse, the authors mapped order events to standardized state transitions and modeled their progression across different facilities of varying complexity levels. After being calibrated with empirical distributions of transition times and validated against observed time-in-system metrics, the simulation demonstrates close alignment with historical performance. Scenario analyses reveal that resource capacity constraints significantly amplify the impact of electronic order surges, which are reflected in the disproportionate growth in backlogs and processing delays. Adjustments in transition probabilities further increased recirculation and extended workflow paths. Network-based analysis identified Reserved, InProgress, and Completed as structurally critical states that function as hubs within the process network but the transitions in-between also act as major bottlenecks. These results showcased the effectiveness of simulation-based approaches in monitoring EHR order processing performance and evaluating consequences of workflow changes on healthcare network resources planning. The proposed simulation framework provides a scalable data-driven tool to support operational decision-making and improve the efficiency of electronic order management in complex healthcare environments.

Engineering↗

Integration of scanning probe microscope with high-performance computing: Fixed-policy and reward-driven workflows implementation

The rapid development of computation power and machine learning algorithms has paved the way for automating scientific discovery with a scanning probe microscope (SPM). The key elements toward operationalization of the automated SPM are the interface to enable SPM control from Python codes, availability of high computing power, and development of workflows for scientific discovery. Here, we build a Python interface library that enables controlling an SPM from either a local computer or a remote high-performance computer, which satisfies the high computation power need of machine learning algorithms in autonomous workflows. We further introduce a general platform to abstract the operations of SPM in scientific discovery into fixed-policy or reward-driven workflows. Furthermore, our work provides a full infrastructure to build automated SPM workflows for both routine operations and autonomous scientific discovery with machine learning.

47 OTHER INSTRUMENTATION↗

A reliable workflow for improving nanoscale X-ray fluorescence tomographic analysis on nanoparticle-treated HeLa cells

Abstract Scanning X-ray fluorescence (XRF) tomography provides powerful characterization capabilities in evaluating elemental distribution and differentiating their inter- and intra-cellular interactions in a three-dimensional (3D) space. Scanning XRF tomography encounters practical challenges from the sample itself, where the range of rotation angles is limited by geometric constraints, involving sample substrates or nearby features either blocking or converging into the field of view. This study aims to develop a reliable and efficient workflow that can (1) expand the experimental window for nanoscale tomographic analysis of local areas of interest within a laterally extended specimen, and (2) bridge 3D analysis at micrometer and nanoscales on the same specimen. We demonstrate the workflow using a specimen of HeLa cells exposed to iron oxide core and titanium dioxide shell (Fe3O4/TiO2) nanocomposites. The workflow utilizes iterative and multiscale XRF data collection with intermediate sample processing by focused ion beam (FIB) sample preparation between measurements at different length scales. Initial assessment combined with precise sample manipulation via FIB allows direct removal of sample regions that are obstacles to both incident X-ray beam and outgoing XRF signals, which considerably improves the subsequent nanoscale tomography analysis. This multiscale analysis workflow has advanced bio-nanotechnology studies by providing deep insights into the interaction between nanocomposites and single cells at a subcellular level as well as statistical assessments from measuring a population of cells.

62 RADIOLOGY AND NUCLEAR MEDICINE↗

RISE: Reducing I/O Contention in Staging-based Extreme-Scale In-situ Workflows

While in-situ workflow formulations have addressed some of the data-related challenges associated with extreme-scale scientific workflows, these workflows involve complex interactions and different modes of data exchange. In the context of increasing system complexity, such workflows present significant resource management challenges, requiring complex cost-performance tradeoffs. This paper presents RISE, an intelligent staging-based data management middleware, which builds on the DataSpaces framework and performs intelligent scheduling of data management operations to reduce I/O contention. In RISE, data are always written immediately to local buffers to reduce the effect of the transfer impact upon application performance. RISE identifies applications’ data access patterns and moves data towards data consumers only when the network is expected to be idle, reducing the impact of asynchronous background data movement upon critical data read/write requests. Here, we experimentally demonstrate that RISE can take advantage of staging nodes to offload data during writes without degrading application data movement performance.

97 MATHEMATICS AND COMPUTING↗

ProvLight: Efficient Workflow Provenance Capture on the Edge-to-Cloud Continuum

Modern scientific workflows require hybrid infrastructures combining numerous decentralized resources on the IoT/Edge interconnected to Cloud/HPC systems (aka the Computing Continuum) to enable their optimized execution. Understanding and optimizing the performance of such complex Edge-to-Cloud workflows is challenging. Capturing the provenance of key performance indicators, with their related data and processes, may assist in understanding and optimizing workflow executions. However, the capture overhead can be prohibitive, particularly in resource-constrained devices, such as the ones on the IoT/Edge.To address this challenge, based on a performance analysis of existing systems, we propose ProvLight, a tool to enable efficient provenance capture on the IoT/Edge. We leverage simplified data models, data compression and grouping, and lightweight transmission protocols to reduce overheads. We further integrate ProvLight into the E2Clab framework to enable workflow provenance capture across the Edge-to-Cloud Continuum. This integration makes E2Clab a promising platform for the performance optimization of applications through reproducible experiments.We validate ProvLight at a large scale with synthetic workloads on 64 real-life IoT/Edge devices in the FIT IoT LAB testbed. Evaluations show that ProvLight outperforms state-of-the-art systems like ProvLake and DfAnalyzer in resource-constrained devices. ProvLight is 26—37x faster to capture and transmit provenance data; uses 5—7x less CPU; 2x less memory; transmits 2x less data; and consumes 2—2.5x less energy. ProvLight [1] and E2Clab [2] are available as open-source tools.

Rosendo, Daniel↗