Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Scientific Workflows”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

NNSA/CEA Workflow Workshop Report 2021

Researchers from three NNSA labs plus CEA recently met for a half-day workshop on scientific and engineering workflows in March of 2021. Tools, projects and use cases from each institution were described. A wide range of unique capabilities and requirements were represented. The workshop highlights the fact that the CEA/NNSA workflows space mirrors the broader open science workflows community, with multiple technologies under development in a number of science domains and under a number of funding streams, with differing capabilities and focuses. Despite the number of tools, presentations and discussions have shown that these tools cover specific mission spaces at the four labs, each with distinctive capabilities that do not completely overlap with each other. We believe there is a strong interest in the short term for sharing lessons learned and collaborating on benchmarks and site evaluation of projects. This report, prepared by the NNSA/CEA Workflows Working Group, briefly summarizes the presentations in the areas of domain specific workflows, end user environments, data management, job and resource management, and infrastructure, and then identifies six broad areas for potential collaboration. A key finding is that users could benefit from greater interoperability, compatibility, and composability of the workflow technologies under development and that point-to-point collaboration opportunities should be identified to explore these aspects.

97 MATHEMATICS AND COMPUTING↗

Universal Workflow Language and Software Enable Geometric Learning and FAIR Scientific Protocol Reporting

Written language and conventional data structures for representing scientific procedures suffer from low process detail, often fail to accurately represent protocols, and lack universality. New strategies for the handling of experimental data are needed to provide viable process information for both humans and machines. In this work, we present the universal workflow language (UWL) and interface (UWLi). UWL is a findable, accessible, interoperable, and reusable (FAIR)-compatible, graph-based data architecture that can capture arbitrary scientific procedures through workflow representation, and UWLi is an accompanying software package for building, manipulating, and interpreting UWL entries. The UWL format was found to be highly effective in identifying deficiencies in the reported process details of high-impact, peer-reviewed scientific journals, and in simulated scenarios, the graph format was shown to be more effective than conventional methods in predictively modeling the outcome of diverse scientific protocols. Implementation of UWL could enable more accurate scientific communication and more impactful process datasets.

14 SOLAR ENERGY↗

A rule-free workflow for the automated generation of databases from scientific literature

Abstract In recent times, transformer networks have achieved state-of-the-art performance in a wide range of natural language processing tasks. Here we present a workflow based on the fine-tuning of BERT models for different downstream tasks, which results in the automated extraction of structured information from unstructured natural language in scientific literature. Contrary to existing methods for the automated extraction of structured compound-property relations from similar sources, our workflow does not rely on the definition of intricate grammar rules. Hence, it can be adapted to a new task without requiring extensive implementation efforts and knowledge. We test our data-extraction workflow by automatically generating a database for Curie temperatures and one for band gaps. These are then compared with manually curated datasets and with those obtained with a state-of-the-art rule-based method. Furthermore, in order to showcase the practical utility of the automatically extracted data in a material-design workflow, we employ them to construct machine-learning models to predict Curie temperatures and band gaps. In general, we find that, although more noisy, automatically extracted datasets can grow fast in volume and that such volume partially compensates for the inaccuracy in downstream tasks.

36 MATERIALS SCIENCE↗

Performance Understanding and Analysis for Exascale Data Management Workflows (Collaboration)

The general approach of the MONA project is depicted in Figure 1. The figure shows performance monitoring applied to an I/O workflow associated with a scientific simulation, with online measurements captured from this workflow informing methods for workflow reconfiguration and dynamic adjustment. The goal is to maintain suitable levels of Quality of Service for workflow execution, by understanding the underlying causes of workflow performance and behavior. One outcome will be workflow performance models able to characterize realistic workflows. Another outcome will be performance ‘mini-apps’ implementing such workflow behavior. Out of scope for this project are advanced methods for online workflow control.

97 MATHEMATICS AND COMPUTING↗

Building the I (Interoperability) of FAIR for performance reproducibility of large-scale composable workflows in RECUP

Abstract-Scientific computing communities increasingly run their experiments using complex data- and compute-intensive workflows that utilize distributed and heterogeneous architectures targeting numerical simulations and machine learning, often executed on the Department of Energy Leadership Computing Facilities (LCFs). We argue that a principled, systematic approach to implementing FAIR principles at scale, including fine-grained metadata extraction and organization, can help with the numerous challenges to performance reproducibility posed by such workflows. We extract workflow patterns, propose a set of tools to manage the entire life cycle of performance metadata, and aggregate them in an HPC-ready framework for reproducibility (RECUP). We describe the challenges in making these tools interoperable, preliminary work, and lessons learned from this experiment.

97 MATHEMATICS AND COMPUTING↗

Enabling Low-Overhead HT-HPC Workflows at Extreme Scale using GNU Parallel

GNU Parallel is a versatile and powerful tool for process parallelization widely used in scientific computing. This paper demonstrates its effective application in high-performance computing (HPC) environments, particularly focusing on its scalability and efficiency in executing large-scale high-throughput high-performance computing (HT-HPC) workflows. Through real-world examples, we highlight GNU Parallel’s performance across various HPC workloads, including GPU computing, container-based workloads, and node-local NVMe storage. Our results on two leading supercomputers, OLCF’s Frontier and NERSC’s Perlmutter, showcase GNU Parallel’s rapid process dispatching ability and its capacity to maintain low overhead even at extreme scales. We explore GNU Parallel’s application in massive parallel file transfers using a scheduled Data Transfer Node (DTN) cluster, emphasizing its broad utility in diverse scientific workflows. Beyond its direct application as a viable workflow manager, GNU Parallel can be employed in conjunction with other workflow systems as a "last-mile" parallelizing driver and as a quick prototyping tool to design and extract parallel profiles from application executions. We then argue that the potential for GNU Parallel to transform workflow management at extreme scales is substantial, paving the way for more efficient and effective scientific discoveries.

Maheshwari, Ketan↗

Advanced Research Directions on AI for Science, Energy, and Security: Report on Summer 2022 Workshops

Over the past decade, fundamental changes in artificial intelligence (AI)—from foundational to applied—have delivered dramatic insights across a wide breadth of U.S. Department of Energy (DOE) mission space. AI is helping to augment and improve scientific and engineering workflows (e.g., for control, design, and dramatic performance gains through surrogate models) in national security, the Office of Science, and DOE’s applied energy programs. The progress and potential for AI in DOE science was captured in the 2020 “AI for Science” report from the DOE laboratory community in collaboration with academia and industry. Specific scientific areas ready to further leverage the power of AI ranged from the scale and performance of computational models to data analysis to creating new classes of observations using computer vision. Since that report, the scale and scope of scientific AI have accelerated, revealing new, emergent properties that yield insights that go beyond enabling opportunities to being potentially transformative in the way that scientific problems are posed and solved. Thus, under the guidance of both the Office of Science (SC) and the National Nuclear Security Administration (NNSA), the DOE national laboratories organized a series of workshops in 2022 to gather input on new and rapidly emerging opportunities and challenges of scientific AI. This 2023 report is a synthesis of those workshops. The scientific community believes AI can have a foundational impact on a broad range of DOE missions, including science, energy, and national security. Further, DOE has unique capabilities that enable the community to drive progress in scientific use of AI, building on long-standing DOE strengths and investments in computation, data, and communications infrastructure, spanning the Energy Sciences Network (ESnet), the Exascale Computing Project (ECP), and integrative programs such as the NNSA Office of Defense Programs Advanced Simulation and Computing (ASC) and the SC Scientific Discovery through Advanced Computing (SciDAC) programs.

97 MATHEMATICS AND COMPUTING↗

Support for the Core Research Activities and Studies of the Computer Science and Telecommunications Board (DE-SC0020446 Final Technical Report)

Supported the core operations of the National Academies' Computer Science and Telecommunications Board (CSTB). Helped support planning and conducting of board meetings, identification of priority topics in computer science and other areas of computing and communications technologies, and oversight for CSTB's portfolio of studies and convenings. Activities shaped and overseen included: a workshop on Al for scientific discovery; collaborative work with other Academies units on a study on foundational research gaps and future directions for digital twins; a study on current capabilities, future prospects, and governance of facial recognition technologies, a study on post- exascale computing; a study on fostering responsible computing research, collaborative work with other Academies units on automated research workflows for accelerated scientific discovery; a study on meeting federal cybersecurity workforce needs, and a study of the ecosystem driving information technology innovation.

97 MATHEMATICS AND COMPUTING↗

Leveraging interpolation models and error bounds for verifiable scientific machine learning

Effective verification and validation techniques for modern scientific machine learning workflows are challenging to devise. Statistical methods are abundant and easily deployed, but often rely on speculative assumptions about the data and methods involved. Error bounds for classical interpolation techniques can provide mathematically rigorous estimates of accuracy, but often are difficult or impractical to determine computationally. Here, in this work, we present a best-of-both-worlds approach to verifiable scientific machine learning by demonstrating that (1) multiple standard interpolation techniques have informative error bounds that can be computed or estimated efficiently; (2) comparative performance among distinct interpolants can aid in validation goals; (3) deploying interpolation methods on latent spaces generated by deep learning techniques enables some interpretability for black-box models. We present a detailed case study of our approach for predicting lift-drag ratios from airfoil images. Code developed for this work is available in a public Github repository.

97 MATHEMATICS AND COMPUTING↗

Enabling discovery data science through cross-facility workflows

Experimental and observational instruments for scientific research (such as light sources, genome sequencers, accelerators, telescopes and electron microscopes) increasingly require High Performance Computing (HPC) scale capabilities for data analysis and workflow processing. Next-generation instruments are being deployed with higher resolutions and faster data capture rates, creating a big data crunch that cannot be handled by modest institutional computing resources. Often these big data analysis pipelines also require near real-time computing and have higher resilience requirements than the simulation and modeling workloads more traditionally seen at HPC centers. While some facilities have enabled workflows to run at a single HPC facility, there is a growing need to integrate capabilities across HPC facilities to enable cross-facility workflows, either to provide resilience to an experiment, increase analysis throughput capabilities, or to better match a workflow to a particular architecture. In this paper we describe the barriers to executing complex data analysis workflows across HPC facilities and propose an architectural design pattern for enabling scientific discovery using cross-facility workflows that includes orchestration services, application programming interfaces (APIs), data access and co-scheduling.

Antypas, Katerina B.↗

Sampling in Long-Screened Wells: Issues, Misconceptions, and Solutions

The issues associated with long-screened wells (LSWs) (and open boreholes) at contaminated sites are well documented in the groundwater literature but are still not fully appreciated in practice. As established in seminal and review papers going back over three decades, the interpretation of sampling results from LSWs is challenging in the presence of vertical hydraulic gradients and borehole flow; furthermore, LSWs allow for vertical redistribution of contamination between aquifer layers. Acknowledgment of these issues has led to the development of new technologies and well designs to enable discrete-zone monitoring (DZM), yet LSWs remain common for many reasons, for example, as multipurpose wells, for geophysical logging, and (or) as legacy installations. Despite the literature on LSWs and despite the adoption of DZM at many sites, the use of LSWs persists and the challenges of interpreting sampling results from LSWs remain. In this issue paper, we provide a conceptual overview of the problems posed by LSWs and review existing literature and past work to improve the interpretation of sampling in LSWs. We draw on experience from previous studies at the Hanford Site in eastern WA, USA, and use synthetic examples to illustrate key concepts and challenges for interpretation. A recently published analytical modeling framework is used to develop illustrative synthetic examples and demonstrate a workflow for building scientific intuition to understand issues around interpreting samples from LSWs, which is critical to effective characterization and groundwater remediation at sites with LSWs.

54 ENVIRONMENTAL SCIENCES↗

Streaming Data from Experimental Facilities to Supercomputers for Real-Time Data Processing

In this paper we demonstrate direct data streaming from instruments and detectors at a large-scale experimental facility to a supercomputer for real-time data processing and feedback. Streaming data to supercomputers introduces the potential for novel scientific applications and workflow models, including the ability to provide real-time feedback from very large datasets during an experiment and the integration of real-time ML training and inference at scale. We discuss a successful demonstration for real-time processing of data from the Advanced Photon Source (APS) on the Polaris supercomputer using an EPICS-based streaming framework. We describe the capabilities of the streaming framework itself, and outline the architecture that allows us to process experimentally derived data on a supercomputer without file-based data transfers. We present throughput measurements that are indicative of system performance capable of sustaining the expected data production rates of the facility, as well as discuss some outstanding challenges and our future directions.

real-time processing↗

CodeScribe Agent

SF-26-086 CodeScribe introduces a structured, multi-stage pipeline that combines deterministic program analysis with LLM-powered translation to enable incremental, testable Fortran-to-C++ migration. First, `code-scribe index` traverses the project directory tree and produces `scribe.yaml` metadata files recording all modules, subroutines, and functions at each level, giving the LLM accurate structural context instead of a hallucinated codebase model. Second, `code-scribe draft` performs the deterministic portion of translation — converting Fortran types to C++ equivalents, replacing `use` statements with `#include` and `using namespace` directives, and detecting constructs requiring special handling — while embedding`scribe-prompt` annotations that guide the LLM through non-trivial cases such as statement-function-to-lambda conversions and `extern "C"` wrapper generation. Third, `code-scribe translate` applies project-specific TOML-based few-shot prompt templates and submits the composed prompt to a pluggable LLM backend (OpenAI, Anthropic, Argonne ARGO, any OpenAI-compatible endpoint, or local Hugging Face checkpoints), producing a C++ source file, a header, and a Fortran-C++ interface file for each translated routine so the codebase compiles and runs correctly throughout the migration. Beyond translation, CodeScribe includes a tool-using coding agent (`code-scribe agent`) with read, bash, edit, and write capabilities, and a bounded loop mode (`code-scribe loop`) that runs repeated stateless agent sessions over a task file with restricted tool access — enabling sustained, auditable software development workflows for broader scientific computing tasks.

Dhruv, Akash [Argonne National Laboratory (ANL), A↗

STNS01-21 BEE - FY21 P6-2: Archive, clone, and re-run workflows [Slide]

BEE will give ECP a tool that great simplifies the deployment of containerized workflows on the next generation of pre-exascale and exascale systems, as well as public and private clouds. BEE allows scientists to describe their workflow using the Common Workflow Language and then deploy that workflow across the entire spectrum of systems without having to learn the specifics of each container runtime, HPC resource manager, or cloud API. BEE also streamlines the curation and sharing of common workflows among the scientific community.

97 MATHEMATICS AND COMPUTING↗

2019 Computing Sciences Strategic Plan

Computing has transformed nearly every aspect of scientific inquiry — across disciplines and across scales — from the behavior of subatomic particles to the formation of structures in the early universe, from the assembly of the human genome to the evolution of earth systems. Over the past two decades, computing has become an integral part of how Berkeley Lab is “Bringing Science Solutions to the World.” Advances in computing and mathematics have been key, with new mathematical models of complex physical phenomena, new methods for analyzing complex data, new algorithms for accuracy and scaling and sophisticated software systems that encapsulate these techniques into open, reusable tools. The performance of NERSC computers and the ESnet network have grown by several orders of magnitude, along with our understanding of how to map scientific computations and workflows onto these systems. From research to facility operations, the passion, talent and dedication of the Computing Sciences Area staff has been the cornerstone of our success. The plan outlined in this document describes the next step in a journey to expand the influence and impact of our efforts, building an increasingly connected global enterprise for science that places more powerful instruments in the hands of scientists, along with more powerful methods and tools for modeling, analysis and prediction.

97 MATHEMATICS AND COMPUTING↗

Data Cards for Standardized Metadata Across DOE-Aligned Data Initiatives: Toward Transparent, Interoperable, and Governed Dataset Documentation

As data-intensive research, advanced computing, and artificial intelligence become increasingly central to scientific and operational workflows, the need for consistent, transparent, and machine-actionable documentation has grown correspondingly. Multiple DOE-aligned communities—including Office of Science, Genesis Mission, American Science Cloud (AmSC), National Nuclear Security Administration (NNSA) stewardship and governance, and related cross-laboratory collaborations—have independently developed metadata practices to support discovery, access, reuse, repository deposit, and compliance.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

ESnet-JLab FPGA Accelerated Transport (data plane) [EJFAT (udplb)] v1.0

The ESnet-JLab FPGA Accelerated Transport system is a solution for streaming high-speed scientific measurement data from Data Acquisition Systems (DAQs) to high-performance computing facilties. It is generally compatible with many science workflows, and makes no assumptions about the specifics of any particular experiment. This program (udplb) implements the data plane portion of the EJFAT system. It is an FPGA design that rewrites and forwards data packets from a UDP-based scientific workflow to high-performance compute nodes. It depends on another program (udplbd, disclosed separately) to implement the control system.

Bengough, Peter [Malleable Networks, Inc.]↗

Performance assessment of ensembles of in situ workflows under resource constraints

Summary Scientific breakthroughs in biomolecular methods and improvements in hardware technology have shifted from a long‐running simulation to a large set of shorter simulations running simultaneously, called an ensemble. In an ensemble, simulations are usually coupled with analyses of data produced by the simulations. In situ methods can be used to analyze large volumes of data generated by scientific simulations at runtime (i.e., simulations and analyses are performed concurrently). In this work, we study the execution of ensemble‐based simulations paired with in situ analyses using in‐memory staging methods. Using an ensemble of molecular dynamics in situ workflows with multiple simulations and analyses, we first show that collecting traditional metrics such as makespan, instructions per cycle, memory usage, or cache miss ratio is not sufficient to characterize complex behaviors of ensembles. We propose a method to evaluate the performance of ensembles of workflows that captures multiple resource usage aspects: resource efficiency, resource allocation, and resource provisioning. Experimental results demonstrate that the proposed method can effectively distinguish the performance of different component placements in an ensemble with up to 32 ensemble members. By evaluating different co‐location scenarios, our proposed performance indicators demonstrate benefits of co‐locating simulation and coupled analyses within a compute node.

Do, Tu Mai Anh↗