Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Scientific workflow”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Report for the ASCR Workshop on the Management and Storage of Scientific Data

The purpose of this workshop is to identify priority research directions in the area of data management for high-performance and scientific computing above and beyond HPC’s traditional "the parallel file system is the data-management system" model. Supporting the breadth of the DOE mission, including the explosion of AI uses and the growing needs of experimental and observational science, motivates revisiting our assumptions about data management. There are many facets of this topic to explore including: (1) Interfaces for accessing data that resides on traditional persistent storage as well as memory devices; (2) Storage-system architecture design that supports scientific workflows on varied hierarchical storage and networking devices; (3) Devising metadata management infrastructure to support FAIR principles (Findability, Accessibility, Interoperability, and Reusability); (4) Capturing provenance information about scientific data; (5) Utilizing AI to learn I/O patterns of emerging workloads for efficient data management; (6) Providing data management support for AI and complex workflows; and (7) Understanding the overlap between traditional storage systems and I/O (SSIO) efforts and data management. While the program committee has identified these topics as important areas for discussion, we welcome position papers from the community that propose additional topics of interest for discussion at the workshop. The workshop agenda will include breakout sessions for discussing these and selected topic areas to inform priority research directions for data management for high-performance and scientific computing.

97 MATHEMATICS AND COMPUTING↗

Toward digital design at the exascale: An overview of project ICECap

High performance computing has entered the Exascale Age. Capable of performing over 1018 floating point operations per second, exascale computers, such as El Capitan, the National Nuclear Security Administration's first, have the potential to revolutionize the detailed in-depth study of highly complex science and engineering systems. However, in addition to these kind of whole machine “hero” simulations, exascale systems could also enable new paradigms in digital design by making petascale hero runs routine. Currently, untenable problems in complex system design, optimization, model exploration, and scientific discovery could all become possible. Motivated by the challenge of uncovering the next generation of robust high-yield inertial confinement fusion (ICF) designs, project ICECap (Inertial Confinement on El Capitan) attempts to integrate multiple advances in machine learning (ML), scientific workflows, high performance computing, GPU-acceleration, and numerical optimization to prototype such a future. Built on a general framework, ICECap is exploring how these technologies could broadly accelerate scientific discovery on El Capitan. In addition to our requirements, system-level design, and challenges, we describe some of the key technologies in ICECap, including ML replacements for multiphysics packages, tools for human-machine teaming, and algorithms for multifidelity design optimization under uncertainty. As a test of our prototype pre-El Capitan system, we advance the state-of-the art for ICF hohlraum design by demonstrating the optimization of a 17-parameter National Ignition Facility experiment and show that our ML-assisted workflow makes design choices that are consistent with physics intuition, but in an automated, efficient, and mathematically rigorous fashion.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Data Management in the Continuum: Cross-facility Object-based Data Transfers

Scientific workflows are evolving from relying on a monolithic storage subsystem at a single High-Performance Computing (HPC) facility to using geographically distributed file systems, repositories, and cloud storage. As a result, storing, accessing, transferring, and managing scientific data have become highly complex and prone to performance inefficiencies. This paper delves into these challenges by exploring an optimized end-to-end interface designed to seamlessly connect various local and remote storage systems, enabling efficient data movement of objects across HPC–Cloud and HPC–HPC environments. We showcase this capability through an object-focused data management runtime system, discuss the effects of relaxed consistency semantics in distributed object scenarios, and illustrate its application in an earthquake simulation workflow. Besides reducing the amount of data by selectively transferring regions of interest, our facility-local results achieved a speedup of 45 × over an optimized HDF5 usage and 15 × over the HDF5 with caching by using the new interface in PDC-XF.

Bez, Jean Luca↗

Novel Proposals for FAIR, Automated, Recommendable, and Robust Workflows

Lightning talks of the Workflows in Support of Large-Scale Science (WORKS) workshop are a venue where the workflow community (researchers, developers, and users) can discuss work in progress, emerging technologies and frameworks, and training and education materials. This paper summarizes the WORKS 2022 lightning talks, which cover five broad topics: data integrity of scientific workflows; a machine learning-based recommendation system; a Python toolkit for running dynamic ensembles of simulations; a cross-platform, high-performance computing utility for processing shell commands; and a meta(data) framework for reproducing hybrid workflows.

Abhinit, Ishan↗

Predictive Modeling and Diagnostic Monitoring of Extreme Science Workflows (Final Report)

This proposal addresses a critical issue of performance prediction identified in the report from the ASCR \Computational Modeling of Big Networks (COMBINE)" workshop: "end-to-end performance is not predictable due to a variety of factors. Even when some performance forecasts or predictions can be made, they often cannot explain the reasons why some predictions fail." We will develop new analytical models to predict the end-to-end performance of scientific workflows on DOE computing infrastructures, and use simulations and experimentation to validate and refine these models, as well as to pinpoint the sources of model inaccuracy. We will also use these models to help diagnose application and infrastructure problems, and to adapt the system based on this diagnosis. This section provides background in the areas relevant to the proposed work. RPI’s specific tasks within the Panorama project are as follows: (1) Develop Aspen-Simulation interface for Workflow Model Driven Simulation. (2) Validate manual performance models of two target workflow scenarios with empirical measurement and simulation. (3) Extend ROSS-Aspen API to simulate workflow descriptions when required. (4) Validate Aspen performance models of two target workflow scenarios with automatic performance model empirical measurement and simulation. (5) Design and implement final system to automatically generate Aspen performance models from workflow descriptions (including methods to compensate for limitations of Aspen analytical models). (6) Validate improved Aspen performance models with target workflow on production infrastructure. To date, all the project milestones assigned to us where reached within the best of our abilities over the course of the project performance period. Below describes the key outcome from our collaborative research in a system named, Durango .

97 MATHEMATICS AND COMPUTING↗

Position Papers for the ASCR Workshop on Cybersecurity and Privacy for Scientific Computing Ecosystems

At the request of the Department of Energy's (DOE) Office of Advanced Scientific Computing Research (ASCR), this program committee has been tasked with organizing a workshop to identify basic research needs in cybersecurity and privacy to better support DOE's science and energy mission. As part of the process, the program committee is soliciting community input in the form of position papers to help identify significant use cases, facility issues, and other barriers to enabling verifiably trustworthy computational science while preserving data confidentiality as appropriate for scientific workflows of interest to DOE. The program committee will review these position papers and based on the fit of their area of expertise and interest, selected contributors will have the opportunity to participate in the workshop currently planned as a virtual event November 3-5th, 2021. The thrust areas that will be explored by this workshop are the following: (1) Algorithms for secure, scalable, privacy-enhancing technologies and frameworks, including: Federated AI/ML, Differential privacy, Randomized algorithms, Adversarial modeling & simulation, Graph algorithms, and Formal methods; (2) Platforms to support the entire scientific-computing ecosystem, including edge computing for large-scale experiments, focusing on heterogeneous systems and distributed systems, including: Heterogeneous computing systems, Distributed computing systems, and Secure data architectures; and (3) Data workflows to allow agile use of data while preserving integrity and privacy, making the important properties verifiable either at runtime or post-computation, including: Integrity and provenance and Data management infrastructure. Topics that are out-of-scope for the workshop include discussing specific proposed solutions or areas that are clearly out of DOE's fundamental and applied-sciences mission scope, e.g., cryptography, enterprise security, and general-operations technology.

97 MATHEMATICS AND COMPUTING↗

DOE Repository Metadata Profile (DRMP): A Metadata Framework for Advancing Interoperability and AI Readiness Across Scientific Repositories

The Department of Energy (DOE) funds a diverse and distributed ecosystem of repositories that steward scientific data, publications, and software across its research programs, user facilities, and national laboratories. While significant progress has been made in standardizing dataset-level metadata, the metadata describing repositories themselves (their identity, governance, access interfaces, policies, and technical capabilities) remains inconsistent and fragmented across DOE-funded systems. This variability limits discoverability, interoperability, automated validation, and AI-driven analysis, all of which are increasingly essential for modern scientific workflows. To address this gap, the DOE Data Curation Working Group (DCWG) developed the DOE Repository Metadata Profile (DRMP). The DRMP is a practical, community-driven framework that defines how repositories can describe themselves in a consistent, machine-actionable, and scalable manner. The DRMP is not a new metadata schema. Instead, it is a mapping profile and structured element set capturing the essential characteristics of DOE repositories. It harmonizes repository-level metadata across six widely adopted community schemas: RE3Data; DCAT-US v3; Schema.org; Dublin Core; DataCite 4.6; and PREMIS 3.0. This harmonization eliminates reinvention and enables interoperability within DOE and across the broader scientific ecosystem. A core objective of the DRMP is to reduce burden on repositories by allowing them to reuse their existing metadata through a Rosetta-style crosswalk rather than redesigning local implementations. The profile introduces a three-level conformance model that supports incremental adoption: • Level 1 – Minimum Viable Record (MVR): foundational identification elements required for workflows, project registration, and basic repository presence. • Level 2 – Interoperable: structured metadata enabling alignment with national and international discovery systems. • Level 3 – AI-Ready: enhanced provenance, policy transparency, fixity, semantic context, and capabilities that support automated reasoning, model training governance, and machine-assisted curation. To support implementation, the DRMP includes JSON Schema definitions, OpenAPI patterns, and MCP templates that allow repositories to publish machine-readable metadata directly within existing platforms. These resources are modular and lightweight, enabling adoption without major architectural change. Adopting the DRMP enables repositories to: • Enhance discoverability and interoperability by aligning identifiers, classifications, and descriptive elements across widely used schema standards. • Support federated discovery and cross-registration across DOE systems, Data.gov, and international catalogs. • Enable AI agents and workflow orchestration systems to interpret repository-level metadata within the American Science Cloud (AmSC) through Model Context Protocol (MCP)-based context publication. • Demonstrate alignment with DOE’s open science, stewardship, and FAIR data priorities. This guidance represents a community-driven step forward. Through voluntary adoption and continued feedback, the DRMP advances a cohesive, machine-actionable description of DOE repositories that supports FAIR data practices, preparing the infrastructure for AI-enabled research, and strengthening the discoverability and reuse of DOE’s scientific outputs.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

An Efficient Storage-Driven Machine Learning Model for Performance in the Era of Multimodal Scientific Data

Scientific workflows are increasingly relying on machine learning (ML), simulation, and hybrid techniques to predict, understand, and optimize the behavior of complex experiments. High-performance computing has greatly improved researchers’ ability to acquire diverse data modalities in these workflows. Recent studies suggest that the performance of machine learning models can be improved by integrating data from various sources. Unfortunately, these workloads pose unprecedent pressure on the network storage to meet the demands associated with accessing these multimodal data. To mitigate the impact of intensive IO, we propose a solution that utilizes a multi-tier High-Performance Computing (HPC) distributed storage and data processing framework, placing computation where the data resides for better performance. By adopting this project, the scientific community will gain new opportunities to explore multimodal storage-driven possibilities, integrating multiple scientific data sources with advanced streaming frameworks. Additionally, our framework effectively utilizes computing resources and bridges the gaps identified by HPC experts. Our proposed approach tackles scalability and persistence challenges by leveraging native persistency, which has posed difficulties in traditional approaches. Furthermore, we seek to enhance fault-tolerance and load-balance of computations by leveraging real-time streaming in diverse scientific computing environments, thereby propelling advanced scientific computing research into the next generation.

97 MATHEMATICS AND COMPUTING↗

Improving climate model coupling through a complete mesh representation: a case study with E3SM (v1) and MOAB (v5.x)

One of the fundamental factors contributing to the spatiotemporal inaccuracy in climate modeling is the mapping of solution field data between different discretizations and numerical grids used in the coupled component models. The typical climate computational workflow involves evaluation and serialization of the remapping weights during the preprocessing step, which is then consumed by the coupled driver infrastructure during simulation to compute field projections. Tools like Earth System Modeling Framework (ESMF) and TempestRemap offer capability to generate conservative remapping weights, while the Model Coupling Toolkit (MCT) that is utilized in many production climate models exposes functionality to make use of the operators to solve the coupled problem. However, such multistep processes present several hurdles in terms of the scientific workflow and impede research productivity. In order to overcome these limitations, we present a fully integrated infrastructure based on the Mesh Oriented datABase (MOAB) library, which allows for a complete description of the numerical grids and solution data used in each submodel. Through a scalable advancing-front intersection algorithm, the supermesh of the source and target grids are computed, which is then used to assemble the high-order, conservative, and monotonicity-preserving remapping weights between discretization specifications. The Fortran-compatible interfaces in MOAB are utilized to directly link the submodels in the Energy Exascale Earth System Model (E3SM) to enable online remapping strategies in order to simplify the coupled workflow process. We demonstrate the superior computational efficiency of the remapping algorithms in comparison with other state-of-the-science tools and present strong scaling results on large-scale machines for computing remapping weights between the spectral element atmosphere and finite volume discretizations on the polygonal ocean grids.

58 GEOSCIENCES↗

Ecosystems for Scientific Computing in the Age of AI

Scientific computing is at an inflection point. Artificial intelligence (AI) is reshaping how scientific software is developed, how teams collaborate, how projects are governed, and how the next generation is trained. Drawing on insights from a 2025 workshop report, this article argues that the future of discovery will depend on agile, robust ecosystems built through socio-technical co-design—the intentional integration of technical and human systems. This perspective is essential for ensuring that future scientific computing remains trustworthy, sustainable, and scalable. It combines advances in AI, high-performance computing, and software with new models for cross-disciplinary collaboration, education, and workforce development. Key recommendations include building modular, trustworthy AI-enabled software ecosystems; enabling teams to integrate AI into scientific workflows while preserving human creativity, integrity, and rigor; and developing adaptive training pathways that keep pace with rapid technological change. By sharing these perspectives, we hope to stimulate broader community dialogue and encourage coordinated action.

AI↗

A KBase Case Study on Genome-wide Transcriptomics and Plant Primary Metabolism in Sorghum

A better understanding of the genetic and metabolic mechanisms that confer stress resistance and tolerance in plants is key to engineering new crops through advanced breeding technologies. This requires a systems biology approach that builds on a genome-wide understanding of the regulation of gene expression, plant metabolism, physiology and growth. In this study, we examine the response to drought stress in Sorghum, as we leverage the tools for transcriptomics and plant metabolic modeling we have implemented at the U.S. Department of Energy Systems Biology Knowledgebase (KBase). KBase enables researchers worldwide to collaborate and advance research by allowing them to upload private or public data into the KBase Narrative Interface, empowering them to analyze this data using a rich, extensible array of computational and data-analytics tools, and allows them to securely share scientific workflows and conclusions. We demonstrate how to use the current RNA-seq tools in KBase, applicable to both plants and microbes, to assemble and quantify long transcripts and identify differentially expressed genes effectively. More specifically, we demonstrate the utility of the platform by identifying key genes that are differentially expressed during drought-stress in Sorghum bicolor, which is an important sustainable production crop plant. We then show how to use KBase tools to predict the membership of genes in metabolic pathways and examine expression data in the context of metabolic subsystems. We demonstrate the power of the platform by making the data analysis and interpretation available to the biologists in the reproducible, re-usable, point-and-click format of a KBase Narrative thus promoting FAIR (Findable, Accessible, Interoperable and Reusable) guiding principles for scientific data management and stewardship.

59 BASIC BIOLOGICAL SCIENCES↗

Investigating User Experiences with Data Abstractions on High Performance Computing Systems

Scientific exploration generates expanding volumes of data that commonly require High Performance Computing (HPC) systems to facilitate research. HPC systems are complex ecosystems of hardware and software that frequently are not user friendly. The Usable Data Abstractions (UDA) project set out to build usable software for scientific workflows in HPC environments by undertaking multiple rounds of qualitative user research. Qualitative research investigates how individuals accomplish their work and our interview-based study surfaced a variety of insights about the experiences of working in and with HPC ecosystems. This report examines multiple facets to the experiences of scientists and developers using and supporting HPC systems. We discuss how stakeholders grasp the design and configuration of these systems, the impacts of abstraction layers on their ability to successfully do work, and the varied perceptions of time that shape this work. Examining the adoption of the Cori HPC at NERSC we explore the anticipations and lived experiences of users interacting with this system’s novel storage feature, the Burst Buffer. We present lessons learned from across these insights to illustrate just some of the challenges HPC facilities and their stakeholders need to account for when procuring and supporting these essential scientific resources to ensure their usability and utility to a variety of scientific practices.

97 MATHEMATICS AND COMPUTING↗

Smart connected worker edge platform for smart manufacturing: Part 2—Implementation and on‐site deployment case study

Abstract In this paper, we describe specific deployments of the Smart Connected Worker (SCW) Edge Platform for Smart Manufacturing through implementation of four instructive real‐world use cases that illustrate the role of people in a Smart Manufacturing paradigm through which affordable, scalable, accessible, and portable (ASAP) information technology (IT) acquires and contextualizes data into information for transmission to operation technologies (OT). For case one, the platform captures the relationships between energy consumption and human workflows for improved energy productivity while workers interact with machines during semiconductor manufacturing. The platform utilizes human cognition to identify anomalous machine behavior for root cause analysis of system faults via neural network (NN) that recognize alarm postures of workers with cameras. For case two, a smart assembly line is demonstrated for state monitoring and fault detection. Machine learning (ML) models are used to recognize system states and identify fault scenarios with human intervention. For case three, the platform monitors human–machine interactions to classify manufacturing machine states for proper operations and energy productivity. Internal energy states of individual or collections of manufacturing equipment are determined via NN based algorithms that disaggregate signals associated with smart metering typically deployed at manufacturing facilities. These methods predict the real time energy profile of each machine from the total energy profile of a manufacturing site. For case four, a software defined sensor system built with scientific workflow engines is demonstrated for contextualizing data from laser surface refraction for characterization, and diagnostics in the processing of additively manufactured titanium alloy.

Donovan, Richard P.↗

Adaptive elasticity policies for staging-based in situ visualization

In situ processing aims to alleviate the growing gap between computation and I/O capabilities by performing data processing close to the data source. In situ processing is widely used to process data generated by multiple data sources, including observation data from edge devices or scientific observational facilities and the simulation data generated by scientific computation on a high-performance computing (HPC) platform. For a scientific workflow that is run on an HPC platform and composed of a simulation program and an in situ data analytics or visualization (abbreviated as ana/vis) task, there is an implicit assumption that the computing resources assigned to the workflow keep static during the workflow execution. However, with the converging trend between the HPC and cloud computing platform, running the in situ ana/vis task in an elastic way is promising to decrease its overhead and improve its resource utilization rate. Resource elasticity represents the ability to change resource configurations such as the number of computing nodes/processes during workflow execution. An elastic job may dynamically adjust resource configurations; it may use a few resources at the beginning and more resources toward the end of the job when interesting data appear. However, it is hard to predict a priori how many computing nodes/processes need to be added/removed during the workflow execution to adapt to changing workflow needs. How to efficiently guide elasticity operations, such as growing or shrinking the number of processes used for in situ analysis during workflow execution, is an open-ended research question. In this article, we present adaptive elasticity policies that adopt workflow runtime information collected during workflow execution to predict how to trigger the addition/removal of processes in order to minimize in situ processing overhead. Taking in situ visualization tasks as an example, we integrate the presented elasticity policies into a staging-based elastic workflow and evaluate its efficiency in multiple elasticity scenarios. Compared with the situation without elasticity or with a static elasticity policy that uses a fixed number of processes for each rescaling operation, the adaptive elasticity policy can save overhead in finding a proper resource configuration and improve resource utilization efficiency. Furthermore, one experiment illustrates that the adaptive elasticity policy saves 41% of core-hours compared with the situation without the resource elasticity.

97 MATHEMATICS AND COMPUTING↗

Introducing NOODLES: Collaborative Visualization in a Flavorful Package

We introduce NOODLES, a lightweight collaborative visualization and analysis protocol. NOODLES was developed to enable new scientific workflows while enhancing existing pipelines. Using a technology-minimal approach to strengthen applicability, this protocol allows researchers to tie disparate software together as well as making domain specific visualizations accessible across various platforms and form factors. In this talk we discuss background, challenges, and gaps in the current tool landscape that lead to the development of the protocol. We then describe the design of the protocol and present several use cases to demonstrate the efficacy of NOODLES in addressing complex visualization challenges and how it co-exists with existing tools. We conclude by outlining in-progress work and potential avenues for future development within the NOODLES framework.

collaborative analysis↗

Performance on HPC Platforms Is Possible Without C++

Computing at large scales has become extremely challenging due to increasing heterogeneity in both hardware and software. More and more scientific workflows must tackle a range of scales and use machine learning and AI intertwined with more traditional numerical modeling methods, placing more demands on computational platforms. These constraints indicate a need to fundamentally rethink the way computational science is done and the tools that are needed to enable these complex workflows. The current set of C++-based solutions may not suffice, and relying exclusively upon C++ may not be the best option, especially because several newer languages and boutique solutions offer more robust design features to tackle the challenges of heterogeneity. In June 2023, we held a mini symposium that explored the use of newer languages and heterogeneity solutions that are not tied to C++ and that offer options beyond template metaprogramming and Parallel. For for performance and portability. In conclusion, we describe some of the presentations and discussion from the mini symposium in this article.

97 MATHEMATICS AND COMPUTING↗

Building an Integrated Ecosystem of Computational and Observational Facilities to Accelerate Scientific Discovery

Future scientific discoveries will rely on flexible ecosystems that incorporate modern scientific instruments, high performance computing resources, parallel distributed data storage, and performant networks across multiple, independent facilities. In addition to connecting physical resources, such an ecosystem presents many challenges in logistics and accessibility, especially in orchestrating computations and experiments that span across leadership computing systems and experimental instruments. Past efforts have typically been application-specific or limited to interfaces for computing resources. This paper proposes a general framework for integrating computation resources and instrument operations, addressing challenges in code development/execution, data staging and collection, software stack, control mechanisms, resource authorization and governance, and hardware integration. We also describe a demonstration use case wherein a Bayesian optimization algorithm running on an edge computing resource guides a scanning probe microscope to autonomously and intelligently characterize a material sample. This science edge ecosystem framework will provide a blueprint for federating multi-institutional, disparate resources and orchestrating scientific workflows across them to enable next-generation discoveries.

Somnath, Suhas↗

On the Feasibility of Simulation-Driven Portfolio Scheduling for Cyberinfrastructure Runtime Systems

Runtime systems that automate the execution of applications on distributed cyberinfrastructures need to make scheduling decisions. Researchers have proposed many scheduling algorithms, but most of them are designed based on analytical models and assumptions that may not hold in practice. The literature is thus rife with algorithms that have been evaluated only within the scope of their underlying assumptions but whose practical effectiveness is unclear. It is thus difficult for developers to decide which algorithm to implement in their runtime systems.To obviate the above difficulty, we propose an approach by which the runtime system executes, throughout application execution, simulations of this very execution. Each simulation is for a different algorithm in a scheduling algorithm portfolio, and the best algorithm is selected based on simulation results. The main objective of this work is to evaluate the feasibility and potential merit of this portfolio scheduling approach, even in the presence of simulation inaccuracy, when compared to the traditional one-algorithm approach. We perform this evaluation via a case study in the context of scientific workflows. Our main finding is that portfolio scheduling can outperform the best one-algorithm approach even in the presence of relatively large simulation inaccuracies.

Casanova, Henri↗