Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Scientific Workflows”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

Quantum Computing Technology Roadmaps and Capability Assessment for Scientific Computing - An analysis of use cases from the NERSC workload

The National Energy Research Scientific Computing Center (NERSC), as the high-performance computing (HPC) facility for the Department of Energy’s Office of Science, recognizes the essential role of quantum computing in its future mission. In this report, we analyze the NERSC workload and identify materials science, quantum chemistry, and high-energy physics as the science domains and application areas that stand to benefit most from quantum computers. These domains jointly make up over 50% of the current NERSC production workload, which is illustrative of the impact quantum computing could have on NERSC’s mission going forward. We perform an extensive literature review and determine the quantum resources required to solve classically intractable problems within these science domains. This review also shows that the quantum resources required have consistently decreased over time due to algorithmic improvements and a deeper understanding of the problems. At the same time, public technology roadmaps from a collection of ten quantum computing companies predict a dramatic increase in capabilities over the next five to ten years. Our analysis reveals a significant overlap emerging in this time frame between the technological capabilities and the algorithmic requirements in these three scientific domains. We anticipate that the execution time of large-scale quantum workflows will become a major performance parameter and propose a simple metric, the Sustained Quantum System Performance (SQSP), to compare system-level performance and throughput for a heterogeneous workload.

97 MATHEMATICS AND COMPUTING↗

Hybrid learning techniques for scientific data reduction with performance guarantees

The research initiatives supported by the U.S. Department of Energy (DOE) Grant DE-SC0022265 are fundamentally aimed at pioneering advanced machine learning (ML) techniques for scientific data compression within high-performance computing (HPC) environments. This comprehensive body of work addresses the critical challenge posed by the exponential growth of data generated by scientific simulations in domains such as fusion energy, climate modeling, and computational fluid dynamics (CFD). A core objective is to develop compression algorithms that achieve substantial data reduction—often by orders of magnitude—while rigorously ensuring the fidelity of both the primary data (PD) and scientifically crucial derived quantities of interest (QoI). The methodologies deployed under this grant integrate sophisticated deep learning architectures, prominently featuring autoencoders, advanced generative models like conditional diffusion, and hybrid learning techniques. Key innovations include the development of Guaranteed Autoencoders (GAE) and the Guaranteed Conditional Diffusion with Tensor Correction (GCDTC) framework, which provide explicit, instance-level error bounds on reconstructed data. Furthermore, specialized strategies such as nonlinear constraint satisfaction are employed to preserve the integrity of QoI, a vital requirement for the trustworthiness of downstream scientific analyses. This research also focuses on the design and implementation of scalable, GPU-accelerated software pipelines that seamlessly integrate into existing HPC workflows, ensuring both computational efficiency and practical applicability. The CAESAR framework, for example, unifies foundation and generative models to create an adaptive and efficient compression solution for spatio-temporal scientific data. Collectively, these efforts represent a significant advancement in mitigating the scientific data deluge, enabling more effective data management, accelerated scientific discovery, and optimized utilization of HPC resources.

97 MATHEMATICS AND COMPUTING↗

Final report- UFL - RAPIDS2: A SciDAC Institute for Computer Science, Data, and Artificial Intelligence

The research initiatives supported by the U.S. Department of Energy (DOE) Grant DE-SC0022265 are fundamentally aimed at pioneering advanced machine learning (ML) techniques for scientific data compression within high-performance computing (HPC) environments. This comprehensive body of work addresses the critical challenge posed by the exponential growth of data generated by scientific simulations in domains such as fusion energy, climate modeling, and computational fluid dynamics (CFD). A core objective is to develop compression algorithms that achieve substantial data reduction—often by orders of magnitude—while rigorously ensuring the fidelity of both the primary data (PD) and scientifically crucial derived quantities of interest (QoI). The methodologies deployed under this grant integrate sophisticated deep learning architectures, prominently featuring autoencoders, advanced generative models like conditional diffusion, and hybrid learning techniques. Key innovations include the development of Guaranteed Autoencoders (GAE) and the Guaranteed Conditional Diffusion with Tensor Correction (GCDTC) framework, which provide explicit, instance-level error bounds on reconstructed data. Furthermore, specialized strategies such as nonlinear constraint satisfaction are employed to preserve the integrity of QoI, a vital requirement for the trustworthiness of downstream scientific analyses. This research also focuses on the design and implementation of scalable, GPU-accelerated software pipelines that seamlessly integrate into existing HPC workflows, ensuring both computational efficiency and practical applicability. The CAESAR framework, for example, unifies foundation and generative models to create an adaptive and efficient compression solution for spatio-temporal scientific data. Collectively, these efforts represent a significant advancement in mitigating the scientific data deluge, enabling more effective data management, accelerated scientific discovery, and optimized utilization of HPC resources.

97 MATHEMATICS AND COMPUTING↗

Exascale workflow applications and middleware: An ExaWorks retrospective

Exascale computers offer transformative capabilities to combine data-driven and learning-based approaches with traditional simulation applications to accelerate scientific discovery and insight. However, these software combinations and integrations are difficult to achieve due to the challenges of coordinating and deploying heterogeneous software components on diverse and massive platforms. Here, we present the ExaWorks project, which addresses many of these challenges. We developed a workflow Software Development Toolkit (SDK), a curated collection of workflow technologies that can be composed and interoperated through a common interface, engineered following current best practices, and specifically designed to work on HPC platforms. ExaWorks also developed PSI/J, a job management abstraction API, to simplify the construction of portable software components and applications that can be used over various HPC schedulers. The PSI/J API is a minimal interface for submitting and monitoring jobs and their execution state across multiple and commonly used HPC schedulers. We also describe several leading and innovative workflow examples of ExaWorks tools used on DOE leadership platforms. Furthermore, we discuss how our project is working with the workflow community, large computing facilities, and HPC platform vendors to address the requirements of workflows sustainably at the exascale.

97 MATHEMATICS AND COMPUTING↗

A Review of Extra-Terrestrial Regolith Excavation Concepts and Prototypes

Regolith is present on many extra-terrestrial bodies, and the crushed rock material it is made of contains many of the resources that are enabling for In-Situ Resource Utilization (ISRU). When extracted these resources can be used to provide consumables such as rocket propellant, human life support, working fluids and gases for industrial processes and feedstocks for manufacturing. In addition, the regolith can also be very beneficial for construction purposes as an aggregate which can be used for construction materials and shielding for radiation protection and micrometeorite impact. Binders for regolith concrete may also be made from geopolymers that may be in the regolith. The regolith can be melted and drawn out into glass fibers and used as reinforcements in a metal, polymer, or concrete matrix. In addition, there is tremendous scientific and geological knowledge that can only be obtained by studying samples of the regolith. However, none of these valuable activities can proceed without first acquiring the regolith granular material with some type of excavation device and method. Excavation is in the critical path of many workflows that will make up the capabilities required to establish a human and robotic presence in our solar system. While scientific in-situ sampling of regolith in small quantities has been achieved since the dawn of the space age in the 1960’s, large scale excavation for mining and construction on extra-terrestrial bodies has only been contemplated for many decades in works of scientific fact and also in fictional stories, but serious development and prototyping of excavation technologies for use in reduced gravity space environments was only started in the late 1990’s.This paper will review and document the evolution of extra-terrestrial excavation concepts and prototypes based on the available literature and the personal experience of the author who has been working on regolith excavation technology development since 1998.

ISRU↗

A Review of Extra-Terrestrial Regolith Excavation Concepts and Prototypes

Regolith is present on many extra-terrestrial bodies, and the crushed rock material it is made of contains many of the resources that are enabling for In-Situ Resource Utilization (ISRU). When extracted, these resources can be used to provide consumables such as rocket propellant, human life support, working fluids and gases for industrial processes and feedstocks for manufacturing. In addition, the regolith can also be very beneficial for construction purposes as an aggregate which can be used for construction materials and shielding for radiation protection and micrometeorite impact. Binders for regolith concrete may also be made from geopolymers that may be in the regolith. The regolith can be melted and drawn out into glass fibers and used as reinforcements in a metal, polymer, or concrete matrix. In addition, there is tremendous scientific and geological knowledge that can only be obtained by studying samples of the regolith. However, none of these valuable activities can proceed without first acquiring the regolith granular material with some type of excavation device and method. Excavation is in the critical path of many workflows that will make up the capabilities required to establish a human and robotic presence in our solar system. While scientific in-situ sampling of regolith in small quantities has been achieved since the dawn of the space age in the 1960’s, large scale excavation for mining and construction on extra-terrestrial bodies has only been contemplated, for many decades, but serious development and prototyping of excavation technologies for use in reduced gravity space environments was only started in the late 1990’s. This paper will review and document the evolution of extra-terrestrial excavation concepts and prototypes based on the available literature and the personal experience of the author who has been working on regolith excavation technology development since 1998.

Regolith↗

A Review of Extra-Terrestrial Regolith Excavation Concepts and Prototype

Regolith is present on many extra-terrestrial bodies, and the crushed rock material it is made of contains many of the resources that are enabling for In-Situ Resource Utilization (ISRU). When extracted, these resources can be used to provide consumables such as rocket propellant, human life support, working fluids and gases for industrial processes and feedstocks for manufacturing. In addition, the regolith can also be very beneficial for construction purposes as an aggregate which can be used for construction materials and shielding for radiation protection and micrometeorite impact. Binders for regolith concrete may also be made from geopolymers that may be in the regolith. The regolith can be melted and drawn out into glass fibers and used as reinforcements in a metal, polymer, or concrete matrix. In addition, there is tremendous scientific and geological knowledge that can only be obtained by studying samples of the regolith. However, none of these valuable activities can proceed without first acquiring the regolith granular material with some type of excavation device and method. Excavation is in the critical path of many workflows that will make up the capabilities required to establish a human and robotic presence in our solar system. While scientific in-situ sampling of regolith in small quantities has been achieved since the dawn of the space age in the 1960’s, large scale excavation for mining and construction on extra-terrestrial bodies has only been contemplated, for many decades, but serious development and prototyping of excavation technologies for use in reduced gravity space environments was only started in the late 1990’s. This paper will review and document the evolution of extra-terrestrial excavation concepts and prototypes based on the available literature and the personal experience of the author who has been working on regolith excavation technology development since 1998.

Regolith↗

DTLMod: A simulation framework for in situ workflow optimization

In situ processing workflows have become essential for coping with the explosion in data volume and velocity in large-scale scientific computing, providing domain scientists with early insights at runtime. Multiple frameworks implement this paradigm through a data transport layer (DTL), offering different data access modes and deployment schemes, but researchers currently lack the appropriate tools to assess design and deployment options before committing to costly real experiments. We introduce DTLMod, an open-source simulated DTL that enables performance evaluation of in situ workflow configurations at scale. Built on SimGrid, it links into any SimGrid-based simulator and is available in C++ and Python. We evaluate DTLMod along four axes: scalability (tens of thousands of simulated processes across interconnected clusters in seconds, with linear memory scaling), versatility (three implementation variants trading fidelity for speed), accuracy (simulated times faithfully reflecting real behavior), and practical utility (two use cases demonstrating evidence-based workflow design decisions).

Suter, Fred [ORNL] (ORCID:0000000319021955)↗

Large language model-driven database for thermoelectric materials

Thermoelectric materials have the ability to convert waste heat into electricity, offering a valuable solution for energy harvesting. However, their widespread use is hindered by low conversion efficiency, the reliance on expensive rare earth elements, and the environmental and regulatory concerns associated with lead-based materials. A fast and cost-effective way to identify highly efficient thermoelectric materials is through data-driven methods. These approaches rely on robust and comprehensive datasets to train models. Although there are several databases on thermoelectric materials, there is still a need to collect and integrate experimental data from peer-reviewed research articles to capture diverse compositions and properties of materials. Here, in this work, we developed a comprehensive database of 7,123 thermoelectric compounds, containing key information such as chemical composition, structural detail, seebeck coefficient, electrical and thermal conductivity, power factor, and figure of merit (ZT). We used the GPTArticleExtractor workflow, powered by large language models (LLM), to extract and curate data automatically from the scientific literature published in Elsevier journals. This process enabled the creation of a structured database that addresses the challenges of manual data collection. The open access database could stimulate data-driven research and advance thermoelectric material analysis and discovery.

Database↗

CalyxFlow

CalyxFlow is a lightweight agentic artificial intelligent workflow. This workflow demonstrates the use of AI LLMs to generate modeling and simulation inputs for a scientific simulation and manage execution and analysis of a suite of simulations.

Shipman, Galen↗

Cryogenic EM Across Length Scales for Li Metal Anode Batteries

An interfacial understanding is necessary for developing strategies to commercialize high-energy density rechargeable lithium metal anode batteries, as currently, the lithium anode/electrolyte interface is unstable with prolonged cycling. We have used several strategies to improve the cycling performance of lithium metal anodes, including reducing the parasitic reactions between lithium metal and the electrolyte, and improving the electrodeposited lithium metal morphology. These strategies have generated unconclusive electrochemical data, that has required the need for nanoscale interfacial characterization of these solid-liquid interfaces. Our team has used the cryogenic transfer workflow developed by Leica in collaboration with cryo-SEM/FIB tools by Thermo Fisher Scientific to cross-section lithium metal anodes and intact coin cell batteries to observe the interfacial structures, lithium morphology, and failure mechanisms relative to changes in electrode contract pressure and electrolyte chemistry. Cross-sectional SEM images and EDS maps of the lithium metal anodes have provided a better understanding of the electrodeposited lithium morphology, quantity of 'dead' lithium metal, and quantity of solid electrolyte interphase material that has formed alongside the lithium metal. In understanding lithium metal battery failure at the system level, we used a cryogenic stage in a laser plasma FIB to cross-section through the coin cell's cap for imaging/mapping the entire battery stack under cryogenic conditions. The tools, methods, and results of these studies will be detailed in this presentation.

characterization↗

Analytical Needs in a Sample Receiving Facility: Input from the MSR Operation Definition Team

The return of scientifically selected samples from Mars would provide a rare opportunity forinvestigation with the full range of the latest technology available, but to take full advantageof this opportunity, it is important to plan ahead to ensure the pristine nature of the samplesupon arrival within the Earth environment until scientific investigations can begin.The NASA/ESA science community-driven MSR Science Planning Group – Phase 2 (MSPG2)delivered recommendations and guidance regarding curation (1) and science (2, 3) activities tobe performed on the samples under containment. High-level requirements for the infrastruc-ture were also developed by MSPG2 (4). In order to prepare infrastructure-targeted input forthe ESA and NASA facility studies planned in the 2022-2023 timeframe, the agency-led MSROperational Scenarios Definition Team (MOSDT) was assembled to conceptualize the sampleoperations that will inform future architecture teams. Emphasis was placed on the respon-sibility of MOSDT to use community-defined requirements and to represent the view of the international scientific community.All necessary and sufficient instruments and analytical needs described in MSPG2 were inte-grated in MOSDT main deliverable, the operational workflow (see Hays et al, this conference).In MSPG2, notional instruments were split between curation analytical needs, and objective-driven (time-sensitive and sterilization-sensitive) science analytical needs. In MOSDT, whilethe first phases of curation, “pre-Basic Characterization” and “Basic Characterization” wererather streamlined and separate from other analytical needs, “Preliminary Examination” and“Science” instruments were not always physically segregated. In addition to the necessary andsufficient instruments described by MSPG2, the MOSDT recommended additional supportequipment for sterilization, cleanliness and contamination monitoring.It was sometimes necessary for the MOSDT to rely on assumptions to integrate instruments inthe activity workflow. In general, the assumptions were very conservative to limit contaminationand cross-contamination risks. It is expected that future work to refine limits of contaminationwill enable optimization of instrumentation.The community was consulted during the course of the MOSDT work. This abstract’s aimis two-fold: on one hand, inform the scientific community and overall MSR stakeholders, tobring their attention on the analytical needs currently considered as necessary and sufficient;on the other hand, to solicit feedback from a larger community audience to optimize and refineanalytical needs during the next phases of MSR ground-segment preparation.Disclaimer: The decision to implement Mars Sample Return will not be finalized until NASA’scompletion of the program’s National Environmental Policy Act (NEPA) process. This docu-ment is being made available for informational purposes only.[1] Tait et al. (2021) Preliminary planning for Mars Sample Return (MSR) curation activities ina Sample Receiving Facility (SRF). Astrobiology in press, doi:10.1089/ast.2021.0105. [2] Toscaet al. (2021) Time-sensitive aspects of Mars Sample Return (MSR) science. Astrobiologyin press, doi:10.1089/ast.2021.0115. [3] Velbel et al. (2021) Planning implications relatedto sterilization-sensitive science investigations associated with Mars Sample Return (MSR).Astrobiology in press, doi:10.1089/ast.2021.0113. [4] Carrier et al. (2021) Science and curationconsiderations for the design of a Mars Sample Return (MSR) Sample Receiving Facility (SRF).Astrobiology in press, doi:10.1089/ast.2021.0110.

Mars Sample Return↗

Accelerating discoveries at DIII-D with the Integrated Research Infrastructure

DIII-D research is being accelerated by leveraging high performance computing (HPC) and data resources available through the National Energy Research Scientific Computing Center (NERSC) Superfacility initiative. As part of this initiative, a high-resolution, fully automated, whole discharge kinetic equilibrium reconstruction workflow was developed that runs at the NERSC for most DIII-D shots in under 20 min. This has eliminated a long-standing research barrier and opened the door to more sophisticated analyses, including plasma transport and stability. These capabilities would benefit from being automated and executed within the larger Department of Energy Advanced Scientific Computing Research program’s Integrated Research Infrastructure (IRI) framework. The goal of IRI is to empower researchers to meld DOE’s world-class research tools, infrastructure, and user facilities seamlessly and securely in novel ways to radically accelerate discovery and innovation. For transport, we are looking at producing flux matched profiles and also using particle tracing to predict fast ion heat deposition from neutral beam injection before a shot takes place. Our starting point for evaluating plasma stability focuses on the pedestal limits that must be navigated to achieve better confinement. This information is meant to help operators run more effective experiments, so it needs to be available rapidly inside the DIII-D control room. So far this has been achieved by ensuring the data is available with existing tools, but as more novel results are produced new visualization tools must be developed. In addition, all of the high-quality data we have generated has been collected into databases that can unlock even deeper insights. This has already been leveraged for model and code validation studies as well as for developing AI/ML surrogates. The workflows developed for this project are intended to serve as prototypes that can be replicated on other experiments and can be run to provide timely and essential information for ITER, as well as next stage fusion power plants.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

ChemGraph as an agentic framework for computational chemistry workflows

Atomistic simulations are essential in chemistry and materials science but remain challenging to run due to the expert knowledge required for the setup, execution, and validation stages of these calculations. We present ChemGraph, an agentic framework powered by artificial intelligence and state-of-the-art simulation tools to streamline and automate computational chemistry and materials science workflows. ChemGraph leverages graph neural network-based foundation models for accurate yet computationally efficient calculations and large language models (LLMs) for natural language understanding, task planning, and scientific reasoning to provide an intuitive and interactive interface. We evaluate ChemGraph across 13 benchmark tasks and demonstrate that smaller LLMs (GPT-4o-mini, Claude-3.5-haiku, Qwen-2.5-14B) perform well on simple workflows, while more complex tasks benefit from using larger models. Importantly, we show that decomposing complex tasks into smaller subtasks through a multi-agent framework enables GPT-4o to reach perfect accuracy and smaller LLMs to match or exceed single-agent GPT-4o's performance in these benchmarks.

Computational chemistry↗

A Brief Survey of Data Streaming Technologies

Streaming data is data that is emitted at variable volumes in a continuous, incremental manner with the goal of low-latency processing often at a different physical location. Network infrastructure is used to facilitate the connection between data sources and sinks, and must be robust to handle the requirements of the workflow. The U.S. Department of Energy Office of Science (DOE SC) a federal agency supporting fundamental scientific research for energy and the Nation’s largest supporter of basic research in the physical sciences. DOE SC has the responsibility for operating $\mathbf{1 0}$ National Laboratories, and 28 scientific user facilities supporting advanced supercomputers, particle accelerators, large x-ray light sources, neutron scattering sources, and other specialized facilities for nanoscience and genomics. This paper investigates the state of streaming data workfows, and details some of the approaches to this challenging problem.

Kissel, Ezra↗

Policy Considerations When Federating Facilities for Experimental and Observational Data Analysis

Today’s computational, experimental, and observational facilities afford us tremendous opportunities to couple theory and experiment at increasingly large scales. Empirical sensing capabilities are growing dramatically with beam line and detector improvements, and with advances in our ability to deploy large-scale data gathering observations of the natural world. The coupling of computational simulations and analysis to process the data from experimental and observational facilities is giving rise to cross-facility workflows. Such federations of facilities are in fact becoming an explicit requirement for large-scale scientific discovery. As we scale up these pipelines of scientific discovery, each participating facility needs to establish and align policies so that the federation can work seamlessly in an end-to-end manner. This chapter outlines specific policy considerations in enabling the federation of facilities for data analysis. Design choices and vital policy decisions cover the areas of data acquisition and storage, data transfer, computational resource allocation and co-scheduling, seamless federated user access, and cross-cutting governance. By highlighting the explicit and implicit interdependencies between facilities, we aim to provide facility designers and policymakers the information on policy issues to address early in a facility’s operations, thus enabling successful cross-facility federation and improved experimental and observational data analysis outcomes.

Shankar, Mallikarjun (Arjun)↗

Distinguishing Provenance Equivalence of Earth Science Data

Reproducibility of scientific research relies on accurate and precise citation of data and the provenance of that data. Earth science data are often the result of applying complex data transformation and analysis workflows to vast quantities of data. Provenance information of data processing is used for a variety of purposes, including understanding the process and auditing as well as reproducibility. Certain provenance information is essential for producing scientifically equivalent data. Capturing and representing that provenance information and assigning identifiers suitable for precisely distinguishing data granules and datasets is needed for accurate comparisons. This paper discusses scientific equivalence and essential provenance for scientific reproducibility. We use the example of an operational earth science data processing system to illustrate the application of the technique of cascading digital signatures or hash chains to precisely identify sets of granules and as provenance equivalence identifiers to distinguish data made in an an equivalent manner.

Tilmes, Curt↗

A Vision for Coupling Operation of US Fusion Facilities with HPC Systems and the Implications for Workflows and Data Management

The operation of large US Department of Energy (DOE) research facilities, like the DIII-D National Fusion Facility, results in the collection of complex multi-dimensional scientific datasets, both experimental and model-generated. In the future, it is envisioned that integrated data analysis coupled with large-scale high performance computing (HPC) simulations will be used to improve experimental planning and operation. Practically, massive data sets from these simulations provide the physics basis for generation of both reduced semi-analytic and machine-learning-based models. Storage of both HPC simulation datasets (generated from US DOE leadership computing facilities) and experimental datasets presents significant challenges. In this paper, we present a vision for a DOE-wide data management workflow that integrates US DOE fusion facilities with leadership computing facilities. Data persistence and long-term availability beyond the length of allocated projects is essential, particularly for verification and recalibration of artificial intelligence and machine learning (AI/ML) models. Because these data sets are often generated and shared among hundreds of users across multiple leadership computing facility centers, they would benefit from cross-platform accessibility, persistent identifiers (e.g. DOI, or digital object identifier), and provenance tracking. Here, the ability to handle different data access patterns suggests that a combination of low cost, high latency (e.g. for storing ML training sets) and high cost, low latency systems (e.g. for real-time, integrated machine control feedback) may be needed.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗