Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Scientific Workflows”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

UWLi (Universal Workflow Language Interface) [SWR-24-12]

UWLi is a user interface for modifying Universal Workflow Language (UWL) files. UWL is a data format used to represent high fidelity scientific procedures in a generalized, field agnostic workflow format. See related publication: https://arxiv.org/pdf/2409.05899

Epps, Robert↗

Reining in an Agentic Harness for High Energy Physics

Agentic systems now address tasks across theoretical, phenomenological, and experimental high energy physics (HEP), but their scientific capabilities remain difficult to reuse across different large language models, providers, and harnesses. We argue that stable parts of these workflows should be promoted into versioned scientific operations and exposed through common protocols. Existing general-purpose harnesses can then be specialized for HEP through task-specific sets of tools and skills, while community-maintained registries would make these capabilities discoverable and citable. We identify mismatches in conventions, assumptions, and domains of validity among independently developed operations as a potential obstacle to their composition, and discuss machine-readable scientific contracts as one possible solution. These design principles and evaluation guidelines provide a near-term path toward a portable and community-maintained agentic harness for HEP.

Menzo, Tony [Alabama U.; Fermilab] (ORCID:00000002↗

Utilizing Distributed Heterogeneous Computing with PanDA in ATLAS

In recent years, advanced and complex analysis workflows have gained increasing importance in the ATLAS experiment at CERN, one of the large scientific experiments at LHC. Support for such workflows has allowed users to exploit remote computing resources and service providers distributed worldwide, overcoming limitations on local resources and services. The spectrum of computing options keeps increasing across the Worldwide LHC Computing Grid (WLCG), volunteer computing, high-performance computing, commercial clouds, and emerging service levels like Platform-as-a-Service (PaaS), Container-as-a-Service (CaaS) and Function-as-a-Service (FaaS), each one providing new advantages and constraints. Users can significantly benefit from these providers, but at the same time, it is cumbersome to deal with multiple providers, even in a single analysis workflow with fine-grained requirements coming from their applications’ nature and characteristics. In this paper, we will first highlight issues in geographically-distributed heterogeneous computing, such as the insulation of users from the complexities of dealing with remote providers, smart workload routing, complex resource provisioning, seamless execution of advanced workflows, workflow description, pseudointeractive analysis, and integration of PaaS, CaaS, and FaaS providers. We will also outline solutions developed in ATLAS with the Production and Distributed Analysis (PanDA) system and future challenges for LHC Run4.

97 MATHEMATICS AND COMPUTING↗

Extreme-scale workflows: A perspective from the JLESC international community

The Joint Laboratory for Extreme-Scale Computing (JLESC) focuses on software challenges in high-performance computing systems to meet the needs of today’s science campaigns, which often require large resources, consist of multiple tasks, and generate vast amounts of data. In this context, extreme-scale workflows have been the key factor in enabling scientific discoveries by helping scientists automate the dependencies and data exchanges between workflow tasks, instead of managing those manually. Here, in this paper, we present representative extreme-scale workflows and feature workflow systems developed by JLESC participating institutions. We present lessons learned while developing these tools, alongside with the open challenges and future research directions in the field of extreme-scale workflows.

97 MATHEMATICS AND COMPUTING↗

Big PanDa Workflow Management on Titan for High Energy and Nuclear Physics and for Future Extreme Scale Scientific Application

Over a three year period, from 2016-2019, this project demonstrated the scientific benefits of integrating the Titan supercomputer at Oak Ridge Leadership Computing Facility into traditional high throughput grid based distributed computing systems managed by PanDA, the workflow management system used for the execution of all distributed computing applications by the ATLAS experiment at the Large Hadron Collider. PanDA manages millions of batch jobs daily at hundreds of clusters worldwide on request by thousands of physicist users, and processes more than an exabyte of data annually using grid middleware. High levels of operational use of Titan was sustained by PanDA in order to meet the physics goals of ATLAS. The success of this project led to the use of other supercomputers worldwide by ATLAS, and to the adoption of PanDA by other experiments and other scientists. Multiple innovative operational and computer science research goals were achieved supporting the use of supercomputers for scientific domains with large scale distributed data and distributed processing needs.

97 MATHEMATICS AND COMPUTING↗

CVEVOLVE

CVEvolve is an agentic AI system for autonomous algorithm discovery for scientific data processing. It creates workflows where large language model agents freely set up and configure development environments and evaluation harnesses, develop and improve data processing algorithms with designed exploration-exploitation balancing mechanisms, log history and findings in a structured database, and run holdout testing to ensure algorithm generalizability. CVEvolve offers a zero-code interface and does not require users to provide structured data and evaluation scripts.

Cherukara, MatthewJoseph [Argonne National Laborat↗

End-to-end codesign of Hessian-aware quantized neural networks for FPGAs

Here, we develop an end-to-end workflow for the training and implementation of co-designed neural networks (NNs) for efficient field-programmable gate array (FPGA) hardware. Our approach leverages Hessian-aware quantization of NNs, the Quantized Open Neural Network Exchange intermediate representation, and the hls4ml tool flow for transpiling NNs into FPGA firmware. This makes efficient NN implementations in hardware accessible to nonexperts in a single open sourced workflow that can be deployed for real-time machine-learning applications in a wide range of scientific and industrial settings. We demonstrate the workflow in a particle physics application involving trigger decisions that must operate at the 40-MHz collision rate of the CERN Large Hadron Collider (LHC). Given the high collision rate, all data processing must be implemented on FPGA hardware within the strict area and latency requirements. Based on these constraints, we implement an optimized mixed-precision NN classifier for high-momentum particle jets in simulated LHC proton-proton collisions.

47 OTHER INSTRUMENTATION↗

Operational Workflow in a Sample Receiving Facility: Input from the MSR Operation Definition Team

The return of scientifically selected samples from Mars would provide a rare opportunity for investigation with the full range of the latest technology available. To take full advantage of this opportunity, it is important to plan ahead to ensure the pristine nature of the samples upon arrival within the Earth environment until scientific investigations can begin. The NASA/ESA science community-driven MSR Science Planning Group – Phase 2 (MSPG2) delivered recommendations and guidance regarding curation (1) and science (2,3) activities to be performed on the samples under containment. High-level requirements for the infrastructure were also developed by MSPG2 (4). In order to prepare infrastructure-targeted input for the ESA and NASA facility studies planned in the 2022-2023 timeframe, the MSR agency-led Operational Scenarios Definition Team (MOSDT) was assembled to conceptualize the sample operations that will inform future architecture teams. Emphasis was placed on the responsibility of MOSDT to use community-defined requirements and to represent the view of the international scientific community. The main deliverable of MOSDT was an operational workflow for a Sample Receiving Facility (SRF). Two other deliverables were produced: a report to narrate the workflow, and a list of instruments (see Hutzler et al., this conference). Activities described in the main sequence of the workflow range from engineering operations to curation to science, with the latter term being used here as the science to be done within a SRF. Side sequences (e.g. engineering inspection of hardware, head gas extraction) were also identified, and detailed when they would have a significant impact on the infrastructure of a SRF. It was necessary for the MOSDT to rely on assumptions for some steps and activities, and though these were kept to a minimum (and are described in both the report supporting the workflow and in the full presentation), in general, the assumptions and overall work were very conservative, as the impact of underestimating the scope of the SRF infrastructure was considered more detrimental than overestimating it. It is expected that future work will be able to confirm or inform these assumptions. The community was consulted during the course of the MOSDT work. This abstract’s aim is two-fold: on one hand, inform the scientific community and overall MSR stakeholders, to make the infrastructure studies and trade-off more understandable; on the other hand, to solicit feedback from a larger community audience for the next iterations planning for SRF design and activities.

Mars Sample Return↗

Enabling machine learning-ready HPC ensembles with Merlin

With the growing complexity of computational and experimental facilities, many scientific researchers are turning to machine learning (ML) techniques to analyze large scale ensemble data. With complexities such as multi-component workflows, heterogeneous machine architectures, parallel file systems, and batch scheduling, care must be taken to facilitate this analysis in a high performance computing (HPC) environment. Here, we present Merlin, a workflow framework to enable large ML-friendly ensembles of scientific HPC simulations. By augmenting traditional HPC with distributed compute technologies, Merlin aims to lower the barrier for scientific subject matter experts to incorporate ML into their analysis. As a producer–consumer workflow model, Merlin enables multi-machine, cross-batch job, dynamically allocated yet persistent workflows capable of utilizing surge-compute resources. Key features of Merlin are a flexible HPC-centric interface, low per-task overhead, multi-tiered fault recovery, and a hierarchical sampling algorithm that allows for $\mathscr{O}$(N) task execution and $\mathscr{O}$(N ln N) task queuing to ensembles of millions of tasks. In addition to Merlin’s design, we test the algorithm’s performance in an HPC center and demonstrate the ability to enqueue 40 million simulations in 100 s, with a 30 millisecond per-task overhead that is independent of ensemble size. Finally, we describe some example applications that Merlin has enabled on leadership-class HPC resources, such as the ML-augmented optimization of nuclear fusion experiments and the calibration of infectious disease models to study the progression of and possible mitigation strategies for COVID-19.

97 MATHEMATICS AND COMPUTING↗

Immersive Scientific Visualization of Molten-Salt Reactor Waste Characteristics Using Virtual Reality

Immersive visualization is changing how we explore, communicate, and understand complex scientific systems. In nuclear energy, an area in which data are often multidimensional, time-dependent, and difficult to interpret, virtual reality (VR) represents a powerful and intuitive informational medium. This work introduces a VR-based platform that visualizes the post-shutdown behavior and waste management lifecycle of molten-salt reactors (MSRs), a next-generation reactor type with unique operational and safety characteristics. The platform, built in Unity, is streamed on the Meta Quest 3 headset. It transforms high-fidelity simulation data into an interactive, immersive experience. Users can explore time-dependent reactor characteristics such as nuclide decay, which is a key factor for evaluating reactor waste strategies. The datasets were generated using the MOOSE (Multiphysics Object-Oriented Simulation Environment) framework and then processed through ParaView scripting for smooth integration into Unity. From a visualization standpoint, the platform emphasizes spatial storytelling, temporal exploration, and user-centered interaction. Users can navigate 3D reactor geometries, slice through volumetric data, and manipulate time to observe how physical phenomena evolve. Real-scale rendering and embodied interaction make the experience feel tangible. The interface is designed to be accessible, even to those without nuclear or simulation expertise. This lowers the barrier for stakeholders, policymakers, and the general public, while still supporting expert analysis and collaborative decision-making. This work shows how immersive visualization can function as both a scientific tool and a communication interface. By integrating simulation, processing, and visualization into a cohesive workflow, we offer a scalable framework for immersive scientific storytelling. The modular design supports future extensions to other reactor types and lifecycle stages, from shutdown to long-term storage, making the platform adaptable for both research and outreach.

99 - GENERAL AND MISCELLANEOUS↗

A Versatile Simulated Data Transport Layer for in Situ Workflows Performance Evaluation

In situ processing does not only allow scientific applications to face the explosion in data volume and velocity but also to address the time constraints of many simulation-analysis workflows by providing scientists with early insights about their applications at runtime. Multiple frameworks implement the concept of a data transport layer (DTL) to enable such in situ workflows. These tools are very versatile, directly or indirectly access the data generated on the same node, another node of the same compute cluster, or a completely distinct node, and allow data publishers and subscribers to run on the same computing resources or not. This versatility puts on researchers the onus of taking key decisions related to resource allocation and how to transport data to ensure the most efficient execution of their in situ workflows. However, domain scientists and workflow practitioners lack the appropriate tools to assess the respective performance of particular design and deployment options. In this paper we introduce a versatile simulated DTL designed to provide researchers with insights on the respective performance of different execution scenarios of in situ workflows. This open-source, standalone library builds on the SimGrid toolkit and can be linked to any SimGrid-based simulator. It facilitates the evaluation of the performance behavior, at scale, of different data transport configurations and the study of the effects of resource allocation strategies. We demonstrate the scalability, versatility, and accuracy of this simulated DTL by reproducing the execution of two synthetic benchmarks and of a real-world in situ workflow composed of an MPI application and a parallel data analysis. Results of simulations run on a single core show that the proposed library can simulate the interactions of tens of thousands of simulated processes deployed on two interconnected commodity clusters in a few seconds, and the execution by a thousand simulated processes of an in situ workflow in less than three minutes.

Suter, Fred [ORNL] (ORCID:0000000319021955)↗

Initial User Study Insights Clarified Community Requirements & Early Feedback on Conceptual Design

The HPDF User Experience (UX) team conducted eight semi-structured interviews with twelve individuals leading data and computing infrastructure work at DOE Office of Science (SC) user facilities or projects. This report presents key takeaways synthesizing community perspectives and needs, along with recommendations for the project. These insights should be considered as conceptual design work continues, early partnerships begin, workflow readiness activities evaluate and enhance key scientific tools, and early access systems are made available to the Office of Science (SC) community. Vigilance in meeting community requirements and ensuring workflow readiness will be needed to ensure HPDF’s success.

97 MATHEMATICS AND COMPUTING↗

FAIR Data Meets FAIR Software

Modern scientific research is increasingly defined by the interplay between data, software, and the workflows that connect them. Yet while the FAIR (Findable, Accessible, Interoperable, Reusable) principles have become foundational for scientific data stewardship, the same level of structure and expectation has only recently begun to extend to research software. This talk covers why and how FAIR principles are being applied to data and software to support data reuse. It outlines the gaps in current sharing norms, the growing federal emphasis on persistent identifiers and public access, and the opportunities created when datasets, computational workflows, code, and models are linked through rich, standardized metadata. Practical implementation pathways for the EIC and JLab communities are described, including datacards for structured dataset documentation and provenance-aware workflows. By aligning data lifecycle management with FAIR-aligned software practices, the scientific community can advance toward autonomous knowledge graphs, generative workflows, and high-quality, AI-ready scientific datasets.

McSpadden, Diana [Thomas Jefferson National Accele↗

Towards Lightweight Data Integration Using Multi-Workflow Provenance and Data Observability

Modern large-scale scientific discovery requires multidisciplinary collaboration across diverse computing facilities, including High Performance Computing (HPC) machines and the Edge-to-Cloud continuum. Integrated data analysis plays a crucial role in scientific discovery, especially in the current AI era, by enabling Responsible AI development, FAIR, Reproducibility, and User Steering. However, the heterogeneous nature of science poses challenges such as dealing with multiple supporting tools, cross-facility environments, and efficient HPC execution. Building on data observability, adapter system design, and provenance, we propose MIDA: an approach for lightweight runtime Multi-workflow Integrated Data Analysis. MIDA defines data observability strategies and adaptability methods for various parallel systems and machine learning tools. With observability, it intercepts the dataflows in the background without requiring instrumentation while integrating domain, provenance, and telemetry data at runtime into a unified database ready for user steering queries. We conduct experiments showing end-to-end multi-workflow analysis integrating data from Dask and MLFlow in a real distributed deep learning use case for materials science that runs on multiple environments with up to 276 GPUs in parallel. We show near-zero overhead running up to 100,000 tasks on 1,680 CPU cores on the Summit supercomputer.

Santos Souza, Renan↗

Visualization techniques for the gyrokinetic tokamak simulation code

Gyrokinetic simulations of plasma microturbulence in tokamaks are challenging to visualize because the compute grid follows the magnetic field lines that spiral around the torus. We have overcome this challenge by developing three new approaches that improve visualization of gyrokinetics. Our techniques work directly with the topology of magnetic flux surfaces where the simulation stores variables in concentric rings on poloidal planes (vertical cross sections of the torus). Our visualization preview step triangulates each consecutive pair of rings to display the data on a poloidal plane. The second visualization technique follows spiral field lines around the torus and constructs polygons to visualize a flux surface. Third, the poloidal triangles are connected between planes to form prisms that compose a 3-D model of the entire torus. The visualization workflow produces detailed geometry that matches the high resolution, irregular compute grid for every time step. The surface and solid models are displayed in scientific visualization programs to effectively explore and communicate the results, including fluctuation of electron density, ion temperature, and electrostatic potential. Highly detailed renderings verify plasma behavior along magnetic field lines over time.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Self-Driving Laboratories for Chemistry and Materials Science

Self-driving laboratories (SDLs) promise an accelerated application of the scientific method. Through the automation of experimental workflows, along with autonomous experimental planning, SDLs hold the potential to greatly accelerate research in chemistry and materials discovery. This review provides an in-depth analysis of the state-of-the-art in SDL technology, its applications across various scientific disciplines, and the potential implications for research and industry. This review additionally provides an overview of the enabling technologies for SDLs, including their hardware, software, and integration with laboratory infrastructure. Most importantly, this review explores the diverse range of scientific domains where SDLs have made significant contributions, from drug discovery and materials science to genomics and chemistry. We provide a comprehensive review of existing real-world examples of SDLs, their different levels of automation, and the challenges and limitations associated with each domain.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Fusion Energy Sciences Network Requirements Review: Mild-cycle Update

The US Department of Energy (DOE) Office of Science (SC) world-class research infrastructure provides the research community with premier observational, experimental, computational, and network capabilities. Each user facility is designed to provide unique capabilities to advance core DOE mission science for its sponsor SC program and to stimulate a rich discovery and innovation ecosystem. Research communities gather and flourish around each user facility, bringing together diverse perspectives. The continual reinvention of the practice of science — as users and staff forge novel approaches expressed in research workflows — unlocks new discoveries and propels scientific progress.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

IMPECCABLE: Integrated Modeling Pipeline for COVID Cure by Assessing Better Leads

ABSTRACT The drug discovery process currently employed in the pharmaceutical industry typically requires about 10 years and $2-3 billion to deliver one new drug. This is both too expensive and too slow, especially in emergencies like the COVID-19 pandemic. In silico methodologies need to be improved both to select better lead compounds, so as to improve the efficiency of later stages in the drug discovery protocol, and to identify those lead compounds more quickly. No known methodological approach can deliver this combination of higher quality and speed. Here, we describe an Integrated Modeling PipEline for COVID Cure by Assessing Better LEads (IMPECCABLE) that employs multiple methodological innovations to overcome this fundamental limitation. We also describe the computational framework that we have developed to support these innovations at scale, and characterize the performance of this framework in terms of throughput, peak performance, and scientific results. We show that individual workflow components deliver 100× to 1000× improvement over traditional methods, and that the integration of methods, supported by scalable infrastructure, speeds up drug discovery by orders of magnitudes. IMPECCABLE has screened ~10 11 ligands and has been used to discover a promising drug candidate. These capabilities have been used by the US DOE National Virtual Biotechnology Laboratory and the EU Centre of Excellence in Computational Biomedicine.

97 MATHEMATICS AND COMPUTING↗