Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “HPC workflows”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

PSI/J: A Portable Interface for Submitting, Monitoring, and Managing Jobs

It is generally desirable for high-performance computing (HPC) applications to be portable between HPC systems, for example to make use of more performant hardware, make effective use of allocations, and to co-locate compute jobs with large datasets. Unfortunately, moving scientific applications between HPC systems is challenging for various reasons, most notably that HPC systems have different HPC schedulers. We introduce PSI/J, a job management abstraction API intended to simplify the construction of software components and applications that are portable over various HPC scheduler implementations. We argue that such a system is both necessary and that no viable alternative currently exists. We analyze similar notable APIs and attempt to determine the factors that influenced their evolution and adoption by the HPC community. We base the design of PSI/J on that analysis. We describe how PSI/J has been integrated in three workflow systems and one application, and also show via experiments that PSI/J imposes minimal overhead.

Hategan, Mihael↗

High Performance Computing Peak Shaving for Microreactor Operation

There are multiple nuclear microreactors currently under development that are designed to provide autonomous power for as many as ten or more years without refueling and are designed to power high performance computing (HPC) datacenters. But the load-follow speeds for a nuclear microreactor will be much slower than grid power and slower than the power variance typical of a HPC system. HPC datacenters experience peak power load variance driven by several factors ranging from the operation of cooling systems to remove heat from the servers to supporting a wide range of user application workflows and architectures each with different power signatures. One mechanism to support the limited load-follow of a microreactor is peak shaving where an energy storage mechanism is used to shed peak load and reduce significant power variance. This work explores peak electrical load shaving using uninterruptible power supply (UPS) systems designed for HPC support in the context of peak shaving when operating using a nuclear microreactor with a load-follow limited to 10% of load per minute. Using a self contained HPC datacenter complete with stand-alone cooling system and provisioned with an x86 cluster, an ARM cluster, and a graphics processing unit (GPU) cluster, peak shaving for microreactor operation using the UPS battery backup is explored while running two classes of typical HPC user applications. HPC architecture suitability for microreactor operation under this type of peak shaving is examined.

97 MATHEMATICS AND COMPUTING↗

Treyson Ricks - Intern Showcase Poster

Quinone-based sorbents offer a tunable, energy-efficient route to electrochemical CO2 capture, but systematic guidance for molecular design is lacking. Here, we report a high-throughput computational workflow that combines density functional theory (DFT) screening with machine-learning (ML) modeling to evaluate CO2 binding thermodynamics across several quinone derivatives, spanning benzoquinones, naphthoquinones, and anthraquinones. In addition to using solvents to stabilize the quinone anion and dianion, we studied the effect of ion-pairing on the reduction potentials and the CO2 binding energy. Automated Python scripts handled geometry optimizations and adduct-formation energies on an HPC cluster, reducing manual effort significantly. This integrated platform can uncover structure–property relationships and enables rapid in silico evaluation of untested candidates. We present one example from our workflow to showcase the capability of using quinones with ion-pairing to effectively capture CO2. Our approach paves the way for the rational selection of optimal quinone sorbents and can be extended with experimental thermochemical and kinetic data, alternative redox cycles, and stability assessments to accelerate development of next-generation electrochemical CO2 capture materials.

37 - INORGANIC, ORGANIC, PHYSICAL AND ANALYTICAL C↗

HPC ODA Commons [SWR-26-003]

HPC ODA Commons is a community-driven platform for standardizing HPC operational data analytics. HPC sites generate enormous volumes of operational data - scheduler logs, accounting records, monitoring streams - but turning that data into actionable insight is needlessly hard. Each site builds bespoke parsers, schemas, and evaluation pipelines. Results can't be compared across institutions. Promising analytics ideas stay siloed because there's no shared language for describing the data, the experiments, or the outcomes. HPC ODA Commons fixes this by establishing community-governed contracts - versioned schemas, canonical artifacts, and benchmark recipes - that make ODA workflows discoverable, reproducible, and comparable. It pairs these standards with a practical, CLI-first toolkit that lets operators and researchers go from raw logs to standardized results without sending data off-cluster.

Menear, Kevin [National Laboratory of the Rockies ↗

Building the I (Interoperability) of FAIR for performance reproducibility of large-scale composable workflows in RECUP

Abstract-Scientific computing communities increasingly run their experiments using complex data- and compute-intensive workflows that utilize distributed and heterogeneous architectures targeting numerical simulations and machine learning, often executed on the Department of Energy Leadership Computing Facilities (LCFs). We argue that a principled, systematic approach to implementing FAIR principles at scale, including fine-grained metadata extraction and organization, can help with the numerous challenges to performance reproducibility posed by such workflows. We extract workflow patterns, propose a set of tools to manage the entire life cycle of performance metadata, and aggregate them in an HPC-ready framework for reproducibility (RECUP). We describe the challenges in making these tools interoperable, preliminary work, and lessons learned from this experiment.

97 MATHEMATICS AND COMPUTING↗

Auto-HPCnet: An Automatic Framework to Build Neural Network-based Surrogate for High-Performance Computing Applications

High-performance computing communities are increasingly adopt- ing Neural Networks (NN) as surrogate models in their applications to generate scientific insights. Replacing an execution phase in the application with NN models can bring significant performance im- provement. However, there is a lack of tools that can help domain scientists automatically apply NN-based surrogate models to HPC applications. We introduce a framework, named Auto-HPCnet, to democratize the usage of NN-based surrogates. Auto-HPCnet is the first end-to-end framework that makes past proposals for the NN-based surrogate model practical and disciplined. Auto-HPCnet introduces a workflow to address unique challenges when apply- ing the approximation, such as feature acquisition and meeting the application-specific constraint on the quality of final computation outcome. We show that Auto-HPCnet can leverage NN for a set of HPC applications and achieve 5.50× speedup on average (up to 16.8× speedup and with data preparation cost included) while meeting the application-specific constraint on the final computation quality.

Dong, Wenqian↗

Scaling SQL to the Supercomputer for Interactive Analysis of Simulation Data

AI and simulation workloads consume and generate large amounts of data that need to be searched, transformed and merged with other data. With the goal of treating data as a first-class citizen inside a traditionally compute-centric HPC environment, we explore how the use of accelerators and high-speed interconnects can speed up tasks which otherwise constitute bottlenecks in computational discovery workflows. BlazingSQL is SQL engine that runs natively on NVIDIA GPUs and supports internode communication for fast analytics on terabyte-scale tabular data sets. We show how a fast interconnect improves query performance if leveraged through the Unified Communication X (UCX) middleware. We envision that future computing platforms will integrate accelerated database query capabilities for immediate and interactive analysis of large simulation data.

Glaser, Jens↗

A Unifying Framework to Enable Artificial Intelligence in High-Performance Computing Workflows

Current trends point to a future where large-scale scientific applications are tightly coupled high-performance computing/artificial intelligence (HPC/AI) hybrids. Hence, we urgently need to invest in creating a seamless, scalable framework where HPC and AI/machine learning can efficiently work together and adapt to novel hardware and vendor libraries without starting from scratch every few years. Finally, the current ecosystem and sparsely connected community are not sufficient to tackle these challenges, and we require a breakthrough catalyst for science similar to what PyTorch enabled for AI.

high-performance computing↗

Data Science and Computation for Rapid and Dynamic Compression Experiment Workflows at Experimental Facilities, September 8-11, 2020. Workshop Report

The application of high pressure to materials has enabled discoveries in scientific fields such as planetary science, materials science, and materials synthesis. Recent advances in X-ray user light sources and other facilities, co-location and integration of user facilities with high-pressure drivers, availability of high-performance computing (HPC) platforms, and the development of new data science techniques have created opportunities for, and challenges in, advancing data analytics for rapid and dynamic compression experiments. To address these challenges, harness the emerging technology now available, and expedite scientific discovery, Los Alamos National Laboratory (LANL) hosted a virtual workshop entitled “Data Science and Computation for Rapid and Dynamic Compression Workflows at Experimental Facilities” from September 8 to 11, 2020. The workshop included 95 registered scientists and analytics experts from 15 universities, 9 United States (US) national laboratories, 5 US and European X-ray light sources, neutron sources such as the Los Alamos Neutron Science Center (LANSCE), other big science facilities such as the National Ignition Facility (NIF), and an industry representative. The workshop included 31 invited talks and 4 lightning talks by students and postdocs.

36 MATERIALS SCIENCE↗

Understanding and Leveraging the I/O Patterns of Emerging Machine Learning Analytics

The scientific community is currently experiencing unprecedented amounts of data generated by cutting-edge science facilities. Soon facilities will be producing up to 1 PB/s which will force scientist to use more autonomous techniques to learn from the data. The adoption of machine learning methods, like deep learning techniques, in large-scale workflows comes with a shift in the workflow’s computational and I/O patterns. These changes often include iterative processes and model architecture searches, in which datasets are analyzed multiple times in different formats with different model configurations in order to find accurate, reliable and efficient learning models. This shift in behavior brings changes in I/O patterns at the application level as well at the system level. These changes also bring new challenges for the HPC I/O teams, since these patterns contain more complex I/O workloads. In this paper we discuss the I/O patterns experienced by emerging analytical codes that rely on machine learning algorithms and highlight the challenges in designing efficient I/O transfers for such workflows. We comment on how to leverage the data access patterns in order to fetch in a more efficient way the required input data in the format and order given by the needs of the application and how to optimize the data path between collaborative processes. We will motivate our work and show performance gains with a study case of medical applications.

Gainaru, Ana↗

Autonomy Loops for Monitoring, Operational Data Analytics, Feedback, and Response in HPC Operations

Many High Performance Computing (HPC) facilities have developed and deployed frameworks in support of continuous monitoring and operational data analytics (MODA) to help improve efficiency and throughput. Because of the complexity and scale of systems and workflows and the need for low-latency response to address dynamic circumstances, automated feedback and response have the potential to be more effective than current human-in-the-loop approaches which are laborious and error prone. Progress has been limited, however, by factors such as the lack of infrastructure and feedback hooks, and successful deployment is often site- and case-specific. In this position paper we report on the outcomes and plans from a recent Dagstuhl Seminar, seeking to carve a path for community progress in the development of autonomous feedback loops for MODA, based on the established formalism of similar (MAPE-K) loops in autonomous computing and self-adaptive systems. By defining and developing such loops for significant cases experienced across HPC sites, we seek to extract commonalities and develop conventions that will facilitate interoperability and interchangeability with system hardware, software, and applications across different sites, and will motivate vendors and others to provide telemetry interfaces and feedback hooks to enable community development and pervasive deployment of MODA autonomy loops.

autonomy loops↗

MARS: Malleable Actor-Critic Reinforcement Learning Scheduler

In this paper, we introduce MARS, a new scheduling system for HPC-cloud infrastructures based on a cost-aware, flexible reinforcement learning approach, which serves as an intermediate layer for next generation HPC-cloud resource manager. MARS ensembles the pre-trained models from heuristic workloads and decides on the most cost-effective strategy for optimization. A whole workflow application would be split into several optimizable dependent sub-tasks, then based on the pre- defined resource management plan, a reward will be generated after executing a scheduled task. Lastly, MARS updates the Deep Neural Network (DNN) model based on the reward. MARS is designed to optimize the existing models through reinforcement mechanisms. MARS adapts to the dynamics of workflow applications, selects the most cost-effective scheduling solution among pre-built scheduling strategies (backfilling, SJF, etc.) and self- learning deep neural network model at run-time. We evaluate MARS with different real-world workflow traces. MARS can achieve 5%-60% increased performance compare to state-of-the- art approaches.

Baheri, Betis↗

Open Reproducible Electron Microscopy Data Analysis

Electron microscopy (EM) is a cornerstone technique in the materials and biological sciences capable of imaging structures at nano- to atomic-scale resolution. Advances in technologies mean that one acquires datasets at increasing data rates and sizes. These advancements present enormous opportunities for researchers to understand complex systems. However, processing the resulting large-scale, complex data in a reproducible and shareable way is a real challenge for researchers. The building, managing, and maintaining complex workflows in a reproducible manner requires extensive knowledge in several areas outside the researchers’ core skill sets, such as software engineering, data science, and high-performance computing (HPC). Our work demonstrates an innovative approach to solving these problems, enabling reproducible EM data analysis through container encapsulated pipelines. Using modern container technologies, we encapsulate processing elements and connect them using shared memory. We expose user-friendly, advanced algorithms and tools to allow end users to utilize without expert programming skills. The platform enables reproducible, scalable, shareable pipelines for the analysis and visualization of EM data. Focusing on interoperability, we leverage the DOE and other agencies’ existing investments to provide a powerful software platform for EM data analysis.

Harris, Christopher↗

Data readiness pipeline patterns for scientific AI at scale: Insights from climate, fusion, life sciences, and materials

This article examines how data readiness for AI principles apply to large scientific datasets used to train foundation models. We analyze archetypal workflows across four representative domains—climate, nuclear fusion, life sciences, and materials—to identify common preprocessing patterns and domain‐specific constraints. We introduce a two‐dimensional readiness model that combines canonical preprocessing patterns with a five‐level operational readiness scale, both tailored to high‐performance computing (HPC) environments. This construct helps outline key challenges in transforming large‐scale scientific data into formats suitable for scalable AI training. Together, these dimensions form a conceptual maturity matrix that characterizes scientific data readiness and guides infrastructure development toward standardized, cross‐domain support for scalable and reproducible AI for science. Finally, we evaluate this maturity matrix in the context of case studies including ClimaX (climate), AFLOW (materials), OpenFold (proteomics), and DIII‐D fusion disruption‐prediction workflows, from which we distill lessons learned and provide recommendations to guide practitioners in developing robust AI‐readiness pipelines. Finally, we discuss remaining cross‐cutting challenges that persist across scientific domains.

97 MATHEMATICS AND COMPUTING↗

VA EDH Advanced Software Pipeline Framework Report: Enhancing Automation and Scalability

The VA Environmental Determinants of Health (EDH) Advanced Software Pipeline Framework is designed to enhance the efficiency, scalability, and security of geospatial data processing workflows. This framework integrates modern data orchestration and containerization technologies, including Prefect for workflow automation, Docker for containerization, and PostgreSQL/PostGIS for geospatial data storage and analysis. It ensures standardized, reproducible, and automated data processing, supporting VA objectives related to substance use risk assessment and recovery research. The pipeline addresses key scalability and performance challenges through horizontal and vertical scaling, high-performance computing (HPC) integration, parallel processing, task caching, and dynamic resource allocation. These optimizations improve throughput and reduce latency, allowing the system to efficiently manage large and complex datasets. Additionally, security and compliance measures—such as data encryption (SSL), Role-Based Access Control (RBAC), and adherence to GDPR and HIPAA standards—safeguard sensitive information throughout data transmission and storage. A key implementation of this framework includes the automation of shelter list geolocation workflows, ensuring that up-to-date data is readily available for VA decision-making. Lessons learned from this project include the transition from in-memory processing to incremental storage writes, improving resource management and reliability. Future enhancements aim to expand automation, integrate AI-driven anomaly detection, and incorporate high-performance computing resources. This framework provides a scalable, secure, and adaptable solution for managing geospatial datasets, reinforcing the VA’s ability to support clinical and strategic initiatives through data-driven decision-making.

97 MATHEMATICS AND COMPUTING↗

Accelerating discoveries at DIII-D with the Integrated Research Infrastructure

DIII-D research is being accelerated by leveraging high performance computing (HPC) and data resources available through the National Energy Research Scientific Computing Center (NERSC) Superfacility initiative. As part of this initiative, a high-resolution, fully automated, whole discharge kinetic equilibrium reconstruction workflow was developed that runs at the NERSC for most DIII-D shots in under 20 min. This has eliminated a long-standing research barrier and opened the door to more sophisticated analyses, including plasma transport and stability. These capabilities would benefit from being automated and executed within the larger Department of Energy Advanced Scientific Computing Research program’s Integrated Research Infrastructure (IRI) framework. The goal of IRI is to empower researchers to meld DOE’s world-class research tools, infrastructure, and user facilities seamlessly and securely in novel ways to radically accelerate discovery and innovation. For transport, we are looking at producing flux matched profiles and also using particle tracing to predict fast ion heat deposition from neutral beam injection before a shot takes place. Our starting point for evaluating plasma stability focuses on the pedestal limits that must be navigated to achieve better confinement. This information is meant to help operators run more effective experiments, so it needs to be available rapidly inside the DIII-D control room. So far this has been achieved by ensuring the data is available with existing tools, but as more novel results are produced new visualization tools must be developed. In addition, all of the high-quality data we have generated has been collected into databases that can unlock even deeper insights. This has already been leveraged for model and code validation studies as well as for developing AI/ML surrogates. The workflows developed for this project are intended to serve as prototypes that can be replicated on other experiments and can be run to provide timely and essential information for ITER, as well as next stage fusion power plants.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

EQSIM—A multidisciplinary framework for fault-to-structure earthquake simulations on exascale computers part I: Computational models and workflow

Computational simulations have become central to the seismic analysis and design of major infrastructure over the past several decades. Most major structures are now “proof tested” virtually through representative simulations of earthquake-induced response. More recently, with the advancement of high-performance computing (HPC) platforms and the associated massively parallel computational ecosystems, simulation is beginning to play a role in increased understanding and prediction of ground motions for earthquake hazard assessments. However, the computational requirements for regional-scale geophysics-based ground motion simulations are extreme, which has restricted the frequency resolution of direct simulations and limited the ability to perform the large number of simulations required to numerically explore the problem parametric space. In this article, recent developments toward an integrated, multidisciplinary earth science-engineering computational framework for the regional-scale simulation of both ground motions and resulting structural response are described with a particular emphasis on advancing simulations to frequencies relevant to engineered systems. This multidisciplinary computational development is being carried out as part of the US Department of Energy (DOE) Exascale Computing Project with the goal of achieving a computational framework poised to exploit emerging DOE exaflop computer platforms scheduled for the 2022–2023 timeframe.

58 GEOSCIENCES↗

Accelerating lattice gauge theory studies with Agentic AI

Lattice gauge theory research, with its computationally intensive simulations and complex multi‑stage workflows, is well positioned to benefit from agentic AI systems. We demonstrate how such tools can support key components of lattice gauge theory research, including novel simulation code development using standard LQCD frameworks, HPC job orchestration, simulation data analysis, and expert‑guided tuning of algorithmic parameters such as Hasenbusch mass preconditioning and multigrid solvers. Our results show that agentic AI can reduce manual effort, improve productivity, and accelerate the research cycle while maintaining essential human oversight.

Ayyar, Venkitesh [Fermilab]↗