Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “HPC training”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Toward an Autonomous Workflow for Single Crystal Neutron Diffraction

The operation of the neutron facility relies heavily on beamline scientists. Some experiments can take one or two days with experts making decisions along the way. Leveraging the computing power of HPC platforms and AI advances in image analyses, here we demonstrate an autonomous workflow for the single-crystal neutron diffraction experiments. The workflow consists of three components: an inference service that provides real-time AI segmentation on the image stream from the experiments conducted at the neutron facility, a continuous integration service that launches distributed training jobs on Summit to update the AI model on newly collected images, and a frontend web service to display the AI tagged images to the expert. Ultimately, the feedback can be directly fed to the equipment at the edge in deciding the next-step experiment without requiring an expert in the loop. With the analyses of the requirements and benchmarks of the performance for each component, this effort serves as the first step toward an autonomous workflow for real-time experiment steering at ORNL neutron facilities.

Yin, Junqi↗

High Flux Isotope Reactor Low Enriched Uranium U-10Mo Fuel Design Parameters

Activities to convert the HFIR from HEU to LEU are ongoing as part of the US Department of Energy (DOE) National Nuclear Security Administration (NNSA) nuclear nonproliferation mission. Design activities to study the conversion of HFIR from HEU to LEU fuel explored different fuel design features and shapes with a uranium-molybdenum (U-10Mo) monolithic alloy fuel. This high-density alloy contains 90 wt % uranium and 10 wt % molybdenum and has a uranium density of 15.318gU/cm 3 . The goal of these studies is to generate several candidate HFIR LEU fuel designs of varying fuel fabrication complexity that meet the current HEU performance metrics and safety requirements. Recent advancements in modeling and simulation tools and design methods enabled a thorough analysis of the available design space with U-10Mo fuel. A surrogate model used this analysis as training data to quickly determine the performance of a design given specific design parameters. An optimization module used this surrogate model to quickly search this multidimensional search space given specific desired performance characteristics. This approach was made possible by the large available design space with U-10Mo fuel. Shift, a Monte Carlo tool optimized for high-performance computing (HPC) architectures, was used for faster calculation and better data management for reactor physics simulations. Once most of these design studies were complete, a new suite called the Python HFIR Analysis and Measurement Engine (PHAME) was developed to connect all fuel design analysis steps, making design studies more efficient and reproducible. The post-processing capabilities of these new tools are leveraged for the information provided herein. Leveraging these tools, several candidate fuel designs were selected with varying levels of feature complexity and reactor performance. This report provides design feature details for four selected HFIR LEU U-10Mo fuel designs and their corresponding performance and safety metrics. Nominal best-estimate design parameters and irradiation conditions, including fission rate densities, power densities, heat fluxes, and cumulative fission densities, are provided. Simulations show that the high uranium density of U-10Mo fuel provides a large potential design space that enables various LEU designs to meet HEU core performance metrics and safety requirements with a power increase from 85 MW (HEU) to 95 MW or 100 MW (LEU).

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Machine Learning Atom Probe Tomography Tool For Automatic And Fast Clustering

The software uses a YOLO11 segmentation model trained on synthetic data to analyze APT datasets. The workflow operates as follows: 1. Data Slicing: The APT dataset is divided into multiple 2D cross-sections of a specified thickness. 2. Segmentation: The model identifies point-dense regions within each 2D slice. 3. 3D Reconstruction: Detected regions (masks) from all slices are combined and reconstructed back into the original 3D space, forming clusters. The integration with HPC resources enables the software to process large-scale APT datasets efficiently. This combination of automation and scalability reduces manual intervention, improves reproducibility, and accelerates the clustering workflow.

Tang, Yalei [Idaho National Laboratory (INL), Idah↗

2020 Budget Request for the DOE Computational Science Graduate Fellowship (CSGF) Grant

The Department of Energy Computational Science Graduate Fellowship (DOE CSGF) is essential for addressing the increasingly complex national workforce demands stemming from the growth of computational science and engineering challenges. Computational science and engineering (CSE) takes a multidisciplinary approach that utilizes scientific computing to tackle practical problems and provide technical tools across the spectrum of scientific discovery. The DOE CSGF specifically highlights high-performance computing (HPC) as a critical enabling technology in CSE, driving advancements in science and engineering that are vital to both the DOE and the broader economy. Over the past half-century, HPC has been an essential tool for DOE’s success. During this period, important missions, such as nuclear stockpile stewardship, have turned to HPC as an essential technology. Entire science disciplines have been transformed through the augmentation of scientific observation via HPC. At government laboratories, academic institutions, and in industry, DOE CSGF alumni are helping push traditional HPC boundaries while contributing to discoveries in high-energy physics, quantum information systems, fusion-reactor design, machine learning, additive manufacturing, nano materials for next-generation batteries and transistors, and advanced nuclear reactor modeling. In addition, HPC is used to address national health needs that will eventually point to cures both by helping cancer researchers manage and analyze huge troves of data, by simulating biological mechanisms, and by accelerating drug development. A 2023 report from the ASCAC Subcommittee on American Competitiveness and Innovation to the ASCR office, “Can the United States Maintain Its Leadership in High-Performance Computing?” says of the Program, “The CSGF program provides a barometer for disciplines that will be of interest to future DOE computing. Computational biology, machine learning, and quantum computing are among the subjects that began to swell in the ranks of CSGF applicants before the labs were hiring as high a percentage of employees in these categories.” The explosion of scientific and technological data has heightened the demand for advanced high-performance computing (HPC) to transform these data into meaningful scientific insights. As access to vast amounts of data increases, the fields of Machine Learning and Artificial Intelligence are experiencing a resurgence, enhancing the established practices of computational modeling and simulation. In its September 2020 subcommittee report on "AI/ML, Data Intensive Science, and High-Performance Computing," the DOE Advanced Scientific Computing Advisory Committee (ASCAC) specifically called for a fellowship program to train computational and data scientists to address exascale and data-intensive computing challenges. This integration of empirical and theoretical modeling will increasingly guide federal policymakers in making decisions that impact American society and future generations. It demands a workforce of highly skilled and intellectually agile computational scientists capable of navigating the rapid advancements in scientific computing within the DOE National Laboratory research environment. The DOE CSGF program has consistently addressed this critical need.

97 MATHEMATICS AND COMPUTING↗

In-Transit Data Transport Strategies for Coupled AI-Simulation Workflow Patterns

Coupled AI-Simulation workflows are becoming the major workloads for HPC facilities, and their increasing complexity necessitates new tools for performance analysis and prototyping of new in-situ workflows. We present SimAI-Bench, a tool designed to both prototype and evaluate these coupled workflows. In this paper, we use SimAI-Bench to benchmark the data transport performance of two common patterns on the Aurora supercomputer: a one-to-one workflow with co-located simulation and AI training instances, and a many-to-one workflow where a single AI model is trained from an ensemble of simulations. For the one-to-one pattern, our analysis shows that node-local and DragonHPC data staging strategies provide excellent performance compared Redis and Lustre file system. For the many-to-one pattern, we find that data transport becomes a dominant bottleneck as the ensemble size grows. Our evaluation reveals that file system is the optimal solution among the tested strategies for the many-to-one pattern.

Tummalapalli, Harikrishna [Argonne National Labora↗

Dataset of Generative AI Workload Power Profiles

This dataset provides a collection of high-resolution (5/10 Hz or every 0.2/0.1 seconds) power consumption profiles for generative artificial intelligence (GenAI) workloads executed on NLR's High Performance Computing (HPC) platform Kestrel. The dataset also includes examples of representative whole-facility power profiles generated using a bottom-up, event-driven, data center energy model . This dataset is designed to support research in energy modeling, infrastructure planning, energy system integration, and sustainability analysis for AI-driven computing systems. The dataset captures time-resolved electrical power measurements across a diverse set of configurations, including variations in job type (inference vs. training), workload (LLM vs. image generation), datasets, and number of compute nodes. Power traces are provided in a standardized format and include both raw/instantaneous and aggregated files. Each profile is accompanied by metadata describing workload parameters, enabling reproducibility and cross-study comparison. The dataset is intended for use in applications such as data center infrastructure planning, energy modeling, demand response and grid impact studies, and development and validation of system-level simulation tools. By making these workload-specific power profiles publicly available, this dataset aims to address the current lack of open, empirical energy data for generative AI systems and to facilitate transparent, reproducible research on the energy and environmental impacts of large-scale AI deployment. If you use this dataset, please cite the associated publication: Vercellino et al., “Measurement of Generative AI Workload Power Profiles for Whole-Facility Data Center Infrastructure Planning,” arXiv:2604.07345 (2026).

97 MATHEMATICS AND COMPUTING↗

Hiperclust

This software leverages transfer learning to analyze atom probe tomography (APT) data. It is trained on synthetic data and then applies this knowledge to predict the optimal number of clusters for a given APT dataset. Initially, the software used preliminary clustering to estimate the general structure of the data. Based on this, it provides suggestions for key parameters like minimum cluster size and minimum number of points. These parameters are critical for algorithms like HDBSCAN, ensuring accurate cluster formation without the need for trial-and-error testing. The software runs on High-Performance computing (HPC) systems, enabling fast, scalable analysis of large APT datasets, ultimately saving time and improving the reliability of clustering outcomes.

Tang, Yalei [Idaho National Laboratory (INL), Idah↗

miniGAN: A Generative Adversarial Network proxy application WBS 2.2.6.08 ECP-2.1.3 (Q3 FY2020 Milestone Report) (V.1.0)

In order to support the machine learning co-design needs of ECP applications in current and future DOE HPC hardware, we have developed a generative adversarial network (GAN) proxy application, miniGAN, that has been released through the ECP proxy application suite. The proxy application is representative of the needs of ExaLearn's target applications, specifically the Cosmoflow and ExaGAN cosmology applications and the ExaWind energy application. The proxy application also demonstrates the first use of performance portable kernels within widely-used machine learning frameworks: PyTorch (Facebook) and Horovod (Uber). We provide performance scaling results for similar workloads to ExaGAN and a profile of individual GAN training components.

97 MATHEMATICS AND COMPUTING↗

Accelerating Collective Communication in Data Parallel Training across Deep Learning Frameworks

This work develops new techniques within Horovod, a generic communication library supporting data parallel training across deep learning frameworks. In particular, we improve the Horovod control plane by implementing a new coordination scheme that takes advantage of the characteristics of the typical data parallel training paradigm, namely the repeated execution of collectives on the gradients of a fixed set of tensors. Using a caching strategy, we execute Horovod’s existing coordinator-worker logic only once during a typical training run, replacing it with a more efficient decentralized orchestration strategy using the cached data and a global intersection of a bitvector for the remaining training duration. Next, we introduce a feature for end users to explicitly group collective operations, enabling finer grained control over the communication buffer sizes. To evaluate our proposed strategies, we conduct experiments on a world-class supercomputer — Summit. We compare our proposals to Horovod’s original design and observe 2x performance improvement at a scale of 6000 GPUs; we also compare them against tf.distribute and torch.DDP and achieve 12% better and comparable performance, respectively, using up to 1536 GPUs; we compare our solution against BytePS in typical HPC settings and achieve about 20% better performance on a scale of 768 GPUs. Finally, we test our strategies on a scientific application (STEMDL) using up to 27,600 GPUs (the entire Summit) and show that we achieve a near-linear scaling of 0.93 with a sustained performance of 1.54 exaflops (with standard error +- 0.02) in FP16 precision.

Romero, Joshua↗

Concurrent Relaxation through Accelerated Deep Learning

CRADL captures performance metrics of machine learning algorithms operating on mesh data from multiphysics codes This proxy application is a tool to explore scalability of inference on HPC platforms, and also gather performance metrics for inference on new machine learning specific hardware. CRADL is designed to give users as fine a control as possible over an inference simulation. Users may select the number of cycles, amount of data, and batch size to pass to the accelerator of choice. Additionally the user may select a number of performance optimization libraries and flags. CRADL comes packaged with a repository of anonymized multi-physics simulation data, as well as a pretrained model for inference. The code allows a user to load their own pre-trained model and data if they wish. The code can operate in multiple parallelization schemes, with performance enhancing options such as half-precision libraries, PyTorch benchmarking, and pinned memory with non-blocking data transfers.

Zieb, KristoferJ.↗

White Box Access to Quantum Testbeds for Co-Design

At Lawrence Livermore National Laboratory (LLNL), we operate and maintain the Quantum Device and Integration Testbed (QuDIT) facility, a small quantum testbed that supports about 10 active research teams (including our own) and over 50 internal and external collaborators. This testbed is designed to give remote white box access to users for research, training, and outreach. A guiding principle behind the development of our testbed infrastructure, software and user interfaces is to empower users to perform experiments at the cutting edge of quantum information science at any level of abstraction, from materials studies, device physics and control and characterization techniques to algorithm development and quantum operating system design. Our testbed targets a multilevel quantum system (qudit) to expand the accessible Hilbert space of a simple-to-manufacture quantum device and focuses on quantum simulation, typically implemented through custom gates designed with quantum optimal control methods, rather than on a universal computing framework with a fixed gate set. We leverage the Lab’s high-performance computing (HPC) program and related expertise to simulate quantum systems, develop hybrid algorithms, and generate gates optimized for given simulations. Additionally, we have adopted a co-design philosophy from the HPC community in designing new hardware, so that the systems we develop are optimized for the specific physics simulations we plan to use them for.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Streaming Data in HPC Workflows Using ADIOS

The “IO Wall” problem, in which the gap between computation rate and data access rate grows continuously, poses significant problems to scientific workflows which have traditionally relied upon using the filesystem for intermediate storage between workflow stages. One way to avoid this problem in scientific workflows is to stream data directly from producers to consumers and avoiding storage entirely. However, the manner in which this is accomplished is key to both performance and usability. This paper presents the Sustainable Staging Transport, an approach which allows direct streaming between traditional file writers and readers with few application changes. SST is an ADIOS “engine”, accessible via standard ADIOS APIs, and because ADIOS allows engines to be chosen at run-time, many existing file-oriented ADIOS workflows can utilize SST for direct application-to-application communication without any source code changes. This paper describes the design of SST and presents performance results from various applications that use SST, for feeding model training with simulation data with substantially higher bandwidth than the theoretical limits of Frontier’s file system, for strong coupling of separately developed applications for multiphysics multiscale simulation, or for in situ analysis and visualization of data to complete all data processing shortly after the simulation finishes.

Podhorszki, Norbert [ORNL] (ORCID:000000019647542X↗

Energy–Performance Trade-offs in Privacy-Preserving Federated Learning on SmartNIC-Enabled HPC Systems

Federated learning (FL) is increasingly deployed on accelerator-rich high-performance computing (HPC) systems, yet the system-level energy cost of privacy-aware FL remains poorly understood, particularly across heterogeneous networking and server-placement options. We present a measurement-driven study of energy–performance trade-offs for FL on GH200-class nodes across three deployment configurations: CPU-Ethernet, CPU-InfiniBand (RDMA-capable), and a DPU-hosted FL server over InfiniBand using a BlueField-3 SmartNIC/DPU. Using NVIDIA FLARE (NVFLARE), we align node-level power telemetry with per-round timing extracted from NVFLARE logs to quantify time-to-solution (TTS), energy-to-solution (ETS), energy-delay product (EDP), and synchronization behavior for three transformer models (ALBERT, DistilBERT, BERT), trained with and without differential privacy (DP). We find that interconnect choice is the dominant driver of runtime and energy: host-managed InfiniBand consistently reduces communication overhead versus Ethernet, yielding lower TTS/ETS/EDP. In contrast, in our NVFLARE deployment, placing the FL server on the DPU does not consistently match CPU-InfiniBand performance and can be slower—especially for larger models—highlighting that server placement alone is not sufficient to guarantee end-to-end gains. Finally, under our fixed-round protocol, DP increases per-round cost and runtime variance; ETS increases largely in proportion to TTS because average node power remains relatively stable across configurations.

Kotevska, Olivera [ORNL] (ORCID:0000000316772243)↗

Toward designing effective exascale scientific computing workflows: experiences and best practices

Many fields within scientific computing have embraced advances in big-data analysis and machine learning, which often requires the deployment of large, distributed and complicated workflows that may combine training neural networks, performing simulations, running inference, and performing database queries and data analysis in asynchronous, parallel and pipelined execution frameworks. Such a shift has brought into focus the need for scalable, efficient workflow management solutions with reproducibility, error and provenance handling, traceability, and checkpoint-restart capabilities, among other needs. Here, we discuss challenges and best-practices for deploying exascale-generation computational science workflows on resources at the Oak Ridge Leadership Computing Facility (OLCF). We present our experiences with large-scale deployment of distributed workflows on the Summit supercomputer, including for bioinformatics and computational biophysics, materials science, and deep learning model optimization. We also present problems and solutions created by working within a Python-centric software base on traditional HPC systems, and discuss steps that will be required before the convergence of HPC, AI, and data science can be fully realized. Our results point to a wealth of exciting new possibilities for harnessing this convergence to tackle new scientific challenges.

Coletti, Mark↗

Potential of the Julia Programming Language for High Energy Physics Computing

Research in high energy physics (HEP) requires huge amounts of computing and storage, putting strong constraints on the code speed and resource usage. To meet these requirements, a compiled high-performance language is typically used; while for physicists, who focus on the application when developing the code, better research productivity pleads for a high-level programming language. A popular approach consists of combining Python, used for the high-level interface, and C++, used for the computing intensive part of the code. A more convenient and efficient approach would be to use a language that provides both high-level programming and high-performance. The Julia programming language, developed at MIT especially to allow the use of a single language in research activities, has followed this path. In this paper the applicability of using the Julia language for HEP research is explored, covering the different aspects that are important for HEP code development: runtime performance, handling of large projects, interface with legacy code, distributed computing, training, and ease of programming. The study shows that the HEP community would benefit from a large scale adoption of this programming language. The HEP-specific foundation libraries that would need to be consolidated are identified.

97 MATHEMATICS AND COMPUTING↗

Experiences Readying Applications for Exascale

The advent of Exascale computing invites an assessment of existing best practices for developing application readiness on the world's largest supercomputers. This work details observations from the last four years in preparing scientific applications to run on the Oak Ridge Leadership Computing Facility's (OLCF) Frontier system. This paper addresses a range of topics in software including programmability, tuning, and portability considerations that are key to moving applications from existing systems to future installations. A set of representative workloads provides case studies for general system and software testing. We evaluate the use of early access systems for development across several generations of hardware. Finally, we discuss how best practices were identified and disseminated to the community through a wide range of activities including user-guides and trainings. We conclude with recommendations for ensuring application readiness on future leadership computing systems.

exascale↗

Opportunities and Challenges from Artificial Intelligence and Machine Learning for the Advancement of Science, Technology, and the Office of Science Missions

In February 2019, the President signed Executive Order 13859, Maintaining American Leadership in Artificial Intelligence. This order launched the American Artificial Intelligence Initiative, a concerted effort to promote and protect AI technology and innovation in the United States. The Initiative implements a government-wide strategy in collaboration and engagement with the private sector, academia, the public, and like-minded international partners. Among other actions, key directives in the Initiative called for Federal agencies to: Prioritize AI research and development investments, Enhance access to high-quality cyberinfrastructure and data, Ensure that the US maintains an international leadership role in the development of technical standards for AI, and Provide education and training opportunities to prepare the American workforce for the new era of AI. The mission of the Department of Energy (DOE) is to ensure America’s security and prosperity by addressing its energy, environmental, and nuclear challenges through transformative science and technology solutions. In terms of Science and Innovation, the DOE’s mission is to maintain a vibrant US effort in science and engineering as a cornerstone of our economic prosperity with clear leadership in strategic areas. From July to October in 2019, the Argonne, Oak Ridge, and Berkeley National Laboratories hosted a series of four AI for Science Town Hall meetings in Chicago, Oak Ridge, Berkeley, and Washington DC. The four meetings were attended by over 1300 scientists from the 17 DOE Labs, 39 companies, and over 90 universities. The goal of the Town Hall series was ‘to examine scientific opportunities in the areas of artificial intelligence, Big Data, and high-performance computing (HPC) in the next decade, and to capture the big ideas, grand challenges, and next steps to realizing these.’ The discussions at the meetings were captured in the final report of the AI for Science Town Hall meetings.

42 ENGINEERING↗

FAIR Surrogate Benchmarks Supporting AI and Simulation Research (Final Report)

Computational Science is being revolutionized by integrating AI and simulation and, in particular, by deep learning surrogate models that can replace all or part of traditional large‐scale HPC computations. Such surrogates can achieve remarkable performance improvements, as much as several orders of magnitude, and save both compute time and energy. The Surrogate Benchmark Initiative (SBI) project creates a community repository and FAIR (Findable, Accessible, Interoperable, and Reusable) data ecosystem for HPC application surrogate benchmarks. The SBI team comes from Argonne National Laboratory (ANL), Indiana University (IU), Rutgers University, the University of Tennessee, Knoxville (UTK), and the University of Virginia (UVA). SBI repositories include data, code, and all relevant collateral artifacts that the science and engineering community need to use and reuse these data sets and surrogates. SBI repositories generate active research from both the participants in SBI and the broad community of AI and domain scientists. This project develops surrogates that use several different neural nets to learn and quickly infer the results of simulations and data systems and captures them as surrogate benchmarks with a rich set of metadata covering: Data; Model; Metrics specification; Machine specification; and Science, Speed, and Power Results. We research FAIR metadata for these benchmarks. We develop application surrogate examples as benchmarks across many fields (ANL, UTK, IU, UVA). We also study non-Surrogate benchmarks that have many common features and similar issues as regards FAIRness. We work with MLCommons (UVA, UTK), which is a major machine learning benchmarking activity where we get metadata ontologies, software, and benchmarks, Benchmarks have datasets, models, and metadata and they need a technical framework developed by UTK and Rutgers and deployed by UVA. We study features of Surrogates including performance, training set size, and uncertainty quantification (Rutgers, UVA and IU).

97 MATHEMATICS AND COMPUTING↗