Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “HPC training”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Towards Low-Overhead Resilience for Data Parallel Deep Learning

Data parallel techniques have been widely adopted both in academia and industry as a tool to enable scalable training of deep learning models. At scale, DL training jobs can fail due to software or hardware bugs, may need to be preempted or terminated due to unexpected events, or may perform suboptimally because they were misconfigured. Under such circumstances, there is a need to recover and/or reconfigure data-parallel DL training jobs on-the-fly, while minimizing the impact on the accuracy of the DNN model and the runtime overhead. In this regard, state-of-art techniques adopted by the HPC community mostly rely on checkpoint-restart, which inevitably leads to loss of progress, thus increasing the runtime overhead. In this paper we explore alternative techniques that exploit the properties of modern deep learning frameworks (overlapping of gradient averaging and weight updates with local gradient computations through pipeline parallelism) to reduce the overhead of resilience/elasticity. To this end we introduce a failure simulation framework and two resilience strategies (immediate mini-batch rollback and lossy forward recovery), which we study compared with checkpoint-restart approaches in a variety of settings in order to understand the trade-offs between the accuracy loss of the DNN model and the runtime overhead.

data-parallel training↗

Can the United States Maintain Its Leadership in High-Performance Computing? - A report from the ASCAC Subcommittee on American Competitiveness and Innovation to the ASCR Office

The United States (U.S.) is no longer the unambiguous leader in the vitally important field of high performance computing (HPC). Japan, the European Union (EU), and China have fielded systems that are on par with our fastest supercomputers. The supply chain for everything from semiconductors to scientific software is globally distributed. Yet our economic future and security depend critically on our ability to innovate faster than our competitors, and the speed of innovation depends increasingly on large-scale computational science and engineering and thus HPC. How should the United States respond to this challenge? This report seeks to initiate a new and potentially transformative national discussion on this vital question. The Department of Energy’s (DOE) Advanced Scientific Computing Research (ASCR) program is well-positioned to make informed, targeted decisions about where the United States should cooperate and where it should compete in the global market for scientific exploration and discovery. By setting its sights on problems critical to our nation and the world, by establishing productive new collaborations, and by making strategic investments, ASCR can restore and maintain U.S. scientific leadership in the critical areas described in this report while strengthening our research infrastructure and training a large, diverse cohort of scientists. In doing so, ASCR and its scientists will pave the way for a secure and prosperous future for America. For more than 30 years, the ASCR program has provided the HPC and networking capabilities and expertise needed to support DOE’s mission to advance the national, economic, and energy security of the United States. The program now faces the challenge of developing and deploying the next generation of HPC systems and technologies, as well as supporting the application of HPC and artificial intelligence (AI) technologies to a wide range of scientific and engineering research problems. Through its research and development efforts, the ASCR program must also advance the state of the art in HPC and accelerate the pace of scientific discovery and technological innovation. Fulfilling this promise will require significantly increased investments, as well as innovative policies and programs. This subcommittee is aware that we are making recommendations and calls for action at a time when federal resources are limited. We understand that a wide range of competing priorities must be balanced by the nation’s leaders and that there is a need to leverage resources in new ways and seek efficiencies in facilities and operations. However, we must not let these realities limit our imagination or silence our advocacy. The ASCR program is a key part of the U.S. research infrastructure and an important component of economic growth and U.S. competitiveness. ASCR has a responsibility to pursue its mission, including advanced scientific computing, applications of AI technologies, and the required advanced research facilities, with determination and enthusiasm. To fulfill the scientific enterprise’s responsibility to the nation, the ASCR program must not only develop and publish a clear vision with an associated list of goals, priorities, and recommendations but also demonstrate scientific leadership by consistently securing long-term funding. This will allow the program to build on its achievements to date, to realize its ambitious vision, and to make lasting contributions to the field.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

AI Applications to Physics Experiments at Jefferson Lab

We survey how AI/ML is being deployed across Jefferson Lab's experimental and accelerator programs. In EPSCI, Hydra applies computer vision to automate real-time data-quality monitoring across all four experimental halls, replacing manual inspection of hundreds to thousands of histograms per shift. AIEC (AI Experiment Controls) uses ML to stabilize drift chamber gains and is now part of standard CEBAF production running, while AI Optimized Polarization (AIOP) targets autonomous control of polarized targets and photon beam angular alignment. In CASA, cavity fault classification models identify faulted cavities and trip types from waveform data with ~85% and ~78% agreement to labeled data, respectively, and are deployed in production; a separate effort applies LLMs and hybrid search to make the CEBAF operations logbook AI-ready. QCD-focused work includes transformer- and GAN-based generative models for particle-level event simulation, with distributed GAN training scaling studies on Polaris. Additional efforts span ML-on-FPGA for the EIC and a new Data Science Department coordinating anomaly detection, uncertainty quantification, and HPC-scalable ML lab-wide. Collectively, these projects illustrate AI's growing role in improving efficiency across JLab's nuclear physics mission.

Mei, Xinxin [Thomas Jefferson National Accelerator↗

OmniFed: A Modular Framework for Configurable Federated Learning from Edge to HPC

Federated Learning (FL) is critical for edge and High Performance Computing (HPC) where data is not centralized and privacy is crucial. We present OmniFed, a modular framework designed around decoupling and clear separation of concerns for configuration, orchestration, communication, and training logic. Its architecture supports configuration-driven prototyping and code-level override-what-you-need customization. We also support different topologies, mixed communication protocols within a single deployment, and popular training algorithms. It also offers optional privacy mechanisms including Differential Privacy (DP), Homomorphic Encryption (HE), and Secure Aggregation (SA), as well as compression strategies. These capabilities are exposed through well-defined extension points, allowing users to customize topology and orchestration, learning logic, and privacy/compression plugins, all while preserving the integrity of the core system. We evaluate multiple models and algorithms to measure various performance metrics. By unifying topology configuration, mixed-protocol communication, and pluggable modules in one stack, OmniFed streamlines FL deployment across heterogeneous environments. Github repository is available at https://github.com/at-aaims/OmniFed.

Tyagi, Sahil [ORNL] (ORCID:0009000783144745)↗

Teaching Software Sustainability for High Performance Computing at ATPESC

The Argonne Training Program in Extreme Scale Computing (ATPESC) was started by Argonne National Laboratory with the objective of expanding the ranks of better-prepared users of high-performance computing (HPC) machines. One of the unique aspects of the program was inclusion of a track on software engineering and community codes. The inclusion was motivated by the observation that the projects with good software processes were better able to meet their scientific goals. Over the years, with greater awareness of software sustainability issues in the community, the track has evolved into a software productivity and sustainability track. In this paper we present our experience in choosing and disseminating the content related to the topic of software engineering in high performance computing science from the beginning of the program until now. We discuss the motivations and the reception of the tracks. We also document the evolution of the track over the years based on student feedback and also the growth of awareness about software productivity in high performance computing.

Dubey, Anshu↗

ChatPORT: Fine-Tuned LLM for Easy Code {PORT}ing

Fine-tuning existing LLMs for specialized tasks has become a very attractive alternative due to its low cost and quick development cycle. With many pre-trained LLMs available, it is an increasingly complex task to choose the correct model as the starting point or base model. In this work we discuss ChatPORT - a specialized fine-tuned LLM geared towards providing correctly translated codes from one programming model to another. We evaluate a number of base models and compare and contrast their features and characteristics that make them a viable starting point. In this paper, we focus on the OpenMP offload porting capabilities of ChatPORT. We build our training data using kernels from the Heterogeneous Computing Benchmarks (HeCBench) [12] and the OpenMP Validation and Verification suite [5] to fine-tune the base models. We then test the model using unseen kernels extracted from the HeCBench benchmark suite. Our results show that: (1) not all open LLMs geared towards HPC are aware of programming models like OpenMP, (2) although all base models benefit from fine-tuning they learn differently and produce different correctness rates, (3) depending on the memory size and compute resource available, different base models can be used for fine-tuning without significantly affecting the quality of transpiled code they generate, (4) fine-tuning improved the correctness rate of the LLM by an average of 43.2%, and (5) feedback-based training data further increased the correctness rate by an average of 6% over the LLMs tested.

Pophale, Swaroop [ORNL] (ORCID:0000000185446367)↗

XPlacer Machine Learning Guided Data Placement

XPlacer Machine Learning Guided Data Placement uses machine learning model to assist programmers to select the data placement advises for application running on HPC system with Nvidia GPUs. XPlacer Machine Learning Guided Data Placement consists three components: (1) 4 revised benchmarks from Rodinia benchmarks; (2) compiler script and execution script to generate variants, and collect training data for the machine learning model; (3) the scripts to post-process the collected data and generate the machine leaning model.

Liao, Chunhua↗

Scalable Deep-Learning-Accelerated Topology Optimization for Additively Manufactured Materials

Topology optimization (TO) is a popular and powerful computational approach for designing novel structures, materials, and devices. Two computational challenges have limited the applicability of TO to a variety of industrial applications. First, a TO problem often involves a large number of design variables to guarantee sufficient expressive power. Second, many TO problems require a large number of expensive physical model simulations, and those simulations cannot be parallelized. To address these issues, we propose a general scalable deep-learning (DL) based TO framework, referred to as SDL-TO, which utilizes parallel schemes in high performance computing (HPC) to accelerate the TO process for designing additively manufactured (AM) materials. Unlike the existing studies of DL for TO, our framework accelerates TO by learning the iterative history data and simultaneously training on the mapping between the given design and its gradient. The surrogate gradient is learned by utilizing parallel computing on multiple CPUs incorporated with a distributed DL training on multiple GPUs. The learned TO gradient enables a fast online update scheme instead of an expensive update based on the physical simulator or solver. Using a local sampling strategy, we achieve to reduce the intrinsic high dimensionality of the design space and improve the training accuracy and the scalability of the SDL-TO framework. The method is demonstrated by benchmark examples and AM materials design for heat conduction. The proposed SDL-TO framework shows competitive performance compared to the baseline methods but significantly reduces the computational cost by a speed up of around 8.6x over the standard TO implementation.

Bi, Sirui↗

BULKI-Store v0.3.2

BULKI-Store is a distributed object storage system optimized for high-performance computing environments. Built with a Rust core and Python bindings, it efficiently manages scientific and machine learning datasets across HPC clusters. The system employs a client-server architecture with MPI integration, enabling seamless scaling on supercomputers like Perlmutter. BULKI-Store's object-oriented approach provides intuitive data organization with rich metadata support, contrasting with traditional file-based solutions. Key optimizations include selective checkpoint loading, unified checkpoint files, and object chunking for large data transfers. For machine learning workloads, BULKI-Store offers advantages through fine-grained access patterns, dynamic data sharing between training instances, and reduced memory pressure. Memory management features include strategic Python GC calls, minimized data copies, and batch processing capabilities. The system leverages Rayon's thread pool for asynchronous data prefetching and supports multiple CPU architectures (ARM64, x86, AMD, RISC-V). By combining performance optimizations with developer-friendly APIs, BULKI-Store addresses the complex data management challenges of modern HPC applications while maintaining compatibility across heterogeneous computing environments.

Zhang, Wei [Lawrence Berkeley National Laboratory ↗

An Extensible High Energy Density Modeling Tool for Extreme Regimes

The overreaching goal of the project is to develop a simulation capability to model an entire High Energy Density (HED) target heated by an X-ray Free Electron Laser (XFEL) beam including the surrounding material with a focus on target material associated with subsequent shots. Our primary objective is to study target fratricide problems, where radiation/debris from the current shot degrade the target for subsequent shots. This work will have a large impact on research being conducted at HED facilities where the quality of data is directly related to the effective repetition rate of the driver, which is often limited by target fratricide problems. This project focuses on a stream of closely spaced droplet targets that are appropriate for new very high repetition rate sources such as the ~1 MHz LCLS-II upgraded XFEL. The capability of the PISALE (Pacific Island Structured-AMR with ALE) code is extended to model this class of problems. We summarize the capability and improvements made to the HED branch of the PISALE code. We discuss our simulations of target fratricide for small (5 micron) hydrogen droplets heated by an XFEL beam at SLAC. We also compare simulations of the initial dynamics of larger (50 micron) water droplets, heated by the same XFEL beam, with experimental data. We briefly discuss training, dissemination, and related activities associated with this project.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

HAPPA: A Modular Platform for HPC Application Resilience Analysis with LLMs Embedded

High-performance computing (HPC) systems are increasingly vulnerable to soft errors, which pose significant challenges in maintaining computational accuracy and reliability. Predicting the resilience of HPC applications to these errors is crucial for robust code protection and detailed resilience analysis. In this study, we present HAppA, a modular platform designed for HPC Application Resilience Analysis. Embedding Large Language Models (LLMs), HAppA addresses understanding the context information of long code sequences typical in HPC applications. HAppA implements a novel code representation module that chunks the code into fixed-size segments and aggregates the embeddings of these segments. Three aggregation methods have been explored: MeanPooling, MaxPooling, and LSTM-based techniques. We built a DAtaset for REsilience analysis using Fault Injection (FI), named DARE. Using our DARE dataset, HAppA is trained for regression prediction tasks. Our evaluation results demonstrate the predictive accuracy of HAppA compared to other models, particularly noting that the LSTM-based aggregation method -- HAppA-LSTM -- achieves a mean squared error (MSE) of 0.078 for SDC prediction, surpassing the existing state-of-the-art PARIS model, which recorded an MSE of 0.1172. Additionally, HAppA with the KeyBERT model extracts a list of keywords representing the source code. A comprehensive importance analysis of these keywords further elucidates the code patterns contributing to the error rate. These findings highlight the effectiveness of HAppA in analyzing the resilience of HPC applications and establish a new benchmark for predictive accuracy in resilience.

Jiang, Hailong [Kent State University]↗

Inference-Engine v0.1.0

Given a pre-trained neural network, Inference-Engine performs maps network inputs to outputs by executing the forward pass through the provided network. Although the predominant programming language for machine-learning is Python, most high-performance computing (HPC) applications are written in Fortran, C, or C++. Inference-Engine aims to support HPC programs and is written in Fortran, a language with a large feature set supporting interoperability with C. This software exposes concurrency in a portable way by using standard language features that some modern Fortran compilers can exploit with various optimizations, including offloading computation to a Graphics Processing Unit (GPU). In particular, this software makes extensive use of Fortran's "do concurrent" parallel loop construct, implicitly parallel array statements, and pure procedures that can be invoked inside "do concurrent" blocks. Inference-Engine also supports dynamic choice of inference methods at runtime. Two current options include one method that uses Fortran's "dot_product" intrinsic function inside "do concurrent" blocks and another method that instead uses Fortran' "matmul" array intrinsic function. We plan to investigate automatic compiler offloading of "do concurrent" calculations to GPUs and compile-time substitution of optimized libraries such as the Basic Linear Algebra Library (BLAS) for "matmul" invocations. We also envision the potential for the choice of which method to use could happen at program launch based on in situ performance measurements on any given platform.

Rouson, Damian↗

HPC Analytics of Fused Thermal Plants Data to Optimize Operating Envelope

In this project, ORNL extensively reviewed the ORAP RAM data, and it guided us to develop machine learning models that can predict time to next failures and forecast failure trends, which will be useful for optimizing power plant operation strategies. More specifically, we trained multiple random forest models and evaluated the model accuracy to validate with 10+ years of historical data. In addition, we implemented a web-based graphical user interface system for the models to show how our models can be used in more intuitive ways. This proof of concept allowed exploration of model use with power plant operators in mind. Developed machine learning models will be helpful for managing risks, planning maintenance and operation, ultimately reducing the down time and increasing the service hours. For future work, there are several interesting research topics including but not limited to model enhancement, creating synergy with traditional failure modeling approaches, and data-driven actionable recommendation and suggestions.

20 FOSSIL-FUELED POWER PLANTS↗

Design and implementation of I/O performance prediction scheme on HPC systems through large-scale log analysis

Abstract Large-scale high performance computing (HPC) systems typically consist of many thousands of CPUs and storage units used by hundreds to thousands of users simultaneously. Applications from large numbers of users have diverse characteristics, such as varying computation, communication, memory, and I/O intensity. A good understanding of the performance characteristics of each user application is important for job scheduling and resource provisioning. Among these performance characteristics, I/O performance is becoming increasingly important as data sizes rapidly increase and large-scale applications, such as simulation and model training, are widely adopted. However, predicting I/O performance is difficult because I/O systems are shared among all users and involve many layers of software and hardware stack, including the application, network interconnect, operating system, file system, and storage devices. Furthermore, updates to these layers and changes in system management policy can significantly alter the I/O behavior of applications and the entire system. To improve the prediction of the I/O performance on HPC systems, we propose integrating information from several different system logs and developing a regression-based approach to predict the I/O performance. Our proposed scheme can dynamically select the most relevant features from the log entries using various feature selection algorithms and scoring functions, and can automatically select the regression algorithm with the best accuracy for the prediction task. The evaluation results show that our proposed scheme can predict the write performance with up to 90% prediction accuracy and the read performance with up to 99% prediction accuracy using the real logs from the Cori supercomputer system at NERSC.

97 MATHEMATICS AND COMPUTING↗

High-Performance Deep Learning Toolbox for Genome-Scale Prediction of Protein Structure and Function

Computational biology is one of many scientific disciplines ripe for innovation and acceleration with the advent of high-performance computing (HPC). In recent years, the field of machine learning has also seen significant benefits from adopting HPC practices. In this work, we present a novel HPC pipeline that incorporates various machine-learning approaches for structure-based functional annotation of proteins on the scale of whole genomes. Our pipeline makes extensive use of deep learning and provides computational insights into best practices for training advanced deep-learning models for high-throughput data such as proteomics data. We showcase methodologies our pipeline currently supports and detail future tasks for our pipeline to envelop, including large-scale sequence comparison using SAdLSA and prediction of protein tertiary structures using AlphaFold2.

Gao, Mu↗

Argonne Leadership Computing Facility: 2021 Operational Assessment Report

This Operational Assessment Report describes how the Argonne Leadership Computing Facility (ALCF) met or exceeded every one of its goals for calendar year (CY) 2021 as an advanced scientific computing center. In CY 2021, the ALCF operated its production resource, Theta, an Intel-based Cray XC40 system (11.7-petaflops) augmented with 24 NVIDIA DGX A100-based nodes (3.9-petaflops) that supports diverse workloads, integrating data analytics with artificial intelligence (AI) training and learning in a single platform. In 2021, we began deploying Polaris, our newest 40- petaflops system, and augmented this powerful testbed system with an additional 28 nodes to support the integration of real-time experiments and HPC resources. We also deployed our two largest storage systems yet, named Grand and Eagle, that will bring new services to our users and will power data-driven research for years to come. Last year, Theta delivered a total of 20.8 million node-hours to 16 Innovative and Novel Computational Impact on Theory and Experiment (INCITE) projects and 7.2 million node-hours to ASCR Leadership Computing Challenge (ALCC) projects (32 awarded during the 2020–2021 ALCC year and 17 awarded during the 2021–2022 ALCC year), as well as substantial support to Director’s Discretionary (DD) projects (5.5 million node-hours). As Table ES.1 shows, Theta performed exceptionally well in terms of overall availability (95.1 percent), scheduled availability (99.4 percent), and utilization (98.1 percent; Table 2.1). As of the submission date of this document, ALCF’s user community has published 249 papers in high-quality, peer-reviewed journals and technical proceedings. At the 2021 International Conference for High Performance Computing, Networking, Storage and Analysis (SC’21), Argonne researchers won two HPCwire Readers’ Choice Awards and were part of a Gordon Bell Prize finalist team recognized for developing an AI-enabled, multi-resolution simulation framework for studying complex biomolecular machines. Their framework was used to observe the SARS-CoV-2 replication-transcription machinery in action, by directly integrating experimental data. ALCF also provided a comprehensive program of high-performance computing (HPC) support services to help our community make productive use of the facility’s diverse and growing collection of resources. We are now entering the exascale era, with exascale machines being planned for national laboratories across the country, including Aurora at Argonne National Laboratory (Argonne) in 2023. ALCF researchers have been leading and guiding numerous strategic activities that will push the boundaries of what’s possible in computational science and engineering and allow us to deliver science on day one.

97 MATHEMATICS AND COMPUTING↗

Impacts of floating-point non-associativity on reproducibility for HPC and deep learning applications

Run to run variability in parallel programs caused by floating-point non-associativity has been known to significantly affect reproducibility in iterative algorithms, due to accumulating errors. Non-reproducibility can critically affect the efficiency and effectiveness of correctness testing for stochastic programs. Recently, the sensitivity of deep learning training and inference pipelines to floating-point non-associativity has been found to sometimes be extreme. It can prevent certification for commercial applications, accurate assessment of robustness and sensitivity, and bug detection. New approaches in scientific computing applications have coupled deep learning models with high-performance computing, leading to an aggravation of debugging and testing challenges. Here we perform an investigation of the statistical properties of floating-point non-associativity within modern parallel programming models, and analyze performance and productivity impacts of replacing atomic operations with deterministic alternatives on GPUs. We examine the recently-added deterministic options in PyTorch within the context of GPU deployment for deep learning, uncovering and quantifying the impacts of input parameters triggering run to run variability and reporting on the reliability and completeness of the documentation. Finally, we evaluate the strategy of exploiting automatic determinism that could be provided by deterministic hardware, using the Groq LPUTM accelerator for inference portions of the deep learning pipeline. We demonstrate the benefits that a hardware-based strategy can provide within reproducibility and correctness efforts.

Shanmugavelu, Sanjif↗

The LCLStream Ecosystem for Multi-Institutional Dataset Exploration

We describe a new end-to-end experimental data streaming framework designed from the ground up to support new types of applications – AI training, extremely high-rate X-ray time-of-flight analysis, crystal structure determination with distributed processing, and custom data science applications and visualizers yet to be created. Throughout, we use design choices merging cloud microservices with traditional HPC batch execution models for security and flexibility. This project makes a unique contribution to the DOE Integrated Research Infrastructure (IRI) landscape. By creating a flexible, API-driven data request service, we address a significant need for high-speed data streaming sources for the X-ray science data analysis community. With the combination of data request API, mutual authentication web security framework, job queue system, high-rate data buffer, and complementary nature to facility infrastructure, the LCLStreamer framework has prototyped and implemented several new paradigms critical for future generation experiments.

Rogers, David [ORNL] (ORCID:0000000251871768)↗