Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “HPC training”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Inference-Engine v0.1.0

Given a pre-trained neural network, Inference-Engine performs maps network inputs to outputs by executing the forward pass through the provided network. Although the predominant programming language for machine-learning is Python, most high-performance computing (HPC) applications are written in Fortran, C, or C++. Inference-Engine aims to support HPC programs and is written in Fortran, a language with a large feature set supporting interoperability with C. This software exposes concurrency in a portable way by using standard language features that some modern Fortran compilers can exploit with various optimizations, including offloading computation to a Graphics Processing Unit (GPU). In particular, this software makes extensive use of Fortran's "do concurrent" parallel loop construct, implicitly parallel array statements, and pure procedures that can be invoked inside "do concurrent" blocks. Inference-Engine also supports dynamic choice of inference methods at runtime. Two current options include one method that uses Fortran's "dot_product" intrinsic function inside "do concurrent" blocks and another method that instead uses Fortran' "matmul" array intrinsic function. We plan to investigate automatic compiler offloading of "do concurrent" calculations to GPUs and compile-time substitution of optimized libraries such as the Basic Linear Algebra Library (BLAS) for "matmul" invocations. We also envision the potential for the choice of which method to use could happen at program launch based on in situ performance measurements on any given platform.

Rouson, Damian↗

HPC Analytics of Fused Thermal Plants Data to Optimize Operating Envelope

In this project, ORNL extensively reviewed the ORAP RAM data, and it guided us to develop machine learning models that can predict time to next failures and forecast failure trends, which will be useful for optimizing power plant operation strategies. More specifically, we trained multiple random forest models and evaluated the model accuracy to validate with 10+ years of historical data. In addition, we implemented a web-based graphical user interface system for the models to show how our models can be used in more intuitive ways. This proof of concept allowed exploration of model use with power plant operators in mind. Developed machine learning models will be helpful for managing risks, planning maintenance and operation, ultimately reducing the down time and increasing the service hours. For future work, there are several interesting research topics including but not limited to model enhancement, creating synergy with traditional failure modeling approaches, and data-driven actionable recommendation and suggestions.

20 FOSSIL-FUELED POWER PLANTS↗

Design and implementation of I/O performance prediction scheme on HPC systems through large-scale log analysis

Abstract Large-scale high performance computing (HPC) systems typically consist of many thousands of CPUs and storage units used by hundreds to thousands of users simultaneously. Applications from large numbers of users have diverse characteristics, such as varying computation, communication, memory, and I/O intensity. A good understanding of the performance characteristics of each user application is important for job scheduling and resource provisioning. Among these performance characteristics, I/O performance is becoming increasingly important as data sizes rapidly increase and large-scale applications, such as simulation and model training, are widely adopted. However, predicting I/O performance is difficult because I/O systems are shared among all users and involve many layers of software and hardware stack, including the application, network interconnect, operating system, file system, and storage devices. Furthermore, updates to these layers and changes in system management policy can significantly alter the I/O behavior of applications and the entire system. To improve the prediction of the I/O performance on HPC systems, we propose integrating information from several different system logs and developing a regression-based approach to predict the I/O performance. Our proposed scheme can dynamically select the most relevant features from the log entries using various feature selection algorithms and scoring functions, and can automatically select the regression algorithm with the best accuracy for the prediction task. The evaluation results show that our proposed scheme can predict the write performance with up to 90% prediction accuracy and the read performance with up to 99% prediction accuracy using the real logs from the Cori supercomputer system at NERSC.

97 MATHEMATICS AND COMPUTING↗

High-Performance Deep Learning Toolbox for Genome-Scale Prediction of Protein Structure and Function

Computational biology is one of many scientific disciplines ripe for innovation and acceleration with the advent of high-performance computing (HPC). In recent years, the field of machine learning has also seen significant benefits from adopting HPC practices. In this work, we present a novel HPC pipeline that incorporates various machine-learning approaches for structure-based functional annotation of proteins on the scale of whole genomes. Our pipeline makes extensive use of deep learning and provides computational insights into best practices for training advanced deep-learning models for high-throughput data such as proteomics data. We showcase methodologies our pipeline currently supports and detail future tasks for our pipeline to envelop, including large-scale sequence comparison using SAdLSA and prediction of protein tertiary structures using AlphaFold2.

Gao, Mu↗

Argonne Leadership Computing Facility: 2021 Operational Assessment Report

This Operational Assessment Report describes how the Argonne Leadership Computing Facility (ALCF) met or exceeded every one of its goals for calendar year (CY) 2021 as an advanced scientific computing center. In CY 2021, the ALCF operated its production resource, Theta, an Intel-based Cray XC40 system (11.7-petaflops) augmented with 24 NVIDIA DGX A100-based nodes (3.9-petaflops) that supports diverse workloads, integrating data analytics with artificial intelligence (AI) training and learning in a single platform. In 2021, we began deploying Polaris, our newest 40- petaflops system, and augmented this powerful testbed system with an additional 28 nodes to support the integration of real-time experiments and HPC resources. We also deployed our two largest storage systems yet, named Grand and Eagle, that will bring new services to our users and will power data-driven research for years to come. Last year, Theta delivered a total of 20.8 million node-hours to 16 Innovative and Novel Computational Impact on Theory and Experiment (INCITE) projects and 7.2 million node-hours to ASCR Leadership Computing Challenge (ALCC) projects (32 awarded during the 2020–2021 ALCC year and 17 awarded during the 2021–2022 ALCC year), as well as substantial support to Director’s Discretionary (DD) projects (5.5 million node-hours). As Table ES.1 shows, Theta performed exceptionally well in terms of overall availability (95.1 percent), scheduled availability (99.4 percent), and utilization (98.1 percent; Table 2.1). As of the submission date of this document, ALCF’s user community has published 249 papers in high-quality, peer-reviewed journals and technical proceedings. At the 2021 International Conference for High Performance Computing, Networking, Storage and Analysis (SC’21), Argonne researchers won two HPCwire Readers’ Choice Awards and were part of a Gordon Bell Prize finalist team recognized for developing an AI-enabled, multi-resolution simulation framework for studying complex biomolecular machines. Their framework was used to observe the SARS-CoV-2 replication-transcription machinery in action, by directly integrating experimental data. ALCF also provided a comprehensive program of high-performance computing (HPC) support services to help our community make productive use of the facility’s diverse and growing collection of resources. We are now entering the exascale era, with exascale machines being planned for national laboratories across the country, including Aurora at Argonne National Laboratory (Argonne) in 2023. ALCF researchers have been leading and guiding numerous strategic activities that will push the boundaries of what’s possible in computational science and engineering and allow us to deliver science on day one.

97 MATHEMATICS AND COMPUTING↗

Impacts of floating-point non-associativity on reproducibility for HPC and deep learning applications

Run to run variability in parallel programs caused by floating-point non-associativity has been known to significantly affect reproducibility in iterative algorithms, due to accumulating errors. Non-reproducibility can critically affect the efficiency and effectiveness of correctness testing for stochastic programs. Recently, the sensitivity of deep learning training and inference pipelines to floating-point non-associativity has been found to sometimes be extreme. It can prevent certification for commercial applications, accurate assessment of robustness and sensitivity, and bug detection. New approaches in scientific computing applications have coupled deep learning models with high-performance computing, leading to an aggravation of debugging and testing challenges. Here we perform an investigation of the statistical properties of floating-point non-associativity within modern parallel programming models, and analyze performance and productivity impacts of replacing atomic operations with deterministic alternatives on GPUs. We examine the recently-added deterministic options in PyTorch within the context of GPU deployment for deep learning, uncovering and quantifying the impacts of input parameters triggering run to run variability and reporting on the reliability and completeness of the documentation. Finally, we evaluate the strategy of exploiting automatic determinism that could be provided by deterministic hardware, using the Groq LPUTM accelerator for inference portions of the deep learning pipeline. We demonstrate the benefits that a hardware-based strategy can provide within reproducibility and correctness efforts.

Shanmugavelu, Sanjif↗

The LCLStream Ecosystem for Multi-Institutional Dataset Exploration

We describe a new end-to-end experimental data streaming framework designed from the ground up to support new types of applications – AI training, extremely high-rate X-ray time-of-flight analysis, crystal structure determination with distributed processing, and custom data science applications and visualizers yet to be created. Throughout, we use design choices merging cloud microservices with traditional HPC batch execution models for security and flexibility. This project makes a unique contribution to the DOE Integrated Research Infrastructure (IRI) landscape. By creating a flexible, API-driven data request service, we address a significant need for high-speed data streaming sources for the X-ray science data analysis community. With the combination of data request API, mutual authentication web security framework, job queue system, high-rate data buffer, and complementary nature to facility infrastructure, the LCLStreamer framework has prototyped and implemented several new paradigms critical for future generation experiments.

Rogers, David [ORNL] (ORCID:0000000251871768)↗

DOE FAIR Surrogate Benchmarks Supporting AI and Simulation Research (SBI Surrogate Benchmark Initiative) (Final Report)

Computational Science is being revolutionized by integrating AI and simulation and, in particular, by deep learning surrogate models that can replace all or part of traditional large‐scale HPC computations. Such surrogates can achieve remarkable performance improvements, as much as several orders of magnitude, and save both compute time and energy. The Surrogate Benchmark Initiative (SBI) project creates a community repository and FAIR (Findable, Accessible, Interoperable, and Reusable) data ecosystem for HPC application surrogate benchmarks. The SBI team comes from Argonne National Laboratory (ANL), Indiana University (IU), Rutgers University, the University of Tennessee, Knoxville (UTK), and the University of Virginia(UVA). SBI repositories include data, code, and all relevant collateral artifacts, that the science and engineering community needs to use and reuse these data sets and surrogates. SBI repositories generate active research from both participants in SBI and the broader AI and domain science communities. This project develops surrogates that use several different neural nets to learn and quickly infer the results of simulations and data systems and capture them as surrogate benchmarks with a rich set of metadata, covering. Data; Model; Metrics specification; Machine specification; Science, Speed, Power Results, We research FAIR metadata for these benchmarks. We develop application surrogate examples as benchmarks across many fields (ANL, UTK, IU, UVA). We also study non Surrogate benchmarks that have many common features and similar issues regarding FAIRness. We work with MLCommons (UVA, UTK), which is a major machine learning benchmarking activity where we get metadata ontologies, software, and benchmarks, benchmarks have datasets, models, and metadata, and they need a technical framework developed by UTK and Rutgers and deployed by UVA. We study features of Surrogates, including performance, training set size, and uncertainty quantification (Rutgers, UVA and IU).

97 MATHEMATICS AND COMPUTING↗

Fault-Tolerant Deep Learning Cache with Hash Ring for Load Balancing in HPC Systems

Large-scale DL on HPC systems like Frontier and Summit uses distributed node-local caching to address scalability and performance challenges. However, as these systems grow more complex, the risk of node failures increases, and current caching approaches lack fault tolerance, jeopardizing large-scale training jobs. We analyzed six months of SLURM job logs from Frontier and found that over 30% of jobs failed after an average of 75 minutes. To address this, we propose fault-tolerance strategies that recache data lost from failed nodes using a hash ring technique for balanced data recaching in the distributed node-local caching, reducing reliance on the PFS. Our extensive evaluations on Frontier showed that the hash ring-based recaching approach reduced training time by approximately 25% compared to the approach that redirects I/O to the PFS after node failures and demonstrated effective load balancing of training data across nodes.

Lee, Seoyeong↗

The U.S. Department of Energy Computational Science Graduate Fellowship, 1991-2021: Follow-Up Study Shows Major Impact on Recipients and the Scientific Workforce

Since 1991, the U.S. Department of Energy Computational Science Graduate Fellowship (DOE CSGF) has addressed DOE National Laboratory needs as well as demands in the national workforce for trained professionals in computational science and engineering. Sponsored by the Department of Energy's Office of Science and the National Nuclear Security Administration, the DOE CSGF supports doctoral students in the pursuit of novel scientific or engineering discoveries using high-performance computing (HPC) resources. To meet the program’s core requirements, recipients participate in multidisciplinary studies, carry out at least one 12-week DOE laboratory research practicum, and contribute to an annual program review where the fellows present their research for sponsor review. The Krell Institute, which as managed the fellowship on behalf of the DOE since 1997, has commissioned several follow-up studies to examine the DOE CSGF recipients’ characteristics, fellows’ outcomes and professional accomplishments, alumni’s career paths and achievements, and recipients’ impact on national priorities through research and education.

97 MATHEMATICS AND COMPUTING↗

The U.S. Department of Energy Computational Science Graduate Fellowship, 1991-2021: Follow-Up Study Shows Major Impact on Recipients and the Scientific Workforce

Since 1991, the U.S. Department of Energy Computational Science Graduate Fellowship (DOE CSGF) has addressed DOE National Laboratory needs as well as demands in the national workforce for trained professionals in computational science and engineering. Sponsored by the Department of Energy's Office of Science and the National Nuclear Security Administration, the DOE CSGF supports doctoral students in the pursuit of novel scientific or engineering discoveries using high-performance computing (HPC) resources. To meet the program’s core requirements, recipients participate in multidisciplinary studies, carry out at least one 12-week DOE laboratory research practicum, and contribute to an annual program review where the fellows present their research for sponsor review. The Krell Institute, which as managed the fellowship on behalf of the DOE since 1997, has commissioned several follow-up studies to examine the DOE CSGF recipients’ characteristics, fellows’ outcomes and professional accomplishments, alumni’s career paths and achievements, and recipients’ impact on national priorities through research and education.

97 MATHEMATICS AND COMPUTING↗

A Vision for Coupling Operation of US Fusion Facilities with HPC Systems and the Implications for Workflows and Data Management

The operation of large US Department of Energy (DOE) research facilities, like the DIII-D National Fusion Facility, results in the collection of complex multi-dimensional scientific datasets, both experimental and model-generated. In the future, it is envisioned that integrated data analysis coupled with large-scale high performance computing (HPC) simulations will be used to improve experimental planning and operation. Practically, massive data sets from these simulations provide the physics basis for generation of both reduced semi-analytic and machine-learning-based models. Storage of both HPC simulation datasets (generated from US DOE leadership computing facilities) and experimental datasets presents significant challenges. In this paper, we present a vision for a DOE-wide data management workflow that integrates US DOE fusion facilities with leadership computing facilities. Data persistence and long-term availability beyond the length of allocated projects is essential, particularly for verification and recalibration of artificial intelligence and machine learning (AI/ML) models. Because these data sets are often generated and shared among hundreds of users across multiple leadership computing facility centers, they would benefit from cross-platform accessibility, persistent identifiers (e.g. DOI, or digital object identifier), and provenance tracking. Here, the ability to handle different data access patterns suggests that a combination of low cost, high latency (e.g. for storing ML training sets) and high cost, low latency systems (e.g. for real-time, integrated machine control feedback) may be needed.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Modeling the Interaction of Laser-Produced Proton Beams with Matter

A major goal of this project is to significantly increase our understanding of isochoric heating of matter using laser produced proton beams, and the associated high energy density (HED) and warm dense matter (WDM) regimes generated. This will benefit research fields such as planetary science, fusion energy, plasma physics, and material science. For example, it will enhance our understanding of WDM properties of iron and silica under conditions encountered in planetary interiors and diagnostic components in fusion devices exposed to high fluxes of energetic plasma ions. The project is motivated by recent experiments that irradiated Si targets with proton beams generated by the 20 TW-laser at the SLAC MEC end-station. The HED/WDM states are probed using the 50 fs hard X-rays available in the 3rd harmonic of the LCLS. As part of this project, results from the phase contrast X-ray imaging, which shows the generation of compression waves that produces rear surface spallation, are compared with results from the 3D multi-physics multi- material code, PISALE, that combines Arbitrary Lagrangian-Eulerian (ALE) hydrodynamics with Adaptive Mesh Refinement (AMR). This comparison required modifications to several physics models in the PISALE (Pacific Island Structured-AMR with ALE) code. An important aspect of this project is the continued training of graduate students in HED physics and in conducting complex multiphysics simulations.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Meta2DB: Curated Shotgun Metagenomic Feature Sets and Metadata for Health State Prediction

Meta2DB is a curated metagenomic and metadata database that provides structurally consistent microbiome taxonomy feature count tables for 13 897 samples across 84 studies, 23 disease states, and 34 geographical locations. All samples were uniformly processed using a streamlined metagenomic classification pipeline that employs a unique and comprehensive reference database indexed to contain all sequences across all kingdoms of life that were present in the NCBI Nucleotide (nt) database retrieved on 4 January 2023. This pipeline leverages high-performance computing (HPC) resources at Lawrence Livermore National Laboratory and was used to process 50TB of publicly available raw metagenomic sequence data. Extensive metadata curation was carried out through a combination of manual curation and automated parsing, producing a consistent inter-study metadata table specifically structured to facilitate training of ML models for prediction of human health.

Kok, C [Lawrence Livermore National Laboratory (LL↗

Focused Ion Beam Tomography of Alloy 617 Corroded in Molten Chloride Salt

Materials qualification of reactor structural materials is a critical step in rapid implementation of advanced nuclear reactor technologies, particularly to assess the corrosion performance in these designs. Accelerated qualification of reactor structural materials requires incorporating powerful computational toolsets, such as phase field modelling in the Multiphysics Object-Oriented Simulation Environment (MOOSE) framework, to predict the evolution of structural materials due to corrosion. Accordingly, computational toolsets will require experimental data generated at appropriate length scales to validate accuracy. Focused ion beam (FIB) provides a high degree of control over manipulation of materials for analytical purposes, including capturing data on the evolution in the microstructure and elemental composition of materials at the mesoscale, an appropriate length scale for phase field modelling of intergranular diffusion phenomena using the MOOSE framework. For instance, the FEI Helios G4 UX dual beam plasma FIB microscope at the Irradiated Materials Characterization Laboratory (IMCL) is capable of backscatter diffraction (EBSD) and energy-dispersive x-ray spectroscopy (EDS) documenting the evolution in the microstructure and elemental composition, respectively. The Helios can perform EDS and EBSD three-dimensionally (3D) using tomography, which is then combined using different software packages to visualize 3D volumes correlating elemental composition to microstructural data. The purpose of this investigation was to develop a streamlined characterization and data processing workflow for 3D tomography studies on the FEI Helios G4 plasma FIB. The investigation is segmented into three parts: 1) Optimizing the data collection workflow, 2) identifying appropriate data processing and visualization software (i.e. DREAM.3D, MIPAR, and VGStudioMax), and 3) establishing an infrastructure for public release. The optimization of the data collection workflow is in collaboration with members of the U220 department to setup formal training on the tomography operation of the G4, through ThermoFisher Scientific, and exploring DREAM.3D, MIPAR, and VGStudioMax data processing/visualization software packages. VGStudioMax currently demonstrates the most promise for future use. Optimization of the data collection and processing workflow is still ongoing. A collaboration with INL High Performance Computing (HPC) established an open-source license for expediting the public release of FIB tomography datasets through HPC. FIB tomography data generated by the G4 will provide comprehensive data for validating 3D phase field mesoscale modelling tools within the MOOSE framework for accelerated qualification of reactor structural materials.

Copeland-Johnson, Trishelle↗

Performance-Aligned LLMs for Generating Fast HPC Code

Optimizing scientific software is a difficult task because codebases are often large and complex, and performance can depend upon several factors including the algorithm, its implementation, and hardware among others. Causes of poor performance can originate from disparate sources and be difficult to diagnose. Recent years have seen a multitude of work that use large language models (LLMs) to assist in software development tasks. However, these tools are trained to model the distribution of code as text, and are not specifically designed to understand performance aspects of code. In this work, we introduce a reinforcement learning based methodology to align the outputs of code LLMs with performance. This allows us to build upon the current code modeling capabilities of LLMs and extend them to generate better performing code. Here, we demonstrate that our fine-tuned model improves the expected speedup of generated code over base models for a set of benchmark tasks from 0.9 to 1.6 for serial code and 1.9 to 4.5 for OpenMP parallel code.

Computer science↗

2019 Budget Request for the DOE Computational Science Graduate Fellowship (CSGF) Grant

The Department of Energy Computational Science Graduate Fellowship (DOE CSGF) is necessary to meet the continual challenging national workforce needs that arise as computational science and engineering problems continue to grow in scope and complexity. Computational science and engineering (CSE) is a multidisciplinary approach that uses scientific computing to solve practical problems methods and to supply technical tools across the scientific discovery spectrum. In particular, the DOE CSGF emphasizes high-performance computing (HPC) that enables CSE that advances science and engineering in directions important to the DOE and the economy in general. Over the past half-century, HPC has been an essential tool for DOE’s success. During this period, important missions, such as nuclear stockpile stewardship, have turned to HPC as an essential technology. Entire science disciplines, such as biology and cosmology, have been transformed through the augmentation of scientific observation via HPC. At government laboratories and in industry, DOE CSGF alumni are helping push traditional HPC boundaries while contributing to discoveries in high-energy physics, renewable energy, fusion-reactor design, additive manufacturing, nanomaterials for next-generation batteries and transistors, and turbine and advanced nuclear reactor modeling. In addition, HPC is used to address national health needs that will eventually point to cures both by helping cancer researchers manage and analyze huge troves of data, by simulating biological mechanisms, and by accelerating drug development — including continuing to rise to the challenge of pandemic-related research. A 2023 report from the ASCAC Subcommittee on American Competitiveness and Innovation to the ASCR office, “Can the United States Maintain Its Leadership in High-Performance Computing?” says of the Program, “The CSGF program provides a barometer for disciplines that will be of interest to future DOE computing.” An explosion in scientific and technological data has driven the need for increasingly sophisticated HPC to transform those data into scientific understanding. With access to more and more data and the proliferation of HPC, Machine Learning and Artificial Intelligence are experiencing a renaissance, complementing the now well-established use of computational simulation. Indeed, in its September 2020 subcommittee report on “AI/ML, Data Intensive Science and High-Performance Computing”, the DOE Advanced Scientific Computing Advisory Committee (ASCAC) explicitly called for a fellowship program to train computational and data scientists to tackle exascale and data-intensive computing challenges. This collaboration of empirical and theory-based modeling will increasingly inform federal policymakers whose decisions affect American society and future generations, and it requires highly skilled and intellectually agile computational scientists who can support the fast-moving DOE National Laboratory research environment. In fact, the DOE CSGF program has explicitly and consistently addressed this need.

97 MATHEMATICS AND COMPUTING↗

Quantum Computing Strategy 2026

Quantum computing (QC) is a rapidly maturing technology with the potential for revolutionary impacts on stockpile stewardship science and national security. Recent developments in fault-tolerant architectures have compressed vendor roadmaps, and predictions of a production-ready quantum computer by the mid-2030s are becoming increasingly credible. This strategy provides a roadmap for integrating QC into the Advanced Simulation and Computing (ASC) program by investing in four strategic focus areas: 1. Develop Capabilities in Mission-Relevant Quantum Applications: ASC will prioritize developing quantum-ready applications in mission areas that have shown significant promise for quantum advantage, including simulations of materials in extreme environments, nuclear dynamics, solving linear and nonlinear partial differential equations, and uncertainty quantification. These applications directly support stockpile stewardship science and modernization objectives. 2. Conduct R&D in Algorithms, Software, and Hardware: Sustained research into quantum algorithms, robust software tools, and quantum hardware is essential. ASC will develop efficient quantum algorithms; invest in quantum compilers, debuggers, and performance tools; and explore specialized quantum hardware tailored to NNSA’s unique requirements. 3. Engage with Vendors and Partners: Early and active collaboration with commercial quantum hardware vendors and academic partners is critical. Through testbeds, co-design agreements, and quantum demonstration facilities, ASC will influence hardware design, gain early access to emerging technologies, and ensure that quantum platforms evolve to meet mission needs. 4. Build Knowledge, Experience, and Workforce: Expanding and upskilling the quantum-trained workforce is essential to long-term success. This includes hiring, internal training, university outreach, and postdoctoral support to ensure ASC maintains the expertise required to operate, program, and integrate quantum systems as they become available. While quantum computing will never replace classical computing, it has the potential to solve certain problems with speed and accuracy that would be unachievable using any conceivable classical high-performance computing (HPC) system. By investing strategically in QC, ASC will help propel the emergent QC industry, maintain U.S. technological leadership, ensure mission readiness, and position itself to rapidly adopt quantum technologies as they mature.

97 MATHEMATICS AND COMPUTING↗