Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “petaflop”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Scalable multiscale modeling of platelets with 100 million particles

Here, we developed the core components of the AI-aided multiple time stepping algorithm for multiscale modeling of cell dynamics. This algorithm was implemented and analyzed on two supercomputer architectures with an application of simulating the aggregation of 250 platelets, or 102 million particles. To scale on these computers with complex memory and network architectures with GPUs, we devised a biomechanics-informed task mapping scheme to optimize load imbalance, communications, and memory utilization. Our simulations, scaling well up to 192 nodes on a Summit-like supercomputer with a peak speed of 11 petaflops, achieved a rate of 423 μs/day which is 500 times faster than the conventional algorithm using static time step and this has enabled studies of record size blood clots at record spatial–temporal resolutions. Additionally, we discovered the sensitive dependence of the scalability and execution time on the methods of decomposition, CPU–GPU coupling, and task mapping.

97 MATHEMATICS AND COMPUTING↗

Study of interconnect errors, network congestion, and applications characteristics for throttle prediction on a large scale HPC system

Today’s High Performance Computing (HPC) systems contain thousand of nodes which work together to provide performance in the order of petaflops. The performance of these systems depends on various components like processors, memory, and interconnect. Among all, interconnect plays a major role as it glues together all the hardware components in an HPC system. A slow interconnect can impact a scientific application running on multiple processes severely as they rely on fast network messages to communicate and synchronize frequently. Unfortunately, the HPC community lacks a study that explores different interconnect errors, congestion events and applications characteristics on a large-scale HPC system. In our previous work, we process and analyze interconnect data of the Titan supercomputer to develop a thorough understanding of interconnects faults, errors, and congestion events. In this work, we first show how congestion events can impact application performance. We then investigate application characteristics interaction with interconnect errors and network congestion to predict applications encountering congestion with more than 90% accuracy.

97 MATHEMATICS AND COMPUTING↗

Toward Exascale: Overview of Large Eddy Simulations and Direct Numerical Simulations of Nuclear Reactor Flows with the Spectral Element Method in Nek5000

At the beginning of the last decade, Petascale supercomputers (i.e., computers capable of more than 1 petaFLOP) emerged. Now, at the dawn of exascale supercomputing, we provide a review of recent landmark simulations of portions of reactor components with turbulence-resolving techniques that this computational power has made possible. In fact, these simulations have provided invaluable insight into flow dynamics, which is difficult or often impossible to obtain with experiments alone. We focus on simulations performed with the spectral element method, as this method has emerged as a powerful tool to deliver massively parallel calculations at high fidelity by using large eddy simulation or direct numerical simulation. We also limit this paper to constant-property incompressible flow of a Newtonian fluid in the absence of other body or external forces, although the method is by no means limited to this class of flows. We briefly review the fundamentals of the method and the reasons it is compelling for the simulation of nuclear engineering flows. We review in detail a series of Petascale simulations, including the simulations of helical coil steam generators, fuel assemblies, and pebble beds. Even with Petascale computing, however, limitations for nuclear modeling and simulation tools remain. In particular, the size and scope of turbulence-resolving simulations are still limited by computing power and resolution requirements, which scale with the Reynolds number. In the final part of this paper, we discuss the future of the field, including recent advancements in emerging architectures such as GPUbased supercomputers, which are expected to power the next generation of high-performance computers.

computational fluid dynamics↗

ORBIT: Oak Ridge Base Foundation Model for Earth System Predictability

Earth system predictability is challenged by the complexity of environmental dynamics and the multitude of variables involved. Current AI foundation models, although advanced by leveraging large and heterogeneous data, are often constrained by their size and data integration, limiting their effectiveness in addressing the full range of Earth system prediction challenges. To overcome these limitations, we introduce the Oak Ridge Base Foundation Model for Earth System Predictability (ORBIT), an advanced vision transformer model that scales up to 113 billion parameters using a novel hybrid tensor-data orthogonal parallelism technique. As the largest model of its kind, ORBIT surpasses the current climate AI foundation model size by a thousandfold. Performance scaling tests conducted on the Frontier supercomputer have demonstrated that ORBIT achieves 684 petaFLOPS to 1.6 exaFLOPS sustained throughput, with scaling efficiency maintained at 41% to 85% across 49,152 AMD GPUs. These breakthroughs establish new advances in AI-driven climate modeling and demonstrate promise to significantly improve the Earth system predictability.

Wang, Xiao↗

Characterization and identification of HPC applications at leadership computing facility

High Performance Computing (HPC) is an important method for scientific discovery via large-scale simulation, data analysis, or artificial intelligence. Leadership-class supercomputers are expensive, but essential to run large HPC applications. The Petascale era of supercomputers began in 2008, with the first machines achieving performance in excess of one petaflops, and with the advent of new supercomputers in 2021 (e.g., Aurora, Frontier), the Exascale era will soon begin. However, the high theoretical computing capability (i.e., peak FLOPS) of a machine is not the only meaningful target when designing a supercomputer, as the resources demand of applications varies. A deep understanding of the characterization of applications that run on a leadership supercomputer is one of the most important ways for planning its design, development and operation. In order to improve our understanding of HPC applications, user demands and resource usage characteristics, we perform correlative analysis of various logs for different subsystems of a leadership supercomputer. This analysis reveals surprising, sometimes counter-intuitive patterns, which, in some cases, conflicts with existing assumptions, and have important implications for future system designs as well as supercomputer operations. For example, our analysis shows that while the applications spend significant time on MPI, most applications spend very little time on file I/O. Combined analysis of hardware event logs and task failure logs show that the probability of a hardware FATAL event causing task failure is low. Combined analysis of control system logs and file I/O logs reveals that pure POSIX I/O is used more widely than higher level parallel I/O. Based on holistic insights of the application gained through combined and co-analysis of multiple logs from different perspectives and general intuition, we engineer features to "fingerprint" HPC applications. We use t-SNE (a machine learning technique for dimensionality reduction) to validate the explainability of our features and finally train machine learning models to identify HPC applications or group those with similar characteristic. To the best of our knowledge, this is the first work that combines logs on file I/O, computing, and inter-node communication for insightful analysis of HPC applications in production.

Liu, Zhengchun↗

Center of Excellence Collaboration Projects: Second Wave

The NEAMS program aims to develop an integrated multi-physics simulation capability “pellet-to-plant” for the design and analysis of future generations of nuclear power plants. In particular, the Reactor Product Line code suite's multi-resolution hierarchy is being designed to ultimately span the full range of length and time scales present in relevant reactor design and safety analyses, as well as scale from desktop to petaflop computing platforms. In particular the NEAMS program is supporting the development of novel thermal-hydraulic codes. The Center of Excellence for Thermal Fluids Application in Nuclear Energy, launched in 2018, has as its key goals to serve as a front door to industry. The Center of Excellence for Thermal Fluids Applications in Nuclear Energy has recently launched a program to start collaborative efforts between the laboratories and industry with the objective of stimulating cooperation and increasing adoption of T/H tools developed under NEAMS by the industry-at-large. In particular, two industry partners agreed to participate in a second wave of short-term collaborations aimed at demonstrating the value of NEAMS tools to their designs. These are Framatome and TerraPower who proposed projects related to LWR (Light Water Reactor) and MSFR (Molten Salt Fast Reactor) designs respectively. In this report we present the results of these two collaborations.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

The Multiphysics on Advanced Platforms Project

In 2015, the Lawrence Livermore National Laboratory started development of next-generation multiphysics simulation capabilities for the National Nuclear Security Administration under the Advanced Technologies Development and Mitigation (ATDM) element of the Advanced Simulation and Computing program in collaboration with the Exascale Computing Project (ECP). A key driver for this effort across the NNSA tri-lab was the emergence of advanced high performance computing (HPC) architectures based on heterogeneous compute capabilities, including GPU based systems, as part of the national drive toward exascale computing platforms at multiple Department of Energy (DOE) facilities. Developing a multiphysics code capable of meeting the various simulation needs of the NNSA as defined by the current generation of integrated codes (or ICs), initially developed as part of the Accelerated Strategic Computing Initiative (ASCI) program beginning in 1996, and able to scale to the current 100 petaflop class pre-exascale systems, as well the forthcoming exaflop class computers, is a daunting challenge. To accomplish this ambitious goal, LLNL has embraced two key themes: use of high-order numerical methods and a modular approach to code development. The LLNL next generation effort is organized under the Multi-Physics on Advanced Platforms Project (MAPP). A foundational component of MAPP is the Axom computer science (CS) toolkit which provides infrastructure for the development of modular, performance portable, multi-physics application codes. MARBL is a next-generation application code built on the Axom base to address the modeling needs of the high energy density physics (HEDP) community for simulating high-explosive, magnetic or laser driven experiments such as inertial confinement fusion (ICF), pulsed-power magneto-hydrodynamics (MHD), equation of state (EOS) and material strength studies as part of the NNSA’s stockpile stewardship program (SSP).

97 MATHEMATICS AND COMPUTING↗

IC reports (2021)

We requested HPC support to continue research on the seismic waves generated by impacts. We used the following codes: (1) the Hybrid Optimization Software Suite (HOSS), developed at LANL. HOSS is based on a combined Finite and Discrete Element Method (FDEM). New material models are developed for the sedimentary rocks. HOSS has been recently benchmarked to iSale and FLAG codes. SPECFEM3D is an open-source code developed since the last 90s. It won the Gordon Bell award for best performance in 2003, was finalist again in 2008 for a run at 0.16 petaflops on 149,784 cores on the ‘Jaguar’ Cray system at Oak Ridge National Laboratory. It also won the BULL Joseph Fourier supercomputing award in 2010.; SW4 is a 4th-order finite difference code developed at LLNL which is currently actively developed to handle complex 3D models and to be ported on future exascale platforms. We assessed our need to a total of 3.9M CPU-hrs for year 1 and and 3.1 M for year 2.

79 ASTRONOMY AND ASTROPHYSICS↗

Advanced Computing Annual Report 2025 [Slides]

In Fiscal Year (FY) 2025, the National Laboratory of the Rockies (NLR) continued to advance computing as a cornerstone of energy innovation, expanding the Kestrel high-performance computing (HPC) system to 56 peak petaflops. This growth strengthened Kestrel's role as a national asset for applied energy research, enabling larger, more complex simulations and accelerating the integration of artificial intelligence (AI) methods across the laboratory's computing portfolio. In FY 2025, AI was a component of most projects running on Kestrel, underscoring its central role in modern energy science and engineering. Kestrel supported a broad and diverse set of 507 modeling and simulation projects, engaging 855 researchers across the U.S. Department of Energy's (DOE's) Office of Critical Minerals and Energy Innovation (CMEI) portfolio and other offices, as well as partners from industry, academia, and utilities. These efforts span critical materials discovery, energy systems modeling, grid modernization, advanced manufacturing, and other areas essential to strengthening U.S. energy security and competitiveness. Together, these collaborations produced 708 technical outputs, including 293 peer-reviewed publications, reflecting both the depth and impact of the science enabled by NLR's computing capabilities. This year's report highlights the growing importance and benefit of AI throughout NLR's research programs and features work by early career researchers who are helping shape the future of computing-enabled energy innovation. Explore these sections and the many project successes captured in the pages that follow.

97 MATHEMATICS AND COMPUTING↗

Estimate of the Mass and Radial Profile of the Orphan–Chenab Stream's Dwarf-galaxy Progenitor Using MilkyWay@home

We fit the mass and radial profile of the Orphan–Chenab Stream's (OCS) dwarf-galaxy progenitor by using turnoff stars in the Sloan Digital Sky Survey and the Dark Energy Camera to constrain N-body simulations of the OCS progenitor falling into the Milky Way on the 1.5 PetaFLOPS MilkyWay@home distributed supercomputer. We infer the internal structure of the OCS's progenitor under the assumption that it was a spherically symmetric dwarf galaxy composed of a stellar system embedded in an extended dark matter halo. We optimize the evolution time, the baryonic and dark matter scale radii, and the baryonic and dark matter masses of the progenitor using a differential evolution algorithm. The likelihood score for each set of parameters is determined by comparing the simulated tidal stream to the angular distribution of OCS stars observed in the sky. We fit the total mass of the OCS's progenitor to (2.0 ± 0.3) × 10 7 M ⊙ with a mass-to-light ratio of γ = 73.5 ± 10.6 and (1.1 ± 0.2) × 10 6 M ⊙ within 300 pc of its center. Within the progenitor's half-light radius, we estimate a total mass of (4.0 ± 1.0) × 10 5 M ⊙ . We also fit the current sky position of the progenitor's remnant to be (α, δ) = ((166.0 ± 0.9)°, (–11.1 ± 2.5)°) and show that it is gravitationally unbound at the present time. The measured progenitor mass is on the low end of previous measurements and, if confirmed, lowers the mass range of ultrafaint dwarf galaxies. Our optimization assumes a fixed Milky Way potential, OCS orbit, and radial profile for the progenitor, ignoring the impact of the Large Magellanic Cloud.

79 ASTRONOMY AND ASTROPHYSICS↗

NLR HPC Eagle Node Power Data

Power time series captured from all Eagle nodes using iLO (Integrated Lights Out) The Eagle HPC operated at NLR from 2019 through 2024. Eagle was a 2,000-node, 8-petaflop system. This dataset is a comprehensive time series of instantaneous snapshots of power usage at 1 minute intervals from all nodes at the node level. Data provided in compressed Hive dataset/Parquet format. iLO Power Time Series Fields ts: Timestamp dv: Device / Node - Rack and Unit - r103u17 == r(ack)103u(nit)17 vl: Value - Value in watts (instantaneous value at sampling time) day month year

97 MATHEMATICS AND COMPUTING↗

NLR HPC Eagle GPU Node Metrics

Ganglia node metrics and iLO (Integrated Lights Out) power data captured from six representative Eagle GPU nodes The Eagle HPC operated at NLR from 2019 through 2024. Eagle was a 2,000-node, 8-petaflop system. This dataset is a representative sample of metrics for 6 of the GPU nodes. Each GPU node contained 2 CPUs and 2 GPUs. Data provided in compressed CSV format. Ganglia and iLO Power Time Series Fields ts: Timestamp dv: Device / Node - Rack and Unit - r103u17 == r(ack)103u(nit)17 mt: Metric (only present for Ganglia) vl: Value - Value in watts for iLO power (instantaneous value at sampling time) or specified Ganglia metric below Ganglia Metrics Metric name -- Metric description -- Unit cpu_aidle -- Percent of time since boot idle CPU -- Percent cpu_idle -- Percent CPU idle -- Percent cpu_nice -- Percent CPU nice -- Percent cpu_speed -- Speed in MHz of CPU -- MHz cpu_user -- Percent CPU user -- Percent cpu_wio -- The percentage of CPU Wait I/O -- Percent gpu0_bar1_memory -- Used GPU bar1 memory -- MB gpu0_decoder_util -- GPU decoder utilization -- Percent gpu0_ecc_db_error -- Total ECC error counts for the GPU -- Number gpu0_encoder_util -- GPU encoder utilization -- Percent gpu0_fan -- Fan speed -- RPM gpu0_fb_memory -- Used GPU framebuffer memory -- MB gpu0_graphics_clock_report -- Current clock speeds for the device -- MHz gpu0_mem_total -- Memory total -- MB gpu0_mem_util -- Memory utilization -- Percent gpu0_power_usage_report -- Power usage report -- Watts gpu0_temp -- GPU 1 temperature -- Celsius gpu1_bar1_memory -- Used GPU bar1 memory -- MB gpu1_decoder_util -- GPU decoder utilization -- Percent gpu1_ecc_db_error -- Total ECC error counts for the GPU -- Number gpu1_encoder_util -- GPU encoder utilization -- Percent gpu1_fan -- Fan speed -- RPM gpu1_fb_memory -- Used GPU framebuffer memory -- MB gpu1_graphics_clock_report -- Current clock speeds for the GPU -- MHz gpu1_mem_total -- Memory total -- MB gpu1_mem_util -- Memory utilization -- MB gpu1_power_usage_report -- Power usage report -- Watts gpu1_temp -- GPU 1 temperature -- Celsius ipmi_cpu1_temp -- CPU 1 temperature -- Celsius ipmi_cpu2_temp -- CPU 2 temperature -- Celsius ipmi_inlet_ambient_temp -- Temperature measured at intake -- Celsius ipmi_vr_p1_temp -- CPU 1 voltage regulator temperature -- Celsius ipmi_vr_p2_temp -- CPU 2 voltage regulator temperature -- Celsius mem_buffers -- Amount of buffered memory -- Bytes mem_cached -- Amount of cached memory -- Bytes mem_free -- Amount of available memory -- Bytes mem_shared -- Amount of shared memory -- Bytes mem_total -- Amount of available memory -- Bytes

97 MATHEMATICS AND COMPUTING↗

NLR HPC Eagle Jobs Data and Additional Energy Metrics

Overview: Anonymized job-level records from the Eagle high-performance computing (HPC) system at the National Laboratory of the Rockies (NLR). Each record represents a Slurm batch job with scheduling metadata, resource requests, resource utilization, CPU/GPU energy consumption, and efficiency metrics. Sensitive fields (user, account, job name) are replaced with cryptographic hashes. System & Timeframe: Eagle was a 2,000-node, 8-petaflop system operated at NLR from 2019–2024. Data covers the full operational lifetime of the system. Slurm data was processed nightly; timestamps are in Mountain Time. Funding provided by the U.S. Department of Energy, EERE. Files: esif.hpc.eagle.job-anon.zip — Core anonymized job records (Hive-partitioned Parquet) esif.hpc.eagle.job-anon-energy-metrics.zip — Same records with additional iLO and Ganglia energy metrics datacard.md — Full dataset documentation ~13.8 million rows, 62 variables. Readable with PyArrow, pandas, DuckDB, Apache Spark, or any Parquet-compatible tool. Data Collection: Jobs collected via sacct through a pipeline: Eagle Jobs API → Redpanda → StreamSets → HPCMON API → PostgreSQL. Node-level power from iLO (HP Integrated Lights-Out); GPU power from Ganglia monitoring, joined to jobs via node lists and time ranges. Preprocessing: Anonymization of name, user, and account fields via cryptographic hashing Derived columns: queue_wait, cpu_eff, max_mem_eff Simplified job state mapping (e.g., "CANCELLED BY 12345" → "CANCELLED") QoS accounting rules (buy-in, standby, or Slurm QoS value) CPU energy estimated from TDP (200W, Intel Xeon Gold 6154, 18 cores) Timezone-aware columns (_tz) sourced from LEX accounting database to correctly handle DST transitions Key Variables: Scheduling: job_id, partition, state_simple, submit_time_tz, start_time_tz, end_time_tz, queue_waitResources: nodes_req/used, processors_req/used, memory_req, wallclock_req/used, gpus_requested Efficiency: cpu_eff, max_mem_eff Energy: cpu_energy_tdp_estimated_max/used_watt_hours, node_energy_total_watt_hours (iLO), gpu0/1_energy_total_watt_hours (Ganglia) Partitions: bigmem, bigmem-8600, bigscratch, csc, dav, ddn, debug, gpu, haswell, long, mono, short, standard Job States: CANCELLED, COMPLETED, FAILED, NODE_FAIL, OUT_OF_MEMORY, PENDING, RUNNING, TIMEOUT QoS Levels: Unknown, normal, buy-in, debug, penalty, high, standby Important Notes: Non-_tz timestamp columns may be off by one hour across DST boundaries; use _tz columns for time difference calculations Energy fields are null for jobs without monitoring coverage Job step records and raw Slurm JSONB fields are excluded from this extract Do not attempt to re-identify individuals from hashed fields

97 MATHEMATICS AND COMPUTING↗

Parallel supercomputing with commodity components

We have implemented a parallel computer architecture based entirely upon commodity personal computer components. Using 16 Intel Pentium Pro microprocessors and switched fast ethernet as a communication fabric, we have obtained sustained performance on scientific applications in excess of one Gigaflop. During one production astrophysics treecode simulation, we performed 1.2 x 10(sup 15) floating point operations (1.2 Petaflops) over a three week period, with one phase of that simulation running continuously for two weeks without interruption. We report on a variety of disk, memory and network benchmarks. We also present results from the NAS parallel benchmark suite, which indicate that this architecture is competitive with current commercial architectures. In addition, we describe some software written to support efficient message passing, as well as a Linux device driver interface to the Pentium hardware performance monitoring registers.

Performance↗

Computational Nanoelectronics and Nanotechnology at NASA ARC

Both physical and economic considerations indicate that the scaling era of CMOS will run out of steam around the year 2010. However, physical laws also indicate that it is possible to compute at a rate of a billion times present speeds with the expenditure of only one Watt of electrical power. NASA has long-term needs where ultra-small semiconductor devices are needed for critical applications: high performance, low power, compact computers for intelligent autonomous vehicles and Petaflop computing technology are some key examples. To advance the design, development, and production of future generation micro- and nano-devices, IT Modeling and Simulation Group has been started at NASA Ames with a goal to develop an integrated simulation environment that addresses problems related to nanoelectronics and molecular nanotechnology. Overview of nanoelectronics and nanotechnology research activities being carried out at Ames Research Center will be presented. We will also present the vision and the research objectives of the IT Modeling and Simulation Group including the applications of nanoelectronic based devices relevant to NASA missions.

Saini, Subhash↗

High End Computing Technologies for Earth Science Applications: Trends, Challenges, and Innovations

Earth science applications of the future will stress the capabilities of even the highest performance supercomputers in the areas of raw compute power, mass storage management, and software environments. These NASA mission critical problems demand usable multi-petaflops and exabyte-scale systems to fully realize their science goals. With an exciting vision of the technologies needed, NASA has established a comprehensive program of advanced research in computer architecture, software tools, and device technology to ensure that, in partnership with US industry, it can meet these demanding requirements with reliable, cost effective, and usable ultra-scale systems. NASA will exploit, explore, and influence emerging high end computing architectures and technologies to accelerate the next generation of engineering, operations, and discovery processes for NASA Enterprises. This article captures this vision and describes the concepts, accomplishments, and the potential payoff of the key thrusts that will help meet the computational challenges in Earth science applications.

Parks, John↗

Computational Nanoelectronics and Nanotechnology at NASA ARC

Both physical and economic considerations indicate that the scaling era of CMOS will run out of steam around the year 2010. However, physical laws also indicate that it is possible to compute at a rate of a billion times present speeds with the expenditure of only one Watt of electrical power. NASA has long-term needs where ultra-small semiconductor devices are needed for critical applications: high performance, low power, compact computers for intelligent autonomous vehicles and Petaflop computing technolpgy are some key examples. To advance the design, development, and production of future generation micro- and nano-devices, IT Modeling and Simulation Group has been started at NASA Ames with a goal to develop an integrated simulation environment that addresses problems related to nanoelectronics and molecular nanotecnology. Overview of nanoelectronics and nanotechnology research activities being carried out at Ames Research Center will be presented. We will also present the vision and the research objectives of the IT Modeling and Simulation Group including the applications of nanoelectronic based devices relevant to NASA missions.

Saini, Subhash↗