Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “computer system benchmarking”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Fast convolutional neural networks on FPGAs with hls4ml

We introduce an automated tool for deploying ultra low-latency, low-power deep neural networks with convolutional layers on field-programmable gate arrays (FPGAs). By extending the hls4ml library, we demonstrate an inference latency of 5 µs using convolutional architectures, targeting microsecond latency applications like those at the CERN Large Hadron Collider. Considering benchmark models trained on the Street View House Numbers Dataset, we demonstrate various methods for model compression in order to fit the computational constraints of a typical FPGA device used in trigger and data acquisition systems of particle detectors. In particular, we discuss pruning and quantization-aware training, and demonstrate how resource utilization can be significantly reduced with little to no loss in model accuracy. We show that the FPGA critical resource consumption can be reduced by 97% with zero loss in model accuracy, and by 99% when tolerating a 6% accuracy degradation.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Enabling topography-resolving structural dynamic contact simulation

Damping of structures and systems is often dominated by frictional dissipation in connections, the prediction of which remains a longstanding scientific challenge. Previous studies have shown that the actual topography of contact interfaces may have a strong effect, especially in the partial slip/liftoff regime. We recently proposed a multi-scale method, which couples finite element and boundary element modeling. The primary benefit of this approach is that it permits to analyze the effect of the actual contact topography on the dynamics of jointed structures. While this multi-scale modeling method was initially developed for quasi-static analysis, we demonstrate herein how it can be used for time step integration and Harmonic Balance analysis. We cross-verify those fully dynamic analysis methods against each other and quasi-static results, for the S4 Beam benchmark. We compare the multi-scale method against state-of-the-art full-FE analysis, in terms of numerical damping and computational performance. Some discrepancy is found to be of physical origin. Depending on the load history, it is shown that the system settles to a slightly different equilibrium. Finally, transient multi-scale simulations enable the prediction of this interesting phenomenon, for the first time, for a structure with bolted joints.

Frictional-unilateral contact↗

Sparse chronology strategy for integrating seasonal energy storage in capacity expansion models

Here, this study develops the sparse chronology method to enhance the representative period framework in capacity expansion models, enabling the effective integration of long-duration energy storage modeling. Traditional representative period methods cannot capture the state of charge of seasonal energy storage systems because they do not establish effective inter-day linkages to connect the state of charge between periods. The sparse chronology approach addresses this limitation by establishing inter-day linkages that allow state of charge to shift inter-seasonally. At the same time, it groups identical representative days into partitions, applying constraints sparsely and implicitly to reduce computational load further. Validation results demonstrate that this method successfully simulates long-duration energy storage patterns, achieving close alignment with a continuous yearly benchmark model, with seasonal trends and state of charge cycles clearly represented. The computational load analysis reveals that the sparse chronology method efficiently applies constraints on maximum and minimum state of charge limits within the representative day framework, eliminating the need for detailed constraints on each individual day. By partitioning representative days and constraining only the start and end of each partition, the method significantly decreases computational requirements. Simulation results show that sparse chronology closely approximates the continuous yearly method's accuracy, even with as few as 20 representative days, achieving correlation values with the benchmark of nearly 0.9 in state of charge plots. Furthermore, it maintains computational efficiency, requiring only 4 % of the solver time compared to the continuous yearly method with 20 representative days. This approach allows capacity expansion models to incorporate long-duration energy storage with high temporal, spatial, and technological resolution, enabling more detailed modeling for large-scale power systems.

24 POWER TRANSMISSION AND DISTRIBUTION↗

GronOR: Massively Parallel and GPU-Accelerated Non-Orthogonal Configuration Interaction for Large Molecular Systems

GronOR is a program package for non-orthogonal configuration interaction calculations for an electronic wave function built in terms of anti-symmetrized products of multi-configuration molecular fragment wave functions. The two-electron integrals that have to be processed may be expressed in terms of atomic orbitals or in terms of an orbital basis determined from the molecular orbitals of the fragments. The code has been specifically designed for execution on distributed memory massively parallel and Graphics Processing Unit (GPU)-accelerated computer architectures, using an MPI+OpenACC/OpenMP programming approach. The task-based execution model used in the implementation allows for linear scaling with the number of nodes on the largest pre-exascale architectures available, provides hardware fault resiliency, and enables effective execution on systems with distinct central processing unit-only and GPU-accelerated partitions. The code interfaces with existing multi-configuration electronic structure codes that provide optimized molecular fragment orbitals, configuration interaction coefficients, and the required integrals. Algorithm and implementation details, parallel and accelerated performance benchmarks, and an analysis of the sensitivity of the accuracy of results and computational performance to thresholds used in the calculations are presented.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

AutonomieAI: An efficient and deployable vehicle energy consumption estimation toolkit

Here, this paper presents AutonomieAI, a novel toolkit designed for efficient energy estimation of vehicles across diverse trip scenarios, routes, and drive cycles, applicable to a broad range of vehicle powertrain technologies. It leverages state-of-the-art Machine Learning techniques to deliver real-time energy prediction of vehicles, enabling co-simulation with transportation level system tools and opening doors for large-scale optimization at city, network or national level. Benchmark results show that AutonomieAI achieves high accuracy, with an average percentage error below 2% for most powertrain types, and computational efficiency capable of processing over 10,000 trips per second. Applications of AutonomieAI have potential to offer the flexibility to assist in solving eco-routing problems, optimize for vehicle and powertrain selection, study charging decision behavior, and optimize for charging station placement. AutonomieAI is the result of large neural network based model architectures, trained on very large and unique high fidelity vehicle simulation data. It is lightweight, deployable, efficient and has accuracy comparable to specialized and complex physics based simulation softwares.

Autonomie↗

Ab Initio Polariton Transport Dynamics with the Classical Path Approximation

We present an ab initio framework for simulating polariton transport dynamics based on the classical path approximation (CPA). The quantum dynamics of polariton transport involves simulating many electronic degrees of freedom, making a fully ab initio dynamics simulation computationally expensive. We demonstrate that the CPA, which removes the need for excited-state nuclear gradients, is well-suited for polaritonic systems because collective light–matter coupling leads to vanishing excited-state forces. Benchmark comparisons between CPA and full evaluation of the excited-state forces show excellent agreement for polariton transport results in model light–matter systems such as polariton group velocities and mean-squared displacements. Ab initio simulations of polariton transport using CPA reproduce key physical trends that are observed in experiments with BODIPY molecules. Our work establishes the CPA as a highly efficient tool for ab initio investigations of transport and energy flow in hybrid light–matter systems.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

On the novel 3-D neutron transport kinetic tRAPID algorithm and its validation

The Real-time Analysis for Particle-transport and In-situ Detection (RAPID) Code System, based on the Multi-stage Response-function Transport (MRT) methodology, allows for real-time simulation of nuclear systems based on 3-D continuous-energy particle transport. RAPID's steady-state (criticality) neutron transport algorithm is based on the Fission Matrix (FM) method, and has been extensively verified and validated against computational benchmarks and experiments. This paper introduces the novel 3-D time-dependent transport algorithm that has been implemented into the code, tRAPID, and its validation using the JSI TRIGA Mark-II reactor. tRAPID accurately and efficiently calculates neutron kinetics parameters (such as β{sub eff}, l{sub eff} , Λ, α{sub Rossi}) and 3-D time-dependent neutron fission source distribution and neutron importances for both prompt and delayed neutrons. tRAPID is used to simulate a rod insertion experiment performed at the JSI TRIGA Mark-II reactor, during which signals from four fission chambers at four different locations in the core were collected. The results demonstrate how tRAPID is capable of calculating detailed and accurate results with only a minimal use computational resources and time.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Energy dataset of Frontier supercomputer for waste heat recovery

The Hewlett Packard Enterprise–Cray EX Frontier is the world’s first and fastest exascale supercomputer, hosted at the Oak Ridge Leadership Computing Facility in Tennessee, United States. Frontier is a significant electricity consumer, drawing 8–30 MW; this massive energy demand produces significant waste heat, requiring extensive cooling measures. Although harnessing this waste heat for campus heating is a sustainability goal at Oak Ridge National Laboratory (ORNL), the 30 °C–38 °C waste heat temperature poses compatibility issues with standard HVAC systems. Heat pump systems, prevalent in residential settings and some industries, can efficiently upgrade low-quality heat to usable energy for buildings. Thus, heat pump technology powered by renewable electricity offers an efficient, cost-effective solution for substantial waste heat recovery. However, a major challenge is the absence of benchmark data on high-performance computing (HPC) heat generation and waste heat profiles. This paper reports power demand and waste heat measurements from an ORNL HPC data centre, aiming to guide future research on optimizing waste heat recovery in large-scale data centres, especially those of HPC calibre.

97 MATHEMATICS AND COMPUTING↗

Using scalable computer vision to automate high-throughput semiconductor characterization

Abstract High-throughput materials synthesis methods, crucial for discovering novel functional materials, face a bottleneck in property characterization. These high-throughput synthesis tools produce 10 4 samples per hour using ink-based deposition while most characterization methods are either slow (conventional rates of 10 1 samples per hour) or rigid (e.g., designed for standard thin films), resulting in a bottleneck. To address this, we propose automated characterization (autocharacterization) tools that leverage adaptive computer vision for an 85x faster throughput compared to non-automated workflows. Our tools include a generalizable composition mapping tool and two scalable autocharacterization algorithms that: (1) autonomously compute the band gaps of 200 compositions in 6 minutes, and (2) autonomously compute the environmental stability of 200 compositions in 20 minutes, achieving 98.5% and 96.9% accuracy, respectively, when benchmarked against domain expert manual evaluation. These tools, demonstrated on the formamidinium (FA) and methylammonium (MA) mixed-cation perovskite system FA 1−x MA x PbI 3 , 0 ≤ x ≤ 1, significantly accelerate the characterization process, synchronizing it closer to the rate of high-throughput synthesis.

Science & Technology - Other Topics↗

Exploring Hilbert space on a budget: Novel benchmark set and performance metric for testing electronic structure methods in the regime of strong correlation

This work explores the ability of classical electronic structure methods to efficiently represent (compress) the information content of full configuration interaction (FCI) wave functions. We introduce a benchmark set of four hydrogen model systems of different dimensionalities and distinctive electronic structures: a 1D chain, a 1D ring, a 2D triangular lattice, and a 3D close-packed pyramid. To assess the ability of a computational method to produce accurate and compact wave functions, we introduce the accuracy volume, a metric that measures the number of variational parameters necessary to achieve a target energy error. Using this metric and the hydrogen models, we examine the performance of three classical deterministic methods: (i) selected configuration interaction (sCI) realized both via an a posteriori (ap-sCI) and variational selection of the most important determinants, (ii) an a posteriori singular value decomposition (SVD) of the FCI tensor (SVD-FCI), and (iii) the matrix product state representation obtained via the density matrix renormalization group (DMRG). We find that the DMRG generally gives the most efficient wave function representation for all systems, particularly in the 1D chain with a localized basis. For the 2D and 3D systems, all methods (except DMRG) perform best with a delocalized basis, and the efficiency of sCI and SVD-FCI is closer to that of DMRG. For larger analogs of the models, the DMRG consistently requires the fewest parameters but still scales exponentially in 2D and 3D systems, and the performance of SVD-FCI is essentially equivalent to that of ap-sCI.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Deep potential generation scheme and simulation protocol for the Li 10 GeP 2 S 12 -type superionic conductors

We report solid-state electrolyte materials with superior lithium ionic conductivities are vital to the next-generation Li-ion batteries. Molecular dynamics could provide atomic scale information to understand the diffusion process of Li-ion in these superionic conductor materials. Here, we implement the deep potential generator to set up an efficient protocol to automatically generate interatomic potentials for Li 10 GeP 2 S 12 -type solid-state electrolyte materials (Li 10 GeP 2 S 12 , Li 10 SiP 2 S 12 , and Li 10 SnP 2 S 12 ). The reliability and accuracy of the fast interatomic potentials are validated. With the potentials, we extend the simulation of the diffusion process to a wide temperature range (300 K–1000 K) and systems with large size (~1000 atoms). Important technical aspects such as the statistical error and size effect are carefully investigated, and benchmark tests including the effect of density functional, thermal expansion, and configurational disorder are performed. The computed data that consider these factors agree well with the experimental results, and we find that the three structures show different behaviors with respect to configurational disorder. Our work paves the way for further research on computation screening of solid-state electrolyte materials.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Accelerated kinetic model for global macro stability studies of high-beta fusion reactors

The field reversed configuration (FRC), such as studied in the C-2W experiment at TAE Technologies, is an attractive candidate for realizing a nuclear fusion reactor. In an FRC, kinetic ion effects play the majority role in macroscopic stability, which allows global stability studies to make use of fluid-kinetic hybrid (also referred to as Ohm's law) models wherein ions are treated kinetically while electrons are treated as a fluid. The development and validation of such a hybrid particle-in-cell algorithm in the Exascale Computing Project code WarpX are reported here. Implementation of this model in the WarpX framework benefits from the numerical efficiency of WarpX as well as its scalability on large HPC systems and portability to different architectures. Performance benchmarks of the new algorithm for large, 3-dimensional, full device simulations from the Perlmutter supercomputer are presented. Results of a series of FRC simulations are discussed in which the impact of two-fluid effects on the tilt-mode growth rate was studied. It was observed that, in agreement with previous Hall-MHD studies, two-fluid effects have a stabilizing impact on the tilt mode.

Physics↗

Unprecedented cloud resolution in a GPU-enabled full-physics atmospheric climate simulation on OLCF’s summit supercomputer

Clouds represent a key uncertainty in future climate projection. While explicit cloud resolution remains beyond our computational grasp for global climate, we can incorporate important cloud effects through a computational middle ground called the Multi-scale Modeling Framework (MMF), also known as Super Parameterization. This algorithmic approach embeds high-resolution Cloud Resolving Models (CRMs) to represent moist convective processes within each grid column in a Global Climate Model (GCM). The MMF code requires no parallel data transfers and provides a self-contained target for acceleration. This study investigates the performance of the Energy Exascale Earth System Model-MMF (E3SM-MMF) code on the OLCF Summit supercomputer at an unprecedented scale of simulation. Hundreds of kernels in the roughly 10K lines of code in the E3SM-MMF CRM were ported to GPUs with OpenACC directives. A high-resolution benchmark using 4600 nodes on Summit demonstrates the computational capability of the GPU-enabled E3SM-MMF code in a full physics climate simulation.

58 GEOSCIENCES↗

Quantifying uncertainties due to irreducible three-body forces in deuteron-nucleus reactions

Deuteron-induced nuclear reactions are an essential tool for probing the structure of nuclei as well as astrophysical information such as (n, γ) cross sections. The deuteron-nucleus system is typically described within a Faddeev three-body model consisting of a neutron (n), a proton (p), and the target nucleus (A) interacting through pairwise phenomenological potentials. While Faddeev techniques enable the exact description of the three-body dynamics, their predictive power is limited in part by the omission of irreducible neutron-proton-nucleus three-body force (n–p–A 3BF). Here, our goal is to quantify systematic uncertainties stemming from the reduction of deuteron-nucleus (d + A) dynamics to a picture of three pointlike nuclear clusters interacting via pairwise nucleon-nucleus forces, using as testing grounds d + α scattering and the 6 Li ground state. We particularly focus on quantifying uncertainties arising from the full antisymmetrization of the (A + 2)-body system with the target nucleus fixed in its ground state. We adopt the ab initio no-core shell model coupled with the resonating group method (NCSM/RGM) to compute microscopic n–α and p–α interactions, and use them in a three-body description of the d + α system by means of momentum-space Faddeev-type equations. Simultaneously, we also carry out ab initio calculations of d + α scattering and 6 Li ground state by means of six-body NCSM/RGM calculations to serve as a benchmark for the three-body model predictions given by the Faddeev calculations. By comparing the Faddeev and NCSM/RGM results, we show that the irreducible n–p–α 3BF has a non-negligible effect on bound state and scattering observables alike. Specifically, the Faddeev approach yields a 6 Li ground state that is approximately 600 keV shallower than the one obtained with the NCSM/RGM. Additionally, the Faddeev calculations for d + α scattering yield a 3 + resonance that is located approximately 400 keV higher in energy compared to the NCSM/RGM result. The shape of the d + α angular distributions computed using the two approaches also differ, owing to the discrepancy in the predictions of the 3 + resonance energy. The Faddeev three-body model predictions for d + α scattering and 6 Li using microscopic n–α and p–α potentials differ from those computed microscopically with the NCSM/RGM. These discrepancies are due to the n–p–α 3BF, which arises from two-nucleon exchange terms in the microscopic d–α interaction and are not accounted for in the three-body model Faddeev calculations. This study lays the foundation for future parametrizations of the 3BF due to Pauli exclusion principle effects in improved three-body calculations of deuteron-induced reactions.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Active space selection with self-healing diffusion Monte Carlo algorithms for periodic solids

Multideterminant Diffusion Monte Carlo (DMC) displays improved accuracy over single determinant DMC. Self-Healing Diffusion Monte Carlo (SHDMC) is a DMC based method that iteratively improves a multideterminant trial wavefunction. Although configuration interaction or complete active space (CAS) methods are very accurate and computationally feasible for many systems, they are not optimal for application to solids. SHDMC is accurate and designed for application to solids, so developing SHDMC based active space selection algorithms is a worthy endeavor. Here, we present and compare active space selection algorithms that are designed for use in conjunction with SHDMC, without relying on external approaches. For benchmarking, we calculated the ground state energy of a small unit cell of graphene and compared the results with a complete basis set extrapolated selected CI and a reference SHDMC trajectory. We found that systematically expanding the active space using an “auto-branching” algorithm optimally balances accuracy with computational practicality. To the best of our knowledge, this is the first work that demonstrates completely self-contained DMC-based active space selection algorithms that do not depend on external methods for determinant selection.

Spanedda, Nicole [ORNL]↗

Summer Internship Report: ARA2 Benchmarking

Over the past decade, the RISC-V Instruction Set Architecture (ISA) has emerged as a significant player in both academic and industrial processor design due to its open-source nature, modular extension system, and versatility across domains ranging from microcontrollers to high-performance computing (HPC). One of its most important recent advancements is the RISC-V Vector Extension (RVV), which enables explicit data-level parallelism through vector registers and vectorized instructions. Unlike traditional SIMD (Single Instruction, Multiple Data) architectures that fix vector lengths at design time, RVV uses the concept of VLEN (vector register length) as a hardware-independent parameter and allows software to adapt dynamically to the available vector width. This flexible approach ensures portability across implementations while enabling scalable performance. The ARA2 core is a parameterizable RISC-V vector processor developed at the Integrated Systems Lab at ETH Zürich and the University of Bologna. Designed as a tightly-coupled accelerator to a scalar RISC-V core, ARA2 implements the RVV 1.0 specification and offers tunable architectural parameters such as the number of vector lanes, VLEN, and cache sizes.

97 MATHEMATICS AND COMPUTING↗

Theoretical framework for new magnetic materials for quantum computing and information storage. Final report for the Award No. DE-SC0018910

The focus of this grant was on molecular magnetic materials for information storage and quantum computing. We have been developing robust, first-principle methods for computing relevant electronic and magnetic properties of molecular building blocks (SMMs) of novel magnetic materials and quantum computers. These tools enable theoretical modeling of SMMs’ behavior, facilitating the interpretation of experimental studies and aiding the design of novel magnetic materials. Our strategy is based on the spin-flip (SF) approach, which extends the hierarchy of black-box single-reference methods to strongly correlated systems. Specifically, we developed general scalable algorithms and computer codes for calculating molecular properties, with an emphasis on spin-related properties, such as zero-field splittings, hyperfine couplings, and g-tensors. While our primary focus was on SF wave functions and SF-TDDFT, the underlying theory and computer codes were formulated using reduced density matrices, such that these tools are applicable to a broader class of methods. To extend the scope of applicability of wave-function-based SF methods to larger systems, we developed reduced-scaling approaches for the equation-of-motion coupled-cluster (EOM-CC) methods and continue developing libtensor (our open-source general tensor contraction library for many-body methods). We carried out extensive benchmarks and also carried out several applications.

36 MATERIALS SCIENCE↗