Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “computer system benchmarking”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19

The NAS parallel benchmarks

A new set of benchmarks was developed for the performance evaluation of highly parallel supercomputers. These benchmarks consist of a set of kernels, the 'Parallel Kernels,' and a simulated application benchmark. Together they mimic the computation and data movement characteristics of large scale computational fluid dynamics (CFD) applications. The principal distinguishing feature of these benchmarks is their 'pencil and paper' specification - all details of these benchmarks are specified only algorithmically. In this way many of the difficulties associated with conventional benchmarking approaches on highly parallel systems are avoided.

Bailey, David↗

Model Checking Degrees of Belief in a System of Agents

Reasoning about degrees of belief has been investigated in the past by a number of authors and has a number of practical applications in real life. In this paper we present a unified framework to model and verify degrees of belief in a system of agents. In particular, we describe an extension of the temporal-epistemic logic CTLK and we introduce a semantics based on interpreted systems for this extension. In this way, degrees of beliefs do not need to be provided externally, but can be derived automatically from the possible executions of the system, thereby providing a computationally grounded formalism. We leverage the semantics to (a) construct a model checking algorithm, (b) investigate its complexity, (c) provide a Java implementation of the model checking algorithm, and (d) evaluate our approach using the standard benchmark of the dining cryptographers. Finally, we provide a detailed case study: using our framework and our implementation, we assess and verify the situational awareness of the pilot of Air France 447 flying in off-nominal conditions.

MAS Verification↗

Sub-microsecond Transformers for Jet Tagging on FPGAs

We present the first sub-microsecond transformer implementation on an FPGA achieving competitive performance for state-of-the-art high-energy physics benchmarks. Transformers have shown exceptional performance on multiple tasks in modern machine learning applications, including jet tagging at the CERN Large Hadron Collider (LHC). However, their computational complexity prohibits use in real-time applications, such as the hardware trigger system of the collider experiments up until now. In this work, we demonstrate the first application of transformers for jet tagging on FPGAs, achieving $\mathcal{O}(100)$ nanosecond latency with superior performance compared to alternative baseline models. We leverage high-granularity quantization and distributed arithmetic optimization to fit the entire transformer model on a single FPGA, achieving the required throughput and latency. Furthermore, we add multi-head attention and linear attention support to hls4ml, making our work accessible to the broader fast machine learning community. This work advances the next-generation trigger systems for the High Luminosity LHC, enabling the use of transformers for real-time applications in high-energy physics and beyond.

Laatu, Lauri [Imperial Coll., London]↗

Validation of Numerical Tools for Calculating Reactivity Feedback in Sodium Fast Reactors Using SEFOR Experimental Data

The Southwest Experimental Fast Oxide Reactor (SEFOR) was an experimental sodium-cooled fast breeder reactor operated from 1969 to 1972 with experiments designed to measure Doppler reactivity feedback in a wide temperature range from around 350 °F to temperatures approaching the melting point of mixed oxide fuel of around 5000 °F, providing valuable data for code validations. Co-supported by the Department of Energy (DOE) Fast Reactor Program (FRP) and the DOE Nuclear Energy Advanced Modeling and Simulation (NEAMS) program, the SEFOR benchmark project focused on using the experimental data to validate numerical tools that are used in industry and academia to design and license sodium-cooled fast reactors (SFRs). By the end of FY-25, substantial progress was achieved in the SEFOR benchmark study. A variety of numerical tools commonly used for modeling SFRs were applied to develop models for SEFOR core configurations I-D, I-E, I-I, and I-J. These included Monte Carlo codes such as MCNP, Serpent, and Shift; deterministic codes such as the legacy Argonne Reactor Computation (ARC) suite and the high-fidelity NEAMS code Griffin; and the system analysis code SAS4A/SASSYS-1 (SAS). Using these models, both SEFOR zero-power experiments and power-ascending tests were successfully simulated. Comparisons were performed against experimental measurements of core criticalities, reflector worth, kinetics parameters (Λ/βeff), isothermal reactivity feedback (from 350 °F to 760 °F at zero power), and power-ascending reactivity feedback (as power increased from 0.4 MW to 17 MW). In general, these comparisons demonstrated very good agreement between numerical results and experimental data. In Fiscal Year 26 (FY-26), the SEFOR benchmark project will continue to address the modeling issues identified in FY-25. Effort will focus on the simulation of reactivity insertion transients in SEFOR core II using the ARC/SAS model. Future work will also focus on incorporating BISON into the SEFOR core modeling process to enable the first Multiphysics simulations of the isothermal tests based on the MOOSE framework.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Issues in ATM Support of High-Performance, Geographically Distributed Computing

This report experimentally assesses the effect of the underlying network in a cluster-based computing environment. The assessment is quantified by application-level benchmarking, process-level communication, and network file input/output. Two testbeds were considered, one small cluster of Sun workstations and another large cluster composed of 32 high-end IBM RS/6000 platforms. The clusters had Ethernet, fiber distributed data interface (FDDI), Fibre Channel, and asynchronous transfer mode (ATM) network interface cards installed, providing the same processors and operating system for the entire suite of experiments. The primary goal of this report is to assess the suitability of an ATM-based, local-area network to support interprocess communication and remote file input/output systems for distributed computing.

Claus, Russell W.↗

Experiences using OpenMP based on Computer Directed Software DSM on a PC Cluster

In this work we report on our experiences running OpenMP programs on a commodity cluster of PCs running a software distributed shared memory (DSM) system. We describe our test environment and report on the performance of a subset of the NAS Parallel Benchmarks that have been automaticaly parallelized for OpenMP. We compare the performance of the OpenMP implementations with that of their message passing counterparts and discuss performance differences.

Hess, Matthias↗

Chimbuko: A Workflow-Level Scalable Performance Trace Analysis Tool

ABSTRACT Due to the sheer volume of data it is typically impractical to analyze the detailed performance of an HPC application running at-scale. While conventional small-scale benchmarking and scaling studies are often sufficient for simple applications, many modern workflow-based applications couple multiple elements with competing resource demands and complex inter-communication patterns for which performance cannot easily be studied in isolation and at small scale. This work discusses Chimbuko, a performance analysis framework that provides real-time, in situ anomaly detection. By focusing specifically on performance anomalies and their origin (aka provenance), data volumes are dramatically reduced without losing necessary details. To the best of our knowledge, Chimbuko is the first online, distributed, and scalable workflow-level performance trace analysis framework. We demonstrate the tool's usefulness on Oak Ridge National Laboratory's Summit system.

97 MATHEMATICS AND COMPUTING↗

Efficient Berry phase calculation via adaptive variational quantum computing approach

We present an adaptive variational quantum algorithm to estimate the Berry phase accumulated by a nondegenerate ground state under cyclic, adiabatic evolution of a time-dependent Hamiltonian. Our method leverages cyclic adiabatic evolution of the Hamiltonian and employs adaptive variational quantum algorithms for state preparation and evolution, optimizing circuit efficiency while maintaining high accuracy. We benchmark our approach on dimerized Fermi–Hubbard chains with four sites, demonstrating precise Berry phase simulations in both noninteracting and interacting regimes. Our results show that circuit depths reach up to 106 layers for noninteracting systems and increase to 279 layers for interacting systems due to added complexity. In addition, we demonstrate the robustness of our scheme across a wide range of parameters governing adiabatic evolution and variational algorithms. These findings highlight the potential of adaptive variational quantum algorithms for advancing quantum simulations of topological materials and computing geometric phases in strongly correlated systems.

Mootz, Martin [Ames Laboratory (AMES), Ames, IA (U↗

ENDF/B-VIII.0 Augmented Covariance Data The first iteration [Slides]

Nuclear data are necessary for reliable modeling and simulation of the next generation of nuclear reactors. However, the variation in the ratio of the computed to experimental values (C/E) for certain types of nuclear systems is much less than predicted by evaluated nuclear data file (ENDF)/B covariances. Figure 1 provides an example of a set of metal-plutonium-fueled, fast-spectrum (PU-MET-FAST) integral experiments from the International Criticality Safety Benchmark Evaluation Project (ICSBEP) in the Oak Ridge National Laboratory (ORNL) VALID database. The variation in the C/E values, shown with one standard deviation error bars, is 100% covered by both the SCALE and ENDB/VIII.0 covariance data. This is due to the comparisons to integral data that are essential during the evaluation process. However, the ENDF evaluations represent uncertainties and correlations in differential data only; they do not reflect the impact of the comparison to integral data in the covariance evaluations.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

FINETUNA: fine-tuning accelerated molecular simulations

Abstract Progress towards the energy breakthroughs needed to combat climate change can be significantly accelerated through the efficient simulation of atomistic systems. However, simulation techniques based on first principles, such as density functional theory (DFT), are limited in their practical use due to their high computational expense. Machine learning approaches have the potential to approximate DFT in a computationally efficient manner, which could dramatically increase the impact of computational simulations on real-world problems. However, they are limited by their accuracy and the cost of generating labeled data. Here, we present an online active learning framework for accelerating the simulation of atomic systems efficiently and accurately by incorporating prior physical information learned by large-scale pre-trained graph neural network models from the Open Catalyst Project. Accelerating these simulations enables useful data to be generated more cheaply, allowing better models to be trained and more atomistic systems to be screened. We also present a method of comparing local optimization techniques on the basis of both their speed and accuracy. Experiments on 30 benchmark adsorbate-catalyst systems show that our method of transfer learning to incorporate prior information from pre-trained models accelerates simulations by reducing the number of DFT calculations by 91%, while meeting an accuracy threshold of 0.02 eV 93% of the time. Finally, we demonstrate a technique for leveraging the interactive functionality built in to Vienna ab initio Simulation Package (VASP) to efficiently compute single point calculations within our online active learning framework without the significant startup costs. This allows VASP to work in tandem with our framework while requiring 75% fewer self-consistent cycles than conventional single point calculations. The online active learning implementation, and examples using the VASP interactive code, are available in the open source FINETUNA package on Github.

97 MATHEMATICS AND COMPUTING↗

Benchmark Comparison of Cloud Analytics Methods Applied to Earth Observations

Cloud computing has the potential to bring high performance computing capabilities to the average science researcher. However, in order to take full advantage of cloud capabilities, the science data used in the analysis must often be reorganized. This typically involves sharding the data across multiple nodes to enable relatively fine-grained parallelism. This can be either via cloud-based file systems or cloud-enabled databases such as Cassandra, Rasdaman or SciDB. Since storing an extra copy of data leads to increased cost and data management complexity, NASA is interested in determining the benefits and costs of various cloud analytics methods for real Earth Observation cases. Accordingly, NASA's Earth Science Technology Office and Earth Science Data and Information Systems project have teamed with cloud analytics practitioners to run a benchmark comparison on cloud analytics methods using the same input data and analysis algorithms. We have particularly looked at analysis algorithms that work over long time series, because these are particularly intractable for many Earth Observation datasets which typically store data with one or just a few time steps per file. This post will present side-by-side cost and performance results for several common Earth observation analysis operations.

science data management↗

Relativistic resolution-of-the-identity with Cholesky integral decomposition

In this study, we present an efficient integral decomposition approach called the restricted-kinetic-balance resolution-of-the-identity (RKB-RI) algorithm, which utilizes a tunable RI method based on the Cholesky integral decomposition for in-core relativistic quantum chemistry calculations. The RKB-RI algorithm incorporates the restricted-kinetic-balance condition and offers a versatile framework for accurate computations. Notably, the Cholesky integral decomposition is employed not only to approximate symmetric large-component electron repulsion integrals but also those involving small-component basis functions. In addition to comprehensive error analysis, we investigate crucial conditions, such as the kinetic balance condition and variational stability, which underlie the applicability of Dirac relativistic electronic structure theory. Here we compare the computational cost of the RKB-RI approach with the full in-core method to assess its efficiency. To evaluate the accuracy and reliability of the RKB-RI method proposed in this work, we employ actinyl oxides as benchmark systems, leveraging their properties for validation purposes. This investigation provides valuable insights into the capabilities and performance of the RKB-RI algorithm and establishes its potential as a powerful tool in the field of relativistic quantum chemistry.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Fast and scalable quantum Monte Carlo simulations of electron-phonon models

We introduce methodologies for highly scalable quantum Monte Carlo simulations of electron-phonon models, and report benchmark results for the Holstein model on the square lattice. The determinant quantum Monte Carlo (DQMC) method is a widely used tool for simulating simple electron-phonon models at finite temperatures, but incurs a computational cost that scales cubically with system size. Alternatively, near-linear scaling with system size can be achieved with the hybrid Monte Carlo (HMC) method and an integral representation of the Fermion determinant. Here, we introduce a collection of methodologies that make such simulations even faster. To combat "stiffness" arising from the bosonic action, we review how Fourier acceleration can be combined with time-step splitting. To overcome phonon sampling barriers associated with strongly-bound bipolaron formation, we design global Monte Carlo updates that approximately respect particle-hole symmetry. To accelerate the iterative linear solver, we introduce a preconditioner that becomes exact in the adiabatic limit of infinite atomic mass. Finally, we demonstrate how stochastic measurements can be accelerated using fast Fourier transforms. Here, these methods are all complementary and, combined, may produce multiple orders of magnitude speedup, depending on model details.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Performance and the NAS Parallel Benchmarks

This talk will describe the NAS (National Aerospace Standards) Parallel Benchmarks, which are now widely cited in the high performance computing field as a measure of sustained performance on realistic scientific applications. The latest performance results will be included. It will be shown that significant progress has been made by several systems during the past year or so, with sustained performance on a par with the best conventional systems, and with performance per dollar significantly exceeding the conventional systems. This talk will also describe many of the pitfalls of performance reporting, and will give advice on how to avoid such pitfalls. The overall state of the field of high performance computing will also be discussed.

Bailey, David H.↗

Acceleration of the particle-in-cell code Osiris with graphics processing units

Fully relativistic particle-in-cell (PIC) simulations are crucial for advancing our knowledge of plasma physics. Modern supercomputers based on graphics processing units (GPUs) offer the potential to perform PIC simulations of unprecedented scale, but require robust and feature-rich codes that can fully leverage their computational resources. In this work, this demand is addressed by adding GPU acceleration to the PIC code Osiris. An overview of the algorithm, which features a CUDA extension to the underlying Fortran architecture, is given. Detailed performance benchmarks for thermal plasmas are presented, which demonstrate excellent weak scaling on NERSC's Perlmutter supercomputer and high levels of absolute performance. The robustness of the code to model a variety of physical systems is demonstrated via simulations of Weibel filamentation and laser-wakefield acceleration run with dynamic load balancing. Finally, measurements and analysis of energy consumption are provided that indicate that the GPU algorithm is up to ~14 times faster and ~7 times more energy efficient than the optimized CPU algorithm on a node-to-node basis. The described development addresses the PIC simulation community's computational demands both by contributing a robust and performant GPU-accelerated PIC code and by providing insight into efficient use of GPU hardware.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Multireference Equation-of-Motion Driven Similarity Renormalization Group: Theoretical Foundations and Applications to Ionized States

We present a formulation and implementation of an equation-of-motion (EOM) extension of the multireference driven similarity renormalization group (MR-DSRG) formalism for ionization potentials (IP-EOM-DSRG). The IP-EOM-DSRG formalism results in a Hermitian generalized eigenvalue problem, delivering accurate ionization potentials for strongly correlated systems. The EOM step scales as O(N 5 ) with the basis set size N, allowing for efficient calculation of spectroscopic properties, such as transition energies and intensities. The IP-EOM-DSRG formalism is combined with three truncation schemes of the parent MR-DSRG theory: an iterative nonperturbative method with up to two-body excitations [MR-LDSRG(2)] and second- and third-order perturbative approximations [DSRG-MRPT2/3]. We benchmark these variants by computing (1) the vertical valence ionization potentials of a series of small molecules at both equilibrium and stretched geometries; (2) the spectroscopic constants of several low-lying electronic states of the OH, CN, N 2 + , and CO + radicals; and (3) the binding curves of low-lying electronic states of the CN radical. A comparison with experimental data and theoretical results shows that all three IP-EOM-DSRG methods accurately reproduce the vertical ionization potentials and spectroscopic constants of these systems. Notably, the DSRG-MRPT3 and MR-LDSRG(2) versions outperform several state-of-the-art multireference methods of comparable or higher cost.

Hamiltonians↗

Accurate numerical simulations of open quantum systems using spectral tensor trains

Decoherence between qubits is a major bottleneck in quantum computations. Decoherence results from intrinsic quantum and thermal fluctuations as well as noise in the external fields that perform the measurement and preparation processes. With prescribed colored noise spectra for intrinsic and extrinsic noise, we present a numerical method, Quantum Accelerated Stochastic Propagator Evaluation (Q-ASPEN), to solve the time-dependent noise-averaged reduced density matrix in the presence of intrinsic and extrinsic noise. Q-ASPEN is arbitrarily accurate and can be applied to provide estimates for the resources needed to error-correct quantum computations. We employ spectral tensor trains, which combine the advantages of tensor networks and pseudospectral methods, as a variational ansatz to the quantum relaxation problem and optimize the ansatz using methods typically used to train neural networks. Here, the spectral tensor trains in Q-ASPEN make accurate calculations with tens of quantum levels feasible. We present benchmarks for Q-ASPEN on the spin-boson model in the presence of intrinsic noise and on a quantum chain of up to 32 sites in the presence of extrinsic noise. In our benchmark, the memory cost of Q-ASPEN scales as a low-order polynomial in the size of the system once the number of system states surpasses the number of basis functions used in the spectral expansion.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

The Gutzwiller conjugate gradient minimization method for correlated electron systems

In this report we review our recent work on the Gutzwiller conjugate gradient minimization method, an ab initio approach developed for correlated electron systems. The complete formalism has been outlined that allows for a systematic understanding of the method, followed by a discussion of benchmark studies of dimers, one- and two-dimensional single-band Hubbard models. In the end, we present some preliminary results of multi-band Hubbard models and large-basis calculations of F 2 to illustrate our efforts to further reduce the computational complexity.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗