Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “computer system benchmarking”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

TRINIDI (Time-of-Flight Resonance Imaging with Neutrons for Isotopic Density Inference)

This software is an open-source Python library that provides tools for processing hyperspectral neutron time-of-flight radiography data. This type of data allows material decomposed reconstructions to be generated with the use of material characteristic spectral responses and the algorithms provided in this code library. The software library will contain tools for pre-processing the neutron measurement data, estimating measurement system parameters, reconstructing material decomposed radiographs, and computing material decomposed computed tomography (CT). Furthermore, it will have capability to generate and process simulated neutron time-of-flight data with the goal of benchmarking and demonstrating the tools that are provided. The software will include thorough documentation and application examples.

Balke, Thilo↗

SpaceCubeX: A Framework for Evaluating Hybrid Multi-Core CPU FPGA DSP Architectures

The SpaceCubeX project is motivated by the need for high performance, modular, and scalable on-board processing to help scientists answer critical 21st century questions about global climate change, air quality, ocean health, and ecosystem dynamics, while adding new capabilities such as low-latency data products for extreme event warnings. These goals translate into on-board processing throughput requirements that are on the order of 100-1,000 more than those of previous Earth Science missions for standard processing, compression, storage, and downlink operations. To study possible future architectures to achieve these performance requirements, the SpaceCubeX project provides an evolvable testbed and framework that enables a focused design space exploration of candidate hybrid CPU/FPGA/DSP processing architectures. The framework includes ArchGen, an architecture generator tool populated with candidate architecture components, performance models, and IP cores, that allows an end user to specify the type, number, and connectivity of a hybrid architecture. The framework requires minimal extensions to integrate new processors, such as the anticipated High Performance Spaceflight Computer (HPSC), reducing time to initiate benchmarking by months. To evaluate the framework, we leverage a wide suite of high performance embedded computing benchmarks and Earth science scenarios to ensure robust architecture characterization. We report on our projects Year 1 efforts and demonstrate the capabilities across four simulation testbed models, a baseline SpaceCube 2.0 system, a dual ARM A9 processor system, a hybrid quad ARM A53 and FPGA system, and a hybrid quad ARM A53 and DSP system.

Hybrid Flight Architectures↗

DFT-based QM/MM with Particle-Mesh Ewald for Direct, Long-Range Electrostatic Embedding

In this work, we present a DFT-based, QM/MM implementation with long-range electrostatic embedding achieved by direct real-space integration of the particle mesh Ewald (PME) computed electrostatic potential. The key transformation is the interpolation of the electrostatic potential from the PME grid to the DFT quadrature grid, from which integrals are easily evaluated utilizing standard DFT machinery. We provide benchmarks of the numerical accuracy with choice of grid size and real-space corrections, and demonstrate that good convergence is achieved while introducing nominal computational overhead. Furthermore, the approach requires only small modification to existing software packages, as is demonstrated with our implementation in the OpenMM and Psi 4 software. After presenting convergence benchmarks, we evaluate the importance of long-range electrostatic embedding in three solute/solvent systems modeled with QM/MM. Water and BMIM/BF 4 ionic liquid were considered as "simple" and "complex" solvents respectively, with water and p-phenylenediamine (PPD) solute molecules treated at QM level of theory. While electrostatic embedding with standard real-space truncation may introduce negligible error for simple systems such as water solute in water solvent, errors become more significant when QM/MM is applied to complex solvents such as ionic liquids. An extreme example is the electrostatic embedding energy for oxidized PPD in BMIM/BF 4 for which real-space truncation produces severe error even at 2-3 nm cutoff distances. This latter example illustrates that utilization of QM/MM to compute redox potentials within concentrated electrolytes/ionic media requires carefully chosen long-range electrostatic embedding algorithms, with our presented algorithm providing a general and robust approach.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Accelerating resonant spectroscopy simulations using multishifted biconjugate gradient

Resonant spectroscopies, which involve intermediate states with finite lifetimes, provide important insights into collective excitations in quantum materials that are otherwise inaccessible. However, theoretical understanding in this area is often limited by the numerical challenges of solving Kramers-Heisenberg-type response functions for large-scale systems. To address this, we introduce a multishifted biconjugate gradient algorithm that exploits the shared structure of Krylov subspaces across spectra with varying incident energies, effectively reducing the computational complexity to that of linear spectroscopies. Both mathematical proofs and numerical benchmarks confirm that this algorithm substantially accelerates spectral simulations, achieving constant complexity independent of the number of incident energies, while ensuring accuracy and stability. This development provides a scalable, versatile framework for simulating advanced spectroscopies in quantum materials.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

VTK-m: Visualization for the Exascale Era and Beyond

A recent trend in modern high-performance computing is the increasing use of hybrid architectures, where the vast majority of performance comes from accelerators. Modern accelerators are based on Graphics Processing Units (GPU) that contain many low power cores that in their aggregate provides an extremely high computation rate. Current and future CPU processors are requiring more explicit parallelism as each successive version of the hardware packs in more cores, and technologies like hyperthreading and vector operations require even more parallel processing to leverage each core’s full potential. As an example, the Frontier supercomputer installed at Oak Ridge National Laboratories recently hit a record breaking 1.1 exaflops1 on the LINPACK HPC benchmark [Shoemaker 2022]. The system contains 37632 AMD MI250x GPUs which requires more than half a billion threads to keep the system fully utilized [Khizeran 2022].VTK-m is a toolkit of scientific visualization algorithms for these emerging processor architectures. VTK-m supports the fine-grained concurrency for data analysis and visualization algorithms required to drive extreme scale computing by providing abstract models for data and execution that can be applied to a variety of algorithms across many different processor architectures.

Bolstad, Mark↗

Comparison of secondary flows predicted by a viscous code and an inviscid code with experimental data for a turning duct

A comparison of the secondary flows computed by the viscous Kreskovsky-Briley-McDonald code and the inviscid Denton code with benchmark experimental data for turning duct is presented. The viscous code is a fully parabolized space-marching Navier-Stokes solver while the inviscid code is a time-marching Euler solver. The experimental data were collected by Taylor, Whitelaw, and Yianneskis with a laser Doppler velocimeter system in a 90 deg turning duct of square cross-section. The agreement between the viscous and inviscid computations was generally very good for the streamwise primary velocity and the radial secondary velocity, except at the walls, where slip conditions were specified for the inviscid code. The agreement between both the computations and the experimental data was not as close, especially at the 60.0 deg and 77.5 deg angular positions within the duct. This disagreement was attributed to incomplete modelling of the vortex development near the suction surface.

Schwab, J. R.↗

Comparison of secondary flows predicted by a viscous code and an inviscid code with experimental data for a turning duct

A comparison of the secondary flows computed by the viscous Kreskovsky-Briley-McDonald code and the inviscid Denton code with benchmark experimental data for turning duct is presented. The viscous code is a fully parabolized space-marching Navier-Stokes solver while the inviscid code is a time-marching Euler solver. The experimental data were collected by Taylor, Whitelaw, and Yianneskis with a laser Doppler velocimeter system in a 90 deg turning duct of square cross-section. The agreement between the viscous and inviscid computations was generally very good for the streamwise primary velocity and the radial secondary velocity, except at the walls, where slip conditions were specified for the inviscid code. The agreement between both the computations and the experimental data was not as close, especially at the 60.0 deg and 77.5 deg angular positions within the duct. This disagreement was attributed to incomplete modeling of the vortex development near the suction surface.

Schwab, J. R.↗

Iterative quantum optimization of spin glass problems with rapidly oscillating transverse fields

In this work, we introduce a new iterative quantum algorithm, called Iterative Symphonic Tunneling for Satisfiability problems (IST-SAT), which solves quantum spin glass optimization problems using high-frequency oscillating transverse fields. IST-SAT operates as a sequence of iterations, in which bitstrings returned from one iteration are used to set spin-dependent phases in oscillating transverse fields in the next iteration. Over several iterations, the novel mechanism of the algorithm steers the system toward the problem ground state. We benchmark IST-SAT on sets of hard MAX-3-XORSAT problem instances with exact state vector simulation, and report polynomial speedups over Trotterized adiabatic quantum computation and the best known semi-greedy classical algorithm. When IST-SAT is seeded with a sufficiently good initial approximation, the algorithm converges to exact solution(s) in a polynomial number of iterations. Our numerical results identify a critical Hamming radius, or quality of initial approximation, where the time-to-solution crosses from exponential to polynomial scaling in problem size. This work proposes IST-SAT a new quantum algorithm, which improves upon solutions obtained from initial classical or quantum optimization algorithms. The steering mechanism we introduce through IST-SAT presents a new path toward achieving quantum advantage in optimization.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Two-Level Sketching Alternating Anderson Acceleration for Complex Physics Applications

We present a novel two-level sketching extension of the Alternating Anderson–Picard (AAP) method for accelerating fixed-point iterations in challenging single- and multiphysics simulations governed by discretized PDEs. Our approach combines a static, physics-based projection that reduces the least-squares (LS) problem to the most informative field (e.g., via Schur-complement insight) with a dynamic, algebraic sketching stage driven by a backward stability analysis under Lipschitz continuity. We introduce inexpensive estimators for stability thresholds and cache-aware randomized selection strategies to balance computational cost against memory access overhead. The resulting algorithm solves reduced LS systems in place, minimizes memory footprints, and seamlessly alternates between low-cost Picard updates and Anderson mixing. Implemented in Julia, our two-level sketching AAP achieves up to 50% time-to-solution reductions compared to standard Anderson acceleration—without degrading convergence rates—on benchmark problems including Stokes, 𝑝-Laplacian, bidomain, and Navier–Stokes formulations at varying problem sizes. These results demonstrate the method’s robustness, scalability, and potential for integration into high-performance scientific computing frameworks. Our implementation is available open source in the AAP.jl library.

Barnafi, Nicolas [University of Chile, Santiago]↗

FFTX-IRIS: Towards Performance Portability and Heterogeneity for SPIRAL Generated Code

FFTX-IRIS is a dynamic system to efficiently utilize novel heterogeneous platforms. This system links two next-generation frameworks, FFTX and IRIS, to navigate the complexity of different hardware architectures. FFTX provides a runtime code generation framework for high-performance Fast Fourier Transform kernels. IRIS runtime provides portability and multi-device heterogeneity, allowing computation on any available compute resource. Together, FFTX-IRIS enables code generation, seamless portability, and performance without user involvement. We show the design of the FFTX-IRIS system along with an evaluation of various small FFT benchmarks. We also demonstrate multi-device heterogeneity of FFTX-IRIS with a larger stencil application.

Rao, Sanil↗

SCALE Input and Result Files Supporting SCALE Inventory and Reactivity Analysis of the gFHR

This dataset contains input and result files of computational simulations with the SCALE code system. The simulations cover radionuclide inventory and reactivity analyses of a fluoride salt-cooled high temperature pebble-bed reactor (PB-FHR), specifically the generic FHR benchmark. Users wanting to reproduce results from this dataset are required to obtain a license to the SCALE code system for which details on the distribution can be found here: https://www.ornl.gov/scale/releases

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Electronic spectroscopy of diatomic molecules

This article provides an overview of the principal computational approaches and their accuracy for the study of electronic spectroscopy of diatomic molecules. We include a number of examples from our work that illustrate the range of application. We show how full configuration interaction benchmark calculations were instrumental in improving the understanding of the computational requirements for obtaining accurate results for diatomic spectroscopy. With this understanding it is now possible to compute radiative lifetimes accurate to within 10% for systems involving first- and second-row atoms. We consider the determination of the infrared vibrational transition probabilities for the ground states of SiO and NO, based on a globally accurate dipole moment function. We show how we were able to assign the a(sup "5)II state of CO as the upper state in the recently observed emission bands of CO in an Ar matrix. We next discuss the assignment of the photoelectron detachment spectra of NO and the alkali oxide negative ions. We then present several examples illustrating the state-of-the-art in determining radiative lifetimes for valence-valence and valence-Rydberg transitions. We next compare the molecular spectroscopy of the valence isoelectronic B2, Al2, and AlB molecules. The final examples consider systems involving transition metal atoms, which illustrate the difficulty in describing states with different numbers of d electrons.

Partridge, Harry↗

Large-scale harmonic balance simulations with Krylov subspace and preconditioner recycling

The multi-harmonic balance method combined with numerical continuation provides an efficient framework to compute a family of time-periodic solutions, or response curves, for large-scale, nonlinear mechanical systems. The predictor and corrector steps repeatedly solve a sequence of linear systems that scale by the model size and number of harmonics in the assumed Fourier series approximation. In this paper, a novel Newton–Krylov iterative method is embedded within the multi-harmonic balance and continuation algorithm to efficiently compute the approximate solutions from the sequence of linear systems that arise during the prediction and correction steps. Further, the method recycles, or reuses, both the preconditioner and the Krylov subspace generated by previous linear systems in the solution sequence. A delayed frequency preconditioner refactorizes the preconditioner only when the performance of the iterative solver deteriorates. The GCRO-DR iterative solver recycles a subset of harmonic Ritz vectors to initialize the solution subspace for the next linear system in the sequence. The performance of the iterative solver is demonstrated on two exemplars with contact-type nonlinearities and benchmarked against a direct solver with traditional Newton–Raphson iterations.

97 MATHEMATICS AND COMPUTING↗

Multi-group Examination of Nickel-Reflected HEU System [Slides]

Researchers noted an unusually large bias for 8-in. nickel reflected HEU sphere in HMF-003 between the 252- group library and the CE library in SCALE 6.2.4. The bias was investigated by reviewing reactions that $k_{eff}$ is sensitive to using TSUNAMI and collecting reaction rate data using tools within SCALE. For this system, bias is primarily due to the elastic scattering in nickel. Researchers compared new multigroup structures in SCALE 6.3. For this system, reduction in CE-to-MG bias seen in the 302-group and further improved in the 1597-group structure. When researchers compared ENDF/B-VIII.0 in SCALE 6.3, library showed improved results with new nickel evaluation. However, the MG-to-CE bias can still be large as large biases can occur in any system. This example highlights the importance of validating results with measured systems.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Integrating quantum computing resources into scientific HPC ecosystems

Quantum Computing (QC) offers significant potential to enhance scientific discovery in fields such as quantum chemistry, optimization, and artificial intelligence. Yet QC faces challenges due to the noisy intermediate-scale quantum era’s inherent external noise issues. Here, this paper discusses the integration of QC as a computational accelerator within classical scientific high-performance computing (HPC) systems. By leveraging a broad spectrum of simulators and hardware technologies, we propose a hardware-agnostic framework for augmenting classical HPC with QC capabilities. Drawing on the HPC expertise of the Oak Ridge National Laboratory (ORNL) and the HPC lifecycle management of the Department of Energy (DOE), our approach focuses on the strategic incorporation of QC capabilities and acceleration into existing scientific HPC workflows. This includes detailed analyses, benchmarks, and code optimization driven by the needs of the DOE and ORNL missions. Our comprehensive framework integrates hardware, software, workflows, and user interfaces to foster a synergistic environment for quantum and classical computing research. This paper outlines plans to unlock new computational possibilities, driving forward scientific inquiry and innovation in a wide array of research domains.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Efficient solvers for hybridized three-field mixed finite element coupled poromechanics

We consider a mixed hybrid finite element formulation for coupled poromechanics. A stabilization strategy based on a macro-element approach is advanced to eliminate the spurious pressure modes appearing in undrained/incompressible conditions. The efficient solution of the stabilized mixed hybrid block system is addressed by developing a class of block triangular preconditioners based on a Schur-complement approximation strategy. Robustness, computational efficiency and scalability of the proposed approach are theoretically discussed and tested using challenging benchmark problems on massively parallel architectures.

42 ENGINEERING↗

Fast convolutional neural networks on FPGAs with hls4ml

We introduce an automated tool for deploying ultra low-latency, low-power deep neural networks with convolutional layers on field-programmable gate arrays (FPGAs). By extending the hls4ml library, we demonstrate an inference latency of 5 µs using convolutional architectures, targeting microsecond latency applications like those at the CERN Large Hadron Collider. Considering benchmark models trained on the Street View House Numbers Dataset, we demonstrate various methods for model compression in order to fit the computational constraints of a typical FPGA device used in trigger and data acquisition systems of particle detectors. In particular, we discuss pruning and quantization-aware training, and demonstrate how resource utilization can be significantly reduced with little to no loss in model accuracy. We show that the FPGA critical resource consumption can be reduced by 97% with zero loss in model accuracy, and by 99% when tolerating a 6% accuracy degradation.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Computer simulation of multigrid body dynamics and control

The objective is to set up and analyze benchmark problems on multibody dynamics and to verify the predictions of two multibody computer simulation codes. TREETOPS and DISCOS have been used to run three example problems - one degree-of-freedom spring mass dashpot system, an inverted pendulum system, and a triple pendulum. To study the dynamics and control interaction, an inverted planar pendulum with an external body force and a torsional control spring was modeled as a hinge connected two-rigid body system. TREETOPS and DISCOS affected the time history simulation of this problem. System state space variables and their time derivatives from two simulation codes were compared.

Swaminadham, M.↗