Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “computer system benchmarking”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Superfluid condensate fraction and pairing wave function of the unitary Fermi gas

The unitary Fermi gas is a many-body system of two-component fermions with zero-range interactions tuned to infinite scattering length. Despite much activity and interest in unitary Fermi gas and its universal properties, there have been great difficulties in performing accurate calculations of the superfluid condensate fraction and pairing wave function. In this paper, we present auxiliary-field lattice Monte Carlo simulations using a lattice interaction which accelerates the approach to the continuum limit, thereby allowing for robust calculations of these difficult observables. As a benchmark test, we compute the ground-state energy of 33 spin-up and 33 spin-down particles. As a fraction of the free Fermi gas energy $E_{\text{FG}}$, we find $E_0/E_{\text{FG}}$ = 0.369(2), 0.372(2), using two different definitions of the finite-system energy ratio, in agreement with the latest theoretical and experimental results. We then determine the condensate fraction by measuring off-diagonal long-range order in the two-body density matrix. We find that the fraction of condensed pairs is α = 0.43(2). Further, we also extract the pairing wave function and find the pair correlation length to be $ζ_pk_F$ = 1.8(3)ℏ, where $k_F$ is the Fermi momentum. Provided that the simulations can be performed without severe sign oscillations, the methods we present here can be applied to superfluid neutron matter as well as more exotic $\textit{P}$-wave and $\textit{D}$-wave superfluids.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Formation of Benzene and Naphthalene through Cyclopentadienyl-Mediated Radical–Radical Reactions

Resonantly stabilized free radicals (RSFRs) have been contemplated as fundamental molecular building blocks and reactive intermediates in molecular mass growth processes leading to polycyclic aromatic hydrocarbons (PAHs) and carbonaceous nanoparticles on Earth and in deep space. Here, by combining molecular beams and computational fluid dynamics simulations, we provide compelling evidence on the formation of benzene via the cyclopentadienyl-methyl reaction and of naphthalene through the cyclopentadienyl self-reaction, respectively. These systems offer benchmarks for the conversion of a five-membered ring to the 6π-aromatic (benzene) and the generation of the simplest 10π-PAH (naphthalene) at elevated temperatures. These results uncover molecular mass growth processes from the “bottom up” via RSFRs in high temperature circumstellar environments and combustion systems expanding our fundamental knowledge of the organic, hydrocarbon chemistry in our universe.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Performance Analysis of Data Processing in Distributed File Systems with Near Data Processing

In the era of big data, the escalating volume and velocity of data generation pose significant challenges in data processing. Traditional systems like Spark and Hadoop manage the increasing amount and velocity of data by improving data placement and processing speeds. However, they face inherent limitations due to the essential data movement required for processing. In this paper, we explore the Skyhook framework, a novel extension of the Ceph distributed system, which significantly reduces the need for data movement. We present an extensive case study using the Skyhook framework, applying it with the TPC-H and K-means clustering algorithms. More specifically, we leverage the TPC-H benchmark to distinguish between CPU-intensive and I/O-intensive tasks. We explore the integration of K-means clustering into SQL, coupled with a near-data processing system to offload the computational burden of the K-means clustering algorithm to storage nodes. We conduct a comprehensive performance evaluation of distributed data processing applications across three processing approaches: traditional layout (baseline), optimized layout, and near-data processing. Additionally, we introduce the use of the FIO tool to simulate real-world system workloads, enabling the measurement of performance metrics such as average latency and CPU utilization. Our research is a significant advance in understanding how to optimize data processing systems to meet the demands of the modern data landscape.

Hou, Shiyue↗

TRINIDI (Time-of-Flight Resonance Imaging with Neutrons for Isotopic Density Inference)

This software is an open-source Python library that provides tools for processing hyperspectral neutron time-of-flight radiography data. This type of data allows material decomposed reconstructions to be generated with the use of material characteristic spectral responses and the algorithms provided in this code library. The software library will contain tools for pre-processing the neutron measurement data, estimating measurement system parameters, reconstructing material decomposed radiographs, and computing material decomposed computed tomography (CT). Furthermore, it will have capability to generate and process simulated neutron time-of-flight data with the goal of benchmarking and demonstrating the tools that are provided. The software will include thorough documentation and application examples.

Balke, Thilo↗

DFT-based QM/MM with Particle-Mesh Ewald for Direct, Long-Range Electrostatic Embedding

In this work, we present a DFT-based, QM/MM implementation with long-range electrostatic embedding achieved by direct real-space integration of the particle mesh Ewald (PME) computed electrostatic potential. The key transformation is the interpolation of the electrostatic potential from the PME grid to the DFT quadrature grid, from which integrals are easily evaluated utilizing standard DFT machinery. We provide benchmarks of the numerical accuracy with choice of grid size and real-space corrections, and demonstrate that good convergence is achieved while introducing nominal computational overhead. Furthermore, the approach requires only small modification to existing software packages, as is demonstrated with our implementation in the OpenMM and Psi 4 software. After presenting convergence benchmarks, we evaluate the importance of long-range electrostatic embedding in three solute/solvent systems modeled with QM/MM. Water and BMIM/BF 4 ionic liquid were considered as "simple" and "complex" solvents respectively, with water and p-phenylenediamine (PPD) solute molecules treated at QM level of theory. While electrostatic embedding with standard real-space truncation may introduce negligible error for simple systems such as water solute in water solvent, errors become more significant when QM/MM is applied to complex solvents such as ionic liquids. An extreme example is the electrostatic embedding energy for oxidized PPD in BMIM/BF 4 for which real-space truncation produces severe error even at 2-3 nm cutoff distances. This latter example illustrates that utilization of QM/MM to compute redox potentials within concentrated electrolytes/ionic media requires carefully chosen long-range electrostatic embedding algorithms, with our presented algorithm providing a general and robust approach.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Accelerating resonant spectroscopy simulations using multishifted biconjugate gradient

Resonant spectroscopies, which involve intermediate states with finite lifetimes, provide important insights into collective excitations in quantum materials that are otherwise inaccessible. However, theoretical understanding in this area is often limited by the numerical challenges of solving Kramers-Heisenberg-type response functions for large-scale systems. To address this, we introduce a multishifted biconjugate gradient algorithm that exploits the shared structure of Krylov subspaces across spectra with varying incident energies, effectively reducing the computational complexity to that of linear spectroscopies. Both mathematical proofs and numerical benchmarks confirm that this algorithm substantially accelerates spectral simulations, achieving constant complexity independent of the number of incident energies, while ensuring accuracy and stability. This development provides a scalable, versatile framework for simulating advanced spectroscopies in quantum materials.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

VTK-m: Visualization for the Exascale Era and Beyond

A recent trend in modern high-performance computing is the increasing use of hybrid architectures, where the vast majority of performance comes from accelerators. Modern accelerators are based on Graphics Processing Units (GPU) that contain many low power cores that in their aggregate provides an extremely high computation rate. Current and future CPU processors are requiring more explicit parallelism as each successive version of the hardware packs in more cores, and technologies like hyperthreading and vector operations require even more parallel processing to leverage each core’s full potential. As an example, the Frontier supercomputer installed at Oak Ridge National Laboratories recently hit a record breaking 1.1 exaflops1 on the LINPACK HPC benchmark [Shoemaker 2022]. The system contains 37632 AMD MI250x GPUs which requires more than half a billion threads to keep the system fully utilized [Khizeran 2022].VTK-m is a toolkit of scientific visualization algorithms for these emerging processor architectures. VTK-m supports the fine-grained concurrency for data analysis and visualization algorithms required to drive extreme scale computing by providing abstract models for data and execution that can be applied to a variety of algorithms across many different processor architectures.

Bolstad, Mark↗

Iterative quantum optimization of spin glass problems with rapidly oscillating transverse fields

In this work, we introduce a new iterative quantum algorithm, called Iterative Symphonic Tunneling for Satisfiability problems (IST-SAT), which solves quantum spin glass optimization problems using high-frequency oscillating transverse fields. IST-SAT operates as a sequence of iterations, in which bitstrings returned from one iteration are used to set spin-dependent phases in oscillating transverse fields in the next iteration. Over several iterations, the novel mechanism of the algorithm steers the system toward the problem ground state. We benchmark IST-SAT on sets of hard MAX-3-XORSAT problem instances with exact state vector simulation, and report polynomial speedups over Trotterized adiabatic quantum computation and the best known semi-greedy classical algorithm. When IST-SAT is seeded with a sufficiently good initial approximation, the algorithm converges to exact solution(s) in a polynomial number of iterations. Our numerical results identify a critical Hamming radius, or quality of initial approximation, where the time-to-solution crosses from exponential to polynomial scaling in problem size. This work proposes IST-SAT a new quantum algorithm, which improves upon solutions obtained from initial classical or quantum optimization algorithms. The steering mechanism we introduce through IST-SAT presents a new path toward achieving quantum advantage in optimization.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Two-Level Sketching Alternating Anderson Acceleration for Complex Physics Applications

We present a novel two-level sketching extension of the Alternating Anderson–Picard (AAP) method for accelerating fixed-point iterations in challenging single- and multiphysics simulations governed by discretized PDEs. Our approach combines a static, physics-based projection that reduces the least-squares (LS) problem to the most informative field (e.g., via Schur-complement insight) with a dynamic, algebraic sketching stage driven by a backward stability analysis under Lipschitz continuity. We introduce inexpensive estimators for stability thresholds and cache-aware randomized selection strategies to balance computational cost against memory access overhead. The resulting algorithm solves reduced LS systems in place, minimizes memory footprints, and seamlessly alternates between low-cost Picard updates and Anderson mixing. Implemented in Julia, our two-level sketching AAP achieves up to 50% time-to-solution reductions compared to standard Anderson acceleration—without degrading convergence rates—on benchmark problems including Stokes, 𝑝-Laplacian, bidomain, and Navier–Stokes formulations at varying problem sizes. These results demonstrate the method’s robustness, scalability, and potential for integration into high-performance scientific computing frameworks. Our implementation is available open source in the AAP.jl library.

Barnafi, Nicolas [University of Chile, Santiago]↗

FFTX-IRIS: Towards Performance Portability and Heterogeneity for SPIRAL Generated Code

FFTX-IRIS is a dynamic system to efficiently utilize novel heterogeneous platforms. This system links two next-generation frameworks, FFTX and IRIS, to navigate the complexity of different hardware architectures. FFTX provides a runtime code generation framework for high-performance Fast Fourier Transform kernels. IRIS runtime provides portability and multi-device heterogeneity, allowing computation on any available compute resource. Together, FFTX-IRIS enables code generation, seamless portability, and performance without user involvement. We show the design of the FFTX-IRIS system along with an evaluation of various small FFT benchmarks. We also demonstrate multi-device heterogeneity of FFTX-IRIS with a larger stencil application.

Rao, Sanil↗

SCALE Input and Result Files Supporting SCALE Inventory and Reactivity Analysis of the gFHR

This dataset contains input and result files of computational simulations with the SCALE code system. The simulations cover radionuclide inventory and reactivity analyses of a fluoride salt-cooled high temperature pebble-bed reactor (PB-FHR), specifically the generic FHR benchmark. Users wanting to reproduce results from this dataset are required to obtain a license to the SCALE code system for which details on the distribution can be found here: https://www.ornl.gov/scale/releases

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Large-scale harmonic balance simulations with Krylov subspace and preconditioner recycling

The multi-harmonic balance method combined with numerical continuation provides an efficient framework to compute a family of time-periodic solutions, or response curves, for large-scale, nonlinear mechanical systems. The predictor and corrector steps repeatedly solve a sequence of linear systems that scale by the model size and number of harmonics in the assumed Fourier series approximation. In this paper, a novel Newton–Krylov iterative method is embedded within the multi-harmonic balance and continuation algorithm to efficiently compute the approximate solutions from the sequence of linear systems that arise during the prediction and correction steps. Further, the method recycles, or reuses, both the preconditioner and the Krylov subspace generated by previous linear systems in the solution sequence. A delayed frequency preconditioner refactorizes the preconditioner only when the performance of the iterative solver deteriorates. The GCRO-DR iterative solver recycles a subset of harmonic Ritz vectors to initialize the solution subspace for the next linear system in the sequence. The performance of the iterative solver is demonstrated on two exemplars with contact-type nonlinearities and benchmarked against a direct solver with traditional Newton–Raphson iterations.

97 MATHEMATICS AND COMPUTING↗

Multi-group Examination of Nickel-Reflected HEU System [Slides]

Researchers noted an unusually large bias for 8-in. nickel reflected HEU sphere in HMF-003 between the 252- group library and the CE library in SCALE 6.2.4. The bias was investigated by reviewing reactions that $k_{eff}$ is sensitive to using TSUNAMI and collecting reaction rate data using tools within SCALE. For this system, bias is primarily due to the elastic scattering in nickel. Researchers compared new multigroup structures in SCALE 6.3. For this system, reduction in CE-to-MG bias seen in the 302-group and further improved in the 1597-group structure. When researchers compared ENDF/B-VIII.0 in SCALE 6.3, library showed improved results with new nickel evaluation. However, the MG-to-CE bias can still be large as large biases can occur in any system. This example highlights the importance of validating results with measured systems.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Integrating quantum computing resources into scientific HPC ecosystems

Quantum Computing (QC) offers significant potential to enhance scientific discovery in fields such as quantum chemistry, optimization, and artificial intelligence. Yet QC faces challenges due to the noisy intermediate-scale quantum era’s inherent external noise issues. Here, this paper discusses the integration of QC as a computational accelerator within classical scientific high-performance computing (HPC) systems. By leveraging a broad spectrum of simulators and hardware technologies, we propose a hardware-agnostic framework for augmenting classical HPC with QC capabilities. Drawing on the HPC expertise of the Oak Ridge National Laboratory (ORNL) and the HPC lifecycle management of the Department of Energy (DOE), our approach focuses on the strategic incorporation of QC capabilities and acceleration into existing scientific HPC workflows. This includes detailed analyses, benchmarks, and code optimization driven by the needs of the DOE and ORNL missions. Our comprehensive framework integrates hardware, software, workflows, and user interfaces to foster a synergistic environment for quantum and classical computing research. This paper outlines plans to unlock new computational possibilities, driving forward scientific inquiry and innovation in a wide array of research domains.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Efficient solvers for hybridized three-field mixed finite element coupled poromechanics

We consider a mixed hybrid finite element formulation for coupled poromechanics. A stabilization strategy based on a macro-element approach is advanced to eliminate the spurious pressure modes appearing in undrained/incompressible conditions. The efficient solution of the stabilized mixed hybrid block system is addressed by developing a class of block triangular preconditioners based on a Schur-complement approximation strategy. Robustness, computational efficiency and scalability of the proposed approach are theoretically discussed and tested using challenging benchmark problems on massively parallel architectures.

42 ENGINEERING↗

Fast convolutional neural networks on FPGAs with hls4ml

We introduce an automated tool for deploying ultra low-latency, low-power deep neural networks with convolutional layers on field-programmable gate arrays (FPGAs). By extending the hls4ml library, we demonstrate an inference latency of 5 µs using convolutional architectures, targeting microsecond latency applications like those at the CERN Large Hadron Collider. Considering benchmark models trained on the Street View House Numbers Dataset, we demonstrate various methods for model compression in order to fit the computational constraints of a typical FPGA device used in trigger and data acquisition systems of particle detectors. In particular, we discuss pruning and quantization-aware training, and demonstrate how resource utilization can be significantly reduced with little to no loss in model accuracy. We show that the FPGA critical resource consumption can be reduced by 97% with zero loss in model accuracy, and by 99% when tolerating a 6% accuracy degradation.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Enabling topography-resolving structural dynamic contact simulation

Damping of structures and systems is often dominated by frictional dissipation in connections, the prediction of which remains a longstanding scientific challenge. Previous studies have shown that the actual topography of contact interfaces may have a strong effect, especially in the partial slip/liftoff regime. We recently proposed a multi-scale method, which couples finite element and boundary element modeling. The primary benefit of this approach is that it permits to analyze the effect of the actual contact topography on the dynamics of jointed structures. While this multi-scale modeling method was initially developed for quasi-static analysis, we demonstrate herein how it can be used for time step integration and Harmonic Balance analysis. We cross-verify those fully dynamic analysis methods against each other and quasi-static results, for the S4 Beam benchmark. We compare the multi-scale method against state-of-the-art full-FE analysis, in terms of numerical damping and computational performance. Some discrepancy is found to be of physical origin. Depending on the load history, it is shown that the system settles to a slightly different equilibrium. Finally, transient multi-scale simulations enable the prediction of this interesting phenomenon, for the first time, for a structure with bolted joints.

Frictional-unilateral contact↗

Sparse chronology strategy for integrating seasonal energy storage in capacity expansion models

Here, this study develops the sparse chronology method to enhance the representative period framework in capacity expansion models, enabling the effective integration of long-duration energy storage modeling. Traditional representative period methods cannot capture the state of charge of seasonal energy storage systems because they do not establish effective inter-day linkages to connect the state of charge between periods. The sparse chronology approach addresses this limitation by establishing inter-day linkages that allow state of charge to shift inter-seasonally. At the same time, it groups identical representative days into partitions, applying constraints sparsely and implicitly to reduce computational load further. Validation results demonstrate that this method successfully simulates long-duration energy storage patterns, achieving close alignment with a continuous yearly benchmark model, with seasonal trends and state of charge cycles clearly represented. The computational load analysis reveals that the sparse chronology method efficiently applies constraints on maximum and minimum state of charge limits within the representative day framework, eliminating the need for detailed constraints on each individual day. By partitioning representative days and constraining only the start and end of each partition, the method significantly decreases computational requirements. Simulation results show that sparse chronology closely approximates the continuous yearly method's accuracy, even with as few as 20 representative days, achieving correlation values with the benchmark of nearly 0.9 in state of charge plots. Furthermore, it maintains computational efficiency, requiring only 4 % of the solver time compared to the continuous yearly method with 20 representative days. This approach allows capacity expansion models to incorporate long-duration energy storage with high temporal, spatial, and technological resolution, enabling more detailed modeling for large-scale power systems.

24 POWER TRANSMISSION AND DISTRIBUTION↗