Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “computer system benchmarking”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

ddcMD: A fully GPU-accelerated molecular dynamics program for the Martini force field

We have implemented the Martini force field within Lawrence Livermore National Laboratory’s molecular dynamics program, ddcMD. The program is extended to a heterogeneous programming model so that it can exploit graphics processing unit (GPU) accelerators. In addition to the Martini force field being ported to the GPU, the entire integration step, including thermostat, barostat, and constraint solver, is ported as well, which speeds up the simulations to 278-fold using one GPU vs one central processing unit (CPU) core. A benchmark study is performed with several test cases, comparing ddcMD and GROMACS Martini simulations. The average performance of ddcMD for a protein–lipid simulation system of 136k particles achieves 1.04 µs/day on one NVIDIA V100 GPU and aggregates 6.19 µs/day on one Summit node with six GPUs. The GPU implementation in ddcMD offloads all computations to the GPU and only requires one CPU core per simulation to manage the inputs and outputs, freeing up remaining CPU resources on the compute node for alternative tasks often required in complex simulation campaigns.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Advancing our understanding of instabilities in high-energy-density systems through the study of benchmark datasets with advanced modeling tools

Numerous types of pulsed power driven inertial confinement fusion (ICF) and high energy density (HED) systems rely on implosion stability to achieve desired temperatures, pressures, and densities. Sandia National Laboratories Pulsed Power Sciences Center’s main ICF platform, Magnetized Liner Inertial Fusion (MagLIF), suffers from implosion instabilities which limit attainable fuel conditions and can compromise fuel confinement. This Truman Fellowship research primarily focused on computationally exploring (a) methods for improving our understanding of hydrodynamic and magnetohydrodynamic instabilities that form during cylindrical liner implosions, (b) methods for mitigating implosion instabilities, particularly those that degrade performance of MagLIF targets, and (c) novel MagLIF target designs intended to improve target performance primarily via enhanced implosion stability. Several multi-dimensional computational tools were used, including the magnetohydrodynamics code ALEGRA, the radiation-magnetohydrodynamics code HYDRA, and the magnetohydrodynamics code KRAKEN. This research succeeded in executing and analyzing simulations of automagnetizing liner implosions, shockless MagLIF implosions, dynamic screw pinch driven cylindrical liner implosions, and cylindrically convergent HED instability studies. The methods and tools explored and developed in this Truman Fellowship research have been published in several peer-reviewed journal articles and will serve as useful contributions to the fields of pulsed power science and engineering, particularly pertaining to pulsed power ICF and HED science.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

An Integrated Framework for Memory-Centric Analysis: From Trace Collection to Co-Design

The memory wall phenomenon—where advances in processor performance significantly outpace those in memory subsystems—poses a fundamental challenge for contemporary computing systems. In memory-bound applications, memory subsystem behavior dominates performance, yet existing analysis approaches present significant limitations: detailed microarchitectural simulators require days to weeks to simulate modest workloads; hardware performance counters provide only aggregate statistics that obscure temporal and spatial access patterns; and scaled simulation approaches face challenges in capturing certain behaviors that emerge at larger scales. These limitations reflect a processor-centric design philosophy increasingly misaligned with memory-bound workloads where detailed understanding of memory access patterns, cache hierarchy interactions, and contention is critical for effective optimization. This paper presents an integrated framework for memory-centric analysis that enables effective hardware-software co-design. We describe practical trace collection techniques, including hardware-assisted processor tracing with minimal overhead and portable software-based instrumentation with statistical sampling. We present multi-perspective analysis methods that examine memory behavior from temporal, sequential, spatial, and relational viewpoints, revealing distinct optimization opportunities invisible in aggregate metrics. We detail an architectural modeling framework that uses sampled traces with temporal interpolation and confidence-based filtering to evaluate cache and memory configurations. Evaluation on representative benchmarks demonstrates that this framework achieves practical accuracy (L2 cache errors of 2.64\%, confidence-filtered L3 errors of 9.92\%, bandwidth errors of 7.33\%) while providing substantial speedup (26.8×) over cycle-accurate simulation, enabling rapid design space exploration. We demonstrate how this integrated framework enables systematic identification of both hardware optimizations (memory controller tuning, bank partitioning, NUMA configuration) and software optimizations (data layout restructuring, prefetching strategies, memory-aware scheduling). Through this comprehensive treatment of the memory-centric analysis pipeline—from trace collection through architectural modeling to co-design application—we provide researchers and practitioners with practical techniques for addressing memory bottlenecks in contemporary computing systems.

Gajaria, Dhruv Mayur↗

Sampling lattices in semi-grand canonical ensemble with autoregressive machine learning

Calculating thermodynamic potentials and observables efficiently and accurately is key for the application of statistical mechanics simulations to materials science. However, naive Monte Carlo approaches, on which such calculations are often dependent, struggle to scale to complex materials in many state-of-the-art disciplines such as the design of high entropy alloys or multi-component catalysts. To address this issue, we adapt sampling tools built upon machine learning-based generative modeling to the materials space by transforming them into the semi-grand canonical ensemble. Furthermore, we show that the resulting models are transferable across wide ranges of thermodynamic conditions and can be implemented with any internal energy model U, allowing integration into many existing materials workflows. We demonstrate the applicability of this approach to the simulation of benchmark systems (AgPd, CuAu) that exhibit diverse thermodynamic behavior in their phase diagrams. Finally, we discuss remaining challenges in model development and promising research directions for future improvements.

36 MATERIALS SCIENCE↗

Resolving femtosecond photoinduced energy flow: capture of nonadiabatic reaction pathway topography and wavepacket dynamics from photoexcitation through the conical intersection seam (Final Technical Report)

The dynamics that take place within just tens to hundreds of femtoseconds following the absorption of light by a molecule can play a critical role in how the absorbed energy is directed, allowing it to be used for a specific function or dissipated harmlessly. The form of chemical change that occurs rapidly in these molecules is called a “nonadiabatic electronic transition.” Such transitions are known to mediate energy flow in natural biological systems such as the ultraviolet photoprotection mechanism of DNA and the first step of the human vision response. Understanding how these mechanisms work precisely may help scientists achieve controlled manipulation of solar energy or optical control of a wide range of energy management functions in artificial systems. Experimental methods, however, have not yet allowed a precisely resolved and complete measurement of nonadiabatic electronic transitions. This constitutes a major obstacle to progress in the field. For progress to occur that would inform a wide body of research aiming to efficiently harness the energy of light for practical purposes, it is especially important to benchmark computational models of the molecules undergoing these rapid changes with experimental measurements, in order to learn which models are accurate. With Dept. of Energy funding, we have made strong progress towards establishing a new optical method for experimentally detecting the full nonadiabatic electronic transition. This requires having coordinated pulses of light covering the visible through the mid-infrared range of the electromagnetic spectrum that last only ten femtoseconds. We have developed a new, relatively simple approach for generating such pulses of laser light, and have incorporated them into a time-resolved spectrometer for measuring rapid changes in molecules. These tools can provide the greater precision and new types of data that are needed to benchmark computational models of molecular change and thus to make progress in the field. Our tools were tested on graphene, an excellent solid-state sample for verifying the capabilities and limitations of our instrumentation. The investment made in these tools by the Dept. of Energy Office of Science will allow new fundamental scientific understanding of energy dynamics in molecules in future studies.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Sizing ramping reserve using probabilistic solar forecasts: A data-driven method

Ramping products have been introduced or proposed in several U.S. power markets to mitigate the impact of load and renewable uncertainties on market efficiency and reliability. Current methods often rely on historical data to estimate the requirements of ramping products and fail to take into account the effects of the latest weather conditions and their uncertainties, which could lead to overly conservative or insufficient requirements. This study proposes a k-nearest-neighbor-based method to give weather-informed estimates of ramping needs based on short-term probabilistic solar irradiance forecasts. Forecasts from multiple sites are employed in conjunction with principal component analysis to derive numerical classifiers to characterize system-level weather conditions. In addition, we develop a data-driven method to optimize the model parameters in a rolling-forward manner. By using real-world data from the California Independent System Operator, we design two metrics to evaluate method performance: 1) frequency of shortage and 2) oversupply of ramping product. Our proposed method presents advantages in comparison with the baseline and a set of benchmark methods: without compromising system reliability, it reduces system ramping requirements by up to 25%, therefore improving both system reliability and economics.

14 SOLAR ENERGY↗

Hindsight logging for model training

In modern Machine Learning, model training is an iterative, experimental process that can consume enormous computation resources and developer time. To aid in that process, experienced model developers log and visualize program variables during training runs. Exhaustive logging of all variables is infeasible, so developers are left to choose between slowing down training via extensive conservative logging, or letting training run fast via minimalist optimistic logging that may omit key information. As a compromise, optimistic logging can be accompanied by program checkpoints; this allows developers to add log statements post-hoc, and "replay" desired log statements from checkpoint---a process we refer to as hindsight logging. Unfortunately, hindsight logging raises tricky problems in data management and software engineering. Done poorly, hindsight logging can waste resources and generate technical debt embodied in multiple variants of training code. In this paper, we present methodologies for efficient and effective logging practices for model training, with a focus on techniques for hindsight logging. Our goal is for experienced model developers to learn and adopt these practices. To make this easier, we provide an open-source suite of tools for Fast Low-Overhead Recovery (flor) that embodies our design across three tasks: (i) efficient background logging in Python, (ii) adaptive periodic checkpointing, and (iii) an instrumentation library that codifies hindsight logging for efficient and automatic record-replay of model-training. Model developers can use each flor tool separately as they see fit, or they can use flor in hands-free mode, entrusting it to instrument their code end-to-end for efficient record-replay. Our solutions leverage techniques from physiological transaction logs and recovery in database systems. Evaluations on modern ML benchmarks demonstrate that flor can produce fast checkpointing with small user-specifiable overheads (e.g. 7%), and still provide hindsight log replay times orders of magnitude faster than restarting training from scratch.

Computer Science↗

Structurally Constrained Evolutionary Algorithm for the Discovery and Design of Metastable Phases

Metastable materials are abundant in nature and technology, showcasing remarkable properties that inspire innovative materials design. However, traditional crystal structure prediction methods, which rely solely on energetic factors to determine a structure’s fitness, are not suitable for predicting the vast number of potentially synthesizable phases that represent a local minimum corresponding to a state in thermodynamic equilibrium. Here, we present a new approach for the prediction of metastable phases with specific structural features, and interface this method with the XTALOPT evolutionary algorithm. Our method relies on structural features that include the local crystalline order (e.g., the coordination number or chemical environment), and symmetry (e.g., Bravais lattice and space group) to filter the breeding pool of an evolutionary crystal structure search. The effectiveness of this approach is benchmarked on three known metastable systems: XeN 8 , with a two-dimensional polymeric nitrogen sublattice, brookite TiO 2 , and a high pressure BaH 4 phase that was recently characterized. Additionally, a newly predicted metastable melaminate salt, P1¯WC 3 N 6 , was found to possess an energy that is lower than two phases proposed in a recent computational study. Here, the method presented here could help in identifying the structures of compounds that have already been synthesized, and developing new synthesis targets with desired properties.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Establish the basis for Breadth-First Search on Frontier System: XBFS on AMD GPUs

Graphics Processing Units (GPUs) offer significant potential for accelerating various computational tasks, including Breadth-First Search (BFS). Numerous efforts have been made to deploy BFS on GPUs effectively. To address the dynamic nature of BFS, XBFS, the state-of-the-art work, employs an adaptive strategy that leverages different optimized frontier queue generation designs, accommodating the varying characteristics of levels in BFS. While XBFS demonstrates excellent performance on NVIDIA Quadro P6000 GPUs, it faces challenges when deployed on AMD GPUs. In this work, we present our efforts to implement XBFS’s adaptive approach on Frontier, the most powerful supercomputer system, by porting XBFS to AMD MI250X GPUs. Through targeted optimizations tailored to the unique features of AMD GPUs, our implementation achieves an average performance of 43 Giga-Traversed Edges Per Second (GTEPS) per Graphics Compute Dies (GCD). Based on these results, we observe potential for surpassing the performance of the official Frontier results from the Graph500 benchmark released in June 2024.

Yang, Haoshen↗

Advancing Concentrating Solar Thermal Modeling Using System Advisor Model (SAM)

Concentrating solar thermal (CST) technologies play a critical role in enabling dispatchable power and high-temperature industrial heat applications. Accurate and flexible modeling tools are essential for evaluating system performance, guiding technology research and development, and informing investment decisions. The National Laboratory of the Rockies's System Advisor Model (SAM) is a widely used techno-economic simulation platform for CST systems, providing detailed performance and financial modeling capabilities for multiple CST system configurations. SAM integrates physics-based performance models with financial analysis to simulate the behavior of complex energy systems under realistic operating conditions. For CST technologies (including tower, parabolic trough, and linear Fresnel), SAM enables hourly simulations using site-specific weather data that ensure feasible operating conditions and convergence of mass and energy between core system components (i.e., solar field, receiver, thermal energy storage, and power cycle). These capabilities allow researchers and developers to evaluate annual energy production, capacity factors, levelized cost of energy (LCOE), and system dispatch strategies. A key advantage of SAM lies in its flexibility for parametric analysis and large-scale computational studies. Users can vary system design parameters such as heliostat field layout, receiver dimensions, thermal energy storage capacity, power block sizing, and installation cost assumptions to investigate their impact on system performance and financial metrics. When combined with automated scripting through LK, SDKTool, or Python interfaces, SAM enables high-throughput simulation workflows that support sensitivity analysis, technology benchmarking, and optimization studies. These approaches are particularly valuable for next-generation CST concepts, where design spaces are large and system interactions are complex. Another important capability of SAM is its support for dispatch optimization and thermal energy storage modeling, which are central to the value proposition of CST technologies. The ability to simulate integrated storage and flexible power generation allows researchers to explore strategies that maximize grid value, improve capacity utilization, and enhance integration with variable resources such as photovoltaic and wind generation. This poster will present an overview of SAM's thermal system modeling capabilities including concentrating solar. Additionally, we will highlight new feature developments including: 1) implementing Google's OR-Tools optimization platform for faster and more robust dispatch optimization, 2) developing a new power load following controller for modeling behind-the-meter applications, 3) enabling direct modeling of CSP-PV hybrid systems with the inclusion of battery storage, and 4) developing a multi-receiver falling particle Gen3 system model.

14 SOLAR ENERGY↗

Impact of Thermal Scattering Law on Similarity Assessment in Light-Water or Polyethylene-Moderated Systems [Slides]

This presentation found validation efforts revealed noticeable differences in c k for light-water and polyethylene moderated systems when compared to PST-002-001. Additionally, although 1 H-H 2 O and 1 H-poly use the same cross section and covariance data, the treatment of TSLs in SCALE lead to no contribution to ck between polyethylene application and benchmark experiment. When H 2 O contribution were removed the difference in c k appears to come from 239 Pu chi. Sensitivity profiles and data-induced uncertainty confirmed that the polyethylene application was more closely resembling the benchmark experiment. Specific differences between similar systems can be accounted for and examined to help understand differences in similarity (c k ) between systems was a conclusion reached by this presentation.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Relativistic Exact Two-Component Theory in the Generalized Pseudospectral Representation

We present a formulation and implementation of exact two-component (X2C) relativistic theory in the generalized pseudospectral representation. When combined with the Hartree-Fock-Slater framework, this approach enables efficient and accurate treatments of scalar-relativistic and spin-orbit effects in atomic electronic structures without the computational overhead of four-component methods. Benchmark calculations across light and heavy elements demonstrate that the our X2C scheme yields substantially more accurate relativistic corrections to core-electron binding energies than perturbational Breit-Pauli treatments while converging more rapidly with respect to basis size. The method provides an improved computational framework for modeling ultrafast x-ray-induced processes in heavy-element systems.

Wang, Xubo↗

High-order matrix-free incompressible flow solvers with GPU acceleration and low-order refined preconditioners

In this work, we present a matrix-free flow solver for high-order finite element discretizations of the incompressible Navier-Stokes and Stokes equations with GPU acceleration. For high polynomial degrees, assembling the matrix for the linear systems resulting from the finite element discretization can be prohibitively expensive, both in terms of computational complexity and memory. For this reason, it is necessary to develop matrix-free operators and preconditioners, which can be used to efficiently solve these linear systems without access to the matrix entries themselves. The matrix-free operator evaluations utilize GPU-accelerated sum-factorization techniques to minimize memory movement and maximize throughput. The preconditioners developed in this work are based on a low-order refined methodology with parallel subspace corrections, as described for diffusion problems in [1]. The saddle-point Stokes system is solved using block-preconditioning techniques, which are robust in mesh size, polynomial degree, time step, and viscosity. For the incompressible Navier-Stokes equations, we make use of projection (fractional step) methods, which require Helmholtz and Poisson solves at each time step. The performance of our flow solvers is assessed on several benchmark problems in two and three spatial dimensions.

97 MATHEMATICS AND COMPUTING↗

Shielding Benchmark Comparison - MCNP6.2

This report documents the calculations performed for a shielding code comparison between various Department of Energy (DOE) sites. These shielding calculations were performed using an experiment drawn from the International Handbook of Evaluated Criticality Safety Benchmark Experiments, published by the Organisation for Economic Cooperation and Development/Nuclear Energy Agency (OECD/NEA). The benchmark selected for comparison is ALARM-CF-AIR-LAB-001 (“Neutron Fields in the Three-Section Concrete Labyrinth from Cf-252 Source, Benchmark ALARM CF AIR LAB-001”). The Y-12 submission for this code comparison was performed with Monte Carlo N Particle (MCNP) Transport Code System, Version 6.2 and Automated Variance Reduction Generator (ADVANTG).

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

HIPLZ: Enabling performance portability for exascale systems

While heterogeneous computing has emerged as a dominant trend in current and future High-Performance Computing (HPC) systems, it is also widely recognized that this shift has led to increased software complexity due to a proliferation of programming systems for different heterogeneous processors. One such example is the Heterogeneous-Compute Interface for Portability from AMD (HIP ), which is composed of a C Runtime API and C++ Kernel Language. Many HPC applications will likely use HIP on future exascale systems (e.g., Frontier and El Capitan), but HIP currently only targets AMD and NVIDIA processors. This limitation creates challenges for users who would also like to run their applications on exascale systems based on other architectures (e.g., Aurora, which is based on Intel hardware) that are currently not targeted by HIP . In this paper, we introduce the design and implementation of HIPLZ , a compiler and runtime system that uses the Intel Level Zero API to support HIP on Intel GPU architectures. We discuss the design of HIPLZ , derived from HIPCL (an implementation of HIP on top of OpenCL ), and portability issues that occur from using the Level Zero runtime as a backend. We evaluate our implementation by running several performance benchmarks and mini-apps written in HIP on Intel architectures using HIPLZ . Our results show that this approach provides competitive performance relative to Intel's OpenCL implementations on Intel Gen9 and UHD Graphics 770 GPUs, while providing good coverage of features needed by HPC applications. Overall, this approach is a promising demonstration of enabling performance portability for exascale systems.

97 MATHEMATICS AND COMPUTING↗

Automated Waterbox Inspection for Nuclear Power Plants Using Computer Vision - Based Change Detection

Nuclear power plant waterboxes require regular inspection for leaks, missing components, and structural damage during maintenance outages. Traditional manual inspection is time-consuming and poses safety risks from confined space entry. We developed an automated computer vision system for drone-based waterbox inspection in partnership with Florida Light and Power. Our approach uses feature detection and matching to identify critical changes between baseline and current inspection images, automatically flagging additions (leaks/debris), removals (missing plugs), and translations (displaced components) while compensating for drone movement and environmental variations. We systematically evaluated six feature matching methods, from classical approaches (SIFT+BF) to state-of-the-art neural networks (SuperPoint+SuperGlue), using both standard benchmarks (HPatches) and waterbox-specific validation with real-world augmentations. SuperPoint+SuperGlue achieved superior performance with 7.82 pixels RMSE and 100% success rate—2.8x better accuracy than our baseline. While the pre-trained model has commercial licensing restrictions for nuclear deployment, our findings validate this architecture for custom training. We implemented a real-time GUI demonstrating the SIFT+BF approach for immediate deployment, processing drone feeds at 30 FPS with color-coded change visualization. Future work includes training a custom SuperPoint+SuperGlue model on waterbox data and integrating Vision-Language Models for automated reporting and maintenance guidance.

46 - INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AN↗

Lessons Learned and Scalability Achieved When Porting Uintah to DOE Exascale Systems

A key challenge faced when preparing codes for Department of Energy (DOE) exascale systems was designing scalable applications for systems featuring hardware and software not yet available at leadership-class scale. With such systems now available, it is important to evaluate scalability of the resulting software solutions on these target systems. One such code designed with the exascale DOE Aurora and DOE Frontier systems in mind is the Uintah Computational Framework, an open-source asynchronous many-task (AMT) runtime system. To prepare for exascale, Uintah adopted a portable MPI+X hybrid parallelism approach using the Kokkos performance portability library (i.e., MPI+Kokkos). This paper complements recent work with additional details and an evaluation of the resulting approach on Aurora and Frontier. Results are shown for a challenging benchmark demonstrating interoperability of 3 portable codes essential to Uintah-related combustion research. These results demonstrate single-source portability across Aurora and Frontier with scaling characteristics shown to 3,072 Aurora nodes and 9,216 Frontier nodes. In addition to showing results run to new scales on new systems, this paper also discusses lessons learned through efforts preparing Uintah for exascale systems.

Holmen, John [ORNL] (ORCID:0000000259342641)↗

Machine learning to alleviate Hubbard-model sign problems

Lattice Monte Carlo calculations of interacting systems on nonbipartite lattices exhibit an oscillatory imaginary phase known as the phase or sign problem, even at zero chemical potential. One method to alleviate the sign problem is to analytically continue the integration region of the state variables into the complex plane via holomorphic flow equations. For asymptotically large flow times, the state variables approach manifolds of constant imaginary phase known as Lefschetz thimbles. Furthermore, flowing such variables and calculating the ensuing Jacobian is a computationally demanding procedure. In this paper, we demonstrate that neural networks can be trained to parametrize suitable manifolds for this class of sign problem and drastically reduce the computational cost for different severely afflicted small volume systems. In particular, we apply our method to the Hubbard model on the triangle and tetrahedron, both of which are nonbipartite. At strong interaction strengths and modest temperatures, the tetrahedron suffers from a severe sign problem that cannot be overcome with standard reweighting techniques, while it quickly yields to our method. We benchmark our results with exact calculations and comment on future directions of this work.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗