Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “computer system benchmarking”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Automated Integration of Continental-Scale Observations in Near-Real Time for Simulation and Analysis of Biosphere–Atmosphere Interactions

The National Ecological Observatory Network (NEON) is a continental-scale observatory with sites across the US collecting standardized ecological observations that will operate for multiple decades. To maximize the utility of NEON data, we envision edge computing systems that gather, calibrate, aggregate, and ingest measurements in an integrated fashion. Edge systems will employ machine learning methods to cross-calibrate, gap-fill and provision data in near-real time to the NEON Data Portal and to High Performance Computing (HPC) systems, running ensembles of Earth system models (ESMs) that assimilate the data. For the first time gridded EC data products and response functions promise to offset pervasive observational biases through evaluating, benchmarking, optimizing parameters, and training new machine learning parameterizations within ESMs all at the same model-grid scale. Leveraging open-source software for EC data analysis, we are already building software infrastructure for integration of near-real time data streams into the International Land Model Benchmarking (ILAMB) package for use by the wider research community. We will present a perspective on the design and integration of end-to-end infrastructure for data acquisition, edge computing, HPC simulation, analysis, and validation, where Artificial Intelligence (AI) approaches are used throughout the distributed workflow to improve accuracy and computational performance.

Durden, David J.↗

Cooling Matters: Benchmarking Large Language Models and Vision-Language Models on Liquid-Cooled Versus Air-Cooled H100 GPU Systems

The unprecedented growth in artificial intelligence (AI) workloads, recently dominated by large language models (LLMs) and vision-language models (VLMs), has intensified power and cooling demands in data centers. This study benchmarks LLMs and VLMs on two HGX nodes, each with 8× NVIDIA H100 graphics processing units (GPUs), using liquid and air cooling. Leveraging GPU Burn, Weights & Biases, and IPMItool, we collect detailed thermal, power, and computation data. Results show that the liquid-cooled systems maintain GPU temperatures between 41-50$^\circ$C, while the air-cooled counterparts fluctuate between 54-72$^\circ$C under load. This thermal stability of liquid-cooled systems yields 17% higher performance (54 TFLOPs/ GPU vs. 46 TFLOPs/GPU), performance-per-watt, reduced energy overhead, and greater system efficiency than the air-cooled counterparts. These findings underscore the energy and sustainability benefits of liquid cooling, offering a compelling path forward for hyperscale data centers seeking to optimize AI infrastructure. https://github.com/iscaas/Cooling-Matters.

Latif, Imran↗

Benchmarking the Performance of Neuromorphic and Spiking Neural Network Simulators

Software simulators play a critical role in the development of new algorithms and system architectures in any field of engineering. Neuromorphic computing, which has shown potential in building brain-inspired energy-efficient hardware, suffers a slow-down in the development cycle due to a lack of flexible and easy-to-use simulators of either neuromorphic hardware itself or of spiking neural networks (SNNs), the type of neural network computation executed on most neuromorphic systems. While there are several openly available neuromorphic or SNN simulation packages developed by a variety of research groups, they have mostly targeted computational neuroscience simulations, and only a few have targeted small-scale machine learning tasks with SNNs. Evaluations or comparisons of these simulators have often targeted computational neuroscience-style workloads. In this work, we seek to evaluate the performance of several publicly available SNN simulators with respect to non-computational neuroscience workloads, in terms of speed, flexibility, and scalability. We evaluate the performance of the NEST, Brian2, Brian2GeNN, BindsNET and Nengo packages under a common front-end neuromorphic framework. Our evaluation tasks include a variety of different network architectures and workload types to mimic the computation common in different algorithms, including feed-forward network inference, genetic algorithms, and reservoir computing. We also study the scalability of each of these simulators when running on different computing hardware, from single core CPU workstations to multi-node supercomputers. Our results show that the BindsNET simulator has the best speed and scalability for most of the SNN workloads (sparse, dense, and layered SNN architectures) on a single core CPU. However, when comparing the simulators leveraging the GPU capabilities, Brian2GeNN outperforms the others for these workloads in terms of scalability. NEST performs the best for small sparse networks and is also the most flexible simulator in terms of reconfiguration capability NEST shows a speedup of at least 2x compared to the other packages when running evolutionary algorithms for SNNs. The multi-node and multi-thread capabilities of NEST show at least 2x speedup compared to the rest of the simulators (single core CPU or GPU based simulators) for large and sparse networks. We conclude our work by providing a set of recommendations on the suitability of employing these simulators for different tasks and scales of operations. We also present the characteristics for a future generic ideal SNN simulator for different neuromorphic computing workloads.

97 MATHEMATICS AND COMPUTING↗

Surrogate models for linear response

Linear response theory is a well-established method in physics and chemistry for exploring excitations of many-body systems. In particular, the quasiparticle random-phase approximation (QRPA) provides a powerful microscopic framework by building excitations on top of the mean-field vacuum; however, its high computational cost limits model calibration and uncertainty quantification studies. Here, we present two complementary QRPA surrogate models and apply them to study response functions of finite nuclei. One is a reduced-order model that exploits the underlying QRPA structure, while the other utilizes the recently developed parametric matrix model algorithm to construct a map between the system’s Hamiltonian and observables. Our benchmark applications, the calculation of the electric dipole polarizability of 180 Yb and the 𝛽-decay half-life of 80 Ni, show that both emulators can achieve 0.1%–1% accuracy while offering a 6–7 orders of magnitude speedup compared to state-of-the-art QRPA solvers. These results demonstrate that the developed QRPA emulators are well positioned to enable Bayesian calibration and large-scale studies of computationally expensive physics models describing the properties of many-body systems.

Beta decay↗

QRCODE: Massively parallelized real-time time-dependent density functional theory for periodic systems

We present a new software module, QRCODE (Quantum Research for Calculating Optically Driven Excitations), for massively parallelized real-time time-dependent density functional theory (RT-TDDFT) calculations of periodic systems in the open-source Qbox software package. Our approach utilizes a custom implementation of a fast Fourier transformation scheme that significantly reduces inter-node message passing interface (MPI) communication of the major computational kernel and shows impressive scaling up to 16,344 CPU cores. In addition to improving computational performance, QRCODE contains a suite of various time propagators for accurate RT-TDDFT calculations. As benchmark applications of QRCODE, we calculate the current density and optical absorption spectra of hexagonal boron nitride (h-BN) and photo-driven reaction dynamics of the ozone-oxygen reaction. We also calculate the second and higher harmonic generation of monolayer and multi-layer boron nitride structures as examples of large material systems. Our optimized implementation of RT-TDDFT in QRCODE enables large-scale calculations of real-time electron dynamics of chemical and material systems with enhanced computational performance and impressive scaling across several thousand CPU cores.

97 MATHEMATICS AND COMPUTING↗

Update on the International Reactor Physics Evaluation Project (IRPhEP) for EGPRS

This set of slides provides an overview of benchmarking work for the International Reactor Physics Evaluation (IRPhE) project, to be presented to the Expert Group on Physics of Reactor Systems (EGPRS), which is an expert group within the OECD/NEA Working Party on Scientific Issues and Uncertainty Analysis of Reactor Systems (WPRS).

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Sensitivity-based Experiment Design Optimization for a Molybdenum Critical Experiment

A lack of intermediate molybdenum benchmarks in the ICSBEP has been identified by LANL, Y-12, and IRSN. This lacking adversely effects criticality safety operations and leaves new differential molybdenum data unvalidated. NCERC is proposing a series of intermediate integral experiments to better the understanding of molybdenum systems. Using MCNP6.2 with the ENDF/B-VIII.0 nuclear data library a single unmoderated and four moderated system designs were identified using a sensitivity optimization method. Each proposed system was found to be at least twice as sensitive to the 95 Mo capture cross section in the URR as the sole existing intermediate molybdenum benchmark in the ISCBEP handbook. The addition of a new molybdenum sensitive intermediate system would improve future nuclear data evaluations.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Statistical Learning for Nonlinear Model Reduction from Local Simulations of Stochastic and Particle- and Agent-Based Systems

Stochastic physical systems across the sciences that have very high-dimensional state spaces, with a large number of fast degrees of freedom that force direct simulators to proceed by integration steps that are orders of magnitude smaller than events of interests (e.g., particle collisions). Examples range from molecular motion to dynamics of large populations of cells. A grand challenge in the simulation and understanding of such systems is the systematic construction of accurate, interpretable, reduced models, enabling faster simulations, revealing fundamental properties of the dynamics, and predicting phenomena of interest that the original simulator could not reached with sufficient accuracy or within a given computational budget. In this projected we developed novel statistical estimation/machine learning techniques for analyzing and building empirical reduced models for important families of high-dimensional stochastic systems, in particular: - we developed techniques for estimating interaction kernels in interacting particle- and agent-based systems, which are ubiquitous in Physics, Biology and many other sciences, given observed trajectories of the system; - we developed techniques for nonlinear model reduction for high-dimensional stochastic systems that have a small number of unknown, nonlinear slow variables, and a large number of fast modes, that are possibly of large magnitude, given observed short trajectories of the system in the form of bursts of trajectories from different initial conditions; - we developed novel techniques for estimating linear dynamical systems on graphs when both the dynamics and the underlying graph are unknown, and we have a sparse set of space-time observations; - we considered the problem of estimating an unknown nonlinear observation function of a standard process (e.g. Brownian motion), so that we can recognized if an observed dynamics is "just" a nonlinear version of a known dynamics; we also developed benchmarks for learning algorithms aimed at learning and classifying diffusion processes.

97 MATHEMATICS AND COMPUTING↗

Evaluation of the Sum-of-Fractions Methodology for Water and Polyethylene Moderated Systems

Sum-of-Fractions is a method intended to assure a subcritical margin for aqueous solutions and slurries of fissionable isotopes. The method indicates that a system is subcritical if the sum of the ratios of the mass of each isotope in a mixture to its individual minimum subcritical mass limit is less than or equal to one. The basis of the Sum-of Fractions has historically been derived from allowances given in ANSI/ANS-8.15-1981. However, the allowance was removed in ANSI/ANS-8.15-2014 due to a lack of technical basis. A methodology was developed to assess the validity of using the Sum-of-Fractions for water or polyethylene moderated systems for the following nuclides: 232 U, 233 U, 234 U, 235 U, 237 Np, 236 Pu, 238 Pu, 239 Pu, 240 Pu, 241 Pu, 242 Pu, 241 Am, 242 mAm, 243 Am, 242 Cm, 243 Cm, 244 Cm, 245 Cm, 246 Cm, 247 Cm, 249 Cf, and 251 Cf. The methodology uses available benchmark data for mixtures of 233 U, 235 U, and 239 Pu to establish the calculational margin, and a mass limit reduction to establish the margin of subcriticality. Water or polyethylene moderated and reflected mixtures containing the nuclides are evaluated with SCALE 6.2.4. Including the calculational margin, subcritical mass limits for each nuclide were computed for optimally water or polyethylene moderated and fully reflected systems. These masses were used to create nuclide mixtures in which the sum of the mass to subcritical mass limit ratios is one. The various nuclide mixtures were modeled over a range of moderation and demonstrate the k eff does not exceed the calculational margin. For additional assurance of subcriticality, a significant mass reduction is applied to each computed minimum critical mass of the nuclides without adequate benchmark data consistent with the method in ANSI/ANS-8.15-2014.

07 ISOTOPE AND RADIATION SOURCES↗

Scalable Predictive Control and Optimization for Grid Integration of Large-Scale Distributed Energy Resources: Preprint

Integration of a large number of distributed energy resources (DERs) into the power grid needs a scalable power balancing method. We formulate the power balancing problem as a look-ahead optimization problem to be solved sequentially by a power distribution system aggregator based on a model predictive control (MPC) framework. Solving large-scale look-ahead control problem requires proper configuration of the control steps. In this paper, to solve large-scale control problems, we propose a variable time granularity where control time steps nearby the current control step have finer resolutions. The aggregator objective includes maximization of power production revenue and minimization of power purchasing expense, renewable power curtailment, and mileage costs for energy storage and electric vehicle (EV) charging stations while satisfying system capacity and operational constraints. The control problem is formulated as a mixed-integer linear program (MILP) and solved using the XpressMP solver. We perform simulations considering a copper plate representation of a large distribution network consisting of 2507 devices (controllable DERs) including curtailable photovoltaics (PVs), energy storage batteries, EV charging stations, and buildings with heating, ventilation, and air conditioning units (HVACs). We show the effectiveness of the proposed approach in managing DERs interactively for maximum energy trading profit and local supply-demand power balancing. Finally, we demonstrate that the proposed method outperformed other benchmark controllers regarding computation time without compromising operational performance.

DER↗

VECTOR Phase 1 Dataset: CAV Trajectory and Energy Consumption Records

This dataset contains benchmark experimental data from Phase 1 of the VECTOR project, focusing on the energy impact of CAV hardware components. The dataset includes vehicle trajectory data (speed and position) and corresponding energy consumption records collected from a CAV platform equipped with lidar, cameras, onboard computation units, and communication modules. The primary objective is to quantify the baseline energy consumption attributable to sensing and computing systems, independent of any advanced cooperative control strategies. During experiments, the leading vehicle followed a predetermined velocity profile, and the following CAV mirrored this trajectory using a basic car-following control to ensure consistent driving behavior. This setup enables a reliable benchmark for assessing the energy cost introduced by onboard CDA hardware (e.g., lidar and GPU-based processing). The dataset is essential for evaluating energy baselines and supports future comparative studies involving additional cooperative strategies. ![system img](system.png) ![vector img](vector.png)

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Simultaneous Optimization of Nuclear–Electronic Orbitals

Accurate modeling of important nuclear quantum effects, such as nuclear delocalization, zero-point energy, and tunneling, as well as non-Born-Oppenheimer effects, requires treatment of both nuclei and electrons quantum mechanically. The nuclear–electronic orbital (NEO) method provides an elegant framework to treat specified nuclei, typically protons, on the same level as the electrons. In conventional electronic structure theory, finding a converged ground state can be a computationally demanding task; converging NEO wavefunctions, due to their coupled electronic and nuclear nature, is even more demanding. Herein, we present an efficient simultaneous optimization method that uses the direct inversion in the iterative subspace method to simultaneously converge wavefunctions for both the electrons and quantum nuclei. In conclusion, benchmark studies show that the simultaneous optimization method can significantly reduce the computational cost compared to the conventional stepwise method for optimizing NEO wavefunctions for multicomponent systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Performance Improvements of the Griffin Solvers in FY24

The Griffin code is a MOOSE-based reactor physics application jointly developed by Idaho National Laboratory and Argonne National Laboratory under the Department of Energy Office of Nuclear Energy Nuclear Energy Advanced Modeling and Simulation Program. This fiscal year, we have made significant efforts to improve the performance of transport solver options and cross-section generation for the efficient use of Griffin in advanced reactor applications. For the HFEM-PN solver, the residual evaluations of HFEM kernels were optimized by utilizing the pre- computed averaged cross sections for individual elements. Numerical integration involving the evaluation of basis functions at quadrature points was bypassed by facilitating precomputed element mass matrices for response matrices. Red-black iterations were improved by introducing a new generalized minimum residual based solver. The memory usage of response matrix storage was significantly reduced by applying basis function rotations on interfaces and calculating volumetric odd-parity moments on the fly. Additionally, the adjoint flux and transient calculation capabilities of the HFEM-PN solver were successfully implemented and verified using the TWIGL benchmark problem. For the DFEM-SN solver, memory footprint and computation time were significantly reduced by not treating angular flux vectors as the MOOSE nonlinear system vectors. Specifically for IQS, scalar adjoint weighting was introduced to further eliminate angular adjoint flux storage in the MOOSE auxiliary system. It was demonstrated through the three-dimensional Advanced Burner Test Reactor core problem that the memory usage for transient calculations with the IQS method was reduced by over 7.5× compared to before the optimizations. For the self-shielding application programming interface, a new double-heterogeneity treatment method, named the Bell Function-Based Analytic Two-Region Slowing Down Method, was developed to efficiently flux-volume homogenize TRISO particles with the matrix. Additionally, optimizations were made to hyper- fine group (HFG) slowing down calculations by pretabulating collision probability coefficients and grouping isotopes, significantly reducing the computational time for calculating scattering sources per HFG. Lastly, the pin power reconstruction module was extended to account for temporal behavior in a microreactor analysis problem, specifically for a control drum transient. Verification tests for each of these improvements demonstrated significant performance enhancements and memory reduction.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

h5bench: A unified benchmark suite for evaluating HDF5 I/O performance on pre‐exascale platforms

Summary Parallel I/O is a critical technique for moving data between compute and storage subsystems of supercomputers. With massive amounts of data produced or consumed by compute nodes, high‐performant parallel I/O is essential. I/O benchmarks play an important role in this process; however, there is a scarcity of I/O benchmarks representative of current workloads on HPC systems. Toward creating representative I/O kernels from real‐world applications, we have created h5bench , a set of I/O kernels that exercise hierarchical data format version 5 (HDF5) I/O on parallel file systems in numerous dimensions. Our focus on HDF5 is due to the parallel I/O library's heavy usage in various scientific applications running on supercomputing systems. The various tests benchmarked in the h5bench suite include I/O operations (read and write), data locality (arrays of basic data types and arrays of structures), array dimensionality (one‐dimensional arrays, two‐dimensional meshes, three‐dimensional cubes), I/O modes (synchronous and asynchronous). In this paper, we present the observed performance of h5bench executed along several of these dimensions on existing supercomputers (Cori and Summit) and pre‐exascale platforms (Perlmutter, Theta, and Polaris). h5bench measurements can be used to identify performance bottlenecks and their root causes and evaluate I/O optimizations. As the I/O patterns of h5bench are diverse and capture the I/O behaviors of various HPC applications, this study will be helpful to the broader supercomputing and I/O community.

97 MATHEMATICS AND COMPUTING↗

How Accurate Are Simulations and Experiments for the Lattice Energies of Molecular Crystals?

Molecular crystals play a central role in a wide range of scientific fields, including pharmaceuticals and organic semiconductor devices. However, they are challenging systems to model accurately with computational approaches because of a delicate interplay of intermolecular interactions such as hydrogen bonding and Van der Waals dispersion forces. Here, by exploiting recent algorithmic developments, we report the first set of diffusion Monte Carlo lattice energies for all 23 molecular crystals in the popular and widely used X23 dataset. Comparisons with previous state-of-the-art lattice energy predictions (on a subset of the dataset) and a careful analysis of experimental sublimation enthalpies reveals that high-accuracy computational methods are now at least as reliable as (computationally derived) experiments for the lattice energies of molecular crystals. Overall, this work demonstrates the feasibility of high-level explicitly correlated electronic structure methods for broad benchmarking studies in complex condensed phase systems, and signposts a route towards closer agreement between experiment and simulation. Published by the American Physical Society 2024

Physics↗

Neural Ordinary Differential Equations for Nonlinear System Identification

Neural ordinary differential equations (NODE) have been recently proposed as a promising approach for nonlinear system identification tasks. In this work, we systematically compare their predictive performance with current state-of-the-art nonlinear and classical linear methods. In particular, we present a quantitative study comparing NODE's performance against neural state-space models and classical linear system identification methods and evaluate their inference speed and prediction performance on open-loop errors across eight different dynamical systems. The experiments show that NODEs can consistently improve the prediction accuracy by order of magnitude compared to benchmark methods. Besides improved accuracy, we also observed that NODEs are less sensitive to hyperparameters compared to neural state-space models by paying the cost of increased computation at the inference time.

machine leaning, system identification, physics in↗

Assessment of Grizzly Capabilities for Reactor Pressure Vessels and Reinforced Concrete Structures

Over the last several years, capabilities to simulate the progression and effects of degradation in critical structures in light water reactor (LWR) nuclear power plants have been under development in the Grizzly code. Age-related material degradation is important for a number of systems in LWRs, but the main focus for Grizzly development has been on reactor pressure vessels (RPVs) and concrete structures because of their central role and the difficulty of replacement of these structures if they are found to be degraded to an unacceptable degree. The capabilities for analyzing both RPVs and concrete structures in Grizzly have reached a point where the feature sets are sufficiently complete to perform credible analyses of those types of structures. Because of this, the emphasis has shifted from foundational development to assessing the accuracy of these modeling capabilities on representative problems of interest and on improving the physical basis of those models to improve their ability to predict the actual response of those structural systems. This report documents a first set of test cases that have been developed to assess the Grizzly code on real-world problems for these two types of structures. For RPVs, these test cases consist of a set of benchmark problems where Grizzly is compared to another code. For concrete structures, a set of models of experimental specimens designed to characterize the multidimensional swelling response of reinforced concrete members due to alkali-silica reaction (ASR) has been developed and compared with experimental results. Both the RPV and concrete test cases developed here are much more computationally intensive than the small regression tests that are used in the testing that Grizzly undergoes every time a proposed set of changes is made to ensure that those changes do not adversely affect previously established behavior. A system for regularly running these large problems to monitor the behavior of the code as it is developed has been instituted. While these test cases generally indicate good comparison with the benchmark results, they also indicate areas where further development is warranted.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Multi-Area Distribution System State Estimation Using Decentralized Physics-Aware Neural Networks

The development of active distribution grids requires more accurate and lower computational cost state estimation. In this paper, the authors investigate a decentralized learning-based distribution system state estimation (DSSE) approach for large distribution grids. The proposed approach decomposes the feeder-level DSSE into subarea-level estimation problems that can be solved independently. The proposed method is decentralized pruned physics-aware neural network (D-P2N2). The physical grid topology is used to parsimoniously design the connections between different hidden layers of the D-P2N2. Monte Carlo simulations based on one-year of load consumption data collected from smart meters for a three-phase distribution system power flow are developed to generate the measurement and voltage state data. The IEEE 123-node system is selected as the test network to benchmark the proposed algorithm against the classic weighted least squares and state-of-the-art learning-based DSSE approaches. Numerical results show that the D-P2N2 outperforms the state-of-the-art methods in terms of estimation accuracy and computational efficiency.

24 POWER TRANSMISSION AND DISTRIBUTION↗