Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “computer system benchmarking”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Dynamic signatures of electronically nonadiabatic coupling in sodium hydride: a rigorous test for the symmetric quasi-classical model applied to realistic, ab initio electronic states in the adiabatic representation

Sodium hydride (NaH) in the gas phase presents a seemingly simple electronic structure making it a potentially tractable system for the detailed investigation of nonadiabatic molecular dynamics from both computational and experimental standpoints. The single vibrational degree of freedom, as well as the strong nonadiabatic coupling that arises from the excited electronic states taking on considerable ionic character, provides a realistic chemical system to test the accuracy of quasi-classical methods to model population dynamics where the results are directly comparable against quantum mechanical benchmarks. Here, using a simulated pump–probe type experiment, this work presents computational predictions of population transfer through the avoided crossings of NaH via symmetric quasi-classical Meyer–Miller (SQC/MM), Ehrenfest, and exact quantum dynamics on realistic, ab initio potential energy surfaces. The main driving force for population transfer arises from the ground vibrational level of the D 1 Σ + adiabatic state that is embedded in the manifold of near-dissociation C 1 Σ + vibrational states. When coupled through a sharply localized first-order derivative coupling most of the population transfers between t = 15 and t = 30 fs depending on the initially excited vibronic wavepacket. While quantum mechanical effects are expected due to the reduced mass of NaH, predictions of the population dynamics from both the SQC/MM and Ehrenfest models perform remarkably well against the quantum dynamics benchmark. Additionally, an analysis of the vibronic structure in the nonadiabatically coupled regime is presented using a variational eigensolver methodology.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Anderson acceleration with approximate calculations: Applications to scientific computing

Here we provide rigorous theoretical bounds for Anderson acceleration (AA) that allow for approximate calculations when applied to solve linear problems. We show that, when the approximate calculations satisfy the provided error bounds, the convergence of AA is maintained while the computational time could be reduced. We also provide computable heuristic quantities, guided by the theoretical error bounds, which can be used to automate the tuning of accuracy while performing approximate calculations. For linear problems, the use of heuristics to monitor the error introduced by approximate calculations, combined with the check on monotonicity of the residual, ensures the convergence of the numerical scheme within a prescribed residual tolerance. Motivated by the theoretical studies, we propose a reduced variant of AA, which consists in projecting the least-squares used to compute the Anderson mixing onto a subspace of reduced dimension. The dimensionality of this subspace adapts dynamically at each iteration as prescribed by the computable heuristic quantities. We numerically show and assess the performance of AA with approximate calculations on: (i) linear deterministic fixed-point iterations arising from the Richardson's scheme to solve linear systems with open-source benchmark matrices with various preconditioners and (ii) non-linear deterministic fixed-point iterations arising from non-linear time-dependent Boltzmann equations.

97 MATHEMATICS AND COMPUTING↗

The multichannel i -propyl + O2 reaction system: A model of secondary alkyl radical oxidation

The i-propyl + O2 reaction mechanism has been investigated by definitive quantum chemical methods to establish this system as a benchmark for the combustion of secondary alkyl radicals. Focal point analyses extrapolating to the ab initio limit were performed based on explicit computations with electron correlation treatments through coupled cluster single, double, triple, and quadruple excitations and basis sets up to cc-pV5Z. The rigorous coupled cluster single, double, and triple excitations/cc-pVTZ level of theory was used to fully optimize all reaction species and transition states, thus, removing some substantial flaws in reference geometries existing in the literature. The vital i-propylperoxy radical (MIN1) and its concerted elimination transition state (TS1) were found 34.8 and 4.4 kcal mol−1 below the reactants, respectively. Two β-hydrogen transfer transition states (TS2, TS2′) lie above the reactants by (1.4, 2.5) kcal mol−1 and display large Born–Oppenheimer diagonal corrections indicative of nearby surface crossings. An α-hydrogen transfer transition state (TS5) is discovered 5.7 kcal mol−1 above the reactants that bifurcates into equivalent α-peroxy radical hanging wells (MIN3) prior to a highly exothermic dissociation into acetone + OH. The reverse TS5 → MIN1 intrinsic reaction path also displays fascinating features, including another bifurcation and a conical intersection of potential energy surfaces. An exhaustive conformational search of two hydroperoxypropyl (QOOH) intermediates (MIN2 and MIN3) of the i-propyl + O2 system located nine rotamers within 0.9 kcal mol−1 of the corresponding lowest-energy minima.

Chemistry↗

hypredrive: high-level interface for solving linear systems with hypre

This software introduces a high-level interface designed to simplify solving linear systems using hypre, a renowned library for such computational challenges. It is crafted to be accessible and user-friendly, making the powerful capabilities of hypre available to a broader audience without requiring in-depth technical knowledge. The interface is characterized by its use of YAML for input, a format celebrated for its structured yet straightforward readability. This choice ensures that users can easily configure the software to meet their specific needs. Additionally, the software boasts an intuitive API that encapsulates hypre's functionalities, making it easier for users to interact with the process of solving linear systems. It is particularly beneficial for prototyping, offering a quick and efficient means to test various solver and preconditioner configurations. Furthermore, the software allows for the creation of an offline testing framework in which predefined linear systems are read from files and benchmarked with user-defined solution strategies. This makes it an invaluable tool for developers and researchers exploring and validating their computational models. Overall, the software serves as a bridge, bringing the advanced computational capabilities of hypre closer to users who may need more specialized technical expertise, thereby facilitating innovation and exploration in the field of numerical linear algebra.

Paludetto Magri, Victor↗

Correlating AGP on a quantum computer

For variational algorithms on the near term quantum computing hardware, it is highly desirable to use very accurate ansatze with low implementation cost. Recent studies have shown that the antisymmetrized geminal power (AGP) wavefunction can be an excellent starting point for ansatze describing systems with strong pairing correlations, as those occurring in superconductors. In this work, we show how AGP can be efficiently implemented on a quantum computer with circuit depth, number of CNOTs, and number of measurements being linear in system size. Using AGP as the initial reference, we propose and implement a unitary correlator on AGP and benchmark it on the ground state of the pairing Hamiltonian. Furthermore, the results show highly accurate ground state energies in all correlation regimes of this model Hamiltonian.

97 MATHEMATICS AND COMPUTING↗

Certifying almost all quantum states with few single-qubit measurements

Certifying that an n -qubit state synthesized in the laboratory is close to a given target state is a fundamental task in quantum information science. However, existing rigorous protocols applicable to general target states have potentially prohibitive resource requirements in the form of either deep quantum circuits or exponentially many single-qubit measurements. Here we prove that almost all n -qubit target states, including those with exponential circuit complexity, can be certified from only O ( n 2 ) single-qubit measurements. Given access to the target state’s amplitudes, our protocol requires only O ( n 3 ) classical computation. This result is established by a technique that relates certification to the mixing time of a random walk. Our protocol has applications for benchmarking quantum systems, for optimizing quantum circuits to generate a desired target state and for learning and verifying neural networks, tensor networks and various other representations of quantum states using only single-qubit measurements. We show that such verified representations can be used to efficiently predict highly non-local properties of a synthesized state that would otherwise require an exponential number of measurements on the state. We demonstrate these applications in numerical experiments with up to 120 qubits and observe an advantage over existing methods such as cross-entropy benchmarking.

information theory and computation↗

System Noise Benchmarks

This project includes benchmarks to assess the presence of system noise on supercomputers. System noise is any activity that interferes with the execution of high-performance computing applications.

Moody, AdamT [Lawrence Livermore National Laborato↗

Verified, Archived, Library of Inputs and Data (VALID) Supporting Files

This dataset contains input, output, and sensitivity data files for computational simulations with the SCALE code system as part of the Verified, Archived Library of Inputs and Data (VALID). The simulations cover critical benchmark experiments from the International Criticality Safety Benchmark Evaluation Project. The files are to be housed in a public directory for distribution. The information contained in the files have been approved for release by the Organisation for Economic Co-operation and Development Nuclear Energy Agency (NEA). Users wanting to reproduce results from this dataset are required to obtain a license to the SCALE code system for which details on the distribution can be found here: https://www.ornl.gov/scale/releases.

keff↗

Particle Swarm Optimization Algorithm for Critical Experiment Design

Nuclear criticality experiments are used to validate nuclear cross section data used by simulation software. This is typically achieved by designing a critical system with a high sensitivity to a certain material’s cross section. Once the experiment has been carried out, a high fidelity model of the system is developed into a benchmark. When this benchmark model is simulated by a transport code, some of the difference between the experimental and computational effective neutron multiplication factor can be attributed to inaccurate nuclear data. Nuclear data evaluators then can make adjustments accordingly to improve cross section data.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Modeling of the Molten Salt Reactor Experiment with SCALE

A SCALE model was developed for the Molten Salt Reactor Experiment (MSRE) benchmark that was recently added to the International Handbook of Evaluated Reactor Physics Benchmark Experiments. This SCALE model served as a basis for criticality calculations and nuclear data sensitivity and uncertainty analyses with the Monte Carlo code Shift and the TSUNAMI computational capabilities in the SCALE code system. The focus of this work is the assessment of the impact of nuclear data on the calculated eigenvalue results in support of the discussion of differences between the calculated and the experimental eigenvalue result. The differences in the eigenvalues obtained using the ENDF/B-VII.0, ENDF/B-VII.1, and ENDF/B-VIII.0 nuclear data libraries cover a relatively small range of ~230 pcm. Since eigenvalue sensitivity of the MSRE is dominated by the neutron multiplicity and neutron capture of 235 U and elastic scattering in graphite, relevant changes in the ENDF/B libraries for nuclear reactions (such as carbon capture) that caused large differences in other graphite-moderated systems did not have a significant impact. Propagation of nuclear data uncertainty results in an eigenvalue uncertainty of ~700 pcm with the major contributors being 235 U neutron multiplicity, graphite elastic scattering, and 7Li neutron capture. All calculations resulted in large differences of ~2000 pcm in eigenvalue compared to the benchmark experimental value. Several potential contributors to this difference—including uncertainties and gaps in the knowledge of the material, geometry, and nuclear data—were identified. Simplified models of the full MSRE core were developed, and similarity assessments were conduced with the full MSRE core model. It was found that simplified models can serve as adequate surrogates of the full-core model such that they can be used for performing selected nuclear data performance assessments with a lower computational burden.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Stochastic tensor contraction for quantum chemistry

Many computational methods in ab initio quantum chemistry are formulated in terms of high-order tensor contractions, whose cost determines the size of system that can be studied. We introduce stochastic tensor contraction to perform such operations with greatly reduced cost, and present its application to the gold-standard quantum chemistry method, coupled cluster theory with up to perturbative triples. For total energy errors more stringent than chemical accuracy, we reduce the computational scaling to that of mean-field theory, while starting to approach the mean-field absolute cost, thereby challenging the existing cost-to-accuracy landscape. Benchmarks against state-of-the-art local correlation approximations further show that we achieve an order-of-magnitude improvement in both total computation time and error, with significantly reduced sensitivity to system dimensionality and electron delocalization. We conclude that stochastic tensor contraction is a powerful computational primitive to accelerate a wide range of quantum chemistry.

Chemical Physics (physics.chem-ph)↗

A multilayer multi-configurational approach to efficiently simulate large-scale circuit-based quantum computers on classical machines

Here, a multilayer multi-configurational theory framework is adapted to simulate circuit-based quantum computers. Quantum addition of superpositions of an exponential number of summands is performed in polynomial time with high accuracy. We demonstrate numerically accurate calculations including up to one million qubits for entangling benchmarks. Simulation cost can be assessed by entropy-based entanglement measures. For the considered systems, we show that the entanglement only grows weakly with the system size. The present simulations demonstrate how quantum algorithms in low-entropy regimes can be used efficiently on classically simulated quantum computers.

97 MATHEMATICS AND COMPUTING↗

The Aerosol Model Benchmarking Repository: A toolkit for model intercomparison

The Aerosol Model Benchmarking Repository and Standards (AMBRS) project was initiated to provide tools and to establish community standards for benchmarking aerosol models. This report describes a set of open-source tools for building, running, and analyzing aerosol box model simulations in a standardized framework. The framework consists of three core components: AMBuilder, a CMake-based build system that compiles supported models consistently; AMBRS, a Python module that defines unified numerical experiments and executes them with aligned inputs; and PyParticle, an aerosol analysis package that standardizes output, computes diagnostics, and visualizes simulation results. Together, these tools enable reproducible intercomparison of aerosol schemes and support process-level evaluation of how model simplifications affect predictions of size distributions, cloud condensation nuclei activity, and other relevant properties relevant for the Earth-Energy system. Beyond its role in benchmarking, AMBRS provides a platform for studying aerosol processes across scales and can be used to generate training data for AI/ML applications in support of a broader hierarchical aerosol modeling strategy.

54 ENVIRONMENTAL SCIENCES↗

Scalable Predictive Control and Optimization for Grid Integration of Large-Scale Distributed Energy Resources

Integrating a large number of distributed energy resources (DERs) into the power grid needs a scalable power balancing method. We formulate the power balancing problem as a look-ahead optimization problem to be solved sequentially by a power distribution system aggregator based on a model predictive control (MPC) framework. Solving large-scale look-ahead control problems requires proper configuration of the control steps. In this paper, to solve large-scale control problems, we propose a variable time granularity where control time steps nearby the current control step have finer resolutions. The aggregator objective includes maximization of power production revenue and minimization of power purchasing expense, renewable power curtailment, and mileage costs for energy storage and electric vehicle (EV) charging stations while satisfying system capacity and operational constraints. The control problem is formulated as a mixed-integer linear program (MILP) and solved using the XpressMP solver. We perform simulations considering a copper plate representation of a large distribution network consisting of 2507 devices (controllable DERs), including curtailable photovoltaics (PVs), energy storage batteries, EV charging stations, and buildings with heating, ventilation, and air conditioning units (HVACs). We show the effectiveness of the proposed approach in managing DERs interactively for maximum energy trading profit and local supply-demand power balancing. Finally, we demonstrate that the proposed method outperforms other benchmark controllers regarding computation time without compromising operational performance.

DER↗

Scalable Predictive Control and Optimization for Grid Integration of Large-Scale Distributed Energy Resources

Integrating a large number of distributed energy resources (DERs) into the power grid needs a scalable power balancing method. We formulate the power balancing problem as a look-ahead optimization problem to be solved sequentially by a power distribution system aggregator based on a model predictive control (MPC) framework. Solving large-scale look-ahead control problems requires proper configuration of the control steps. In this paper, to solve large-scale control problems, we propose a variable time granularity where control time steps nearby the current control step have finer resolutions. The aggregator objective includes maximization of power production revenue and minimization of power purchasing expense, renewable power curtailment, and mileage costs for energy storage and electric vehicle (EV) charging stations while satisfying system capacity and operational constraints. The control problem is formulated as a mixed-integer linear program (MILP) and solved using the XpressMP solver. We perform simulations considering a copper plate representation of a large distribution network consisting of 2507 devices (control-lable DERs), including curtailable photovoltaics (PVs), energy storage batteries, EV charging stations, and buildings with heating, ventilation, and air conditioning units (HVACs). We show the effectiveness of the proposed approach in managing DERs interactively for maximum energy trading profit and local supply-demand power balancing. Finally, we demonstrate that the proposed method outperforms other benchmark controllers regarding computation time without compromising operational performance.

DER↗

Automated Integration of Continental-Scale Observations in Near-Real Time for Simulation and Analysis of Biosphere–Atmosphere Interactions

The National Ecological Observatory Network (NEON) is a continental-scale observatory with sites across the US collecting standardized ecological observations that will operate for multiple decades. To maximize the utility of NEON data, we envision edge computing systems that gather, calibrate, aggregate, and ingest measurements in an integrated fashion. Edge systems will employ machine learning methods to cross-calibrate, gap-fill and provision data in near-real time to the NEON Data Portal and to High Performance Computing (HPC) systems, running ensembles of Earth system models (ESMs) that assimilate the data. For the first time gridded EC data products and response functions promise to offset pervasive observational biases through evaluating, benchmarking, optimizing parameters, and training new machine learning parameterizations within ESMs all at the same model-grid scale. Leveraging open-source software for EC data analysis, we are already building software infrastructure for integration of near-real time data streams into the International Land Model Benchmarking (ILAMB) package for use by the wider research community. We will present a perspective on the design and integration of end-to-end infrastructure for data acquisition, edge computing, HPC simulation, analysis, and validation, where Artificial Intelligence (AI) approaches are used throughout the distributed workflow to improve accuracy and computational performance.

Durden, David J.↗

Cooling Matters: Benchmarking Large Language Models and Vision-Language Models on Liquid-Cooled Versus Air-Cooled H100 GPU Systems

The unprecedented growth in artificial intelligence (AI) workloads, recently dominated by large language models (LLMs) and vision-language models (VLMs), has intensified power and cooling demands in data centers. This study benchmarks LLMs and VLMs on two HGX nodes, each with 8× NVIDIA H100 graphics processing units (GPUs), using liquid and air cooling. Leveraging GPU Burn, Weights & Biases, and IPMItool, we collect detailed thermal, power, and computation data. Results show that the liquid-cooled systems maintain GPU temperatures between 41-50$^\circ$C, while the air-cooled counterparts fluctuate between 54-72$^\circ$C under load. This thermal stability of liquid-cooled systems yields 17% higher performance (54 TFLOPs/ GPU vs. 46 TFLOPs/GPU), performance-per-watt, reduced energy overhead, and greater system efficiency than the air-cooled counterparts. These findings underscore the energy and sustainability benefits of liquid cooling, offering a compelling path forward for hyperscale data centers seeking to optimize AI infrastructure. https://github.com/iscaas/Cooling-Matters.

Latif, Imran↗

Benchmarking the Performance of Neuromorphic and Spiking Neural Network Simulators

Software simulators play a critical role in the development of new algorithms and system architectures in any field of engineering. Neuromorphic computing, which has shown potential in building brain-inspired energy-efficient hardware, suffers a slow-down in the development cycle due to a lack of flexible and easy-to-use simulators of either neuromorphic hardware itself or of spiking neural networks (SNNs), the type of neural network computation executed on most neuromorphic systems. While there are several openly available neuromorphic or SNN simulation packages developed by a variety of research groups, they have mostly targeted computational neuroscience simulations, and only a few have targeted small-scale machine learning tasks with SNNs. Evaluations or comparisons of these simulators have often targeted computational neuroscience-style workloads. In this work, we seek to evaluate the performance of several publicly available SNN simulators with respect to non-computational neuroscience workloads, in terms of speed, flexibility, and scalability. We evaluate the performance of the NEST, Brian2, Brian2GeNN, BindsNET and Nengo packages under a common front-end neuromorphic framework. Our evaluation tasks include a variety of different network architectures and workload types to mimic the computation common in different algorithms, including feed-forward network inference, genetic algorithms, and reservoir computing. We also study the scalability of each of these simulators when running on different computing hardware, from single core CPU workstations to multi-node supercomputers. Our results show that the BindsNET simulator has the best speed and scalability for most of the SNN workloads (sparse, dense, and layered SNN architectures) on a single core CPU. However, when comparing the simulators leveraging the GPU capabilities, Brian2GeNN outperforms the others for these workloads in terms of scalability. NEST performs the best for small sparse networks and is also the most flexible simulator in terms of reconfiguration capability NEST shows a speedup of at least 2x compared to the other packages when running evolutionary algorithms for SNNs. The multi-node and multi-thread capabilities of NEST show at least 2x speedup compared to the rest of the simulators (single core CPU or GPU based simulators) for large and sparse networks. We conclude our work by providing a set of recommendations on the suitability of employing these simulators for different tasks and scales of operations. We also present the characteristics for a future generic ideal SNN simulator for different neuromorphic computing workloads.

97 MATHEMATICS AND COMPUTING↗