Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “computer system benchmarking”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

An evaluative model of system performance in manned teleoperational systems

Manned teleoperational systems are used in aerospace operations in which humans must interact with machines remotely. Manual guidance of remotely piloted vehicles, controling a wind tunnel, carrying out a scientific procedure remotely are examples of teleoperations. A four input parameter throughput (Tp) model is presented which can be used to evaluate complex, manned, teleoperations-based systems and make critical comparisons among candidate control systems. The first two parameters of this model deal with nominal (A) and off-nominal (B) predicted events while the last two focus on measured events of two types, human performance (C) and system performance (D). Digital simulations showed that the expression A(1-B)/C+D) produced the greatest homogeneity of variance and distribution symmetry. Results from a recently completed manned life science telescience experiment will be used to further validate the model. Complex, interacting teleoperational systems may be systematically evaluated using this expression much like a computer benchmark is used.

Haines, Richard F.↗

How Accurate Are Approximate Density Functionals for Noncovalent Interaction of Very Large Molecular Systems?

Noncovalent intermolecular interactions are very important in many research areas. Therefore, it is vital to understand the extent to which approximate density functionals give a proper description of noncovalent interactions. Previous research has demonstrated that some approximate density functionals can predict usefully accurate interaction energies for many noncovalent systems; however, most of that work is limited to small and moderate-sized molecules. Very recently though, accurate benchmarks have become available for some very large molecules. Here, the present work applies 21 approximate density functionals to compute the binding energies of seven large molecular systems that have a number of atoms ranging from 200 to 910. The results are judged by comparison to the recently published CIM-DLPNO-CCSD(T) results, which are assumed to provide a reliable benchmark. The five most accurate methods among those tested are found to be PW6B95-D4, PW6B95-D3(BJ), revM11, M06-L, and MN15.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Open‐Source Anaerobic Digestion Modeling Platform, Anaerobic Digestion Model No. 1 Fast (ADM1F)

An open‐source modeling platform, called Anaerobic Digestion Model No. 1 Fast (ADM1F), is introduced to achieve fast and numerically stable simulations of anaerobic digestion processes. ADM1F is compatible with an iPython interface to facilitate model configuration, simulation, data analysis, and visualization. Faster simulations and more stable results are accomplished by implementing an advanced open‐source library of numerical methods called Portable Extensive Toolkit for Scientific Computation (PETSc) to solve the ADM1 system of equations. Leveraging PETSc, ADM1F can consistently complete a steady‐state simulation under 0.2 s, over 99% faster than a benchmark ADM1 model implemented with MATLAB while achieving agreement of model outputs within 1% of those obtained with the benchmark model. For dynamic simulations, however, ADM1F has a computational speed advantage only when the influent characteristics update more frequently than every 4 h. The ability of ADM1F to be useful as a tool to study anaerobic digestion systems is demonstrated through two example implementations of ADM1F: (1) a two‐phase co‐digestion scenario evaluating the impact of the organic loading rate and the substrate composition on reactor performance and stability, and (2) a conventional digester scenario assessing the effectiveness of recovery strategies after disruptions that led to instability. These examples demonstrate how the high simulation speed and the convenience of the iPython interface allow ADM1F to complete complex analyses within minutes, much faster than computational strategies currently reported in the literature.

anaerobic co-digestion↗

Review of Experimental Data for Validating Computer Codes Used in Shielding Calculations for Spent Fuel Storage and Transportation Systems

This report presents a review of available radiochemical assay data and shielding benchmarks applicable to spent nuclear fuel (SNF) shielding calculations. The relevant information reviewed herein includes the Spent Fuel Composition (SFCOMPO) database, the Shielding Integral Benchmark Archive and Database (SINBAD), the International Handbook of Evaluated Criticality Safety Benchmark Experiments, and published measurements of external dose rates of casks loaded with SNF. The relevant experimental data identified in this report may be used to support verification and validation of computer codes used in SNF cask/transport shielding applications, as well as development of calculation uncertainties. It should be noted that a relatively small subset of the identified experimental data (e.g., criticality alarm experiments) is available in a standard format established by the international community participating in experimental isotopic and shielding data evaluations. An effort of the SFCOMPO Technical Review Group (TRG) is underway to publish first isotopic evaluations of individual assay data using a standard data evaluation format. The SINBAD TRG has recently initiated benchmark evaluations and modernization of the database. Therefore, more relevant information is expected in the future that will enable users to select quality experimental data in depletion code and shielding code validations for SNF applications.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

CommBench: Micro-Benchmarking Hierarchical Networks with Multi-GPU, Multi-NIC Nodes

Modern high-performance computing systems have multiple GPUs and network interface cards (NICs) per node. The resulting network architectures have multilevel hierarchies of subnetworks with different interconnect and software technologies. These systems offer multiple vendor-provided communication capabilities and library implementations (IPC, MPI, NCCL, RCCL, OneCCL) with APIs providing varying levels of performance across the different levels. Understanding this performance is currently difficult because of the wide range of architectures and programming models (CUDA, HIP, OneAPI). We present CommBench, a library with cross-system portability and a high-level API that enables developers to easily build microbenchmarks relevant to their use cases and gain insight into the performance (bandwidth & latency) of multiple implementation libraries on different networks. We demonstrate CommBench with three sets of microbenchmarks that profile the performance of six systems. Our experimental results reveal the effect of multiple NICs on optimizing the bandwidth across nodes and also present the performance characteristics of four available communication libraries within and across nodes of NVIDIA, AMD, and Intel GPU networks.

Hidayetoglu, Mert↗

Benchmark Testing on the IBM-Q Network

The goal of this project is to evaluate a proposed set of candidate benchmarks being developed by the Standards and Performance Metrics Technical Advisory Group (TAG) of the Quantum Economic Development Consortium (QED-C), of which LANL is a member. The QED-C Standards and Performance Metrics TAG has developed implementations of several benchmark codes with the hope and expectation that these codes will be helpful to QED-C members interested in investigating various quantum computer platforms. The benchmark set will be run on LANL’s access to the IBM-Q system with the intent to evaluate the set for scalability, correct execution, and coverage of the application space. The goal is for the QED-C Standards and Performance Metrics TAC to be able to produce a coherent, consistent, scalable set of benchmarks that will run on multiple quantum computing platforms and will be available to members of QED-C, including LANL.

97 MATHEMATICS AND COMPUTING↗

Turnkey CAD/CAM selection and evaluation

The methodology to be followed in evaluating and selecting a computer system for manufacturing applications is discussed. Main frames and minicomputers are considered. Benchmark evaluations, demonstrations, and contract negotiations are discussed.

Moody, T.↗

Initial Performance Results on IBM POWER6

The POWER5+ processor has a faster memory bus than that of the previous generation POWER5 processor (533 MHz vs. 400 MHz), but the measured per-core memory bandwidth of the latter is better than that of the former (5.7 GB/s vs. 4.3 GB/s). The reason for this is that in the POWER5+, the two cores on the chip share the L2 cache, L3 cache and memory bus. The memory controller is also on the chip and is shared by the two cores. This serializes the path to memory. For consistently good performance on a wide range of applications, the performance of the processor, the memory subsystem, and the interconnects (both latency and bandwidth) should be balanced. Recognizing this, IBM has designed the Power6 processor so as to avoid the bottlenecks due to the L2 cache, memory controller and buffer chips of the POWER5+. Unlike the POWER5+, each core in the POWER6 has its own L2 cache (4 MB - double that of the Power5+), memory controller and buffer chips. Each core in the POWER6 runs at 4.7 GHz instead of 1.9 GHz in POWER5+. In this paper, we evaluate the performance of a dual-core Power6 based IBM p6-570 system, and we compare its performance with that of a dual-core Power5+ based IBM p575+ system. In this evaluation, we have used the High- Performance Computing Challenge (HPCC) benchmarks, NAS Parallel Benchmarks (NPB), and four real-world applications--three from computational fluid dynamics and one from climate modeling.

Saini, Subbash↗

Qutrit Randomized Benchmarking

Ternary quantum processors offer significant potential computational advantages over conventional qubit technologies, leveraging the encoding and processing of quantum information in qutrits (three-level systems). Therefore, to evaluate and compare the performance of such emerging quantum hardware it is essential to have robust benchmarking methods suitable for a higher-dimensional Hilbert space. We demonstrate extensions of industry standard randomized benchmarking (RB) protocols, developed and used extensively for qubits, suitable for ternary quantum logic. Using a superconducting five-qutrit processor, we find an average single-qutrit process infidelity of 3.8×10 -3 . Through interleaved RB, we characterize a few relevant gates, and employ simultaneous RB to fully characterize crosstalk errors. Finally, we apply cycle benchmarking to a two-qutrit CSUM gate and obtain a two-qutrit process fidelity of 0.85. Our results present and demonstrate RB-based tools to characterize the performance of a qutrit processor, and a general approach to diagnose control errors in future qudit hardware.

97 MATHEMATICS AND COMPUTING↗

SMR safety through HTTF modeling and benchmark efforts for code validation for gas-cooled reactor applications

Accurate modeling and simulation tools for thermal-hydraulics calculations are a key element needed to design and license new advanced reactors including Small Modular Reactors (SMR) and Microreactors. Uncertainties in modeling and simulation can have significant safety and economic implications. The High Temperature Test Facility (HTTF) at Oregon State University (OSU) is a scaled integral effects experiment designed to investigate transient behavior in high-temperature gas-cooled prismatic-block nuclear reactors. High-quality measurement data is available from the HTTF that is suitable for a thermal-hydraulics code validation benchmark for gas-cooled reactor simulations. Here, this paper summarizes individual HTTF modeling efforts to date for tool validation at Idaho National Laboratory (INL), Argonne National Laboratory (ANL), Oregon State University (OSU) and Canadian Nuclear Laboratories (CNL) using system thermal-hydraulics codes, Computational Fluid Dynamics (CFD) codes and system-CFD code couplings. Also, the paper introduces the ongoing OECD Nuclear Energy Agency (NEA) High Temperature Gas Reactor Thermal-Hydraulics (HTGR T/H) benchmark that allows for better comparisons of results between different international modeling teams. The benchmark provides well defined computational problems that include code-to-code comparisons and comparisons to measured data. These problems provide an avenue for quantifying accuracy and identifying sources of uncertainty in thermal-hydraulics calculations, including in measured thermophysical properties, as part of validation for gas-cooled reactor simulation tools.

21 - SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLAN↗

Entanglement Benchmarking in Quantum Simulations of Spin Systems

We simulate quantum spin systems and measure entanglement using circuits tailored for near-term quantum computers. Traditional tools like entanglement entropy are limited to pure states and require full state tomography, making them impractical on current hardware. Instead, we employ the novel approach, Positive Partial Transpose (PPT) criterion to efficiently detect pairwise entanglement from two-spin reduced density matrices, applicable to both pure and mixed states. This method enables scalable entanglement detection, providing a practical route to study quantum correlations, phase transitions, and benchmark quantum devices.

Baul, Anshumitra [ORNL] (ORCID:0000000268947191)↗

Validation of time-dependent shift using the pulsed sphere benchmarks

The detailed behavior of neutrons in a rapidly changing time-dependent physical system is a challenging computational physics problem, particularly when using Monte Carlo methods on heterogeneous high-performance computing architectures. A small number of algorithms and code implementations have been shown to be performant for time-independent (fixed source and k-eigenvalue) Monte Carlo, and there are existing simulation tools that successfully solve the time-dependent Monte Carlo problem on smaller computing platforms. To bridge this gap, a time-dependent version of ORNL’s Shift code has been recently developed. Shift’s history-based algorithm on CPUs, and its event-based algorithm on GPUs, have both been observed to scale well to very large numbers of processors, which motivated the extension of this code to solve time-dependent problems. The validation of this new capability requires a comparison with time-dependent neutron experiments. Lawrence Livermore National Laboratory’s (LLNL) pulsed sphere benchmark experiments were simulated in Shift to validate both the time-independent as well as new time-dependent features recently incorporated into Shift. A suite of pulsed-sphere models was simulated using Shift and compared to the available experimental data and simulations with MCNP. Overall results indicate that Shift accurately simulates the pulsed sphere benchmarks, and that the new time-dependent modifications of Shift are working as intended. Validated exascale neutron transport codes are essential for a wide variety of future multiphysics applications.

Palmer, Camille J.↗

RISC Processors and High Performance Computing

This tutorial will discuss the top five RISC microprocessors and the parallel systems in which they are used. It will provide a unique cross-machine comparison not available elsewhere. The effective performance of these processors will be compared by citing standard benchmarks in the context of real applications. The latest NAS Parallel Benchmarks, both absolute performance and performance per dollar, will be listed. The next generation of the NPB will be described. The tutorial will conclude with a discussion of future directions in the field. Technology Transfer Considerations: All of these computer systems are commercially available internationally. Information about these processors is available in the public domain, mostly from the vendors themselves. The NAS Parallel Benchmarks and their results have been previously approved numerous times for public release, beginning back in 1991.

Bailey, David H.↗

Comparison of Origin 2000 and Origin 3000 Using NAS Parallel Benchmarks

This report describes results of benchmark tests on the Origin 3000 system currently being installed at the NASA Ames National Advanced Supercomputing facility. This machine will ultimately contain 1024 R14K processors. The first part of the system, installed in November, 2000 and named mendel, is an Origin 3000 with 128 R12K processors. For comparison purposes, the tests were also run on lomax, an Origin 2000 with R12K processors. The BT, LU, and SP application benchmarks in the NAS Parallel Benchmark Suite and the kernel benchmark FT were chosen to determine system performance and measure the impact of changes on the machine as it evolves. Having been written to measure performance on Computational Fluid Dynamics applications, these benchmarks are assumed appropriate to represent the NAS workload. Since the NAS runs both message passing (MPI) and shared-memory, compiler directive type codes, both MPI and OpenMP versions of the benchmarks were used. The MPI versions used were the latest official release of the NAS Parallel Benchmarks, version 2.3. The OpenMP versiqns used were PBN3b2, a beta version that is in the process of being released. NPB 2.3 and PBN 3b2 are technically different benchmarks, and NPB results are not directly comparable to PBN results.

Turney, Raymond D.↗

Benchmarking highly entangled states on a 60-atom analogue quantum simulator

Abstract Quantum systems have entered a competitive regime in which classical computers must make approximations to represent highly entangled quantum states 1,2 . However, in this beyond-classically-exact regime, fidelity comparisons between quantum and classical systems have so far been limited to digital quantum devices 2–5 , and it remains unsolved how to estimate the actual entanglement content of experiments 6 . Here, we perform fidelity benchmarking and mixed-state entanglement estimation with a 60-atom analogue Rydberg quantum simulator, reaching a high-entanglement entropy regime in which exact classical simulation becomes impractical. Our benchmarking protocol involves extrapolation from comparisons against an approximate classical algorithm, introduced here, with varying entanglement limits. We then develop and demonstrate an estimator of the experimental mixed-state entanglement 6 , finding our experiment is competitive with state-of-the-art digital quantum devices performing random circuit evolution 2–5 . Finally, we compare the experimental fidelity against that achieved by various approximate classical algorithms, and find that only the algorithm we introduce is able to keep pace with the experiment on the classical hardware we use. Our results enable a new model for evaluating the ability of both analogue and digital quantum devices to generate entanglement in the beyond-classically-exact regime, and highlight the evolving divide between quantum and classical systems.

Science & Technology - Other Topics↗

Numerically exact configuration interaction at quadrillion-determinant scale

The combinatorial growth of configuration interaction (CI) has long limited this formally exact quantum chemistry method to only the smallest molecules. Here, we report a numerically exact CI calculation exceeding one quadrillion (10 15 ) determinants, made possible by a lossless categorical compression strategy within the small-tensor-product distributed active space (STP-DAS) framework. This approach overcomes the traditional memory bottlenecks of CI by a numerically exact compression of the wavefunction representation and reformulating the most computationally demanding matrix–vector operations. Using this method, we performed a fully relativistic CI calculation of the ground state of HBrTe with over 10 15 complex-valued determinants in just 34.5 h on 1000 computing nodes—the largest CI calculation ever reported. We further achieved fast computation for systems with hundreds of billions of determinants on only a few compute nodes. Extensive benchmarks confirm that the method retains full numerical exactness while cutting memory and computational cost by orders of magnitude. Compared to previous state-of-the-art CI calculations, this work achieves a 1000 times increase in CI space, a 10 6 -fold increase in floating-point operations performed, and a 10 6 -fold improvement in computational speed.

Computational chemistry↗

Application Experiences on a GPU-Accelerated Arm-based HPC Testbed

This paper assesses and reports the experience of ten teams working to port, validate, and benchmark several High Performance Computing applications on a novel GPU-accelerated Arm testbed system. The testbed consists of eight NVIDIA Arm HPC Developer Kit systems, each one equipped with a server-class Arm CPU from Ampere Computing and two data center GPUs from NVIDIA Corp. The systems are connected together using InfiniBand interconnect. The selected applications and mini-apps are written using several programming languages and use multiple accelerator-based programming models for GPUs such as CUDA, OpenACC, and OpenMP offloading. Working on application porting requires a robust and easy-to-access programming environment, including a variety of compilers and optimized scientific libraries. The goal of this work is to evaluate platform readiness and assess the effort required from developers to deploy well-established scientific workloads on current and future generation Arm-based GPU-accelerated HPC systems. The reported case studies demonstrate that the current level of maturity and diversity of software and tools is already adequate for large-scale production deployments.

Elwasif, Wael↗

PETSc/TAO developments for GPU-based early exascale systems

The Portable Extensible Toolkit for Scientific Computation (PETSc) library provides scalable solvers for nonlinear time-dependent differential and algebraic equations and for numerical optimization via the Toolkit for Advanced Optimization (TAO). PETSc is used in dozens of scientific fields and is an important building block for many simulation codes. During the U.S. Department of Energy’s Exascale Computing Project, the PETSc team has made substantial efforts to enable efficient utilization of the massive fine-grain parallelism present within exascale compute nodes and to enable performance portability across exascale architectures. We recap some of the challenges that designers of numerical libraries face in such an endeavor, and then discuss the many developments we have made, which include the addition of new GPU backends, features supporting efficient on-device matrix assembly, better support for asynchronicity and GPU kernel concurrency, and new communication infrastructure. In conclusion, we evaluate the performance of these developments on some pre-exascale systems as well as the early exascale systems Frontier and Aurora, using compute kernel, communication layer, solver, and mini-application benchmark studies, and then close with a few observations drawn from our experiences on the tension between portable performance and other goals of numerical libraries.

Exascale Computing Project (ECP)↗