Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “computer system benchmarking”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

A Look at the Impact of High-End Computing Technologies on NASA Missions

From its bold start nearly 30 years ago and continuing today, the NASA Advanced Supercomputing (NAS) facility at Ames Research Center has enabled remarkable breakthroughs in the space agency s science and engineering missions. Throughout this time, NAS experts have influenced the state-of-the-art in high-performance computing (HPC) and related technologies such as scientific visualization, system benchmarking, batch scheduling, and grid environments. We highlight the pioneering achievements and innovations originating from and made possible by NAS resources and know-how, from early supercomputing environment design and software development, to long-term simulation and analyses critical to design safe Space Shuttle operations and associated spinoff technologies, to the highly successful Kepler Mission s discovery of new planets now capturing the world s imagination.

Biswas, Rupak↗

Opportunities for enhancing MLCommons efforts while leveraging insights from educational MLCommons earthquake benchmarks efforts

MLCommons is an effort to develop and improve the artificial intelligence (AI) ecosystem through benchmarks, public data sets, and research. It consists of members from start-ups, leading companies, academics, and non-profits from around the world. The goal is to make machine learning better for everyone. In order to increase participation by others, educational institutions provide valuable opportunities for engagement. In this article, we identify numerous insights obtained from different viewpoints as part of efforts to utilize high-performance computing (HPC) big data systems in existing education while developing and conducting science benchmarks for earthquake prediction. As this activity was conducted across multiple educational efforts, we project if and how it is possible to make such efforts available on a wider scale. This includes the integration of sophisticated benchmarks into courses and research activities at universities, exposing the students and researchers to topics that are otherwise typically not sufficiently covered in current course curricula as we witnessed from our practical experience across multiple organizations. As such, we have outlined the many lessons we learned throughout these efforts, culminating in the need for benchmark carpentry for scientists using advanced computational resources. The article also presents the analysis of an earthquake prediction code benchmark while focusing on the accuracy of the results and not only on the runtime; notedly, this benchmark was created as a result of our lessons learned. Energy traces were produced throughout these benchmarks, which are vital to analyzing the power expenditure within HPC environments. Additionally, one of the insights is that in the short time of the project with limited student availability, the activity was only possible by utilizing a benchmark runtime pipeline while developing and using software to generate jobs from the permutation of hyperparameters automatically. It integrates a templated job management framework for executing tasks and experiments based on hyperparameters while leveraging hybrid compute resources available at different institutions. The software is part of a collection called cloudmesh with its newly developed components, cloudmesh-ee (experiment executor) and cloudmesh-cc (compute coordinator).

58 GEOSCIENCES↗

Capillary Driven Flows Along Differentially Wetted Interior Corners

Closed-form analytic solutions useful for the design of capillary flows in a variety of containers possessing interior corners were recently collected and reviewed. Low-g drop tower and aircraft experiments performed at NASA to date show excellent agreement between theory and experiment for perfectly wetting fluids. The analytical expressions are general in terms of contact angle, but do not account for variations in contact angle between the various surfaces within the system. Such conditions may be desirable for capillary containment or to compute the behavior of capillary corner flows in containers consisting of different materials with widely varying wetting characteristics. A simple coordinate rotation is employed to recast the governing system of equations for flows in containers with interior corners with differing contact angles on the faces of the corner. The result is that a large number of capillary driven corner flows may be predicted with only slightly modified geometric functions dependent on corner angle and the two (or more) contact angles of the system. A numerical solution is employed to verify the new problem formulation. The benchmarked computations support the use of the existing theoretical approach to geometries with variable wettability. Simple experiments to confirm the theoretical findings are recommended. Favorable agreement between such experiments and the present theory may argue well for the extension of the analytic results to predict fluid performance in future large length scale capillary fluid systems for spacecraft as well as for small scale capillary systems on Earth.

Golliher, Eric L.↗

Benchmarking quantum computers

The rapid pace of development in quantum computing technology has sparked a proliferation of benchmarks to assess the performance of quantum computing hardware and software. However, not all benchmarks are of equal merit. Good ones empower scientists, engineers, programmers and users to understand the power of a computing system, whereas bad ones can misdirect research and inhibit progress. In this Perspective, we survey the science of quantum computer benchmarking. Here, we discuss the role of benchmarks and benchmarking and how good benchmarks can drive and measure progress towards the long-term goal of useful quantum computations, known as quantum utility. We explain how different kinds of benchmark quantify the performance of different parts of a quantum computer, discuss existing benchmarks, examine recent trends in benchmarking, and highlight important open research questions in this field.

Proctor, Timothy James [Sandia National Laboratori↗

Assessing and advancing the potential of quantum computing: A NASA case study

Quantum computing is one of the most enticing computational paradigms with the potential to revolutionize diverse areas of future-generation computational systems. While quantum computing hardware has advanced rapidly, from tiny laboratory experiments to quantum chips that can outperform even the largest supercomputers on specialized computational tasks, these noisy-intermediate scale quantum (NISQ) processors are still too small and non-robust to be directly useful for any real-world applications. In this paper, we describe NASA’s work in assessing and advancing the potential of quantum computing. We discuss advances in algorithms, both near- and longer-term, and the results of our explorations on current hardware as well as with simulations, including illustrating the benefits of algorithm-hardware co-design in the NISQ era. This work also includes physics-inspired classical algorithms that can be used at application scale today. We discuss innovative tools supporting the assessment and advancement of quantum computing and describe improved methods for simulating quantum systems of various types on high-performance computing systems that incorporate realistic error models. We provide an overview of recent methods for benchmarking, evaluating, and characterizing quantum hardware for error mitigation, as well as insights into fundamental quantum physics that can be harnessed for computational purposes.

Rieffel, Eleanor G.↗

In situ temperature measurements in sooting methane/air flames using synchrotron x-ray fluorescence of seeded krypton atoms

Synchrotron x-ray fluorescence has been used to measure temperatures in optically dense gases where traditional methods would fail. These data provide a benchmark for stringent tests of computational fluid dynamics models for complex systems where physical and chemical processes are intimately linked. The experiments measured krypton number densities in a sooting, atmospheric pressure, nonpremixed coflow flame that is widely used in combustion research. The experiments not only form targets for the models, but the simulations also identify potential sources of uncertainties in the measurements, allowing for future improvements.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Advances and trends in the development of computational models for tires

Status and some recent developments of computational models for tires are summarized. Discussion focuses on a number of aspects of tire modeling and analysis including: tire materials and their characterization; evolution of tire models; characteristics of effective finite element models for analyzing tires; analysis needs for tires; and impact of the advances made in finite element technology, computational algorithms, and new computing systems on tire modeling and analysis. An initial set of benchmark problems has been proposed in concert with the U.S. tire industry. Extensive sets of experimental data will be collected for these problems and used for evaluating and validating different tire models. Also, the new Aircraft Landing Dynamics Facility (ALDF) at NASA Langley Research Center is described.

Noor, A. K.↗

Virtualized Infiniband: Enabling HPC in the Cloud

Presentation describes virtualized infiniband and how it enables doing high performance computing level processing and file systems in a cloud environment. Generic benchmarks are included.

virtualization↗

Performance Characteristics of the Multi-Zone NAS Parallel Benchmarks

We describe a new suite of computational benchmarks that models applications featuring multiple levels of parallelism. Such parallelism is often available in realistic flow computations on systems of grids, but had not previously been captured in bench-marks. The new suite, named NPB Multi-Zone, is extended from the NAS Parallel Benchmarks suite, and involves solving the application benchmarks LU, BT and SP on collections of loosely coupled discretization meshes. The solutions on the meshes are updated independently, but after each time step they exchange boundary value information. This strategy provides relatively easily exploitable coarse-grain parallelism between meshes. Three reference implementations are available: one serial, one hybrid using the Message Passing Interface (MPI) and OpenMP, and another hybrid using a shared memory multi-level programming model (SMP+OpenMP). We examine the effectiveness of hybrid parallelization paradigms in these implementations on three different parallel computers. We also use an empirical formula to investigate the performance characteristics of the multi-zone benchmarks.

Jin, Haoqiang↗

Insights from Optimizing HPL Performance on Exascale Systems: A Comparative Analysis of Panel Factorization

High performance LINPACK (HPL) remains the primary benchmark for evaluating supercomputing performance. It includes many parts with substantial internal complexity, and its performance is affected by a large number of parameters that interact in ways that are difficult to predict on large-scale heterogeneous supercomputer systems. We present a comprehensive performance analysis of HPL on Frontier, the world’s first exascale supercomputer, which achieved HPL performance of 1.35 exaflops. Through empirical parameter tuning, detailed modeling, and comparative evaluation, we uncover critical performance insights, share lessons learned, and outline best practices for effective parameter tuning on exascale systems. We introduce and evaluate two novel PDFACT strategies: a dedicated-thread (DT) variant and a GPU-based variant (GPUPDFACT) implementation using HIP cooperative groups, demonstrating that GPU-based factorization outperforms conventional CPU-based PDFACT on Frontier’s architecture. Our findings establish key performance factors for HPL on exascale systems and offer valuable guidance for future high-performance computing and benchmarking efforts.

Lu, Hao [ORNL] (ORCID:000000018941870X)↗

NCCS High Performance GMRES Mixed Precision

HPG-MxP is a software package that performs a fixed number of multigrid preconditioned (using a Gauss-Seidel smoother) Generalized minimal residual (PGMRES) iterations in order to solve a possibly nonsymmetric large sparse linear system of equations. It is designed to be a benchmark to measure a computer's performance for sparse linear algebra workloads typical in scientific computing while allowing the use of mixed precision methods. The solution is required to have convergence characteristics and accuracy similar to double precision GMRES. It is based on the High Performance Conjugate Gradient Benchmark (HPCG) which restricts all implementations to use only the IEEE double precision format (FP64). The original implementation (https://github.com/hpg-mxp/hpg-mxp) was written by Ichitaro Yamazaki, Jennifer Loe, Christian Glusa, Sivasankaran Rajamanickam, Piotr Luszczek, and Jack Dongarra. Please refer to that repository for documentation on the original implementation. This version is maintained by the National Center for Computational Sciences at Oak Ridge National Laboratory. It is highly scalable and optimized for Oak Ridge Leadership Computing Facility (OLCF) systems, particularly Frontier.

Kashi, Aditya [Oak Ridge National Laboratory (ORNL↗

Heterogeneous Distributed Computing for Computational Aerosciences

The research supported under this award focuses on heterogeneous distributed computing for high-performance applications, with particular emphasis on computational aerosciences. The overall goal of this project was to and investigate issues in, and develop solutions to, efficient execution of computational aeroscience codes in heterogeneous concurrent computing environments. In particular, we worked in the context of the PVM[1] system and, subsequent to detailed conversion efforts and performance benchmarking, devising novel techniques to increase the efficacy of heterogeneous networked environments for computational aerosciences. Our work has been based upon the NAS Parallel Benchmark suite, but has also recently expanded in scope to include the NAS I/O benchmarks as specified in the NHT-1 document. In this report we summarize our research accomplishments under the auspices of the grant.

Sunderam, Vaidy S.↗

Unimolecular Dynamics of Partially Deuterated Hydroperoxyalkyl Intermediates (•QOOD) in Cyclohexane and Cyclopentane Oxidation

The hydroperoxyalkyl radicals (•QOOH) formed in cyclohexane and cyclopentane oxidation with carbon radical center at the β site relative to the hydroperoxy group have recently been observed through their infrared fingerprint and time- and energy-resolved unimolecular dissociation dynamics to hydroxyl (OH) radical and cyclic ether products. Partial deuteration shifts the IR transitions associated with OH/OD motion to lower frequencies, providing a new window on the IR spectroscopy and unimolecular dynamics of these transient intermediates. Specifically, the overtone OD stretch (2ν OD ) has been identified at 5215.0 cm –1 and 5216.0 cm –1 for the β-QOOD intermediates in cyclohexane and cyclopentane oxidation, respectively. Unimolecular decay rates of the •QOOD intermediates at these energies have been obtained from the time-resolved appearance of OD products. The experimental rates are compared with statistical microcanonical rates evaluated using RRKM theory, including heavy-atom tunneling associated with simultaneous O–O bond elongation and C–O–O angle contraction along the reaction pathway. The experimental rates provide a further test of the previously determined transition state (TS) barriers, which were computed utilizing a benchmarking approach that builds on higher-level reference calculations for the smaller ethane oxidation system. Here, the experimental rate measurements for both •QOOD intermediates agree well with the computed rates after accounting for small changes in zero-point energy and a minor empirical adjustment of the TS barriers in both systems.

Ethers↗

Enhancing f -Element Separations with ADAAM-EH: The Impact of Phase Modifiers and a DGA Aqueous Complexant

Recent investigations have used a 2-ethylhexyl diamide amine (ADAAM-EH) for Am/Cm separations in combination with N,N,N ',N '-tetraethyldiglycolamide as an aqueous complexant to achieve an unprecedented separation factor of 41. The aim of this research effort is to understand the speciation of trivalent lanthanide (Ln) and actinide (An) ions in the organic phase of an ADAAM-EH extraction system, both with and without phase modifiers (PM) (1-octanol and tri-n-butyl phosphate (TBP)). Leveraging spectroscopic techniques in combination with distribution ratio measurements provides an understanding of organic phase f-element ligand complexation. In the absence of PM, Ln is extracted in a stoichiometric 1:1 [M(ADAAM-EH) 1 (NO 3 ) x (H 2 O) 1 ](NO 3 ) 3-x complex. The addition of 1-octanol at 20 vol % results in multiple species present. One of the species is the same as the no PM case, and the other species results in an increased -OH coordination to the inner sphere, potentially displacing some NO 3 . In the case of TBP, increasing concentration results in additional red-shifted bands in the UV-visible spectra, suggesting the complexation of additional ligands of either ADAAM-EH or TBP. Finally, the new system knowledge obtained by and spectroscopic experiments will provide benchmarking information for computational studies of the inner- and outer-sphere coordination environments of f-element cations and insights into ADAAM-EH adduct formation with PM, like 1-octanol and TBP.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Electronic Structure and Bonding of US, SUO, and US 2

Anion photoelectron spectra of US – and US 2 – were recorded using the third (355 nm) and fourth (266 nm) harmonics of an Nd:YAG laser, which yielded vertical detachment energies (VDEs) of 1.71 and 2.02 eV, respectively. The experimental results are supported by extensive relativistic ab initio calculations, primarily at the coupled cluster level of theory, with systematic sequences of correlation consistent basis sets. Calculations include the closely related SUO and SUO – molecules, as well as the oxide congeners UO/UO – and UO 2 /UO 2 – which are well-known experimentally and provide benchmark systems for the sulfide calculations. Adiabatic electron detachment energies (ADEs) are computed for UO – , UO 2 – , US – , SUO – , and US 2 – using the Feller–Peterson–Dixon (FPD) composite approach. Additionally, ADEs are determined for UO – and US – using a spinor-based coupled cluster approach where spin–orbit coupling is included at the orbital level. VDEs are derived from the ab initio results from Franck–Condon simulations of the photoelectron spectra.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

The accuracies of effective interactions in downfolding coupled-cluster approaches for small-dimensionality active spaces

Here, this paper evaluates the accuracy of the Hermitian form of the downfolding procedure using the double unitary coupled cluster (DUCC) ansatz on the benchmark systems of linear chains of hydrogen atoms, H6 and H8. The computational infrastructure employs the occupation-number-representation codes to construct the matrix representation of arbitrary second-quantized operators, allowing for the exact representation of exponentials of various operators. The tests demonstrate that external amplitudes from standard single-reference coupled cluster methods that sufficiently describe external (out-of-active-space) correlations reliably parameterize the Hermitian downfolded effective Hamiltonians in the DUCC formalism. The results show that this approach can overcome the problems associated with losing the variational character of corresponding energies in the corresponding SR-CC theories.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Empirically-calibrated H100 node power models for accurate AI training energy estimation

Accurately quantifying the energy use of artificial intelligence (AI) training is critical for infrastructure planning, carbon accounting, and sustainable data center operation, but few studies have directly measured the power consumption of production workloads on contemporary hardware. By combining empirical measurements from Brookhaven National Laboratory during AI training on 8-graphics-processing-unit H100 systems with open-source benchmarking data, we develop statistical models relating computational intensity to node-level power consumption. We measure the gap between manufacturer-rated thermal design power (TDP) and actual power demand during AI training. Our analysis reveals that even computationally intensive workloads operate at only 76% of the 10.2 kW TDP rating. Our architecture-specific model, calibrated to floating-point operations, predicts energy consumption with 11.4% mean absolute percentage error, significantly outperforming TDP-based approaches (27%–37% error). We identified distinct power signatures between transformer and convolutional neural network architectures, with transformers showing characteristic fluctuations that may impact grid stability. These results provide a measurement-grounded basis for improving AI training energy estimates, enabling more reliable infrastructure sizing, cost projections, and environmental impact assessments.

Newkirk, Alex C↗

An evaluative model of system performance in manned teleoperational systems

Manned teleoperational systems are used in aerospace operations in which humans must interact with machines remotely. Manual guidance of remotely piloted vehicles, controling a wind tunnel, carrying out a scientific procedure remotely are examples of teleoperations. A four input parameter throughput (Tp) model is presented which can be used to evaluate complex, manned, teleoperations-based systems and make critical comparisons among candidate control systems. The first two parameters of this model deal with nominal (A) and off-nominal (B) predicted events while the last two focus on measured events of two types, human performance (C) and system performance (D). Digital simulations showed that the expression A(1-B)/C+D) produced the greatest homogeneity of variance and distribution symmetry. Results from a recently completed manned life science telescience experiment will be used to further validate the model. Complex, interacting teleoperational systems may be systematically evaluated using this expression much like a computer benchmark is used.

Haines, Richard F.↗