Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Performance benchmark”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23

The NAS Parallel Benchmarks 2.1 Results

We present performance results for version 2.1 of the NAS Parallel Benchmarks (NPB) on the following architectures: IBM SP2/66 MHz; SGI Power Challenge Array/90 MHz; Cray Research T3D; and Intel Paragon. The NAS Parallel Benchmarks are a widely-recognized suite of benchmarks originally designed to compare the performance of highly parallel computers with that of traditional supercomputers.

Saphir, William↗

Comparison of UNL laser imaging and sizing system and a phase Doppler system for analyzing sprays from a NASA nozzle

Research was conducted on characteristics of aerosol sprays using a P/DPA and a laser imaging/video processing system on a NASA MOD-1 air assist nozzle being evaluated for use in aircraft icing research. Benchmark tests were performed on monodispersed particles and on the NASA MOD-1 nozzle under identical lab operating conditions. The laser imaging/video processing system and the P/DPA showed agreement on a calibration tests in monodispersed aerosol sprays of + or - 2.6 micron with a standard deviation of + or - 2.6 micron. Benchmark tests were performed on the NASA MOD-1 nozzle on the centerline and radially at 0.5 inch increments to the outer edge of the spray plume at a distance 2 ft downstream from the exit nozzle. Comparative results at two operation conditions of the nozzle are presented for the two instruments. For the 1st case studied, the deviation in arithmetic mean diameters determined by the two instruments was in a range of 0.1 to 2.8 micron, and the deviation in Sauter mean diameters varied from 0 to 2.2 micron. Severe operating conditions in the 2nd case resulted in the arithmetic mean diameter deviating from 1.4 to 7.1 micron and the deviation in the Sauter mean diameters ranging from 0.4 to 6.7 micron.

Alexander, Dennis R.↗

RISC Processors and High Performance Computing

This tutorial will discuss the top five RISC microprocessors and the parallel systems in which they are used. It will provide a unique cross-machine comparison not available elsewhere. The effective performance of these processors will be compared by citing standard benchmarks in the context of real applications. The latest NAS Parallel Benchmarks, both absolute performance and performance per dollar, will be listed. The next generation of the NPB will be described. The tutorial will conclude with a discussion of future directions in the field. Technology Transfer Considerations: All of these computer systems are commercially available internationally. Information about these processors is available in the public domain, mostly from the vendors themselves. The NAS Parallel Benchmarks and their results have been previously approved numerous times for public release, beginning back in 1991.

Bailey, David H.↗

SEED Platform for Building Performance Standards Implementation Guide (Spanish Translation)

This guide provides an overview of the Standard Energy Efficiency Data (SEED) Platform. The SEED Platform developed by the U.S. Department of Energy (DOE) to provide a low-cost, user-friendly tool for jurisdictions to launch and manage energy benchmarking and Building Performance Standard (BPS) programs. It has been translated into Spanish. This is the Spanish translation of NREL/FS-5500-90691.

benchmarking↗

SEED Platform for Building Performance Standards Implementation Guide (French Translation)

This guide provides an overview of the Standard Energy Efficiency Data (SEED) Platform. The SEED Platform developed by the U.S. Department of Energy (DOE) to provide a low-cost, user-friendly tool for jurisdictions to launch and manage energy benchmarking and Building Performance Standard (BPS) programs. It has been translated into Spanish. This is the French translation of NREL/FS-5500-90691.

benchmarking↗

SEED Platform for Building Performance Standards Implementation Guide (Arabic Translation)

This guide provides an overview of the Standard Energy Efficiency Data (SEED) Platform. The SEED Platform developed by the U.S. Department of Energy (DOE) to provide a low-cost, user-friendly tool for jurisdictions to launch and manage energy benchmarking and Building Performance Standard (BPS) programs. It has been translated into Spanish. This is the Arabic translation of NREL/FS-5500-90691.

benchmarking↗

SEED Platform for Building Performance Standards Implementation Guide (Mandarin Translation)

This guide provides an overview of the Standard Energy Efficiency Data (SEED) Platform. The SEED Platform developed by the U.S. Department of Energy (DOE) to provide a low-cost, user-friendly tool for jurisdictions to launch and manage energy benchmarking and Building Performance Standard (BPS) programs. It has been translated into Spanish. This is the Mandarin translation of NREL/FS-5500-90691.

benchmarking↗

Criticality Accident Alarm System Shielding Benchmark: Integral Experiment Request 498, Critical Engineering Decision 2 Report

A workable design to perform a CAAS benchmark experiment is detailed herein. Key dimensions, materials, source intensity levels, and detectors are listed in this report. Sensitivity to 21 perturbations was determined to be acceptable. The next step will be for NCSP management to determine whether procurement should occur and if the experiment should proceed. The perturbation study suggests that the room return shield cavity radius and runout should be maintained to within a millimeter, the room return shield should be positioned carefully (perhaps with a laser range finder), and that the detectors should be mounted in a lightweight fixture such as aluminum, so their positioning is assured. Using a 3D scanner or photogrammetry to record part shapes may be beneficial. Further work is also needed to verify source reproducibility.

36 MATERIALS SCIENCE↗

wa-hls4ml and lui-gnn: A benchmark and GNN-based surrogate model for hls4ml resource and latency estimation

As machine learning (ML) increasingly serves as a tool for addressing real-time challenges in scientific applications, the development of advanced tooling has significantly reduced the time required to iterate on various designs. These advancements have solved major obstacles, but also exposed new challenges. For example, processes that were not previously considered bottlenecks, such as model synthesis, are now becoming limiting factors in the rapid iteration of designs. To reduce these emerging constraints, multiple efforts are being launched toward designing an ML-based surrogate model that estimates resource usage of synthesized accelerator architectures. This model would reduce the design iteration time, especially when designing within a set of given hardware constraints. This approach shows considerable potential, but as it stands, the effort is early and would benefit from coordination and standardization to assist future work as it emerges. We introduce wa-hls4ml, a benchmark for ML accelerator resource and latency estimation, and its corresponding initial dataset of more than 100,000 fully connected neural networks, all synthesized using hls4ml and targeting Xilinx FPGAs. In addition to the resource utilization and latency data provided, the dataset includes generated artifacts and log files for many of the synthesized neural networks, in order to support future research in ML-based code generation. The benchmark evaluates the performance of resource and latency predictors against several common ML model architectures, primarily originating from scientific domains, as exemplar models, as well as the average performance across a subset of the dataset. We measure the performance of a given predictor model through multiple metrics, including $R^2$ score and SMAPE on regression tasks, as well as inference time to further characterize the estimator under test. Additionally, we introduce the latency/utilization inference graph neural network (lui-gnn), a surrogate model that uses a graph neural network to represent input architectures in the form of a directed graph. This graph representation allows for a diverse set of model architectures to all be effectively handled by a surrogate model. We present the architecture and performance of the model, as evaluated by the new proposed benchmark, including SMAPE, $R^2$ score, and inference times, and find that lui-gnn generally predicts latency and utilization for the 75\% quantile within several percent of the synthesized resources on the synthetic test dataset, indicating that this approach of estimating resource and latency via a surrogate models has promise and warrants further research.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

HipBone: A performance-portable graphics processing unit-accelerated C++ version of the NekBone benchmark

We present hipBone, an open-source performance-portable proxy application for the Nek5000 (and NekRS) computational fluid dynamics applications. HipBone is a fully GPU-accelerated C++ implementation of the original NekBone CPU proxy application with several novel algorithmic and implementation improvements which optimize its performance on modern fine-grain parallel GPU accelerators. Our optimizations include a conversion to store the degrees of freedom of the problem in assembled form in order to reduce the amount of data moved during the main iteration and a portable implementation of the main Poisson operator kernel. We demonstrate near-roofline performance of the operator kernel on three different modern GPU accelerators from two different vendors. We present a novel algorithm for splitting the application of the Poisson operator on GPUs which aggressively hides MPI communication required for both halo exchange and assembly. Our implementation of nearest-neighbor MPI communication then leverages several different routing algorithms and GPU-Direct RDMA capabilities, when available, which improves scalability of the benchmark. We demonstrate the performance of hipBone on three different clusters housed at Oak Ridge National Laboratory, namely, the Summit supercomputer and the Frontier early-access clusters, Spock and Crusher. Our tests demonstrate both portability across different clusters and very good scaling efficiency, especially on large problems.

Computer Science↗

Random circuit block-encoded matrix and a proposal of quantum LINPACK benchmark

The LINPACK benchmark reports the performance of a computer for solving a system of linear equations with dense random matrices. Although this task was not designed with a real application directly in mind, the LINPACK benchmark has been used to define the list of TOP500 supercomputers since the debut of the list in 1993. We propose that a similar benchmark, called the quantum LINPACK benchmark, could be used to measure the whole machine performance of quantum computers. The success of the quantum LINPACK benchmark should be viewed as the minimal requirement for a quantum computer to perform a useful task of solving linear algebra problems, such as linear systems of equations. We propose an input model called the Random Circuit Block-Encoded Matrix (RACBEM), which is a proper generalization of a dense random matrix in the quantum setting. The RACBEM model is efficient to be implemented on a quantum computer and can be designed to optimally adapt to any given quantum architecture, with relying on a black-box quantum compiler. Besides solving linear systems, the RACBEM model can be used to perform a variety of linear algebra tasks relevant to many physical applications, such as computing spectral measures, time series generated by a Hamiltonian simulation, and thermal averages of the energy. We implement these linear algebra operations on IBM Q quantum devices as well as quantum virtual machines, and demonstrate their performance in solving scientific computing problems.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Leveraging observed soil heterotrophic respiration fluxes as a novel constraint on global-scale models

Microbially-explicit models may improve understanding and projections of carbon dynamics in response to future climate change, but their fidelity in simulating global-scale soil heterotrophic respiration (RH), a stringent test for soil biogeochemical models, has never been evaluated. We used statistical global RH products, as well as 7,821 daily site-scale RH measurements, to evaluate the spatio-temporal performance of one first-order decay model (CASA-CNP) and two microbially-explicit biogeochemical models (CORPSE and MIMICS) that were forced by two different climate datasets. CORPSE and MIMICS did not provide any measurable performance improvement; instead, the models were highly sensitive to the meteorological input data used to drive them. Spatial RH variability was generally well simulated except in the northern middle latitudes (~50°N) and arid regions; models captured the seasonal variability of RH well, but showed more divergence in tropic and arctic regions. Our results demonstrate that the next generation of biogeochemical models shows promise, but also needs to be improved for realistic spatio-temporal variability of RH. Finally, we emphasize the importance of net primary production, soil moisture, and soil temperature inputs, and that jointly evaluating soil models for their spatial (global scale) and temporal (site scale) performance provides crucial benchmarks for improving biogeochemical models.

Jian, Jinshi↗

Deep-Learning-Derived Evaluation Metrics Enable Effective Benchmarking of Computational Tools for Phosphopeptide Identification

Tandem mass spectrometry (MS/MS)-based phosphoproteomics is a powerful technology for global phosphorylation analysis. However, applying four computational pipelines to a typical mass spectrometry (MS)-based phosphoproteomic dataset from a human cancer study, we observed a large discrepancy among the reported phosphopeptide identification and phosphosite localization results, underscoring a critical need for benchmarking. While efforts have been made to compare performance of computational pipelines using data from synthetic phosphopeptides, evaluations involving real application data have been largely limited to comparing the numbers of phosphopeptide identifications due to the lack of appropriate evaluation metrics. We investigated three deep learning-derived features as potential evaluation metrics: phosphosite probability, Delta RT and spectral similarity. Predicted phosphosite probability is computed by MusiteDeep, which provides high accuracy as previously reported; Delta RT is defined as the absolute retention time (RT) difference between RTs observed and predicted by AutoRT; and spectral similarity is defined as the Pearson’s correlation coefficient between spectra observed and predicted by pDeep2. Using a synthetic peptide dataset, we found that both Delta RT and spectral similarity provided excellent discrimination between correct and incorrect peptide-spectrum matches (PSMs) both when incorrect PSMs involved wrong peptide sequences and even when incorrect PSMs were caused by only incorrect phosphosite localization. Based on these results, we used all the three deep learning-derived features as evaluation metrics to compare different computational pipelines on diverse set of phosphoproteomic datasets and showed their utility in benchmarking performance of the pipelines. The benchmark metrics demonstrated in this study will enable users to select computational pipelines and parameters for routine analysis of phosphoproteomics data and will offer guidance for developers to improve computational methods.

59 BASIC BIOLOGICAL SCIENCES↗

Expanding Benchmarks for Nuclear Data Validation [Slides]

The presentation discusses benchmarking in comparison to experimental truth, the integral experiments that were performed, and that benchmarks are evaluated experiments. It also discusses validation testing, and that the ultimate goal of the validation end product is to improve evaluated nuclear data for applications It does state that validation highlights errors in the Nuclear Data Pipeline that affect applications, for example: Missing Cd Capture Gammas in ENDF/B-VII.1. It acknowledges that additional types of experiments are needed to test data used in applications, and that validation experiments do not have to be complicated and expensive. The presentation also discusses different types of measurements including activation foil and fission chamber measurements, reactor kinetics measurements, and subcritical measurements. The presentation concludes with an experiment summary and states that all applications need validation.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Benchmarking a Proof-of-Concept Performance Portable SYCL-based Fast Fourier Transformation Library

ABSTRACT In this paper, we present an early version of a SYCL-based FFT library, capable of running on all major vendor hardware, including CPUs and GPUs from AMD, ARM, Intel and NVIDIA. The current limitations of our library is it supports single-dimension FFTs up to 211 in length and base-2 input sequences. Although preliminary, the aim of this work is to seed further developments for a rich set of features for calculating FFTs. The library has the advantage over existing portable FFT libraries in that it is single-source, and there- fore removes the complexities that arise due to abundant use of pre-processor macros and auto-generated kernels to target different architectures. We exercise two SYCL-enabled compilers, Codeplay ComputeCpp and Intel's open-source LLVM project, to evaluate performance portability of our SYCL-based FFT on various hetero- geneous architectures.We provide studies comparing our portable library with highly optimized vendor-specific FFT libraries, and discuss potential sources hindering performance.

97 MATHEMATICS AND COMPUTING↗

LLM Benchmarking with LLaMA2: Evaluating Code Development Performance Across Multiple Programming Languages

The rapid evolution of large language models (LLMs) has opened new possibilities for automating various tasks in software development. This paper evaluates the capabilities of the LLaMA 2-70B model in automating these tasks for scientific applications written in commonly used programming languages. Using representative test problems, we assess the model's capacity to generate code, documentation, and unit tests, as well as its ability to translate existing code between commonly used programming languages. Our comprehensive analysis evaluates the compilation, runtime behavior, and correctness of the generated and translated code. Additionally, we assess the quality of automatically generated code, documentation, and unit tests. Here, our results indicate that while LLaMA 2-70B frequently generates syntactically correct and functional code for simpler numerical tasks, it encounters substantial difficulties with more complex, parallelized, or distributed computations, requiring considerable manual corrections. We identify key limitations and suggest areas for future improvements to better leverage AI-driven automation in scientific computing workflows.

97 MATHEMATICS AND COMPUTING↗

RISC Processors and High Performance Computing

In this tutorial, we will discuss top five current RISC microprocessors: The IBM Power2, which is used in the IBM RS6000/590 workstation and in the IBM SP2 parallel supercomputer, the DEC Alpha, which is in the DEC Alpha workstation and in the Cray T3D; the MIPS R8000, which is used in the SGI Power Challenge; the HP PA-RISC 7100, which is used in the HP 700 series workstations and in the Convex Exemplar; and the Cray proprietary processor, which is used in the new Cray J916. The architecture of these microprocessors will first be presented. The effective performance of these processors will then be compared, both by citing standard benchmarks and also in the context of implementing a real applications. In the process, different programming models such as data parallel (CM Fortran and HPF) and message passing (PVM and MPI) will be introduced and compared. The latest NAS Parallel Benchmark (NPB) absolute performance and performance per dollar figures will be presented. The next generation of the NP13 will also be described. The tutorial will conclude with a discussion of general trends in the field of high performance computing, including likely future developments in hardware and software technology, and the relative roles of vector supercomputers tightly coupled parallel computers, and clusters of workstations. This tutorial will provide a unique cross-machine comparison not available elsewhere.

Saini, Subhash↗