Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Performance benchmark”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32

NAS parallel benchmark results

The NAS (Numerical Aerodynamic Simulation) parallel benchmarks have been developed at NASA Ames Research Center to study the performance of parallel supercomputers. The eight benchmark problems are specified in a 'pencil and paper' fashion. The performance results of various systems using the NAS parallel benchmarks are presented. These results represent the best results that have been reported to the authors for the specific systems listed. They represent implementation efforts performed by personnel in both the NAS Applied Research Branch of NASA Ames Research Center and in other organizations.

Bailey, D. H.↗

Benchmarking for AI for Science

AI has been instrumental for recent developments in a number of domains of the sciences. With several hundred machine learning (ML) algorithms and models, and numerous AI-specific hardware platforms, a common quest for all scientists working on AI for Science is around the selection of machine learning algorithm(s) to solve their domain-specific scientific problems. A number of different initiatives around AI Benchmarking have been set up and have been useful in understanding the benefits of different ML algorithms for different tasks.However, with the majority of these AI Benchmarking initiatives focusing on the conventional notions of benchmarking, where the focus is purely runtime performance (such as training time or inference time), their suitability for benchmarking different ML algorithms for solving scientific problems has been viewed as a performance problem even though both are hardly the same. To make reasonable, explainable, and justifiable advancements in science using AI, it is critical to focus on the merits of these algorithms in handling different domain science problems. In other words, more emphasis must be given on Benchmarking for AI for Science than AI Benchmarking. The vision of the former is not only to assess the performance of ML algorithms, but also to assess, and understand the benefits and merits of different ML algorithms in handling scientific problems. Benchmarking for AI for Science, instead of pure performance focused AI Benchmarking, has several benefits: (i) it has the potential to offer advances in the sciences, much more rapidly than through pure performance-based AI methods, (ii) it will encourage the community to focus on developing better domain-specific AI techniques, particularly given the provision for being able to benchmark different techniques, and (iii) it will encourage hardware manufacturers to focus on developing science-specific hardware subsystems.

Thiyagalingam, Jeyan↗

The QICK (Quantum Instrumentation Control Kit): Readout and control for qubits and detectors

We introduce a Xilinx RF System-on-Chip (RFSoC)-based qubit controller (called the Quantum Instrumentation Control Kit, or QICK for short), which supports the direct synthesis of control pulses with carrier frequencies of up to 6 GHz. The QICK can control multiple qubits or other quantum devices. The QICK consists of a digital board hosting an RFSoC field-programmable gate array, custom firmware, and software and an optional companion custom-designed analog front-end board. We characterize the analog performance of the system as well as its digital latency, important for quantum error correction and feedback protocols. We benchmark the controller by performing standard characterizations of a transmon qubit. We achieve an average gate fidelity of ℱ avg =99.93%. All of the schematics, firmware, and software are open-source.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

An Open-Source Coupling for Depletion During Fuel Cycle Modeling

Fuel depletion is an important aspect of fuel cycle modeling, allowing a user to account for how loaded fuel compositions affect in-core and spent fuel compositions and their related fuel cycle metrics. Therefore, multiple methods have been developed to account for depletion within fuel cycle simulations. This work adds to that list of methods by introducing an open-source coupling between C$\scriptsize{YCLUS}$ and OpenMC to perform fuel depletion during a fuel cycle simulation, called OpenMCyclus. This work explains the methodology of OpenMCyclus and presents a benchmark comparison between the performance of OpenMCyclus and another C$\scriptsize{YCLUS}$ archetype that uses recipes to define spent fuel compositions. In conclusion, the development of this coupling expands the functionalities possible through C$\scriptsize{YCLUS}$ by providing real-time fuel depletion that is reactor agnostic and open source.

C$\scriptsize{YCLUS}$↗

Discriminating Quantum States with Quantum Machine Learning

Quantum machine learning (QML) algorithms have obtained great relevance in the machine learning (ML) field due to the promise of quantum speedups when performing basic linear algebra subroutines (BLAS), a fundamental element in most ML algorithms. By making use of BLAS operations, we propose, implement and analyze a quantum k-means (qk-means) algorithm with a low time complexity of O(NKlog(D)I/C) to apply it to the fundamental problem of discriminating quantum states at readout. Discriminating quantum states allows the identification of quantum states |0⟩ and |1⟩ from low-level in-phase and quadrature signal (IQ) data, and can be done using custom ML models. In order to reduce dependency on a classical computer, we use the qk-means to perform state discrimination on the IBMQ Bogota device and managed to find assignment fidelities of up to 98.7% that were only marginally lower than that of the k-means algorithm. We also performed a cross-talk benchmark on the quantum device by applying both algorithms to perform state discrimination on a combination of quantum states and using Pearson Correlation coefficients and assignment fidelities of discrimination results to conclude on the presence of cross-talk on qubits. Evidence shows cross-talk in the (1, 2) and (2, 3) neighboring qubit couples for the analyzed device.

Quiroga, David↗

MLCommons Science Benchmarks

Benchmarks are a cornerstone of modern machine learning practice, providing standardized eval- uations that enable reproducibility, comparison, and scientific progress. Yet, as AI systems particularly deep learning models become increasingly dynamic, traditional static benchmarking approaches are losing their relevance. Models rapidly evolve in architecture, scale, and capability; datasets shift; and deployment contexts continuously change, creating a moving target for evaluation. Without adaptive benchmarking frame- works, both scientific assessment and real-world de- ployment risk becoming misaligned with actual system behavior. Drawing on our experience from MLCommons, educa- tional initiatives, and government programs such as the DOE s Million Parameter Consortium, we identify key barriers that hinder the broader adoption and utility of benchmarking in AI. These include substantial resource demands, limited access to specialized hardware, lack of expertise in benchmark design, and uncertainty among practitioners about how to relate benchmark results to their own application domains. Moreover, current benchmarks often emphasize peak performance on leadership-class hardware, offering limited guidance for more diverse, real-world deployment scenarios. We argue that benchmarking itself must become dy- namic in order to incorporate evolving models, updated data, and heterogeneous computational platforms while maintaining transparency, reproducibility, and inter- pretability. Democratizing this process requires not only technical innovation, but also systematic educational efforts spanning undergraduate to professional levels to develop sustained expertise in benchmark design and use. Finally, benchmarks should be framed and com- municated to support application-relevant comparisons, enabling both developers and users to make informed, context-sensitive decisions. Advancing dynamic and inclusive benchmarking practices will be essential to ensure that evaluation keeps pace with the evolving AI landscape and supports responsible, reproducible, and accessible AI deployment.

Hawks, Benjamin G. [Fermilab]↗

Design of the EO-1 Pulsed Plasma Thruster Attitude Control Experiment

The Pulsed Plasma Thruster (PPT) Experiment on the Earth Observing 1 (EO-1) spacecraft has been designed to demonstrate the capability of a new generation PPT to perform spacecraft attitude control. The PPT is a small, self-contained pulsed electromagnetic Propulsion system capable of delivering high specific impulse (900-1200 s), very small impulse bits (10-1000 micro N-s) at low average power (less than 1 to 100 W). EO-1 has a single PPT that can produce torque in either the positive or negative pitch direction. For the PPT in-flight experiment, the pitch reaction wheel will be replaced by the PPT during nominal EO-1 nadir pointing. A PPT specific proportional-integral-derivative (PID) control algorithm was developed for the experiment. High fidelity simulations of the spacecraft attitude control capability using the PPT were conducted. The simulations, which showed PPT control performance within acceptable mission limits, will be used as the benchmark for on-orbit performance. The flight validation will demonstrate the ability of the PPT to provide precision pointing resolution. response and stability as an attitude control actuator.

Zakrzwski, Charles↗

Unstructured Adaptive (UA) NAS Parallel Benchmark

We present a complete specification of a new benchmark for measuring the performance of modern computer systems when solving scientific problems featuring irregular, dynamic memory accesses. It complements the existing NAS Parallel Benchmark suite. The benchmark involves the solution of a stylized heat transfer problem in a cubic domain, discretized on an adaptively refined, unstructured mesh.

Feng, Huiyu↗

Resolved resonance region evaluations of n+ 206,207,208 Pb for fast spectrum applications

Resolved resonance region evaluations of the major isotopes of natural lead, 206 Pb, 207 Pb, 208 Pb, have been per formed to support the development of Generation IV reactors. In this study, validation of nuclear data for lead fast reactors was performed with simulations of shielding benchmarks, integral critical benchmarks, and quasi-differential scattering measurements. Sensitivity analyses of these systems showed that elastic scattering reactions above 100 keV were the dominant reactions driving system performance. The resolved resonance regions (RRRs) of the lead isotopes extend past 100 keV, making the RRR an ideal starting point to evaluate lead cross sections. Since the R-matrix requires knowledge of bound, distant, and observed resonances, it was necessary to evaluate from 10 -5 eV up to the respective limit of the RRR. The 208 Pb RRR evaluation was extended to 1.5 MeV in order to obtain resonance parameters used to calculate new elastic scattering angular distributions up to 1.5 MeV. Resonance parameter uncertainties and covariance were generated using the R-matrix code SAMMY. The new RRR parameters show a direct improvement to the scattering kernel below 1.5 MeV which in turn greatly improves fast critical experiments over ENDF/B-VIII.0.

206Pb↗

Implementation of BT, SP, LU, and FT of NAS Parallel Benchmarks in Java

A number of Java features make it an attractive but a debatable choice for High Performance Computing. We have implemented benchmarks working on single structured grid BT,SP,LU and FT in Java. The performance and scalability of the Java code shows that a significant improvement in Java compiler technology and in Java thread implementation are necessary for Java to compete with Fortran in HPC applications.

Schultz, Matthew↗

Comparative evaluation of deep learning workloads for leadership-class systems

Deep learning (DL) workloads and their performance at scale are becoming important factors to consider as we design, develop and deploy next-generation high-performance computing systems. Since DL applications rely heavily on DL frameworks and underlying compute (CPU/GPU) stacks, it is essential to gain a holistic understanding from compute kernels, models, and frameworks of popular DL stacks, and to assess their impact on science-driven, mission-critical applications. At Oak Ridge Leadership Computing Facility (OLCF), we employ a set of micro and macro DL benchmarks established through the Collaboration of Oak Ridge, Argonne, and Livermore (CORAL) to evaluate the AI readiness of our next-generation supercomputers. In this paper, we present our early observations and performance benchmark comparisons between the Nvidia V100 based Summit system with its CUDA stack and an AMD MI100 based testbed system with its ROCm stack. We take a layered perspective on DL benchmarking and point to opportunities for future optimizations in the technologies that we consider.

Yin, Junqi↗

MPI, HPF or OpenMP: A Study with the NAS Benchmarks

Porting applications to new high performance parallel and distributed platforms is a challenging task. Writing parallel code by hand is time consuming and costly, but this task can be simplified by high level languages and would even better be automated by parallelizing tools and compilers. The definition of HPF (High Performance Fortran, based on data parallel model) and OpenMP (based on shared memory parallel model) standards has offered great opportunity in this respect. Both provide simple and clear interfaces to language like FORTRAN and simplify many tedious tasks encountered in writing message passing programs. In our study, we implemented the parallel versions of the NAS Benchmarks with HPF and OpenMP directives. Comparison of their performance with the MPI implementation and pros and cons of different approaches will be discussed along with experience of using computer-aided tools to help parallelize these benchmarks. Based on the study, potentials of applying some of the techniques to realistic aerospace applications will be presented.

Jin, H.↗

MPI, HPF or OpenMP: A Study with the NAS Benchmarks

Porting applications to new high performance parallel and distributed platforms is a challenging task. Writing parallel code by hand is time consuming and costly, but the task can be simplified by high level languages and would even better be automated by parallelizing tools and compilers. The definition of HPF (High Performance Fortran, based on data parallel model) and OpenMP (based on shared memory parallel model) standards has offered great opportunity in this respect. Both provide simple and clear interfaces to language like FORTRAN and simplify many tedious tasks encountered in writing message passing programs. In our study we implemented the parallel versions of the NAS Benchmarks with HPF and OpenMP directives. Comparison of their performance with the MPI implementation and pros and cons of different approaches will be discussed along with experience of using computer-aided tools to help parallelize these benchmarks. Based on the study,potentials of applying some of the techniques to realistic aerospace applications will be presented

Jin, Hao-Qiang↗

Benchmark Testing on the IBM-Q Network

The goal of this project is to evaluate a proposed set of candidate benchmarks being developed by the Standards and Performance Metrics Technical Advisory Group (TAG) of the Quantum Economic Development Consortium (QED-C), of which LANL is a member. The QED-C Standards and Performance Metrics TAG has developed implementations of several benchmark codes with the hope and expectation that these codes will be helpful to QED-C members interested in investigating various quantum computer platforms. The benchmark set will be run on LANL’s access to the IBM-Q system with the intent to evaluate the set for scalability, correct execution, and coverage of the application space. The goal is for the QED-C Standards and Performance Metrics TAC to be able to produce a coherent, consistent, scalable set of benchmarks that will run on multiple quantum computing platforms and will be available to members of QED-C, including LANL.

97 MATHEMATICS AND COMPUTING↗

Turbine Electrified Energy Management with Model Predictive Control

Affordability, sustainability, and efficiency are primary motivators driving the future of NASA aeronautics research. These factors are realized, in part, through the development and implementation of new technologies and strategies enabling efficient, affordable, and safe hybrid-electric aircraft. Research supporting electrified aircraft propulsion control systems exemplifies such new methodologies, offering varied opportunities to integrate electric machines with gas-based turbine engines. For hybrid-electric propulsion systems, current conceptual architectures seek to introduce energy storage and exploit electrical power system components to assist gas-based system components. Capitalizing on the electric machines in hybridized engines, Turbine Electrified Energy Management (TEEM) is a control approach that enhances transient operability to improve overall propulsion and vehicle efficiency by injecting or extracting power from engine shafts. Traditionally implemented with proportional-integral (PI) control, this study expands the application of TEEM by presenting model predictive control (MPC) schemes to execute the TEEM concept. Via cost function design and constraint selection, the transient operability goals for TEEM are considered in the controller designs. The proposed MPCs are simulated on a nonlinear turbofan engine model at two environmental conditions, with comparisons drawn to a baseline PI. Performance is evaluated using compressor maps and two TEEM-specific metrics: transient stack usage and transient excursion integral. Simulation results reveal the developed schemes perform comparably to the benchmark controller and can be implemented in two distinct configurations. Potential modifications for future investigations include cost function measures that optimize energy use, additional performance effectiveness measures, and battery storage capabilities.

Elyse D. Hill↗

Studying CPU and memory utilization of applications on Fujitsu A64FX and Nvidia Grace Superchip

ARM-based manycore CPU architectures are well-positioned to provide the rising memory throughput requirements of modern data intensive scientific applications in High Performance Computing (HPC). The Fujitsu A64FX CPU platform is based on the ARM v8.2A architecture, and is the processor of the flagship Japanese supercomputer - "Fugaku", which was previously ranked as the #1 supercomputer in the world according to the Top500 list. The Nvidia Grace superchip features 144 Neoverse V2 cores based on the ARMv9 architecture with 4x128b SVE2, providing exceptional computational power. The chip supports up to 480GB of memory, making it ideal for AI, machine learning, and scientific computing workloads. In this paper, we conduct a thorough performance exploration of a variety of parallel bandwidth-sensitive benchmarks and applications compiled with the native Fujitsu compiler on a Fugaku A64FX compute node and ARM (LLVM) Compiler on an NVIDIA Grace superchip compute node, engaging all the computational cores per cluster using OpenMP multithreading (assuming the cores can drive the available bandwidth). Our ultimate goals are to study the resource utilization of scientific applications and benchmarks on A64FX and Grace superchip, considering graph application scenarios ( GAP Benchmark suite) and eleven appli- cation proxies from the Rodinia heterogeneous benchmark suite (considering domains such as Data Mining, Bioinformatics, Fluid Dynamics, Pattern Recognition, etc.). Through exhaustive performance monitoring, we quantify the resource utilization of diverse OpenMP-based HPC applications on both the Fujitsu A64FX and the Nvidia Grace Superchip platforms.

benchmarking, Performance Analysis, High performan↗

Sub-microsecond Transformers for Jet Tagging on FPGAs

We present the first sub-microsecond transformer implementation on an FPGA achieving competitive performance for state-of-the-art high-energy physics benchmarks. Transformers have shown exceptional performance on multiple tasks in modern machine learning applications, including jet tagging at the CERN Large Hadron Collider (LHC). However, their computational complexity prohibits use in real-time applications, such as the hardware trigger system of the collider experiments up until now. In this work, we demonstrate the first application of transformers for jet tagging on FPGAs, achieving $\mathcal{O}(100)$ nanosecond latency with superior performance compared to alternative baseline models. We leverage high-granularity quantization and distributed arithmetic optimization to fit the entire transformer model on a single FPGA, achieving the required throughput and latency. Furthermore, we add multi-head attention and linear attention support to hls4ml, making our work accessible to the broader fast machine learning community. This work advances the next-generation trigger systems for the High Luminosity LHC, enabling the use of transformers for real-time applications in high-energy physics and beyond.

Laatu, Lauri [Imperial Coll., London]↗

BACT Simulation User Guide (Version 7.0)

This report documents the structure and operation of a simulation model of the Benchmark Active Control Technology (BACT) Wind-Tunnel Model. The BACT system was designed, built, and tested at NASA Langley Research Center as part of the Benchmark Models Program and was developed to perform wind-tunnel experiments to obtain benchmark quality data to validate computational fluid dynamics and computational aeroelasticity codes, to verify the accuracy of current aeroservoelasticity design and analysis tools, and to provide an active controls testbed for evaluating new and innovative control algorithms for flutter suppression and gust load alleviation. The BACT system has been especially valuable as a control system testbed.

Waszak, Martin R.↗