Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Performance benchmark”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

An Evaluation of Representation Learning Methods in Particle Physics Foundation Models

We present a systematic evaluation of representation learning objectives for particle physics within a unified framework. Our study employs a shared transformer-based particle-cloud encoder with standardized preprocessing, matched sampling, and a consistent evaluation protocol on a jet classification dataset. We compare contrastive (supervised and self-supervised), masked particle modeling, and generative reconstruction objectives under a common training regimen. In addition, we introduce targeted supervised architectural modifications that achieve state-of-the-art performance on benchmark evaluations. This controlled comparison isolates the contributions of the learning objective, highlights their respective strengths and limitations, and provides reproducible baselines. We position this work as a reference point for the future development of foundation models in particle physics, enabling more transparent and robust progress across the community.

Chen, Michael [Caltech]↗

AI Benchmark Democratization and Carpentry

Benchmarks are a cornerstone of modern machine learning, enabling reproducibility, comparison, and scientific progress. However, AI benchmarks are increasingly complex, requiring dynamic, AI-focused workflows. Rapid evolution in model architectures, scale, datasets, and deployment contexts makes evaluation a moving target. Large language models often memorize static benchmarks, causing a gap between benchmark results and real-world performance. Beyond traditional static benchmarks, continuous adaptive benchmarking frameworks are needed to align scientific assessment with deployment risks. This calls for skills and education in AI Benchmark Carpentry. From our experience with MLCommons, educational initiatives, and programs like the DOE's Trillion Parameter Consortium, key barriers include high resource demands, limited access to specialized hardware, lack of benchmark design expertise, and uncertainty in relating results to application domains. Current benchmarks often emphasize peak performance on top-tier hardware, offering limited guidance for diverse, real-world scenarios. Benchmarking must become dynamic, incorporating evolving models, updated data, and heterogeneous platforms while maintaining transparency, reproducibility, and interpretability. Democratization requires both technical innovation and systematic education across levels, building sustained expertise in benchmark design and use. Benchmarks should support application-relevant comparisons, enabling informed, context-sensitive decisions. Dynamic, inclusive benchmarking will ensure evaluation keeps pace with AI evolution and supports responsible, reproducible, and accessible AI deployment. Community efforts can provide a foundation for AI Benchmark Carpentry.

von Laszewski, Gregor [Virginia U.]↗

Advances in building data management for building performance standards using the SEED platform

Reducing energy consumption and greenhouse gas emissions in the built environment is a critical step in achieving emission goals to mitigate climate change impacts. Local, federal, and international jurisdictions are deploying several methods to reduce energy and emissions such as voluntary and mandatory benchmarking and building performance standards, requiring building owners to reach energy and emission targets. Jurisdictions leveraging benchmarking and building performance standards require knowledge of the buildings covered; which is a large task due to staffing constraints, limited information on building characteristics and tax parcel data, and the need for advanced data management techniques to align datasets. This paper describes an open-source platform's recent advances to create consistent taxonomies, identify erroneous data, enable auditability, and track building performance. The paper concludes with two use cases on how the platform has been used by jurisdictions.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

FY24 Progress Report on Viscosity and Thermal Conductivity Measurements of Nuclear Industry Relevant Chloride Salts: An Experimental and Computational Study

As presented in this report, experimental and computational techniques were performed to assess the viscosity and thermal conductivity of key alkali and actinide chloride mixtures for molten salt reactor developers. These mixtures were pure LiCl, NaCl-KCl, LiCl-NaCl, LiCl-KCl, LiCl-NaCl-KCl, and NaCl-UCl 3 . Experimental measurements of viscosity were performed with a rolling ball viscometer, whereas experimental measurements of thermal conductivity were performed with a variable gap apparatus. Additional benchmarking work was performed using both property measurement systems to prepare for x-ray radiography in stainless-steel crucibles for viscosity and to ensure that calibration methods were accurate for thermal conductivity before assessing the NaCl-UCl 3 system. Validation data for the NaCl-UCl 3 in literature are minimal. Details on the calibration methods, salt measurement processes, and sources of error and uncertainty are discussed in detail for both property measurements. The computational methods described herein involved ab-initio molecular dynamics (AIMD) calculations using CP2K. The calculations were performed for the LiCl-KCl-NaCl and NaCl-UCl 3 systems. These calculations not only provided thermophysical property estimations for comparison to experimental data, but they also allowed for the determination of diffusion coefficients, coordination numbers, and radial distribution functions to provide insight into ion mobility and local coordination environments, which is linked to macroscopic property trends.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Benchmark of numerical modeling approaches on the systematic performance evaluation of wave energy converters

Different numerical modeling methods have been developed and applied to evaluate a variety of performance indicators of wave energy converters (WECs), including the power performance, structural loads, levelized cost of energy, etc. Based on the modeling fidelity, the commonly used numerical modeling approaches can be classified as linear modeling, weakly nonlinear modeling and fully nonlinear modeling approaches. Each method differs in accuracy and computational efficiency, making them suitable for different stages of WEC design. However, the selection of modeling approach could significantly impact evaluation outcomes. For instance, simplified linear models may underestimate structural loads or overestimate energy production in some operational conditions, potentially leading to less cost-effective designs. Given the widespread utilization of these models, it is essential to understand the uncertainties brought by them in performance evaluations. This work is dedicated to benchmarking different linear-potential-flow-based numerical models for evaluating the systematic performance of WECs. Three representative numerical modeling approaches are considered in this work, including linear frequency-domain modeling, statistically linearized spectral-domain modeling and Cummins equation-based nonlinear time-domain modeling. A generic point absorber WEC is considered as the research reference in this work, and different sea sites are taken into account. The numerical models are utilized to predict critical performance indicators, including power performance, the annual energy production, the capacity factor, the levelized cost of energy and the PTO fatigue loads. By comparing the results, this work identifies the uncertainties associated with different modeling approaches in evaluating WEC performance.

Fatigue↗

A robust synthetic data generation framework for machine learning in high-resolution transmission electron microscopy (HRTEM)

Machine learning techniques are attractive options for developing highly-accurate analysis tools for nanomaterials characterization, including high-resolution transmission electron microscopy (HRTEM). However, successfully implementing such machine learning tools can be difficult due to the challenges in procuring sufficiently large, high-quality training datasets from experiments. In this work, we introduce Construction Zone, a Python package for rapid generation of complex nanoscale atomic structures which enables fast, systematic sampling of realistic nanomaterial structures and can be used as a random structure generator for large, diverse synthetic datasets. Using Construction Zone, we develop an end-to-end machine learning workflow for training neural network models to analyze experimental atomic resolution HRTEM images on the task of nanoparticle image segmentation purely with simulated databases. Further, we study the data curation process to understand how various aspects of the curated simulated data—including simulation fidelity, the distribution of atomic structures, and the distribution of imaging conditions—affect model performance across three benchmark experimental HRTEM image datasets. Using our workflow, we are able to achieve state-of-the-art segmentation performance on these experimental benchmarks and, further, we discuss robust strategies for consistently achieving high performance with machine learning in experimental settings using purely synthetic data. Construction Zone and its documentation are available at https://github.com/lerandc/construction_zone.

36 MATERIALS SCIENCE↗

Benchmarking quantum computers

The rapid pace of development in quantum computing technology has sparked a proliferation of benchmarks to assess the performance of quantum computing hardware and software. However, not all benchmarks are of equal merit. Good ones empower scientists, engineers, programmers and users to understand the power of a computing system, whereas bad ones can misdirect research and inhibit progress. In this Perspective, we survey the science of quantum computer benchmarking. Here, we discuss the role of benchmarks and benchmarking and how good benchmarks can drive and measure progress towards the long-term goal of useful quantum computations, known as quantum utility. We explain how different kinds of benchmark quantify the performance of different parts of a quantum computer, discuss existing benchmarks, examine recent trends in benchmarking, and highlight important open research questions in this field.

Proctor, Timothy James [Sandia National Laboratori↗

Updates to the n+ 63,65 Cu Angular Distributions [Slides]

The performance of the benchmark suite is very sensitive to changes in the angular distribution data. The quasi-differential measurements performed by Blain et al at RPI provide a valuable constraint. Furthermore, ENDF/B-VIII.0 disagrees with their measurements consistently at 300 keV (the transition between RRR and high energy). The next step is to conservatively adjust the Legendre coefficients near 300 keV, validating performance against both critical benchmarks and the RPI quasi-differential measurements.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Results of a Geant4 benchmarking study for bio‐medical applications, performed with the G4‐Med system

Geant4, a Monte Carlo Simulation Toolkit extensively used in bio-medical physics, is in continuous evolution to include newest research findings to improve its accuracy and to respond to the evolving needs of a very diverse user community. In 2014, the G4-Med benchmarking system was born from the effort of the Geant4 Medical Simulation Benchmarking Group, to benchmark and monitor the evolution of Geant4 for medical physics applications. The G4-Med system was first described in our Medical Physics Special Report published in 2021. Results of the tests were reported for Geant4 10.5. Purpose In this work, we describe the evolution of the G4-Med benchmarking system. Methods The G4-Med benchmarking suite currently includes 23 tests, which benchmark Geant4 from the calculation of basic physical quantities to the simulation of more clinically relevant set-ups. New tests concern the benchmarking of Geant4-DNA physics and chemistry components for regression testing purposes, dosimetry for brachytherapy with a 125 I source, dosimetry for external x-ray and electron FLASH radiotherapy, experimental microdosimetry for proton therapy, and in vivo PET for carbon and oxygen beams. Regression testing has been performed between Geant4 10.5 and 11.1. Finally, a simple Geant4 simulation has been developed and used to compare Geant4 EM physics constructors and physics lists in terms of execution times. Results In summary, our EM tests show that the parameters of the multiple scattering in the Geant4 EM constructor G4EmStandardPhysics_option3 in Geant4 11.1, while improving the modeling of the electron backscattering in high atomic number targets, are not adequate for dosimetry for clinical x-ray and electron beams. Therefore, these parameters have been reverted back to those of Geant4 10.5 in Geant4 11.2.1. The x-ray radiotherapy test shows significant differences in the modeling of the bremsstrahlung process, especially between G4EmPenelopePhysics and the other constructors under study (G4EmLivermorePhysics, G4EmStandardPhysics_option3, and G4EmStandardPhysics_option4). These differences will be studied in an in-depth investigation within our Group. Improvement in Geant4 11.1 has been observed for the modeling of the proton and carbon ion Bragg peak with energies of clinical interest, thanks to the adoption of ICRU90 to calculate the low energy proton stopping powers in water and of the Linhard–Sorensen ion model, available in Geant4 since version 11.0. Nuclear fragmentation tests of interest for carbon ion therapy show differences between Geant4 10.5 and 11.1 in terms of fragment yields. In particular, a higher production of boron fragments is observed with Geant4 11.1, leading to a better agreement with reference data for this fragment. Conclusions Based on the overall results of our tests, we recommend to use G4EmStandardPhysics_option4 as EM constructor and QGSP_BIC_HP with G4EmStandardPhysics_option4, for hadrontherapy applications. The Geant4-DNA physics lists report differences in modeling electron interactions in water, however, the tests have a pure regression testing purpose so no recommendation can be formulated.

62 RADIOLOGY AND NUCLEAR MEDICINE↗

Implementing and Improving CBMZ-MAM3 Chemistry and Aerosol Modules in the Regional Climate Model WRF-CAM5: An Evaluation over the Western US and Eastern North Pacific

The representation of aerosols in climate-chemistry models is important for air quality and climate change research, but it can require significant computational resources. The objective of this study was to improve the representation of aerosols in climate–chemistry models, specifically in the carbon bond mechanism, version Z (CBMZ), and modal aerosol modules with three lognormal modes (MAM3) in the WRF-CAM5 model. The study aimed to enhance the model’s chemistry capabilities by incorporating biomass burning emissions, establishing a conversion mechanism between volatile organic compounds (VOCs) and secondary organic carbons (SOCs), and evaluating its performance against observational benchmarks. The results of the study demonstrated the effectiveness of the enhanced chemistry capabilities in the WRF-CAM5 model. Six simulations were conducted over the western U.S. and northeastern Pacific region, comparing the model’s performance with observational benchmarks such as reanalysis, ground-based, and satellite data. The findings revealed a significant reduction in root-mean-square errors (RMSE) for surface concentrations of black carbon (BC) and organic carbon (OC). Specifically, the model exhibited a 31% reduction in RMSE for BC concentrations and a 58% reduction in RMSE for OC concentrations. These outcomes underscored the importance of accurate aerosol representation in climate-chemistry models and emphasized the potential for improving simulation accuracy and reducing errors through the incorporation of enhanced chemistry modules in such models.

54 ENVIRONMENTAL SCIENCES↗

Reactor physics benchmark experiments at the JSI TRIGA MARK II reactor - Current status and future outlook

Full text of publication follows. With the development of new high-fidelity computational methods, improvement of nuclear data, and multiphysics modelling, there is an increased need for benchmark experiments to experimentally validate the models, methods and input data. Many of the nuclear facilities designed to perform reactor physics benchmark experiments have been shut down. Therefore, research reactors offer a great opportunity for benchmark experiments, if they are well designed and performed with great care and accuracy. In this presentation we provide an overview of the past and ongoing activities related to benchmark experiments at the Jozef Stefan Institute TRIGA Mark II research reactor. The following experiments have been performed: criticality with fresh fuel, {sup 197}Au(n,γ) and {sup 27}Al(n,α) reaction rates in irradiation channels, absolute and relative {sup 197}Au(n,γ), {sup 235}U(n,f) and {sup 238}U(n,f) reaction rates in the core, burnup, kinetic parameters, control rod worth, isothermal reactivity coefficient, self-shielding, slow and fast (pulse) transients, nuclear heating, delayed and prompt gamma ray production, temperature profiles for multi-physics. Since the existing fleet of research reactors is ageing very rapidly and new experiments are needed, new research reactors should be designed and built to meet the needs of future advanced reactors, education and training, and other technologies in the coming years. We will review planned activities at the JSI TRIGA reactors and plans for the new research reactor in Slovenia. (authors)

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

SEED Platform for Building Performance Standards Implementation Guide (Spanish Translation)

This guide provides an overview of the Standard Energy Efficiency Data (SEED) Platform. The SEED Platform developed by the U.S. Department of Energy (DOE) to provide a low-cost, user-friendly tool for jurisdictions to launch and manage energy benchmarking and Building Performance Standard (BPS) programs. It has been translated into Spanish. This is the Spanish translation of NREL/FS-5500-90691.

benchmarking↗

SEED Platform for Building Performance Standards Implementation Guide (French Translation)

This guide provides an overview of the Standard Energy Efficiency Data (SEED) Platform. The SEED Platform developed by the U.S. Department of Energy (DOE) to provide a low-cost, user-friendly tool for jurisdictions to launch and manage energy benchmarking and Building Performance Standard (BPS) programs. It has been translated into Spanish. This is the French translation of NREL/FS-5500-90691.

benchmarking↗

SEED Platform for Building Performance Standards Implementation Guide (Arabic Translation)

This guide provides an overview of the Standard Energy Efficiency Data (SEED) Platform. The SEED Platform developed by the U.S. Department of Energy (DOE) to provide a low-cost, user-friendly tool for jurisdictions to launch and manage energy benchmarking and Building Performance Standard (BPS) programs. It has been translated into Spanish. This is the Arabic translation of NREL/FS-5500-90691.

benchmarking↗

SEED Platform for Building Performance Standards Implementation Guide (Mandarin Translation)

This guide provides an overview of the Standard Energy Efficiency Data (SEED) Platform. The SEED Platform developed by the U.S. Department of Energy (DOE) to provide a low-cost, user-friendly tool for jurisdictions to launch and manage energy benchmarking and Building Performance Standard (BPS) programs. It has been translated into Spanish. This is the Mandarin translation of NREL/FS-5500-90691.

benchmarking↗

Criticality Accident Alarm System Shielding Benchmark: Integral Experiment Request 498, Critical Engineering Decision 2 Report

A workable design to perform a CAAS benchmark experiment is detailed herein. Key dimensions, materials, source intensity levels, and detectors are listed in this report. Sensitivity to 21 perturbations was determined to be acceptable. The next step will be for NCSP management to determine whether procurement should occur and if the experiment should proceed. The perturbation study suggests that the room return shield cavity radius and runout should be maintained to within a millimeter, the room return shield should be positioned carefully (perhaps with a laser range finder), and that the detectors should be mounted in a lightweight fixture such as aluminum, so their positioning is assured. Using a 3D scanner or photogrammetry to record part shapes may be beneficial. Further work is also needed to verify source reproducibility.

36 MATERIALS SCIENCE↗

wa-hls4ml and lui-gnn: A benchmark and GNN-based surrogate model for hls4ml resource and latency estimation

As machine learning (ML) increasingly serves as a tool for addressing real-time challenges in scientific applications, the development of advanced tooling has significantly reduced the time required to iterate on various designs. These advancements have solved major obstacles, but also exposed new challenges. For example, processes that were not previously considered bottlenecks, such as model synthesis, are now becoming limiting factors in the rapid iteration of designs. To reduce these emerging constraints, multiple efforts are being launched toward designing an ML-based surrogate model that estimates resource usage of synthesized accelerator architectures. This model would reduce the design iteration time, especially when designing within a set of given hardware constraints. This approach shows considerable potential, but as it stands, the effort is early and would benefit from coordination and standardization to assist future work as it emerges. We introduce wa-hls4ml, a benchmark for ML accelerator resource and latency estimation, and its corresponding initial dataset of more than 100,000 fully connected neural networks, all synthesized using hls4ml and targeting Xilinx FPGAs. In addition to the resource utilization and latency data provided, the dataset includes generated artifacts and log files for many of the synthesized neural networks, in order to support future research in ML-based code generation. The benchmark evaluates the performance of resource and latency predictors against several common ML model architectures, primarily originating from scientific domains, as exemplar models, as well as the average performance across a subset of the dataset. We measure the performance of a given predictor model through multiple metrics, including $R^2$ score and SMAPE on regression tasks, as well as inference time to further characterize the estimator under test. Additionally, we introduce the latency/utilization inference graph neural network (lui-gnn), a surrogate model that uses a graph neural network to represent input architectures in the form of a directed graph. This graph representation allows for a diverse set of model architectures to all be effectively handled by a surrogate model. We present the architecture and performance of the model, as evaluated by the new proposed benchmark, including SMAPE, $R^2$ score, and inference times, and find that lui-gnn generally predicts latency and utilization for the 75\% quantile within several percent of the synthesized resources on the synthetic test dataset, indicating that this approach of estimating resource and latency via a surrogate models has promise and warrants further research.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

HipBone: A performance-portable graphics processing unit-accelerated C++ version of the NekBone benchmark

We present hipBone, an open-source performance-portable proxy application for the Nek5000 (and NekRS) computational fluid dynamics applications. HipBone is a fully GPU-accelerated C++ implementation of the original NekBone CPU proxy application with several novel algorithmic and implementation improvements which optimize its performance on modern fine-grain parallel GPU accelerators. Our optimizations include a conversion to store the degrees of freedom of the problem in assembled form in order to reduce the amount of data moved during the main iteration and a portable implementation of the main Poisson operator kernel. We demonstrate near-roofline performance of the operator kernel on three different modern GPU accelerators from two different vendors. We present a novel algorithm for splitting the application of the Poisson operator on GPUs which aggressively hides MPI communication required for both halo exchange and assembly. Our implementation of nearest-neighbor MPI communication then leverages several different routing algorithms and GPU-Direct RDMA capabilities, when available, which improves scalability of the benchmark. We demonstrate the performance of hipBone on three different clusters housed at Oak Ridge National Laboratory, namely, the Summit supercomputer and the Frontier early-access clusters, Spock and Crusher. Our tests demonstrate both portability across different clusters and very good scaling efficiency, especially on large problems.

Computer Science↗