Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Performance benchmark”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

DAISY Benchmark Performance Data

This repository contains the underlying data from benchmark experiments for Drifting Acoustic Instrumentation SYstems (DAISYs) in waves and currents described in "Performance of a Drifting Acoustic Instrumentation SYstem (DAISY) for Characterizing Radiated Noise from Marine Energy Converters" (https://link.springer.com/article/10.1007/s40722-024-00358-6). DAISYs consist of a surface expression connected to a hydrophone recording package by a tether. Both elements are instrumented to provide metadata (e.g., position, orientation, and depth). Information about how to build DAISYs is available at https://www.pmec.us/research-projects/daisy. The repository's primary content is three compressed archives (.zip format), each containing multiple MATLAB binary data files (.mat format). A table relating individual data files to figures in the paper, as well as the structure of each file, is included in the repository as a Word document (Data Description MHK-DR.docx). Most of the files contain time series information for a single DAISY deployment (file naming convention: [site]_DAISY_[Drift #].mat) consisting of processed hydrophone data and associated metadata. For a limited number of DAISY deployments, the hydrophone package was replaced with an acoustic Doppler velocimeter (file naming convention: [site]_DAISY_[Drift #]_ADV.mat). Data were collected over several years at three locations: (1) Sequim Bay at Pacific Northwest National Laboratory's Marine & Coastal Research Laboratory (MCRL) in Sequim, WA, the energetic tidal channel in Admiralty Inlet, WA (Admiralty Inlet), and the U.S. Navy's Wave Energy Test Site (WETS) in Kaneohe, HI. Brief descriptions of data files at each location follow. - MCRL - (1) Drift #4 and #16 contrast the performance of a DAISY and a reference hydrophone (icListen HF Reson), respectively, in the quiescent interior of Sequim Bay (September 2020). (2) Drift #152 and #153 are velocity measurements for a drifting acoustic Doppler velocimeter in in the tidally-energetic entrance channel inside a flow shield and exposed to the flow, respectively (January 2018). (3) Two non-standard files are also included: DAISY_data.mat corresponds to a subset of a DAISY drift over an Adaptable Monitoring Package (AMP) and AMP_data.mat corresponds to approximately co-temporal data for a stationary hydrophone on the AMP (February 2019). - Admiralty Inlet - (1) Drift #1-12 correspond to tests with flow shielded DAISYs, unshielded DAISYs, a reference hydrophone, and drifting acoustic Doppler velocimeter with 5, 10, and 15 m tether lengths between surface expression and hydrophone recording package (July 2022). (2) Drift #13-20 correspond to tests of flow shielded DAISYs with three different tether materials (rubber cord, nylon line, and faired nylon line) in lengths of 5, 10, and 15 m (July 2022). - WETS - (1) Drift #30-32 correspond to tests with a heave plate incorporated into the tether (standard configuration for wave sites), rubber cord only, and rubber cord, but with a flow shielded hydrophone (November 2022). (2) Drift #49-58 and Drift #65-68 correspond to measurements around mooring infrastructure at the 60 m berth where time-delay-of-arrival localization was demonstrated for different DAISY arrangements and hydrophone depths (November 2022).

16 TIDAL AND WAVE POWER↗

Two MCNP Models for Computational Performance Benchmarking

The purpose of this report is to describe two non-trivial MCNP models and execution configurations that are suitable to characterizing computational performance. To that end, this report uses a constructive solid geometry (CSG) representation of the Oak Ridge National Laboratory (ORNL) Pool Critical Assembly (PCA) [1–3] to perform a k-eigenvalue calculation and an unstructured mesh (UM) representation of the International Commission on Radiological Protection (ICRP) publication 145 (ICRP145) male human phantom [4, 5] to perform a fixed-source calculation assuming that a 1 MeV photon source is distributed throughout the phantom’s liver.

97 MATHEMATICS AND COMPUTING↗

Performance Benchmark of Commercial and Developmental Fission Chambers in Elevated Temperatures

This report documents the testing of two in-core fission chamber technologies for high-temperature irradiation environments. This work is in collaboration with the French Alternative Energies and Atomic Energy Commission (CEA). The fission chamber evaluated by Idaho National Laboratory is the micro-pocket fission detector (MPFD). The fission chamber evaluated by the CEA are the 3 mm miniaturized fission chamber and the 7 mm high-temperature fission chamber. Demonstrations of the MPFD were performed at the Neutron Radiography Facility and the Massachusetts Institute of Technology Reactor. Demonstrations of the CEA fission chambers were performed at the Ohio State University Research Reactor. All demonstrations were conducted with a heated experiment rig up to 850°C.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

High-throughput calculations of charged point defect properties with semi-local density functional theory—performance benchmarks for materials screening applications

Abstract Calculations of point defect energetics with Density Functional Theory (DFT) can provide valuable insight into several optoelectronic, thermodynamic, and kinetic properties. These calculations commonly use methods ranging from semi-local functionals with a-posteriori corrections to more computationally intensive hybrid functional approaches. For applications of DFT-based high-throughput computation for data-driven materials discovery, point defect properties are of interest, yet are currently excluded from available materials databases. This work presents a benchmark analysis of automated, semi-local point defect calculations with a-posteriori corrections, compared to 245 “gold standard” hybrid calculations previously published. We consider three different a-posteriori correction sets implemented in an automated workflow, and evaluate the qualitative and quantitative differences among four different categories of defect information: thermodynamic transition levels, formation energies, Fermi levels, and dopability limits. We highlight qualitative information that can be extracted from high-throughput calculations based on semi-local DFT methods, while also demonstrating the limits of quantitative accuracy.

36 MATERIALS SCIENCE↗

Quantum-Hardware Focused Application Performance Benchmarks (Final Technical Report)

Quantum computers promise to transform how we do scientific calculations. In this project, we benchmark different current quantum computers for solving chemistry problems and find ways to build noise-resilient implementations. We also use the characterization techniques developed to test ion trap quantum computers.

74 ATOMIC AND MOLECULAR PHYSICS↗

Using Likwid and Byfl to Benchmark Hardware Performance [Poster]

Benchmark Study conducted focusing on CPU and program performance analysis. Performance data gathered using 2 different programs and comparisons made based on performance. After comparisons are made, conclusions can be drawn and improvements are made upon hardware and software.

97 MATHEMATICS AND COMPUTING↗

Using Likwid and Byfl to Benchmark Hardware Performance

This paper outlines a benchmarking study conducted during my internship at LANL, focusing on CPU (Computer Processing Unit) and program performance assessment. The primary goal was to gather memory access data using three methods across five polybench kernels The data gathered would then be used to compare and contrast to one another and calculate operational intensity for performance comparisons. Benchmarking tools like Byfl and Likwid were employed, with Byfl offering hardware-independent data through LLVM compiler communication and Likwid directly interacting with computer hardware. The study considered various benchmarking factors, including optimization levels, Big O notation ((n)), CPU diversity and specific kernel equations. Big O notation was utilized to simplify code complexity, with detailed breakdwons of operations and memory components for each polybench application. Specific O(n) equations enabled nuanced kernel compariosns, facilitating the identification of performance variations. CPU efficiency assessments were conducted using Likwid tests on two CPUs. The central focus on code optimization aimed at achieving higher speeds and reduced memory usage through streamlined code. Future work propsoes creating a roofline model, synthesizing benchmarking data into a comprehensive data graph to assist in optimizing code and improving hardware performance. The potential impact on the laboratory or national mission was underscored, emphasizing the importance of optimizing applications and hardware to conserve resources and accelerate program execution. The specific relevance to LANL’s operations in math-intensive fields such as Nuclear Fission, Space Exploration, and Nanotechnology highlights the necessity of efficient benchmarking for resource conservation and proram speed. Overall, this study contributes to the understanding of CPU and program performance, providing insights for future optimization efforts in a laboratory setting

97 MATHEMATICS AND COMPUTING↗

Development of a Annual Air Handling Unit Fault Dataset for FDD Tools: Lessons Learned and Considerations for FDD Developers

As energy management and information systems (e.g., automated fault detection and diagnostics [AFDD] tools) become more prevalent in the commercial building stock, it is important to determine the effectiveness of these technologies by benchmarking their performance. The authors have been working to develop the largest publicly available dataset of HVAC fault data for performance benchmarking applications, covering the most common HVAC systems and designs including chiller plants, rooftop packaged units, dual duct air handling units and single duct air handling units. This study covers the development, modeling, and validation of a synthetic fault dataset for a single duct air handling unit (AHU), one of the most common HVAC configurations found in the commercial building stock. Despite this being a common system, real-world time series data are scarce and usually do not span a wide range of weather conditions. Due to this limitation, a detailed AHU model was employed to carry out annual simulations of numerous common sensor and mechanical faults, which were then validated by comparing their effects on system performance to expected symptoms. We summarize the nature of each fault and their impacts under different weather and operation conditions. Finally, we highlight considerations for FDD developers that may want to use this dataset to assess their algorithms’ performance and their improvement over time.

Casillas, Armando↗

Modeling Air Handling Units to Create a Diverse Fault Dataset for FDD Innovation: Lessons Learned and Recommendations

As energy management and information systems (e.g., automated fault detection and diagnostics [AFDD] tools) become more prevalent in the commercial building stock, it is important to determine the effectiveness of these technologies by benchmarking their performance. The authors have been working to develop the largest publicly available dataset of HVAC fault datasets for performance benchmarking applications, covering the most common HVAC systems and designs including chiller plants, rooftop packaged units, dual duct air handling unit and single duct air handling units. This study covers the development, modeling, and validation of a synthetic fault dataset for the air handling unit (AHU), one of the most common HVAC configurations found in the commercial building stock. Despite this being a common system, real-world time series data are scarce and usually do not span a wide range of weather conditions. Due to this limitation, two detailed AHU models, which included the single duct AHU and dual duct AHU developed in the Modelica language and HVACSIM+ were employed to carry out annual simulations of numerous common sensor faults, mechanical faults, and control sequence faults. The fault inclusive data were then validated by comparing fault effects on system performance to expected symptoms. We summarize the nature of each fault and their impacts under different weather and operation conditions. We report some lessons learnt during the efforts of validating the high volumes of the FDD data sets. Finally, we highlight considerations for FDD developers that may want to use this dataset to assess their algorithms’ performance and their improvement over time.

Casillas, Armando↗

Benchmark Tracking System for Performance Monitoring

Benchmarking is essential for high-performance software development, particularly for monitoring performance across code iterations. This project focused on enhancing the benchmarking process for Lamellar, an asynchronous runtime for High-Performance Computing (HPC) systems developed at Pacific Northwest National Laboratory. Prior to this work, benchmark results were difficult to track and compare across code versions, presenting significant challenges in identifying performance regressions and long-term trends. The primary objective was to establish a systematic, reproducible approach for measuring performance and detecting regressions following code commits. Our methodology involved three key components: standardizing benchmark outputs, implementing data versioning, and developing analysis tools. We standardized the benchmark output format to JSON Line records containing specific fields (execution time, hardware specifications, and environmental variables). To address data management challenges, we evaluated several options and eventually chose a git repository dedicated to benchmark data. We developed a suite of Python tools that processed benchmark results, enriched them with metadata, and facilitated search in the repository. The resulting system enables more efficient filtering and comparison of performance metrics across commit histories, hardware configurations, and benchmark variants through a unified query interface. Our implementation reduces computational overhead by first checking for existing results through configuration matching before initiating new benchmark runs, thereby conserving resources. The system has been validated by Lamellar developers. It organizes results by benchmark type and build configurations for efficient retrieval. Future developments include a planned Large Language Model interface for predicting benchmark performance, incorporating the criterion package for statistical analysis, which will enable automated detection of statistically significant performance changes, and integration with continuous integration pipelines. Despite these enhancements being reserved for future work, this project has successfully provided the Lamellar development team with a framework for maintaining consistent performance standards and identifying optimization opportunities across workloads and hardware environments.

97 MATHEMATICS AND COMPUTING↗

Updates and Validation for the n+ 63,65 Cu Cross Sections [Abstract]

The neutron induced total, elastic, and capture cross sections of 63,65 Cu isotopes were selected for evaluation in the resolved and unresolved resonance energy ranges by the National Criticality Safety Program to resolve discrepancies related to benchmark performance. This is especially evident for the series of ZEUS benchmarks in which copper is used as a reflector. Because copper is also used as structural material in both fission and fusion reactors, the need to address benchmark discrepancies linked to nuclear data deficiencies is a task of primary importance. The aim of this work is to describe the steps of evaluation work towards a consistent improvement of the benchmark performance. The R-matrix analysis with the SAMMY code focused on the 63 Cu(n,γ) reaction channel between 100-300 keV coupled to unresolved resonance region parameters up to 650 keV to fit average cross section data from a recent experiment. Due to the high sensitivity of many benchmarks to elastic scattering angular distribution data, especially for the 65 Cu isotope, the impact of these data was tested by generating Legendre coefficients from both resonance parameters and the Hauser-Feshbach model. Guided by the findings of Shaw et al., the performance of the current evaluation for 65 Cu was compared to that of ENDF/B-VII.1 and ENDF/B-VIII.0 by testing the reactivity coefficients corresponding to the validation suite of experimental criticality benchmarks for thermal, intermediate, and fast systems taken from the International Criticality Safety Benchmark Experiments Project Handbook. The benchmark performance is especially sensitive to 63 Cu(n,γ) and 65 Cu elastic scattering for neutron energies in the 100–500 keV region, whereas 100 keV is the upper limit of the resolved resonance region in the ENDF/B-VIII.0 evaluations for 63,65 Cu. The results highlight the need to handle the transition from the resolved resonance region to the high energy region carefully.

07 ISOTOPE AND RADIATION SOURCES↗

The Case for and Against a Gadolinium Bias in SCALE: Round 2

The “Opening Arguments” for and against a gadolinium bias in SCALE were presented at the American Nuclear Society Annual Meeting in Philadelphia, Pennsylvania, in June, 2018. Some critical experiments included in the Oak Ridge National Laboratory Verified, Archived Library of Inputs and Data (VALID) indicate a significant bias as a function of gadolinium concentration. Other experiments indicate that no significant bias exists. The work presented here develops a larger suite of gadolinium-bearing benchmark models to further examine code, data, and benchmark performance. The new benchmark models have been reviewed for accuracy, but documentation and review for addition to the VALID library have not been completed. The problematic benchmarks included in VALID are HEU-SOL-THERM-014 and -016. These are two evaluations from a series of experiments from the Institute for Physics and Power Engineering (IPPE), Russia, documented in the International Criticality Safety Benchmark Evaluation Project (ICSBEP) Handbook. The entire set of evaluations also includes HEU-SOL-THERM-015, -017, -018, -019, and -025. Each evaluation contains a different uranium concentration, and different configurations within each evaluation include different gadolinium concentrations. These 7 evaluations contain a total of 52 configurations and form the largest subset of experiments considered, and they allow for a more complete assessment of the performance of these benchmarks than has historically been possible using just the HEU-SOL-THERM-014 and -016 results. Additional solution experiments are considered, including MIX-SOL-THERM-006 and -007 and PU-SOL-THERM-034. MIX-SOL-THERM-007 is in the VALID library, whereas MIX-SOL-THERM-006 and PU-SOL-THERM-034 are not. The results for MIX-SOL THERM-007 have not shown a gadolinium trend. These mixed- and plutonium-fueled solutions include 28 configurations. Some experiments including solid fuel and solid gadolinium are also included. These experiments include highly enriched uranium (HEU) foils moderated with polyethylene in HEU-MET-THERM-010, -016, and -034, and low enriched uranium (LEU) pin arrays with gadolinia absorber rods in LEU-COMP THERM-036 and -043. A total 32 cases with solid fuel are included. The IPPE solution benchmarks show a fairly high degree of variability, but no clear trend as a function of gadolinium concentration can be observed. The mixed- and plutonium-fueled solutions show less variability than the HEU solutions and also no trend relative to gadolinium concentration. The solid-fueled experiments also show no trend as a function of gadolinium content. The entire set of benchmarks shows no clear trend on gadolinium content or the energy of the average neutron lethargy causing fission.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Editorial: Neuroscience, computing, performance, and benchmarks: Why it matters to neuroscience how fast we can compute

At the turn of the millennium the computational neuroscience community realized that neuroscience was in a software crisis: software development was no longer progressing as expected and reproducibility declined. The International Neuroinformatics Coordinating Facility (INCF) was inaugurated in 2007 as an initiative to improve this situation. The INCF has since pursued its mission to help the development of standards and best practices. In a community paper published this very same year, Brette et al. tried to assess the state of the field and to establish a scientific approach to simulation technology, addressing foundational topics, such as which simulation schemes are best suited for the types of models we see in neuroscience. In 2015, a Frontiers Research Topic “Python in neuroscience” by Muller et al. triggered and documented a revolution in the neuroscience community, namely in the usage of the scripting language Python as a common language for interfacing with simulation codes and connecting between applications. The review by Einevoll et al. documented that simulation tools have since further matured and become reliable research instruments used by many scientific groups for their respective questions. Open source and community standard simulators today allow research groups to focus on their scientific questions and leave the details of the computational work to the community of simulator developers. A parallel development has occurred, which has been barely visible in neuroscientific circles beyond the community of simulator developers: Supercomputers used for large and complex scientific calculations have increased their performance from ~10 TeraFLOPS (10 13 floating point operations per second) in the early 2000s to above 1 ExaFLOPS (10 18 floating point operations per second) in the year 2022. This represents a 100,000-fold increase in our computational capabilities, or almost 17 doublings of computational capability in 22 years. Moore's law (the observation that it is economically viable to double the number of transistors in an integrated circuit every other 18–24 months) explains a part of this; our ability and willingness to build and operate physically larger computers, explains another part. It should be clear, however, that such a technological advancement requires software adaptations and under the hood, simulators had to reinvent themselves and change substantially to embrace this technological opportunity. It actually is quite remarkable that—apart from the change in semantics for the parallelization—this has mostly happened without the users knowing. The current Research Topic was motivated by the wish to assemble an update on the state of neuroscientific software (mostly simulators) in 2022, to assess whether we can see more clearly which scientific questions can (or cannot) be asked due to our increased capability of simulation, and also to anticipate whether and for how long we can expect this increase of computational capabilities to continue.

biophysically detailed models↗

Benchmarking the performance of a high-Q cavity qudit using random unitaries

High-coherence cavity resonators are excellent resources for encoding quantum information in higher-dimensional Hilbert spaces, moving beyond traditional qubit-based platforms. A natural strategy is to use the Fock basis to encode information in qudits. One can perform quantum operations on the cavity mode qudit by coupling the system to a non-linear ancillary transmon qubit. However, the performance of the cavity-transmon device is limited by the noisy transmons. It is, therefore, important to develop practical benchmarking tools for these qudit systems in an algorithm-agnostic manner. We gauge the performance of these qudit platforms using sampling tests such as the heavy output generation test as well as the linear cross-entropy benchmark, by way of simulations of such a system subject to realistic dominant noise channels. We use selective number-dependent arbitrary phase and unconditional displacement gates as our universal gateset. Our results show that contemporary transmons comfortably enable controlling a few tens of Fock levels of a cavity mode. This framework allows benchmarking even higher dimensional qudits as those become accessible with improved transmons.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Benchmarking the performance of uncertainty quantification methods for neural network-based interatomic potentials

Machine-learned interatomic potentials (ML-IAPs) continue to gain popularity as accurate, computationally efficient replacements for traditional, physics-based interatomic potentials and expensive ab initio methods. Uncertainty quantification (UQ) of ML-IAPs is a growing area of research as UQ is critical in many applications of IAPs, such as developing curated datasets, active learning-based data augmentation, self-improving models, and estimating the uncertainty of molecular dynamics simulations. In this paper, we construct and benchmark a series of different neural network potentials (NNPs) with varying network architectures to determine the performance of these models with respect to both the mean and uncertainty calibration error. Each NNP method is specifically designed to predict either epistemic or aleatoric uncertainty with particular focus on the differences in behavior between the epistemic and aleatoric uncertainty estimates. We benchmark these methods using multiple datasets common in the ML-IAP literature. The results show that the aleatoric uncertainty from single-shot model architectures is a competitive alternative to ensemble-based epistemic uncertainty predictions in regions of sufficient data-density. However, in regions where the representative data is sparse, aleatoric uncertainty models tend to overpredict and epistemic methods tend to underpredict the actual model error. We conclude that the type of UQ is crucial when discussing performance of probabilistic model results as different methods have different performance characteristics depending on the regime in which they are evaluated. Therefore, the type of UQ method should be carefully evaluated against both the data characteristics and requirements for the intended application.

97 MATHEMATICS AND COMPUTING↗