Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Performance benchmark”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Modeling Air Handling Units to Create a Diverse Fault Dataset for FDD Innovation: Lessons Learned and Recommendations

As energy management and information systems (e.g., automated fault detection and diagnostics [AFDD] tools) become more prevalent in the commercial building stock, it is important to determine the effectiveness of these technologies by benchmarking their performance. The authors have been working to develop the largest publicly available dataset of HVAC fault datasets for performance benchmarking applications, covering the most common HVAC systems and designs including chiller plants, rooftop packaged units, dual duct air handling unit and single duct air handling units. This study covers the development, modeling, and validation of a synthetic fault dataset for the air handling unit (AHU), one of the most common HVAC configurations found in the commercial building stock. Despite this being a common system, real-world time series data are scarce and usually do not span a wide range of weather conditions. Due to this limitation, two detailed AHU models, which included the single duct AHU and dual duct AHU developed in the Modelica language and HVACSIM+ were employed to carry out annual simulations of numerous common sensor faults, mechanical faults, and control sequence faults. The fault inclusive data were then validated by comparing fault effects on system performance to expected symptoms. We summarize the nature of each fault and their impacts under different weather and operation conditions. We report some lessons learnt during the efforts of validating the high volumes of the FDD data sets. Finally, we highlight considerations for FDD developers that may want to use this dataset to assess their algorithms’ performance and their improvement over time.

Casillas, Armando↗

Benchmark Tracking System for Performance Monitoring

Benchmarking is essential for high-performance software development, particularly for monitoring performance across code iterations. This project focused on enhancing the benchmarking process for Lamellar, an asynchronous runtime for High-Performance Computing (HPC) systems developed at Pacific Northwest National Laboratory. Prior to this work, benchmark results were difficult to track and compare across code versions, presenting significant challenges in identifying performance regressions and long-term trends. The primary objective was to establish a systematic, reproducible approach for measuring performance and detecting regressions following code commits. Our methodology involved three key components: standardizing benchmark outputs, implementing data versioning, and developing analysis tools. We standardized the benchmark output format to JSON Line records containing specific fields (execution time, hardware specifications, and environmental variables). To address data management challenges, we evaluated several options and eventually chose a git repository dedicated to benchmark data. We developed a suite of Python tools that processed benchmark results, enriched them with metadata, and facilitated search in the repository. The resulting system enables more efficient filtering and comparison of performance metrics across commit histories, hardware configurations, and benchmark variants through a unified query interface. Our implementation reduces computational overhead by first checking for existing results through configuration matching before initiating new benchmark runs, thereby conserving resources. The system has been validated by Lamellar developers. It organizes results by benchmark type and build configurations for efficient retrieval. Future developments include a planned Large Language Model interface for predicting benchmark performance, incorporating the criterion package for statistical analysis, which will enable automated detection of statistically significant performance changes, and integration with continuous integration pipelines. Despite these enhancements being reserved for future work, this project has successfully provided the Lamellar development team with a framework for maintaining consistent performance standards and identifying optimization opportunities across workloads and hardware environments.

97 MATHEMATICS AND COMPUTING↗

The Case for and Against a Gadolinium Bias in SCALE: Round 2

The “Opening Arguments” for and against a gadolinium bias in SCALE were presented at the American Nuclear Society Annual Meeting in Philadelphia, Pennsylvania, in June, 2018. Some critical experiments included in the Oak Ridge National Laboratory Verified, Archived Library of Inputs and Data (VALID) indicate a significant bias as a function of gadolinium concentration. Other experiments indicate that no significant bias exists. The work presented here develops a larger suite of gadolinium-bearing benchmark models to further examine code, data, and benchmark performance. The new benchmark models have been reviewed for accuracy, but documentation and review for addition to the VALID library have not been completed. The problematic benchmarks included in VALID are HEU-SOL-THERM-014 and -016. These are two evaluations from a series of experiments from the Institute for Physics and Power Engineering (IPPE), Russia, documented in the International Criticality Safety Benchmark Evaluation Project (ICSBEP) Handbook. The entire set of evaluations also includes HEU-SOL-THERM-015, -017, -018, -019, and -025. Each evaluation contains a different uranium concentration, and different configurations within each evaluation include different gadolinium concentrations. These 7 evaluations contain a total of 52 configurations and form the largest subset of experiments considered, and they allow for a more complete assessment of the performance of these benchmarks than has historically been possible using just the HEU-SOL-THERM-014 and -016 results. Additional solution experiments are considered, including MIX-SOL-THERM-006 and -007 and PU-SOL-THERM-034. MIX-SOL-THERM-007 is in the VALID library, whereas MIX-SOL-THERM-006 and PU-SOL-THERM-034 are not. The results for MIX-SOL THERM-007 have not shown a gadolinium trend. These mixed- and plutonium-fueled solutions include 28 configurations. Some experiments including solid fuel and solid gadolinium are also included. These experiments include highly enriched uranium (HEU) foils moderated with polyethylene in HEU-MET-THERM-010, -016, and -034, and low enriched uranium (LEU) pin arrays with gadolinia absorber rods in LEU-COMP THERM-036 and -043. A total 32 cases with solid fuel are included. The IPPE solution benchmarks show a fairly high degree of variability, but no clear trend as a function of gadolinium concentration can be observed. The mixed- and plutonium-fueled solutions show less variability than the HEU solutions and also no trend relative to gadolinium concentration. The solid-fueled experiments also show no trend as a function of gadolinium content. The entire set of benchmarks shows no clear trend on gadolinium content or the energy of the average neutron lethargy causing fission.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Benchmarking and performance analysis of the CM-2

A suite of benchmarking routines testing communication, basic arithmetic operations, and selected kernel algorithms written in LISP and PARIS was developed for the CM-2. Experiment runs are automated via a software framework that sequences individual tests, allowing for unattended overnight operation. Multiple measurements are made and treated statistically to generate well-characterized results from the noisy values given by cm:time. The results obtained provide a comparison with similar, but less extensive, testing done on a CM-1. Tests were chosen to aid the algorithmist in constructing fast, efficient, and correct code on the CM-2, as well as gain insight into what performance criteria are needed when evaluating parallel processing machines.

Myers, David W.↗

Editorial: Neuroscience, computing, performance, and benchmarks: Why it matters to neuroscience how fast we can compute

At the turn of the millennium the computational neuroscience community realized that neuroscience was in a software crisis: software development was no longer progressing as expected and reproducibility declined. The International Neuroinformatics Coordinating Facility (INCF) was inaugurated in 2007 as an initiative to improve this situation. The INCF has since pursued its mission to help the development of standards and best practices. In a community paper published this very same year, Brette et al. tried to assess the state of the field and to establish a scientific approach to simulation technology, addressing foundational topics, such as which simulation schemes are best suited for the types of models we see in neuroscience. In 2015, a Frontiers Research Topic “Python in neuroscience” by Muller et al. triggered and documented a revolution in the neuroscience community, namely in the usage of the scripting language Python as a common language for interfacing with simulation codes and connecting between applications. The review by Einevoll et al. documented that simulation tools have since further matured and become reliable research instruments used by many scientific groups for their respective questions. Open source and community standard simulators today allow research groups to focus on their scientific questions and leave the details of the computational work to the community of simulator developers. A parallel development has occurred, which has been barely visible in neuroscientific circles beyond the community of simulator developers: Supercomputers used for large and complex scientific calculations have increased their performance from ~10 TeraFLOPS (10 13 floating point operations per second) in the early 2000s to above 1 ExaFLOPS (10 18 floating point operations per second) in the year 2022. This represents a 100,000-fold increase in our computational capabilities, or almost 17 doublings of computational capability in 22 years. Moore's law (the observation that it is economically viable to double the number of transistors in an integrated circuit every other 18–24 months) explains a part of this; our ability and willingness to build and operate physically larger computers, explains another part. It should be clear, however, that such a technological advancement requires software adaptations and under the hood, simulators had to reinvent themselves and change substantially to embrace this technological opportunity. It actually is quite remarkable that—apart from the change in semantics for the parallelization—this has mostly happened without the users knowing. The current Research Topic was motivated by the wish to assemble an update on the state of neuroscientific software (mostly simulators) in 2022, to assess whether we can see more clearly which scientific questions can (or cannot) be asked due to our increased capability of simulation, and also to anticipate whether and for how long we can expect this increase of computational capabilities to continue.

biophysically detailed models↗

Benchmarking Memory Performance with the Data Cube Operator

Data movement across a computer memory hierarchy and across computational grids is known to be a limiting factor for applications processing large data sets. We use the Data Cube Operator on an Arithmetic Data Set, called ADC, to benchmark capabilities of computers and of computational grids to handle large distributed data sets. We present a prototype implementation of a parallel algorithm for computation of the operatol: The algorithm follows a known approach for computing views from the smallest parent. The ADC stresses all levels of grid memory and storage by producing some of 2d views of an Arithmetic Data Set of d-tuples described by a small number of integers. We control data intensity of the ADC by selecting the tuple parameters, the sizes of the views, and the number of realized views. Benchmarking results of memory performance of a number of computer architectures and of a small computational grid are presented.

Frumkin, Michael A.↗

Benchmarking the performance of a high-Q cavity qudit using random unitaries

High-coherence cavity resonators are excellent resources for encoding quantum information in higher-dimensional Hilbert spaces, moving beyond traditional qubit-based platforms. A natural strategy is to use the Fock basis to encode information in qudits. One can perform quantum operations on the cavity mode qudit by coupling the system to a non-linear ancillary transmon qubit. However, the performance of the cavity-transmon device is limited by the noisy transmons. It is, therefore, important to develop practical benchmarking tools for these qudit systems in an algorithm-agnostic manner. We gauge the performance of these qudit platforms using sampling tests such as the heavy output generation test as well as the linear cross-entropy benchmark, by way of simulations of such a system subject to realistic dominant noise channels. We use selective number-dependent arbitrary phase and unconditional displacement gates as our universal gateset. Our results show that contemporary transmons comfortably enable controlling a few tens of Fock levels of a cavity mode. This framework allows benchmarking even higher dimensional qudits as those become accessible with improved transmons.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Predicting Cost/Performance Trade-offs For Whitney: A Commodity Computing Cluster

Recent advances in low-end processor and network technology have made it possible to build a "supercomputer" out of commodity components. We develop simple models of the NAS Parallel Benchmarks version 2 (NPB 2) to explore the cost/performance trade-offs involved in building a balanced parallel computer supporting a scientific workload. By measuring single processor benchmark performance, network latency, and network bandwidth, and using closed form expressions detailing the number and size of messages sent by each benchmark, our models predict benchmark performance to within 30%. A comparison based on total system cost reveals that current commodity technology (200 MHz Pentium Pros with 100baseT Ethernet) is well balanced for the NPBs up to a total system cost of around $ 1,000,000.

Becker, Jeffrey C.↗

Benchmarking the performance of uncertainty quantification methods for neural network-based interatomic potentials

Machine-learned interatomic potentials (ML-IAPs) continue to gain popularity as accurate, computationally efficient replacements for traditional, physics-based interatomic potentials and expensive ab initio methods. Uncertainty quantification (UQ) of ML-IAPs is a growing area of research as UQ is critical in many applications of IAPs, such as developing curated datasets, active learning-based data augmentation, self-improving models, and estimating the uncertainty of molecular dynamics simulations. In this paper, we construct and benchmark a series of different neural network potentials (NNPs) with varying network architectures to determine the performance of these models with respect to both the mean and uncertainty calibration error. Each NNP method is specifically designed to predict either epistemic or aleatoric uncertainty with particular focus on the differences in behavior between the epistemic and aleatoric uncertainty estimates. We benchmark these methods using multiple datasets common in the ML-IAP literature. The results show that the aleatoric uncertainty from single-shot model architectures is a competitive alternative to ensemble-based epistemic uncertainty predictions in regions of sufficient data-density. However, in regions where the representative data is sparse, aleatoric uncertainty models tend to overpredict and epistemic methods tend to underpredict the actual model error. We conclude that the type of UQ is crucial when discussing performance of probabilistic model results as different methods have different performance characteristics depending on the regime in which they are evaluated. Therefore, the type of UQ method should be carefully evaluated against both the data characteristics and requirements for the intended application.

97 MATHEMATICS AND COMPUTING↗

Agilent CRADA (Abstract)

The CRADA between Agilent Technologies Inc. and Battelle will focus on five software components as listed below: Prototype 4D Feature Finding functionality with a particular focus on recovering low level features and extending the bottom end dynamic range of IM-MS technology. Compare and contrast developments to current 4D Feature Finding capabilities. Highlight important algorithmic aspects employed. Implement the PNNL saturation correction algorithm. Agilent will give PNNL the needed data file access API and assistance in understanding it implementation and any needed instrumental aspects. Supported high resolution products to include Agilent’s TOF, QTOF and IM-QTOF mass spectrometers. PNNL will then work with Agilent to benchmark performance. Implementation of the PNNL Hadamard de-multiplexing algorithm. Agilent will give provide PNNL the needed date file access API access and as needed assistance in understanding the current Agilent multiplexed IM offering. PNNL will then work with Agilent on benchmark performance. Add ion mobility collision cross sections to existing and new metabolomic libraries for data analysis with Agilent’s informatics program MPP/ID Browser. PNNL will work with Agilent to create a software pipeline that takes data from chemical and metabolic standards and properly formats it for inclusion in MPP accessible libraries, using the collision cross section as a new separation dimension. Improvements of MPP multidimensional matching to identify metabolomic features using multiple characteristics beyond retention time and accurate mass. Most significantly matching will include analyte collision cross section with proposed support for sample fraction or RapidFire cartridge and fragmentation spectra. PNNL will work with Agilent to modify and improve the current MPP analysis pipeline to allow for creating, aligning, and identifying MS features defined by accurate mass, collision cross section and chromatographic retention time. As additional criteria such as fraction or RapidFire cartridge type are supported in the identification process, then they also will become part of the automation workflow. This includes the automation of said system to work with command line program (i.e. not a GUI) sufficient for programmatic execution in a pipeline.

97 MATHEMATICS AND COMPUTING↗

A full CI treatment of Ne atom - A benchmark calculation performed on the NAS CRAY 2

Full CI calculations are performed for Ne atom using Gaussian basis sets of up to triple-zeta plus double polarization quality. The total valence correlation energy through double, triple, quadruple and octuple excitations is compared for eight different basis sets. These results are expected to be an important benchmark for calibrating methods for estimating the importance of higher excitations.

Bauschlicher, C. W., Jr.↗

Sintered Cathodes for All-Solid-State Structural Lithium-Ion Batteries

All-solid-state structural lithium ion batteries serve as both structural load-bearing components and as electrical energy storage devices to achieve system level weight savings in aerospace and other transportation applications. This multifunctional design goal is critical for the realization of next generation hybrid or all-electric propulsion systems. Additionally, transitioning to solid state technology improves upon battery safety from previous volatile architectures. This research established baseline solid state processing conditions and performance benchmarks for intercalation-type layered oxide materials for multifunctional application. Under consideration were lithium cobalt oxide and lithium nickel manganese cobalt oxide. Pertinent characteristics such as electrical conductivity, strength, chemical stability, and microstructure were characterized for future application in all-solid-state structural battery cathodes. The study includes characterization by XRD, ICP, SEM, ring-on-ring mechanical testing, and electrical impedance spectroscopy to elucidate optimal processing parameters, material characteristics, and multifunctional performance benchmarks. These findings provide initial conditions for implementing existing cathode materials in load bearing applications.

Ceramics↗

Influence of Pt-Metal Alloy Catalysts with Various Ionomers on Oxygen Reduction Reaction in Fuel Cell Application

Pt-M/C (M = Co, Ni, Mn, etc.) alloy catalysts exhibit superior oxygen reduction reaction (ORR) activity compared to pure Pt/C, leading to a high energy efficiency in hydrogen fuel cells. However, many Pt-M/C alloy catalysts were synthesized and evaluated at the lab scale in model test-bed systems like rotating disc electrodes, which don't always correlate to performance within a fuel cell system; there is a clear need to evaluate catalysts in electrodes that can be prepared at industrially relevant scales to evaluate how factors like ink formulation can greatly affect device-level of fuel cell performance. Herein, three commercial Pt-M/C alloy catalysts (two Pt-Co/C and one Pt-Ni/C) were comprehensively characterized by various techniques. The results show that the average particle sizes of the three catalysts are close to 5 nm; the atomic ratio of Pt/M is around 4; and the M was successfully embedded into Pt lattice, resulting in the positive shift of Pt 4f in XPS spectra and XRD patterns. These catalytic materials were incorporated into 9 different cathode catalyst layers (CCLs) with three kinds of ionomers (Nafion D2020, high oxygen permeability ionomer (HOPI), and Aquivion D79-25BS), and their performance in proton exchange membrane fuel cells (PEMFCs) were investigated. The results demonstrate that the Pt-Co/C catalysts possess a higher mass activity (MA) than Pt-Ni/C; the cathodes with Nafion ionomer provide the highest MA while electrodes with Aquivion ionomer showed the lowest activity, attributed to poor H+ conductivity resulting from suboptimal ionomer incorporation. Finally, these alloys were shown to exceed DOE targets for MA and H2/Air performance reported in the recent publications at beginning of life and after 90k cycle catalyst AST protocol. This study provides valuable performance benchmarks for these materials guiding future Pt-M/C catalyst design and material integration for heavy duty PEMFC applications.

08 HYDROGEN↗

Airport Ground Support Equipment Infrastructure & Logistics Electrification Assessment Tool: 2025 Data Development, Modeling and Analysis for DFW

The aviation industry is increasingly turning to modernize freight facilities by integrating electric Ground Support Equipment (eGSE) to enhance operational efficiency of freight facility moving vehicles and equipment. Airports worldwide are adopting eGSE to streamline cargo movement, reduce fuel and maintenance costs, and improve logistics coordination.1 North America, with its advanced aviation infrastructure, leads this transition, leveraging Internet of things (IoT)-enabled automation and zero emission technologies to boost reliability and reduce human errors.2 Electrification of freight facility moving vehicles and equipment boosts turnaround times, improves equipment reliability, and optimizes logistics coordination, giving operators a competitive advantage. With rising fuel price volatility and the pressure to meet stringent performance benchmarks, airports are focusing on cost-effective, scalable solutions for long-term financial and operational gains. To further accelerate electrification, airports are integrating Zero Emission Vehicles (ZEVs) into rental car fleets and deploying electric baggage carts, requiring strategic investments in charging infrastructure. 3 The shift, however, presents challenges, such as limited technical expertise, high capital costs, and complex procurement processes. By forging strategic partnerships, leveraging advanced technologies, and optimizing infrastructure investments, airports can create a resilient, future-ready ecosystem that enhances the movement of people and goods through electrification-driven efficiency. Supported by the U.S. Department of Energy (DOE) Vehicle Technologies Office (VTO), this electrification effort provides a scalable, cost-effective solution to improve airport freight operations. Through targeted investments and innovation, airports enhance efficiency, reduce costs, and meet performance benchmarks while advancing toward a resilient, electrified future.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Real-time inference and extrapolation with Time-Conditioned UNet: Applications in hypersonic flows, incompressible flows, and global temperature forecasting

Neural Operators are fast and accurate surrogates for nonlinear mappings between functional spaces within training domains. Extrapolation beyond the training domain remains a grand challenge across all application areas. We present Time-Conditioned UNet (TC-UNet) as an operator learning method to solve time-dependent PDEs continuously in time without any temporal discretization, including in extrapolation scenarios. TC-UNet incorporates the temporal evolution of the PDE into its architecture by combining a parameter conditioning approach with the attention mechanism from the Transformer architecture. After training, TC-UNet makes real-time inferences on an arbitrary temporal grid. We demonstrate its extrapolation capability on a climate problem by estimating the global temperature for several years and also for inviscid hypersonic flow around a double cone. We propose different training strategies involving temporal bundling and sub-sampling. We demonstrate performance improvements for several benchmarks, performing extrapolation for long time intervals and zero-shot super-resolution time.

Deep learning↗

Laser Beam Welding Benchmark Experiments Performed in Reduced Gravity and Vacuum

Laser beam welding (LBW) is affected by the extreme temperatures, reduced pressure, and reduced gravity present in space environments. Gravity and pressure especially influence its melt pool and solidification dynamics. A compact, modular vacuum chamber adaptable to flight platforms from parabolic to orbital currently hosts an experiment to investigate the combined influence of reduced gravity and pressure on LBW. A swappable cartridge contains a rotating platen on which customizable workpieces can be welded under vacuum, greatly increasing experimental throughput. Instrumentation includes weld and thermal cameras observing the process, thermocouples placed on workpieces, accelerometers, and vacuum sensors. Experimental data gathered during the welding process will be combined with post-flight nondestructive evaluation, metallography, and mechanical testing to provide validation datasets for computational modeling. Phase I of this effort involves a parabolic flight campaign in low gravity while an anticipated Phase II would proceed to in-space demonstration to access extended duration microgravity.

in-space welding↗