Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Performance benchmark”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Preliminary RAM-RODD results for the MUSiC subcritical configurations

The Measurement of Uranium Subcritical and Critical (MUSiC) was performed at the DOE’s National Criticality Experiments Research Center (NCERC) located in the Nevada National Security Site (NNSS). The measurement utilized the Rocky Flats shells to perform benchmark measurements of similar highly enriched uranium (HEU) systems that span a wide range of reactivities. The Rocky Flats (RF) shells are 93.16% U-235 enriched metal hemishells that can be stacked concentrically. Ten configurations were measured with effective multiplication factors spanning between deeply subcritical (~ 0.64) through delayed critical. Details of the measured configurations are listed in Table 1. This unique set of measurements with its large span of reactivies is being used to determine the range over which neutron noise techniques such as Feynman variance-to-mean, Rossi-alpha, and pulsed neutron source techniques can be accurately employed for a bare HEU system. The results of the measurements will be published as a benchmark in The International Criticality Safety Benchmark Evaluation Project (ICSBEP) Handbook. The results will support the growing amount of subcritical benchmark data, such as the SCRaP measurements, that is available to the community. The measurements were performed using several different detector systems for the purposes of cross-validation and obtaining detector independent results. Four detector systems were deployed, three by the NCERC team and one from the University of Michigan. The detector systems included a Neutron Multiplicity Array Detector (NoMAD) system (similar to the MC-15), four small volume 0.635 cm (Ø) × 7.59 cm 3 He detectors (ideal for measuring prompt neutron decay constants due to their fast recovery speed), the Rossi Alpha Measurements – Rapid Organic (n, γ) Discrimination Detector (RAM-RODD), and the Organic Scintillator Array (OSCAR). RAM-RODD is an array of eight 5.08 cm (Ø) × 5.08 cm EJ-309 organic scintillator detectors. OSCAR is a University of Michigan system and is an array of twelve 5.08 cm (Ø) × 5.08 cm stilbene detectors. Details of measurements performed with OSCAR will be discussed in a separate talk. The focus of this work is preliminary results obtained by RAM-RODD for the 8 subcritical configurations. Additional details on the measurements and on the setup and deployment of RAM-RODD will be discussed. Preliminary neutron noise analysis results including Rossi-alpha and Feynman-alpha will also be presented.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

A new control score concept for building performance assessment

In buildings, performance assessment often focuses on energy use with metrics such as energy use intensity (EUI) used to benchmark performance. However, energy performance of a building is fundamentally determined by the control system that engages the energy-using systems. There are two aspects of control that are of particular importance: (1) the ability to regulate process variables to their setpoints; and (2) whether the setpoints are at the right levels and/or following desired profiles. Most buildings do not reach their energy efficiency potential due to deficiencies in control performance and operators do not have access to metrics that can illuminate these deficiencies. Here this paper addresses this problem by providing novel techniques that combine these two aspects of control performance into a single standardized score on the scale of 0-10. The concept of a standardized control scores enables all systems in a building to be compared on the same scale and also for scores to be rolled up to different levels in the building and system hierarchy for system-wide analysis. The paper presents the theory for the method, describes a prototype tool for displaying scores, and presents results from application to a large building in Minneapolis.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Membrane Pretreatment and Cell Conditioning for Proton Exchange Membrane Water Electrolysis

As the research community studying proton exchange membrane water electrolysis (PEMWE) grows, it is important to develop methods that achieve transparent, reproducible research. Reproducibility of performance and durability is a challenge facing the PEMWE field, as the published literature includes a wide spread of results obtained with nominally similar materials. Prior round-robin performance benchmarking efforts [1] have identified inadequate cell conditioning as a major source of variation in apparent cell performance. Inadequate pre-treatment and conditioning can lead to instability in initial performance and adds ambiguity to durability measurements. However, excessively prolonged conditioning procedures limit the throughput of testing and the pace of research. Pretreatment and conditioning methods vary significantly across the research literature, but little systematic investigation is available into the mechanisms of these procedures or how procedural differences may impact results. This presentation will discuss investigations into the effects of membrane pre-treatment and operating procedures during cell conditioning on initial performance and catalyst-specific accelerated stress tests, with the aim of recommending procedures to enable clear, reproducible, and high-throughput research. Methods investigated include the use of hydrogen peroxide, acids, and hydration at elevated temperature for membrane pretreatment, and conditioning procedures such as current or voltage holds and cycling. The investigations cover both the impacts of these procedures on cell performance and stability as well as underlying mechanisms and processes taking place in the cell materials.

cell↗

Using FIPD and OPTD to Benchmark Metallic Fuel Performance

This report serves as an introduction, tutorial, and benchmark specification for out-of-pile tests on metallic fuel. It introduces a new user to the EBR-II legacy fuel performance test program and the fast reactor fuel performance databases built to preserve the records. It then details the information stored in each database and how to find it. A benchmark specification is included for a small set of out-of-pile tests on U-10Zr fuel to function as a tutorial demonstrating how the legacy fuel performance data sets stored in the FIPD and OPTD databases can be used together to benchmark fuel performance models for steady-state and transient performance.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Charged point defect benchmark of Hybrid and GGA-PBE

Data used by the publication "High-throughput calculations of charged point defect properties with semi-local density functional theory - performance benchmarks for materials screening applications." This work presented an in-depth benchmark analysis of automated, semi-local point defect calculations with a-posteriori corrections, compared to 245 “gold standard” hybrid calculations previously published. We considered three different a-posteriori correction sets for semi-local calculations, implemented in a fully automated workflow, and consider the qualitative and quantitative differences for four different categories of defect information: thermodynamic transition levels, formation energies, fermi levels, and dopability limits. We highlighted the type of qualitative information about point defect properties that can be extracted from high-throughput calculations based on semi-local DFT methods, while also demonstrating the limits of quantitative accuracy that can be achieved by these approaches.

Broberg, Danny↗

HIPLZ: Enabling performance portability for exascale systems

While heterogeneous computing has emerged as a dominant trend in current and future High-Performance Computing (HPC) systems, it is also widely recognized that this shift has led to increased software complexity due to a proliferation of programming systems for different heterogeneous processors. One such example is the Heterogeneous-Compute Interface for Portability from AMD (HIP ), which is composed of a C Runtime API and C++ Kernel Language. Many HPC applications will likely use HIP on future exascale systems (e.g., Frontier and El Capitan), but HIP currently only targets AMD and NVIDIA processors. This limitation creates challenges for users who would also like to run their applications on exascale systems based on other architectures (e.g., Aurora, which is based on Intel hardware) that are currently not targeted by HIP . In this paper, we introduce the design and implementation of HIPLZ , a compiler and runtime system that uses the Intel Level Zero API to support HIP on Intel GPU architectures. We discuss the design of HIPLZ , derived from HIPCL (an implementation of HIP on top of OpenCL ), and portability issues that occur from using the Level Zero runtime as a backend. We evaluate our implementation by running several performance benchmarks and mini-apps written in HIP on Intel architectures using HIPLZ . Our results show that this approach provides competitive performance relative to Intel's OpenCL implementations on Intel Gen9 and UHD Graphics 770 GPUs, while providing good coverage of features needed by HPC applications. Overall, this approach is a promising demonstration of enabling performance portability for exascale systems.

97 MATHEMATICS AND COMPUTING↗

Acceleration of the particle-in-cell code Osiris with graphics processing units

Fully relativistic particle-in-cell (PIC) simulations are crucial for advancing our knowledge of plasma physics. Modern supercomputers based on graphics processing units (GPUs) offer the potential to perform PIC simulations of unprecedented scale, but require robust and feature-rich codes that can fully leverage their computational resources. In this work, this demand is addressed by adding GPU acceleration to the PIC code Osiris. An overview of the algorithm, which features a CUDA extension to the underlying Fortran architecture, is given. Detailed performance benchmarks for thermal plasmas are presented, which demonstrate excellent weak scaling on NERSC's Perlmutter supercomputer and high levels of absolute performance. The robustness of the code to model a variety of physical systems is demonstrated via simulations of Weibel filamentation and laser-wakefield acceleration run with dynamic load balancing. Finally, measurements and analysis of energy consumption are provided that indicate that the GPU algorithm is up to ~14 times faster and ~7 times more energy efficient than the optimized CPU algorithm on a node-to-node basis. The described development addresses the PIC simulation community's computational demands both by contributing a robust and performant GPU-accelerated PIC code and by providing insight into efficient use of GPU hardware.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Impact of Coating Defects on Performance of Coated Zirconium Cladding

Research on accident tolerant fuels (ATF) has started after the Fukushima accident [1–3]. While efforts have been expended on both fuel and cladding ATF concepts, the bulk of work has been devoted to improved cladding. The overarching goal of these approaches is to extend the coping time available during a severe accident before any event would result in release of radioactivity to the public. The most basic ATF cladding concept is obtained by applying a thin coating of highly corrosion-resistant material on the surface of a licensed zirconium cladding alloy. This thin coating is intended to not interfere with the neutronic or mechanical performance of the base cladding under normal operating conditions. Different coating materials, thicknesses, coating processes, process parameters, and testing methods have an impact on the microstructure and mechanical properties and therefore on the results of the applied investigation methods. These challenges have motivated an initial focus on demonstrating that the presence of coatings do not perturb the critical performance benchmarks of uncoated material. Ongoing lead test assembly irradiations of coated zirconium concepts in commercial reactions is intended to establish baseline performance in this regard in the coming years.

36 MATERIALS SCIENCE↗

TChem-atm (v2.0.0): scalable performance-portable multiphase atmospheric chemistry

We present TChem-atm, a performance-portable approach that enables efficient simulation of chemically detailed and multiphase atmospheric chemistry on modern heterogeneous computing architectures. Unlike previous efforts that rely on architecture-specific code or focus exclusively on gas-phase chemistry, TChem-atm supports fully coupled gas–aerosol systems with execution across CPUs, NVIDIA GPUs, and AMD GPUs through the Kokkos programming model. It integrates the flexible multiphase capabilities of the Community Atmospheric Model Chemistry Package (CAMP) with the high-performance kinetic routines of TChem, and includes automatic Jacobian construction with support for a range of stiff ODE solvers. In a proof-of-concept integration with the particle-resolved model PartMC, TChem-atm reproduces the existing PartMC–CAMP implementation within solver tolerances and delivers substantial GPU speedups, especially for large particle populations. Performance benchmarks reveal substantial speedups on GPU platforms, particularly for large particle populations, with consistent results across hardware backends. TChem-atm enables performance-portable execution across CPUs and GPUs, though optimal efficiency may require modest architecture-specific tuning (e.g., team and vector sizes), with up to a twofold improvement on the NVIDIA H100. It directly supports sectional and particle-resolved host models, while modal aerosol schemes require minor adaptation to provide particle-scale quantities such as representative diameters. By enabling chemically detailed, multiphase simulations with performance portability and host-model flexibility, TChem-atm facilitates the incorporation of advanced chemistry into atmospheric models.

Díaz-Ibarra, Oscar Homero [Sandia National Laborat↗

On the predictability of turbulent fluxes from land: PLUMBER2 MIP experimental description and preliminary results

Accurate representation of the turbulent exchange of carbon, water, and heat between the land surface and the atmosphere is critical for modelling global energy, water, and carbon cycles in both future climate projections and weather forecasts. Evaluation of models' ability to do this is performed in a wide range of simulation environments, often without explicit consideration of the degree of observational constraint or uncertainty and typically without quantification of benchmark performance expectations. We describe a Model Intercomparison Project (MIP) that attempts to resolve these shortcomings, comparing the surface turbulent heat flux predictions of around 20 different land models provided with in situ meteorological forcing evaluated with measured surface fluxes using quality-controlled data from 170 eddy-covariance-based flux tower sites. Predictions from seven out-of-sample empirical models are used to quantify the information available to land models in their forcing data and so the potential for land model performance improvement. Sites with unusual behaviour, complicated processes, poor data quality, or uncommon flux magnitude are more difficult to predict for both mechanistic and empirical models, providing a means of fairer assessment of land model performance. When examining observational uncertainty, model performance does not appear to improve in low-turbulence periods or with energy-balance-corrected flux tower data, and indeed some results raise questions about whether the energy balance correction process itself is appropriate. In all cases the results are broadly consistent, with simple out-of-sample empirical models, including linear regression, comfortably outperforming mechanistic land models. In all but two cases, latent heat flux and net ecosystem exchange of CO 2 are better predicted by land models than sensible heat flux, despite it seeming to have fewer physical controlling processes. Land models that are implemented in Earth system models also appear to perform notably better than stand-alone ecosystem (including demographic) models, at least in terms of the fluxes examined here. The approach we outline enables isolation of the locations and conditions under which model developers can know that a land model can improve, allowing information pathways and discrete parameterisations in models to be identified and targeted for future model development.

54 ENVIRONMENTAL SCIENCES↗

AMVOS: Additive Manufacturing Video Object Segmentation Dataset

This dataset provides labeled video frames from four additive manufacturing (AM) processes for video object segmentation (VOS) tasks. It contains 90 video segments comprising 900 individually annotated frames across five AM datasets: laser hot-wire directed energy deposition (LHW-DED), tungsten inert gas wire arc additive manufacturing (TIG-WAAM), plasma arc welding (PAW), visible-light polymer extrusion (visPolymer), and near-infrared polymer extrusion (irPolymer). Each video segment consists of 10 contiguous frames with corresponding pixel-level object instance annotations. Depending on the process, two of four object classes are labeled per frame: Melt Pool, Feed Wire, Nozzle, or Material. Raw frames are provided as .jpg files and annotations as palettized .png files. The dataset follows the directory structure of established VOS benchmarks (DAVIS, YouTube-VOS, MOSE), enabling direct integration into VOS model training and evaluation pipelines for foundation model fine-tuning, domain adaptation, or zero-shot performance benchmarking. Data was collected at Oak Ridge National Laboratory's Manufacturing Demonstration Facility.

Wetzel, Jon [ORNL]↗

Disentangling the gap between pure and mixed-gas performance of thin film composite membranes through improved cell design and testing methods

Testing thin film composite (TFC) membrane coupons at low stage-cuts (≤5%) in a sweep-gas permeation system is a common practice to obtain mixed-gas separation properties for benchmarking performance and making scale-up decisions. However, even under these idealized conditions, mixed-gas permeance and selectivity can be more than 30% lower than their pure-gas values, partially due to concentration polarization, an effect that typically intensifies with increased membrane permeance. This study investigates the effect of cell design on mixed-gas testing using PolyActive TM TFC membranes with pure-gas CO 2 permeance of 1700 – 3100 gas permeance unit (GPU), covering the permeance range of most state-of-the-art CO 2 /N 2 separation membranes. Here, we designed and 3D-printed a counter-current permeation cell with enhanced feed and sweep flow efficiency, resulting in a 33 – 41% increase in mixed-gas CO 2 permeance compared to traditional permeation cells. Furthermore, we compared sweep-gas and vacuum permeation methods using traditional permeation cells, revealing that the latter delivers 41% higher mixed-gas CO 2 permeance, because vacuuming effectively minimizes the downstream concentration polarization. These findings highlight the importance of cell design and permeation apparatus selection in lab-scale mixed-gas testing, with strong implications for module design and process optimization at the industrial scale.

mixed gas performance↗

A Data-Driven Framework for Automated Detection of Aircraft-Generated Signals in Seismic Array Data Using Machine Learning

Abstract Ground motions associated with aircraft overflights can cover a significant portion of the seismic data collected by shallowly emplaced seismometers, such as new nodal and Distributed Acoustic Sensing systems. This article describes the first published framework for automated detection of aircraft on single channel and multichannel seismic data. The seismic data are converted to spectrograms in a sliding time window and classified as aircraft or nonaircraft in each window using a deep convolutional neural network trained with analyst-labeled data. A majority voting scheme is used to convert the output from the sequence of sliding time windows onto a decision time sequence for each channel and to combine the binary classifications on the decision time sequences across multiple channels. Precision, recall, and F-score are used to quantify the detection performance of the algorithm on nodal data using fourfold time-series cross validation. By applying our framework to data from the Sage Brush Flats nodal array in Southern California, we provide a benchmark performance and demonstrate the advantage of using an array of sensors.

Geochemistry & Geophysics↗

Satellite solar-induced chlorophyll fluorescence and near-infrared reflectance capture complementary aspects of dryland vegetation productivity dynamics

Mounting evidence indicates dryland ecosystems play an important role in driving the interannual variability and trend of the terrestrial carbon sink. Nevertheless, our understanding of the seasonal dynamics of dryland ecosystem carbon uptake through photosynthesis [gross primary productivity (GPP)] remains relatively limited due in part to the limited availability of long-term data and unique challenges associated with satellite remote sensing across dryland ecosystems. Here, we comprehensively evaluated longstanding and emerging satellite vegetation proxies in their ability to capture seasonal dryland GPP dynamics. Specifically, we evaluated: 1) reflectance-based proxies normalized difference vegetation index (NDVI), soil adjusted vegetation index (SAVI), near infrared reflectance index (NIR v ), and kernel NDVI (kNDVI) from the MODerate resolution Imaging Spectroradiometer (MODIS); and 2) newly available physiologically-based proxy solar-induced chlorophyll fluorescence (SIF) from the TROPOspheric Monitoring Instrument (TROPOMI). As a performance benchmark, we used GPP estimates from a robust network of 21 western United States eddy covariance tower sites that span representative gradients in dryland ecosystem climate and functional composition. We found that NIR v and SIF were the best performing GPP proxies and captured complementary aspects of seasonal GPP dynamics across dryland ecosystem types. NIR v offered better performance than the other proxies across relatively low-productivity, sparsely non-evergreen vegetated sites (R 2 = 0.59 ± 0.13); whereas SIF best captured seasonal dynamics across relatively high-productivity sites, including evergreen-dominated sites (R 2 = 0.74 ± 0.07). Notably, across grass-dominated sites, all reflectance-based proxies (NDVI, SAVI, NIRv and kNDVI) showed significant seasonal bias (hysteresis) that strengthened with the total fraction of woody vegetation cover, likely due to seasonal patterns in woody vegetation reflectance that are unrelated to or decoupled from GPP. In conclusion, future efforts to fully integrate the complementary strengths of NIR v and SIF could significantly improve our understanding and representation of dryland GPP dynamics in satellite-based models.

54 ENVIRONMENTAL SCIENCES↗

BCSR on GPU: A Way Forward Extreme-scale Graph Processing on Accelerator-enabled Frontier Supercomputer

Handling large graphs in a distributed environment requires effective partitioning across processors and efficient management of local partitions. In 2D partitioning, local graphs often become too sparse, making memory-efficient data structures crucial. Using the Compressed Sparse Row (CSR) format wastes space, especially for > 83% of vertices with empty edges for the sparse graphs. This study explores bit-CSR (BCSR), a modified CSR representation, on GPUs to reduce memory usage in graph computations. We achieved 16.67% memory savings on a sparse rmat dataset with 268 million vertices and 357 million edges, without performance degradation, supported by both theoretical and experimental storage savings of 33%. However, we observed a 1.7× slowdown in degree lookup times due to bitwise operations on AMD CPUs. This analysis highlights the potential of BCSR on GPUs for improving Graph500 benchmark performance on GPU-accelerated systems, such as the Frontier supercomputer.

Sattar, Naw Safrin↗

ExaTN: Scalable GPU-Accelerated High-Performance Processing of General Tensor Networks at Exascale

We present ExaTN (Exascale Tensor Networks), a scalable GPU-accelerated C++ library which can express and process tensor networks on shared- as well as distributed-memory high-performance computing platforms, including those equipped with GPU accelerators. Specifically, ExaTN provides the ability to build, transform, and numerically evaluate tensor networks with arbitrary graph structures and complexity. It also provides algorithmic primitives for the optimization of tensor factors inside a given tensor network in order to find an extremum of a chosen tensor network functional, which is one of the key numerical procedures in quantum many-body theory and quantum-inspired machine learning. Numerical primitives exposed by ExaTN provide the foundation for composing rather complex tensor network algorithms. We enumerate multiple application domains which can benefit from the capabilities of our library, including condensed matter physics, quantum chemistry, quantum circuit simulations, as well as quantum and classical machine learning, for some of which we provide preliminary demonstrations and performance benchmarks just to emphasize a broad utility of our library.

97 MATHEMATICS AND COMPUTING↗

Benchmarking Optimizers for Qumode State Preparation with Variational Quantum Algorithms

Quantum state preparation involves preparing a target state from an initial system, a process integral to applications such as quantum machine learning and solving systems of linear equations. Recently, there has been a growing interest in qumodes due to advancements in the field and their potential applications. However there is a notable gap in the literature specifically addressing this area. This paper aims to bridge this gap by providing performance benchmarks of various optimizers used in state preparation with Variational Quantum Algorithms. We conducted extensive testing across multiple scenarios, including different target states, both ideal and sampling simulations, and varying numbers of basis gate layers. Our evaluations offer insights into the complexity of learning each type of target state and demonstrate that some optimizers perform better than others in this context. Notably, the Powell optimizer was found to be exceptionally robust against sampling errors, making it a preferred choice in scenarios prone to such inaccuracies. Additionally, the Simultaneous Perturbation Stochastic Approximation optimizer was distinguished for its efficiency and ability to handle increased parameter dimensionality effectively.

Kan, Shuwen [Fordham University]↗