NAS Grid Benchmarks: A Tool for Measuring Performance of Computational Grids
This viewgraph presentation includes a brief history of benchmarking computational grids at NASA Ames, measuring grid performance and preliminary results of recent studies.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
This viewgraph presentation includes a brief history of benchmarking computational grids at NASA Ames, measuring grid performance and preliminary results of recent studies.
SAND2024-08539O The High Performance GMRES Mixed-Precision (HPG-MxP) is a benchmark for ranking high-performance supercomputers, allowing use of mixed-precision. Similar to HPCG benchmark, it is designed to profile the computers' capabilities to perform the computational and communication tasks that are commonly found in important classes of real-world applications. At the same time, like HPL-MxP benchmark, it allows the use of mixed-precision arithmetic, while ensuring the double-precision accuracy of the computed solution. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.
The Institute for the Design of Advanced Energy Systems (IDAES) Integrated Platform is a versatile computational environment offering extensive process systems engineering (PSE) capabilities for optimizing the design and operation of complex, interacting technologies and systems. IDAES enables users to efficiently search vast, complex design spaces to discover the lowest cost, most environmentally sustainable solutions while supporting the full process modeling lifecycle, from conceptual design to dynamic optimization and control. The extensible, open platform empowers users to create models of novel processes and rapidly develop custom analyses, workflows, and end-user applications. IDAES-PSE 2.0.0 Release Highlights Removal of deprecated features from IDAES v1 Update to Pyomo v6.5 – this required a number of updates to support the new NL solver writer and to address some changes in Pyomo Creation of new testing suite for backward compatibility, model robustness and verification More general implementation of the Helmholtz EoS. This brings some new features like standard property diagrams, choice of mass or mole basis, and new state variable options Standardizing names in Heat Exchanger models (breaking change from v2.0.0a2): Control Volumes named hot_side and cold_side Ports names hot_side_inlet, hot_side_outlet, cold_side_inlet and cold_side_outlet Config Blocks names hot_side_config and cold_side_config Config arguments for user provided names for each side: hot_side_name and cold_side_name. Updating Keras surrogate tool to use v1.1 of OMLT New prototype API for model initialization (idaes.core.initialization) The new API uses "Model Initializer" objects instead of class methods, allowing for the definition of multiple initialization routines for a single model A number of common, model agnostic initialization routines have also been defined, including initialization from data, block-decomposition and a general hierarchical approach equivalent to the existing method for common unit models New metadata for thermophysical properties – valid_range This can be used to record the range of values over which a property value can be trusted, such as the range of experimental data used to regress parameters A number of new utility functions have been added to check for properties with values outside the valid range and to set bounds based on this metadata Updated construction of balance expressions in Control Volumes to remove unneeded terms In the past, unneeded terms were added as a constant 0 term, however they will now be dropped entirely from the expression This was necessary due to more strict unit checking in the new Pyomo solver writer which no longer ignores 0 terms Updates to metadata for thermophysical properties to better define known properties and units of measurement This results in more strict enforcement of standard naming for thermophysical and reaction properties Users can still define custom properties, but these must be done explicitly using the define_custom_properties() method instead of being implicitly created by add_property() Updated convergence tester utility tool to support definition of benchmark files (JSON format) and comparison of performance to benchmarks Set default iteration limit for IPOPT in IDAES config to 200 iterations Update scaling of example models to work with new Pyomo NL solver writer Improve testing of extensions and examples infrastructure to avoid need for downloading files Updated distillation column to centralize common functionality and remove a number of Pyomo warnings
Containers have taken over large swaths of cloud computing as the most convenient way of packaging and deploying applications. The features that containers offer for packaging and deploying applications translate to high performance computing (HPC) as well. At The National Oceanic and Atmospheric Administration, containers provide an easy way to build and distribute complex HPC applications, allowing faster collaboration, portability, and experiment computer environment reproducibility amongst the scientific community. The challenge arises when applications rely on message passing interface (MPI). This necessitates investigation into how to properly run these applications with their own unique requirements and produce performance on par with native runs. We investigate the MPI performance for benchmarks and containerized climate models for various containers covering selection of compiler and MPI library combinations from the Cray provided programming environments on the Cray XC supercomputer GAEA. Performance from the benchmarks and the climate models shows that for the most part containerized applications perform on par with the natively built applications when the system optimized Cray MPICH libraries are bound into the container, and the hybrid model containers have poor performance in comparison. We also describe several challenges and our solutions in running these containers, particularly challenges with heterogeneous jobs for the containerized model runs.
The optical degradation of encapsulants from ultraviolet (UV) radiation has historically resulted in a significant loss in performance throughout the life of a photovoltaic (PV) module. International Electrotechnical Commission (IEC) test methods have recently been developed to screen for PV encapsulants prone to loss in optical performance. The present study was performed to benchmark polymeric packaging materials relative to IEC 62788-1-4 (covering the measurement of optical transmittance) and IEC 62788-1-7 (on the durability of transmittance), provide feedback toward improvement of the methods, and develop insight regarding optical degradation. Contemporary materials were examined, including poly(ethylene-co-vinyl acetate) (EVA), thermoplastic polyolefin (TPO), polyolefin elastomer (POE), and polyvinyl butyral (PVB) encapsulants; a poly(ethene-co-tetrafluoroethene)/poly(ethylene terephthalate) (ETFE/PET) transparent backsheet; and a polystyrene (PS) working reference material. The use of silica-, specialty-, and rolled-glass was also compared in laminated coupons. Specimen size was separately examined from 2.5 to 12.5 cm. Weathering was performed with a xenon source, using IEC TS 62788-7-2 methods A2, A3, A4, and A5 (chamber temperature of 55 degrees C, 65 degrees C, 75 degrees C, or 85 degrees C), respectively. Characterizations were made using a UV-visible-near-infrared (UV-VIS-NIR) spectrophotometer (transmittance and reflectance, with and without an integrating sphere), a UV-VIS fluorescence spectrophotometer, a camera, and an optical microscope. Performance was analyzed, including solar weighted transmittance, yellowness index, UV cut-off wavelength, and haze (scattering). Separate Arrhenius analyses were performed to assess retention of transmittance and changes in yellowness index. The activation energy for both characteristics was found to range from 15-80 kJ mol-1, with an average of 48 kJ mol-1, similar to the average of 45 kJ mol-1 identified in the previous international PV Quality Assurance Task Force (PVQAT) Task Group 5 (TG5) study of more traditional encapsulants. The separate degradation modes of discoloration and scattering were distinguished in the encapsulants using a comprehensive spectral characterization. Based on these results, the IEC 62788-1-7 pass/fail criteria of 5% change in transmittance was confirmed to identify a known bad encapsulant.
Inverter-based resources (IBRs) such as photovoltaics (PVs), wind turbines, and battery energy storage systems (BESSs) are widely deployed in low-carbon power systems. However, these resources typically do not provide the inertia needed for grid stability, resulting in a low-inertia power system. IBRs and lack of inertia have been known to cause anomalies such as waveform distortions and wideband oscillations in power systems due to the limited inertia level, leading to increased generation trips and load shedding. Here, to achieve effective anomaly identification, this paper proposes a synchro-waveform-based algorithm utilizing real-time synchronized voltage waveform measurements from waveform measurement units (WMUs). In the proposed method, different physical characteristics, as well as statistical features, are extracted from synchronized voltage waveform measurements to filter anomalies. Then, the anomaly identification approach based on the random forest is developed and deployed into the FNET/GridEye system considering trade-offs among accuracy, computational burden, and deployment cost. Moreover, four WMUs are specially designed and deployed on Kauai Island to receive instantaneous synchronized voltage waveform measurements. To verify the performance of the proposed algorithm, different experiments are carried out with collected field test data. The result demonstrates that the performance of the proposed synchro-waveform-based anomaly categorization algorithm can accurately identify anomalies 95.35% of the time, which has comparable performance among benchmarking algorithms.
This paper provides results from steady state HTR-10 benchmark calculations performed using AGREE and Serpent that are compared to experimental results as well as calculations performed by INET. The purpose of completing this benchmark is to validate AGREE and Serpent for the prediction of HTGR operation so that they may ultimately be used to support licensing and deployment efforts of advanced reactors. The benchmark consists of several problems ranging from control rod worth calculations to k{sub eff} calculations at various temperatures. Overall, AGREE and Serpent show good agreement with the reference solutions and are effectively able to predict HTGR operation for a variety of steady state cases. The largest difference from the reference was for the initial core single rod worth, which is possibly due to the larger core helium cavity causing inaccuracies in the diffusion calculation whereas the control rod worth determined by Serpent and AGREE is much closer to the reference result for the full core loading. Moreover, in every benchmark problem, increasing the number of energy groups in the cross sections results in improved agreement of the AGREE result with the reference, although it also causes an increase in computation time. For problems with experimental data available, the accuracy of results generated by Serpent and AGREE is comparable to results obtained by other benchmark participants. (authors)
This report documents a study performed to investigate the requirements for criticality safety benchmark experiments for high-assay, low-enriched uranium (HALEU) fuel in transportation applications. In this work, an exploratory application model, the “Pebble Tanker,” was developed to represent TRISO fuel in a transportation scenario for an analysis of the validation basis in industrial quantities. An aspect of the criticality validation process involves assessing the “similarity” between application and experimental benchmark systems through an integral index parameter evaluation. Here, this includes propagating nuclear data uncertainties and calculating a correlation coefficient (hereinafter referred to as “c k ”) to evaluate the similarity of benchmark experiments compared with the application Pebble Tanker model. Finding sufficient critical benchmark experiments allows for the evaluation of bias and bias uncertainty, thus determining the upper subcritical limit (USL) of the transportation package. A target k eff of ~0.94 was used in this work to establish appropriate modeling conditions, reflecting a reasonable estimate for a USL. Two container models were investigated: one with the Hermes-type pebble and one with the Pebble Bed Modular Reactor (PBMR)–type pebble. The models were simplified, considering only fuel, containment structure, and either water or air. This allows a focus on the underlying physics of applications involving TRISO fuel pebbles using the Pebble Tanker model. A crucial consideration is the transport package's ability to safely hold pebbles while flooded, maintaining subcritical conditions. Tools available in the SCALE 6.3.1 suite—the CSAS6-Shift, TSUNAMI-3D-Shift, and TSUNAMI-IP sequences—were employed for neutronics and sensitivity and uncertainty (S/U) analysis of the Pebble Tanker. Findings demonstrated sufficient available critical experiment benchmarks to perform a validation of the Pebble Tanker in the most reactive state, i.e., when the Tanker is flooded.
Understandinrelationships between magnetic field and plasma conditions
Abstract not provided.
Abstract not provided.
Summary Parallel I/O is a critical technique for moving data between compute and storage subsystems of supercomputers. With massive amounts of data produced or consumed by compute nodes, high‐performant parallel I/O is essential. I/O benchmarks play an important role in this process; however, there is a scarcity of I/O benchmarks representative of current workloads on HPC systems. Toward creating representative I/O kernels from real‐world applications, we have created h5bench , a set of I/O kernels that exercise hierarchical data format version 5 (HDF5) I/O on parallel file systems in numerous dimensions. Our focus on HDF5 is due to the parallel I/O library's heavy usage in various scientific applications running on supercomputing systems. The various tests benchmarked in the h5bench suite include I/O operations (read and write), data locality (arrays of basic data types and arrays of structures), array dimensionality (one‐dimensional arrays, two‐dimensional meshes, three‐dimensional cubes), I/O modes (synchronous and asynchronous). In this paper, we present the observed performance of h5bench executed along several of these dimensions on existing supercomputers (Cori and Summit) and pre‐exascale platforms (Perlmutter, Theta, and Polaris). h5bench measurements can be used to identify performance bottlenecks and their root causes and evaluate I/O optimizations. As the I/O patterns of h5bench are diverse and capture the I/O behaviors of various HPC applications, this study will be helpful to the broader supercomputing and I/O community.
Abstract The IEA PVPS Task 13 group, experts who focus on photovoltaic performance, operation, and reliability from several leading R&D centers, universities, and industrial companies, is developing a framework for the calculation of performance loss rates of a large number of commercial and research photovoltaic (PV) power plants and their related weather data coming across various climatic zones. The general steps to calculate the performance loss rate are (i) input data cleaning and grading; (ii) data filtering; (iii) performance metric selection, corrections, and aggregation; and finally, (iv) application of a statistical modeling method to determine the performance loss rate value. In this study, several high‐quality power and irradiance datasets have been shared, and the participants of the study were asked to calculate the performance loss rate of each individual system using their preferred methodologies. The data are used for benchmarking activities and to define capabilities and uncertainties of all the various methods. The combination of data filtering, metrics (performance ratio or power based), and statistical modeling methods are benchmarked in terms of (i) their deviation from the average value and (ii) their uncertainty, standard error, and confidence intervals. It was observed that careful data filtering is an essential foundation for reliable performance loss rate calculations. Furthermore, the selection of the calculation steps filter/metric/statistical method is highly dependent on one another, and the steps should not be assessed individually.
Quantum diamond microscope (QDM) magnetic field imaging is an emerging interrogation and diagnostic technique for integrated circuits (ICs). To date, the ICs measured with a QDM have been either too complex for us to predict the expected magnetic fields and benchmark the QDM performance or too simple to be relevant to the IC community. In this paper, we establish a 555 timer IC as a “model system” to optimize QDM measurement implementation, benchmark performance, and assess IC device functionality. To validate the magnetic field images taken with a QDM, we use a spice electronic circuit simulator and finite-element analysis (FEA) to model the magnetic fields from the 555 die for two functional states. Furthermore, we compare the advantages and the results of three IC-diamond measurement methods, confirm that the measured and simulated magnetic images are consistent, identify the magnetic signatures of current paths within the device, and discuss using this model system to advance QDM magnetic imaging as an IC diagnostic tool.
This talk will describe the NAS (National Aerospace Standards) Parallel Benchmarks, which are now widely cited in the high performance computing field as a measure of sustained performance on realistic scientific applications. The latest performance results will be included. It will be shown that significant progress has been made by several systems during the past year or so, with sustained performance on a par with the best conventional systems, and with performance per dollar significantly exceeding the conventional systems. This talk will also describe many of the pitfalls of performance reporting, and will give advice on how to avoid such pitfalls. The overall state of the field of high performance computing will also be discussed.
There are two common ways to evaluate algorithms: performance on benchmark problems derived from real applications and analysis of performance on parametrized families of problems. The two approaches complement each other, each having its advantages and disadvantages. The planning community has concentrated on the first approach, with few ways of generating parametrized families of hard problems known prior to this work. Our group's main interest is in comparing approaches to solving planning problems using a novel type of computational device - a quantum annealer - to existing state-of-the-art planning algorithms. Because only small-scale quantum annealers are available, we must compare on small problem sizes. Small problems are primarily useful for comparison only if they are instances of parametrized families of problems for which scaling analysis can be done. In this technical report, we discuss our approach to the generation of hard planning problems from classes of well-studied NP-complete problems that map naturally to planning problems or to aspects of planning problems that many practical planning problems share. These problem classes exhibit a phase transition between easy-to-solve and easy-to-show-unsolvable planning problems. The parametrized families of hard planning problems lie at the phase transition. The exponential scaling of hardness with problem size is apparent in these families even at very small problem sizes, thus enabling us to characterize even very small problems as hard. The families we developed will prove generally useful to the planning community in analyzing the performance of planning algorithms, providing a complementary approach to existing evaluation methods. We illustrate the hardness of these problems and their scaling with results on four state-of-the-art planners, observing significant differences between these planners on these problem families. Finally, we describe two general, and quite different, mappings of planning problems to QUBOs, the form of input required for a quantum annealing machine such as the D-Wave II.
The fluoride-salt-cooled high-temperature reactor (FHR) is one of the advanced reactors that has been attracting considerable interest from both the research community and the nuclear industry. To help facilitate the nuclear community's familiarity with the FHR, Kairos Power has developed a generic FHR (gFHR) benchmark. In the research performed here, this benchmark was used to assess innovative modeling methods that combine stochastic and deterministic computer codes to perform the design and analysis of the gFHR. Further, the Monte Carlo code Serpent 2 was used to generate few-group cross sections that were then used in the neutron diffusion and thermal-fluids code AGREE to perform full-core neutronics and thermal-fluids steady-state and transient core analysis. The Argonne National Laboratory code SAM was then used to model the gFHR system and to simulate the load-follow operation of the gFHR.
Large Language Models (LLMs) have propelled groundbreaking advancements across several domains and are commonly used for text generation applications. However, the computational demands of these complex models pose significant challenges, requiring efficient hardware acceleration. Benchmarking the performance of LLMs across diverse hardware platforms is crucial to understanding their scalability and throughput characteristics. We introduce LLM-Inference-Bench, a comprehensive benchmarking suite to evaluate the hardware inference performance of LLMs. We thoroughly analyze diverse hardware platforms, including GPUs from Nvidia and AMD and specialized AI accelerators, Intel Habana and SambaNova. Our evaluation includes several LLM inference frameworks and models from LLaMA, Mistral, and Qwen families with 7B and 70B parameters. Our benchmarking results reveal the strengths and limitations of various models, hardware platforms, and inference frameworks. We provide an interactive dashboard to help identify configurations for optimal performance for a given hardware platform.