Benchmarking the performance of quantum computing software for quantum circuit creation, manipulation and compilation
Not Available
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Not Available
There are several different calculation approaches and tools that can be used to evaluate the risk of hydrogen energy applications. A comparative study of Air Liquide’s ALDEA (Air Liquide Dispersion and Explosion Assessment) tools suite and Sandia’s HyRAM (Hydrogen Risk Assessment Models) toolkit has been conducted. The purpose of this study was to understand and evaluate the differences between the two calculation approaches, and identify areas for model improvements. There were several scenarios examined in this effort regarding hydrogen release dynamics. These scenarios include free jet release cases at varying pressures, vessel blowdown, and hydrogen build-up scenarios with and without ventilation. For each scenario, the input and output of the HyRAM calculations are documented, along with a comparison to the ALDEA results. Generally, the results from the two different tools were reasonably aligned. However, there were fundamental differences in evaluation methodology and functional limitations in HyRAM that caused discrepancies in some calculations.
In order to characterize and benchmark computational hardware, software, and algorithms, it is essential to have many problem instances on-hand. This is no less true for quantum computation, where a large collection of real-world problem instances would allow for benchmarking studies that in turn help to improve both algorithms and hardware designs. To this end, here we present a large dataset of qubit-based quantum Hamiltonians. The dataset, called HamLib (for Hamiltonian Library), is freely available online and contains problem sizes ranging from 2 to 1000 qubits. HamLib includes problem instances of the Heisenberg model, Fermi-Hubbard model, Bose-Hubbard model, molecular electronic structure, molecular vibrational structure, MaxCut, Max- k -SAT, Max- k -Cut, QMaxCut, and the traveling salesperson problem. The goals of this effort are (a) to save researchers time by eliminating the need to prepare problem instances and map them to qubit representations, (b) to allow for more thorough tests of new algorithms and hardware, and (c) to allow for reproducibility and standardization across research studies.
The field of High-Performance Computing (HPC) is defined by providing computing devices with highest performance for a variety of demanding scientific users. The tight co-design relationship between HPC providers and users propels the field forward, paired with technological improvements, achieving continuously higher performance and resource utilization. A key device for system architects, architecture researchers, and scientific users are benchmarks, allowing for well-defined assessment of hardware, software, and algorithms. Many benchmarks exist in the community, from individual niche benchmarks testing specific features, to large-scale benchmark suites for whole procurements. We survey the available HPC benchmarks, summarizing them in table form with key details and concise categorization, also through an interactive website. For categorization, we present a benchmark taxonomy for well-defined characterization of benchmarks.
The rapid pace of development in quantum computing technology has sparked a proliferation of benchmarks to assess the performance of quantum computing hardware and software. However, not all benchmarks are of equal merit. Good ones empower scientists, engineers, programmers and users to understand the power of a computing system, whereas bad ones can misdirect research and inhibit progress. In this Perspective, we survey the science of quantum computer benchmarking. Here, we discuss the role of benchmarks and benchmarking and how good benchmarks can drive and measure progress towards the long-term goal of useful quantum computations, known as quantum utility. We explain how different kinds of benchmark quantify the performance of different parts of a quantum computer, discuss existing benchmarks, examine recent trends in benchmarking, and highlight important open research questions in this field.
In this era of large and complex astronomical survey data, interpreting, validating, and comparing inference techniques becomes increasingly difficult. This is particularly critical for emerging inference methods like Simulation-Based Inference (SBI), which offer significant speedup potential and posterior modeling flexibility, especially when deep learning is incorporated. We present a study to assess and compare the performance and uncertainty prediction capability of Bayesian inference algorithms – from traditional MCMC sampling of analytic functions to deep learning-enabled SBI. We focus on testing the capacity of hierarchical inference modeling in those scenarios. Before we extend this study to cosmology, we first use astrophysical simulation data to ensure interpretability. We demonstrate a probabilistic programming implementation of hierarchical and non-hierarchical Bayesian inference using simulations derived from the DeepBench software library, a benchmarking tool developed by our group that generates simple and controllable astrophysical objects from first principles. This study will enable astronomers and physicists to harness the inference potential of these methods with confidence.
This presentation discusses the benchmark of the IER 498: Godiva-IV criticality accident alarm system (CAAS). It details the goal of the project which is to create a CAAS benchmark capability for Godiva-IV in order to create a shielding benchmark for NCSP software and data, to add geometrically diverse benchmark/benchmarks to the ICSBEP Handbook, to enable more precise measurements at NCERC, and to help assure CAAS performance is as-stated. It also states that work was completed on CED-1 in 2019 and that future work includes the completion of CED-2 in 2020.
We created a set of standards to efficiently evaluate AI benchmarks through objective means. Although prevalent, especially in recent times, AI benchmarks have no single way to measure their effectiveness. The MLCommons team provided a set of criteria for evaluating benchmarks, although the criteria lacks a clearly defined set of evaluation rules. We created a rubric with preset factors to efficiently and objectively evaluate a benchmark s quality. We created a software framework for processing lists of benchmarks for visualization. The framework and rating system allows researchers to quickly check if their benchmarks are effective.
Metagenome binning is a key step, downstream of metagenome assembly, to group scaffolds by their genome of origin. Although accurate binning has been achieved on datasets containing multiple samples from the same community, the completeness of binning is often low in datasets with a small number of samples due to a lack of robust species co-abundance information. In this study, we exploited the chromatin conformation information obtained from Hi-C sequencing and developed a new reference-independent algorithm, Metagenome Binning with Abundance and Tetra-nucleotide frequencies—Long Range (metaBAT-LR), to improve the binning completeness of these datasets. This self-supervised algorithm builds a model from a set of high-quality genome bins to predict scaffold pairs that are likely to be derived from the same genome. Then, it applies these predictions to merge incomplete genome bins, as well as recruit unbinned scaffolds. We validated metaBAT-LR’s ability to bin-merge and recruit scaffolds on both synthetic and real-world metagenome datasets of varying complexity. Benchmarking against similar software tools suggests that metaBAT-LR uncovers unique bins that were missed by all other methods.
We created a set of standards to efficiently evaluate AI benchmarks through objective means. Although prevalent, especially in recent times, AI benchmarks have no single way to measure their effectiveness. The MLCommons team provided a set of criteria for evaluating benchmarks, although the criteria lacks a clearly defined set of evaluation rules. We created a rubric with preset factors to efficiently and objectively evaluate a benchmark’s quality. We created a software framework for processing lists of benchmarks for visualization. The framework and rating system allows researchers to quickly check if their benchmarks are effective.
Confidently designing safe, new nuclear criticality experiments requires expert judgement, which could take years of experience. Sensitivity/uncertainty (S/U) analysis can be utilized by less experienced individuals to conservatively estimate uncertainties in important parameters, such as k eff , in newly proposed nuclear applications. This type of analysis relies on matching new nuclear applications with existing benchmark experiments. The Whisper-1.1 software package included in MCNP6.2 ®1 contains more than 1,100 International Criticality Safety Benchmark Evaluation Project (ICSBEP) benchmarks. These benchmarks however rarely match new nuclear applications. The number of benchmarks available to match a given set of materials or geometric configurations varies significantly. Furthermore, recently performed benchmark experiments may not have had enough time to be properly documented and published. Benchmarks are vital for determining the accuracy of nuclear data and can assist nuclear physics and evaluators in improving nuclear data libraries. Exhaustively exploring the parameter space using simulations with software such as MCNP is too computationally expensive. In this work, Gaussian process optimization was implemented to reduce the number of simulations needed for optimization over multiple parameters. This optimization scheme was designed to selectively generate new benchmarks with high sensitivity-based similarity metrics to user-defined nuclear applications. Two benchmark models of spherically nested shells containing plutonium, uranium, tantalum, and water were used in the optimization to match an application containing plutonium plates, stacked in a tantalum reflector, surrounded by water.
While confidence in photovoltaic (PV) modeling software has always been essential, the rapid pace of new PV plant developments makes accuracy and credibility more critical than ever. Independent assessments, particularly through blind modeling comparisons, are therefore necessary to ensure unbiased benchmarking across PV modeling software. Previous studies have been limited by a narrow range of models compared, anonymized results, or system size. This study presents results from the first-ever onymous blind modeling comparison, evaluated using both lab- and utility-scale fixed-tilt, monofacial, south-facing systems at sub-hourly time intervals. Seven commercially used PV software tools were compared: 3E SynaptiQ, PlantPredict, PVsyst, RatedPower, SAM, SolarFarmer, and Solargis Evaluate. Predictions were submitted directly by software representatives, providing unique insights into each software’s implementation and resulting prediction behavior. Notable features, including plane-of-array (POA) transposition model, module temperature model, shading model, and performance model were analyzed and compared. Four summary tables compile these features of the software, serving as a resource to help users understand the methodological differences and select the most suitable software for their applications. The software tools show deviations from mean error in annual yield up to 2.5 % in the lab-scale system, increasing to 6.0 % for the utility-scale system. These differences arise from a combination of user decisions and the inherent behavior of the software, indicating the need for continuous and rigorous validation of modeling methods using these software tools against complex, real-world systems.
In this work, we present a DFT-based, QM/MM implementation with long-range electrostatic embedding achieved by direct real-space integration of the particle mesh Ewald (PME) computed electrostatic potential. The key transformation is the interpolation of the electrostatic potential from the PME grid to the DFT quadrature grid, from which integrals are easily evaluated utilizing standard DFT machinery. We provide benchmarks of the numerical accuracy with choice of grid size and real-space corrections, and demonstrate that good convergence is achieved while introducing nominal computational overhead. Furthermore, the approach requires only small modification to existing software packages, as is demonstrated with our implementation in the OpenMM and Psi 4 software. After presenting convergence benchmarks, we evaluate the importance of long-range electrostatic embedding in three solute/solvent systems modeled with QM/MM. Water and BMIM/BF 4 ionic liquid were considered as "simple" and "complex" solvents respectively, with water and p-phenylenediamine (PPD) solute molecules treated at QM level of theory. While electrostatic embedding with standard real-space truncation may introduce negligible error for simple systems such as water solute in water solvent, errors become more significant when QM/MM is applied to complex solvents such as ionic liquids. An extreme example is the electrostatic embedding energy for oxidized PPD in BMIM/BF 4 for which real-space truncation produces severe error even at 2-3 nm cutoff distances. This latter example illustrates that utilization of QM/MM to compute redox potentials within concentrated electrolytes/ionic media requires carefully chosen long-range electrostatic embedding algorithms, with our presented algorithm providing a general and robust approach.
The EQ_phase_detection software is designed to scan continuous daily waveforms to detect earthquake phase arrivals from local to regional (150 km) events. The detections are made with a deep learning encoder-decoder model. When the model detects an earthquake in the waveforms, a second model is implemented to classify the first arriving motions. Both deep learning models are trained with the Tensorflow package using publicly available benchmark data sets. The software input is a path to a directory that contains waveforms in mseed format and the associated response files in xml format. The output is a data table of time stamped detections, signal amplitude, signal-to-noise ratio, and softmax probability of the detection in a generic format applicable to post-processing association algorithms for event locations. Additionally, the p-wave and s-wave waveforms are saved in a data table for rapid access when producing improved locations using correlation-based techniques. The software is designed for multiprocessing with multiple GPU’s for rapid processing of large data sets. The configuration file provides flexibility in the trained models implemented and allows access to multiple models trained for different sampling rates or input dimensions. This is particularly useful for regions with multiple networks that do not have the same data parameters.
Resonant ultrasound spectroscopy (RUS) is an efficient, nondestructive technique to study the elastic properties of solids. A low-temperature (2-300 K) probe has been assembled and tested at the NOMAD beamline at the Spallation Neutron Source at Oak Ridge National Laboratory to assess the probe's neutronic properties, data acquisition system, and compatibility with existing sample environment. A case study on a bulk metallic glass, La 65 Cu 20 Al 10 Co 5, served to benchmark both hardware and software developments. The elastic constants of the metallic glass were determined as a function of temperature between 4 and 300 K and were used to guide neutron diffraction measurements at NOMAD. Tracking of a specific RUS peak center frequency and width enabled live monitoring of the sample temperature evolution that lagged thermometry by upwards of 40 K near room temperature. Assembly of a high-temperature (300-875 K) probe is underway and both probes are scheduled to be available to users by late-2021. Our aim is to provide users with live monitoring of an intrinsic variable at the neutron scattering beamlines, in addition to existing controls, to monitor the state of their samples and its elastic moduli and make informed decisions in real time.
During the first phase of the Consortium for Advanced Simulation of Light Water Reactors (CASL) program, the Virtual Environment for Reactor Applications (VERA) was developed with a focus on capabilities for high-fidelity, multiphysics simulation of pressurized water reactors (PWRs). During this development effort, a set of progression problems was created ranging from smaller pin cell calculations to larger 3D full-core calculations. These progression problems helped to guide the development of the software and served as benchmarks against which to test VERA. Since 2019, efforts have been made to extend the capabilities of VERA to model boiling water reactors (BWRs). Because BWR simulations come with many unique challenges, a set of BWR progression problems was developed to aid in this new effort. The BWR progression problems range from 2D lattice calculations to 3D mini-core problems, and reference neutronic solutions were computed using continuous-energy Monte Carlo codes. MPACT, one of the neutronics code in VERA, was benchmarked using the BWR progression problems. The code is capable of computing solutions to all problems. The eigenvalues computed by MPACT agree well with the Monte Carlo reference solutions. (authors)
Many advanced reactor concept designs rely on high-assay low-enriched uranium (HALEU) fuel, enriched up to approximately 19.75% 235 U by weight. Efforts are underway by the US government to increase HALEU production in the United States to meet anticipated needs. However, very few data exist for validation of computational models that include HALEU, beyond a few fresh fuel benchmark specifications in the International Reactor Physics Experiment Evaluation Project. Nevertheless, there are other data with potential value available for developing into quality benchmarks for use in data- and software-validation efforts. This paper reviews the available evaluated HALEU fuel benchmarks and some of the potentially relevant benchmarks for fresh highly enriched uranium. It then introduces experimental data for HALEU fuel irradiated at Idaho National Laboratory, from relatively recent irradiation programs at the Advanced Test Reactor. Such data should be evaluated and, if valuable, collected into detailed benchmark specifications to meet the needs of HALEU-based reactor designers.
It is challenging to quantitatively predict shearing of intersecting fractures/faults because of dynamic frictional contacts accompanied by possible nonlinear rock deformation. To address such challenges, a new conceptual model—the simplified DFN model—was proposed and validated by Hu et al. 46 to use major paths (MPs) to represent complicated DFNs for calculation of shearing. In this work, we conducted a benchmark study for three examples that involve different levels of complexity of intersecting fractures, and correspondingly different numbers of MPs. The codes and software that were used in the benchmark cover a range of continuum, discontinuum and hybrid numerical methods: NMM (LBNL), FLAC3D (LBNL), GBDEM (KIGAM), FRACOD (DynaFrax), and CASRock (CAS). The general consistency between DFN and MP cases as predicted by all the codes/software demonstrates that major paths can be used to simplify the geometry of DFNs in a wide range of software. Disagreement in results made by some software and potential future improvements are discussed. We show that (1) shearing of one or multiple major fractures can be reduced if there are multiple smaller intersecting fractures in that area, which is a useful basis for understanding and controlling induced seismicity and merits further analysis, and (2) the agreement achieved in the benchmark examples provide confidence that the simplified DFN model is a promising conceptual model that can be used for different types of numerical approaches and software for simplifying the analysis of the shearing of intersecting fractures and faults.