Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “computer system benchmarking”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Systematic Crosstalk Mitigation for Superconducting Qubits via Frequency-Aware Compilation

One of the key challenges in current Noisy Intermediate-Scale Quantum (NISQ) computers is to control a quantum system with high-fidelity quantum gates. There are many reasons a quantum gate can go wrong - for superconducting transmon qubits in particular, one major source of gate error is the unwanted crosstalk between neighboring qubits due to a phenomenon called frequency crowding. We motivate a systematic approach for understanding and mitigating the crosstalk noise when executing near-term quantum programs on superconducting NISQ computers. Here, we present a general software solution to alleviate frequency crowding by systematically tuning qubit frequencies according to input programs, trading parallelism for higher gate fidelity when necessary. The net result is that our work dramatically improves the crosstalk resilience of tunable-qubit, fixed-coupler hardware, matching or surpassing other more complex architectural designs such as tunable-coupler systems. On NISQ benchmarks, we improve worst-case program success rate by 13.3x on average, compared to existing traditional serialization strategies.

Computer architecture↗

Elevating SolTrace's Capabilities for the Next Generation of Concentrating Solar Analysis

SolTrace is an open-source Monte Carlo ray tracing software developed at NREL. SolTrace can characterize concentrating solar thermal (CST) collector optical performance and is CST technology agnostic. Shown in Fig. 1, SolTrace is a foundational tool in NREL's CST system and component modeling suite. SolTrace's generic surface elements can flexibly model novel collector and receiver designs to predict spatial and temporal flux distributions - critical to understand for CST component design, performance prediction, and system integration. Since its initial development, SolTrace has over 1,650 references on Google Scholar, over 9,800 downloads since 2017, and has served the CST research and development community as a benchmark of 3rd party verification. SolTrace provides users with many options for defining surface shape and boundaries. However, SolTrace provides limited documentation which can result in a steep learning curve for new users. Additionally, SolTrace lacks the computational performance required to evaluate optical performance of a CST system over the course of a year and/or iteratively over design parameters in a timely manner. To address this, we are working towards a new release of SolTrace that enables increased computational throughput by implementing ray tracing acceleration structures and enabling GPU parallelization. Additionally, we are working to improve SolTrace's usability, accessibility, and maintainability by (1) automating solar position time-dependent simulation processes, (2) creating general CST collector templates of grouped elements, (3) updating the user interface to better visualize model inputs and outputs, and (4) creating a user support network through forums, "how to" videos, and documentation.

14 SOLAR ENERGY↗

On the Computational Viability of Quantum Optimization for PMU Placement

Using optimal phasor measurement unit placement as a prototypical problem, we assess the computational viability of the current generation D-Wave Systems 2000Q quantum annealer for power systems design problems. We reformulate minimum dominating set for the annealer hardware, solve the reformulation for a standard set of IEEE test systems, and benchmark solution quality and time to solution against the CPLEX optimizer and simulated annealing. For some problem instances the 2000Q outpaces CPLEX. For instances where the 2000Q underperforms with respect to CPLEX and simulated annealing, we suggest hardware improvements for the next generation of quantum annealers.

hardware↗

3D CFD Model Validation Using Benchmark Data of 1/16th Scaled VHTR Upper Plenum and Development of Wall Heat-Transfer Correlation For Laminar Flow

With support from the U.S. Department of Energy-Office of Nuclear Energy’s (DOE-NE’s) Nuclear Energy Advanced Modeling and Simulation (NEAMS) program, an effort has been pursued to support high-temperature gas-cooled reactor (HTGR) technology development and its modeling and simulation needs. There is a particular need for advanced modeling and simulation tools to predict thermal-fluid behavior in the nuclear reactor primary system, especially in the core and the lower and upper plena, during safety-related transients. In this report, two main such activities are presented relevant to the HTGRs: (1) three-dimensional (3D) computational fluid dynamics (CFD) validation using benchmark data from the upper plenum of Texas A&M University’s 1/16th scaled very-high-temperature gas-cooled reactor (VHTR), and (2) development of wall heat-transfer correlation for laminar flow in a wall-heated pipe. The CFD tool validation exercises can be helpful to choose the models and CFD tools to simulate and design specific components of the HTRGs such as upper plenum where jet mixing is a complex phenomenon. In a loss of forced circulation event, the laminar flow can be observed during the development of natural circulation flow. This work includes the development and validation of heat transfer correlations for laminar flow using the Nek5000 CFD code due to limited available experimental data for laminar flow conditions to guide low-order models (1D). In this report, the flow characteristics of a single isothermal jet discharging into the upper plenum was investigated using the Nek5000 Large-Eddy Simulation (LES) CFD tool. Several numerical simulations were performed for various jet-discharged Reynolds numbers ranging from 3,413 to 12,819. A grid-independent study was performed. The numerical results of mean velocity, root-mean-square fluctuating velocity, and Reynolds stress were compared against the benchmark data. Good agreement was obtained between simulated and measured data for axial mean velocities, except near the upper plenum hemisphere. The maximum predicted errors for axial mean velocities at various normalized coolant channel diameter heights of 1, 5, and 10 are 1.56%, 1.88%, and 3.82%, respectively. In addition, the predicted root-mean-square fluctuating velocity and Reynolds stress are qualitatively in agreement with the experimental data. The Nek5000 code was used to develop wall-heat transfer correlation for laminar flow in a cylindrical tube. Several simulations were performed for various Reynolds flow and wall-heat fluxes. A new heat transfer correlation was developed using data from Nek5000 simulation results and regression functions in Matlab. The developed heat transfer correlation is valid for various Reynolds flows from 200 to 2000. The predicted R² value for model fit was 0.875, which ensures that 87.5% of the model data lies on the Nek5000 data. Moreover, a machine learning (ML) tool was used to train and test the Nek5000 data. A good fit of the ML-based model was observed with the test data.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

DFT-FE 1.0: A massively parallel hybrid CPU-GPU density functional theory code using finite-element discretization

In this work, we present DFT-FE 1.0, building on DFT-FE 0.6 [Comput. Phys. Commun. 246, 106853 (2020)], to conduct fast and accurate large-scale density functional theory (DFT) calculations (reaching ~ 100,000 electrons) on both many-core CPU and hybrid CPU-GPU computing architectures. This work involves improvements in the real-space formulation—via an improved treatment of the electrostatic interactions that substantially enhances the computational efficiency—as well high-performance computing aspects, including the GPU acceleration of all the key compute kernels in DFT-FE. We demonstrate the accuracy by comparing the ground-state energies, ionic forces and cell stresses on a wide-range of benchmark systems against those obtained from widely used DFT codes. Further, we demonstrate the numerical efficiency of our implementation, which yields ~ 20× CPU-GPU speed-up by using GPU acceleration on hybrid CPU-GPU nodes. Notably, owing to the parallel-scaling of the GPU implementation, we obtain wall-times of 80–140 seconds for full ground-state calculations, with stringent accuracy, on benchmark systems containing ~ 6, 000 – 15,000 electrons.

pseudopotential↗

Emergent temperature sensitivity of soil organic carbon driven by mineral associations

Abstract Soil organic matter decomposition and its interactions with climate depend on whether the organic matter is associated with soil minerals. However, data limitations have hindered global-scale analyses of mineral-associated and particulate soil organic carbon pools and their benchmarking in Earth system models used to estimate carbon cycle–climate feedbacks. Here we analyse observationally derived global estimates of soil carbon pools to quantify their relative proportions and compute their climatological temperature sensitivities as the decline in carbon with increasing temperature. We find that the climatological temperature sensitivity of particulate carbon is on average 28% higher than that of mineral-associated carbon, and up to 53% higher in cool climates. Moreover, the distribution of carbon between these underlying soil carbon pools drives the emergent climatological temperature sensitivity of bulk soil carbon stocks. However, global models vary widely in their predictions of soil carbon pool distributions. We show that the global proportion of model pools that are conceptually similar to mineral-protected carbon ranges from 16 to 85% across Earth system models from the Coupled Model Intercomparison Project Phase 6 and offline land models, with implications for bulk soil carbon ages and ecosystem responsiveness. To improve projections of carbon cycle–climate feedbacks, it is imperative to assess underlying soil carbon pools to accurately predict the distribution and vulnerability of soil carbon.

54 ENVIRONMENTAL SCIENCES↗

Supercomputing '91; Proceedings of the 4th Annual Conference on High Performance Computing, Albuquerque, NM, Nov. 18-22, 1991

Various papers on supercomputing are presented. The general topics addressed include: program analysis/data dependence, memory access, distributed memory code generation, numerical algorithms, supercomputer benchmarks, latency tolerance, parallel programming, applications, processor design, networks, performance tools, mapping and scheduling, characterization affecting performance, parallelism packaging, computing climate change, combinatorial algorithms, hardware and software performance issues, system issues. (No individual items are abstracted in this volume)

Source record↗

Spaceborne autonomous multiprocessor systems

The goal of this task is to provide technology for the specification and integration of advanced processors into the Space Station Freedom data management system environment through computer performance measurement tools, simulators, and an extended testbed facility. The approach focuses on five categories: (1) user requirements--determine the suitability of existing computer technologies and systems for real-time requirements of NASA missions; (2) system performance analysis--characterize the effects of languages, architectures, and commercially available hardware on real-time benchmarks; (3) system architecture--expand NASA's capability to solve problems with integrated numeric and symbolic requirements using advanced multiprocessor architectures; (4) parallel Ada technology--extend Ada software technology to utilize parallel architectures more efficiently; and (5) testbed--extend in-house testbed to support system performance and system analysis studies.

Fernquist, Alan↗

Experiences Using OpenMP Based on Compiler Directed Software DSM on a PC Cluster

In this work we report on our experiences running OpenMP (message passing) programs on a commodity cluster of PCs (personal computers) running a software distributed shared memory (DSM) system. We describe our test environment and report on the performance of a subset of the NAS (NASA Advanced Supercomputing) Parallel Benchmarks that have been automatically parallelized for OpenMP. We compare the performance of the OpenMP implementations with that of their message passing counterparts and discuss performance differences.

Hess, Matthias↗

Parallel 3D Mortar Element Method for Adaptive Nonconforming Meshes

High order methods are frequently used in computational simulation for their high accuracy. An efficient way to avoid unnecessary computation in smooth regions of the solution is to use adaptive meshes which employ fine grids only in areas where they are needed. Nonconforming spectral elements allow the grid to be flexibly adjusted to satisfy the computational accuracy requirements. The method is suitable for computational simulations of unsteady problems with very disparate length scales or unsteady moving features, such as heat transfer, fluid dynamics or flame combustion. In this work, we select the Mark Element Method (MEM) to handle the non-conforming interfaces between elements. A new technique is introduced to efficiently implement MEM in 3-D nonconforming meshes. By introducing an "intermediate mortar", the proposed method decomposes the projection between 3-D elements and mortars into two steps. In each step, projection matrices derived in 2-D are used. The two-step method avoids explicitly forming/deriving large projection matrices for 3-D meshes, and also helps to simplify the implementation. This new technique can be used for both h- and p-type adaptation. This method is applied to an unsteady 3-D moving heat source problem. With our new MEM implementation, mesh adaptation is able to efficiently refine the grid near the heat source and coarsen the grid once the heat source passes. The savings in computational work resulting from the dynamic mesh adaptation is demonstrated by the reduction of the the number of elements used and CPU time spent. MEM and mesh adaptation, respectively, bring irregularity and dynamics to the computer memory access pattern. Hence, they provide a good way to gauge the performance of computer systems when running scientific applications whose memory access patterns are irregular and unpredictable. We select a 3-D moving heat source problem as the Unstructured Adaptive (UA) grid benchmark, a new component of the NAS Parallel Benchmarks (NPB). In this paper, we present some interesting performance results of ow OpenMP parallel implementation on different architectures such as the SGI Origin2000, SGI Altix, and Cray MTA-2.

Feng, Huiyu↗

Predictions of Slat Noise from the 30P30N at High Angles of Attack Using Zonal Hybrid RANS-LES

Aeroacoustic predictions of slat noise from the 30P30N three-element high-lift system at high angles of attack are presented using a zonal hybrid RANS-LES method. The simulations are part of the 5th AIAA Benchmark problems for Airframe Noise Computations (BANC-V) Workshop. An economical approach utilizing structured overset grids with spatially varying span-wise grid resolution and a high-order accurate finite difference method is described. The method is utilized for near-field predictions at three angles of attack: α = 5.5, 9.5, and 14.0 degrees. Far-field noise is obtained by propagating the near-field solution using a permeable surface Ffowcs Williams-Hawkings (FWH) method. Good agreement is obtained with both near-field and far-field Power Spectral Density (PSD) data from an experimental study of the 30P30N in the 2m x 2m Kevlar-wall wind tunnel at the Japan Aerospace Exploration Agency (JAXA). Specifically, the reduction in narrow band peaks and overall broadband noise levels with increasing angle of attack is captured well using the zonal hybrid RANS-LES method.

Housman, Jeffrey A.↗

A numerical evaluation of the ambient air temperature in the Electron-Ion Collider tunnel

The Electron-Ion Collider (EIC) is a next-generation collider-accelerator that may require consistent operating temperature conditions for the beams within the accelerator tunnels to maintain stable operation. Variations in ambient temperature within the tunnel can cause thermal expansion of beampipe and component supports and can negatively affect the tunnel equipment, impacting the stability of the beamline. Modifications will be made to the Relativistic Heavy Ion Collider (RHIC) at Brookhaven National Laboratory (BNL) to create the EIC, which necessitates a temperature model that addresses these modifications. To approach this problem, the consistency of temperature changes in different tunnel sections was first evaluated by plotting RHIC tunnel temperature data at various times of the day and year. From this data, a tunnel section was selected and a 2D temperature model was created for RHIC, EIC, and EIC with added cooling configurations. Soil temperature data was analyzed to determine the maximum, average, and mode soil temperatures, which were used as boundary conditions in different temperature scenarios. Computational fluid dynamics modeling was used to create 2D temperature profiles for the configurations. From this model, the predicted temperatures indicate that further analysis is required to validate the boundary conditions and benchmark the current conditions to allow the prediction of the tunnel ambient conditions at EIC. This research can be used as a preliminary model to create an EIC tunnel cooling system that will increase the operational stability of the EIC. As a result of my work this summer, I have become familiar with computational fluid dynamics, including creating fluid dynamic simulations using ANSYS Fluent and related software. I have also learned about the project process required for planning large-scale engineering projects.

43 PARTICLE ACCELERATORS↗

Quantum reservoir computing implementation on coherently coupled quantum oscillators

Quantum reservoir computing is a promising approach for quantum neural networks, capable of solving hard learning tasks on both classical and quantum input data. However, current approaches with qubits suffer from limited connectivity. We propose an implementation for quantum reservoir that obtains a large number of densely connected neurons by using parametrically coupled quantum oscillators instead of physically coupled qubits. We analyze a specific hardware implementation based on superconducting circuits: with just two coupled quantum oscillators, we create a quantum reservoir comprising up to 81 neurons. We obtain state-of-the-art accuracy of 99% on benchmark tasks that otherwise require at least 24 classical oscillators to be solved. Our results give the coupling and dissipation requirements in the system and show how they affect the performance of the quantum reservoir. Beyond quantum reservoir computing, the use of parametrically coupled bosonic modes holds promise for realizing large quantum neural network architectures, with billions of neurons implemented with only 10 coupled quantum oscillators.

97 MATHEMATICS AND COMPUTING↗

Testing New Programming Paradigms with NAS Parallel Benchmarks

Over the past decade, high performance computing has evolved rapidly, not only in hardware architectures but also with increasing complexity of real applications. Technologies have been developing to aim at scaling up to thousands of processors on both distributed and shared memory systems. Development of parallel programs on these computers is always a challenging task. Today, writing parallel programs with message passing (e.g. MPI) is the most popular way of achieving scalability and high performance. However, writing message passing programs is difficult and error prone. Recent years new effort has been made in defining new parallel programming paradigms. The best examples are: HPF (based on data parallelism) and OpenMP (based on shared memory parallelism). Both provide simple and clear extensions to sequential programs, thus greatly simplify the tedious tasks encountered in writing message passing programs. HPF is independent of memory hierarchy, however, due to the immaturity of compiler technology its performance is still questionable. Although use of parallel compiler directives is not new, OpenMP offers a portable solution in the shared-memory domain. Another important development involves the tremendous progress in the internet and its associated technology. Although still in its infancy, Java promisses portability in a heterogeneous environment and offers possibility to "compile once and run anywhere." In light of testing these new technologies, we implemented new parallel versions of the NAS Parallel Benchmarks (NPBs) with HPF and OpenMP directives, and extended the work with Java and Java-threads. The purpose of this study is to examine the effectiveness of alternative programming paradigms. NPBs consist of five kernels and three simulated applications that mimic the computation and data movement of large scale computational fluid dynamics (CFD) applications. We started with the serial version included in NPB2.3. Optimization of memory and cache usage was applied to several benchmarks, noticeably BT and SP, resulting in better sequential performance. In order to overcome the lack of an HPF performance model and guide the development of the HPF codes, we employed an empirical performance model for several primitives found in the benchmarks. We encountered a few limitations of HPF, such as lack of supporting the "REDISTRIBUTION" directive and no easy way to handle irregular computation. The parallelization with OpenMP directives was done at the outer-most loop level to achieve the largest granularity. The performance of six HPF and OpenMP benchmarks is compared with their MPI counterparts for the Class-A problem size in the figure in next page. These results were obtained on an SGI Origin2000 (195MHz) with MIPSpro-f77 compiler 7.2.1 for OpenMP and MPI codes and PGI pghpf-2.4.3 compiler with MPI interface for HPF programs.

Jin, H.↗

Adaptive Sampling-Based Bi-Fidelity Stochastic Trust Region Method for Stochastic Derivative-Free Optimization

Bi-fidelity stochastic optimization has gained increasing attention as an efficient approach to reduce computational costs by leveraging a low-fidelity (LF) model to optimize an expensive high-fidelity (HF) objective. In this paper, we propose ASTRO-BFDF, an adaptive sampling trust-region method specifically designed for unconstrained bi-fidelity stochastic derivative-free optimization problems. In ASTRO-BFDF, the LF function serves two purposes: (i) to identify better iterates for the HF function when the optimization process indicates a high correlation between them and (ii) to reduce the variance of the HF function estimates using bi-fidelity Monte Carlo (BFMC). The algorithm dynamically determines sample sizes while adaptively choosing between crude Monte Carlo and BFMC to balance the trade-off between optimization and sampling errors. We prove that the iterates generated by ASTRO-BFDF converge to a first-order stationary point almost surely. Additionally, we demonstrate the effectiveness of the proposed algorithm through numerical experiments on synthetic benchmarks and simulation optimization problems involving discrete event systems.

97 MATHEMATICS AND COMPUTING↗

Scaling Superconducting Quantum Computers with Chiplet Architectures

Fixed-frequency transmon quantum computers (QCs) have advanced in coherence times, addressability, and gate fidelities. Unfortunately, these devices are restricted by the number of on-chip qubits, capping processing power and slowing progress toward fault-tolerance. Although emerging transmon devices feature over 100 qubits, building QCs large enough for meaningful demonstrations of quantum advantage requires overcoming many design challenges. For example, today’s transmon qubits suffer from significant variation due to limited precision in fabrication. As a result, barring significant improvements in current fabrication techniques, scaling QCs by building ever larger individual chips with more qubits is hampered by device variation. Severe device variation that degrades QC performance is referred to as a defect. Here, we focus on a specific defect known as a frequency collision. When transmon frequencies collide, their difference falls within a range that limits two-qubit gate fidelity. Frequency collisions occur with greater probability on larger QCs, causing collision-free yields to decline as the number of on-chip qubits increases. As a solution, we propose exploiting the higher yields associated with smaller QCs by integrating quantum chiplets within quantum multi-chip modules (MCMs). Yield, gate performance, and application-based analysis show the feasibility of QC scaling through modularity. Our results demonstrate that chiplet architectures, relative to monolithic designs, benefit from average yield improvements ranging from 9.6 – 92.6 × for ≲5 qubit machines. In addition, our simulations explore the design space of chiplet systems and discover configurations that demonstrate average two-qubit gate infidelity reductions that are at best 0.815 × their monolithic counterpart. Lastly, we observe that carefully-selected modular systems achieve fidelity improvements on a range of benchmark circuits.

quantum architecture↗

Nuclide Inventory Benchmark for BWR Spent Nuclear Fuel: Challenges in Evaluation of Modeling Data Assumptions and Uncertainties

This work discusses challenges and approaches to uncertainty analyses associated with the development of a nuclide inventory benchmark for fuel irradiated in a boiling water reactor. The benchmark under consideration is being developed based on experimental data from the SFCOMPO international database. The focus herein is on how to address missing data in fuel design and operating conditions that are important for adequately simulating the time-dependent changes in fuel during irradiation in the reactor. The effects of modeling assumptions and uncertainties in modeling parameters on the calculated nuclide inventory were analyzed and quantified through computational models developed using capabilities in the SCALE code system. Particular attention was given to the impact of the power history and water coolant density on the calculated nuclide inventory, as well as to the effect of geometry modeling considerations not usually addressed in a nuclide inventory benchmark. These considerations include gap closure, channel bow, and channel corner radius, which do not usually apply to regular reactor operation but are relevant for assessing impacts of potential anomalous operating scenarios.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Classical Benchmarks for Variational Quantum Eigensolver Simulations of the Hubbard Model

Simulating the Hubbard model is of great interest to a wide range of applications within condensed matter physics, however its solution on classical computers remains challenging in dimensions larger than one. The relative simplicity of this model, embodied by the sparseness of the Hamiltonian matrix, allows for its efficient implementation on quantum computers, and for its approximate solution using variational algorithms such as the variational quantum eigensolver. While these algorithms have been shown to reproduce the qualitative features of the Hubbard model, their quantitative accuracy in terms of producing true ground state energies and other properties, and the dependence of this accuracy on the system size and interaction strength, the choice of variational ansatz, and the degree of spatial inhomogeneity in the model, remains unknown. Here we present a rigorous classical benchmarking study, demonstrating the potential impact of these factors on the accuracy of the variational solution of the Hubbard model on quantum hardware, for systems with up to 32 qubits. We find that even when using the most accurate wavefunction ansätze for the Hubbard model, the error in its ground state energy and wavefunction plateaus for larger lattices, while stronger electronic correlations magnify this issue. Concurrently, spatially inhomogeneous parameters and the presence of off-site Coulomb interactions only have a small effect on the accuracy of the computed ground state energies. Our study highlights the capabilities and limitations of current approaches for solving the Hubbard model on quantum hardware, and we discuss potential future avenues of research.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗