Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “computer system benchmarking”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

Accelerated kinetic model for global macro stability studies of high-beta fusion reactors

The field reversed configuration (FRC), such as studied in the C-2W experiment at TAE Technologies, is an attractive candidate for realizing a nuclear fusion reactor. In an FRC, kinetic ion effects play the majority role in macroscopic stability, which allows global stability studies to make use of fluid-kinetic hybrid (also referred to as Ohm's law) models wherein ions are treated kinetically while electrons are treated as a fluid. The development and validation of such a hybrid particle-in-cell algorithm in the Exascale Computing Project code WarpX are reported here. Implementation of this model in the WarpX framework benefits from the numerical efficiency of WarpX as well as its scalability on large HPC systems and portability to different architectures. Performance benchmarks of the new algorithm for large, 3-dimensional, full device simulations from the Perlmutter supercomputer are presented. Results of a series of FRC simulations are discussed in which the impact of two-fluid effects on the tilt-mode growth rate was studied. It was observed that, in agreement with previous Hall-MHD studies, two-fluid effects have a stabilizing impact on the tilt mode.

Physics↗

Unprecedented cloud resolution in a GPU-enabled full-physics atmospheric climate simulation on OLCF’s summit supercomputer

Clouds represent a key uncertainty in future climate projection. While explicit cloud resolution remains beyond our computational grasp for global climate, we can incorporate important cloud effects through a computational middle ground called the Multi-scale Modeling Framework (MMF), also known as Super Parameterization. This algorithmic approach embeds high-resolution Cloud Resolving Models (CRMs) to represent moist convective processes within each grid column in a Global Climate Model (GCM). The MMF code requires no parallel data transfers and provides a self-contained target for acceleration. This study investigates the performance of the Energy Exascale Earth System Model-MMF (E3SM-MMF) code on the OLCF Summit supercomputer at an unprecedented scale of simulation. Hundreds of kernels in the roughly 10K lines of code in the E3SM-MMF CRM were ported to GPUs with OpenACC directives. A high-resolution benchmark using 4600 nodes on Summit demonstrates the computational capability of the GPU-enabled E3SM-MMF code in a full physics climate simulation.

58 GEOSCIENCES↗

Validation of CFD/Heat Transfer Software for Turbine Blade Analysis

I am an intern in the Turbine Branch of the Turbomachinery and Propulsion Systems Division. The division is primarily concerned with experimental and computational methods of calculating heat transfer effects of turbine blades during operation in jet engines and land-based power systems. These include modeling flow in internal cooling passages and film cooling, as well as calculating heat flux and peak temperatures to ensure safe and efficient operation. The branch is research-oriented, emphasizing the development of tools that may be used by gas turbine designers in industry. The branch has been developing a computational fluid dynamics (CFD) and heat transfer code called GlennHT to achieve the computational end of this analysis. The code was originally written in FORTRAN 77 and run on Silicon Graphics machines. However the code has been rewritten and compiled in FORTRAN 90 to take advantage of more modem computer memory systems. In addition the branch has made a switch in system architectures from SGI's to Linux PC's. The newly modified code therefore needs to be tested and validated. This is the primary goal of my internship. To validate the GlennHT code, it must be run using benchmark fluid mechanics and heat transfer test cases, for which there are either analytical solutions or widely accepted experimental data. From the solutions generated by the code, comparisons can be made to the correct solutions to establish the accuracy of the code. To design and create these test cases, there are many steps and programs that must be used. Before a test case can be run, pre-processing steps must be accomplished. These include generating a grid to describe the geometry, using a software package called GridPro. Also various files required by the GlennHT code must be created including a boundary condition file, a file for multi-processor computing, and a file to describe problem and algorithm parameters. A good deal of this internship will be to become familiar with these programs and the structure of the GlennHT code. Additional information is included in the original extended abstract.

Kiefer, Walter D.↗

Flow and Heat Transfer in 180-Degree Turn Square Ducts: Effects of Turning Configuration and System Rotation

Forced flow through channels connected by sharp bends is frequently encountered in various rocket and gas turbine engines. For example, the transfer ducts, the coolant channels surround the combustion chamber, the internal cooling passage in a blade or vane, the flow path in the fuel element of a nuclear rocket engine, the flow around a pressure relieve valve piston, and the recirculated base flow of multiple engine clustered nozzles. Transport phenomena involved in such a flow passage are complex and considered to be very different from those of conventional turning flow with relatively mild radii of curvature. While previous research pertaining to this subject has been focused primarily on the experimental heat transfer, very little analytical work is directed to understanding the flowfield and energy transport in the passage. Therefore, the primary goal of this paper is to benchmark the predicted wall heat fluxes using a state-of-the-art computational fluid dynamics (CFD) formulation against those of measurement for a rectangular turn duct. Other secondary goals include studying the effects of turning configurations, e.g., the semi-circular turn, and the rounded-corner turn, and the effect of system rotation. The computed heat fluxes for the rectangular turn duct compared favorably with those of the experimental data. The results show that the flow pattern, pressure drop, and heat transfer characteristics are different among the three turning configurations, and are substantially different with system rotation. Also demonstrated in this work is that the present computational approach is quite effective and efficient and will be suitable for flow and thermal modeling in rocket and turbine engine applications.

Wang, Ten-See↗

Quantifying uncertainties due to irreducible three-body forces in deuteron-nucleus reactions

Deuteron-induced nuclear reactions are an essential tool for probing the structure of nuclei as well as astrophysical information such as (n, γ) cross sections. The deuteron-nucleus system is typically described within a Faddeev three-body model consisting of a neutron (n), a proton (p), and the target nucleus (A) interacting through pairwise phenomenological potentials. While Faddeev techniques enable the exact description of the three-body dynamics, their predictive power is limited in part by the omission of irreducible neutron-proton-nucleus three-body force (n–p–A 3BF). Here, our goal is to quantify systematic uncertainties stemming from the reduction of deuteron-nucleus (d + A) dynamics to a picture of three pointlike nuclear clusters interacting via pairwise nucleon-nucleus forces, using as testing grounds d + α scattering and the 6 Li ground state. We particularly focus on quantifying uncertainties arising from the full antisymmetrization of the (A + 2)-body system with the target nucleus fixed in its ground state. We adopt the ab initio no-core shell model coupled with the resonating group method (NCSM/RGM) to compute microscopic n–α and p–α interactions, and use them in a three-body description of the d + α system by means of momentum-space Faddeev-type equations. Simultaneously, we also carry out ab initio calculations of d + α scattering and 6 Li ground state by means of six-body NCSM/RGM calculations to serve as a benchmark for the three-body model predictions given by the Faddeev calculations. By comparing the Faddeev and NCSM/RGM results, we show that the irreducible n–p–α 3BF has a non-negligible effect on bound state and scattering observables alike. Specifically, the Faddeev approach yields a 6 Li ground state that is approximately 600 keV shallower than the one obtained with the NCSM/RGM. Additionally, the Faddeev calculations for d + α scattering yield a 3 + resonance that is located approximately 400 keV higher in energy compared to the NCSM/RGM result. The shape of the d + α angular distributions computed using the two approaches also differ, owing to the discrepancy in the predictions of the 3 + resonance energy. The Faddeev three-body model predictions for d + α scattering and 6 Li using microscopic n–α and p–α potentials differ from those computed microscopically with the NCSM/RGM. These discrepancies are due to the n–p–α 3BF, which arises from two-nucleon exchange terms in the microscopic d–α interaction and are not accounted for in the three-body model Faddeev calculations. This study lays the foundation for future parametrizations of the 3BF due to Pauli exclusion principle effects in improved three-body calculations of deuteron-induced reactions.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Active space selection with self-healing diffusion Monte Carlo algorithms for periodic solids

Multideterminant Diffusion Monte Carlo (DMC) displays improved accuracy over single determinant DMC. Self-Healing Diffusion Monte Carlo (SHDMC) is a DMC based method that iteratively improves a multideterminant trial wavefunction. Although configuration interaction or complete active space (CAS) methods are very accurate and computationally feasible for many systems, they are not optimal for application to solids. SHDMC is accurate and designed for application to solids, so developing SHDMC based active space selection algorithms is a worthy endeavor. Here, we present and compare active space selection algorithms that are designed for use in conjunction with SHDMC, without relying on external approaches. For benchmarking, we calculated the ground state energy of a small unit cell of graphene and compared the results with a complete basis set extrapolated selected CI and a reference SHDMC trajectory. We found that systematically expanding the active space using an “auto-branching” algorithm optimally balances accuracy with computational practicality. To the best of our knowledge, this is the first work that demonstrates completely self-contained DMC-based active space selection algorithms that do not depend on external methods for determinant selection.

Spanedda, Nicole [ORNL]↗

Summer Internship Report: ARA2 Benchmarking

Over the past decade, the RISC-V Instruction Set Architecture (ISA) has emerged as a significant player in both academic and industrial processor design due to its open-source nature, modular extension system, and versatility across domains ranging from microcontrollers to high-performance computing (HPC). One of its most important recent advancements is the RISC-V Vector Extension (RVV), which enables explicit data-level parallelism through vector registers and vectorized instructions. Unlike traditional SIMD (Single Instruction, Multiple Data) architectures that fix vector lengths at design time, RVV uses the concept of VLEN (vector register length) as a hardware-independent parameter and allows software to adapt dynamically to the available vector width. This flexible approach ensures portability across implementations while enabling scalable performance. The ARA2 core is a parameterizable RISC-V vector processor developed at the Integrated Systems Lab at ETH Zürich and the University of Bologna. Designed as a tightly-coupled accelerator to a scalar RISC-V core, ARA2 implements the RVV 1.0 specification and offers tunable architectural parameters such as the number of vector lanes, VLEN, and cache sizes.

97 MATHEMATICS AND COMPUTING↗

Theoretical framework for new magnetic materials for quantum computing and information storage. Final report for the Award No. DE-SC0018910

The focus of this grant was on molecular magnetic materials for information storage and quantum computing. We have been developing robust, first-principle methods for computing relevant electronic and magnetic properties of molecular building blocks (SMMs) of novel magnetic materials and quantum computers. These tools enable theoretical modeling of SMMs’ behavior, facilitating the interpretation of experimental studies and aiding the design of novel magnetic materials. Our strategy is based on the spin-flip (SF) approach, which extends the hierarchy of black-box single-reference methods to strongly correlated systems. Specifically, we developed general scalable algorithms and computer codes for calculating molecular properties, with an emphasis on spin-related properties, such as zero-field splittings, hyperfine couplings, and g-tensors. While our primary focus was on SF wave functions and SF-TDDFT, the underlying theory and computer codes were formulated using reduced density matrices, such that these tools are applicable to a broader class of methods. To extend the scope of applicability of wave-function-based SF methods to larger systems, we developed reduced-scaling approaches for the equation-of-motion coupled-cluster (EOM-CC) methods and continue developing libtensor (our open-source general tensor contraction library for many-body methods). We carried out extensive benchmarks and also carried out several applications.

36 MATERIALS SCIENCE↗

A Sparse Distributed Gigascale Resolution Material Point Method

In this paper, we present a four-layer distributed simulation system and its adaptation to the Material Point Method (MPM). The system is built upon a performance portable C++ programming model targeting major High-Performance-Computing (HPC) platforms. A key ingredient of our system is a hierarchical block-tile-cell sparse grid data structure that is distributable to an arbitrary number of Message Passing Interface (MPI) ranks. We additionally propose strategies for efficient dynamic load balance optimization to maximize the efficiency of MPI tasks. Our simulation pipeline can easily switch among backend programming models, including OpenMP and CUDA, and can be effortlessly dispatched onto supercomputers and the cloud. Finally, we construct benchmark experiments and ablation studies on supercomputers and consumer workstations in a local network to evaluate the scalability and load balancing criteria. We demonstrate massively parallel, highly scalable, and gigascale resolution MPM simulations of up to 1.01 billion particles for less than 323.25 seconds per frame with 8 OpenSSH-connected workstations.

97 MATHEMATICS AND COMPUTING↗

An LDA investigation of three-dimensional normal shock-boundary layer interactions in a corner

Nonintrusive, three-dimensional, measurements have been made of a normal shock wave-turbulent boundary layer interaction. The measurements were made in the corner of the test section of a continuous supersonic wind tunnel in which a normal shock wave had been stabilized. LDA, surface pressure measurement and flow visualization techniques were employed for two freestream Mach number test cases: 1.6 and 1.3. The former contained separated flow regions and a system of shock waves. The latter was found to be far less complicated. The reported results are believed to accurately define the flow physics of each case and may be used as benchmark data to verify three-dimensional computer codes.

Chriss, R. M.↗

Dispatch Manager for NEML2 Constitutive Model Calculations Embedded in MOOSE

This report describes the extended capabilities of the NEML2 constitutive modeling library, including a flexible and efficient work dispatching system designed to leverage both CPU and GPU resources. This enhancement addresses one of the primary computational challenges in large-scale simulations: the ability to distribute and execute batches of material model evaluations across heterogeneous computing devices. The new dispatch system introduces a modular set of dispatcher and scheduler classes that coordinate the flow of data and execution between devices. The dispatcher is responsible for efficiently packaging work, managing device-specific memory operations, and synchronizing results. This modularity allows for extensibility, making it straightforward to integrate additional computing backends in the future. From an implementation standpoint, the dispatcher system interfaces seamlessly with NEML2's existing models. They handle device-aware tensor operations, optimize memory transfers, and support asynchronous execution when applicable. This design ensures that batches of material points can be evaluated concurrently, substantially improving throughput compared to previous single-device or serial implementations. These improvements not only enhance the raw performance of NEML2 but also improve its usability in multiscale and high-fidelity simulations, where the simultaneous evaluation of large material point batches is critical. Benchmarks included in the report demonstrate the system’s scalability, highlighting its effectiveness when leveraging modern GPU architectures.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

FY24 Progress Report on Viscosity and Thermal Conductivity Measurements of Nuclear Industry Relevant Chloride Salts: An Experimental and Computational Study

As presented in this report, experimental and computational techniques were performed to assess the viscosity and thermal conductivity of key alkali and actinide chloride mixtures for molten salt reactor developers. These mixtures were pure LiCl, NaCl-KCl, LiCl-NaCl, LiCl-KCl, LiCl-NaCl-KCl, and NaCl-UCl 3 . Experimental measurements of viscosity were performed with a rolling ball viscometer, whereas experimental measurements of thermal conductivity were performed with a variable gap apparatus. Additional benchmarking work was performed using both property measurement systems to prepare for x-ray radiography in stainless-steel crucibles for viscosity and to ensure that calibration methods were accurate for thermal conductivity before assessing the NaCl-UCl 3 system. Validation data for the NaCl-UCl 3 in literature are minimal. Details on the calibration methods, salt measurement processes, and sources of error and uncertainty are discussed in detail for both property measurements. The computational methods described herein involved ab-initio molecular dynamics (AIMD) calculations using CP2K. The calculations were performed for the LiCl-KCl-NaCl and NaCl-UCl 3 systems. These calculations not only provided thermophysical property estimations for comparison to experimental data, but they also allowed for the determination of diffusion coefficients, coordination numbers, and radial distribution functions to provide insight into ion mobility and local coordination environments, which is linked to macroscopic property trends.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Virtualized Logical Qubits: A 2.5D Architecture for Error-Corrected Quantum Computing

Current, near-term quantum devices have shown great progress in the last several years culminating recently with a demonstration of quantum supremacy. In the medium-term, however, quantum machines will need to transition to greater reliability through error correction, likely through promising techniques like surface codes which are well suited for near-term devices with limited qubit connectivity. We discover quantum memory, particularly resonant cavities with transmon qubits arranged in a 2.5D architecture, can efficiently implement surface codes with substantial hardware savings and performance/fidelity gains. Specifically, we virtualize logical qubits by storing them in layers of qubit memories connected to each transmon. Surprisingly, distributing each logical qubit across many memories has a minimal impact on fault tolerance and results in substantially more efficient operations. Our design permits fast transversal application of CNOT operations between logical qubits sharing the same physical address (same set of cavities) which are 6x faster than standard lattice surgery CNOTs. We develop a novel embedding which saves approximately 10x in transmons with another 2x savings from an additional optimization for compactness. Although qubit virtualization pays a 10x penalty in serialization, advantages in the transversal CNOT and in area efficiency result in fault-tolerance and performance comparable to conventional 2D transmon-only architectures. Our simulations show our system can achieve fault tolerance comparable to conventional two-dimensional grids while saving substantial hardware. Furthermore, our architecture can produce magic states at 1.22x the baseline rate given a fixed number of transmon qubits. Here, this is a critical benchmark for future fault-tolerant quantum computers as magic states are essential and machines will spend the majority of their resources continuously producing them. This architecture substantially reduces the hardware requirements for fault-tolerant quantum computing and puts within reach a proof-of-concept experimental demonstration of around 10 logical qubits, requiring only 11 transmons and 9 attached cavities in total.

quantum computing↗

Issue Summary of INL Phase IV Transient Results for IAEA CRP on HTGR UAM Benchmark

This report details the Parallel and Highly Innovative Simulation for Idaho National Laboratory (INL) Code System (PHISICS)/Reactor Excursions and Leak Analysis Program (RELAP5)-3D results obtained for the transient core exercises defined for Phase IV of the International Atomic Energy Agency (IAEA) Coordinated Research Project (CRP) on high-temperature gas cooled reactor (HTGR) uncertainty analysis in modeling (UAM). The Phase III models and results are linked to the earlier Standardized Computer Analyses for Licensing Evaluation (SCALE)/Sampler/New ESC-based Weighting Transport (NEWT) data generated for the lattice physics (lattice) stage Phase I of the CRP. The focus of this report is the Uncertainty/Sensitivity Assessment (U/SA) of the prismatic modular high-temperature gas cooled reactor (MHTGR)-350 design, and specifically for Exercises IV-1 and IV-2 of the benchmark: the Control Rod Withdrawal (CRW) and Pressurised Loss of Cooling (PLOFC) events. The statistical U/SA methodology is implemented and demonstrated using the RAVEN code, based on perturbed cross-section libraries obtained from the SCALE/Sampler sequence. Uncertainties in nuclear data (cross-sections and the average number of neutrons produced per fission, 235U[¯v ]) lead to standard deviations (uncertainties of one s) of approximately 0.5% in the core eigenvalues of the MHTGR-350 and core models. For the coupled neutronics/thermal fluid model, local power density uncertainties up to 3.6% were observed in the colder regions of the core, while the local maximum fuel temperature uncertainties reached 1.5% for the models that included thermal fluid uncertainties. The addition of thermal fluid uncertainties dominated the impacts of nuclear data uncertainties in all cases. The main contributors to uncertainties in the power density and fuel temperatures during the transients were uncertainties in the reactor operating conditions (total power, inlet mass flow rate and inlet gas temperature). Variations in the bypass flows did not have significant impact on any of the output variables. For the nuclear data uncertainties it was found that the 235U(¯v ) / 235U(¯v ) covariance produced the largest sensitivities in terms of its impact on the eigenvalue and peak reactor power. It was also observed that the impact of any nuclear data uncertainties on the maximum fuel temperature was much less significant that the impact on eigenvalue and power. Another important finding was that although the use of eight or more energy groups is recommended for best-estimate HTGR simulation, two-group models produced acceptable uncertainty and sensitivity results for most FOMs. Since the statistical U/SA methodology is computationally expensive, and most transient solver requirements will scale directly with the number of energy groups, two energy groups could be used by HTGR developers during the early stages of design when larger uncertainty margins can be tolerated.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Experimental test of model predictive control in a variable air volume system

Model predictive control (MPC) has been widely studied as a promising approach for improving energy efficiency and operational flexibility in buildings, yet its real-world performance for commercial variable air volume (VAV) systems remains insufficiently characterized. In particular, the impacts of model mismatch on control robustness, real-time computational burden, and device-level operation are rarely evaluated using long-term field data. Here, this study presents a comprehensive experimental evaluation of MPC applied to a full-scale VAV system in Oak Ridge National Laboratory’s Flexible Research Platform-2 building with constant cooling/heating temperature setpoints and no occupancy. The study offers three key advantages over existing work: (1) it uses a representative building in a full-scale experimental test, capturing realistic system dynamics and complexity; (2) it evaluates a relatively sophisticated MPC formulation using two different optimization solvers (Gurobi and PSO), fully accounting for computational complexity and methodological diversity; and (3) it systematically assesses potential negative impacts on various building devices, benchmark against a well-established baseline, ASHRAE Guideline 36 (G36). To isolate zone- and air-handling-unit–level supervisory control effects, the supply fan was operated with a fixed static pressure setpoint under all strategies, and the trim-and-response static pressure reset in G36 was not enabled. Results show that MPC maintained thermal comfort while improving energy efficiency. Abrupt solar radiation variations degraded performance. Computation times ranged from ∼1 s (Gurobi) to ∼ 70 s (PSO). Compared with G36, MPC achieves 33% energy savings and reduces median reheat coil output by approximately a factor of 5–10 for a representative cooling day under matched weather conditions. However, it increases the maximum discomfort deviation from 0.5 to 1°C and results in a 32% increase in staging frequency. In addition, PSO-based MPC introduced damper oscillations, also affecting actuator longevity.

ASHRAE guideline 36↗

STEPS: A Portable Numerical Simulation Toolkit for Electrical Power System Dynamic Studies

Numerical simulation is the key technique for large scale power system analysis. Redistribution of global renewable power via international interconnections requires new simulation tools to study the interconnected systems with different nominal frequencies as a whole. In this paper we introduce an open source simulation toolkit for electrical power systems (STEPS) which is hosted at Github. Its kernel is coded in C++ with major functions of power flow and electro-mechanical dynamic simulation. Flexible options are provided and configurable to improve power flow solution and dynamic simulation. Common devices and models are supported in STEPS for AC/DC hybrid system studies. Studies of interconnected systems with different nominal frequencies is supported in STEPS for research of international interconnection. Application program interfaces are provided and wrapped with Python to enable high-level interfaces for general applications. STEPS is thread safe and parallel computation is supported in both kernel and script levels to accelerate simulation. It is portable and works on Windows and GNU/Linux platforms. Cases from small to large scale systems are thoroughly tested to validate the toolkit with commercial packages as benchmarks.

42 ENGINEERING↗

The NAS parallel benchmarks

A new set of benchmarks has been developed for the performance evaluation of highly parallel supercomputers in the framework of the NASA Ames Numerical Aerodynamic Simulation (NAS) Program. These consist of five 'parallel kernel' benchmarks and three 'simulated application' benchmarks. Together they mimic the computation and data movement characteristics of large-scale computational fluid dynamics applications. The principal distinguishing feature of these benchmarks is their 'pencil and paper' specification-all details of these benchmarks are specified only algorithmically. In this way many of the difficulties associated with conventional benchmarking approaches on highly parallel systems are avoided.

Bailey, D. H.↗