Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “computer system benchmarking”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Radiant Energy Measurements from a Scaled Jet Engine Axisymmetric Exhaust Nozzle for a Baseline Code Validation Case

A non-flowing, electrically heated test rig was developed to verify computer codes that calculate radiant energy propagation from nozzle geometries that represent aircraft propulsion nozzle systems. Since there are a variety of analysis tools used to evaluate thermal radiation propagation from partially enclosed nozzle surfaces, an experimental benchmark test case was developed for code comparison. This paper briefly describes the nozzle test rig and the developed analytical nozzle geometry used to compare the experimental and predicted thermal radiation results. A major objective of this effort was to make available the experimental results and the analytical model in a format to facilitate conversion to existing computer code formats. For code validation purposes this nozzle geometry represents one validation case for one set of analysis conditions. Since each computer code has advantages and disadvantages based on scope, requirements, and desired accuracy, the usefulness of this single nozzle baseline validation case can be limited for some code comparisons.

Baumeister, Joseph F.↗

Evaluation of the Sum-of-Fractions Methodology for Water and Polyethylene Moderated Systems

Sum-of-Fractions is a method intended to assure a subcritical margin for aqueous solutions and slurries of fissionable isotopes. The method indicates that a system is subcritical if the sum of the ratios of the mass of each isotope in a mixture to its individual minimum subcritical mass limit is less than or equal to one. The basis of the Sum-of Fractions has historically been derived from allowances given in ANSI/ANS-8.15-1981. However, the allowance was removed in ANSI/ANS-8.15-2014 due to a lack of technical basis. A methodology was developed to assess the validity of using the Sum-of-Fractions for water or polyethylene moderated systems for the following nuclides: 232 U, 233 U, 234 U, 235 U, 237 Np, 236 Pu, 238 Pu, 239 Pu, 240 Pu, 241 Pu, 242 Pu, 241 Am, 242 mAm, 243 Am, 242 Cm, 243 Cm, 244 Cm, 245 Cm, 246 Cm, 247 Cm, 249 Cf, and 251 Cf. The methodology uses available benchmark data for mixtures of 233 U, 235 U, and 239 Pu to establish the calculational margin, and a mass limit reduction to establish the margin of subcriticality. Water or polyethylene moderated and reflected mixtures containing the nuclides are evaluated with SCALE 6.2.4. Including the calculational margin, subcritical mass limits for each nuclide were computed for optimally water or polyethylene moderated and fully reflected systems. These masses were used to create nuclide mixtures in which the sum of the mass to subcritical mass limit ratios is one. The various nuclide mixtures were modeled over a range of moderation and demonstrate the k eff does not exceed the calculational margin. For additional assurance of subcriticality, a significant mass reduction is applied to each computed minimum critical mass of the nuclides without adequate benchmark data consistent with the method in ANSI/ANS-8.15-2014.

07 ISOTOPE AND RADIATION SOURCES↗

Scalable Predictive Control and Optimization for Grid Integration of Large-Scale Distributed Energy Resources: Preprint

Integration of a large number of distributed energy resources (DERs) into the power grid needs a scalable power balancing method. We formulate the power balancing problem as a look-ahead optimization problem to be solved sequentially by a power distribution system aggregator based on a model predictive control (MPC) framework. Solving large-scale look-ahead control problem requires proper configuration of the control steps. In this paper, to solve large-scale control problems, we propose a variable time granularity where control time steps nearby the current control step have finer resolutions. The aggregator objective includes maximization of power production revenue and minimization of power purchasing expense, renewable power curtailment, and mileage costs for energy storage and electric vehicle (EV) charging stations while satisfying system capacity and operational constraints. The control problem is formulated as a mixed-integer linear program (MILP) and solved using the XpressMP solver. We perform simulations considering a copper plate representation of a large distribution network consisting of 2507 devices (controllable DERs) including curtailable photovoltaics (PVs), energy storage batteries, EV charging stations, and buildings with heating, ventilation, and air conditioning units (HVACs). We show the effectiveness of the proposed approach in managing DERs interactively for maximum energy trading profit and local supply-demand power balancing. Finally, we demonstrate that the proposed method outperformed other benchmark controllers regarding computation time without compromising operational performance.

DER↗

VECTOR Phase 1 Dataset: CAV Trajectory and Energy Consumption Records

This dataset contains benchmark experimental data from Phase 1 of the VECTOR project, focusing on the energy impact of CAV hardware components. The dataset includes vehicle trajectory data (speed and position) and corresponding energy consumption records collected from a CAV platform equipped with lidar, cameras, onboard computation units, and communication modules. The primary objective is to quantify the baseline energy consumption attributable to sensing and computing systems, independent of any advanced cooperative control strategies. During experiments, the leading vehicle followed a predetermined velocity profile, and the following CAV mirrored this trajectory using a basic car-following control to ensure consistent driving behavior. This setup enables a reliable benchmark for assessing the energy cost introduced by onboard CDA hardware (e.g., lidar and GPU-based processing). The dataset is essential for evaluating energy baselines and supports future comparative studies involving additional cooperative strategies. ![system img](system.png) ![vector img](vector.png)

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Parallel Ada benchmarks for the SVMS

The use of parallel processing paradigm to design and develop faster and more reliable computers appear to clearly mark the future of information processing. NASA started the development of such an architecture: the Spaceborne VHSIC Multi-processor System (SVMS). Ada will be one of the languages used to program the SVMS. One of the unique characteristics of Ada is that it supports parallel processing at the language level through the tasking constructs. It is important for the SVMS project team to assess how efficiently the SVMS architecture will be implemented, as well as how efficiently Ada environment will be ported to the SVMS. AUTOCLASS II, a Bayesian classifier written in Common Lisp, was selected as one of the benchmarks for SVMS configurations. The purpose of the R and D effort was to provide the SVMS project team with the version of AUTOCLASS II, written in Ada, that would make use of Ada tasking constructs as much as possible so as to constitute a suitable benchmark. Additionally, a set of programs was developed that would measure Ada tasking efficiency on parallel architectures as well as determine the critical parameters influencing tasking efficiency. All this was designed to provide the SVMS project team with a set of suitable tools in the development of the SVMS architecture.

Collard, Philippe E.↗

Simultaneous Optimization of Nuclear–Electronic Orbitals

Accurate modeling of important nuclear quantum effects, such as nuclear delocalization, zero-point energy, and tunneling, as well as non-Born-Oppenheimer effects, requires treatment of both nuclei and electrons quantum mechanically. The nuclear–electronic orbital (NEO) method provides an elegant framework to treat specified nuclei, typically protons, on the same level as the electrons. In conventional electronic structure theory, finding a converged ground state can be a computationally demanding task; converging NEO wavefunctions, due to their coupled electronic and nuclear nature, is even more demanding. Herein, we present an efficient simultaneous optimization method that uses the direct inversion in the iterative subspace method to simultaneously converge wavefunctions for both the electrons and quantum nuclei. In conclusion, benchmark studies show that the simultaneous optimization method can significantly reduce the computational cost compared to the conventional stepwise method for optimizing NEO wavefunctions for multicomponent systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Performance Improvements of the Griffin Solvers in FY24

The Griffin code is a MOOSE-based reactor physics application jointly developed by Idaho National Laboratory and Argonne National Laboratory under the Department of Energy Office of Nuclear Energy Nuclear Energy Advanced Modeling and Simulation Program. This fiscal year, we have made significant efforts to improve the performance of transport solver options and cross-section generation for the efficient use of Griffin in advanced reactor applications. For the HFEM-PN solver, the residual evaluations of HFEM kernels were optimized by utilizing the pre- computed averaged cross sections for individual elements. Numerical integration involving the evaluation of basis functions at quadrature points was bypassed by facilitating precomputed element mass matrices for response matrices. Red-black iterations were improved by introducing a new generalized minimum residual based solver. The memory usage of response matrix storage was significantly reduced by applying basis function rotations on interfaces and calculating volumetric odd-parity moments on the fly. Additionally, the adjoint flux and transient calculation capabilities of the HFEM-PN solver were successfully implemented and verified using the TWIGL benchmark problem. For the DFEM-SN solver, memory footprint and computation time were significantly reduced by not treating angular flux vectors as the MOOSE nonlinear system vectors. Specifically for IQS, scalar adjoint weighting was introduced to further eliminate angular adjoint flux storage in the MOOSE auxiliary system. It was demonstrated through the three-dimensional Advanced Burner Test Reactor core problem that the memory usage for transient calculations with the IQS method was reduced by over 7.5× compared to before the optimizations. For the self-shielding application programming interface, a new double-heterogeneity treatment method, named the Bell Function-Based Analytic Two-Region Slowing Down Method, was developed to efficiently flux-volume homogenize TRISO particles with the matrix. Additionally, optimizations were made to hyper- fine group (HFG) slowing down calculations by pretabulating collision probability coefficients and grouping isotopes, significantly reducing the computational time for calculating scattering sources per HFG. Lastly, the pin power reconstruction module was extended to account for temporal behavior in a microreactor analysis problem, specifically for a control drum transient. Verification tests for each of these improvements demonstrated significant performance enhancements and memory reduction.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

High-Speed On-Board Data Processing for Science Instruments

A new development of on-board data processing platform has been in progress at NASA Langley Research Center since April, 2012, and the overall review of such work is presented in this paper. The project is called High-Speed On-Board Data Processing for Science Instruments (HOPS) and focuses on a high-speed scalable data processing platform for three particular National Research Council's Decadal Survey missions such as Active Sensing of CO2 Emissions over Nights, Days, and Seasons (ASCENDS), Aerosol-Cloud-Ecosystems (ACE), and Doppler Aerosol Wind Lidar (DAWN) 3-D Winds. HOPS utilizes advanced general purpose computing with Field Programmable Gate Array (FPGA) based algorithm implementation techniques. The significance of HOPS is to enable high speed on-board data processing for current and future science missions with its reconfigurable and scalable data processing platform. A single HOPS processing board is expected to provide approximately 66 times faster data processing speed for ASCENDS, more than 70% reduction in both power and weight, and about two orders of cost reduction compared to the state-of-the-art (SOA) on-board data processing system. Such benchmark predictions are based on the data when HOPS was originally proposed in August, 2011. The details of these improvement measures are also presented. The two facets of HOPS development are identifying the most computationally intensive algorithm segments of each mission and implementing them in a FPGA-based data processing board. A general introduction of such facets is also the purpose of this paper.

Beyon, Jeffrey Y.↗

Parallelization of Lower-Upper Symmetric Gauss-Seidel Method for Chemically Reacting Flow

Development of technologies for exploration of the solar system has revived an interest in computational simulation of chemically reacting flows since planetary probe vehicles exhibit non-equilibrium phenomena during the atmospheric entry of a planet or a moon as well as the reentry to the Earth. Stability in combustion is essential for new propulsion systems. Numerical solution of real-gas flows often increases computational work by an order-of-magnitude compared to perfect gas flow partly because of the increased complexity of equations to solve. Recently, as part of Project Columbia, NASA has integrated a cluster of interconnected SGI Altix systems to provide a ten-fold increase in current supercomputing capacity that includes an SGI Origin system. Both the new and existing machines are based on cache coherent non-uniform memory access architecture. Lower-Upper Symmetric Gauss-Seidel (LU-SGS) relaxation method has been implemented into both perfect and real gas flow codes including Real-Gas Aerodynamic Simulator (RGAS). However, the vectorized RGAS code runs inefficiently on cache-based shared-memory machines such as SGI system. Parallelization of a Gauss-Seidel method is nontrivial due to its sequential nature. The LU-SGS method has been vectorized on an oblique plane in INS3D-LU code that has been one of the base codes for NAS Parallel benchmarks. The oblique plane has been called a hyperplane by computer scientists. It is straightforward to parallelize a Gauss-Seidel method by partitioning the hyperplanes once they are formed. Another way of parallelization is to schedule processors like a pipeline using software. Both hyperplane and pipeline methods have been implemented using openMP directives. The present paper reports the performance of the parallelized RGAS code on SGI Origin and Altix systems.

Yoon, Seokkwan↗

h5bench: A unified benchmark suite for evaluating HDF5 I/O performance on pre‐exascale platforms

Summary Parallel I/O is a critical technique for moving data between compute and storage subsystems of supercomputers. With massive amounts of data produced or consumed by compute nodes, high‐performant parallel I/O is essential. I/O benchmarks play an important role in this process; however, there is a scarcity of I/O benchmarks representative of current workloads on HPC systems. Toward creating representative I/O kernels from real‐world applications, we have created h5bench , a set of I/O kernels that exercise hierarchical data format version 5 (HDF5) I/O on parallel file systems in numerous dimensions. Our focus on HDF5 is due to the parallel I/O library's heavy usage in various scientific applications running on supercomputing systems. The various tests benchmarked in the h5bench suite include I/O operations (read and write), data locality (arrays of basic data types and arrays of structures), array dimensionality (one‐dimensional arrays, two‐dimensional meshes, three‐dimensional cubes), I/O modes (synchronous and asynchronous). In this paper, we present the observed performance of h5bench executed along several of these dimensions on existing supercomputers (Cori and Summit) and pre‐exascale platforms (Perlmutter, Theta, and Polaris). h5bench measurements can be used to identify performance bottlenecks and their root causes and evaluate I/O optimizations. As the I/O patterns of h5bench are diverse and capture the I/O behaviors of various HPC applications, this study will be helpful to the broader supercomputing and I/O community.

97 MATHEMATICS AND COMPUTING↗

How Accurate Are Simulations and Experiments for the Lattice Energies of Molecular Crystals?

Molecular crystals play a central role in a wide range of scientific fields, including pharmaceuticals and organic semiconductor devices. However, they are challenging systems to model accurately with computational approaches because of a delicate interplay of intermolecular interactions such as hydrogen bonding and Van der Waals dispersion forces. Here, by exploiting recent algorithmic developments, we report the first set of diffusion Monte Carlo lattice energies for all 23 molecular crystals in the popular and widely used X23 dataset. Comparisons with previous state-of-the-art lattice energy predictions (on a subset of the dataset) and a careful analysis of experimental sublimation enthalpies reveals that high-accuracy computational methods are now at least as reliable as (computationally derived) experiments for the lattice energies of molecular crystals. Overall, this work demonstrates the feasibility of high-level explicitly correlated electronic structure methods for broad benchmarking studies in complex condensed phase systems, and signposts a route towards closer agreement between experiment and simulation. Published by the American Physical Society 2024

Physics↗

Sensitivity Analysis of Multidisciplinary Rotorcraft Simulations

A multidisciplinary sensitivity analysis of rotorcraft simulations involving tightly coupled high-fidelity computational fluid dynamics and comprehensive analysis solvers is presented and evaluated. An unstructured sensitivity-enabled Navier-Stokes solver, FUN3D, and a nonlinear flexible multibody dynamics solver, DYMORE, are coupled to predict the aerodynamic loads and structural responses of helicopter rotor blades. A discretely-consistent adjoint-based sensitivity analysis available in FUN3D provides sensitivities arising from unsteady turbulent flows and unstructured dynamic overset meshes, while a complex-variable approach is used to compute DYMORE structural sensitivities with respect to aerodynamic loads. The multidisciplinary sensitivity analysis is conducted through integrating the sensitivity components from each discipline of the coupled system. Numerical results verify accuracy of the FUN3D/DYMORE system by conducting simulations for a benchmark rotorcraft test model and comparing solutions with established analyses and experimental data. Complex-variable implementation of sensitivity analysis of DYMORE and the coupled FUN3D/DYMORE system is verified by comparing with real-valued analysis and sensitivities. Correctness of adjoint formulations for FUN3D/DYMORE interfaces is verified by comparing adjoint-based and complex-variable sensitivities. Finally, sensitivities of the lift and drag functions obtained by complex-variable FUN3D/DYMORE simulations are compared with sensitivities computed by the multidisciplinary sensitivity analysis, which couples adjoint-based flow and grid sensitivities of FUN3D and FUN3D/DYMORE interfaces with complex-variable sensitivities of DYMORE structural responses.

Wang, Li↗

A Quantum Volume Metric for Qudit Quantum Computers

Qudit-based quantum computers, such as the high-Q superconducting radio frequency resonator cavities being pursued by the SQMS Center hosted at Fermilab, may provide computational advantages relative to qubit-based machines for some problems. However, the novel system’s hardware and software will present different errors that must be understood and mitigated, along with other unforeseen inefficiencies across the entire stack. It is essential to benchmark these qudit platforms against existing systems. We propose such a benchmark by generalizing the qubit-based quantum volume metric to qudit devices. In particular, we show that the performance depends on the available gateset used to compile unitary operations, as well as realistic noise sources.

Quantum Computing↗

Neural Ordinary Differential Equations for Nonlinear System Identification

Neural ordinary differential equations (NODE) have been recently proposed as a promising approach for nonlinear system identification tasks. In this work, we systematically compare their predictive performance with current state-of-the-art nonlinear and classical linear methods. In particular, we present a quantitative study comparing NODE's performance against neural state-space models and classical linear system identification methods and evaluate their inference speed and prediction performance on open-loop errors across eight different dynamical systems. The experiments show that NODEs can consistently improve the prediction accuracy by order of magnitude compared to benchmark methods. Besides improved accuracy, we also observed that NODEs are less sensitive to hyperparameters compared to neural state-space models by paying the cost of increased computation at the inference time.

machine leaning, system identification, physics in↗

Assessment of Grizzly Capabilities for Reactor Pressure Vessels and Reinforced Concrete Structures

Over the last several years, capabilities to simulate the progression and effects of degradation in critical structures in light water reactor (LWR) nuclear power plants have been under development in the Grizzly code. Age-related material degradation is important for a number of systems in LWRs, but the main focus for Grizzly development has been on reactor pressure vessels (RPVs) and concrete structures because of their central role and the difficulty of replacement of these structures if they are found to be degraded to an unacceptable degree. The capabilities for analyzing both RPVs and concrete structures in Grizzly have reached a point where the feature sets are sufficiently complete to perform credible analyses of those types of structures. Because of this, the emphasis has shifted from foundational development to assessing the accuracy of these modeling capabilities on representative problems of interest and on improving the physical basis of those models to improve their ability to predict the actual response of those structural systems. This report documents a first set of test cases that have been developed to assess the Grizzly code on real-world problems for these two types of structures. For RPVs, these test cases consist of a set of benchmark problems where Grizzly is compared to another code. For concrete structures, a set of models of experimental specimens designed to characterize the multidimensional swelling response of reinforced concrete members due to alkali-silica reaction (ASR) has been developed and compared with experimental results. Both the RPV and concrete test cases developed here are much more computationally intensive than the small regression tests that are used in the testing that Grizzly undergoes every time a proposed set of changes is made to ensure that those changes do not adversely affect previously established behavior. A system for regularly running these large problems to monitor the behavior of the code as it is developed has been instituted. While these test cases generally indicate good comparison with the benchmark results, they also indicate areas where further development is warranted.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Ground-based PIV and numerical flow visualization results from the surface tension driven convection experiment

The Surface Tension Driven Convection Experiment (STDCE) is a Space Transportation System flight experiment to study both transient and steady thermocapillary fluid flows aboard the United States Microgravity Laboratory-1 (USML-1) Spacelab mission planned for June, 1992. One of the components of data collected during the experiment is a video record of the flow field. This qualitative data is then quantified using an all electric, two dimensional Particle Image Velocimetry (PIV) technique called Particle Displacement Tracking (PDT), which uses a simple space domain particle tracking algorithm. Results using the ground based STDCE hardware, with a radiant flux heating mode, and the PDT system are compared to numerical solutions obtained by solving the axisymmetric Navier Stokes equations with a deformable free surface. The PDT technique is successful in producing a velocity vector field and corresponding stream function from the raw video data which satisfactorily represents the physical flow. A numerical program is used to compute the velocity field and corresponding stream function under identical conditions. Both the PDT system and numerical results were compared to a streak photograph, used as a benchmark, with good correlation.

Pline, Alexander D.↗

Performance Evaluation and Modeling Techniques for Parallel Processors

In practice, the performance evaluation of supercomputers is still substantially driven by singlepoint estimates of metrics (e.g., MFLOPS) obtained by running characteristic benchmarks or workloads. With the rapid increase in the use of time-shared multiprogramming in these systems, such measurements are clearly inadequate. This is because multiprogramming and system overhead, as well as other degradations in performance due to time varying characteristics of workloads, are not taken into account. In multiprogrammed environments, multiple jobs and users can dramatically increase the amount of system overhead and degrade the performance of the machine. Performance techniques, such as benchmarking, which characterize performance on a dedicated machine ignore this major component of true computer performance. Due to the complexity of analysis, there has been little work done in analyzing, modeling, and predicting the performance of applications in multiprogrammed environments. This is especially true for parallel processors, where the costs and benefits of multi-user workloads are exacerbated. While some may claim that the issue of multiprogramming is not a viable one in the supercomputer market, experience shows otherwise. Even in recent massively parallel machines, multiprogramming is a key component. It has even been claimed that a partial cause of the demise of the CM2 was the fact that it did not efficiently support time-sharing. In the same paper, Gordon Bell postulates that, multicomputers will evolve to multiprocessors in order to support efficient multiprogramming. Therefore, it is clear that parallel processors of the future will be required to offer the user a time-shared environment with reasonable response times for the applications. In this type of environment, the most important performance metric is the completion of response time of a given application. However, there are a few evaluation efforts addressing this issue.

Dimpsey, Robert Tod↗

Ground-based PIV and numerical flow visualization results from the Surface Tension Driven Convection Experiment

The Surface Tension Driven Convection Experiment (STDCE) is a Space Transportation System flight experiment to study both transient and steady thermocapillary fluid flows aboard the United States Microgravity Laboratory-1 (USML-1) Spacelab mission planned for June, 1992. One of the components of data collected during the experiment is a video record of the flow field. This qualitative data is then quantified using an all electric, two dimensional Particle Image Velocimetry (PIV) technique called Particle Displacement Tracking (PDT), which uses a simple space domain particle tracking algorithm. Results using the ground based STDCE hardware, with a radiant flux heating mode, and the PDT system are compared to numerical solutions obtained by solving the axisymmetric Navier Stokes equations with a deformable free surface. The PDT technique is successful in producing a velocity vector field and corresponding stream function from the raw video data which satisfactorily represents the physical flow. A numerical program is used to compute the velocity field and corresponding stream function under identical conditions. Both the PDT system and numerical results were compared to a streak photograph, used as a benchmark, with good correlation.

Pline, Alexander D.↗