Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “computer system benchmarking”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20

Progress Towards the Validation of a new RELAP5-3D model of the High Temperature Test Facility

Validation is a key step in the development of any type of systems model. As the next generation of reactors approaches, the need for codes that have been validated for these new types of systems continues to grow. An example of a prominent option is the Reactor Excursion Leak Analysis Program (RELAP5-3D), developed by Idaho National Laboratory. This code was developed for the purpose of systems level thermal-hydraulic modeling of light water reactors (LWRS) and postulated transients that can occur in LWRS.RELAP5-3D has been substantially validated against LWR data. Due to its long history as a reactor safety analysis tool, there has been an effort to adapt RELAP5-3D for the purposes of advanced reactor concepts such as prismatic high-temperature gas-cooled reactors (HTGRs). However, RELAP5-3D has not nearly been validated and verified for HTGRs to the degree of LWRs, warranting verification and validation opportunities with computational benchmarks and existing experimental facilities. Examples of such facilities include the modular high-temperature gas-cooled reactor (MHTGR) 350 and the high temperature engineering test reactor (HTTR) from Japan. The MHTGR 350 is a benchmark design concept for code-to-code verification purposes; therefore, it does not provide any experimental data for validation opportunities The HTTR provides useful multiphysics validation data but does not have the in-core instruments to generate thermal-hydraulic experimental data to help with RELAP5-3D validation. Consequently, a facility that could provide key in-core temperatures for thermal-hydraulic validation was still needed. The High Temperature Test Facility (HTTF) is an integral effects facility for HTGR thermal hydraulics developed and operated by Oregon State University. HTTF represents ¼ length scale of the General Atomics MHTGR and is rated for a total power of 2.2 MW. Axially, the core consists of an upper and lower reflector and 10 blocks, numbered from bottom to top (Block 1 is right above lower reflector). The core is heated via graphite resistive heater rods, with respective channels distributed throughout the core. The primary coolant is helium and heat can radiate out of the core to the reactor cavity cooling system (RCCS), which is cooled by water. The primary purpose of the facility is to investigate pressurized conduction cooldown (PCC) and depressurized conduction cooldown (DCC) transients, which are also referred to as the pressurized and depressurized loss of forced cooling respectively. Two experiments were chosen to perform the validation study with a RELAP5-3D model of HTTF. These experiments are PG-27 (PCC) and PG-29 (DCC). These were chosen based off of the quality of available experimental data before and during the experiment which led to their inclusion in the HTGR Thermal Hydraulics Benchmark.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

A Model-Free Voltage Control Approach to Mitigate Motor Stalling and FIDVR for Smart Grids

Electric power networks are large and highly nonlinear dynamical systems that present unique challenges to control design. Though there is a large number of dynamic models for power system stability and control, many models are only useful with right assumptions and wrong for other tasks. Moreover, the dynamic behavior of the grid is increasingly complex under the banner of smart grids. These lead to the difficulty of developing appropriate dynamic modeling, and thus an efficient control strategy. To avoid such modeling challenges, here we present a novel dynamic voltage control strategy based on a model-free control (MFC) approach, requiring no modeling procedure. In particular, it focuses on fault-induced delayed voltage recovery (FIDVR) events, which require complex and accurate dynamic load models to replicate such events. This work utilizes MFC as an online controller to achieve the desired voltage stability under the FIDVR event. The proposed MFC strategy allows simple implementation and low computational cost for efficient mitigation of FIDVR. For benchmarking, a reasonably accurate dynamic performance model is explored. Simulation results with the IEEE 57 bus test network demonstrate the enhanced dynamic voltage profile for load buses having induction motors with the support of reactive power resources.

24 POWER TRANSMISSION AND DISTRIBUTION↗

A massively parallel and scalable multi-CPU material point method

Harnessing the power of modern multi-GPU architectures, we present a massively parallel simulation system based on the Material Point Method (MPM) for simulating physical behaviors of materials undergoing complex topological changes, self-collision, and large deformations. Our system makes three critical contributions. First, we introduce a new particle data structure that promotes coalesced memory access patterns on the GPU and eliminates the need for complex atomic operations on the memory hierarchy when writing particle data to the grid. Second, we propose a kernel fusion approach using a new Grid-to-Particles-to-Grid (G2P2G) scheme, which efficiently reduces GPU kernel launches, improves latency, and significantly reduces the amount of global memory needed to store particle data. Finally, we introduce optimized algorithmic designs that allow for efficient sparse grids in a shared memory context, enabling us to best utilize modern multi-GPU computational platforms for hybrid Lagrangian-Eulerian computational patterns. We demonstrate the effectiveness of our method with extensive benchmarks, evaluations, and dynamic simulations with elastoplasticity, granular media, and fluid dynamics. In comparisons against an open-source and heavily optimized CPU-based MPM codebase [Fang et al. 2019] on an elastic sphere colliding scene with particle counts ranging from 5 to 40 million, our GPU MPM achieves over 100x per-time-step speedup on a workstation with an Intel 8086K CPU and a single Quadro P6000 GPU, exposing exciting possibilities for future MPM simulations in computer graphics and computational science. Moreover, compared to the state-of-the-art GPU MPM method [Hu et al. 2019a], we not only achieve 2x acceleration on a single GPU but our kernel fusion strategy and Array-of-Structs-of-Array (AoSoA) data structure design also generalizes to multi-GPU systems. Our multi-GPU MPM exhibits near-perfect weak and strong scaling with 4 GPUs, enabling performant and large-scale simulations on a 10243 grid with close to 100 million particles with less than 4 minutes per frame on a single 4-GPU workstation and 134 million particles with less than 1 minute per frame on an 8-GPU workstation.

Wang, Xinlei↗

Multinode Multi-GPU Two-Electron Integrals: Code Generation Using the Regent Language

The computation of two-electron repulsion integrals (ERIs) is often the most expensive step of integral-direct self-consistent field methods. Formally it scales as O(N 4 ), where N is the number of Gaussian basis functions used to represent the molecular wave function. In practice, this scaling can be reduced to O(N 2 ) or less by neglecting small integrals with screening methods. The contributions of the ERIs to the Fock matrix are of Coulomb (J) and exchange (K) type and require separate algorithms to compute matrix elements efficiently. We previously implemented highly efficient GPU-accelerated J-matrix and K-matrix algorithms in the electronic structure code TeraChem. Although these implementations supported the use of multiple GPUs on a node, they did not support the use of multiple nodes. This presents a key bottleneck to cutting-edge ab initio simulations of large systems, e.g., excited state dynamics of photoactive proteins. We present our implementation of multinode multi-GPU J- and K-matrix algorithms in TeraChem using the Regent programming language. Regent directly supports distributed computation in a task-based model and can generate code for a variety of architectures, including NVIDIA GPUs. We demonstrate multinode scaling up to 45 GPUs (3 nodes) and benchmark against hand-coded TeraChem integral code. Finally, we also outline our metaprogrammed Regent implementation, which enables flexible code generation for integrals of different angular momenta.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Introducing the embedded random phase approximation: H 2 dissociative adsorption on Cu(111) as an exemplar

The random phase approximation (RPA) as a means of treating electron correlation recently has been shown to outperform standard density functional theory (DFT) approximations in a variety of cases. However, the computational cost of the RPA is substantially more than DFT, especially when aiming to study extended surfaces. Properly accounting for sufficient surface ensemble size, Brillouin zone sampling, and vacuum separation of periodic images in standard periodic-planewave-based DFT code raises the cost to achieve converged results. Here, we show that sub-system embedding schemes enable use of the RPA for modeling heterogeneous reactions at reduced computational cost. Further, we explore two different embedded RPA (emb-RPA) approaches, periodic emb-RPA and cluster emb-RPA. We use the (experimentally and theoretically) well-studied H 2 dissociative adsorption on Cu(111) as our exemplar, and first perform full periodic RPA calculations as a benchmark. The full RPA results match well the semi-empirical barrier fit to experimental observables and others derived from high-level computations, e.g., from recent embedded n-electron valence second order perturbation theory [Zhao et al., J. Chem. Theory Comput. 16(11), 7078–7088 (2020)] and quantum Monte Carlo [Doblhoff-Dier et al., J. Chem. Theory Comput. 13(7), 3208–3219 (2017)] simulations. Among the two emb-RPA approaches tested, the cluster emb-RPA accurately reproduces the energy profile (maximum error of 50 meV along the reaction pathway) while reducing the computational cost by approximately two orders of magnitude. We therefore expect that the embedded cluster approach will enable wider RPA implementation in heterogeneous catalysis.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

SPIKANs: separable physics-informed Kolmogorov–Arnold networks

Physics-Informed Neural Networks (PINNs) have emerged as a promising method for solving partial differential equations (PDEs) in scientific computing. While PINNs typically use multilayer perceptrons (MLPs) as their underlying architecture, recent advancements have explored alternative neural network structures. One such innovation is the Kolmogorov–Arnold Network (KAN), which has demonstrated benefits over traditional MLPs, including faster neural scaling and better interpretability. The application of KANs to physics-informed learning has led to the development of Physics-Informed KANs (PIKANs), enabling the use of KANs to solve PDEs. However, despite their advantages, KANs often suffer from slower training speeds, particularly in higher-dimensional problems where the number of collocation points grows exponentially with the dimensionality of the system. To address this challenge, we introduce Separable Physics-Informed Kolmogorov–Arnold Networks (SPIKANs). This novel architecture applies the principle of separation of variables to PIKANs, decomposing the problem such that each dimension is handled by an individual KAN. This approach drastically reduces the computational complexity of training without sacrificing accuracy, facilitating their application to higher-dimensional PDEs. Through a series of benchmark problems, we demonstrate the effectiveness of SPIKANs, showcasing their superior scalability and performance compared to PIKANs and highlighting their potential for solving complex, high-dimensional PDEs in scientific computing.

Kolmogorov-Arnold networks↗

CityBES v2021

City Buildings, Energy, and Sustainability (CityBES) is a web-based data and computing platform, focusing on energy modeling and analysis of a city's building stock to support district or city-scale building energy efficiency programs. CityBES uses an international open data standard, CityGML, to represent and exchange 3D city models. CityBES employs EnergyPlus to simulate building energy use and savings from energy efficient retrofits. Other CityBES features include energy benchmarking, district heating and cooling system modeling, rooftop PV analysis, building performance visualization, heat resilience modeling, as well as urban scale mapping of microclimate and heat vulnerability at census tract level. Different from other tools, CityBES uses integrated open and standard 3D city building data and models each individual building using EnergyPlus. CityBES can be used by urban planners, city energy managers, building owners, utilities, energy consultants and researchers.

Hong, Tianzhen↗

QTRAJ 1.0: A Lindblad equation solver for heavy-quarkonium dynamics

Here we introduce an open-source package called QTraj that solves the Lindblad equation for heavy-quarkonium dynamics using the quantum trajectories algorithm. The package allows users to simulate the suppression of heavy-quarkonium states using externally-supplied input from 3+1D hydrodynamics simulations. The code uses a split-step pseudo-spectral method for updating the wave-function between jumps, which is implemented using the open-source multi-threaded FFTW3 package. This allows one to have manifestly unitary evolution when using real-valued potentials. In this paper, we provide detailed documentation of QTraj 1.0, installation instructions, and present various tests and benchmarks of the code.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Bringing HPE Slingshot 11 support to Open MPI

The Cray HPE Slingshot 11 network is used on the new exascale systems arriving at the U.S. Department of Energy (DoE) laboratories (e.g., Frontier, Aurora, Perlmutter). As such, the support of this network is an important capability to meet the needs of exascale applications. Here, this article highlights recent work to develop supporting infrastructure to enable Open MPI to efficiently support these new platforms. A key component of this effort involves development of a new Open Fabrics Interface (OFI) provider, LinkX. We discuss the design and development of enhancements that take advantage of the new Slingshot 11 network and AMD GPUs. We include performance data from tests on the Frontier supercomputer using synthetic communication benchmarks, and the vendor provided MPI as a baseline for comparison. The tests demonstrate full functionality of Open MPI on the system and initial results show favorable performance when compared to the highly tuned vendor implementation.

97 MATHEMATICS AND COMPUTING↗

A data-driven network optimisation approach to coordinated control of distributed photovoltaic systems and smart buildings in distribution systems

The increasing integration of distributed energy resources, including demand-side resources and distributed photovoltaics (PVs), into distribution systems has resulted in more complicated power system operation. A data-driven network optimisation approach is proposed to coordinate the control of distributed PVs and smart buildings in distribution networks considering the uncertainties of solar power, outdoor temperature and heat gain associated with building thermal dynamics. These uncertain parameters have a significant impact on the operation and control of distributed PVs and smart buildings, bringing challenges to the distribution system operation. In the proposed data-driven distributionally robust optimisation (DRO) approach, the Wasserstein ball is used to construct an ambiguity set for the uncertain parameters, which does not require the probability distributions to be known. Furthermore, a conditional value-at-risk is incorporated into the Wasserstein-based DRO model and converted into a computationally tractable mixed-integer convex optimisation problem. Benchmarked with robust optimisation and chance-constrained programming, the proposed data-driven model can give a less conservative robust solution.

24 POWER TRANSMISSION AND DISTRIBUTION↗

A Deep Generative Model for Non-Intrusive Identification of EV Charging Profiles

The proliferation of electric vehicles (EVs) brings environmental benefits and technical challenges to power grids. An identification algorithm which can accurately extract individual EV charging profiles out of widely available smart meter measurements has attracted great interests. This paper proposes a non-intrusive identification framework for EV charging profile extraction, which is driven by deep generative models (DGM). First, the proposed DGM is designed as a representation layer embedded into the Markov process and used to model the joint probability distribution of available time-series data. A novel contribution is to approximate posterior distributions by neural networks whose parameters are obtained by variational inference and supervised learning. Second, the EV charging status is inferred from the DGM via dynamic programming. Lastly, the desired EV charging profile can be reconstructed by the rated power of EV models and inferred status. Compared with the benchmark Hidden Markov Models, the proposed framework can better handle noise in data with less computational complexity and better overall accuracy performances with smaller recall. The proposed framework is validated by numerical experiments on the Pecan Street dataset.

33 ADVANCED PROPULSION SYSTEMS↗

gdess: A framework for evaluating simulated atmospheric CO 2 in Earth System Models

Atmospheric carbon dioxide (CO 2 ) plays a key role in the global carbon cycle and global warming. Climate-carbon feedbacks are often studied and estimated using Earth System Models (ESMs), which couple together multiple model components—including the atmosphere, ocean, terrestrial biosphere, and cryosphere—to jointly simulate mass and energy exchanges within and between these components. Despite tremendous advances, model intercomparisons and benchmarking are aspects of ESMs that warrant further improvement (Fer et al., 2021; Smith et al., 2014). Such benchmarking is critical because comparing the value of state variables in these simulations against observed values provides evidence for appropriately refining model components; moreover, researchers can learn much about Earth system dynamics in the process (Randall et al., 2019). We introduce `gdess` (a.k.a., Greenhouse gas Diagnostics for Earth System Simulations), which parses observational datasets and ESM simulation output, combines them to be in a consistent structure, computes statistical metrics, and generates diagnostic visualizations. In its current incarnation, `gdess` facilitates evaluating a model's ability to reproduce observed temporal and spatial variations of atmospheric CO 2 . The diagnostics implemented modularly in `gdess` support more rapid assessment and improvement of model-simulated global CO 2 sources and sinks associated with land and ocean ecosystem processes. We intend for this set of automated diagnostics to form an extensible, open source framework for future comparisons of simulated and observed concentrations of various greenhouse gases across Earth system models.

97 MATHEMATICS AND COMPUTING↗

Probabilistic evolution of stochastic dynamical systems: A meso-scale perspective

Stochastic dynamical systems arise naturally across nearly all areas of science and engineering. Typically, a dynamical system model is based on some prior knowledge about the underlying dynamics of interest in which probabilistic features are used to quantify and propagate uncertainties associated with the initial conditions, external excitations, etc. From a probabilistic modeling standing point, two broad classes of methods exist, i.e. macro-scale methods and micro-scale methods. Classically, macro-scale methods such as statistical moments-based strategies are usually too coarse to capture the multi-mode shape or tails of a non-Gaussian distribution. Micro-scale methods such as random samples-based approaches, on the other hand, become computationally very challenging in dealing with high-dimensional stochastic systems. In view of these potential limitations, a meso-scale scheme is proposed here that utilizes a meso-scale statistical structure to describe the dynamical evolution from a probabilistic perspective. The significance of this statistical structure is twofold. First, it can be tailored to any arbitrary random space. Second, it not only maintains the probability evolution around sample trajectories but also requires fewer meso-scale components than the micro-scale samples. To demonstrate the efficacy of the proposed meso-scale scheme, a set of examples of increasing complexity are provided. Connections to the benchmark stochastic models as conservative and Markov models along with practical implementation guidelines are presented.

97 MATHEMATICS AND COMPUTING↗

PRISMS-PF: A general framework for phase-field modeling with a matrix-free finite element method

Abstract A new phase-field modeling framework with an emphasis on performance, flexibility, and ease of use is presented. Foremost among the strategies employed to fulfill these objectives are the use of a matrix-free finite element method and a modular, application-centric code structure. This approach is implemented in the new open-source PRISMS-PF framework. Its performance is enabled by the combination of a matrix-free variant of the finite element method with adaptive mesh refinement, explicit time integration, and multilevel parallelism. Benchmark testing with a particle growth problem shows PRISMS-PF with adaptive mesh refinement and higher-order elements to be up to 12 times faster than a finite difference code employing a second-order-accurate spatial discretization and first-order-accurate explicit time integration. Furthermore, for a two-dimensional solidification benchmark problem, the performance of PRISMS-PF meets or exceeds that of phase-field frameworks that focus on implicit/semi-implicit time stepping, even though the benchmark problem’s small computational size reduces the scalability advantage of explicit time-integration schemes. PRISMS-PF supports an arbitrary number of coupled governing equations. The code structure simplifies the modification of these governing equations by separating their definition from the implementation of the numerical methods used to solve them. As part of its modular design, the framework includes functionality for nucleation and polycrystalline systems available in any application to further broaden the phenomena that can be used to study. The versatility of this approach is demonstrated with examples from several common types of phase-field simulations, including coarsening subsequent to spinodal decomposition, solidification, precipitation, grain growth, and corrosion.

36 MATERIALS SCIENCE↗

Isomer-resolved unimolecular dynamics of the hydroperoxyalkyl intermediate (•QOOH) in cyclohexane oxidation

The oxidation of cycloalkanes is important in the combustion of transportation fuels and in atmospheric secondary organic aerosol formation. A transient carbon-centered radical intermediate (•QOOH) in the oxidation of cyclohexane is identified through its infrared fingerprint and time- and energy-resolved unimolecular dissociation dynamics to hydroxyl (OH) radical and bicyclic ether products. Although the cyclohexyl ring structure leads to three nearly degenerate •QOOH isomers (β-, γ-, and δ-QOOH), their transition state (TS) barriers to OH products are predicted to differ considerably. Selective characterization of the β-QOOH isomer is achieved at excitation energies associated with the lowest TS barrier, resulting in rapid unimolecular decay to OH products that are detected. A benchmarking approach is employed for the calculation of high-accuracy stationary point energies, in particular TS barriers, for cyclohexane oxidation (C 6 H 11 O 2 ), building on higher-level reference calculations for the smaller ethane oxidation (C 2 H 5 O 2 ) system. The isomer-specific characterization of β-QOOH is validated by comparison of experimental OH product appearance rates with computed statistical microcanonical rates, including significant heavy-atom tunneling, at energies in the vicinity of the TS barrier. Master-equation modeling is utilized to extend the results to thermal unimolecular decay rate constants at temperatures and pressures relevant to cyclohexane combustion.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Accelerated Modeling of Lithium Diffusion in Solid State Electrolytes using Artificial Neural Networks

Abstract Previous efforts to understand structure‐function relationships in high ionic conductivity materials for solid state batteries have predominantly relied on density functional theory (DFT‐) based ab initio molecular dynamics (MD). Such simulations, however, are computationally demanding and cannot be reasonably applied to large systems containing more than a hundred atoms. Here, an artificial neural network (ANN) is trained to accelerate the calculation of high accuracy atomic forces and energies used during such MD simulations. After carefully training a robust ANN for four and five element systems, nearly identical lithium ion diffusivities are obtained for Li 10 GeP 2 S 12 (LGPS) when benchmarking the ANN‐MD results with DFT‐MD. Applying the ANN‐MD approach, the effect of chlorine doping on the lithium diffusivity is calculated in an LGPS‐like structure and it is found that a dopant concentration of 1.3% maximizes ionic conductivity. The optimal concentration balances the competing consequences of effective atomic radii and dielectric constants on lithium diffusion and agrees with the experimental composition. Performing simulations at the resolution necessary to model experimentally relevant and optimal concentrations would be infeasible with traditional DFT‐MD. Systems that require a large number of simulated atoms can be studied more efficiently while maintaining high accuracy with the proposed ANN‐MD framework.

Rao, Karun K.↗

A zeroth-order active-space frozen-orbital embedding scheme for multireference calculations

Multireference computations of large-scale chemical systems are typically limited by the computational cost of quantum chemistry methods. In this work, we develop a zeroth-order active space embedding theory [ASET(0)], a simple and automatic approach for embedding any multireference dynamical correlation method based on a frozen-orbital treatment of the environment. ASET(0) is combined with the second-order multireference driven similarity renormalization group and tested on several benchmark problems, including the excitation energy of 1-octene and bond-breaking in ethane and pentyldiazene. Finally, we apply ASET(0) to study the singlet–triplet gap of p-benzyne and 9,10-anthracyne diradicals adsorbed on a NaCl surface. Our results show that despite its simplicity, ASET(0) is a powerful and sufficiently accurate embedding scheme applicable when the coupling between the fragment and the environment is in the weak to medium regime.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

BCSR on GPU: A Way Forward Extreme-scale Graph Processing on Accelerator-enabled Frontier Supercomputer

Handling large graphs in a distributed environment requires effective partitioning across processors and efficient management of local partitions. In 2D partitioning, local graphs often become too sparse, making memory-efficient data structures crucial. Using the Compressed Sparse Row (CSR) format wastes space, especially for > 83% of vertices with empty edges for the sparse graphs. This study explores bit-CSR (BCSR), a modified CSR representation, on GPUs to reduce memory usage in graph computations. We achieved 16.67% memory savings on a sparse rmat dataset with 268 million vertices and 357 million edges, without performance degradation, supported by both theoretical and experimental storage savings of 33%. However, we observed a 1.7× slowdown in degree lookup times due to bitwise operations on AMD CPUs. This analysis highlights the potential of BCSR on GPUs for improving Graph500 benchmark performance on GPU-accelerated systems, such as the Frontier supercomputer.

Sattar, Naw Safrin↗