Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “computer system benchmarking”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19

Verification and Performance Impact of the New Parallel MCNP6.3 Particle Track Output Capability for Subcritical Multiplication Simulations

The MCNP6® code, version 6.3, has several new features that are intended to ultimately replace legacy features that are now marked for deprecation. One of these features is the new particle track output (PTRAC) format and capability, where the legacy PTRAC capability still exists alongside the modern PTRAC capability in MCNP6.3. While the MCNP6.3 code has been extensively verified and validated for many applications, the PTRAC feature is not exercised in any of the typical verification and validation (V&V) applications studied during the course of a typical MCNP code release. The primary goal of this paper is to verify that the legacy and modern PTRAC feature produces equivalent results for subcritical multiplication benchmarks previously studied. In the process of verifying that the simulated benchmark results are equivalent, the computational performance is compared between the legacy and modern PTRAC uses. In addition to verification of the update, which is important to the community as a whole, this effort also supports advances in the simulation of recent subcritical neutron noise measurements that require higher computational effort per second of real-time measurement than that of systems typically measured.

97 MATHEMATICS AND COMPUTING↗

Implementation and verification of PyNE R2S with DAG-OpenMC

The mesh-based Rigorous-Two-Step (R2S) method has been widely used in the accurate estimation of the shutdown dose rate (SDR) of fusion systems. Several mesh-based R2S code has been implemented based on Monte Carlo particle transport code MCNP5 or DAG-MCNP5 and inventory calculation code such as FISPACT-II, ALARA or ACAB. Recently, the OpenMC, a community-developed Monte Carlo particle transport code, implemented the photon particle transport capability, which enables OpenMC to be used in R2S. In this paper, PyNE R2S is extended with the support of DAG-OpenMC, i.e., OpenMC with DAGMC. Modifications that allow PyNE to read the flux result in the state point file of OpenMC has been implemented to PyNE R2S workflow. A custom source routine of OpenMC that reads the photon source file from PyNE R2S and sampling photon particle in OpenMC photon transport has been implemented in OpenMC. With these modifications, PyNE R2S is now capable of performing R2S calculations with both DAG-MCNP5 and DAG-OpenMC. The ITER FNG dose rate benchmark problem has been used to validate the code. The FNG neutron source has been modified for neutron transport with DAG-OpenMC. The shutdown dose rate of 19 cooling times has been calculated and compared with the experimental data and the computational results of other R2S codes. The results of PyNE R2S with DAG-OpenMC show satisfactory agreement with the experimental data and other computational results. Therefore, we consider PyNE R2S with DAG-OpenMC is a reliable code to calculate the SDR of fusion systems.

fusion↗

FY23 Progress on Computational Modeling of the Water-Based NSTF

This report summarizes the system-level modeling effort by Argonne National Laboratory (Argonne) of the Natural convection Shutdown heat removal Test Facility (NSTF) in FY23. As a continuation of the modeling effort from FY22, this year’s work focuses on improving the RELAP5-3D model developed previously for two-phase flow simulations. The RELAP5-3D model is updated to more accurately capture the heat loss experienced by the facility. The updated model is compared against experimental data for benchmarking purposes of the RELAP5-3D input model. By correctly accounting for heat loss, the updated RELAP5-3D model can now predict the two-phase baseline case more accurately. The onset and the duration of instability are captured well by the model. Furthermore, analyses are performed to better understand the instability mechanism experienced by the flow where the expansion of the boiling boundary in the chimney is studied in details and the fundamental frequencies of the oscillations are obtained. The updated RELAP5-3D model is further compared against four fault conditions, namely the reduction of riser header inlet flow area, depletion of system inventory, blocked riser channels, and static boiling scenario. For each fault condition, minor modifications and tuning are performed to improve the predictions of the model. The purpose of the analyses is to investigate the capability of RELAP5-3D in predicting complex two-phase flows in possible accident scenarios in actual Reactor Cavity Cooling System (RCCS). Overall, the model is able to capture the behaviors and trends of these fault conditions relatively well. Some discrepancies remain between the experimental data and the predictions, many of which are likely due to the differences in the predicted and experimental vapor generation rate. Future work will focus on the continued development of the current RELAP5-3D input model of the NSTF to both improve the accuracy of the model’s predictive capability and continue supporting the experimental program needs. The mutually beneficial relationship between analysis and experimental efforts has become integral to the parent NSTF program, and the greater objective to fully understand and accurately predict the heat removal performance of a full scale RCCS concept.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

A time-parallel method for scalable heat transfer simulations of additive manufacturing

Here, a major challenge in simulating the thermal behavior in additive manufacturing processes is the disparate length and time scales between transport phenomena occurring in the melt pool and the component. A common simulation approach relies on spatial decomposition for parallel computing, but due to the nature of heat transfer in AM, where most of the computational expenditure is localized near the melt pool, the computational speedup from spatial parallelization saturates quickly. Therefore, additional parallelism by means of time-domain decomposition is needed to fully take advantage of high-performance computing (HPC) resources. This work introduces a time-parallel method to improve the computational scalability of additive manufacturing simulations on HPC systems, while maintaining high temporal resolution of heat transfer near the melt pool. The method, inspired by the nonlinear paraexp formalism, performs an iterative superposition of nonlinear solutions to the initial value problem, integrating the heat equation across overlapping time-parallel intervals. For a single layer of the NIST AMB2018–01 L7 benchmark problem, the method achieves a 38.51x speedup in wall-clock time with a maximum error in the global temperature solution of 0.99%. This reduces the total solution time from 196.72 min to 5.11 min on 128 nodes of the ORNL Frontier supercomputer. The tradeoff between accuracy and total wall-clock time is investigated and recommendations for time-parallel deployment for AM problems are made.

Additive manufacturing↗

Development of a Pulsed Slowing-Down-Time Benchmark of Neutron Thermalization in Graphite

Graphite is a classic neutronic material that has been used as both a reactor reflector and moderator in various nuclear reactor systems. The ability to accurately predict the slowing down and thermalization of neutrons in graphite can have significant implications on the safety and operation of such reactor systems. In reactors, the neutron thermalization process is quantified using the thermal scattering law (TSL) and related cross sections for a given moderator. An ideal approach to assess the validity of TSL data is using benchmark measurement based on the pulsed Slowing-Down-Time technique and its comparison with the appropriate graphite nuclear library. In this work, experimental measurement and computational Monte Carlo simulations were performed to benchmark the slowing down characteristics and thermalization of neutrons in nuclear (reactor-grade) graphite. Given the density of graphite, various graphite libraries were selected for the benchmark analysis.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Simple and efficient algorithms for training machine learning potentials to force data

Machine learning models, trained on data from ab initio quantum simulations, are yielding molecular dynamics potentials with unprecedented accuracy. One limiting factor is the quantity of available training data, which can be expensive to obtain. A quantum simulation often provides all atomic forces, in addition to the total energy of the system. These forces provide much more information than the energy alone. It may appear that training a model to this large quantity of force data would introduce significant computational costs. Actually, training to all available force data should only be a few times more expensive than training to energies alone. Here, we present a new algorithm for efficient force training, and benchmark its accuracy by training to forces from real-world datasets for organic chemistry and bulk aluminum.

74 ATOMIC AND MOLECULAR PHYSICS↗

Towards interpretable Cryo-EM: disentangling latent spaces of molecular conformations

Molecules are essential building blocks of life and their different conformations (i.e., shapes) crucially determine the functional role that they play in living organisms. Cryogenic Electron Microscopy (cryo-EM) allows for acquisition of large image datasets of individual molecules. Recent advances in computational cryo-EM have made it possible to learn latent variable models of conformation landscapes. However, interpreting these latent spaces remains a challenge as their individual dimensions are often arbitrary. The key message of our work is that this interpretation challenge can be viewed as an Independent Component Analysis (ICA) problem where we seek models that have the property of identifiability. That means, they have an essentially unique solution, representing a conformational latent space that separates the different degrees of freedom a molecule is equipped with in nature. Thus, we aim to advance the computational field of cryo-EM beyond visualizations as we connect it with the theoretical framework of (nonlinear) ICA and discuss the need for identifiable models, improved metrics, and benchmarks. Moving forward, we propose future directions for enhancing the disentanglement of latent spaces in cryo-EM, refining evaluation metrics and exploring techniques that leverage physics-based decoders of biomolecular systems. Moreover, we discuss how future technological developments in time-resolved single particle imaging may enable the application of nonlinear ICA models that can discover the true conformation changes of molecules in nature. The pursuit of interpretable conformational latent spaces will empower researchers to unravel complex biological processes and facilitate targeted interventions. This has significant implications for drug discovery and structural biology more broadly. More generally, latent variable models are deployed widely across many scientific disciplines. Thus, the argument we present in this work has much broader applications in AI for science if we want to move from impressive nonlinear neural network models to mathematically grounded methods that can help us learn something new about nature.

59 BASIC BIOLOGICAL SCIENCES↗

An MLCommons Scientific Benchmarks Ontology

Scientific machine learning research spans diverse domains and data modalities, yet existing benchmark efforts remain siloed and lack standardization. This makes novel and transformative applications of machine learning to critical scientific use-cases more fragmented and less clear in pathways to impact. This paper introduces an ontology for scientific benchmarking developed through a unified, community-driven effort that extends the MLCommons ecosystem to cover physics, chemistry, materials science, biology, climate science, and more. Building on prior initiatives such as XAI-BENCH, FastML Science Benchmarks, PDEBench, and the SciMLBench framework, our effort consolidates a large set of disparate benchmarks and frameworks into a single taxonomy of scientific, application, and system-level benchmarks. New benchmarks can be added through an open submission workflow coordinated by the MLCommons Science Working Group and evaluated against a six-category rating rubric that promotes and identifies high-quality benchmarks, enabling stakeholders to select benchmarks that meet their specific needs. The architecture is extensible, supporting future scientific and AI/ML motifs, and we discuss methods for identifying emerging computing patterns for unique scientific workloads. The MLCommons Science Benchmarks Ontology provides a standardized, scalable foundation for reproducible, cross-domain benchmarking in scientific machine learning. A companion webpage for this work has also been developed as the effort evolves: https://mlcommons-science.github.io/benchmark/

Hawks, Ben [Fermilab] (ORCID:0000000157000288)↗

Simulating molecular polaritons in the collective regime using few-molecule models

The study of molecular polaritons beyond simple quantum emitter ensemble models (e.g., Tavis–Cummings) is challenging due to the large dimensionality of these systems and the complex interplay of molecular electronic and nuclear degrees of freedom. This complexity constrains existing models to either coarse-grain the rich physics and chemistry of the molecular degrees of freedom or artificially limit the description to a small number of molecules. In this work, we exploit permutational symmetries to drastically reduce the computational cost of ab initio quantum dynamics simulations for large N . Furthermore, we discover an emergent hierarchy of timescales present in these systems, that justifies the use of an effective single molecule to approximately capture the dynamics of the entire ensemble, an approximation that becomes exact as N → ∞. We also systematically derive finite N corrections to the dynamics and show that addition of k extra effective molecules is enough to account for phenomena whose rates scale as 𝒪( N − k ). Based on this result, we discuss how to seamlessly modify existing single-molecule strong coupling models to describe the dynamics of the corresponding ensemble. We call this approach collective dynamics using truncated equations (CUT-E), benchmark it against well-known results of polariton relaxation rates, and apply it to describe a universal cavity-assisted energy funneling mechanism between different molecular species. Beyond being a computationally efficient tool, this formalism provides an intuitive picture for understanding the role of bright and dark states in chemical reactivity, necessary to generate robust strategies for polariton chemistry.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A Particle-in-Cell Method for Plasmas with a Generalized Momentum Formulation, Part I: Model Formulation

Here, this paper formulates a new particle-in-cell method for the Vlasov–Maxwell system. Under the Lorenz gauge condition, Maxwell’s equations for the electromagnetic fields can be written as a collection of scalar and vector wave equations. The use of potentials for the fields motivates the adoption of a Hamiltonian formulation for particles that employs the generalized (conjugate) momentum. A notable advantage offered by the Hamiltonian formulation is the elimination of time derivatives in the Lorenz gauge formulation that are required by the standard Newton–Lorentz treatment of the particles. This allows the fields to retain the full time-accuracy guaranteed by the field solver. The resulting updates for particles require only knowledge of the fields and their spatial derivatives. An analytical method for constructing these spatial derivatives is presented that exploits the underlying integral solution used in the field solver for the wave equations. Moreover, these derivatives are demonstrated to converge at the same rate as the fields in both time and space. The Method of Lines Transpose field solver we consider in this work is globally first-order accurate in time and high-order accurate in space (e.g., fourth- and fifth-order) and belongs to a larger class of methods which are unconditionally stable, can address geometry, and leverage $\mathcal {O}(N)$ fast summation methods for efficiency. We demonstrate the method on several well-established benchmark problems on bounded domains, including a plasma sheath as well as a relativistic particle beam. The efficacy of the proposed formulation is established by comparing with a second-order accurate finite-difference time-domain method that employs a leapfrog time advance for particles and a charge conserving map suitable for bounded domains. The new method shows mesh-independent numerical heating properties even in cases where the plasma Debye length is smaller than the grid spacing. This is an important feature of the new method for problems defined on bounded domains, because it permits the use of coarser grids in space in the representation of the fields. Such a capability has significant implications for the simulation of plasmas in bounded domains with complex geometry, where the ratio between the largest and smallest cells can vary significantly. The use of high-order spatial approximations in the new method also means that fewer grid points are required in order to achieve a fixed accuracy. Our results also suggest that the new method can be used with fewer simulation particles per cell compared to the benchmark explicit method, which permits further computational savings.

97 MATHEMATICS AND COMPUTING↗

Bidding Curve Design for Hybrid Power Plants with Uncertain Solar Forecast

This paper presents a novel bidding curve design algorithm tailored for hybrid power plants (HPPs) to participate in the wholesale electricity market. Utilizing forecasts for photovoltaic (PV) generation and available battery power, our algorithm strategically computes the bidding curve to maximize HPP profit while adeptly managing the inherent uncertainty associated with PV power generation. In addition, the introduction of the penalty cost in HPP bidding curves provides the system operator a tool to effectively manage the system-level uncertainty that caused by HPPs. Numerical analysis through Monte Carlo simulations confirms that our bidding curve methodology outperforms the benchmark across various scenarios.

bidding curve↗

Iteration-based Linearized Distribution-level Locational Marginal Price for Three-phase Unbalanced Distribution Systems

Distributed energy resources (DERs) are rocking the utilities’ business landscape. It calls for competitive market environments that incentivize DERs to form maximum operating efficiency. Among proposed pricing schemes, distribution-level locational marginal price (DLMP) is effective in signaling the marginal generation cost differences driven by energy losses and network constraints. It can be derived from a distribution-level optimal power flow (OPF) framework, as it essentially presents the sensitivity of optimized generation cost towards incremental loads. However, due to the high resistance-to-inductance ratio and unbalanced characteristics of distribution networks, computational affordable DLMPs are highly challenged. This article provides a linear-approximated DLMP that can be solved efficiently and generalized to account for reactive power flow, three-phase unbalanced loads and meshed network structure. The successive linear programming technique is introduced to enhance the model accuracy. Case studies on an IEEE 123-Bus system validate its accuracy against a nonlinear benchmark and capability in offering proper incentives.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Bidding Curve Design for Hybrid Power Plants with Uncertain Solar Forecast: Preprint

This paper presents a novel bidding curve design algorithm tailored for hybrid power plants (HPPs) to participate in the wholesale electricity market. Utilizing forecasts for photovoltaic (PV) generation and available battery power, our algorithm strategically computes the bidding curve to maximize HPP profit while adeptly managing the inherent uncertainty associated with PV power generation. In addition, the introduction of the penalty cost in HPP bidding curves provides the system operator a tool to effectively manage the system-level uncertainty that caused by HPPs. Numerical analysis through Monte Carlo simulations confirms that our bidding curve methodology outperforms the benchmark across various scenarios.

bidding curve↗

Manifold Learning-Based Polynomial Chaos Expansions for High-Dimensional Surrogate Models

In this work we introduce a manifold learning-based method for uncertainty quantification (UQ) in systems describing complex spatiotemporal processes. Our first objective is to identify the embedding of a set of high-dimensional data representing quantities of interest of the computational or analytical model. For this purpose, we employ Grassmannian diffusion maps, a two-step nonlinear dimension reduction technique which allows us to reduce the dimensionality of the data and identify meaningful geometric descriptions in a parsimonious and inexpensive manner. Polynomial chaos expansion is then used to construct a mapping between the stochastic input parameters and the diffusion coordinates of the reduced space. An adaptive clustering technique is proposed to identify an optimal number of clusters of points in the latent space. The similarity of points allows us to construct a number of geometric harmonic emulators which are finally utilized as a set of inexpensive pretrained models to perform an inverse map of realizations of latent features to the ambient space and thus perform accurate out-of-sample predictions. Thus, the proposed method acts as an encoder-decoder system which is able to automatically handle very high-dimensional data while simultaneously operating successfully in the small-data regime. The method is demonstrated on two benchmark problems and on a system of advection-diffusion-reaction equations which model a first-order chemical reaction between two species. In all test cases, the proposed method is able to achieve highly accurate approximations which ultimately lead to the significant acceleration of UQ tasks.

42 ENGINEERING↗

Regen: An object layout regenerator on large-scale production HPC systems

This article proposes an object layout regenerator called Regen which regenerates and removes the object layout dynamically to improve the read performance of applications. Regen first detects frequent access patterns from the I/O requests of the applications. Second, Regen reorganizes the objects and regenerates or preallocates new object layouts according to the identified access patterns. Finally, Regen removes or reuses the obsolete or regenerated object layouts as necessary. As a result, Regen accelerates access to objects by providing a flexible object layout. We implement Regen as a framework on top of Proactive Data Container (PDC) and evaluate it on Cori supercomputer, a production-scale HPC system, by using realistic HPC I/O benchmarks. The experimental results show that Regen improves the I/O performance by up to 16.92 × compared with an existing system.

Distributed file system↗

Report on G4-Med, a Geant4 benchmarking system for medical physics applications developed by the Geant4 Medical Simulation Benchmarking Group

Geant4 is a Monte Carlo code extensively used in medical physics for a wide range of applications, such as dosimetry, micro- and nanodosimetry, imaging, radiation protection, and nuclear medicine. Geant4 is continuously evolving, so it is crucial to have a system that benchmarks this Monte Carlo code for medical physics against reference data and to perform regression testing. In this work, to respond to these needs, we developed G4-Med, a benchmarking and regression testing system of Geant4 for medical physics. G4-Med currently includes 18 tests. They range from the benchmarking of fundamental physics quantities to the testing of Monte Carlo simulation setups typical of medical physics applications. Both electromagnetic and hadronic physics processes and models within the prebuilt Geant4 physics lists are tested. The tests included in G4-Med are executed on the CERN computing infrastructure via the use of the geant-val web application, developed at CERN for Geant4 testing. The physical observables can be compared to reference data for benchmarking and to results of previous Geant4 versions for regression testing purposes. This paper describes the tests included in G4-Med and shows the results derived from the benchmarking of Geant4 10.5 against reference data.

60 APPLIED LIFE SCIENCES↗

Evaluation of Best Practices in Mitigating Startup Costs on Leadership-Class Supercomputers

Supercomputers at Department of Energy (DOE) National Laboratories face a widening range of workloads, from traditional modeling and simulation to Artificial Intelligence model training or complex multi-stage workflows, and beyond. At DOE Leadership Computing Facilities like the Oak Ridge Leadership Computing Facility (OLCF), these workloads demand concurrent access to large portions of the supercomputer’s resources. Launching a job across massive supercomputers is challenging from the start; the file system struggles with a large backlog of metadata requests as tens of thousands of processes read thousands of the same files, and the compute job cannot start until this is completed. There are multiple existing approaches to calm this metadata storm, ranging from vendor-developed tools like sbcast to National Laboratory-developed tools like Spindle and Copper. In this paper, we benchmark and discuss three common approaches to improving compute job launch latencies on Frontier: Slurm’s sbcast tool, Spindle, and Copper. We evaluate these tools by measuring the launch latencies of four workloads: OSU Microbenchmark’s osu_init, Pynamic, Python import mpi4py, and Python import torch. We provide discussion of the results, highlighting data that meet expectations and that do not meet expectations.

Hagerty, Nick [ORNL] (ORCID:0000000330014414)↗

Progress Towards the Validation of a new RELAP5-3D model of the High Temperature Test Facility

Validation is a key step in the development of any type of systems model. As the next generation of reactors approaches, the need for codes that have been validated for these new types of systems continues to grow. An example of a prominent option is the Reactor Excursion Leak Analysis Program (RELAP5-3D), developed by Idaho National Laboratory. This code was developed for the purpose of systems level thermal-hydraulic modeling of light water reactors (LWRS) and postulated transients that can occur in LWRS.RELAP5-3D has been substantially validated against LWR data. Due to its long history as a reactor safety analysis tool, there has been an effort to adapt RELAP5-3D for the purposes of advanced reactor concepts such as prismatic high-temperature gas-cooled reactors (HTGRs). However, RELAP5-3D has not nearly been validated and verified for HTGRs to the degree of LWRs, warranting verification and validation opportunities with computational benchmarks and existing experimental facilities. Examples of such facilities include the modular high-temperature gas-cooled reactor (MHTGR) 350 and the high temperature engineering test reactor (HTTR) from Japan. The MHTGR 350 is a benchmark design concept for code-to-code verification purposes; therefore, it does not provide any experimental data for validation opportunities The HTTR provides useful multiphysics validation data but does not have the in-core instruments to generate thermal-hydraulic experimental data to help with RELAP5-3D validation. Consequently, a facility that could provide key in-core temperatures for thermal-hydraulic validation was still needed. The High Temperature Test Facility (HTTF) is an integral effects facility for HTGR thermal hydraulics developed and operated by Oregon State University. HTTF represents ¼ length scale of the General Atomics MHTGR and is rated for a total power of 2.2 MW. Axially, the core consists of an upper and lower reflector and 10 blocks, numbered from bottom to top (Block 1 is right above lower reflector). The core is heated via graphite resistive heater rods, with respective channels distributed throughout the core. The primary coolant is helium and heat can radiate out of the core to the reactor cavity cooling system (RCCS), which is cooled by water. The primary purpose of the facility is to investigate pressurized conduction cooldown (PCC) and depressurized conduction cooldown (DCC) transients, which are also referred to as the pressurized and depressurized loss of forced cooling respectively. Two experiments were chosen to perform the validation study with a RELAP5-3D model of HTTF. These experiments are PG-27 (PCC) and PG-29 (DCC). These were chosen based off of the quality of available experimental data before and during the experiment which led to their inclusion in the HTGR Thermal Hydraulics Benchmark.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗