Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “computer system benchmarking”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Frequency-domain vs time-domain SP{sub N} equations to simulate neutron noise

Inside the reactor core mechanical vibrations of fuel assemblies can produce high fluctuations around a steady-state configuration, known as neutron noise. This effect can cause the triggering of power reduction measures. Classically, diffusion theory has been used to simulate this behavior. However, this equation has some limitations if the materials of the reactor have strong variations. In this work, we use the diffusive time-dependent simplified spherical harmonics equations that improve the previous results without the necessity of using high computational requirements. In particular, two types of analyses with these equations (SP{sub 3}) are made: a frequency-domain and a time-domain. A numerical neutron noise benchmark tests the methodology and compare both formulations. First, numerical results show a good agreement between the amplitudes and phases of the SP3 equations computed with the frequency-domain and time-domain. Therefore, as the frequency-domain computation only requires to solve a linear system, it is a recommendable option for neutron noise computations. Second, one can conclude that for this type of nuclear systems, where the assemblies are not homogenized, the SP{sub 3} approximation results improve considerably the accuracy of the diffusion theory.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Analysis and Benchmarking of feature reduction for classification under computational constraints

Abstract Machine learning is most often expensive in terms of computational and memory costs due to training with large volumes of data. Current computational limitations of many computing systems motivate us to investigate practical approaches, such as feature selection and reduction, to reduce the time and memory costs while not sacrificing the accuracy of classification algorithms. In this work, we carefully review, analyze, and identify the feature reduction methods that have low costs/overheads in terms of time and memory. Then, we evaluate the identified reduction methods in terms of their impact on the accuracy, precision, time, and memory costs of traditional classification algorithms. Specifically, we focus on the least resource intensive feature reduction methods that are available in Scikit-Learn library. Since our goal is to identify the best performing low-cost reduction methods, we do not consider complex expensive reduction algorithms in this study. In our evaluation, we find that at quadratic-scale feature reduction, the classification algorithms achieve the best trade-off among competitive performance metrics. Results show that the overall training times are reduced 61%, the model sizes are reduced 6×, and accuracy scores increase 25% compared to the baselines on average with quadratic scale reduction.

97 MATHEMATICS AND COMPUTING↗

Quantum many-body calculations using body-centered cubic lattices

It is often computationally advantageous to model space as a discrete set of points forming a lattice grid. This technique is particularly useful for computationally difficult problems such as quantum many-body systems. For reasons of simplicity and familiarity, nearly all quantum many-body calculations have been performed on simple cubic lattices. Since the removal of lattice artifacts is often an important concern, it would be useful to perform calculations using more than one lattice geometry. In this paper we show how to perform quantum many-body calculations using auxiliary-field Monte Carlo simulations on a three-dimensional body-centered cubic (BCC) lattice. As a benchmark test we compute the ground state energy of 33 spin-up and 33 spin-down neutrons in the unitary limit, which is an idealized limit where the interaction range is zero and scattering length is infinite. As a fraction of the free Fermi gas energy E FG , we find that the ground state energy is E 0 /E FG =0.369(2),0.371(2), using two different definitions of the finite-system energy ratio. This is in excellent agreement with recent results obtained on a cubic lattice [He et al., Phys. Rev. A 101, 063615 (2020)]. We find that the computational effort and performance on a BCC lattice is approximately the same as that for a cubic lattice with the same number of lattice points. We discuss how the lattice simulations with different geometries can be used to constrain the size of lattice artifacts in simulations of continuum quantum many-body systems.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Finding the missing pieces: filling gaps that impede the translation of omics data into models

High-throughput omics technologies such as DNA sequencing have made the sequencing and computational assembly of microbial genomes recovered from the environment relatively routine. Computational inference of the protein products encoded by these genomes, and the associated biochemical functions, should enable the accurate prediction and modeling of microbial metabolism, organismal interactions, and ecosystem processes. However, a lack of scalable, probabilistic protein annotation tools limits the full potential of modeling for understanding the metabolism and biogeochemical cycles of microbial communities. Our approach to improve inference of protein annotations and metabolic models relied on learning from and emulating expert manual curation, leveraging software engineering and data science best practices to scale up the throughput and accuracy of annotations and metabolic model construction, building software to objectively evaluate different annotation strategies, and more closely linking the protein annotation and metabolic model inference process. Outcomes of this research include several improved or new computational tools, including DRAM (Distilled and Refined Annotation of Metabolism) for annotating microbial genomes with protein function and metabolic traits, CAMPER (Curated Annotations for Microbial Polyphenol Enzymes and Reactions) for annotating key polyphenol metabolisms, EC-Bench for comprehensive and unbiased benchmarking of annotation tools, and several apps available via the DOE Systems Biology Knowledgebase (KBase) for building genome-scale metabolic models. We demonstrate that these tools allow us to scalably annotate and understand thousands of genomes for microbial communities from a variety of systems and test cases, including rivers, thawing permafrost, and gut microbiomes. All of these computational tools are available as open-source software, with most broadly and easily accessible to the scientific community via KBase apps.

59 BASIC BIOLOGICAL SCIENCES↗

A graphics processing unit accelerated sparse direct solver and preconditioner with block low rank compression

We present the GPU implementation efforts and challenges of the sparse solver package STRUMPACK. The code is made publicly available on github with a permissive BSD license. STRUMPACK implements an approximate multifrontal solver, a sparse LU factorization which makes use of compression methods to accelerate time to solution and reduce memory usage. Multiple compression schemes based on rank-structured and hierarchical matrix approximations are supported, including hierarchically semi-separable, hierarchically off-diagonal butterfly, and block low rank. Here, in this paper, we present the GPU implementation of the block low rank (BLR) compression method within a multifrontal solver. Our GPU implementation relies on highly optimized vendor libraries such as cuBLAS and cuSOLVER for NVIDIA GPUs, rocBLAS and rocSOLVER for AMD GPUs and the Intel oneAPI Math Kernel Library (oneMKL) for Intel GPUs. Additionally, we rely on external open source libraries such as SLATE (Software for Linear Algebra Targeting Exascale), MAGMA (Matrix Algebra on GPU and Multi-core Architectures), and KBLAS (KAUST BLAS). SLATE is used as a GPU-capable ScaLAPACK replacement. From MAGMA we use variable sized batched dense linear algebra operations such as GEMM, TRSM and LU with partial pivoting. KBLAS provides efficient (batched) low rank matrix compression for NVIDIA GPUs using an adaptive randomized sampling scheme. The resulting sparse solver and preconditioner runs on NVIDIA, AMD and Intel GPUs. Interfaces are available from PETSc, Trilinos and MFEM, or the solver can be used directly in user code. We report results for a range of benchmark applications, using the Perlmutter system from NERSC, Frontier from ORNL, and Aurora from ALCF. For a high frequency wave equation on a regular mesh, using 32 Perlmutter compute nodes, the factorization phase of the exact GPU solver is about 6.5× faster compared to the CPU-only solver. The BLR-enabled GPU solver is about 13.8× faster than the CPU exact solver. For a collection of SuiteSparse matrices, the STRUMPACK exact factorization on a single GPU is on average 1.9× faster than NVIDIA’s cuDSS solver.

97 MATHEMATICS AND COMPUTING↗

Design and Performance of Kokkos Staging Space toward Scalable Resilient Application Couplings

With the growing number of applications designed for heterogeneous HPC devices, application programmers and users are finding it challenging to compose scalable workflows as ensembles of these applications, that are portable, performant and resilient. The Kokkos C++ library has been designed to simplify this cumbersome procedure by providing an intra-application uniform programming model and portable performance. However, assembling multiple Kokkos-enabled applications into a complex workflow is still a challenge. Although Kokkos enables a uniform programming model, the inter-application data exchange still remains a challenge from both performance and software development cost perspectives. In order to address this issue, we propose Kokkos data staging memory space, an extension of Kokkos' data abstraction (memory space) for heterogeneous computing systems. This new abstraction allows to express data on a virtual shared-space for multiple Kokkos applications, thus extending Kokkos to support inter-application data exchange to build an efficient application workflow. Additionally, we study the effectiveness of asynchronous data layout conversions for applications requiring different memory access patterns for the shared data. Our preliminary evaluation with a synthetic benchmark indicate the effectiveness of this conversion adapted to three different scenarios representing access frequency and use patterns of the shared data.

97 MATHEMATICS AND COMPUTING↗

Accelerating self-consistent field iterations in Kohn-Sham density functional theory using a low-rank approximation of the dielectric matrix

We present an efficient preconditioning technique for accelerating the fixed-point iteration in real-space Kohn-Sham density functional theory (DFT) calculations. The preconditioner uses a low-rank approximation of the dielectric matrix (LRDM) based on Gâteaux derivatives of the residual of fixed-point iteration along appropriately chosen direction functions. We develop a computationally efficient method to evaluate these Gâteaux derivatives in conjunction with the Chebyshev filtered subspace iteration procedure, an approach widely used in large-scale Kohn-Sham DFT calculations. Further, we propose a variant of LRDM preconditioner based on adaptive accumulation of low-rank approximations from previous self-consistent field iterations, and also extend the LRDM preconditioner to spin-polarized Kohn-Sham DFT calculations. We demonstrate the robustness and efficiency of the LRDM preconditioner against other widely used preconditioners on a range of benchmark systems with sizes ranging from ~100 to 1100 atoms (~500–20,000 electrons). The benchmark systems include various combinations of metal-insulating-semiconducting heterogeneous material systems, nanoparticles with localized d orbitals near the Fermi energy, nanofilm with metal dopants, and magnetic systems. In all benchmark systems, the LRDM preconditioner converges robustly within 20–30 iterations. In contrast, other widely used preconditioners show slow convergence in many cases, as well as divergence of the fixed-point iteration in some cases. Lastly, we demonstrate the computational efficiency afforded by the LRDM method, with up to 3.4-fold reduction in computational cost for the total ground-state calculation compared to other preconditioners.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Reassessing the Origins and Contemporary Relevance of ck Acceptability Parameters: Evolving Perspectives on Similarity

“Sensitivity and Uncertainty Analyses Applied to Criticality Safety Validation,” introduces sensitivity and uncertainty methods to address challenges in defining and extending areas of applicability for criticality safety validation. These areas are traditionally defined by the bounds or limits on key parameters, but establishing valid ranges and managing complex parameter variations remain challenging. NUREG/CR-6655 introduces ck and other integral indices, as well as concepts such as the completeness of benchmark coverage, to better quantify system similarities. The work proposed herein seeks to evaluate these foundational concepts to ensure that the bounds remain effective in guiding the assessment of similarity and applicability in modern applications. The concept of completeness, along with other parameters envisioned within the framework, serves as an example of the foundational ideas that have been established, though their effectiveness in practice may not be fully understood. Advancements in scripting tools, coupled with the speed and efficiency of modern computing and statistical models, now allow for faster and more thorough assessments than previously possible. These advancements also enable the identification of trends within the data, which could provide additional insight into system behavior and further broaden the scope of previously performed benchmarks. By leveraging these capabilities, we will revisit and expand the scope of these foundational methods to determine whether the necessary elements for robust similarity evaluation are already embedded, partially realized, or remain untapped.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Comparative evaluation of deep learning workloads for leadership-class systems

Deep learning (DL) workloads and their performance at scale are becoming important factors to consider as we design, develop and deploy next-generation high-performance computing systems. Since DL applications rely heavily on DL frameworks and underlying compute (CPU/GPU) stacks, it is essential to gain a holistic understanding from compute kernels, models, and frameworks of popular DL stacks, and to assess their impact on science-driven, mission-critical applications. At Oak Ridge Leadership Computing Facility (OLCF), we employ a set of micro and macro DL benchmarks established through the Collaboration of Oak Ridge, Argonne, and Livermore (CORAL) to evaluate the AI readiness of our next-generation supercomputers. In this paper, we present our early observations and performance benchmark comparisons between the Nvidia V100 based Summit system with its CUDA stack and an AMD MI100 based testbed system with its ROCm stack. We take a layered perspective on DL benchmarking and point to opportunities for future optimizations in the technologies that we consider.

Yin, Junqi↗

Real-Space Constrained Density Functional Theory Investigation of Site-Specific, Interfacial Charge Recombination Dynamics Across the Au Nanoparticle/TiO 2 Heterojunction

Au nanoparticle (NP)/TiO 2 heterojunction is a representative system to study interfacial charge transfer in photocatalysis and photovoltaics, where suppressing recombination from TiO 2 to Au can enhance hot carrier extraction. We apply real-space constrained density functional theory (CDFT) with Marcus theory to quantify charge recombination time scales across Au/TiO 2 . This approach enables direct control and visualization of charge-separated states, aligning with site-specific probes like time-resolved X-ray photoelectron spectroscopy (trXPS). We find that the charge-separated state features a bipolaron, with recombination dominated by TiO 2 LUMO to Au HOMO transitions, primarily at interfacial Au sites. Marcus rate predictions are benchmarked with surface hopping methods, quantifying differences in time scales and computational efficiency. Lastly, we examine how the Au cluster size affects the free energy change (ΔG) and reorganization energy (λ), explaining trends in closed-shell systems and highlighting challenges for open-shell extrapolations. Overall, CDFT + Marcus theory provides efficient, mechanistically transparent interfacial charge transfer modeling, and we clearly defined its applicability and limitation.

Glenna, Drew M. [Univ. of Idaho, Idaho Falls, ID (↗

Systematic Crosstalk Mitigation for Superconducting Qubits via Frequency-Aware Compilation

One of the key challenges in current Noisy Intermediate-Scale Quantum (NISQ) computers is to control a quantum system with high-fidelity quantum gates. There are many reasons a quantum gate can go wrong - for superconducting transmon qubits in particular, one major source of gate error is the unwanted crosstalk between neighboring qubits due to a phenomenon called frequency crowding. We motivate a systematic approach for understanding and mitigating the crosstalk noise when executing near-term quantum programs on superconducting NISQ computers. Here, we present a general software solution to alleviate frequency crowding by systematically tuning qubit frequencies according to input programs, trading parallelism for higher gate fidelity when necessary. The net result is that our work dramatically improves the crosstalk resilience of tunable-qubit, fixed-coupler hardware, matching or surpassing other more complex architectural designs such as tunable-coupler systems. On NISQ benchmarks, we improve worst-case program success rate by 13.3x on average, compared to existing traditional serialization strategies.

Computer architecture↗

Elevating SolTrace's Capabilities for the Next Generation of Concentrating Solar Analysis

SolTrace is an open-source Monte Carlo ray tracing software developed at NREL. SolTrace can characterize concentrating solar thermal (CST) collector optical performance and is CST technology agnostic. Shown in Fig. 1, SolTrace is a foundational tool in NREL's CST system and component modeling suite. SolTrace's generic surface elements can flexibly model novel collector and receiver designs to predict spatial and temporal flux distributions - critical to understand for CST component design, performance prediction, and system integration. Since its initial development, SolTrace has over 1,650 references on Google Scholar, over 9,800 downloads since 2017, and has served the CST research and development community as a benchmark of 3rd party verification. SolTrace provides users with many options for defining surface shape and boundaries. However, SolTrace provides limited documentation which can result in a steep learning curve for new users. Additionally, SolTrace lacks the computational performance required to evaluate optical performance of a CST system over the course of a year and/or iteratively over design parameters in a timely manner. To address this, we are working towards a new release of SolTrace that enables increased computational throughput by implementing ray tracing acceleration structures and enabling GPU parallelization. Additionally, we are working to improve SolTrace's usability, accessibility, and maintainability by (1) automating solar position time-dependent simulation processes, (2) creating general CST collector templates of grouped elements, (3) updating the user interface to better visualize model inputs and outputs, and (4) creating a user support network through forums, "how to" videos, and documentation.

14 SOLAR ENERGY↗

On the Computational Viability of Quantum Optimization for PMU Placement

Using optimal phasor measurement unit placement as a prototypical problem, we assess the computational viability of the current generation D-Wave Systems 2000Q quantum annealer for power systems design problems. We reformulate minimum dominating set for the annealer hardware, solve the reformulation for a standard set of IEEE test systems, and benchmark solution quality and time to solution against the CPLEX optimizer and simulated annealing. For some problem instances the 2000Q outpaces CPLEX. For instances where the 2000Q underperforms with respect to CPLEX and simulated annealing, we suggest hardware improvements for the next generation of quantum annealers.

hardware↗

DFT-FE 1.0: A massively parallel hybrid CPU-GPU density functional theory code using finite-element discretization

In this work, we present DFT-FE 1.0, building on DFT-FE 0.6 [Comput. Phys. Commun. 246, 106853 (2020)], to conduct fast and accurate large-scale density functional theory (DFT) calculations (reaching ~ 100,000 electrons) on both many-core CPU and hybrid CPU-GPU computing architectures. This work involves improvements in the real-space formulation—via an improved treatment of the electrostatic interactions that substantially enhances the computational efficiency—as well high-performance computing aspects, including the GPU acceleration of all the key compute kernels in DFT-FE. We demonstrate the accuracy by comparing the ground-state energies, ionic forces and cell stresses on a wide-range of benchmark systems against those obtained from widely used DFT codes. Further, we demonstrate the numerical efficiency of our implementation, which yields ~ 20× CPU-GPU speed-up by using GPU acceleration on hybrid CPU-GPU nodes. Notably, owing to the parallel-scaling of the GPU implementation, we obtain wall-times of 80–140 seconds for full ground-state calculations, with stringent accuracy, on benchmark systems containing ~ 6, 000 – 15,000 electrons.

pseudopotential↗

Emergent temperature sensitivity of soil organic carbon driven by mineral associations

Abstract Soil organic matter decomposition and its interactions with climate depend on whether the organic matter is associated with soil minerals. However, data limitations have hindered global-scale analyses of mineral-associated and particulate soil organic carbon pools and their benchmarking in Earth system models used to estimate carbon cycle–climate feedbacks. Here we analyse observationally derived global estimates of soil carbon pools to quantify their relative proportions and compute their climatological temperature sensitivities as the decline in carbon with increasing temperature. We find that the climatological temperature sensitivity of particulate carbon is on average 28% higher than that of mineral-associated carbon, and up to 53% higher in cool climates. Moreover, the distribution of carbon between these underlying soil carbon pools drives the emergent climatological temperature sensitivity of bulk soil carbon stocks. However, global models vary widely in their predictions of soil carbon pool distributions. We show that the global proportion of model pools that are conceptually similar to mineral-protected carbon ranges from 16 to 85% across Earth system models from the Coupled Model Intercomparison Project Phase 6 and offline land models, with implications for bulk soil carbon ages and ecosystem responsiveness. To improve projections of carbon cycle–climate feedbacks, it is imperative to assess underlying soil carbon pools to accurately predict the distribution and vulnerability of soil carbon.

54 ENVIRONMENTAL SCIENCES↗

A numerical evaluation of the ambient air temperature in the Electron-Ion Collider tunnel

The Electron-Ion Collider (EIC) is a next-generation collider-accelerator that may require consistent operating temperature conditions for the beams within the accelerator tunnels to maintain stable operation. Variations in ambient temperature within the tunnel can cause thermal expansion of beampipe and component supports and can negatively affect the tunnel equipment, impacting the stability of the beamline. Modifications will be made to the Relativistic Heavy Ion Collider (RHIC) at Brookhaven National Laboratory (BNL) to create the EIC, which necessitates a temperature model that addresses these modifications. To approach this problem, the consistency of temperature changes in different tunnel sections was first evaluated by plotting RHIC tunnel temperature data at various times of the day and year. From this data, a tunnel section was selected and a 2D temperature model was created for RHIC, EIC, and EIC with added cooling configurations. Soil temperature data was analyzed to determine the maximum, average, and mode soil temperatures, which were used as boundary conditions in different temperature scenarios. Computational fluid dynamics modeling was used to create 2D temperature profiles for the configurations. From this model, the predicted temperatures indicate that further analysis is required to validate the boundary conditions and benchmark the current conditions to allow the prediction of the tunnel ambient conditions at EIC. This research can be used as a preliminary model to create an EIC tunnel cooling system that will increase the operational stability of the EIC. As a result of my work this summer, I have become familiar with computational fluid dynamics, including creating fluid dynamic simulations using ANSYS Fluent and related software. I have also learned about the project process required for planning large-scale engineering projects.

43 PARTICLE ACCELERATORS↗

Quantum reservoir computing implementation on coherently coupled quantum oscillators

Quantum reservoir computing is a promising approach for quantum neural networks, capable of solving hard learning tasks on both classical and quantum input data. However, current approaches with qubits suffer from limited connectivity. We propose an implementation for quantum reservoir that obtains a large number of densely connected neurons by using parametrically coupled quantum oscillators instead of physically coupled qubits. We analyze a specific hardware implementation based on superconducting circuits: with just two coupled quantum oscillators, we create a quantum reservoir comprising up to 81 neurons. We obtain state-of-the-art accuracy of 99% on benchmark tasks that otherwise require at least 24 classical oscillators to be solved. Our results give the coupling and dissipation requirements in the system and show how they affect the performance of the quantum reservoir. Beyond quantum reservoir computing, the use of parametrically coupled bosonic modes holds promise for realizing large quantum neural network architectures, with billions of neurons implemented with only 10 coupled quantum oscillators.

97 MATHEMATICS AND COMPUTING↗

Adaptive Sampling-Based Bi-Fidelity Stochastic Trust Region Method for Stochastic Derivative-Free Optimization

Bi-fidelity stochastic optimization has gained increasing attention as an efficient approach to reduce computational costs by leveraging a low-fidelity (LF) model to optimize an expensive high-fidelity (HF) objective. In this paper, we propose ASTRO-BFDF, an adaptive sampling trust-region method specifically designed for unconstrained bi-fidelity stochastic derivative-free optimization problems. In ASTRO-BFDF, the LF function serves two purposes: (i) to identify better iterates for the HF function when the optimization process indicates a high correlation between them and (ii) to reduce the variance of the HF function estimates using bi-fidelity Monte Carlo (BFMC). The algorithm dynamically determines sample sizes while adaptively choosing between crude Monte Carlo and BFMC to balance the trade-off between optimization and sampling errors. We prove that the iterates generated by ASTRO-BFDF converge to a first-order stationary point almost surely. Additionally, we demonstrate the effectiveness of the proposed algorithm through numerical experiments on synthetic benchmarks and simulation optimization problems involving discrete event systems.

97 MATHEMATICS AND COMPUTING↗