Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “computer system benchmarking”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Ground-based PIV and numerical flow visualization results from the surface tension driven convection experiment

The Surface Tension Driven Convection Experiment (STDCE) is a Space Transportation System flight experiment to study both transient and steady thermocapillary fluid flows aboard the United States Microgravity Laboratory-1 (USML-1) Spacelab mission planned for June, 1992. One of the components of data collected during the experiment is a video record of the flow field. This qualitative data is then quantified using an all electric, two dimensional Particle Image Velocimetry (PIV) technique called Particle Displacement Tracking (PDT), which uses a simple space domain particle tracking algorithm. Results using the ground based STDCE hardware, with a radiant flux heating mode, and the PDT system are compared to numerical solutions obtained by solving the axisymmetric Navier Stokes equations with a deformable free surface. The PDT technique is successful in producing a velocity vector field and corresponding stream function from the raw video data which satisfactorily represents the physical flow. A numerical program is used to compute the velocity field and corresponding stream function under identical conditions. Both the PDT system and numerical results were compared to a streak photograph, used as a benchmark, with good correlation.

Pline, Alexander D.↗

Performance Evaluation and Modeling Techniques for Parallel Processors

In practice, the performance evaluation of supercomputers is still substantially driven by singlepoint estimates of metrics (e.g., MFLOPS) obtained by running characteristic benchmarks or workloads. With the rapid increase in the use of time-shared multiprogramming in these systems, such measurements are clearly inadequate. This is because multiprogramming and system overhead, as well as other degradations in performance due to time varying characteristics of workloads, are not taken into account. In multiprogrammed environments, multiple jobs and users can dramatically increase the amount of system overhead and degrade the performance of the machine. Performance techniques, such as benchmarking, which characterize performance on a dedicated machine ignore this major component of true computer performance. Due to the complexity of analysis, there has been little work done in analyzing, modeling, and predicting the performance of applications in multiprogrammed environments. This is especially true for parallel processors, where the costs and benefits of multi-user workloads are exacerbated. While some may claim that the issue of multiprogramming is not a viable one in the supercomputer market, experience shows otherwise. Even in recent massively parallel machines, multiprogramming is a key component. It has even been claimed that a partial cause of the demise of the CM2 was the fact that it did not efficiently support time-sharing. In the same paper, Gordon Bell postulates that, multicomputers will evolve to multiprocessors in order to support efficient multiprogramming. Therefore, it is clear that parallel processors of the future will be required to offer the user a time-shared environment with reasonable response times for the applications. In this type of environment, the most important performance metric is the completion of response time of a given application. However, there are a few evaluation efforts addressing this issue.

Dimpsey, Robert Tod↗

Ground-based PIV and numerical flow visualization results from the Surface Tension Driven Convection Experiment

The Surface Tension Driven Convection Experiment (STDCE) is a Space Transportation System flight experiment to study both transient and steady thermocapillary fluid flows aboard the United States Microgravity Laboratory-1 (USML-1) Spacelab mission planned for June, 1992. One of the components of data collected during the experiment is a video record of the flow field. This qualitative data is then quantified using an all electric, two dimensional Particle Image Velocimetry (PIV) technique called Particle Displacement Tracking (PDT), which uses a simple space domain particle tracking algorithm. Results using the ground based STDCE hardware, with a radiant flux heating mode, and the PDT system are compared to numerical solutions obtained by solving the axisymmetric Navier Stokes equations with a deformable free surface. The PDT technique is successful in producing a velocity vector field and corresponding stream function from the raw video data which satisfactorily represents the physical flow. A numerical program is used to compute the velocity field and corresponding stream function under identical conditions. Both the PDT system and numerical results were compared to a streak photograph, used as a benchmark, with good correlation.

Pline, Alexander D.↗

Multi-Area Distribution System State Estimation Using Decentralized Physics-Aware Neural Networks

The development of active distribution grids requires more accurate and lower computational cost state estimation. In this paper, the authors investigate a decentralized learning-based distribution system state estimation (DSSE) approach for large distribution grids. The proposed approach decomposes the feeder-level DSSE into subarea-level estimation problems that can be solved independently. The proposed method is decentralized pruned physics-aware neural network (D-P2N2). The physical grid topology is used to parsimoniously design the connections between different hidden layers of the D-P2N2. Monte Carlo simulations based on one-year of load consumption data collected from smart meters for a three-phase distribution system power flow are developed to generate the measurement and voltage state data. The IEEE 123-node system is selected as the test network to benchmark the proposed algorithm against the classic weighted least squares and state-of-the-art learning-based DSSE approaches. Numerical results show that the D-P2N2 outperforms the state-of-the-art methods in terms of estimation accuracy and computational efficiency.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Frequency-domain vs time-domain SP{sub N} equations to simulate neutron noise

Inside the reactor core mechanical vibrations of fuel assemblies can produce high fluctuations around a steady-state configuration, known as neutron noise. This effect can cause the triggering of power reduction measures. Classically, diffusion theory has been used to simulate this behavior. However, this equation has some limitations if the materials of the reactor have strong variations. In this work, we use the diffusive time-dependent simplified spherical harmonics equations that improve the previous results without the necessity of using high computational requirements. In particular, two types of analyses with these equations (SP{sub 3}) are made: a frequency-domain and a time-domain. A numerical neutron noise benchmark tests the methodology and compare both formulations. First, numerical results show a good agreement between the amplitudes and phases of the SP3 equations computed with the frequency-domain and time-domain. Therefore, as the frequency-domain computation only requires to solve a linear system, it is a recommendable option for neutron noise computations. Second, one can conclude that for this type of nuclear systems, where the assemblies are not homogenized, the SP{sub 3} approximation results improve considerably the accuracy of the diffusion theory.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Analysis and Benchmarking of feature reduction for classification under computational constraints

Abstract Machine learning is most often expensive in terms of computational and memory costs due to training with large volumes of data. Current computational limitations of many computing systems motivate us to investigate practical approaches, such as feature selection and reduction, to reduce the time and memory costs while not sacrificing the accuracy of classification algorithms. In this work, we carefully review, analyze, and identify the feature reduction methods that have low costs/overheads in terms of time and memory. Then, we evaluate the identified reduction methods in terms of their impact on the accuracy, precision, time, and memory costs of traditional classification algorithms. Specifically, we focus on the least resource intensive feature reduction methods that are available in Scikit-Learn library. Since our goal is to identify the best performing low-cost reduction methods, we do not consider complex expensive reduction algorithms in this study. In our evaluation, we find that at quadratic-scale feature reduction, the classification algorithms achieve the best trade-off among competitive performance metrics. Results show that the overall training times are reduced 61%, the model sizes are reduced 6×, and accuracy scores increase 25% compared to the baselines on average with quadratic scale reduction.

97 MATHEMATICS AND COMPUTING↗

Quantum many-body calculations using body-centered cubic lattices

It is often computationally advantageous to model space as a discrete set of points forming a lattice grid. This technique is particularly useful for computationally difficult problems such as quantum many-body systems. For reasons of simplicity and familiarity, nearly all quantum many-body calculations have been performed on simple cubic lattices. Since the removal of lattice artifacts is often an important concern, it would be useful to perform calculations using more than one lattice geometry. In this paper we show how to perform quantum many-body calculations using auxiliary-field Monte Carlo simulations on a three-dimensional body-centered cubic (BCC) lattice. As a benchmark test we compute the ground state energy of 33 spin-up and 33 spin-down neutrons in the unitary limit, which is an idealized limit where the interaction range is zero and scattering length is infinite. As a fraction of the free Fermi gas energy E FG , we find that the ground state energy is E 0 /E FG =0.369(2),0.371(2), using two different definitions of the finite-system energy ratio. This is in excellent agreement with recent results obtained on a cubic lattice [He et al., Phys. Rev. A 101, 063615 (2020)]. We find that the computational effort and performance on a BCC lattice is approximately the same as that for a cubic lattice with the same number of lattice points. We discuss how the lattice simulations with different geometries can be used to constrain the size of lattice artifacts in simulations of continuum quantum many-body systems.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Finding the missing pieces: filling gaps that impede the translation of omics data into models

High-throughput omics technologies such as DNA sequencing have made the sequencing and computational assembly of microbial genomes recovered from the environment relatively routine. Computational inference of the protein products encoded by these genomes, and the associated biochemical functions, should enable the accurate prediction and modeling of microbial metabolism, organismal interactions, and ecosystem processes. However, a lack of scalable, probabilistic protein annotation tools limits the full potential of modeling for understanding the metabolism and biogeochemical cycles of microbial communities. Our approach to improve inference of protein annotations and metabolic models relied on learning from and emulating expert manual curation, leveraging software engineering and data science best practices to scale up the throughput and accuracy of annotations and metabolic model construction, building software to objectively evaluate different annotation strategies, and more closely linking the protein annotation and metabolic model inference process. Outcomes of this research include several improved or new computational tools, including DRAM (Distilled and Refined Annotation of Metabolism) for annotating microbial genomes with protein function and metabolic traits, CAMPER (Curated Annotations for Microbial Polyphenol Enzymes and Reactions) for annotating key polyphenol metabolisms, EC-Bench for comprehensive and unbiased benchmarking of annotation tools, and several apps available via the DOE Systems Biology Knowledgebase (KBase) for building genome-scale metabolic models. We demonstrate that these tools allow us to scalably annotate and understand thousands of genomes for microbial communities from a variety of systems and test cases, including rivers, thawing permafrost, and gut microbiomes. All of these computational tools are available as open-source software, with most broadly and easily accessible to the scientific community via KBase apps.

59 BASIC BIOLOGICAL SCIENCES↗

A graphics processing unit accelerated sparse direct solver and preconditioner with block low rank compression

We present the GPU implementation efforts and challenges of the sparse solver package STRUMPACK. The code is made publicly available on github with a permissive BSD license. STRUMPACK implements an approximate multifrontal solver, a sparse LU factorization which makes use of compression methods to accelerate time to solution and reduce memory usage. Multiple compression schemes based on rank-structured and hierarchical matrix approximations are supported, including hierarchically semi-separable, hierarchically off-diagonal butterfly, and block low rank. Here, in this paper, we present the GPU implementation of the block low rank (BLR) compression method within a multifrontal solver. Our GPU implementation relies on highly optimized vendor libraries such as cuBLAS and cuSOLVER for NVIDIA GPUs, rocBLAS and rocSOLVER for AMD GPUs and the Intel oneAPI Math Kernel Library (oneMKL) for Intel GPUs. Additionally, we rely on external open source libraries such as SLATE (Software for Linear Algebra Targeting Exascale), MAGMA (Matrix Algebra on GPU and Multi-core Architectures), and KBLAS (KAUST BLAS). SLATE is used as a GPU-capable ScaLAPACK replacement. From MAGMA we use variable sized batched dense linear algebra operations such as GEMM, TRSM and LU with partial pivoting. KBLAS provides efficient (batched) low rank matrix compression for NVIDIA GPUs using an adaptive randomized sampling scheme. The resulting sparse solver and preconditioner runs on NVIDIA, AMD and Intel GPUs. Interfaces are available from PETSc, Trilinos and MFEM, or the solver can be used directly in user code. We report results for a range of benchmark applications, using the Perlmutter system from NERSC, Frontier from ORNL, and Aurora from ALCF. For a high frequency wave equation on a regular mesh, using 32 Perlmutter compute nodes, the factorization phase of the exact GPU solver is about 6.5× faster compared to the CPU-only solver. The BLR-enabled GPU solver is about 13.8× faster than the CPU exact solver. For a collection of SuiteSparse matrices, the STRUMPACK exact factorization on a single GPU is on average 1.9× faster than NVIDIA’s cuDSS solver.

97 MATHEMATICS AND COMPUTING↗

Performance Measurement, Visualization and Modeling of Parallel and Distributed Programs

This paper presents a methodology for debugging the performance of message-passing programs on both tightly coupled and loosely coupled distributed-memory machines. The AIMS (Automated Instrumentation and Monitoring System) toolkit, a suite of software tools for measurement and analysis of performance, is introduced and its application illustrated using several benchmark programs drawn from the field of computational fluid dynamics. AIMS includes (i) Xinstrument, a powerful source-code instrumentor, which supports both Fortran77 and C as well as a number of different message-passing libraries including Intel's NX Thinking Machines' CMMD, and PVM; (ii) Monitor, a library of timestamping and trace -collection routines that run on supercomputers (such as Intel's iPSC/860, Delta, and Paragon and Thinking Machines' CM5) as well as on networks of workstations (including Convex Cluster and SparcStations connected by a LAN); (iii) Visualization Kernel, a trace-animation facility that supports source-code clickback, simultaneous visualization of computation and communication patterns, as well as analysis of data movements; (iv) Statistics Kernel, an advanced profiling facility, that associates a variety of performance data with various syntactic components of a parallel program; (v) Index Kernel, a diagnostic tool that helps pinpoint performance bottlenecks through the use of abstract indices; (vi) Modeling Kernel, a facility for automated modeling of message-passing programs that supports both simulation -based and analytical approaches to performance prediction and scalability analysis; (vii) Intrusion Compensator, a utility for recovering true performance from observed performance by removing the overheads of monitoring and their effects on the communication pattern of the program; and (viii) Compatibility Tools, that convert AIMS-generated traces into formats used by other performance-visualization tools, such as ParaGraph, Pablo, and certain AVS/Explorer modules.

Yan, Jerry C.↗

Applications of Flow Control to Wing High-Lift Leading Edge Devices on a Commercial Aircraft

Active flow control was applied to the leading edge region of a representative future short/medium-range twin-engine airplane to improve aerodynamic performance during high-lift operations. The study is aimed at enhanced lift over the practical angle of attack range, including stall, and at reduced drag. These benefits translate to airplane performance improvements, such as longer range or larger payload. Various flow control applications were explored using Computational Fluid Dynamics and the aerodynamic performance enhancements were benchmarked against the baseline configuration. The computational analyses are used to quantify aerodynamic benefits, as well as the input required for actuation. The results were used in a system integration study for identifying potential practical implementations, which are described in a companion paper. Combined with the integration analysis, the objective of this project is to identify the most promising flow control candidates that potentially provide material net airplane level enhancements using onboard fluidic sources. Depending on the implementation of active flow control, the current study indicates that up to 1.5% net improvement in L/D at takeoff and 4% increase in maximum lift during landing are potentially achievable, after accounting for factors of system integration.

CFD↗

Flow Control for Enhanced Aileron Effectiveness on a Commercial Aircraft

Active flow control was applied to the ailerons of a representative future short/medium-range twin-engine airplane to improve aerodynamic performance during high-lift operations. The study is aimed at reduced drag and enhanced lift over the range of practical angles of attack, including stall. These benefits translate to airplane performance improvements, such as longer range or larger payload. Various flow control techniques were explored using Computational Fluid Dynamics and the aerodynamic performance enhancements were benchmarked against the baseline configuration. The computational analyses are used to quantify aerodynamic benefits, as well as the input required for actuation. The results were used in a system integration study for identifying potential practical implementations, which are described in a companion paper. Combined with the integration analysis, the objective of this project is to identify the most promising flow control candidates that potentially provide material net airplane level enhancements using onboard fluidic sources. The current study indicates that up to 5% net improvement in L/D at takeoff is potentially achievable using active flow control on the aileron, after accounting for factors of system integration.

CFD↗

Design and Performance of Kokkos Staging Space toward Scalable Resilient Application Couplings

With the growing number of applications designed for heterogeneous HPC devices, application programmers and users are finding it challenging to compose scalable workflows as ensembles of these applications, that are portable, performant and resilient. The Kokkos C++ library has been designed to simplify this cumbersome procedure by providing an intra-application uniform programming model and portable performance. However, assembling multiple Kokkos-enabled applications into a complex workflow is still a challenge. Although Kokkos enables a uniform programming model, the inter-application data exchange still remains a challenge from both performance and software development cost perspectives. In order to address this issue, we propose Kokkos data staging memory space, an extension of Kokkos' data abstraction (memory space) for heterogeneous computing systems. This new abstraction allows to express data on a virtual shared-space for multiple Kokkos applications, thus extending Kokkos to support inter-application data exchange to build an efficient application workflow. Additionally, we study the effectiveness of asynchronous data layout conversions for applications requiring different memory access patterns for the shared data. Our preliminary evaluation with a synthetic benchmark indicate the effectiveness of this conversion adapted to three different scenarios representing access frequency and use patterns of the shared data.

97 MATHEMATICS AND COMPUTING↗

Accelerating self-consistent field iterations in Kohn-Sham density functional theory using a low-rank approximation of the dielectric matrix

We present an efficient preconditioning technique for accelerating the fixed-point iteration in real-space Kohn-Sham density functional theory (DFT) calculations. The preconditioner uses a low-rank approximation of the dielectric matrix (LRDM) based on Gâteaux derivatives of the residual of fixed-point iteration along appropriately chosen direction functions. We develop a computationally efficient method to evaluate these Gâteaux derivatives in conjunction with the Chebyshev filtered subspace iteration procedure, an approach widely used in large-scale Kohn-Sham DFT calculations. Further, we propose a variant of LRDM preconditioner based on adaptive accumulation of low-rank approximations from previous self-consistent field iterations, and also extend the LRDM preconditioner to spin-polarized Kohn-Sham DFT calculations. We demonstrate the robustness and efficiency of the LRDM preconditioner against other widely used preconditioners on a range of benchmark systems with sizes ranging from ~100 to 1100 atoms (~500–20,000 electrons). The benchmark systems include various combinations of metal-insulating-semiconducting heterogeneous material systems, nanoparticles with localized d orbitals near the Fermi energy, nanofilm with metal dopants, and magnetic systems. In all benchmark systems, the LRDM preconditioner converges robustly within 20–30 iterations. In contrast, other widely used preconditioners show slow convergence in many cases, as well as divergence of the fixed-point iteration in some cases. Lastly, we demonstrate the computational efficiency afforded by the LRDM method, with up to 3.4-fold reduction in computational cost for the total ground-state calculation compared to other preconditioners.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Reassessing the Origins and Contemporary Relevance of ck Acceptability Parameters: Evolving Perspectives on Similarity

“Sensitivity and Uncertainty Analyses Applied to Criticality Safety Validation,” introduces sensitivity and uncertainty methods to address challenges in defining and extending areas of applicability for criticality safety validation. These areas are traditionally defined by the bounds or limits on key parameters, but establishing valid ranges and managing complex parameter variations remain challenging. NUREG/CR-6655 introduces ck and other integral indices, as well as concepts such as the completeness of benchmark coverage, to better quantify system similarities. The work proposed herein seeks to evaluate these foundational concepts to ensure that the bounds remain effective in guiding the assessment of similarity and applicability in modern applications. The concept of completeness, along with other parameters envisioned within the framework, serves as an example of the foundational ideas that have been established, though their effectiveness in practice may not be fully understood. Advancements in scripting tools, coupled with the speed and efficiency of modern computing and statistical models, now allow for faster and more thorough assessments than previously possible. These advancements also enable the identification of trends within the data, which could provide additional insight into system behavior and further broaden the scope of previously performed benchmarks. By leveraging these capabilities, we will revisit and expand the scope of these foundational methods to determine whether the necessary elements for robust similarity evaluation are already embedded, partially realized, or remain untapped.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Comparative evaluation of deep learning workloads for leadership-class systems

Deep learning (DL) workloads and their performance at scale are becoming important factors to consider as we design, develop and deploy next-generation high-performance computing systems. Since DL applications rely heavily on DL frameworks and underlying compute (CPU/GPU) stacks, it is essential to gain a holistic understanding from compute kernels, models, and frameworks of popular DL stacks, and to assess their impact on science-driven, mission-critical applications. At Oak Ridge Leadership Computing Facility (OLCF), we employ a set of micro and macro DL benchmarks established through the Collaboration of Oak Ridge, Argonne, and Livermore (CORAL) to evaluate the AI readiness of our next-generation supercomputers. In this paper, we present our early observations and performance benchmark comparisons between the Nvidia V100 based Summit system with its CUDA stack and an AMD MI100 based testbed system with its ROCm stack. We take a layered perspective on DL benchmarking and point to opportunities for future optimizations in the technologies that we consider.

Yin, Junqi↗

Automated Instrumentation, Monitoring and Visualization of PVM Programs Using AIMS

We present views and analysis of the execution of several PVM codes for Computational Fluid Dynamics on a network of Sparcstations, including (a) NAS Parallel benchmarks CG and MG (White, Alund and Sunderam 1993); (b) a multi-partitioning algorithm for NAS Parallel Benchmark SP (Wijngaart 1993); and (c) an overset grid flowsolver (Smith 1993). These views and analysis were obtained using our Automated Instrumentation and Monitoring System (AIMS) version 3.0, a toolkit for debugging the performance of PVM programs. We will describe the architecture, operation and application of AIMS. The AIMS toolkit contains (a) Xinstrument, which can automatically instrument various computational and communication constructs in message-passing parallel programs; (b) Monitor, a library of run-time trace-collection routines; (c) VK (Visual Kernel), an execution-animation tool with source-code clickback; and (d) Tally, a tool for statistical analysis of execution profiles. Currently, Xinstrument can handle C and Fortran77 programs using PVM 3.2.x; Monitor has been implemented and tested on Sun 4 systems running SunOS 4.1.2; and VK uses X11R5 and Motif 1.2. Data and views obtained using AIMS clearly illustrate several characteristic features of executing parallel programs on networked workstations: (a) the impact of long message latencies; (b) the impact of multiprogramming overheads and associated load imbalance; (c) cache and virtual-memory effects; and (4significant skews between workstation clocks. Interestingly, AIMS can compensate for constant skew (zero drift) by calibrating the skew between a parent and its spawned children. In addition, AIMS' skew-compensation algorithm can adjust timestamps in a way that eliminates physically impossible communications (e.g., messages going backwards in time). Our current efforts are directed toward creating new views to explain the observed performance of PVM programs. Some of the features planned for the near future include: (a) ConfigView, showing the physical topology of the virtual machine, inferred using specially formatted IP (Internet Protocol) packets; and (b) LoadView, synchronous animation of PVM-program execution and resource-utilization patterns.

Mehra, Pankaj↗

Real-Space Constrained Density Functional Theory Investigation of Site-Specific, Interfacial Charge Recombination Dynamics Across the Au Nanoparticle/TiO 2 Heterojunction

Au nanoparticle (NP)/TiO 2 heterojunction is a representative system to study interfacial charge transfer in photocatalysis and photovoltaics, where suppressing recombination from TiO 2 to Au can enhance hot carrier extraction. We apply real-space constrained density functional theory (CDFT) with Marcus theory to quantify charge recombination time scales across Au/TiO 2 . This approach enables direct control and visualization of charge-separated states, aligning with site-specific probes like time-resolved X-ray photoelectron spectroscopy (trXPS). We find that the charge-separated state features a bipolaron, with recombination dominated by TiO 2 LUMO to Au HOMO transitions, primarily at interfacial Au sites. Marcus rate predictions are benchmarked with surface hopping methods, quantifying differences in time scales and computational efficiency. Lastly, we examine how the Au cluster size affects the free energy change (ΔG) and reorganization energy (λ), explaining trends in closed-shell systems and highlighting challenges for open-shell extrapolations. Overall, CDFT + Marcus theory provides efficient, mechanistically transparent interfacial charge transfer modeling, and we clearly defined its applicability and limitation.

Glenna, Drew M. [Univ. of Idaho, Idaho Falls, ID (↗