Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “scalable performance”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Proton Tunable Analog Transistor for Low Power Computing

This project was broadly motivated by the need for new hardware that can process information such as images and sounds right at the point of where the information is sensed (e.g. edge computing). The project was further motivated by recent discoveries by group demonstrating that while certain organic polymer blends can be used to fabricate elements of such hardware, the need to mix ionic and electronic conducting phases imposed limits on performance, dimensional scalability and the degree of fundamental understanding of how such devices operated. As an alternative to blended polymers containing distinct ionic and electronic conducting phases, in this LDRD project we have discovered that a family of mixed valence coordination compounds called Prussian blue analogue (PBAs), with an open framework structure and ability to conduct both ionic and electronic charge, can be used for inkjet-printed flexible artificial synapses that reversibly switch conductance by more than four orders of magnitude based on electrochemically tunable oxidation state. Retention of programmed states is improved by nearly two orders of magnitude compared to the extensively studied organic polymers, thus enabling in-memory compute and avoiding energy costly off-chip access during training. We demonstrate dopamine detection using PBA synapses and biocompatibility with living neurons, evoking prospective application for brain - computer interfacing. By application of electron transfer theory to in-situ spectroscopic probing of intervalence charge transfer, we elucidate a switching mechanism whereby the degree of mixed valency between N-coordinated Ru sites controls the carrier concentration and mobility, as supported by density functional theory (DFT) .

97 MATHEMATICS AND COMPUTING↗

Celeritas Midterm SciDAC Report

Celeritas is a new Monte Carlo (MC) code that helps satisfy the increasing demand for high energy physics (HEP) detector simulation, using Graphics Processing Unit (GPU) hardware on high performance computing (HPC) systems to model Large Hadron Collider (LHC) experiments and beyond. This report details the project’s progress midway through its SciDAC funding period, highlighting the first complete implementation of standard electromagnetic (EM) physics on GPUs, initial results for performance and scalability on Leadership Computing Facilities (LCFs), and preliminary integration into the CMS and ATLAS experiments. By integrating HEP domain knowledge with expertise in MC transport, Celeritas has catalyzed a shift in the HEP community’s perception of GPU platforms as the future for HPC simulations.

97 MATHEMATICS AND COMPUTING↗

Automatically parallelizing batch inference on deep neural networks using Fiats and Fortran 2023 `do concurrent`

This paper introduces novel programming strategies that leverage features of the Fortran 2023 standard of the International Standards Organization (ISO) to automatically parallelize computations on deep neural networks. The paper focuses on the interplay of object-oriented, parallel, and functional programming paradigms in the Fiats deep learning library. We demonstrate how several infrequently used language features play a role in enabling efficient, parallel execution. Specifically, the ability to explicitly declare that a procedure is pure facilitates inference in the context of the language’s loop-parallelism construct `do concurrent`. Also, explicitly prohibiting the overriding of a parent type’s type-bound procedures eliminates the need for dynamic dispatch in performance-critical code. Finally, this paper uses batch inference calculations on a neural network surrogate for atmospheric aerosol dynamics to demonstrate that LLVM Flang compiler’s automatic parallelization of `do concurrent` achieves roughly the same performance and scalability as achieved by OpenMP compiler directives. We also demonstrate that double-precision inference costs 37–72% longer runtime than default-real precision with most values in the range 57-60%.

Rouson, Damian↗

GR-Athena++: Puncture Evolutions on Vertex-centered Oct-tree Adaptive Mesh Refinement

Numerical relativity is central to the investigation of astrophysical sources in the dynamical and strong-field gravity regime, such as binary black hole and neutron star coalescences. Current challenges set by gravitational-wave and multimessenger astronomy call for highly performant and scalable codes on modern massively parallel architectures. We present GR-Athena++, a general-relativistic, high-order, vertex-centered solver that extends the oct-tree, adaptive mesh refinement capabilities of the astrophysical (radiation) magnetohydrodynamics code Athena++. To simulate dynamical spacetimes, GR-Athena++ uses the Z4c evolution scheme of numerical relativity coupled to the moving puncture gauge. We demonstrate stable and accurate binary black hole merger evolutions via extensive convergence testing, cross-code validation, and verification against state-of-the-art effective-one-body waveforms. GR-Athena++ leverages the task-based parallelism paradigm of Athena++ to achieve excellent scalability. We measure strong-scaling efficiencies above 95% for up to ~1.2 × 10 4 CPUs and excellent weak scaling is shown up to ~10 5 CPUs in a production binary black hole setup with adaptive mesh refinement. GR-Athena++ thus allows for the robust simulation of compact binary coalescences and offers a viable path toward numerical relativity at exascale.

79 ASTRONOMY AND ASTROPHYSICS↗

DMTN-293: Current status of APDB and PPDB implementation

The design of the Alert Production DataBase (APDB) and Prompt Products DataBase (PPDB) has been evolving for the past few years as a result of various tests and updated requirements. This technical note describes the current state of these systems, as well as issues related to scalability and performance. It also discusses unresolved issues in the system's architecture and possible future directions.

79 ASTRONOMY AND ASTROPHYSICS↗

mpi-sppy: Optimization Under Uncertainty for Pyomo

We introduce a new Pyomo extension for optimization under uncertainty, mpi-sppy. This extension allows for the mixing of, and sharing of information between, various algorithmic approaches for optimization under uncertainty. We will discuss the underlying architecture and assess the performance and scalability of mpi-sppy on various systems, including HPC clusters.

27 ARPA - Advanced Research Projects Agency-Energy↗

Architecture for Web-Based Visualization of Large-Scale Energy Domains: Preprint

With the growing penetration of inverter-based distributed energy resources and increased loads through electrification, power systems analyses are becoming more important and more complex. Moreover, these analyses increasingly involve the combination of interconnected energy domains with data that are spatially and temporally increasing in scale by orders of magnitude, surpassing the capabilities of many existing analysis and decision-support systems. We present the architectural design, development, and application of a high-resolution web-based visualization environment capable of cross-domain analysis of tens of millions of energy assets, focusing on scalability and performance. Our system supports the exploration, navigation, and analysis of large data from diverse domains such as electrical transmission and distribution systems, mobility and electric vehicle charging networks, communications networks, cyber assets, and other supporting infrastructure. We evaluate this system across multiple use cases, describing the capabilities and limitations of a web-based approach for high-resolution energy system visualizations.

grid modernization↗

Dual Channel Dual Staging: Hierarchical and Portable Staging for GPU-Based In-Situ Workflow

In-situ workflows have emerged as an attractive approach for addressing data movement challenges at very large scales. Since GPU-based architectures dominate the HPC landscapes, porting these in-situ workflows, and, specifically, the inter-application data exchange, to GPU-based systems can be challenging. Technologies such as GPUDirect RDMA (GDR), which is typically used for I/O in GPU applications as an optimization that circumvents the CPU overhead, can be leveraged to support bulk data exchanges between GPU applications. However, current GDR design often lacks performance portability across HPC clusters built with different hardware configurations. Furthermore, the local CPU may also be effectively used as an auxiliary communication mechanism to offload data exchanges. In this paper, we present a dual channel dual staging approach for efficient, scalable, and performance-portable inter-application data exchange for in-situ workflows. This approach exploits the data access pattern within in-situ workflows along with the inherent execution asynchrony to accelerate data exchanges and, at the same time, improve performance portability. Specifically, the dual channel dual staging method leverages both the local CPU and the remote data staging server to build a hierarchical joint staging area and uses this staging area to transform blocking inter-application bulk data exchanges into best-effort local data movements between GPU and CPU. The dual channel dual staging is implemented as a portability extension of the Dataspaces-GPU staging framework. We present an experimental evaluation of its performance, portability, and scalability using this implementation on three leadership GPU clusters. The evaluation results demonstrate that the dual channel dual staging method saves up to 75% in data-exchange time compared to host-based, GDR, and alternate portable designs, while maintaining scalability (up to 512 GPUs) and performance portability across the three platforms.

Zhang, Bo [University of Utah]↗

Extending Petsc's Composable Hierarchically Nested Linear Solvers

The Contributions from the RELACS group at both Rice University and the University at Buffalo in this phase of the PETSc Composable Solvers effort have centered around four main areas: scalable mesh processing, mesh adaptivity, solvers for subsurface flow, and performance modeling. The prominence of mesh processing demonstrates the tight relationship between meshing and discretization on the one hand, and optimal solvers on the other. All optimal solvers that we consider depend on some notion of hierarchy, and we express this using the DMPlex abstraction in PETSc. This relationship demands tight integration between the DM and SNES/TS components in PETSc that is the foundation of much of this work. In addition, interpretation of performance results for scalable solvers necessitates that information from the discretization and solver enter the performance model. Without this, comparing different solvers can be a fruitless exercise. Some major accomplishment of the past three years in these areas include: scalable mesh loading in PETSc on more than 10K cores, integrated mesh adaptivity using both p4est and Pragmatic, scalable multigrid for DG discretizations of subsurface flow, and predictive performance modeling incorporating error estimates.

79 ASTRONOMY AND ASTROPHYSICS↗

Benchmarking the Performance of Neuromorphic and Spiking Neural Network Simulators

Software simulators play a critical role in the development of new algorithms and system architectures in any field of engineering. Neuromorphic computing, which has shown potential in building brain-inspired energy-efficient hardware, suffers a slow-down in the development cycle due to a lack of flexible and easy-to-use simulators of either neuromorphic hardware itself or of spiking neural networks (SNNs), the type of neural network computation executed on most neuromorphic systems. While there are several openly available neuromorphic or SNN simulation packages developed by a variety of research groups, they have mostly targeted computational neuroscience simulations, and only a few have targeted small-scale machine learning tasks with SNNs. Evaluations or comparisons of these simulators have often targeted computational neuroscience-style workloads. In this work, we seek to evaluate the performance of several publicly available SNN simulators with respect to non-computational neuroscience workloads, in terms of speed, flexibility, and scalability. We evaluate the performance of the NEST, Brian2, Brian2GeNN, BindsNET and Nengo packages under a common front-end neuromorphic framework. Our evaluation tasks include a variety of different network architectures and workload types to mimic the computation common in different algorithms, including feed-forward network inference, genetic algorithms, and reservoir computing. We also study the scalability of each of these simulators when running on different computing hardware, from single core CPU workstations to multi-node supercomputers. Our results show that the BindsNET simulator has the best speed and scalability for most of the SNN workloads (sparse, dense, and layered SNN architectures) on a single core CPU. However, when comparing the simulators leveraging the GPU capabilities, Brian2GeNN outperforms the others for these workloads in terms of scalability. NEST performs the best for small sparse networks and is also the most flexible simulator in terms of reconfiguration capability NEST shows a speedup of at least 2x compared to the other packages when running evolutionary algorithms for SNNs. The multi-node and multi-thread capabilities of NEST show at least 2x speedup compared to the rest of the simulators (single core CPU or GPU based simulators) for large and sparse networks. We conclude our work by providing a set of recommendations on the suitability of employing these simulators for different tasks and scales of operations. We also present the characteristics for a future generic ideal SNN simulator for different neuromorphic computing workloads.

97 MATHEMATICS AND COMPUTING↗

Performance portable ice-sheet modeling with MALI

High-resolution simulations of polar ice sheets play a crucial role in the ongoing effort to develop more accurate and reliable Earth system models for probabilistic sea-level projections. These simulations often require a massive amount of memory and computation from large supercomputing clusters to provide sufficient accuracy and resolution; therefore, it has become essential to ensure performance on these platforms. Many of today’s supercomputers contain a diverse set of computing architectures and require specific programming interfaces in order to obtain optimal efficiency. In an effort to avoid architecture-specific programming and maintain productivity across platforms, the ice-sheet modeling code known as MPAS-Albany Land Ice (MALI) uses high-level abstractions to integrate Trilinos libraries and the Kokkos programming model for performance portable code across a variety of different architectures. In this article, we analyze the performance portable features of MALI via a performance analysis on current CPU-based and GPU-based supercomputers. The analysis highlights not only the performance portable improvements made in finite element assembly and multigrid preconditioning within MALI with speedups between 1.26 and 1.82x across CPU and GPU architectures but also identifies the need to further improve performance in software coupling and preconditioning on GPUs. We perform a weak scalability study and show that simulations on GPU-based machines perform 1.24–1.92x faster when utilizing the GPUs. The best performance is found in finite element assembly, which achieved a speedup of up to 8.65x and a weak scaling efficiency of 82.6% with GPUs. We additionally describe an automated performance testing framework developed for this code base using a changepoint detection method. The framework is used to make actionable decisions about performance within MALI. We provide several concrete examples of scenarios in which the framework has identified performance regressions, improvements, and algorithm differences over the course of 2 years of development.

54 ENVIRONMENTAL SCIENCES↗

Towards Superior Software Portability with SHAD and HPX C++ Libraries

As hardware architectures and software stacks complexity grows, development productivity, performance and software portability, quickly evolve from desirable features to actual needs. SHAD, the Scalable High-performance Algorithms and Data-structures C++ library is designed to mitigate these issues: it provides general purpose building blocks as well as high-level custom utilities, and offers a shared-memory programming abstraction which facilitates the programming of complex systems, scaling up to High Performance Computing clusters. SHAD’s portability is achieved through an abstract runtime interface, which decouples the upper layers of the library and hides the low level details of the underlying architecture. This layer enables SHAD to interface with different runtime/threading systems, e.g. Intel TBB and Global Memory and Threading (GMT). However, current backends targeting distributed systems, rely on a centralized controller which may possibly limit scalability up to hundreds of nodes and creates a network hot spot due to all to one communication for synchronization, and possibly resulting in degraded performance at high process counts. In this research, we explore HPX, the C++ standard library for parallelism and concurrency, as an additional backend in support of the SHAD library, and present the methodologies in support of local and remote task executions in SHAD with respect to HPX. Finally, we evaluate the proposed system by comparing against existing backends of SHAD and analyzing their performance on C++ Standard Template Library algorithms.

Wu, Nanmiao↗

Demonstration of low-density, high-performance operation of sustained spheromaks and favorable scalability toward compact, low-cost fusion power plants (Final Scientific/Technical Report)

This project worked to advance the technical viability of a novel method for efficiently sustaining stable, high-performance spheromak plasma configurations to serve as the basis of compact, low-cost fusion power plants. In particular, our group worked to improve the method of Steady Inductive Helicity Injection (SIHI) with Imposed-Dynamo Current Drive (IDCD) for spheromak plasma sustainment. Prior to this project demonstrations of this plasma sustainment technology have achieved plasma performance consistent with entry milestone 3 of the BETHE FOA. Research and development (R&D) activities for this project were focused on increasing plasma performance toward a level consistent with exit milestone 4. To do this the PI and his group worked to increase the performance of sustained spheromaks produced in an existing experimental prototype (HIT-SIU) while improving confidence in projections to and design of future, higher performance devices through three primary R&D activities: 1) Improved control over the density of plasma in the device throughout a discharge to provide a pathway for demonstration of spheromaks Ohmically heating to the Mercier beta limit via: a. Fueling the device directly with plasma through the installation of pre-ionized source on the injectors b. Optimization of electrical current waveforms in the driver circuits to enable low-density plasma formation with a lower fueling rate 2) Computational demonstration of a validated, realistic injector circuit coupled to a dynamic plasma model capable of use as a design tool for SIHI drivers and associated circuits for new experimental design points on the pathway to commercial reactors. The improvements in plasma performance achieved during research activity 1), and the computational projections performed in research activity 2) increased the technological readiness level (TRL) of this fusion energy concept toward a level sufficient to attract early-stage private investment and/or other forms of follow-on investment to pursue required R&D activities required for the eventual fusion power plants based on this novel technical approach.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A Fine-grained Asynchronous Bulk Synchronous parallelism model for PGAS applications

The Partitioned Global Address Space (PGAS) model is well suited for executing irregular applications on cluster-based systems, due to its efficient support for short, one-sided messages. Separately, the actor model has been gaining popularity as a productive asynchronous message-passing approach for distributed objects in enterprise and cloud computing platforms, typically implemented in languages such as Erlang, Scala or Rust. To the best of our knowledge, there has been no past work on using the actor model to deliver both productivity and scalability to irregular PGAS applications with large number of small messages. In this paper, we introduce a new programming system for PGAS applications, in which point-to-point remote operations can be expressed as fine-grained asynchronous actor messages. In our approach, the programmer does not need to worry about programming complexities related to message aggregation and termination detection. Our approach can be viewed as extending the classical Bulk Synchronous Parallelism model with fine-grained asynchronous communications within a phase or superstep. Here, we believe that our approach offers a desirable point in the productivity-performance space for PGAS applications, with more scalable performance and higher productivity relative to past approaches. Specifically, for seven irregular mini-applications from the Bale Kernels and three graph kernels executed using 2048 cores in the NERSC Cori system, our approach shows geometric mean performance improvements of ≥ 20X relative to standard PGAS versions (UPC and OpenSHMEM) while maintaining comparable productivity to those versions.

97 MATHEMATICS AND COMPUTING↗

Direct numerical simulations of turbulent reacting flows with shock waves and stiff chemistry using many-core/GPU acceleration

Compressible reacting flows may display sharp spatial variation related to shocks, contact discontinuities or reactive zones embedded within relatively smooth regions. The presence of such phenomena emphasizes the relevance of shock-capturing schemes such as the weighted essentially non-oscillatory (WENO) scheme as an essential ingredient of the numerical solver. However, these schemes are complex and have more computational cost than the simple high-order compact or non-compact schemes. In this paper, we present the implementation of a seventh-order, minimally-dissipative mapped WENO (WENO7M) scheme in a newly developed direct numerical simulation (DNS) code called KAUST Adaptive Reactive Flows Solver (KARFS). In order to make efficient use of the computer resources and reduce the solution time, without compromising the resolution requirement, the WENO routines are accelerated via graphics processing unit (GPU) computation. The performance characteristics and scalability of the code are studied using different grid sizes and block decomposition. Furthermore, the performance portability of KARFS is demonstrated on a variety of architectures including NVIDIA Tesla P100 GPUs and NVIDIA Kepler K20X GPUs. In addition, the capability and potential of the newly implemented WENO7M scheme in KARFS to perform DNS of compressible flows is also demonstrated with model problems involving shocks, isotropic turbulence, detonations and flame propagation into a stratified mixture with complex chemical kinetics.

97 MATHEMATICS AND COMPUTING↗

Large area transparent refractory aerogels with high solar thermal performance

Application of transparent silica aerogels in low-temperature solar thermal systems has led to major improvements in performance. In high temperature concentrating solar thermal (CST) systems, aerogels have yet to demonstrate the necessary scalability, durability, and performance to support their widespread deployment. Here, large-area transparent refractory aerogel tiles are synthesized and shown to achieve a record-high receiver figure-of-merit (FOM) at high temperatures. The work leverages a scaled-up process for sol–gel synthesis to control the density of the aerogels for improved solar transmittance and adapts a previous atomic layer deposition (ALD) technique with the aid of predictive reaction-transport modeling. After aging for 10 days at 700 °C, the large-area tiles exhibit a solar-weighted transmittance of 95.6 % and a thermal emittance of 0.31, corresponding to a FOM of 80 % at 100 suns and 700 °C. The observed sintering rates at 700 °C are comparably low to earlier one-inch aerogels, suggesting long-term stability under relevant operating conditions. Furthermore, the study indicates that refractory aerogels are scalable materials for efficient photothermal conversion at high temperatures.

Aerogels↗

Massively parallel modeling and inversion of electrical resistivity tomography data using PFLOTRAN

Abstract. Electrical resistivity tomography (ERT) is a broadly accepted geophysical method for subsurface investigations. Interpretation of field ERT data usually requires the application of computationally intensive forward modeling and inversion algorithms. For large-scale ERT data, the efficiency of these algorithms depends on the robustness, accuracy, and scalability on high-performance computing resources. In this regard, we present a robust and highly scalable implementation of forward modeling and inversion algorithms for ERT data. The implementation is publicly available and developed within the framework of PFLOTRAN, an open-source, state-of-the-art massively parallel subsurface flow and transport simulation code. The forward modeling is based on a finite-volume discretization of the governing differential equations, and the inversion uses a Gauss–Newton optimization scheme. To evaluate the accuracy of the forward modeling, two examples are first presented by considering layered (1D) and 3D earth conductivity models. The computed numerical results show good agreement with the analytical solutions for the layered earth model and results from a well-established code for the 3D model. Inversion of ERT data, simulated for a 3D model, is then performed to demonstrate the inversion capability by recovering the conductivity of the model. To demonstrate the parallel performance of PFLOTRAN's ERT process model and inversion capabilities, large-scale scalability tests are performed by using up to 131 072 processes on a leadership class supercomputer. These tests are performed for the two most computationally intensive steps of the ERT inversion: forward modeling and Jacobian computation. For the forward modeling, we consider models with up to 122 ×106 degrees of freedom (DOFs) in the resulting system of linear equations and demonstrate that the code exhibits almost linear scalability on up to 10 000 DOFs per process. On the other hand, the code shows superlinear scalability for the Jacobian computation, mainly because all computations are fairly evenly distributed over each process with no parallel communication.

58 GEOSCIENCES↗

Scalable molecular dynamics on CPU and GPU architectures with NAMD

NAMD is a molecular dynamics program designed for high-performance simulations of very large biological objects on CPU- and GPU-based architectures. NAMD offers scalable performance on petascale parallel supercomputers consisting of hundreds of thousands of cores, as well as on inexpensive commodity clusters commonly found in academic environments. It is written in C++ and leans on Charm++ parallel objects for optimal performance on low-latency architectures. NAMD is a versatile, multipurpose code that gathers state-of-the-art algorithms to carry out simulations in apt thermodynamic ensembles, using the widely popular CHARMM, AMBER, OPLS, and GROMOS biomolecular force fields. Here, we review the main features of NAMD that allow both equilibrium and enhanced-sampling molecular dynamics simulations with numerical efficiency. We describe the underlying concepts utilized by NAMD and their implementation, most notably for handling long-range electrostatics; controlling the temperature, pressure, and pH; applying external potentials on tailored grids; leveraging massively parallel resources in multiple-copy simulations; and hybrid quantum-mechanical/molecular-mechanical descriptions. We detail the variety of options offered by NAMD for enhanced-sampling simulations aimed at determining free-energy differences of either alchemical or geometrical transformations and outline their applicability to specific problems. Last, we discuss the roadmap for the development of NAMD and our current efforts toward achieving optimal performance on GPU-based architectures, for pushing back the limitations that have prevented biologically realistic billion-atom objects to be fruitfully simulated, and for making large-scale simulations less expensive and easier to set up, run, and analyze. NAMD is distributed free of charge with its source code at www.ks.uiuc.edu.

high-performance computing↗