Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parallel processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Multi-fidelity is the new annealing: Gradient-free learning of posterior densities via transport maps

To tackle concentrated, multi-modal Bayesian inference problems, we propose using an annealed importance sampling procedure. To do this, we form a sequence of annealed distributions and employ transport maps to act as a surrogate of each distribution. This process is demonstrated on a few examples, favorably showing its potential for efficiently parallelizing the process of PDE evaluations and allowing for surrogates that we can sample from exactly.

van Bloemen Waanders, Bart G [Sandia National Labo↗

Parallel interior-point solver for block-structured nonlinear programs on SIMD/GPU architectures

Here, we investigate how to port the standard interior-point method to new exascale architectures for block-structured nonlinear programs with state equations. Computationally, we decompose the interior-point algorithm into two successive operations: the evaluation of the derivatives and the solution of the associated Karush-Kuhn-Tucker (KKT) linear system. Our method accelerates both operations using two levels of parallelism. First, we distribute the computations on multiple processes using coarse parallelism. Second, each process uses SIMD/GPU accelerators locally to accelerate the operations using fine-grained parallelism. The KKT system is reduced by eliminating the inequalities and the state variables from the corresponding equations. We demonstrate our method's capability on the supercomputer Polaris, a testbed for the future exascale Aurora system. Each node is equipped with four GPUs, a setup amenable to our two-level approach. Our experiments on the stochastic optimal power flow problem show that the reduction method is 50x faster than the sparse linear solver HSL MA57 running in serial on the CPU, and 6x faster than Pardiso running in parallel on CPU on the same number of processes.

97 MATHEMATICS AND COMPUTING↗

A Flexible Forwarding Scheme to Improve Latency-Bound Irregular P2P Communication in MPI

We propose an algorithm to efficiently perform latency-bound communication scenarios that consist of many small messages. In these parallel scenarios, processes typically pass around a lot of small-sized messages of a few KBs of size. Performing communication operations with P2P MPI routines or collective MPI routines (including neighborhood collectives) in such scenarios may not always yield the optimal results and may not resolve the latency bottleneck. To this end, we develop a regular structure called virtual process topology (VPT) on which the messages can be communicated in a structured and controlled manner. Using parameters of this topology, one can tune the rate of aggression in tackling the latency costs. We demonstrate that our communication algorithm is preferable to MPI P2P and collective routines for latency-bound communication and it can easily be adapted only by replacing calls to MPI routines in a parallel application. We show how to adapt existing topology-aware mapping heuristics to address the volume overhead due to communicating messages on the VPT. Moreover, we propose a novel swap-based mapping heuristic to address this overhead by optimizing the maximum volume handled by a process. Experiments on synthetic communication graphs as well as real-world applications such as parallel Canonical Polyadic sparse tensor decomposition and parallel sparse matrix-dense matrix multiplication show that our approach is a powerful way of overcoming the bottlenecks posed by sparse and latency-bound irregular communication.

communication algorithm↗

Human Host Cellular Response to HCoV-229E Infection Proteomics (ACS-JM-DP2)

The purpose of this experiment was to evaluate the human host cellular response to wild-type Human coronavirus strain 229E (HCoV-229E) infection. Sample data was obtained for mock and infected immortalized human lung epithelial cells (A549) (MOI 5) nuclear extracts, immortalized human lung fibroblasts cells (MRC5) (MOI5) nuclear extracts, and primary human airway epithelial (HAE) (MOI 3) cells from lung tissue and processed for proteome analysis. Processed datasets are openly accessible from the download button and contain secondary processed proteomic results files and supporting metadata materials. Experimental proteomics samples were prepared using Limited Proteolysis (LiP) methods for Label-free quantification (LFQ) and global proteomic evaluation. Sample data was acquired using a Q-Exactive HF-X mass spectrometer and was processed and compiled using MaxQuant software (v.1.6.17.0). Processed proteomic data downloads include a sample naming key, processed MaxQuant results/parameters, and protein annotated relative abundance files. See corresponding primary data accessions below and Viral Experiment LiP Analysis source code supporting data transparency and reuse. Experimental transcriptomics samples were collected in parallel and processed for RNA sequencing (RNA-Seq) as summarized under ACS-DP1 (https://data.pnnl.gov/group/nodes/dataset/34069).

59 BASIC BIOLOGICAL SCIENCES↗

Global ion heating/transport during merging spherical tokamak formation

Here we report global ion heating/transport characteristics of magnetic reconnection during merging spherical tokamak formation experiment on TS-6 (TS-3U). Using the 96CH/320CH ultra high resolution ion Doppler tomography diagnostics, the full- 2 D imaging measurement clearly revealed that magnetic reconnection initially forms localized hot spots in the downstream region of outflow jet with inboard/outboard asymmetry (more deposition in the high field side) but the continuous accumulation of the heating coupled with transport process expands the high temperature region globally and forms characteristic poloidally ring-like structure aligned with field lines. The dynamic ion heating/transport process is also affected by the polarity of toroidal field and poloidally tilted/rotating global structure has experimentally been found both during and after merging. The characteristic poloidal asymmetry gets flipped when toroidal field direction is reversed and it was found that higher temperature appears in the positive potential side, which is opposite to the conventional understanding/prediction of guide field reconnection. Through the parallel acceleration process coupled with global heat transport, poloidally asymmetric non-classical feature has experimentally been found for the first time.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Does plasma jet sintering follow an Arrhenius-type expression?

Atmospheric pressure, ambient temperature plasma jets have become a promising candidate for material processing in parallel with developments in additive manufacturing. Recent work has shown that plasma jets can be used to sinter printed nanoparticles at temperatures much lower than typically required for conventional thermal sintering. In this report we conduct a mechanistic study on plasma jet sintering that correlates specific energy input with the electrical conductivity of printed silver films after sintering. Increasing the specific energy input accelerated the sintering process following an Arrhenius-like exponential trend across a large range of conditions, including both helium and argon plasma jets. Although an exponential relationship is also found with the plasma heated substrate temperature, independent studies indicate that heating is not the primary mechanism. These results suggest there is a general behavior that couples the plasma jet with the surface.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Trust: Triangle Counting Reloaded on GPUs

Triangle counting is a building block for a wide range of graph applications. Here, traditional wisdom suggests that i) hashing is not suitable for triangle counting, ii) edge-centric triangle counting beats vertex-centric design, and iii) communication-free and workload balanced graph partitioning is a grand challenge for triangle counting. On the contrary, we advocate that i) hashing can help the key operations for scalable triangle counting on Graphics Processing Units (GPUs), i.e., list intersection and graph partitioning, ii) vertex-centric option reduces both hash table construction cost and memory consumption, which is limited on GPUs. In addition, iii) we exploit graph and workload collaborative, and hash-based 2D partitioning to scale vertex-centric triangle counting over 1,000 GPUs with sustained scalability. In this work, we present TRUST, which performs triangle counting with the hash operation and vertex-centric paradigm. To the best of our knowledge, TRUST is the first work that achieves over one trillion Traversed Edges Per Second (TEPS) rate for triangle counting.

97 MATHEMATICS AND COMPUTING↗

Phlex: Parallel, Hierarchical, and Layered EXecution of data-processing algorithms

Phlex is a computing framework supporting the parallel, hierarchical, and layered execution of data-processing algorithms. It is based on the functional-programming paradigm, thus guaranteeing thread-safety when invoking user-defined pure functions. Phlex allows users to specify arbitrary graph-based hierarchies of data organization, enabling more flexible processing of data as required by the constraints of the program.

Knoepfel, KyleJ. [Fermi National Accelerator Labor↗

Performance of Julia for High Energy Physics Analyses

We argue that the Julia programming language is a compelling alternative to currently more common implementations in Python and C++ for common data analysis workflows in high energy physics. We compare the speed of implementations of different workflows in Julia with those in Python and C++. Furthermore, our studies show that the Julia implementations are competitive for tasks that are dominated by computational load rather than data access. For work that is dominated by data access, we demonstrate an application with concurrent file reading and parallel data processing.

97 MATHEMATICS AND COMPUTING↗

3-D Simulations of earthquakes rupture jumps: 1. Homogeneous pre-stress conditions

SUMMARY Observational and modelling studies indicate that earthquake ruptures can jump between fault sections as large as ∼3 and ∼5 km for compressional and extensional offsets, respectively. Here, we compare characteristics of the rupture jump process on parallel but offset fault sections from traditional 3-D dynamic rupture simulations governed by slip weakening friction using the finite element code, FaultMod, to those from quasi-dynamic simulations governed by rate- and state-dependent friction (rate-state friction) using the code RSQSim. These simulations use spatially uniform initial stresses. For a variety of measures the rupture renucleation position on the offset fault, the rate-state friction and slip weakening friction models produce very similar results. The principal difference is the additional occurrence of delayed rupture jumps that arise from the time- and stress-dependent nucleation that is characteristic of rate-state friction. For immediate rupture jumps, models with slip weakening friction span greater offsets than those with rate-state friction. However the jump distances are nearly identical when delayed rupture jumps are included in the comparisons. We propose that delayed rupture jumps are the likely mechanism for adjacent large-earthquake pairs and clusters. Based on the similarity of renucleation positions with both dynamic and quasi-dynamic models, we conclude that the renucleation positions for rupture initiation on the receiver fault (separated by less than ∼3 km from the source fault) are primarily controlled by static stress changes induced by slip on the initiating fault. However, in light of the slightly greater maximum jump distances (>3 km) seen with the dynamic slip weakening friction model, dynamic stress changes from seismic waves play an increasingly important role as offset distances increase.

, RSQSim↗

Communication-Avoiding and Memory-Constrained Sparse Matrix-Matrix Multiplication at Extreme Scale

Sparse matrix-matrix multiplication (SpGEMM) is a widely used kernel in various graph, scientific computing and machine learning algorithms. In this paper, we consider SpGEMMs performed on hundreds of thousands of processors generating trillions of nonzeros in the output matrix. Distributed SpGEMM at this extreme scale faces two key challenges: (1) high communication cost and (2) inadequate memory to generate the output. Furthermore, we address these challenges with an integrated communication-avoiding and memory-constrained SpGEMM algorithm that scales to 262,144 cores (more than 1 million hardware threads) and can multiply sparse matrices of any size as long as inputs and a fraction of output fit in the aggregated memory. As we go from 16,384 cores to 262,144 cores on a Cray XC40 supercomputer, the new SpGEMM algorithm runs 10x faster when multiplying large-scale protein-similarity matrices.

97 MATHEMATICS AND COMPUTING↗

Sparse Binary Matrix-Vector Multiplication on Neuromorphic Computers

Neuromorphic computers offer the opportunity for low-power, efficient computation. Though they have been primarily applied to neural network tasks, there is also the opportunity to leverage the inherent characteristics of neuromorphic computers (low power, massive parallelism, collocated processing and memory) to perform non-neural network tasks. Here, we demonstrate how an approach for performing sparse binary matrix-vector multiplication on neuromorphic computers. We describe the approach, which relies on the connection between binary matrix-vector multiplication and breadth first search, and we introduce the algorithm for performing this calculation in a neuromorphic way. We validate the approach in simulation. Finally, we provide a discussion of the runtime of this algorithm and discuss where neuromorphic computers in the future may have a computational advantage when performing this computation.

Schuman, Catherine↗

New Horizons for High-Performance Computing

Here we provide an overview of the past, present, and a diverse collection of future computer architecture alternatives for HPC. The end of Moore’s Law influenced the current HPC architecture focus on accelerated compute nodes composed of CPU and GPU computing components integrated into massively parallel processor architecture systems. There are many alternatives for future HPC directions, with different technologies, computing ecosystems, opportunities for lead user application-driven customization, and the role of open innovation business models. This paper provides an overview of these different new horizons for HPC, an organizing principle to focus future computing research, different public-private partnership models, and the critical role of workforce development.

97 MATHEMATICS AND COMPUTING↗

The U.S. High-Performance Computing Consortium in the Fight Against COVID-19

U.S. computing leaders, including Department of Energy National Laboratories, have partnered with universities, government agencies, and the private sector to research responses to COVID-19, providing an unprecedented collection of resources that include some of the fastest computers in the world. For HPC users, these leadership machines will drive the AI to accelerate the discovery of promising treatments, enable at-scale simulations to understand the virus’s protein structure and attack mechanisms, and help inform policymakers to deploy resources effectively.

60 APPLIED LIFE SCIENCES↗

Structured Adaptive Mesh Refinement Adaptations to Retain Performance Portability With Increasing Heterogeneity

Adaptive mesh refinement (AMR) is an important method that enables many mesh-based applications to run at effectively higher resolution within limited computing resources by allowing high resolution only where really needed. This advantage comes at a cost, however: greater complexity in the mesh management machinery and challenges with load distribution. With the current trend of increasing heterogeneity in hardware architecture, AMR presents an orthogonal axis of complexity. Additionally, the usual techniques, such as asynchronous communication and hierarchy management for parallelism and memory that are necessary to obtain reasonable performance are very challenging to reason about with AMR. Different groups working with AMR are bringing different approaches to this challenge. Here, we examine the design choices of several AMR codes and also the degree to which demands placed on them by their users influence these choices.

42 ENGINEERING↗

It’s Time to Talk About HPC Storage: Perspectives on the Past and Future

High-performance computing (HPC) storage systems are a key component of the success of HPC to date. Recently, we have seen major developments in storage-related technologies, as well as changes to how HPC platforms are used, especially in relation to artificial intelligence and experimental data analysis workloads. Additionally, these developments merit a revisit of HPC storage system architectural designs. In this article, we discuss the drivers, identify key challenges to status quo posed by these developments, and discuss directions future research might take to unlock the potential of new technologies for the breadth of HPC applications.

97 MATHEMATICS AND COMPUTING↗

A Robust Parallel Distributed State Estimation for Large Scale Distribution Systems

The growing need and interest in real-time monitoring of large distribution networks motivated by the rapid population of renewable sources, EVs and etc. demand a computationally efficient state estimation framework. Furthermore, this paper presents an improved computational framework for implementing a robust state estimator using a multi-core processor. The main contribution of the paper is the proposed computational framework along with two partitioning strategies which enable fast and robust state estimation for large scale radial and/or meshed distribution systems. Formulation of the proposed method and its implementation are described in detail. Performance of the estimator is tested by simulations first using a small 84-bus radial distribution system. Then the method’s scalability is demonstrated by simulations on two very large scale distribution networks one configured radially and the other meshed each containing over 12,500 buses.

42 ENGINEERING↗