Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “memory mapping”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Picasso: Memory-Efficient Graph Coloring Using Palettes With Applications in Quantum Computing

A coloring of a graph is an assignment of colors to vertices such that no two neighboring vertices have the same color. The need for memory-efficient coloring algorithms is motivated by their application in computing clique partitions of graphs arising in quantum computations where the objective is to map a large set of Pauli strings into a compact set of unitaries. We present Picasso, a randomized memory-efficient iterative parallel graph coloring algorithm with theoretical sublinear space guarantees under practical assumptions. The parameters of our algorithm provide a trade-off between coloring quality and resource consumption. To assist the user, we also propose a machine learning model to predict the coloring algorithm’s parameters considering these trade-offs. We provide a sequential and a parallel implementation of the proposed algorithm. We perform an experimental evaluation on a 64-core AMD CPU equipped with 512 GB of memory and an Nvidia A100 GPU with 40GB of memory. For a small dataset where existing coloring algorithms can be executed within the 512 GB memory budget, we show up to 68× memory savings. On massive datasets we demonstrate that GPU-accelerated Picasso can process inputs with 49.5× more Pauli strings (vertex set in our graph) and 2,478× more edges than state-of-the-art parallel approaches.

artificial intelligence, quantum computing↗

CryoTEN: efficiently enhancing cryo-EM density maps using transformers

Abstract Motivation Cryogenic electron microscopy (cryo-EM) is a core experimental technique used to determine the structure of macromolecules such as proteins. However, the effectiveness of cryo-EM is often hindered by the noise and missing density values in cryo-EM density maps caused by experimental conditions such as low contrast and conformational heterogeneity. Although various global and local map-sharpening techniques are widely employed to improve cryo-EM density maps, it is still challenging to efficiently improve their quality for building better protein structures from them. Results In this study, we introduce CryoTEN—a 3D UNETR++ style transformer to improve cryo-EM maps effectively. CryoTEN is trained using a diverse set of 1295 cryo-EM maps as inputs and their corresponding simulated maps generated from known protein structures as targets. An independent test set containing 150 maps is used to evaluate CryoTEN, and the results demonstrate that it can robustly enhance the quality of cryo-EM density maps. In addition, automatic de novo protein structure modeling shows that protein structures built from the density maps processed by CryoTEN have substantially better quality than those built from the original maps. Compared to the existing state-of-the-art deep learning methods for enhancing cryo-EM density maps, CryoTEN ranks second in improving the quality of density maps, while running >10 times faster and requiring much less GPU memory than them. Availability and implementation The source code and data are freely available at https://github.com/jianlin-cheng/cryoten.

Biochemistry & Molecular Biology↗

Map reduce using coordination namespace hardware acceleration

A system and method for supporting data MapReduce operations in a tuple space/coordinated namespace (CNS) extended memory storage architecture. The system-wide CNS provides an efficient means for storing and communicating data generated by local processes running at the nodes, and coordinated to provide MapReduce operations in a multi-nodal system. A hardware accelerated mechanism supports map reduce sorting/shuffle operations and reduce operations according to an aggregate function. Local processes running at a node generate a tuple corresponding to data generated by a process, each tuple having a tuple name and tuple data value corresponding to the generated data. Each tuple is processed and stored at the node or another node, dependent upon its tuple name. Tuple records associated with a tuple name are accumulated at one or more nodes according to a linked list structure at each that is accessible via a hash table index pointer at the node.

Jacob, Philip↗

Brief Announcement: Communication Optimal Sparse LU Factorization for Planar Matrices

We introduce a new parallel algorithm for solving sparse LU factorization of planar matrices, which commonly arise in the finite element method for 2D PDEs. Existing scalable methods, such as the multifrontal approach with subtree-to-subcube mapping by Gupta et al. [1] and right-looking with 3D mapping by Sao et al. [2] fail to achieve optimal communication costs for these matrices. Our new algorithm combines 3D mapping and subtree-to-subcube mapping to minimize communication costs while allowing trade-offs between extra memory and reduced communication. We demonstrate that our proposed algorithm attains the communication lower bound up to a factor of O(log log n) in the memory-optimal case and up to a factor of O(log P) in the memory-independent case for an n-dimensional planar sparse matrix on P processors.

Sao, Piyush↗

Optical quantum memory for noble-gas spins based on spin-exchange collisions

Optical quantum memories, which store and preserve the quantum state of photons, rely on a coherent mapping of the photonic state onto matter states that are optically accessible. Here we outline and characterize schemes to map the state of photons onto long-lived but optically inaccessible collective states of noble-gas spins. The mapping employs coherent spin-exchange interaction arising from random collisions with alkali vapor. We propose efficient storage strategies in two operating regimes and analyze their performance for several proposed experimental configurations.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Experimental Characterization of OpenMP Offloading Memory Operations and Unified Shared Memory Support

The OpenMP specification recently introduced support for unified shared memory, allowing implementation to leverage underlying system software to provide a simpler GPU offloading model where explicit mapping of variables is optional. Support for this feature is becoming more available in different OpenMP implementations on several hardware platforms. A deeper understanding of the different implementation’s execution profile and performance is crucial for applications as they consider the performance portability implications of adopting a unified memory offloading programming style. This work introduces a benchmark tool to characterize unified memory support in several OepnMP compilers and runtimes, with emphasis on identifying discrepancies between different OpenMP implementations as to how they various memory allocation strategies interact with unified shared memory. The benchmark tool is used to characterize OpenMP compilers on three leading High Performance Computing platforms supporting different CPU and device architectures. The benchmark tool is used to assess the impact of enabling unified shared memory on the performance of memory-bound code, highlighting implementation differences that should be accounted for when applications consider performance portability across platforms and compilers.

Elwasif, Wael↗

APNN-TC: Accelerating Arbitrary Precision Neural Networks on Ampere GPU Tensor Cores

Over the years, accelerating neural networks with quantization has been widely studied. Unfortunately, prior efforts with diverse precisions (e.g., 1-bit weights and 2-bit activations) are usually restricted by limited precision support on GPUs (e.g., int1 and int4). To break such restrictions, we introduce the first Arbitrary Precision Neural Network framework (APNN-TC) to fully exploit quantization benefits on Ampere GPU Tensor Cores. Specifically, APNN-TC first incorporates a novel emulation algorithm to support arbitrary short bit-width computation with int1 compute primitives and XOR/AND Boolean operations. Second, APNN-TC integrates arbitrary precision layer designs to efficiently map our emulation algorithm to Tensor Cores with novel batching strategies and specialized memory organization. Third, APNN-TC embodies a novel arbitrary precision NN design to minimize memory access across layers and further improve performance. Extensive evaluations show that APNN-TC can achieve significant speedup over CUTLASS kernels and various NN models, such as ResNet and VGG.

Feng, Boyuan↗

AXI4MLIR: User-Driven Automatic Host Code Generation for Custom AXI-Based Accelerators

Tensor algebra operations represent an important class of algorithms used across many applications, including machine learning, scientific computing, and data analytics. As a result, the efficient generation of custom accelerators for tensor operations has received increased attention. Previous efforts have produced automated tools enabling users to prototype and explore optimized accelerators. However, little effort has been focused on the host-accelerator interaction in these tools. Efficient use of hardware accelerators requires knowledge about the accelerator's capabilities (operations, data formats, and opcode support), the host CPU microarchitecture (e.g., memory hierarchy), the host-accelerator interface, and the application's features (which code regions should be mapped onto an accelerator). Manually rewriting the original applications to facilitate improved custom accelerator mapping is an error-prone and time-consuming endeavor. To cope with this, we propose AXI4MLIR, a new framework to automatically generate and optimize the communication between the host CPU and arbitrary accelerators that implement linear algebra algorithms. AXI4MLIR extends the MLIR compiler framework to automatically generate efficient host-accelerator driver code for accelerators with AXI-based interfaces. Our compiler extensions enable automatic driver code generation while carefully considering the host's memory hierarchy and target accelerator features. To demonstrate the flexibility and utility of AXI4MLIR, we test it with diverse use cases that include different types of accelerators, tiling scenarios, and dataflow schemes. We compare our experimental results to manual implementations of host-accelerator driver code and find that our approach can reduce CPU cache references by 56% and deliver up to a 1.65x speedup.

Bohm Agostini, Nicolas↗

Understanding the Design Space of Sparse/Dense Multiphase Dataflows for Mapping Graph Neural Networks on Spatial Accelerators

Graph Neural Networks (GNNs) have garnered a lot of recent interest because of their success in learning representations from graph-structured data across several critical applications in cloud and HPC. Owing to their unique compute and memory characteristics that come from an interplay between dense and sparse phases of computations, the emergence of reconfigurable dataflow (aka spatial) accelerators offers promise for acceleration by mapping optimized dataflows (i.e., computation order and parallelism) for both phases. The goal of this work is to characterize and understand the design-space of dataflow choices for running GNNs on spatial accelerators in order for the compilers to optimize the dataflow based on the workload. Specifically, we propose a taxonomy to describe all possible choices for mapping the dense and sparse phases of GNNs spatially and temporally over a spatial accelerator, capturing both the intra-phase dataflow and the inter-phase (pipelined) dataflow. Using this taxonomy, we do deep-dives into the cost and benefits of several dataflows and perform case studies on implications of hardware parameters for dataflows and value of flexibility to support pipelined execution.

97 MATHEMATICS AND COMPUTING↗

Patch2Self2: Self-supervised Denoising on Coresets via Matrix Sketching

Diffusion MRI (dMRI) non-invasively maps brain white matter yet necessitates denoising due to low signal-to-noise ratios. Patch2Self (P2S) employing self-supervised techniques and regression on a Casorati matrix effectively denoises dMRI images and has become the new de-facto standard in this field. P2S however is resource intensive both in terms of running time and memory usage as it uses all voxels (n) from all-but-one held-in volumes (d-1) to learn a linear mapping Phi : \mathbb R ^ n x(d-1) \mapsto \mathbb R ^ n for denoising the held-out volume. The increasing size and dimensionality of higher resolution dMRI acquisitions can make P2S infeasible for large-scale analyses. This work exploits the redundancy imposed by P2S to alleviate its performance issues and inspect regions that influence the noise disproportionately. Specifically this study makes a three-fold contribution: (1) We present Patch2Self2 (P2S2) a method that uses matrix sketching to perform self-supervised denoising. By solving a sub-problem on a smaller sub-space so called coreset we show how P2S2 can yield a significant speedup in training time while using less memory. (2) We present a theoretical analysis of P2S2 focusing on determining the optimal sketch size through rank estimation a key step in achieving a balance between denoising accuracy and computational efficiency. (3) We show how the so-called statistical leverage scores can be used to interpret the denoising of dMRI data a process that was traditionally treated as a black-box. Experimental results on both simulated and real data affirm that P2S2 maintains denoising quality while significantly enhancing speed and memory efficiency achieved by training on a reduced data subset.

Fadnavis, Shreyas↗

MrHyDE v.1.0

SAND2024-01324O MrHyDE, which stands for Multi-resolution Hybridized Differential Equations, is a general-purpose C++ package for the solution of coupled multiphysics and multiscale systems on massively parallel computing systems. MrHyDE is designed to enable moving beyond forward simulation for multiscale applications which includes optimization, control, uncertainty quantification, and stochastic inversion. The framework provides interfaces to several packages within the Trilinos framework and leverages automatic differentiation to enable adjoint capabilities for large-scale, gradient-based optimization. MrHyDE provides automated multiscale capabilities through a subgrid model interface and multiscale Dirichlet-to-Neumann maps. For extreme-scale applications, MrHyDE provides in situ data-compression algorithms to reduce memory requirements while maintaining performance. MrHyDE is a general-purpose, computational framework for the solution of multiscale and multiphysics applications. It uses a combination of structure-preserving, physics-compatible discretizations, fully implicit methods, multi-resolution schemes, or fully explicit methods. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

SciDAC↗

Crystallographic variant mapping using precession electron diffraction data

In this work, we developed three methods to map crystallographic variants of samples at the nanoscale by analyzing precession electron diffraction data using a high-temperature shape memory alloy and a VO2 thin film on sapphire as the model systems. The three methods are (I) a user-selecting-reference pattern approach, (II) an algorithm-selecting-reference-pattern approach, and (III) a k-means approach. In the first two approaches, Euclidean distance, Cosine, and Structural Similarity (SSIM) algorithms were assessed for the diffraction pattern similarity quantification. We demonstrated that the Euclidean distance and SSIM methods outperform the Cosine algorithm. We further revealed that the random noise in the diffraction data can dramatically affect similarity quantification. Denoising processes could improve the crystallographic mapping quality. With the three methods mentioned above, we were able to map the crystallographic variants in different materials systems, thus enabling fast variant number quantification and clear variant distribution visualization. The advantages and disadvantages of each approach are also discussed. We expect these methods to benefit researchers who work on martensitic materials, in which the variant information is critical to understand their properties and functionalities.

Crystallographic variant mapping↗

CORE-BFS: Communication-Optimized REctangular-partitioned BFS Achieving 160.845 TeraTEPS on Frontier Supercomputer

Distributed Breadth-First Search (BFS) is fundamental to many large-scale graph applications, but its performance on parallel systems is often limited by high communication overhead. This paper presents CORE-BFS, an extremely scalable GPU-based BFS implementation that introduces a unique rectangular 2D partitioning-based design for Frontier supercomputer. To further improve performance, we propose four key optimizations: (1) Rectangular 2D-partition specific data formats that use two compressed row and one compressed column status array bitmaps combined with a Double Compressed Sparse Row (DCSR) format per partition, reducing memory footprint and inter-rank traffic; (2) Adaptive frontier & communication strategy that unifies top-down and bottom-up traversal on the rectangular layout, uses lazy synchronization in top-down levels, and switches variants based on frontier size to minimize communication overhead; (3) Frontier-split degree-aware update that maps frontier vertices to thread-centric, wavefront-centric, and block-centric kernels based on their degree to improve GPU utilization and memory coalescing; (4) Row-reduction pipeline that overlaps bottom-up adjacency list processing with row-wise bitmap reduction to hide inter-rank latency. Together, these techniques increase parallelism while reducing memory and communication overhead. On the Graph500 benchmark, CORE - BFS scales up to 9,248 Frontier nodes with scale-42 graphs and reaches 160.845 TTEPS, delivering a 5.42 × speedup over our previous Frontier implementation.

Yang, Haoshen [Rutgers University]↗

Registration and fusion of large-scale melt pool temperature and morphology monitoring data demonstrated for surface topography prediction in LPBF

In-situ monitoring technologies for laser powder bed fusion (LPBF) additive manufacturing often face one key challenge, extracting the ultrafast melt pool (MP) signatures for understanding the localized part properties. Further, the spatial information of each monitored MP signature is essential for correlating the MP – part property. This spatial information is often unavailable especially from commercial LPBF printers. Many MP monitoring methods have been reported and utilized. However, very few of these have the MP’s spatial information. To overcome this challenge, in this work we report a method for spatially registering the key MP signatures (MP intensity, temperature, and area) to the monitored print parts. The MP signatures are obtained from our coaxial high-speed single-camera based two-wavelength imaging pyrometry (STWIP) system and the MP spatial information is obtained from an off-axis camera system. A machine learning aided image analysis method is employed to retrieve the spatial distribution of MPs within the corresponding part’s coordinates system. Then, the MP signature maps (MPSMs) are reconstructed by mapping the STWIP measured MP signatures to the registered MP coordinates. Further, a long short-term memory (LSTM) neural network is developed for estimating the layer surface topography from the registered MPSMs. The obtained results indicate that the layer surface topography can be more accurately estimated by using MP temperature signature rather than MP intensity and/or area signatures as in common practice. Finally, our developed methods for MP monitoring, registration, and MP-surface topography prediction offer advanced capabilities for the online detection of process anomalies and part defects.

36 MATERIALS SCIENCE↗

Wavelet and Deep-Learning-Based Approach for Generation System Problematic Parameters Identification and Calibration

Accurate models of generation systems are critical for maintaining reliable and secure grid operations. In this paper, a novel and systematic approach is proposed to identify and calibrate the generation system problematic parameters using continuous wavelet transform (CWT) and advanced deep-learning technology. The phasor measurement unit (PMU) data are used through “event playback” to check whether the parameter calibration is required, and if yes, a group of suspicious parameters will be identified as the primary problematic parameter candidates (PPCs). These primary PPCs are randomly perturbed to generate the event playback simulation data, which are used by the CWT and convolutional neural networks (CNNs) to further narrow down the primary PPCs into a smaller set of candidates. Then, the identified candidates are perturbed again to generate massive event playback simulation data for training a parameter calibration neural network. Here, we designed a multi-output neural network structure to find the mappings between the perturbed parameters and the simulation data using both CNN and long short-term memory (LSTM) models. Finally, the well-trained and tested CNN-LSTM model is used to estimate the accurate value of the suspicious parameters with actual PMU measurements. The proposed CNN-LSTM network can accurately and reliably estimate the generation-system problematic parameters, and has better performance when compared to other machine-learning methods, such as the multilayer perceptron network and the conditional variational autoencoder method. The accuracy and effectiveness of the proposed approach have been validated through simulation and real-world data.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Multi-frequency electrical impedance tomography

Apparatus includes a plurality of geological subsurface electrical line sensors spaced apart from each other proximate a predetermined geological subsurface region of interest, with at least one of the electrical line sensors situated as a line source to produce a multi-frequency electrical impedance tomography source signal, and with at least one of the electrical line sensors situated as a line detector to receive the multi-frequency electrical impedance tomography response signal associated with the source signal that propagates through the predetermined geological subsurface region of interest, and a controller including a processor and a memory configured with instructions that, when executed by the processor, cause the processor to determine an electrical mapping over the predetermined geological subsurface region of interest based on the multi-frequency electrical impedance tomography source signal, response signal, and the spatial positions of the geological subsurface electrical line sensors.

Karra, Satish↗

Scalable In Situ Computation of Lagrangian Representations via Local Flow Maps

In situ computation of Lagrangian flow maps to enable post hoc time-varying vector field analysis has recently become an active area of research. However, the current literature is largely limited to theoretical settings and lacks a solution to address scalability of the technique in distributed memory. To improve scalability, we propose and evaluate the benefits and limitations of a simple, yet novel, performance optimization. Our proposed optimization is a communication-free model resulting in local Lagrangian flow maps, requiring no message passing or synchronization between processes, intrinsically improving scalability, and thereby reducing overall execution time and alleviating the encumbrance placed on simulation codes from communication overheads. To evaluate our approach, we computed Lagrangian flow maps for four time-varying simulation vector fields and investigated how execution time and reconstruction accuracy are impacted by the number of GPUs per compute node, the total number of compute nodes, particles per rank, and storage intervals. Our study consisted of experiments computing Lagrangian flow maps with up to 67M particle trajectories over 500 cycles and used as many as 2048 GPUs across 512 compute nodes. In all, our study contributes an evaluation of a communication-free model as well as a scalability study of computing distributed Lagrangian flow maps at scale using in situ infrastructure on a modern supercomputer.

Sane, Sudhanshu↗

Physics in the Machine: Integrating Physical Knowledge in Autonomous Phase-Mapping

Application of artificial intelligence (AI), and more specifically machine learning, to the physical sciences has expanded significantly over the past decades. In particular, science-informed AI, also known as scientific AI or inductive bias AI, has grown from a focus on data analysis to now controlling experiment design, simulation, execution and analysis in closed-loop autonomous systems. The CAMEO (closed-loop autonomous materials exploration and optimization) algorithm employs scientific AI to address two tasks: learning a material system’s composition-structure relationship and identifying materials compositions with optimal functional properties. By integrating these, accelerated materials screening across compositional phase diagrams was demonstrated, resulting in the discovery of a best-in-class phase change memory material. Key to this success is the ability to guide subsequent measurements to maximize knowledge of the composition-structure relationship, or phase map. In this work we investigate the benefits of incorporating varying levels of prior physical knowledge into CAMEO’s autonomous phase-mapping. This includes the use of ab-initio phase boundary data from the AFLOW repositories, which has been shown to optimize CAMEO’s search when used as a prior.

97 MATHEMATICS AND COMPUTING↗