Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “memory mapping”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Experimental Characterization of OpenMP Offloading Memory Operations and Unified Shared Memory Support

The OpenMP specification recently introduced support for unified shared memory, allowing implementation to leverage underlying system software to provide a simpler GPU offloading model where explicit mapping of variables is optional. Support for this feature is becoming more available in different OpenMP implementations on several hardware platforms. A deeper understanding of the different implementation’s execution profile and performance is crucial for applications as they consider the performance portability implications of adopting a unified memory offloading programming style. This work introduces a benchmark tool to characterize unified memory support in several OepnMP compilers and runtimes, with emphasis on identifying discrepancies between different OpenMP implementations as to how they various memory allocation strategies interact with unified shared memory. The benchmark tool is used to characterize OpenMP compilers on three leading High Performance Computing platforms supporting different CPU and device architectures. The benchmark tool is used to assess the impact of enabling unified shared memory on the performance of memory-bound code, highlighting implementation differences that should be accounted for when applications consider performance portability across platforms and compilers.

Elwasif, Wael↗

APNN-TC: Accelerating Arbitrary Precision Neural Networks on Ampere GPU Tensor Cores

Over the years, accelerating neural networks with quantization has been widely studied. Unfortunately, prior efforts with diverse precisions (e.g., 1-bit weights and 2-bit activations) are usually restricted by limited precision support on GPUs (e.g., int1 and int4). To break such restrictions, we introduce the first Arbitrary Precision Neural Network framework (APNN-TC) to fully exploit quantization benefits on Ampere GPU Tensor Cores. Specifically, APNN-TC first incorporates a novel emulation algorithm to support arbitrary short bit-width computation with int1 compute primitives and XOR/AND Boolean operations. Second, APNN-TC integrates arbitrary precision layer designs to efficiently map our emulation algorithm to Tensor Cores with novel batching strategies and specialized memory organization. Third, APNN-TC embodies a novel arbitrary precision NN design to minimize memory access across layers and further improve performance. Extensive evaluations show that APNN-TC can achieve significant speedup over CUTLASS kernels and various NN models, such as ResNet and VGG.

Feng, Boyuan↗

Vienna FORTRAN: A FORTRAN language extension for distributed memory multiprocessors

Exploiting the performance potential of distributed memory machines requires a careful distribution of data across the processors. Vienna FORTRAN is a language extension of FORTRAN which provides the user with a wide range of facilities for such mapping of data structures. However, programs in Vienna FORTRAN are written using global data references. Thus, the user has the advantage of a shared memory programming paradigm while explicitly controlling the placement of data. The basic features of Vienna FORTRAN are presented along with a set of examples illustrating the use of these features.

Chapman, Barbara↗

CD-ROM publication of the Mars digital cartographic data base

The recently completed Mars mosaicked digital image model (MDIM) and the soon-to-be-completed Mars digital terrain model (DTM) are being transcribed to optical disks to simplify distribution to planetary investigators. These models, completed in FY 1991, provide a cartographic base to which all existing Mars data can be registered. The digital image map of Mars is a cartographic extension of a set of compact disk read-only memory (CD-ROM) volumes containing individual Viking Orbiter images now being released. The data in these volumes are pristine in the sense that they were processed only to the extent required to view them as images. They contain the artifacts and the radiometric, geometric, and photometric characteristics of the raw data transmitted by the spacecraft. This new set of volumes, on the other hand, contains cartographic compilations made by processing the raw images to reduce radiometric and geometric distortions and to form geodetically controlled MDIM's. It also contains digitized versions of an airbrushed map of Mars as well as a listing of all feature names approved by the International Astronomical Union. In addition, special geodetic and photogrammetric processing has been performed to derive rasters of topographic data, or DTM's. The latter have a format similar to that of MDIM, except that elevation values are used in the array instead of image brightness values. The set consists of seven volumes: (1) Vastitas Borealis Region of Mars; (2) Xanthe Terra of Mars; (3) Amazonis Planitia Region of Mars; (4) Elysium Planitia Region of Mars; (5) Arabia Terra of Mars; (6) Planum Australe Region of Mars; and (7) a digital topographic map of Mars.

Batson, R. M.↗

Single-Frame Terrain Mapping Software for Robotic Vehicles

This software is a component in an unmanned ground vehicle (UGV) perception system that builds compact, single-frame terrain maps for distribution to other systems, such as a world model or an operator control unit, over a local area network (LAN). Each cell in the map encodes an elevation value, terrain classification, object classification, terrain traversability, terrain roughness, and a confidence value into four bytes of memory. The input to this software component is a range image (from a lidar or stereo vision system), and optionally a terrain classification image and an object classification image, both registered to the range image. The single-frame terrain map generates estimates of the support surface elevation, ground cover elevation, and minimum canopy elevation; generates terrain traversability cost; detects low overhangs and high-density obstacles; and can perform geometry-based terrain classification (ground, ground cover, unknown). A new origin is automatically selected for each single-frame terrain map in global coordinates such that it coincides with the corner of a world map cell. That way, single-frame terrain maps correctly line up with the world map, facilitating the merging of map data into the world map. Instead of using 32 bits to store the floating-point elevation for a map cell, the vehicle elevation is assigned to the map origin elevation and reports the change in elevation (from the origin elevation) in terms of the number of discrete steps. The single-frame terrain map elevation resolution is 2 cm. At that resolution, terrain elevation from 20.5 to 20.5 m (with respect to the vehicle's elevation) is encoded into 11 bits. For each four-byte map cell, bits are assigned to encode elevation, terrain roughness, terrain classification, object classification, terrain traversability cost, and a confidence value. The vehicle s current position and orientation, the map origin, and the map cell resolution are all included in a header for each map. The map is compressed into a vector prior to delivery to another system.

Rankin, Arturo L.↗

AXI4MLIR: User-Driven Automatic Host Code Generation for Custom AXI-Based Accelerators

Tensor algebra operations represent an important class of algorithms used across many applications, including machine learning, scientific computing, and data analytics. As a result, the efficient generation of custom accelerators for tensor operations has received increased attention. Previous efforts have produced automated tools enabling users to prototype and explore optimized accelerators. However, little effort has been focused on the host-accelerator interaction in these tools. Efficient use of hardware accelerators requires knowledge about the accelerator's capabilities (operations, data formats, and opcode support), the host CPU microarchitecture (e.g., memory hierarchy), the host-accelerator interface, and the application's features (which code regions should be mapped onto an accelerator). Manually rewriting the original applications to facilitate improved custom accelerator mapping is an error-prone and time-consuming endeavor. To cope with this, we propose AXI4MLIR, a new framework to automatically generate and optimize the communication between the host CPU and arbitrary accelerators that implement linear algebra algorithms. AXI4MLIR extends the MLIR compiler framework to automatically generate efficient host-accelerator driver code for accelerators with AXI-based interfaces. Our compiler extensions enable automatic driver code generation while carefully considering the host's memory hierarchy and target accelerator features. To demonstrate the flexibility and utility of AXI4MLIR, we test it with diverse use cases that include different types of accelerators, tiling scenarios, and dataflow schemes. We compare our experimental results to manual implementations of host-accelerator driver code and find that our approach can reduce CPU cache references by 56% and deliver up to a 1.65x speedup.

Bohm Agostini, Nicolas↗

Mobile Thread Task Manager

The Mobile Thread Task Manager (MTTM) is being applied to parallelizing existing flight software to understand the benefits and to develop new techniques and architectural concepts for adapting software to multicore architectures. It allocates and load-balances tasks for a group of threads that migrate across processors to improve cache performance. In order to balance-load across threads, the MTTM augments a basic map-reduce strategy to draw jobs from a global queue. In a multicore processor, memory may be "homed" to the cache of a specific processor and must be accessed from that processor. The MTTB architecture wraps access to data with thread management to move threads to the home processor for that data so that the computation follows the data in an attempt to avoid L2 cache misses. Cache homing is also handled by a memory manager that translates identifiers to processor IDs where the data will be homed (according to rules defined by the user). The user can also specify the number of threads and processors separately, which is important for tuning performance for different patterns of computation and memory access. MTTM efficiently processes tasks in parallel on a multiprocessor computer. It also provides an interface to make it easier to adapt existing software to a multiprocessor environment.

Clement, Bradley J.↗

Using Long-Short Term Memory Models to Predict Solar Active Regions Emergence

We train Long Short-Term Memory (LSTM) models that predict the formation of active regions (ARs), the main source of eruptive solar activity. Using the Doppler shift velocity, the continuum intensity and the magnetic field full-disk maps from SDO/HMI we have created time-series datasets of acoustic power and magnetic flux which are used to train Long Short-Term Memory (LSTM) models on predicting decreases in continuum intensity 12 hours in advance. Testing of the models' performance was done on data from 5 ARs, unseen from the model during training. The model predicted the emergence of AR11726, AR13165 and AR13179, 10, 29 and 5 hours in advance, and variations of this model achieved average RMSE values of 0.11 for both active and quiet parts of the solar disc, showing the ability of the model to capture acoustic power anomalies and predict continuum intensity variations. This work sets the foundations for the very first ML-aided prediction of solar ARs.

SMD↗

Understanding the Design Space of Sparse/Dense Multiphase Dataflows for Mapping Graph Neural Networks on Spatial Accelerators

Graph Neural Networks (GNNs) have garnered a lot of recent interest because of their success in learning representations from graph-structured data across several critical applications in cloud and HPC. Owing to their unique compute and memory characteristics that come from an interplay between dense and sparse phases of computations, the emergence of reconfigurable dataflow (aka spatial) accelerators offers promise for acceleration by mapping optimized dataflows (i.e., computation order and parallelism) for both phases. The goal of this work is to characterize and understand the design-space of dataflow choices for running GNNs on spatial accelerators in order for the compilers to optimize the dataflow based on the workload. Specifically, we propose a taxonomy to describe all possible choices for mapping the dense and sparse phases of GNNs spatially and temporally over a spatial accelerator, capturing both the intra-phase dataflow and the inter-phase (pipelined) dataflow. Using this taxonomy, we do deep-dives into the cost and benefits of several dataflows and perform case studies on implications of hardware parameters for dataflows and value of flexibility to support pipelined execution.

97 MATHEMATICS AND COMPUTING↗

Gaussian Mixture Models for Temporal Depth Fusion

Sensing the 3D environment of a moving robot is essential for collision avoidance. Most 3D sensors produce dense depth maps, which are subject to imperfections due to various environmental factors. Temporal fusion of depth maps is crucial to overcome those. Temporal fusion is traditionally done in 3D space with voxel data structures, but it can be approached by temporal fusion in image space, with potential benefits in reduced memory and computational cost for applications like reactive collision avoidance for micro air vehicles. In this paper, we present an efficient Gaussian Mixture Models based depth map fusion approach, introducing an online update scheme for dense representations. The environment is modeled from an ego-centric point of view, where each pixel is represented by a mixture of Gaussian inverse-depth models. Consecutive frames are related to each other by transformations obtained from visual odometry. This approach achieves better accuracy than alternative image space depth map fusion techniques at lower computational cost.

Matthies, Larry↗

Programming in Vienna Fortran

Exploiting the full performance potential of distributed memory machines requires a careful distribution of data across the processors. Vienna Fortran is a language extension of Fortran which provides the user with a wide range of facilities for such mapping of data structures. In contrast to current programming practice, programs in Vienna Fortran are written using global data references. Thus, the user has the advantages of a shared memory programming paradigm while explicitly controlling the data distribution. In this paper, we present the language features of Vienna Fortran for FORTRAN 77, together with examples illustrating the use of these features.

Chapman, Barbara↗

Collector-Output Analysis Program

Collector-Output Analysis Program (COAP) programmer's aid for analyzing output produced by UNIVAC collector (MAP processor). COAP developed to aid in design of segmentation structures for programs with large memory requirements and numerous elements but of value in understanding relationships among components of any program. Crossreference indexes and supplemental information produced. COAP written in FORTRAN 77.

Glandorf, D. R.↗

A Domain-Decomposed Multi-Level Method for Adaptively Refined Cartesian Grids with Embedded Boundaries

The work presents a new method for on-the-fly domain decomposition technique for mapping grids and solution algorithms to parallel machines, and is applicable to both shared-memory and message-passing architectures. It will be demonstrated on the Cray T3E, HP Exemplar, and SGI Origin 2000. Computing time has been secured on all these platforms. The decomposition technique is an outgrowth of techniques used in computational physics for simulations of N-body problems and the event horizons of black holes, and has not been previously used by the CFD community. Since the technique offers on-the-fly partitioning, it offers a substantial increase in flexibility for computing in heterogeneous environments, where the number of available processors may not be known at the time of job submission. In addition, since it is dynamic it permits the job to be repartitioned without global communication in cases where additional processors become available after the simulation has begun, or in cases where dynamic mesh adaptation changes the mesh size during the course of a simulation. The platform for this partitioning strategy is a completely new Cartesian Euler solver tarcreted at parallel machines which may be used in conjunction with Ames' "Cart3D" arbitrary geometry simulation package.

Aftosmis, M. J.↗

Design and Application of Hybrid Magnetic Field-Eddy Current Probe

The incorporation of magnetic field sensors into eddy current probes can result in novel probe designs with unique performance characteristics. One such example is a recently developed electromagnetic probe consisting of a two-channel magnetoresistive sensor with an embedded single-strand eddy current inducer. Magnetic flux leakage maps of ferrous materials are generated from the DC sensor response while high-resolution eddy current imaging is simultaneously performed at frequencies up to 5 megahertz. In this work the design and optimization of this probe will be presented, along with an application toward analysis of sensory materials with embedded ferromagnetic shape-memory alloy (FSMA) particles. The sensory material is designed to produce a paramagnetic to ferromagnetic transition in the FSMA particles under strain. Mapping of the stray magnetic field and eddy current response of the sample with the hybrid probe can thereby image locations in the structure which have experienced an overstrain condition. Numerical modeling of the probe response is performed with good agreement with experimental results.

Wincheski, Buzz↗

Patch2Self2: Self-supervised Denoising on Coresets via Matrix Sketching

Diffusion MRI (dMRI) non-invasively maps brain white matter yet necessitates denoising due to low signal-to-noise ratios. Patch2Self (P2S) employing self-supervised techniques and regression on a Casorati matrix effectively denoises dMRI images and has become the new de-facto standard in this field. P2S however is resource intensive both in terms of running time and memory usage as it uses all voxels (n) from all-but-one held-in volumes (d-1) to learn a linear mapping Phi : \mathbb R ^ n x(d-1) \mapsto \mathbb R ^ n for denoising the held-out volume. The increasing size and dimensionality of higher resolution dMRI acquisitions can make P2S infeasible for large-scale analyses. This work exploits the redundancy imposed by P2S to alleviate its performance issues and inspect regions that influence the noise disproportionately. Specifically this study makes a three-fold contribution: (1) We present Patch2Self2 (P2S2) a method that uses matrix sketching to perform self-supervised denoising. By solving a sub-problem on a smaller sub-space so called coreset we show how P2S2 can yield a significant speedup in training time while using less memory. (2) We present a theoretical analysis of P2S2 focusing on determining the optimal sketch size through rank estimation a key step in achieving a balance between denoising accuracy and computational efficiency. (3) We show how the so-called statistical leverage scores can be used to interpret the denoising of dMRI data a process that was traditionally treated as a black-box. Experimental results on both simulated and real data affirm that P2S2 maintains denoising quality while significantly enhancing speed and memory efficiency achieved by training on a reduced data subset.

Fadnavis, Shreyas↗

Deep Neural Network Based Unsteady Flamelet Progress Variable Approach in a Supersonic Combustor

Higher dimensional flamelet manifolds are essential in capturing the coupled effects of pressure gradients and unsteady chemical kinetics observed in supersonic combustion applications. Previous studies have validated the feasibility of using deep neural networks as an alternative to computation-ally intensive multidimensional flamelet table storage and lookup. This approach has demonstrated a significant reduction in memory footprint and enabled the use of larger dimensional tabulated manifolds for supersonic combustion in canonical problems. In this study, the Unsteady Flamelet Progress Variable (UFPV)-ANN model implemented in the VULCAN-CFD code is validated by the Burrows-Kurkov supersonic mixing/combustion configuration. The well characterized experimental problem consists of hydrogen injection into a supersonic vitiated crossflow that results in a lifted flame structure. The initial model consists of a 4-dimensional table where the independent variables Z, C, Xst, P are tabulated using an unsteady flamelet code with boundary conditions corresponding to the vitiated air conditions. The results show the development of a lifted flame structure and over-all acceptable agreement with finite-rate chemistry (FRC) simulation and the experimental data. Moreover, direct mapping between the independent variables and the flamelet table is replaced by a deep neural network for significant memory reduction. The results indicate that the UFPV-ANN approach can retrieve the same solution as the memory intensive lookup table approach.

Flamelet↗

MrHyDE v.1.0

SAND2024-01324O MrHyDE, which stands for Multi-resolution Hybridized Differential Equations, is a general-purpose C++ package for the solution of coupled multiphysics and multiscale systems on massively parallel computing systems. MrHyDE is designed to enable moving beyond forward simulation for multiscale applications which includes optimization, control, uncertainty quantification, and stochastic inversion. The framework provides interfaces to several packages within the Trilinos framework and leverages automatic differentiation to enable adjoint capabilities for large-scale, gradient-based optimization. MrHyDE provides automated multiscale capabilities through a subgrid model interface and multiscale Dirichlet-to-Neumann maps. For extreme-scale applications, MrHyDE provides in situ data-compression algorithms to reduce memory requirements while maintaining performance. MrHyDE is a general-purpose, computational framework for the solution of multiscale and multiphysics applications. It uses a combination of structure-preserving, physics-compatible discretizations, fully implicit methods, multi-resolution schemes, or fully explicit methods. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

SciDAC↗

The relationship between interstellar dust and the isotopic anomalies in meteorites

Work on the ways in which the isotopic anomalies found in meteorites can be regarded as the chemical memory of even larger anomalies found in interstellar dust is outlined. This approach constitutes one theory of the isotopic anomalies, standing in contrast to the idea of a spatial inhomogeneity in the early solar system owing to inhomogeneous admixture from a neighboring supernova. The four mechanisms of isotopic chemical memory in interstellar dust are: (1) thermal condensation within expanding events of nucleosynthesis; (2) different isotopic mappings onto the grain size spectrum; (3) dust components of differing age; and (4) isotope-dependent interstellar chemistry. Specific examples of each mechanism are given to illustrate how each may have contributed to known isotopic anomalies.

Clayton, D. D.↗