Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Read Only Memory”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Field Programmable Gate Array Apparatus, Method, and Computer Program

An apparatus is provided that includes a plurality of modules, a plurality of memory banks, and a multiplexor. Each module includes at least one agent that interfaces between a module and a memory bank. Each memory bank includes an arbiter that interfaces between the at least one agent of each module and the memory bank. The multiplexor is configured to assign data paths between the at least one agent of each module and a corresponding arbiter of each memory bank based on the assigned data path. The at least one agent of each module is configured to read data from the corresponding arbiter of the memory bank or write modified data to the corresponding arbiter of the memory bank.

Morfopoulos, Arin C.↗

Out-of-Core Streamline Visualization on Large Unstructured Meshes

It's advantageous for computational scientists to have the capability to perform interactive visualization on their desktop workstations. For data on large unstructured meshes, this capability is not generally available. In particular, particle tracing on unstructured grids can result in a high percentage of non-contiguous memory accesses and therefore may perform very poorly with virtual memory paging schemes. The alternative of visualizing a lower resolution of the data degrades the original high-resolution calculations. This paper presents an out-of-core approach for interactive streamline construction on large unstructured tetrahedral meshes containing millions of elements. The out-of-core algorithm uses an octree to partition and restructure the raw data into subsets stored into disk files for fast data retrieval. A memory management policy tailored to the streamline calculations is used such that during the streamline construction only a very small amount of data are brought into the main memory on demand. By carefully scheduling computation and data fetching, the overhead of reading data from the disk is significantly reduced and good memory performance results. This out-of-core algorithm makes possible interactive streamline visualization of large unstructured-grid data sets on a single mid-range workstation with relatively low main-memory capacity: 5-20 megabytes. Our test results also show that this approach is much more efficient than relying on virtual memory and operating system's paging algorithms.

Ueng, Shyh-Kuang↗

NOR Flash Memory Scrubbing Application for Boot File Preservation of NASA’s Descent and Landing Computer (DLC)

Progress on NASA’s Safe and Precise Landing Integrated Capabilities Evolution (SPLICE)project continues, specifically with this development of the Descent and Landing Computer(DLC). One of the DLC’s primary contributions as a SPLICE technology is its implementationof algorithms and operation of sensors to autonomously guide a spacecraft in performing moreprecise and safer landings on celestial bodies such as the Moon and Mars. The second iterationof the DLC is known as the Engineering Test Unit (ETU) and one of its desired functionalitiesis the ability to preserve the fidelity of the system’s boot file through the use of memoryscrubbing. The ETU has two primary boards, one for housing a Multi-Processor System ona Chip (MPSoC) and the other for housing a Xilinx Kintex Ultrascale FPGA3. To emulate amemory scrubbing function implemented on the ETU’s FPGA board, the design and testingof a software application was performed on a Xilinx KCU105 FPGA evaluation board. Thememory scrubbing application had to meet certain key criteria such as (1) properly utilize withthe flash memory’s Serial Peripheral Interface (SPI) to read, write, and erase flash memoryproperly, (2) be able to detect arbitrarily large or small amounts of bit-errors, (3) be able tocorrect all detected errors, and (4) perform memory scrubbing indefinitely and autonomously.A prototype implementation was constructed and tested, demonstrating successful detectionand correction of bit errors in multiple configurations. In the form of burst errors or singularbit flips, and in amounts of errors ranging from one to fifteen (per 256 Bytes), the applicationwas successful in preserving memory fidelity.

Radiation tolerant↗

Performance Potential of Mixed Data Management Modes for Heterogeneous Memory Systems

Many high-performance systems now include different types of memory devices within the same compute platform to meet strict performance and cost constraints. Such heterogeneous memory systems often include an upper-level tier with better performance, but limited capacity, and lower-level tiers with higher capacity, but less bandwidth and longer latencies for reads and writes. To utilize the different memory layers efficiently, current systems rely on hardware-directed, memory -side caching or they provide facilities in the operating system (OS) that allow applications to make their own data-tier assignments. Since these data management options each come with their own set of trade-offs, many systems also include mixed data management configurations that allow applications to employ hardware- and software-directed management simultaneously, but for different portions of their address space. Despite the opportunity to address limitations of stand-alone data management options, such mixed management modes are under-utilized in practice, and have not been evaluated in prior studies of complex memory hardware. In this work, we develop custom program profiling, configurations, and policies to study the potential of mixed data management modes to outperform hardware- or software-based management schemes alone. Our experiments, conducted on an Intel ® Knights Landing platform with high-bandwidth memory, demonstrate that the mixed data management mode achieves the same or better performance than the best stand-alone option for five memory intensive benchmark applications (run separately and in isolation), resulting in an average speedup compared to the best stand-alone policy of over 10 %, on average.

Effler, Chad↗

ECRAM Materials, Devices, Circuits and Architectures: A Perspective

Abstract Non‐von‐Neumann computing using neuromorphic systems based on two‐terminal resistive nonvolatile memory elements has emerged as a promising approach, but its full potential has not been realized due to the lack of materials and devices with the appropriate attributes. Unlike memristors, which require large write currents to drive phase transformations or filament growth, electrochemical random access memory (ECRAM) decouples the “write” and “read” operations using a “gate” electrode to tune the conductance state through charge‐transfer reactions, and every electron transferred through the external circuit in ECRAM corresponds to the migration of ≈1 ion used to store analogue information. Like static dopants in traditional semiconductors, electrochemically inserted ions modulate the conductivity by locally perturbing a host's electronic structure; however, ECRAM does so in a dynamic and reversible manner. The resulting change in conductance can span orders of magnitude, from gradual increments needed for analog elements, to large, abrupt changes for dynamically reconfigurable adaptive architectures. In this in‐depth perspective, the history of ECRAM, the recent progress in devices spanning organic, inorganic, and 2D materials, circuits, architectures, the rich portfolio of challenging, fundamental questions, and how ECRAM can be harnessed to realize a new paradigm for low‐power neuromorphic computing are discussed.

Talin, A. Alec↗

Benchmarking Operators in Deep Neural Networks for Improving Performance Portability of SYCL

SYCL is a portable programming model for heterogeneous computing, so it is important to obtain reasonable performance portability of SYCL. Towards the goal of better understanding and improving performance portability of SYCL for machine learning workloads, we have been developing benchmarks for basic operators in deep neural networks (DNNs). These operators could be offloaded to heterogeneous computing devices such as graphics processing units (GPUs) to speed up computation. In this paper, we introduce the benchmarks, evaluate the performance of the operators on GPU-based systems, and describe the causes of the performance gap between the SYCL and Compute Unified Device Architecture (CUDA) kernels. We find that the causes are related to the utilization of the texture cache for read-only data, optimization of the memory accesses with strength reduction, use of local memory, and register usage per thread. We hope that the efforts of developing benchmarks for studying performance portability will stimulate discussion and interactions within the community.

Jin, Zheming [ORNL] (ORCID:000000027197780X)↗

LowFive v1.0

LowFive is a new data transport layer based on the HDF5 data model, for in situ workflows. Executables using LowFive can communicate in situ (using in-memory data and MPI message passing), reading and writing traditional HDF5 files to physical storage, and combining the two modes. Minimal and often no source-code modification is needed for programs that already use HDF5. LowFive maintains deep copies or shallow references of datasets, configurable by the user. More than one task can produce (write) data, and more than one task can consume (read) data, accommodating fan-in and fan-out in the workflow task graph. LowFive supports data redistribution from n producer processes to m consumer processes.

Morozov, Dmitriy↗

Evaluating Operators in Deep Neural Networks for Improving Performance Portability of SYCL

SYCL is a portable programming model for heterogeneous computing, so it is important to obtain reasonable performance portability of SYCL. Towards the goal of better understanding and improving performance portability of SYCL for machine learning workloads, we have been developing benchmarks for basic operators in deep neural networks (DNNs). These operators could be offloaded to heterogeneous computing devices such as graphics processing units (GPUs) to speed up computation. In this work, we introduce the benchmarks, evaluate the performance of the operators on GPU-based systems, and describe the causes of the performance gap between the SYCL and Compute Unified Device Architecture (CUDA) kernels. We find that the causes are related to the utilization of the texture cache for read-only data, optimization of the memory accesses with strength reduction, shared local memory accesses, and register usage per thread. We hope that the efforts of developing benchmarks for studying performance portability will stimulate discussion and interactions within the community.

97 MATHEMATICS AND COMPUTING↗

Accelerating shared file checkpoint with local burst buffers

A data management system and method for accelerating shared file checkpointing. Written application data is aggregated in an application data file created in a local burst buffer memory at a compute node, and an associated data mapping built index to maintain information related to the offsets into a shared file at which segments of the application data is to be stored in a parallel file system, and where in the buffer those segments are located. The node asynchronously transfers a data file containing the application data and the associated data mapping index to a file server for shared file storage. The data management system and method further accelerates shared file checkpointing in which a shared file, together with a map file that specifies how the shared file is to be distributed, is asynchronously transferred to local burst buffer memories at the nodes to accelerate reading of the shared file.

Gooding, Thomas↗

Performance Optimization Methods for a Memory-Bound, Unstructured-Grid CFD Application on Massively Parallel GPU Platforms

Computational performance of the FUN3D unstructured-grid computational fluid dynamics (CFD) application on massively parallel GPU environments is memory-bound and highly dependent upon efficient reads from and atomic updates to the irregular cell-, edge-, and node-based data structures. In this talk, we present recent efforts into optimizing select performance-critical kernels on NVIDIA Tesla V100 and A100 GPUs and AMD CDNA MI100 GPUs. A novel use of L2 cache residency controls and asynchronous loads into on-chip shared memory are explored on the A100 GPU for the sparse iterative solver, which is dominated by mixed-precision, sparse matrix vector multiplication. Demonstrations show that these methods improve global memory bandwidth utilization by 13.5% on the A100 GPU. Several techniques are also presented that use registers and/or shared memory to facilitate array transposition and aggregation which combine to reduce the frequency and increase the cache efficiency of floating-point atomic updates to the irregular data structures. These methods are demonstrated to improve the kernel throughput by nearly 500% on select kernels on the AMD MI100 over atomic updates directly to global memory. Overall, both V100 and A100 GPUs outperformed the MI100 GPU on kernels dominated by double-precision atomic updates; however, the techniques demonstrated here reduced the performance gap and improved the MI100 performance.

GPU CPU unstructured CFD memory↗

A design approach to real-time formatting of high speed multispectral image data

A design approach to formatting multispectral image data in real time at very high data rates is presented for future onboard processing applications. The approach employs a microprocessor-based alternating buffer memory configuration whose formatting function is completely programmable. Data are read from an output buffer in the desired format by applying the proper sequence of addresses to the buffer via a lookup table memory. Sensor data can be processed using this approach at rates limited by the buffer memory access time and the buffer switching process delay time. This design offers flexible high speed data processing and benefits from continuing increases in the performance of digital memories.

Meredith, B. D.↗

First Demonstration of Vertical 2T-nC FeRAM Hybrid Cell and its Scalability for High-Density 3D Ferroelectric Capacitor Memory

In this article, we perform a comprehensive experimental and modeling study into the scaling of vertical 2T-nC ferroelectric random-access memory (FeRAM) hybrid cell to demonstrate a high performance and high-density 3D capacitor memory. We demonstrate: i) first time successful integration of the vertical 2T-3C FeRAM cell by stacking the vertical metal-ferroelectricmetal (MFM) stack on top of Si CMOS transistors; ii) successful experimental operation of the memory cell, including the quasi-nondestructive read out (QNRO) of the polarization without write back after 106 read cycles; iii) the write bit line (WBL) heavily screens the coupling between neighboring strings, making it a minor concern; V ) aggressive stacking of the WBLs, i.e., number of MFMs in a string, could facilitate the self-boosting during write operation due to ferroelectric linear capacitance (CFE), which allows self-boosted inhibition for Vw/2 scheme and worsens the Vw/3 scheme as disturb increases to intolerable 2Vw/3; v) aggressive horizontal scaling significantly increases the read disturb to cells on neighboring planes due to capacitance between two WBLs (Cz).

42 ENGINEERING↗

Algorithms for Efficient Reproducible Floating Point Summation

We define “reproducibility” as getting bitwise identical results from multiple runs of the same program, perhaps with different hardware resources or other changes that should not affect the answer. Many users depend on reproducibility for debugging or correctness. However, dynamic scheduling of parallel computing resources, combined with nonassociative floating point addition, makes reproducibility challenging even for summation, or operations like the BLAS. We describe a “reproducible accumulator” data structure (the “binned number”) and associated algorithms to reproducibly sum binary floating point numbers, independent of summation order. We use a subset of the IEEE Floating Point Standard 754-2008 and bitwise operations on the standard representations in memory. Our approach requires only one read-only pass over the data, and one reduction in parallel, using a 6-word reproducible accumulator (more words can be used for higher accuracy), enabling standard tiling optimization techniques. Summing n words with a 6-word reproducible accumulator requires approximately 9 n floating point operations (arithmetic, comparison, and absolute value) and approximately 3 n bitwise operations. The final error bound with a 6-word reproducible accumulator and our default settings can be up to 2 29 times smaller than the error bound for conventional (recursive) summation on ill-conditioned double-precision inputs.

Computer Science↗

Error latency measurements in symbolic architectures

Error latency, the time that elapses between the occurrence of an error and its detection, has a significant effect on reliability. In computer systems, failure rates can be elevated during a burst of system activity due to increased detection of latent errors. A hybrid monitoring environment is developed to measure the error latency distribution of errors occurring in main memory. The objective of this study is to develop a methodology for gauging the dependability of individual data categories within a real-time application. The hybrid monitoring technique is novel in that it selects and categorizes a specific subset of the available blocks of memory to monitor. The precise times of reads and writes are collected, so no actual faults need be injected. Unlike previous monitoring studies that rely on a periodic sampling approach or on statistical approximation, this new approach permits continuous monitoring of referencing activity and precise measurement of error latency.

Young, L. T.↗

Optically Addressable, Ferroelectric Memory With NDRO

For readout, memory cells addressed via on-chip semiconductor lasers. Proposed thin-film ferroelectric memory device features nonvolatile storage, optically addressable, nondestructive readout (NDRO) with fast access, and low vulnerability to damage by ionizing radiation. Polarization switched during recording and erasure, but not during readout. As result, readout would not destroy contents of memory, and operating life in specific "read-intensive" applications increased up to estimated 10 to the 16th power cycles.

Thakoor, Sarita↗

A Byzantine resilient processor with an encoded fault-tolerant shared memory

The memory requirements for ultra-reliable computers are expected to increase due to future increases in mission functionality and operating-system requirements. This increase will have a negative effect on the reliability and cost of the system. Increased memory size will also reduce the ability to reintegrate a channel after a transient fault, since the time required to reintegrate a channel in a conventional fault-tolerant processor is dominated by memory realignment time. A Byzantine Resilient Fault-Tolerant Processor with Fault-Tolerant Shared Memory (FTP/FTSM) is presented as a solution to these problems. The FTSM uses an encoded memory system, which reduces the memory requirement by one-half compared to a conventional quad-FTP design. This increases the reliability and decreases the cost of the system. The realignment problem is also addressed by the FTSM. Because any single error is corrected upon a read from the FTSM, a faulty channel's corrupted memory does not need realignment before reintegration of the faulty channel. A combination of correct-on-access and background scrubbing is proposed to prevent the accumulation of transient errors in the memory. With a hardware-implemented scrubber, the scrubbing cycle time, and therefore the memory fault latency, can be upper-bounded at a small value. This technique increases the reliability of the memory system and facilitates validation of its reliability model.

Butler, Bryan↗

Thrifty Array Format (TAF) file specifications

Thrifty Array Foram (TAF) files store numeric data in a binary format, minimizing storage requirements while preserving quick read access. Real data of any size and dimensionality can be stored in this format at varying degrees of numeric precision. Implicit array are associated with each dimension, eliminating the need to explicitly store uniformly-spaced grid vectors. Unlimited text comments may be included with the array for user documentation, and every file begins with a text synopsis of the binary structure. The format is deliberately designed for memory mapping, where portions of the array can be read without loading the entire file at once.

97 MATHEMATICS AND COMPUTING↗