Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Memory systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20

Modeling and measuring multiprogramming and system overheads on a shared-memory multiprocessor - Case study

The present discussion of methods for quantifying multiprogramming (MP) overhead on a computer system illustrates two such techniques, respectively for quantifying MP overheads' lower bound and determining the MP overload of real workloads, in light of the percentage of parallel processing time that is consumed by MP overhead on Alliant multiprocessors. Kernel lock spinning is found to be a major factor in MP overhead, which accounts for more than half of total system overhead. It is noted that parallel environments' MP overhead is not statistically dependent on the number of parallel jobs undergoing multiprogramming.

Dimpsey, Robert T.↗

Energy-efficient Mott activation neuron for full-hardware implementation of neural networks

To circumvent the von Neumann bottleneck, substantial progress has been made towards in-memory computing with synaptic devices. However, compact nanodevices implementing non-linear activation functions are required for efficient full-hardware implementation of deep neural networks. Here, in this work, we present an energy-efficient and compact Mott activation neuron based on vanadium dioxide and its successful integration with a conductive bridge random access memory (CBRAM) crossbar array in hardware. The Mott activation neuron implements the rectified linear unit function in the analogue domain. The neuron devices consume substantially less energy and occupy two orders of magnitude smaller area than those of analogue complementary metal–oxide semiconductor implementations. The LeNet-5 network with Mott activation neurons achieves 98.38% accuracy on the MNIST dataset, close to the ideal software accuracy. We perform large-scale image edge detection using the Mott activation neurons integrated with a CBRAM crossbar array. Our findings provide a solution towards large-scale, highly parallel and energy-efficient in-memory computing systems for neural networks.

electrical and electronic engineering↗

Integration of Ag-CBRAM crossbars and Mott ReLU neurons for efficient implementation of deep neural networks in hardware

In-memory computing with emerging non-volatile memory devices (eNVMs) has shown promising results in accelerating matrix-vector multiplications. However, activation function calculations are still being implemented with general processors or large and complex neuron peripheral circuits. Here, we present the integration of Ag-based conductive bridge random access memory (Ag-CBRAM) crossbar arrays with Mott rectified linear unit (ReLU) activation neurons for scalable, energy and area-efficient hardware (HW) implementation of deep neural networks. We develop Ag-CBRAM devices that can achieve a high ON/OFF ratio and multi-level programmability. Compact and energy-efficient Mott ReLU neuron devices implementing ReLU activation function are directly connected to the columns of Ag-CBRAM crossbars to compute the output from the weighted sum current. We implement convolution filters and activations for VGG-16 using our integrated HW and demonstrate the successful generation of feature maps for CIFAR-10 images in HW. Our approach paves a new way toward building a highly compact and energy-efficient eNVMs-based in-memory computing system.

Mott insulators↗

Noise-induced stabilization of dynamical states with broken time-reversal symmetry

Under a high frequency drive, Josephson junctions demonstrate Shapiro steps of quantized voltage. These are dynamically stabilized states in which the phase across the junction locks to the external drive. We explore the stochastic switching between two symmetric steps at $\frac{hw}{2e}$ and –$\frac{hw}{2e}$. Surprisingly, the switching rate exhibits a pronounced nonmonotonicity as a function of temperature, violating the general expectation that transitions should become faster with temperature. As a result, we explain this behavior by realizing that the system retains memory of the dynamic state from which it is switching, thereby breaking the conventional simplifying assumptions about separations of timescales.

36 MATERIALS SCIENCE↗

Accelerating Scientific Applications With SambaNova Reconfigurable Dataflow Architecture

Here, our exploratory work finds that the SambaNova Reconfigurable Dataflow Architecture (RDA) along with the SambaFlow software stack provides for an attractive system and solution to accelerate AI for science workloads. We have observed the efficacy of using the system with a diverse set of science applications and reasoned their suitability for performance gains over traditional hardware. As the Data-Scale system provides for a very large memory capacity, the system can be used to train models that typically do not fit in a GPU. The architecture also provides for deeper integration with upcoming supercomputers at the Argonne Leadership Computing Facility (ALCF), a US Department of Energy Office of Science user facility, to help advance science insights.

97 MATHEMATICS AND COMPUTING↗

Chalcogenide phase-change material advances programmable terahertz metamaterials: a non-volatile perspective for reconfigurable intelligent surfaces

Terahertz (THz) waves have gained considerable attention in the rising 6G communication due to their large bandwidth. However, the cost and power consumption become the major constraints for the commercialization of 6G THz systems as the frequency increases. Reconfigurable intelligent surface (RIS) comprising active metasurfaces and digital controllers has been proposed for beamforming in the 6G multiple-input-multiple-output systems, showing good potential to suppress the system size, weight, and power consumption (SWaP). Currently, their controlling diodes can hardly work up to THz frequencies. Therefore, several active stimuli have been investigated as alternatives. Among them, chalcogenide phase-change material Ge 2 Sb 2 Te 5 (GST) addresses large modulation depth, picosecond switching speed, and non-volatile properties. Notably, the non-volatile GST may enable RIS systems with memory and low control power. This work briefly reviews the advances of GST-tuned THz metamaterials (MTMs), discusses the current obstacles to overcome, and gives a perspective of GST applications in the rising 6G communications.

6G↗

Data storage, image tube type

Method and apparatus for the storage of digital or analog electrical signals are provided by a memory storage system employing a conventional vidicon tube. At the beginning of an operating cycle, the vidicon is conditioned to accept electrical data input by exposing its photosensitive target to a short, high intensity light flash. A first electron beam scan of the photosensitive surface then sets up a charge pattern on the photosensitive target. A second electron beam scan of the photosensitive surface by an unmodulated electron beam then develops an output signal across an output resistor by using capacitive currents. The conditioning and scanning steps are operated repetitively at high speed using conventional television camera scan, sync, and power supply circuitry to provide a low cost data storage system.

Lipoma, P. C.↗

Prevention of design flaws in multicomputer systems

Report summarizes research on failure mode analysis for multicomputer systems where two or more computers may serve as redundant set. Failure modes such as data bus monopolization, shutdown due to transients, loss of control system equalization, memory alteration, and software errors are discussed.

Romberg, J. M.↗

NASTRAN computer resource management for the matrix decomposition modules

Detailed computer resource measurements of the NASTRAN matrix decomposition spill logic were made using a software input/output monitor. These measurements showed that, in general, job cost can be reduced by avoiding spill. The results indicated that job cost can be minimized by using dynamic memory management. A prototype memory management system is being implemented and evaluated for the CDC Cyber computer.

Bolz, C. W.↗

On the impact of communication complexity in the design of parallel numerical algorithms

This paper describes two models of the cost of data movement in parallel numerical algorithms. One model is a generalization of an approach due to Hockney, and is suitable for shared memory multiprocessors where each processor has vector capabilities. The other model is applicable to highly parallel nonshared memory MIMD systems. In the second model, algorithm performance is characterized in terms of the communication network design. Techniques used in VLSI complexity theory are also brought in, and algorithm independent upper bounds on system performance are derived for several problems that are important to scientific computation.

Gannon, D.↗

On the impact of communication complexity on the design of parallel numerical algorithms

This paper describes two models of the cost of data movement in parallel numerical alorithms. One model is a generalization of an approach due to Hockney, and is suitable for shared memory multiprocessors where each processor has vector capabilities. The other model is applicable to highly parallel nonshared memory MIMD systems. In this second model, algorithm performance is characterized in terms of the communication network design. Techniques used in VLSI complexity theory are also brought in, and algorithm-independent upper bounds on system performance are derived for several problems that are important to scientific computation.

Gannon, D. B.↗

Partitioning problems in parallel, pipelined and distributed computing

The problem of optimally assigning the modules of a parallel program over the processors of a multiple computer system is addressed. A Sum-Bottleneck path algorithm is developed that permits the efficient solution of many variants of this problem under some constraints on the structure of the partitions. In particular, the following problems are solved optimally for a single-host, multiple satellite system: partitioning multiple chain structured parallel programs, multiple arbitrarily structured serial programs and single tree structured parallel programs. In addition, the problems of partitioning chain structured parallel programs across chain connected systems and across shared memory (or shared bus) systems are also solved under certain constraints. All solutions for parallel programs are equally applicable to pipelined programs. These results extend prior research in this area by explicitly taking concurrency into account and permit the efficient utilization of multiple computer architectures for a wide range of problems of practical interest.

Bokhari, S.↗

Parafrase restructuring of FORTRAN code for parallel processing

Parafrase transforms a FORTRAN code, subroutine by subroutine, into a parallel code for a vector and/or shared-memory multiprocessor system. Parafrase is not a compiler; it transforms a code and provides information for a vector or concurrent process. Parafrase uses a data dependency to reveal parallelism among instructions. The data dependency test distinguishes between recurrences and statements that can be directly vectorized or parallelized. A number of transformations are required to build a data dependency graph.

Wadhwa, Atul↗