Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Memory systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Semiconductor Cubing

Through Goddard Space Flight Center and Jet Propulsion Laboratory Small Business Innovation Research contracts, Irvine Sensors developed a three-dimensional memory system for a spaceborne data recorder and other applications for NASA. From these contracts, the company created the Memory Short Stack product, a patented technology for stacking integrated circuits that offers higher processing speeds and levels of integration, and lower power requirements. The product is a three-dimensional semiconductor package in which dozens of integrated circuits are stacked upon each other to form a cube. The technology is being used in various computer and telecommunications applications.

Source record↗

Distributed memory, GPU accelerated Fock construction for hybrid, Gaussian basis density functional theory

With the growing reliance of modern supercomputers on accelerator-based architecture such a graphics processing units (GPUs), the development and optimization of electronic structure methods to exploit these massively parallel resources has become a recent priority. While significant strides have been made in the development GPU accelerated, distributed memory algorithms for many modern electronic structure methods, the primary focus of GPU development for Gaussian basis atomic orbital methods has been for shared memory systems with only a handful of examples pursing massive parallelism. Here in this work, we present a set of distributed memory algorithms for the evaluation of the Coulomb and exact exchange matrices for hybrid Kohn–Sham DFT with Gaussian basis sets via direct density-fitted (DF-J-Engine) and seminumerical (sn-K) methods, respectively. The absolute performance and strong scalability of the developed methods are demonstrated on systems ranging from a few hundred to over one thousand atoms using up to 128 NVIDIA A100 GPUs on the Perlmutter supercomputer.

97 MATHEMATICS AND COMPUTING↗

Electronic implementation of associative memory based on neural network models

An electronic embodiment of a neural network based associative memory in the form of a binary connection matrix is described. The nature of false memory errors, their effect on the information storage capacity of binary connection matrix memories, and a novel technique to eliminate such errors with the help of asymmetrical extra connections are discussed. The stability of the matrix memory system incorporating a unique local inhibition scheme is analyzed in terms of local minimization of an energy function. The memory's stability, dynamic behavior, and recall capability are investigated using a 32-'neuron' electronic neural network memory with a 1024-programmable binary connection matrix.

Moopenn, A.↗

Benchmark characterization

An abstract system of benchmark characteristics that makes it possible, in the beginning of the design stage, to design with benchmark performance in mind is presented. The benchmark characteristics for a set of commonly used benchmarks are then shown. The benchmark set used includes some benchmarks from the Systems Performance Evaluation Cooperative (SPEC). The SPEC programs are industry-standard applications that use specific inputs. Processor, memory-system, and operating-system characteristics are addressed.

Conte, Thomas M.↗

A History of High-Performance Computing

Faster than most speedy computers. More powerful than its NASA data-processing predecessors. Able to leap large, mission-related computational problems in a single bound. Clearly, it s neither a bird nor a plane, nor does it need to don a red cape, because it s super in its own way. It's Columbia, NASA s newest supercomputer and one of the world s most powerful production/processing units. Named Columbia to honor the STS-107 Space Shuttle Columbia crewmembers, the new supercomputer is making it possible for NASA to achieve breakthroughs in science and engineering, fulfilling the Agency s missions, and, ultimately, the Vision for Space Exploration. Shortly after being built in 2004, Columbia achieved a benchmark rating of 51.9 teraflop/s on 10,240 processors, making it the world s fastest operational computer at the time of completion. Putting this speed into perspective, 20 years ago, the most powerful computer at NASA s Ames Research Center, home of the NASA Advanced Supercomputing Division (NAS), ran at a speed of about 1 gigaflop (one billion calculations per second). The Columbia supercomputer is 50,000 times faster than this computer and offers a tenfold increase in capacity over the prior system housed at Ames. What s more, Columbia is considered the world s largest Linux-based, shared-memory system. The system is offering immeasurable benefits to society and is the zenith of years of NASA/private industry collaboration that has spawned new generations of commercial, high-speed computing systems.

Source record↗

Unveiling the nature of Ga-based chalcogenides for electrical switching selectors

Three-dimensional phase-change memory with stackable crossbar architecture is a promising technology to meet the urgent demands for high-density storage and rapid information processing in the era of explosive data growth. The performance depends strongly on the properties of ovonic threshold switching (OTS) selectors, which control the on/off states of memory units. Amorphous GaS serves as an outstanding OTS material, distinguished by its sizable mobility gap and high crystallization temperature, while the underlying mechanism continues to be inadequately comprehended. Here, in this work, we systematically studied the structural and electronic properties of amorphous Ga-X (X = S/Se/Te) using first-principles calculations. The results show that Ga atoms adopt tetrahedral motifs, while S/Se/Te atoms predominantly exhibit the structure of a distorted triangular pyramid. This structural arrangement is ascribed to the substantial dative bonds formed by the lone-pair electrons of the anions and the vacant sp3 orbitals around Ga atoms. Large mobility gaps (e.g., GaS: 2.43 eV, GaSe: 1.76 eV, GaTe: 1.26 eV) and distinct mid-gap states (e.g., ∼0.66 eV above valence band tail) ensure that these three chalcogenide glasses can be switched on under an external electric field while effectively suppressing leakage current without a bias, and the defect electronic states originate from short, robust Ga-Ga bonds due to the formation of distorted chain-like local structures. Our research elucidates the mechanisms of amorphous Ga-X as OTS materials, enriching the spectrum of electrical switching selectors by incorporating III-VI chalcogenides. This inclusion offers novel opportunities for the refinement and optimization of high-density integrated memory systems.

36 MATERIALS SCIENCE↗

Formal Definitions and Performance Comparison of Consistency Models for Parallel File Systems

The semantics of HPC storage systems are defined by the consistency models to which they abide. Storage consistency models have been less studied than their counterparts in memory systems, with the exception of the POSIX standard and its strict consistency model. The use of POSIX consistency imposes a performance penalty that becomes more significant as the scale of parallel file systems increases and the access time to storage devices, such as node-local solid storage devices, decreases. While some efforts have been made to adopt relaxed storage consistency models, these models are often defined informally and ambiguously as by-products of a particular implementation. Here in this work, we establish a connection between memory consistency models and storage consistency models and revisit the key design choices of storage consistency models from a high-level perspective. Further, we propose a formal and unified framework for defining storage consistency models and a layered implementation that can be used to easily evaluate their relative performance for different I/O workloads. Finally, we conduct a comprehensive performance comparison of two relaxed consistency models on a range of commonly seen parallel I/O workloads, such as checkpoint/restart of scientific applications and random reads of deep learning applications. We demonstrate that for certain I/O scenarios, a weaker consistency model can significantly improve the I/O performance. For instance, in small random reads that are typically found in deep learning applications, session consistency achieved a 5x improvement in I/O bandwidth compared to commit consistency, even at small scales.

97 MATHEMATICS AND COMPUTING↗

Kinetics of the xanthophyll cycle and its role in photoprotective memory and response

Efficiently balancing photochemistry and photoprotection is crucial for survival and productivity of photosynthetic organisms in the rapidly fluctuating light levels found in natural environments. The ability to respond quickly to sudden changes in light level is clearly advantageous. In the alga Nannochloropsis oceanica we observed an ability to respond rapidly to sudden increases in light level which occur soon after a previous high-light exposure. This ability implies a kind of memory. In this work, we explore the xanthophyll cycle in N. oceanica as a short-term photoprotective memory system. By combining snapshot fluorescence lifetime measurements with a biochemistry-based quantitative model, we show that short-term memory arises from the xanthophyll cycle. In addition, the model enables us to characterize the relative quenching abilities of the three xanthophyll cycle components. Given the ubiquity of the xanthophyll cycle in photosynthetic organisms the model described here will be of utility in improving our understanding of vascular plant and algal photoprotection with important implications for crop productivity.

59 BASIC BIOLOGICAL SCIENCES↗

Relaxing consistency in recoverable distributed shared memory

Relaxed memory consistency models tolerate increased memory access latency in both hardware and software distributed shared memory systems. In recoverable systems, relaxing consistency has the added benefit of reducing the number of checkpoints needed to avoid rollback propagation. In this paper, we introduce new checkpointing algorithms that take advantage of relaxed consistency to reduce the performance overhead of checkpointing. We also introduce a scheme based on lazy relaxed consistency, that reduces both checkpointing overhead and the overhead of avoiding error propagation in systems with error latency. We use multiprocessor address traces to evaluate the relaxed consistency approach to checkpointing with distributed shared memory.

Janssens, Bob↗

The ijk forms of factorization methods. II - Parallel systems

The paper considers the 'ijk forms' of LU and Cholesky factorization on certain parallel computers. This extends an earlier analysis for vector computers. Attention is restricted to local memory systems with processors that may or may not have vector capability. Special attention is given to bus architectures but qualitative analyses are given for other interconnection systems.

Ortega, J. M.↗

Ropes: Support for collective opertions among distributed threads

Lightweight threads are becoming increasingly useful in supporting parallelism and asynchronous control structures in applications and language implementations. Recently, systems have been designed and implemented to support interprocessor communication between lightweight threads so that threads can be exploited in a distributed memory system. Their use, in this setting, has been largely restricted to supporting latency hiding techniques and functional parallelism within a single application. However, to execute data parallel codes independent of other threads in the system, collective operations and relative indexing among threads are required. This paper describes the design of ropes: a scoping mechanism for collective operations and relative indexing among threads. We present the design of ropes in the context of the Chant system, and provide performance results evaluating our initial design decisions.

Haines, Matthew↗

Fixing Amdahl's Law within the Limits of Accelerated Systems: FALLACY

Closeout report for FALLACY project. The performance of Data Model Convergence Initiative (DMC) applications on parallel machines is far below the limit set by Amdahl’s law. Whether the machine is based on many-core, GPUs, FPGAs, or a heterogeneous combination, usually the most significant bottleneck is accessing data from the memory system. Aligning with DMC’s HW/architecture thrust, this project developed a set of memory-centric tools called ‘MemGaze’ that inform the HW/SW stack about an application’s memory behavior, including data access latency and diagnosing poor data layout and data composition. Our approach uses architectural modeling and analysis of workload data accesses.

97 MATHEMATICS AND COMPUTING↗

Bias control for a memory device

Methods, systems, and devices for bias control for a memory device are described. A memory system may store indication of whether data is coherent. In some examples, the indication may be stored as metadata, where a first value indicates that the data is not coherent and a second value or a third value indicate that the data is coherent. When a processing unit or other component of the memory system processes a command to access data, the memory system may operate according to a device bias mode when the indication is the first value, and according to a host bias mode when the indication is the second value or the third value.

97 MATHEMATICS AND COMPUTING↗

Bias control for a memory device

Methods, systems, and devices for bias control for a memory device are described. A memory system may store indication of whether data is coherent. In some examples, the indication may be stored as metadata, where a first value indicates that the data is not coherent and a second value or a third value indicate that the data is coherent. When a processing unit or other component of the memory system processes a command to access data, the memory system may operate according to a device bias mode when the indication is the first value, and according to a host bias mode when the indication is the second value or the third value.

Walker, Dean↗

Monitoring Top-of-Atmosphere Radiative Energy Imbalance for Climate Prediction

Large climate feedback uncertainties limit the prediction accuracy of the Earth s future climate with an increased CO2 atmosphere. One potential to reduce the feedback uncertainties using satellite observations of top-of-atmosphere (TOA) radiative energy imbalance is explored. Instead of solving the initial condition problem in previous energy balance analysis, current study focuses on the boundary condition problem with further considerations on climate system memory and deep ocean heat transport, which is more applicable for the climate. Along with surface temperature measurements of the present climate, the climate feedbacks are obtained based on the constraints of the TOA radiation imbalance. Comparing to the feedback factor of 3.3 W/sq m/K of the neutral climate system, the estimated feedback factor for the current climate system ranges from -1.3 to -1.0 W/sq m/K with an uncertainty of +/-0.26 W/sq m/K. That is, a positive climate feedback is found because of the measured TOA net radiative heating (0.85 W/sq m) to the climate system. The uncertainty is caused by the uncertainties in the climate memory length. The estimated time constant of the climate is large (70 to approx. 120 years), implying that the climate is not in an equilibrium state under the increasing CO2 forcing in the last century.

Lin, Bing↗

Compact holographic memory using E - O beam steering

An innovative holographic memory system has been developed at JPL for high-density and high speed data storage in a space environment. This system ulitlizes a newly developed electro-optic (E-O) beam steering technology for beam steering to enable high-speed random access memory read/write without moving parts. Recently, a compact CD-sized holographic memory broadboard has been developed and demonstrated for holographic data storage adn retrieval. Detail technical progress will be presented in this paper.

electo - optic (E - O ) beam↗

Man machine interactive imaging and data processing using high speed digital mass storage

Attention is given to general questions regarding television systems, aspects of image distortion in TV systems, approaches of image restoration, noise limitations, and the digital image integrator. The digital image recorder has been built around a high-speed digital memory system. It is shown that television systems can be used to augment man's capabilities in the control loop of remote manipulation systems.

Alsberg, H.↗

Asymmetric soft-error resistant memory

A memory system is provided, of the type that includes an error-correcting circuit that detects and corrects, that more efficiently utilizes the capacity of a memory formed of groups of binary cells whose states can be inadvertently switched by ionizing radiation. Each memory cell has an asymmetric geometry, so that ionizing radiation causes a significantly greater probability of errors in one state than in the opposite state (e.g., an erroneous switch from '1' to '0' is far more likely than a switch from '0' to'1'. An asymmetric error correcting coding circuit can be used with the asymmetric memory cells, which requires fewer bits than an efficient symmetric error correcting code.

Buehler, Martin G.↗