Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Memory systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Integration of Ag-CBRAM crossbars and Mott ReLU neurons for efficient implementation of deep neural networks in hardware

In-memory computing with emerging non-volatile memory devices (eNVMs) has shown promising results in accelerating matrix-vector multiplications. However, activation function calculations are still being implemented with general processors or large and complex neuron peripheral circuits. Here, we present the integration of Ag-based conductive bridge random access memory (Ag-CBRAM) crossbar arrays with Mott rectified linear unit (ReLU) activation neurons for scalable, energy and area-efficient hardware (HW) implementation of deep neural networks. We develop Ag-CBRAM devices that can achieve a high ON/OFF ratio and multi-level programmability. Compact and energy-efficient Mott ReLU neuron devices implementing ReLU activation function are directly connected to the columns of Ag-CBRAM crossbars to compute the output from the weighted sum current. We implement convolution filters and activations for VGG-16 using our integrated HW and demonstrate the successful generation of feature maps for CIFAR-10 images in HW. Our approach paves a new way toward building a highly compact and energy-efficient eNVMs-based in-memory computing system.

Mott insulators↗

Noise-induced stabilization of dynamical states with broken time-reversal symmetry

Under a high frequency drive, Josephson junctions demonstrate Shapiro steps of quantized voltage. These are dynamically stabilized states in which the phase across the junction locks to the external drive. We explore the stochastic switching between two symmetric steps at $\frac{hw}{2e}$ and –$\frac{hw}{2e}$. Surprisingly, the switching rate exhibits a pronounced nonmonotonicity as a function of temperature, violating the general expectation that transitions should become faster with temperature. As a result, we explain this behavior by realizing that the system retains memory of the dynamic state from which it is switching, thereby breaking the conventional simplifying assumptions about separations of timescales.

36 MATERIALS SCIENCE↗

Accelerating Scientific Applications With SambaNova Reconfigurable Dataflow Architecture

Here, our exploratory work finds that the SambaNova Reconfigurable Dataflow Architecture (RDA) along with the SambaFlow software stack provides for an attractive system and solution to accelerate AI for science workloads. We have observed the efficacy of using the system with a diverse set of science applications and reasoned their suitability for performance gains over traditional hardware. As the Data-Scale system provides for a very large memory capacity, the system can be used to train models that typically do not fit in a GPU. The architecture also provides for deeper integration with upcoming supercomputers at the Argonne Leadership Computing Facility (ALCF), a US Department of Energy Office of Science user facility, to help advance science insights.

97 MATHEMATICS AND COMPUTING↗

Chalcogenide phase-change material advances programmable terahertz metamaterials: a non-volatile perspective for reconfigurable intelligent surfaces

Terahertz (THz) waves have gained considerable attention in the rising 6G communication due to their large bandwidth. However, the cost and power consumption become the major constraints for the commercialization of 6G THz systems as the frequency increases. Reconfigurable intelligent surface (RIS) comprising active metasurfaces and digital controllers has been proposed for beamforming in the 6G multiple-input-multiple-output systems, showing good potential to suppress the system size, weight, and power consumption (SWaP). Currently, their controlling diodes can hardly work up to THz frequencies. Therefore, several active stimuli have been investigated as alternatives. Among them, chalcogenide phase-change material Ge 2 Sb 2 Te 5 (GST) addresses large modulation depth, picosecond switching speed, and non-volatile properties. Notably, the non-volatile GST may enable RIS systems with memory and low control power. This work briefly reviews the advances of GST-tuned THz metamaterials (MTMs), discusses the current obstacles to overcome, and gives a perspective of GST applications in the rising 6G communications.

6G↗

Dynamic adaptation of memory page management policy

Systems, apparatuses, and methods for determining preferred memory page management policies by software are disclosed. Software executing on one or more processing units generates a memory request. Software determines the preferred page management policy for the memory request based at least in part on the data access size and data access pattern of the memory request. Software conveys an indication of a preferred page management policy to a memory controller. Then, the memory controller accesses memory for the memory request using the preferred page management policy specified by software.

97 MATHEMATICS AND COMPUTING↗

Processable, tunable thio-ene crosslinked polyurethane shape memory polymers

An embodiment includes a platform shape memory polymer system. Such an embodiment exhibits a blend of tunable, high performance mechanical attributes in combination with advanced processing capabilities and good biocompatibility. A post-polymerization crosslinking synthetic approach is employed that combines polyurethane and thiol-ene synthetic processes. Other embodiments are described herein.

36 MATERIALS SCIENCE↗

Register-Like Block RAM: Implementation, Testing in FPGA and Applications for High Energy Physics Trigger Systems

In high energy physics experiment trigger systems, block memories are utilized for various purposes, especially in indexed searching algorithms. It is often demanded to globally reset all memory locations between different events which is a feature not supported in regular block memories. Another common demand is to be able to update the contents in any memory location in a single clock cycle. These two demands can be fulfilled with registers but the cost of using registers for large memory is unaffordable. In this paper, a register-like block memory design scheme is described, which allows updating memory locations in single clock cycle and effectively refreshing entire memory within a single clock. The implementation and test results are presented.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Processable, tunable thiol-ene crosslinked polyurethane shape memory polymers

An embodiment includes a platform shape memory polymer system. Such an embodiment exhibits a blend of tunable, high performance mechanical attributes in combination with advanced processing capabilities and good biocompatibility. A post-polymerization crosslinking synthetic approach is employed that combines polyurethane and thiol-ene synthetic processes. Other embodiments are described herein.

Hearon, Keith↗

Selective consolidation of learning and memory via recall-gated plasticity

In a variety of species and behavioral contexts, learning and memory formation recruits two neural systems, with initial plasticity in one system being consolidated into the other over time. Moreover, consolidation is known to be selective; that is, some experiences are more likely to be consolidated into long-term memory than others. Here, we propose and analyze a model that captures common computational principles underlying such phenomena. The key component of this model is a mechanism by which a long-term learning and memory system prioritizes the storage of synaptic changes that are consistent with prior updates to the short-term system. This mechanism, which we refer to as recall-gated consolidation, has the effect of shielding long-term memory from spurious synaptic changes, enabling it to focus on reliable signals in the environment. We describe neural circuit implementations of this model for different types of learning problems, including supervised learning, reinforcement learning, and autoassociative memory storage. These implementations involve synaptic plasticity rules modulated by factors such as prediction accuracy, decision confidence, or familiarity. We then develop an analytical theory of the learning and memory performance of the model, in comparison to alternatives relying only on synapse-local consolidation mechanisms. We find that recall-gated consolidation provides significant advantages, substantially amplifying the signal-to-noise ratio with which memories can be stored in noisy environments. We show that recall-gated consolidation gives rise to a number of phenomena that are present in behavioral learning paradigms, including spaced learning effects, task-dependent rates of consolidation, and differing neural representations in short- and long-term pathways.

59 BASIC BIOLOGICAL SCIENCES↗

Selective consolidation of learning and memory via recall-gated plasticity

In a variety of species and behavioral contexts, learning and memory formation recruits two neural systems, with initial plasticity in one system being consolidated into the other over time. Moreover, consolidation is known to be selective; that is, some experiences are more likely to be consolidated into long-term memory than others. Here, we propose and analyze a model that captures common computational principles underlying such phenomena. The key component of this model is a mechanism by which a long-term learning and memory system prioritizes the storage of synaptic changes that are consistent with prior updates to the short-term system. This mechanism, which we refer to as recall-gated consolidation, has the effect of shielding long-term memory from spurious synaptic changes, enabling it to focus on reliable signals in the environment. We describe neural circuit implementations of this model for different types of learning problems, including supervised learning, reinforcement learning, and autoassociative memory storage. These implementations involve synaptic plasticity rules modulated by factors such as prediction accuracy, decision confidence, or familiarity. We then develop an analytical theory of the learning and memory performance of the model, in comparison to alternatives relying only on synapse-local consolidation mechanisms. We find that recall-gated consolidation provides significant advantages, substantially amplifying the signal-to-noise ratio with which memories can be stored in noisy environments. We show that recall-gated consolidation gives rise to a number of phenomena that are present in behavioral learning paradigms, including spaced learning effects, task-dependent rates of consolidation, and differing neural representations in short- and long-term pathways.

Lindsey, Jack W. (ORCID:0000000309307327)↗

An Integrated Framework for Memory-Centric Analysis: From Trace Collection to Co-Design

The memory wall phenomenon—where advances in processor performance significantly outpace those in memory subsystems—poses a fundamental challenge for contemporary computing systems. In memory-bound applications, memory subsystem behavior dominates performance, yet existing analysis approaches present significant limitations: detailed microarchitectural simulators require days to weeks to simulate modest workloads; hardware performance counters provide only aggregate statistics that obscure temporal and spatial access patterns; and scaled simulation approaches face challenges in capturing certain behaviors that emerge at larger scales. These limitations reflect a processor-centric design philosophy increasingly misaligned with memory-bound workloads where detailed understanding of memory access patterns, cache hierarchy interactions, and contention is critical for effective optimization. This paper presents an integrated framework for memory-centric analysis that enables effective hardware-software co-design. We describe practical trace collection techniques, including hardware-assisted processor tracing with minimal overhead and portable software-based instrumentation with statistical sampling. We present multi-perspective analysis methods that examine memory behavior from temporal, sequential, spatial, and relational viewpoints, revealing distinct optimization opportunities invisible in aggregate metrics. We detail an architectural modeling framework that uses sampled traces with temporal interpolation and confidence-based filtering to evaluate cache and memory configurations. Evaluation on representative benchmarks demonstrates that this framework achieves practical accuracy (L2 cache errors of 2.64\%, confidence-filtered L3 errors of 9.92\%, bandwidth errors of 7.33\%) while providing substantial speedup (26.8×) over cycle-accurate simulation, enabling rapid design space exploration. We demonstrate how this integrated framework enables systematic identification of both hardware optimizations (memory controller tuning, bank partitioning, NUMA configuration) and software optimizations (data layout restructuring, prefetching strategies, memory-aware scheduling). Through this comprehensive treatment of the memory-centric analysis pipeline—from trace collection through architectural modeling to co-design application—we provide researchers and practitioners with practical techniques for addressing memory bottlenecks in contemporary computing systems.

Gajaria, Dhruv Mayur↗

ZFP Hardware Implementation

As core counts increase in new HPC systems with comparatively little increase in memory bandwidth, the trend is an effective decrease in memory bandwidth per core. Other bandwidth limitations in HPC systems exist between CPU and GPU memory, between system nodes, and between node memory and storage. Compression of floating-point data has the potential to reduce data movement and pressure across these communication channels. Furthermore, it has the potential to reduce the footprint of floating-point arrays stored in memory. ZFP, implemented in software, is gaining traction as an effective method in floating-point compression; however, performance gains are limited to the spare compute cycles available before reaching the bandwidth limitations of the communication channel. A hardware implementation of ZFP has the potential to raise the bar on performance. From the inception of ZFP, it was designed to accommodate a hardware implementation.

42 ENGINEERING↗

TriC: Distributed-memory Triangle Counting by Exploiting the Graph Structure

Graph analytics has emerged as an important tool in the analysis of large scale data from diverse application domains such as social networks, cyber security and bioinformatics. Counting the number of triangles in a graph is a fundamental kernel with several applications such as detecting the community structure of a graph or in identifying important vertices in a graph. The ubiquity of massive datasets is driving the need to scale graph analytics on parallel systems. However, numerous challenges exist in efficiently parallelizing graph algorithms, especially on distributed-memory systems. Irregular memory accesses and communication patterns, low computation to communication ratios, and the need for frequent synchronization are some of the leading challenges. In this paper, we present TriC, our distributed-memory implementation of triangle counting in graphs using the Message Passing Interface (MPI), as a submission to the 2020 GraphChallenge competition. Using a set of synthetic and real-world inputs from the challenge, we demonstrate a speedup of up to 90x relative to previous work on 32 processor-cores of a NERSC Cori node. We also provide details from distributed runs with up to8192 processes along with strong scaling results. The observations presented in this work provide an understanding of the system-level bottlenecks at scale that specifically impact sparse-irregular workloads and will therefore benefit other efforts to parallelize graph algorithms.

Halappanavar, Mahantesh↗

Malicious activity detection in a memory

A method and apparatus for monitoring a volatile memory in a computer system. Samples of compressed data from locations in the volatile memory in the computer system are read. Data in the volatile memory is reconstructed using the samples of compressed data. The data is an image of the volatile memory. The image enables determining whether an undesired process is present in the volatile memory.

Wheeler, Jason W.↗

JTAG-based PLC memory acquisition framework for industrial control systems

In industrial control systems (ICS), programmable logic controllers (PLC) are the embedded devices that directly control and monitor critical industrial infrastructure processes such as nuclear plants and power grid stations. Cyberattacks often target PLCs to sabotage a physical process. A memory forensic analysis of a suspect PLC can answer questions about an attack, including compromised firmware and manipulation of PLC control logic code and I/O devices. Given physical access to a PLC, collecting forensic information from the PLC memory at the hardware-level is risky and challenging. It may cause the PLC to crash or hang since PLCs have proprietary, legacy hardware with heterogeneous architecture. This paper addresses this research problem and proposes a novel JTAG (Joint Test Action Group)-based framework, Kyros, for reliable PLC memory acquisition. Kyros systematically creates a JTAG profile of a PLC through hardware assessment, JTAG pins identification, memory map creation, and optimizing acquisition parameters. It also facilitates the community of interest (such as ICS owners, operators, and vendors) to develop the JTAG profiles of PLCs. Further, we present a case study of Kyros implementation over Allen-Bradley 1756-A10/B to help understand the framework's application on a real-world PLC used in industry settings. The sample PLC memory dumps are shared with the research community to facilitate further research.

Rais, Muhammad Haris↗

Multimode Strong Coupling in Cavity Optomechanics

Optomechanical systems show great potential as quantum transducers and information storage devices for use in future hybrid quantum networks. In this context, optomechanical strong coupling can enable efficient, high-bandwidth, and deterministic transfer of quantum states. While optomechanical strong coupling has been realized at optical frequencies, it has proven difficult to identify a robust optomechanical system that features the low loss and high coupling rates required for more sophisticated control of mechanical motion. In this paper, we demonstrate strong coupling in a Brillouin-based bulk cavity optomechanical system in both the single-mode and the multimode strong-coupling regime, which leads to a useful device both for applications in quantum information and for investigating decoherence phenomena in bulk acoustic wave resonators. Using nontrivial mode hybridizations in the strong-coupling regime, we create hybridized photonic-phononic modes with lifetimes that are significantly longer than those of the uncoupled system. This surprising lifetime enhancement, which results from the interference of decay channels, showcases the use of multimode strong coupling as a general strategy to control extrinsic decoherence mechanisms. Moreover, phonons supported by such bulk-acoustic-wave resonators have a collection of properties, including high frequencies, long coherence times, and robustness against thermal decoherence, that make this optomechanical system particularly enticing for applications such as quantum transduction and memories. Hence, this system provides access to phenomena in a previously unexplored regime of optomechanical interactions and could serve as an important building block for future quantum devices.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Detection and Diagnosis of Data Integrity Attacks in Solar Farms Based on Multilayer Long Short-Term Memory Network

Photovoltaic (PV) systems are becoming more vulnerable to cyber threats. In response to this emerging concern, developing cyber-secure power electronics converters has received increased attention from the IEEE Power Electronics Society that recently launched a cyber-physical-security initiative. Here this letter proposes a deep sequence learning based diagnosis solution for data integrity attacks on PV systems in smart grids, including dc–dc and dc–ac converters. The multilayer long short-term memory networks are used to leverage time-series electric waveform data from current and voltage sensors in PV systems. The proposed method has been evaluated in a PV smart grid benchmark model with extensive quantitative analysis. As a comparison, we have evaluated classic data-driven methods, including K-nearest neighbor, decision tree, support vector machine, artificial neural network, and convolutional neural network. Comparison results verify performances of the proposed method for detection and diagnosis of various data integrity attacks on PV systems.

42 ENGINEERING↗