Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Memory systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Application of mesh refinement to relativistic magnetic reconnection

During relativistic magnetic reconnection, antiparallel magnetic fields undergo a rapid change in topology, releasing a large amount of energy in the form of non-thermal particle acceleration. This work explores the application of mesh refinement to 2D reconnection simulations to efficiently model the inherent disparity in length-scales. We have systematically investigated the effects of mesh refinement and determined necessary modifications to the algorithm required to mitigate non-physical artifacts at the coarse–fine interface. We have used the ultrahigh-order pseudo-spectral analytical time-domain Maxwell solver to analyze how its use can mitigate the numerical dispersion that occurs with the finite-difference time-domain (or “Yee”) method. Absorbing layers are introduced at the coarse–fine interface to eliminate spurious effects that occur with mesh refinement. We also study how damping the electromagnetic fields and current density in the absorbing layer can help prevent the non-physical accumulation of charge and current density at the coarse–fine interface. Using a mesh refinement ratio of 8 for two-dimensional magnetic reconnection simulations, we obtained good agreement with the high-resolution baseline simulation, using only 36% of the macroparticles and 71% of the node-hours needed for the baseline. The methods presented here are especially applicable to 3D systems where higher memory savings are expected than in 2D, enabling comprehensive, computationally efficient 3D reconnection studies in the future.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Differential Power Processing for Ultra-Efficient Data Storage

Here this paper presents the hardware, software, and power codesign of an ultra-efficient data storage server with differential power processing (DPP). DPP can reduce the power conversion stress, improve the efficiency, and enhance the functionality of modular power electronics systems. The power inputs of a large number of hard disk drives (HDDs) were connected in series and supported by a multiport ac-coupled differential power processing (MAC-DPP) converter through a multiwinding transformer. Methods for controlling the multi-input multi-output power flow in the multiwinding transformer while avoiding core saturation were investigated. A ten-port MAC-DPP prototype with 700-W/in 3 power density was built to support a 450-W HDD storage system with ten series-stacked voltage domains. The prototype was tested on a 50-HDD server testbench, and the overall system loss is below 1 W (99.77% system efficiency). The server was able to maintain high-speed reading and writing operation of all 50 HDDs against the worst hot-swapping scenarios. A variety of hardware/software configurations and many cloud storage techniques were tested on the fully functioning server. Experimental results show that the energy efficiency of large-scale information systems (CPU/GPU clusters, memory banks, HDD arrays, etc.) can be greatly improved by software, hardware, and power codesign.

42 ENGINEERING↗

pyDRESCALk

Modern data scientists are tasked to analyze ever-growing data sets with increasingly complex relationships. Tensor decompositions have come to play a central role in identifying underlying latent structures in higher-order data. The problem of fitting tensor models to different distributions is complicated by the combinations of size, dimensionality, and sparsity present in real world data. The situation demands efficient algorithms designed for shared-memory and distributed systems. This work will present new research that tackles these challenges on several different fronts, leveraging optimizations in numerical algorithms and sparse tensor representations in heterogeneous high performance computing environments.

Bhattarai, Manish↗

The Exploitation of Data Reduction for Visualization

The disparity between the computational speed and storage bandwidth, as demonstrated in Figure 1, is a well known problem that grows with each successive generation. The visualization community is principally responding to this issue by using in situ to reduce which data must be written to storage. However, other communities are taking different, possibly complementary approaches. In particular, data compression is a common general approach to reduce storage demands. Data compression technologies are typically not designed with post processing in mind. The principal metrics measured are compression ratio, the improved bandwidth to storage, and the error introduced. It is assumed that data is inflated to its full size before any post processing can happen. Although when talking about bandwidth disparities, HPC’s dirty little secret is that no part of the memory nor interconnect hardware is increasing at the rate of computation. For example, the Summit supercomputer has a peak computation rate almost 10 times its predecessor, Titan, but only about 4 times the memory, less than twice the aggregate memory bandwidth, and almost no improvement in the interconnect bisection bandwidth. Naively inflating data for post processing does not help with limitations in the memory and interconnect systems.

97 MATHEMATICS AND COMPUTING↗

Locality-aware and sharing-aware cache coherence for collections of processors

A cache coherence technique for operating a multi-processor system including shared memory includes allocating a cache line of a cache memory of a processor to a memory address in the shared memory in response to execution of an instruction of a program executing on the processor. The technique includes encoding a shared information state of the cache line to indicate whether the memory address is a shared memory address shared by the processor and a second processor, or a private memory address private to the processor, in response to whether the instruction is included in a critical section of the program, the critical section being a portion of the program that confines access to shared, writeable data.

Farmahini Farahani, Amin↗

Error containment for enabling local checkpoint and recovery

Various embodiments include a parallel processing computer system that detects memory errors as a memory client loads data from memory and disables the memory client from storing data to memory, thereby reducing the likelihood that the memory error propagates to other memory clients. The memory client initiates a stall sequence, while other memory clients continue to execute instructions and the memory continues to service memory load and store operations. When a memory error is detected, a specific bit pattern is stored in conjunction with the data associated with the memory error. When the data is copied from one memory to another memory, the specific bit pattern is also copied, in order to identify the data as having a memory error.

Cherukuri, Naveen↗

Study of carbon nanotube embedded honey as a resistive switching material

In this paper, natural organic honey embedded with carbon nanotubes (CNTs) was studied as a resistive switching material for biodegradable nonvolatile memory in emerging neuromorphic systems. CNTs were dispersed in a honey-water solution with the concentration of 0.2 wt% CNT and 30 wt% honey. The final honey-CNT-water mixture was spin-coated and dried into a thin film sandwiched in between Cu bottom electrode and Al top electrode to form a honey-CNT based resistive switching memory (RSM). Surface morphology, electrical characteristics and current conduction mechanism were investigated. The results show that although CNTs formed agglomerations in the dried honey-CNT film, both switching speed and the stability in SET and RESET process of honey-CNT RSM were improved. The mechanism of current conduction in CNT is governed by Ohm’s law in low-resistance state and the low-voltage range in high-resistance state, but transits to the space charge limited conduction at high voltages approaching the SET voltage.

36 MATERIALS SCIENCE↗

Performance Impact and Trade-Offs for Tuning Key Architectural Parameters on CPU+GPU Systems

In this work, we performed an initial design space exploration of an accelerated processing unit (APU)—a hybrid CPU+GPU architecture that integrates both compute units (CUs) and memory into a unified system. This integration aims to reduce data movement, enhance memory locality, and improve energy efficiency by enabling the CPU and GPU to share memory directly. This effort focused on the interplay of key design components—cache line size, the number of CUs, and main memory technology—and the trade-offs of each configuration were analyzed. This paper highlights the various configurations’ impact on memory accesses, data reuse, and power utilization. The results provide valuable insights that can be leveraged to optimize APU architectures for high-performance and energy-efficient computing and thus create a balanced architecture. This optimization can be achieved by adopting dynamic cache management, runtime CU scaling, and advanced memory integration, highlighting the potential of APUs to address critical challenges in compute, data movement, and memory power consumption.

Asifuzzaman, Kazi [ORNL] (ORCID:0000000240044791)↗

LC-MEMENTO: A Memory Model for Accelerated Architectures

With the advent of heterogeneous architectures, in particular, with the ubiquity of multi-GPU systems, it is becoming increasingly important to manage device memory efficiently in order to reap the benefits of the additional core count. To date, such responsibility mainly falls on the programmer where device-to-host data communication (and vice versa), if not done properly, may incur costly memory transfer operations and synchronization. The problem may be compounded by additional requirement to maintain system-wide memory consistency that may involve expensive synchronization overhead. In this paper, we present Location Consistency Memory Model for Enhanced Transfer Operations (LC-MEMENTO). This framework considers incorporating runtime techniques for multi-GPU memory management to support relaxed synchronization semantics and memory transfer operations automatically. Specifically, we implement a relaxed form of a memory consistency model based on the Location Consistency (LC) in an Asynchronous Many-Task Runtime (ARTS) and demonstrate that, this memory model enables additional optimization opportunities for the three representative applications encompassing different computational patterns (scientific computation, graphs, data streaming, etc.).

Memory Models, Accelerators, Adaptive Optimization↗

Noisy Intermediate-Scale Quantum Applications on a Pathfinder System

Work performed under this one-year LDRD was concerned with estimating resource requirements for small quantum test beds that are expected to be available in the near future. This work represents a preliminary demonstration of our ability to leverage quantum hardware for solving small quantum simulation problems in areas of interest to the DOE. The algorithms enabling such studies are hybrid quantum-classical variational algorithms, in particular the widely-used variational quantum eigensolver (VQE). Employing this hybrid algorithm, in which the quantum computer complements the classical one, we implemented an end-to-end application-level toolchain that allows the user to specify a molecule of interest and compute the ground state energy using the VQE approach. We found significant limitations attributable to the classical portion of the hybrid system, including a greater than greater-than-quartic power scaling of the classical memory requirements with the system size. Current VQE approaches would require an exascale machine for solving any molecule with size greater than 150 nuclei. Our findings include several improvements that we implemented into the VQE toolchain, including a new classical optimizer that is decades old but hadn't been considered before in the VQE ecosystem. Our findings suggest limitations to variational hybrid approaches to simulation that further motivate the need for a gate-based fault-tolerant quantum processor that can implement larger problems using the fully digital quantum phase estimation algorithm.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Towards Superior Software Portability with SHAD and HPX C++ Libraries

As hardware architectures and software stacks complexity grows, development productivity, performance and software portability, quickly evolve from desirable features to actual needs. SHAD, the Scalable High-performance Algorithms and Data-structures C++ library is designed to mitigate these issues: it provides general purpose building blocks as well as high-level custom utilities, and offers a shared-memory programming abstraction which facilitates the programming of complex systems, scaling up to High Performance Computing clusters. SHAD’s portability is achieved through an abstract runtime interface, which decouples the upper layers of the library and hides the low level details of the underlying architecture. This layer enables SHAD to interface with different runtime/threading systems, e.g. Intel TBB and Global Memory and Threading (GMT). However, current backends targeting distributed systems, rely on a centralized controller which may possibly limit scalability up to hundreds of nodes and creates a network hot spot due to all to one communication for synchronization, and possibly resulting in degraded performance at high process counts. In this research, we explore HPX, the C++ standard library for parallelism and concurrency, as an additional backend in support of the SHAD library, and present the methodologies in support of local and remote task executions in SHAD with respect to HPX. Finally, we evaluate the proposed system by comparing against existing backends of SHAD and analyzing their performance on C++ Standard Template Library algorithms.

Wu, Nanmiao↗

Memory hierarchy using page-based compression

A system includes a device coupleable to a first memory. The device includes a second memory to cache data from the first memory. The second memory is to store a set of compressed pages of the first memory and a set of page descriptors. Each compressed page includes a set of compressed data blocks. Each page descriptor represents a corresponding page and includes a set of location identifiers that identify the locations of the compressed data blocks of the corresponding page in the second memory. The device further includes compression logic to compress data blocks of a page to be stored to the second memory and decompression logic to decompress compressed data blocks of a page accessed from the second memory.

Loh, Gabriel H.↗

The Evolution of Volatile Memory Forensics

The collection and analysis of volatile memory is a vibrant area of research in the cybersecurity community. The ever-evolving and growing threat landscape is trending towards fileless malware, which avoids traditional detection but can be found by examining a system’s random access memory (RAM). Additionally, volatile memory analysis offers great insight into other malicious vectors. It contains fragments of encrypted files’ contents, as well as lists of running processes, imported modules, and network connections, all of which are difficult or impossible to extract from the file system. For these compelling reasons, recent research efforts have focused on the collection of memory snapshots and methods to analyze them for the presence of malware. However, to the best of our knowledge, no current reviews or surveys exist that systematize the research on both memory acquisition and analysis. We fill that gap with this novel survey by exploring the state-of-the-art tools and techniques for volatile memory acquisition and analysis for malware identification. For memory acquisition methods, we explore the trade-offs many techniques make between snapshot quality, performance overhead, and security. For memory analysis, we examined the traditional forensic methods used, including signature-based methods, dynamic methods performed in a sandbox environment, as well as machine learning-based approaches. We summarize the currently available tools, and suggest areas for more research.

Nyholm, Hannah↗

Pushing the limits of NAND technology scaling with ferroelectrics

Artificial intelligence (AI) continues to drive transformative advancements across various industries. The data-intensive nature of AI training (and inferencing) has resulted in the generation of unprecedented volumes of data with machine-generated content surpassing human-generated data by more than 100-fold in 2025. Efficiently managing this data influx necessitates advanced digital storage technologies. However, traditional NAND flash memory, which is critical for supporting data flows in AI systems—alongside high-bandwidth memory, for AI training—faces fundamental scaling limitations as it approaches the 1000-layer milestone, encompassing more than 40 trillion transistors. This article delves into the potential of hafnia-based ferroelectric materials as a breakthrough solution to these challenges. Recent advancements indicate that the intrinsic limitations of ferroelectric field-effect transistors (FEFETs) can be mitigated through material and device-level engineering. These advancements enable FEFETs to meet the stringent density, reliability, and scalability requirements of future three-dimensional NAND technology. The role of ferroelectrics in addressing NAND scaling challenges and expanding storage capabilities presents a promising avenue for meeting the storage demands of the AI-driven era.

3D NAND↗

Emergent disorder and mechanical memory in periodic metamaterials

Ordered mechanical systems typically have one or only a few stable rest configurations, and hence are not considered useful for encoding memory. Multistable and history-dependent responses usually emerge from quenched disorder, for example in amorphous solids or crumpled sheets. In contrast, due to geometric frustration, periodic magnetic systems can create their own disorder and espouse an extensive manifold of quasi-degenerate configurations. Inspired by the topological structure of frustrated artificial spin ices, we introduce an approach to design ordered, periodic mechanical metamaterials that exhibit an extensive set of spatially disordered states. While our design exploits the correspondence between frustration in magnetism and incompatibility in meta-mechanics, our mechanical systems encompass continuous degrees of freedom, and thus generalize their magnetic counterparts. We show how such systems exhibit non-Abelian and history-dependent responses, as their state can depend on the order in which external manipulations were applied. We demonstrate how this richness of the dynamics enables to recognize, from a static measurement of the final state, the sequence of operations that an extended system underwent. Thus, multistability and potential to perform computation emerge from geometric frustration in ordered mechanical lattices that create their own disorder.

36 MATERIALS SCIENCE↗

Telecom Networking with a Diamond Quantum Memory

Practical quantum networks require interfacing quantum memories with existing channels and systems that operate in the telecom band. Here we demonstrate low-noise, bidirectional quantum frequency conversion that enables a solid-state quantum memory to directly interface with telecom-band systems. In particular, we demonstrate conversion of visible-band single photons emitted from a silicon-vacancy ( Si V ) center in diamond to the telecom O band, maintaining low noise ( g 2 ( 0 ) < 0.1 ) and high indistinguishability ( V = 89 ± 8 % ). We further demonstrate the utility of this system for quantum networking by converting telecom-band time-bin pulses, sent across a lossy and noisy 50-km deployed fiber link, to the visible band and entangling them with a diamond quantum memory with fidelity F ≥ 87 ± 2.5 % . These results demonstrate the viability of Si V quantum memories integrated with telecom-band systems for scalable quantum networking applications. Published by the American Physical Society 2024

Physics↗

Extensions to the SENSEI In situ Framework for Heterogeneous Architectures

The proliferation of GPUs and accelerators in recent supercomputing systems, so called heterogeneous architectures, has led to increased complexity in execution environments and programming models as well as to deeper memory hierarchies on these systems. In this work, we discuss challenges that arise in in situ code coupling on these heterogeneous architectures. In particular, we present data and execution model extensions to the SENSEI in situ framework that are targeted at the effective use of systems with heterogeneous architectures. We then use these new data and execution model extensions to investigate several in situ placement and execution configurations and to analyze the impact these choices have on overall performance.

Loring, Burlen↗

Seasonal forecasting skill for the High Mountain Asia region in the Goddard Earth Observing System

Seasonal variability of the global hydrologic cycle directly impacts human activities, including hazard assessment and mitigation, agricultural decisions, and water resources management. This is particularly true across the High Mountain Asia (HMA) region, where availability of water resources can change depending on local seasonality of the hydrologic cycle. Forecasting the atmospheric states and surface conditions, including hydrometeorologically relevant variables, at subseasonal-to-seasonal (S2S) lead times of weeks to months is an area of active research and development. NASA's Goddard Earth Observing System (GEOS) S2S prediction system has been developed with this research goal in mind. Here, we benchmark the forecast skill of GEOS-S2S (version 2) hydrometeorological forecasts at 1–3-month lead times in the HMA region, including a portion of the Indian subcontinent, during the retrospective forecast period, 1981–2016. To assess forecast skill, we evaluate 2 m air temperature, total precipitation, fractional snow cover, snow water equivalent, surface soil moisture, and terrestrial water storage forecasts against the Modern-Era Retrospective analysis for Research and Applications, Version 2 (MERRA-2) and independent reanalysis data, satellite observations, and data fusion products. Anomaly correlation is highest when the forecasts are evaluated against MERRA-2 and particularly in variables with long memory in the climate system, likely due to the similar initial conditions and model architecture used in GEOS-S2S and MERRA-2. When compared to MERRA-2, results for the 1-month forecast skill range from an anomaly correlation of R anom =0.18 for precipitation to R anom =0.62 for soil moisture. Anomaly correlations are consistently lower when forecasts are evaluated against independent observations; results for the 1-month forecast skill range from R anom =0.13 for snow water equivalent to R anom =0.24 for fractional snow cover. We find that, generally, hydrometeorological forecast skill is dependent on the forecast lead time, the memory of the variable within the physical system, and the validation dataset used. Overall, these results benchmark the GEOS-S2S system's ability to forecast HMA hydrometeorology.

54 ENVIRONMENTAL SCIENCES↗