Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Memory systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Two-Level Sketching Alternating Anderson Acceleration for Complex Physics Applications

We present a novel two-level sketching extension of the Alternating Anderson–Picard (AAP) method for accelerating fixed-point iterations in challenging single- and multiphysics simulations governed by discretized PDEs. Our approach combines a static, physics-based projection that reduces the least-squares (LS) problem to the most informative field (e.g., via Schur-complement insight) with a dynamic, algebraic sketching stage driven by a backward stability analysis under Lipschitz continuity. We introduce inexpensive estimators for stability thresholds and cache-aware randomized selection strategies to balance computational cost against memory access overhead. The resulting algorithm solves reduced LS systems in place, minimizes memory footprints, and seamlessly alternates between low-cost Picard updates and Anderson mixing. Implemented in Julia, our two-level sketching AAP achieves up to 50% time-to-solution reductions compared to standard Anderson acceleration—without degrading convergence rates—on benchmark problems including Stokes, 𝑝-Laplacian, bidomain, and Navier–Stokes formulations at varying problem sizes. These results demonstrate the method’s robustness, scalability, and potential for integration into high-performance scientific computing frameworks. Our implementation is available open source in the AAP.jl library.

Barnafi, Nicolas [University of Chile, Santiago]↗

Systems and methods for shock-resistant memory devices

A shock-resistant memory device comprises a housing and a memory module. The memory module is disposed within the housing and surrounded by potting material to protect the memory module from damage during a shock event. The housing can include a port that accommodates a data connection between the memory module and a sensor from which data is desirably received by the memory module. During a shock event the connection between the memory module and the sensor may be severed, but data stored in the memory module can be retained in the memory module which is protected by the housing. To facilitate retrieval of the memory device subsequent to a shock event, a balloon can be affixed to the housing. The balloon can be configured to inflate subsequent to the shock event so that the shock-resistant memory device does not sink in water and to make the memory device more visible for recovery.

42 ENGINEERING↗

Reliability processing of remote direct memory access

Methods and systems for monitoring remote transmissions of messages among a plurality of nodes are described. A processing element in a first node may allocate a sequence number to a request to read and/or update data in a second node. The processing element may be different from main processors of the first node. The processing element may send the message and the sequence number to the second node. The processing element may modify a status of the sequence number to an active state, indicating a transmission of the message is pending. The processing element may, in response to a response from the second node, modify the status of the sequence number to an inactive state, indicating a completed transmission of the message. The processing element may, in response to no response from the second node within a time period, resend the message and the sequence number to the second node.

Kumar, Sameer↗

Ergodicity, lack thereof, and the performance of reservoir computing with memristive networks and nanowire

Networks composed of nanoscale memristive components, such as nanowire and nanoparticle networks, have recently received considerable attention because of their potential use as neuromorphic devices. In this study, we explore ergodicity in memristive networks, showing that the performance on machine leaning tasks improves when these networks are tuned to operate at the edge between two global stability points. We find this lack of ergodicity is associated with the emergence of memory in the system. We measure the level of ergodicity using the Thirumalai-Mountain metric, and we show that in the absence of ergodicity, two different memristive network systems show improved performance when utilized as reservoir computers (RC). We highlight that it is also important to let the system synchronize to the input signal in order for the performance of the RC to exhibit improvements over the baseline.

97 MATHEMATICS AND COMPUTING↗

GPU Direct I/O with HDF5

Exascale HPC systems are being designed with accelerators, such as GPUs, to accelerate parts of applications. In machine learning workloads as well as large-scale simulations that use GPUs as accelerators, the CPU (or host) memory is currently used as a buffer for data transfers between GPU (or device) memory and the file system. If the CPU does not need to operate on the data, then this is sub-optimal because it wastes host memory by reserving space for duplicated data. Furthermore, this “bounce buffer” approach wastes CPU cycles spent on transferring data. A new technique, NVIDIA GPUDirect Storage (GDS), can eliminate the need to use the host memory as a bounce buffer. Thereby, it becomes possible to transfer data directly between the device memory and the file system. This direct data path shortens latency by omitting the extra copy and enables higher-bandwidth. To take full advantage of GDS in existing applications, it is necessary to provide support with existing I/O libraries, such as HDF5 and MPI-IO, which are heavily used in applications. In this paper, we describe our effort of integrating GDS with HDF5, the top I/O library at NERSC and at DOE leadership computing facilities. We design and implement this integration using a HDF5 Virtual File Driver (VFD). The GDS VFD provides a file system abstraction to the application that allows HDF5 applications to perform I/O without the need to move data between CPUs and GPUs explicitly. We compare performance of the HDF5 GDS VFD with explicit data movement approaches and demonstrate superior performance with the GDS method.

Ravi, J↗

Ergodicity, lack thereof, and the performance of reservoir computing with memristive networks

Abstract Networks composed of nanoscale memristive components, such as nanowire and nanoparticle networks, have recently received considerable attention because of their potential use as neuromorphic devices. In this study, we explore ergodicity in memristive networks, showing that the performance on machine leaning tasks improves when these networks are tuned to operate at the edge between two global stability points. We find this lack of ergodicity is associated with the emergence of memory in the system. We measure the level of ergodicity using the Thirumalai-Mountain metric, and we show that in the absence of ergodicity, two different memristive network systems show improved performance when utilized as reservoir computers (RC). We highlight that it is also important to let the system synchronize to the input signal in order for the performance of the RC to exhibit improvements over the baseline.

97 MATHEMATICS AND COMPUTING↗

Memory page access counts based on page refresh

A processing system tracks counts of accesses to memory pages using a set of counters located at the memory module that stores the pages, wherein the counts are adjusted at least in part based on refreshes of the memory pages. This approach allows a processing system to efficiently maintain the counts with relatively small counters and with relatively low overhead. Furthermore, the rate at which the counters are adjusted, relative to the page refreshes, is adjustable, so that the access counts are useful for a wide variety of application types.

97 MATHEMATICS AND COMPUTING↗

Exceptional magnetic and magnetoelastic behavior of rare-earth non-centrosymmetric Sm 7 Pd 3

Magnetic compounds possessing an intrinsic combination of near-zero magnetization with high magnetic anisotropy are highly desirable for spinronic applications and memory recording. A comprehensive study of Sm 7 Pd 3 binary compound uncovered a unique combination of strong magnetoelastic behavior, very low net magnetization, and exceptionally high magnetic coercivity. The temperature-dependent X-ray synchrotron powder diffraction study indicates the abrupt changes in the compound's lattice parameters at the magnetic ordering temperature of T C =169 K, although the crystal structure remains non-centrosymmetric hexagonal Th 7 Fe 3 -type down to 6 K. Density functional theory calculations confirm high intrinsic magnetocrystalline anisotropy of Sm 7 Pd 3 , which explains the extremely large coercivity of the polycrystalline sample, up to H cr = 130 kOe at 2 K. In conclusion, this discovery brings to life a novel class of highly anisotropic materials that are distinctly different from known spintronic materials, making them interesting future systems for magnetic memory research.

36 MATERIALS SCIENCE↗

Co-design of Advanced Architectures for Graph Analytics using Machine Learning

A graph is an excellent way of representing relationships among entities. We can use graph analytics to synthesize and analyze such relational data, and extract relevant features that are useful for various tasks such as machine learning. Considering the crucial role of graph analytics in various domains, it is important and timely to investigate the right hardware configurations that can achieve optimal performance for graph workloads on future high-performance computing systems. Design space exploration studies facilitate the selection of appropriate configurations (e.g. memory) to achieve a desired system performance. Recently, the approach of accelerating graph analytics using persistent non-volatile memory has gained a lot of attention. Traditional system simulators such as Gem5 and NVMain can be used to explore the design space of these advanced memory architectures for graph workloads. However, these simulators are slow in execution thus limiting the efficiency of design space exploration studies. To overcome this challenge, we proposed a machine learning based approach to co-design advanced memory architectures for graph workloads. We tested our approach with DRAM, non-volatile memory, and hybrid memory (DRAM+NVM) using a breadth first search benchmark algorithm. Our results showed the applicability of the proposed machine learning based approach to the co-design of the advanced memory architectures. In this paper, we provide recommendations on selecting advanced memory architectures to achieve desired performance for graph workloads. We also discuss the performances of different machine learning models that were considered in this study.

Kurte, Kuldeep↗

Tunable Interfacial to Filamentary Resistive Switching Mechanism in Room-Temperature-Grown Amorphous YBa 2 Cu 3 O x with Excess Cu Addition

Resistive switching technologies have the potential not only to create large efficiency gains in computer memory but also to revolutionize emerging fields such as neuromorphic computing. In this paper, we report on novel resistive switching behavior in devices made from room-temperature-grown Cu-rich amorphous YBa 2 Cu 3 O x (YBCO) films, a material otherwise well-known as a high-temperature superconductor. In Nb:STO substrate/amorphous YBCO film (≈200 nm)/metallic Cu (15 nm)/metallic Pt (15 nm) devices, we demonstrate that the resistive switching can be tuned between mechanisms involving extended areas of the YBCO/electrode interface and a single-point filamentary mechanism simply by changing the Cu content of the deposition target and hence in the films. Changing the Cu content can also be used to optimize the properties of the devices further, with devices with an added 15 mol % of Cu in YBCO initially providing an on/off ratio >100, switching endurance potential >6500 cycles, and state retention >2 × 10 4 s, all at low switching fields of 0.3 MV/cm. The amalgam of promising resistive switching properties, fast growth (150 nm/min) at room temperature, and tuneability of the switching mechanism indicates the strong potential of this proof-of-concept amorphous system for future memory applications.

Cu↗

Effect of Gamma Radiation on TaOₓ ECRAM

Electrochemical random access memory (ECRAM) is an emerging three-terminal nonvolatile memory (NVM) with highly controllable channel conductance which is promising for use as an analog memory (or synapse) in analog in-memory computing (IMC) systems. Energy-efficient analog IMC computing is particularly desirable for power-constrained, high-radiation environments such as satellites. However, little is known about the suitability of ECRAM for use in a total ionizing dose (TID) environment. Here, this work investigates the effect of Co-60 gamma radiation on the channel conductance and noise—two properties critical for analog IMC systems—of a TaO x -based ECRAM up to 17.3 Mrad(SiO 2 ) for both low- and high-channel-conductance state devices. A transient increase in conductance is observed in response to radiation which consists of two elements: an immediate increase in conductivity due to photocurrent and a secondary increase in conductivity, which has a slower rise and saturation and can persist for hours after exposure. This secondary, persistent photoconductivity is attributed to charging caused by hole trapping. These transient effects would not likely occur in a space environment due to the low dose rate compared with this experiment. No permanent change is found in the low conductance state (LCS) following exposure and the minor shift in the high conductance change would be less significant than the regular retention decay in this state. A permanent increase in the random telegraph noise is observed, possibly due to increased traps created in the channel. This work demonstrates that TaO x -based ECRAM is suitable for use in spaceborne analog IMC systems that are subject to significant TID.

ECRAM↗

Braiding for the win: Harnessing braiding statistics in topological states to play quantum games

Nonlocal quantum games provide proof of principle that quantum resources can confer an advantage at certain tasks. They also provide a compelling way to explore the computational utility of phases of matter on quantum hardware. In a recent paper [O. Hart et al., Phys. Rev. Lett. 134, 130602 (2025)], we demonstrated that a toric code resource state conferred advantage at a certain nonlocal game, which remained robust to small deformations of the resource state. In this paper we demonstrate that this robust advantage is a generic property of resource states drawn from topological or fracton ordered phases of quantum matter. To this end, we illustrate how several other states from paradigmatic topological and fracton ordered phases can function as resources for suitably defined nonlocal games, notably the three-dimensional toric-code phase, the X-cube fracton phase, and the double-semion phase. The key in every case is to design a nonlocal game that harnesses the characteristic braiding processes of a quantum phase as a source of contextuality. We unify the strategies that take advantage of mutual statistics by relating the operators to be measured to order and disorder parameters of an underlying generalized symmetry-breaking phase transition. Additionally, by connecting the win probability to twist products, we show that success at the game serves as a many-body entanglement witness. Namely, if the players implement a perfect quantum strategy on large length scales, the quantum state they share cannot be connected to a trivial product state via a constant-depth local unitary circuit. Lastly, we massively generalize the family of games that admit perfect strategies when codewords of homological quantum error-correcting codes are used as resources.

Fractons↗

Analysis of Vector Particle-In-Cell (VPIC) memory usage optimizations on cutting-edge computer architectures

Vector Particle-In-Cell (VPIC) is one of the fastest plasma simulation codes in the world, with particle numbers ranging from one trillion on the first petascale system, Roadrunner, to ten trillion particles on the more recent Blue Waters supercomputer. As supercomputers continue to grow rapidly in size, so too does the gap between computing capability and memory capability. Current memory systems limit VPIC simulations greatly as the maximum number of particles that can be simulated directly depends on the available memory. In this study, we present a suite of VPIC memory optimizations (i.e., particle weight, half-precision, and fixed-point optimizations) that enable a significant increase in the number of particles in VPIC simulations. Here, we assess the optimizations’ impact on memory and runtime performance for a suite of cutting-edge computer architectures such has the NVIDIA V100 GPU, the IBM Power9, and the Fujitsu A64FX architectures. Our optimizations enable a 31.25% reduction in memory usage and up to 40% increase in the number of particles. This paper extends our work on developing particle storage format optimizations Tan et al.

97 MATHEMATICS AND COMPUTING↗

Diurnal cycle of precipitation over the tropics and central United States: intercomparison of general circulation models

Diurnal precipitation is a fundamental mode of variability that climate models have difficulty in accurately simulating. Here, in this work, the diurnal cycle of precipitation (DCP) in participating climate models from the Global Energy and Water Exchanges' DCP project is evaluated over the tropics and central United States. Common model biases such as excessive precipitation over the tropics, too frequent light-to-moderate rain, and the failure to capture propagating convection in the central United States still exist. Over the central United States, the issues of too weak rainfall intensity in climate runs is well improved in their hindcast runs with initial conditions from numerical weather prediction analyses. But the improvement is minimal over the central Amazon. Incorporating the role of the large-scale environment in convective triggering processes helps resolve the phase-locking issue in many models where precipitation often incorrectly peaks near noon due to maximum insolation over land. Allowing air parcels to be lifted above the boundary layer improves the simulation of nocturnal precipitation which is often associated with the propagation of mesoscale systems. Including convective memory in cumulus parameterizations acts to suppress light-to-moderate rain and promote intense rainfall; however, it also weakens the diurnal variability. Simply increasing model resolution (with cumulus parameterizations still used) cannot fully resolve the biases of low-resolution climate models in DCP. The hierarchy modeling framework from this study is useful for identifying the missing physics in models and testing new development of model convective processes over different convective regimes.

58 GEOSCIENCES↗

Low-temperature grapho-epitaxial La-substituted BiFeO 3 on metallic perovskite

Bismuth ferrite has garnered considerable attention as a promising candidate for magnetoelectric spin-orbit coupled logic-in-memory. As model systems, epitaxial BiFeO 3 thin films have typically been deposited at relatively high temperatures (650–800 °C), higher than allowed for direct integration with silicon-CMOS platforms. Here, we circumvent this problem by growing lanthanum-substituted BiFeO 3 at 450 °C (which is reasonably compatible with silicon-CMOS integration) on epitaxial BaPb 0.75 Bi 0.25 O 3 electrodes. Notwithstanding the large lattice mismatch between the La-BiFeO 3 , BaPb 0.75 Bi 0.25 O 3 , and SrTiO 3 (001) substrates, all the layers in the heterostructures are well ordered with a [001] texture. Polarization mapping using atomic resolution STEM imaging and vector mapping established the short-range polarization ordering in the low temperature grown La-BiFeO 3 . Current-voltage, pulsed-switching, fatigue, and retention measurements follow the characteristic behavior of high-temperature grown La-BiFeO 3 , where SrRuO 3 typically serves as the metallic electrode. These results provide a possible route for realizing epitaxial multiferroics on complex-oxide buffer layers at low temperatures and opens the door for potential silicon-CMOS integration.

36 MATERIALS SCIENCE↗

Application of mesh refinement to relativistic magnetic reconnection

During relativistic magnetic reconnection, antiparallel magnetic fields undergo a rapid change in topology, releasing a large amount of energy in the form of non-thermal particle acceleration. This work explores the application of mesh refinement to 2D reconnection simulations to efficiently model the inherent disparity in length-scales. We have systematically investigated the effects of mesh refinement and determined necessary modifications to the algorithm required to mitigate non-physical artifacts at the coarse–fine interface. We have used the ultrahigh-order pseudo-spectral analytical time-domain Maxwell solver to analyze how its use can mitigate the numerical dispersion that occurs with the finite-difference time-domain (or “Yee”) method. Absorbing layers are introduced at the coarse–fine interface to eliminate spurious effects that occur with mesh refinement. We also study how damping the electromagnetic fields and current density in the absorbing layer can help prevent the non-physical accumulation of charge and current density at the coarse–fine interface. Using a mesh refinement ratio of 8 for two-dimensional magnetic reconnection simulations, we obtained good agreement with the high-resolution baseline simulation, using only 36% of the macroparticles and 71% of the node-hours needed for the baseline. The methods presented here are especially applicable to 3D systems where higher memory savings are expected than in 2D, enabling comprehensive, computationally efficient 3D reconnection studies in the future.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Differential Power Processing for Ultra-Efficient Data Storage

Here this paper presents the hardware, software, and power codesign of an ultra-efficient data storage server with differential power processing (DPP). DPP can reduce the power conversion stress, improve the efficiency, and enhance the functionality of modular power electronics systems. The power inputs of a large number of hard disk drives (HDDs) were connected in series and supported by a multiport ac-coupled differential power processing (MAC-DPP) converter through a multiwinding transformer. Methods for controlling the multi-input multi-output power flow in the multiwinding transformer while avoiding core saturation were investigated. A ten-port MAC-DPP prototype with 700-W/in 3 power density was built to support a 450-W HDD storage system with ten series-stacked voltage domains. The prototype was tested on a 50-HDD server testbench, and the overall system loss is below 1 W (99.77% system efficiency). The server was able to maintain high-speed reading and writing operation of all 50 HDDs against the worst hot-swapping scenarios. A variety of hardware/software configurations and many cloud storage techniques were tested on the fully functioning server. Experimental results show that the energy efficiency of large-scale information systems (CPU/GPU clusters, memory banks, HDD arrays, etc.) can be greatly improved by software, hardware, and power codesign.

42 ENGINEERING↗

pyDRESCALk

Modern data scientists are tasked to analyze ever-growing data sets with increasingly complex relationships. Tensor decompositions have come to play a central role in identifying underlying latent structures in higher-order data. The problem of fitting tensor models to different distributions is complicated by the combinations of size, dimensionality, and sparsity present in real world data. The situation demands efficient algorithms designed for shared-memory and distributed systems. This work will present new research that tackles these challenges on several different fronts, leveraging optimizations in numerical algorithms and sparse tensor representations in heterogeneous high performance computing environments.

Bhattarai, Manish↗