Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “limited memory”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

UltraLiM: In-Memory Boolean Logic Architecture Using UltraRAM

Conventional computing architectures encounter ‘von Neumann’ and ‘memory wall’ bottlenecks which arise due to the back-and-forth data movement between the physically separate memory and processing units and the speed mismatch between them, respectively. These bottlenecks hurt both energy efficiency and the throughput of computing systems. To address these challenges, in-memory computing architectures have emerged as a promising alternative. They reduce the need for frequent data movement by executing different computing tasks inside the memory system. Here, we present UltraLiM, a logic-in-memory architecture using the UltraRAM-based memory system. UltraRAM holds the promise of developing a ‘universal memory’, overcoming the limitations of charge-based memories thanks to their non-volatile behavior with lower operating voltage. This work presents an in-memory computing architecture that integrates an UltraRAM-based memory array with a custom-designed peripheral circuitry. With this architecture, we can perform various in-memory Boolean logic operations (such as NOT, NAND, NOR, and XOR) in a single cycle. Leveraging the separate read-write paths in the UltraRAM-based memory array, we optimize read operations without encountering design conflicts. This optimization enhances the sense margin, enabling the use of simpler peripheral circuitry for in-memory logic operations.

Alam, Shamiul [University of Tennessee, Knoxville ↗

Linear complexity

We present factorization and solution phases for a new linear complexity direct solver designed for concurrent batch operations on fine-grained parallel architectures, for matrices amenable to hierarchical representation. We focus on the strong-admissibility-based $\mathscr{H}^{2}$ format, where strong recursive skeletonization factorization compresses remote interactions. We build upon previous implementations of $\mathscr{H}^{2}$ matrix construction for efficient factorization and solution algorithm design, which are illustrated graphically in stepwise detail. The algorithms are ‘blackbox’ in the sense that the only inputs are the matrix and right-hand side, without analytical or geometrical information about the origin of the system. We demonstrate linear complexity scaling in both time and memory on four representative families of dense matrices up to one million in size. Parallel scaling up to 16 threads is enabled by a multi-level matrix graph coloring and avoidance of dynamic memory allocations thanks to prefix-sum memory management. An experimental backward error analysis is included. We break down the timings of different phases, identify phases that are memory-bandwidth limited, and discuss alternatives for phases that may be sensitive to the trend to employ lower precisions for performance.

Boukaram, Wajih↗

Scaling the memory wall using mixed-precision - HPG-MxP on an exascale-class machine

Mixed-precision algorithms have been proposed as a way for scientific computing to benefit from some of the gains seen for AI on recent high performance computing (HPC) platforms. A few applications dominated by dense matrix operations have seen substantial speedups by utilizing low precision formats such as FP16. However, a majority of scientific simulation applications are memory bandwidth limited. Beyond preliminary studies, the practical gain from using mixed-precision algorithms on a given high-performance computing (HPC) system is largely unclear. The High Performance GMRES Mixed Precision (HPG-MxP) benchmark has been proposed to measure the useful performance of a HPC system on sparse matrix-based mixed-precision applications. In this work, we present an implementation of the HPG-MxP benchmark for an exascale system and describe our algorithm enhancements. We show for the first time a speedup of 1.6x using a combination of double- and single-precision keeping the same residual level on modern GPU-based supercomputers.

Kashi, Aditya [ORNL] (ORCID:0000000325893792)↗

A GPU-based compressible combustion solver for applications exhibiting disparate space and time scales

High-speed chemically active flows pose significant computational challenges due to their disparate space and time scales, with stiff chemistry often dominating simulation time. While modern scientific computing programs achieve exascale performance by leveraging graphics processing units (GPUs), existing GPU-based compressible combustion solvers face critical limitations in memory management, load balancing, and handling the highly localized nature of chemical reactions. To this end, we present a high-performance compressible reacting flow solver built on the AMReX framework and optimized for multi-GPU settings. Here, our approach addresses three GPU performance bottlenecks: memory access patterns through column-major storage optimization, computational workload variability via a bulk-sparse integration strategy for chemical kinetics, and multi-GPU load distribution for adaptive mesh refinement applications. The solver adapts existing matrix-based chemical kinetics formulations to multi-grid contexts. Using representative combustion applications, including 2D and 3D detonations and a 3D jet-in-crossflow configuration, we demonstrate 1.4–5× performance improvements over initial implementations on an in-house cluster of NVIDIA H100 GPUs, and near-ideal weak scaling on the Frontier supercomputer (Oak Ridge Leadership Computing Facility) with up to 1024 AMD Instinct MI250X GPUs. Roofline analysis reveals substantial improvements in arithmetic intensity for both convection (∼ 10 ×) and chemistry (∼ 4 ×) routines, confirming efficient utilization of GPU memory bandwidth and computational resources.

42 ENGINEERING↗

Evaluating the potential of disaggregated memory systems for HPC applications

Summary Disaggregated memory is a promising approach that addresses the limitations of traditional memory architectures by enabling memory to be decoupled from compute nodes and shared across a data center. Cloud platforms have deployed such systems to improve overall system memory utilization, but performance can vary across workloads. High‐performance computing (HPC) is crucial in scientific and engineering applications, where HPC machines also face the issue of underutilized memory. As a result, improving system memory utilization while understanding workload performance is essential for HPC operators. Therefore, learning the potential of a disaggregated memory system before deployment is a critical step. This paper proposes a methodology for exploring the design space of a disaggregated memory system. It incorporates key metrics that affect performance on disaggregated memory systems: memory capacity, local and remote memory access ratio, injection bandwidth, and bisection bandwidth, providing an intuitive approach to guide machine configurations based on technology trends and workload characteristics. We apply our methodology to analyze thirteen diverse workloads, including AI training, data analysis, genomics, protein, fusion, atomic nuclei, and traditional HPC bookends. Our methodology demonstrates the ability to comprehend the potential and pitfalls of a disaggregated memory system and provides motivation for machine configurations. Our results show that eleven of our thirteen applications can leverage injection bandwidth disaggregated memory without affecting performance, while one pays a rack bisection bandwidth penalty and two pay the system‐wide bisection bandwidth penalty. In addition, we also show that intra‐rack memory disaggregation would meet the application's memory requirement and provide enough remote memory bandwidth.

Ding, Nan↗

A randomized sketching trust-region secant method for low-memory dynamic optimization

The numerical solution of dynamic optimization problems is often limited by the memory required to store the state trajectory, which is used to evaluate the objective function and its derivatives. Recently, [R. Muthukumar et al., SIAM Journal on Optimization 31(2), pp. 1242–1275 (2021)] introduced a trust-region method for dynamic optimization that employs randomized sketching to compress the state trajectory, resulting in inexact derivative computations. By adaptively learning the sketch rank, the trust-region algorithm achieves rigorous convergence guarantees. Here, we extend this approach to use secant Hessian approximations. Due to the randomness introduced by the sketch, the traditional secant update formulae can produce poor Hessian approximations. In particular, the difference of two gradients, computed from two different sketches, may be inconsistent. To overcome this, we employ a sketched approximation of the Hessian application, in lieu of computing the gradient difference. We numerically demonstrate the improved stability of this approach on an example from PDE-constrained optimization.

dynamic optimization↗

Design Space Exploration of Emerging Memory Technologies for Machine Learning Applications

Memory design space exploration methods study memory systems’ performances and limitations before implementation. The computer memory design space has grown exponentially because of the enormous growth of memory types, memory controllers, and application software. Computer simulators are commonly used for memory design space exploration. However, complex memory simulations take an enormous amount of time. Hence, in this paper, we proposed a machine learning-based design space exploration method for dynamic random-access memory and non-volatile memory systems. We applied our method to the CosmoGAN and LeNet applications to predict the following six memory response parameters: (i) bandwidth, (ii) power, (iii) average latency, (iv) average total latency, (v) memory reads, and (vi) memory writes. Our experimental results show that machine learning models can predict memory response parameter values faster than simulations. We used support vector machine, random forest, and gradient boosting machine learning models. We observed that the support vector machine provides better performance for bandwidth, average latency, and average total latency. The random forest model works better for memory reads and writes. The gradient boosting model provides superior prediction performance for power. We provide a detailed discussion on learning curve characteristics, error analysis, and memory type recommendation.

Hasan, S M Shamimul↗

HAM: Hotspot-Aware Manager for Improving Communications with 3D-Stacked Memory

merging High-Performance Computing (HPC) workloads, such as graph analytics, machine learning, and big data science, are data-intensive. Data-intensive workloads usually present fine-grained memory accesses with limited or no data locality, and thus incur frequent cache misses and low utilization of memory bandwidth. 3D-stacked memory devices such as Hybrid Memory Cube (HMC) and High Bandwidth Memory (HBM) can provide significantly higher bandwidth than conventional memory modules. However, the traditional interfaces and optimization methods for JEDEC DDR devices do not allow to fully exploit the potential performance of 3D-stacked memory with the massive amount of irregular memory accesses of data-intensive applications. In this paper, we propose a novel Hotspot-Aware Manager (HAM) infrastructure for 3D-stacked memory devices capable of optimizing memory access streams via request aggregation, hotspot detection, and in-memory prefetching. %and an associated hotspot-aware page policy. We present the HAM design and implementation, and simulate it on a system using RISC-V embedded cores with attached HMC devices. We extensively evaluate HAM with over 12 benchmarks and applications representing diverse irregular memory access patterns. The results show that, on average, HAM reduces redundant requests by 37.51\% and increases the prefetch buffer hit rate by 4.2 times, compared to a baseline streaming prefetcher. On the selected benchmark set, HAM provides performance gains of 21.81\% in average (up to 34.28\%) and power savings of 35.07\% over a standard 3D-stacked memory.

Wang, Xi↗

The Exploitation of Data Reduction for Visualization

The disparity between the computational speed and storage bandwidth, as demonstrated in Figure 1, is a well known problem that grows with each successive generation. The visualization community is principally responding to this issue by using in situ to reduce which data must be written to storage. However, other communities are taking different, possibly complementary approaches. In particular, data compression is a common general approach to reduce storage demands. Data compression technologies are typically not designed with post processing in mind. The principal metrics measured are compression ratio, the improved bandwidth to storage, and the error introduced. It is assumed that data is inflated to its full size before any post processing can happen. Although when talking about bandwidth disparities, HPC’s dirty little secret is that no part of the memory nor interconnect hardware is increasing at the rate of computation. For example, the Summit supercomputer has a peak computation rate almost 10 times its predecessor, Titan, but only about 4 times the memory, less than twice the aggregate memory bandwidth, and almost no improvement in the interconnect bisection bandwidth. Naively inflating data for post processing does not help with limitations in the memory and interconnect systems.

97 MATHEMATICS AND COMPUTING↗

Advances in relaxation and memory effects of magnetic nanoparticles for biomedical applications

Functionalized magnetic nanoparticles are pivotal in magnetic resonance imaging, computed tomography, controlled drug delivery, and hyperthermia treatments due to their exceptional magnetic relaxation and functional properties. The magnetic core composition and structure significantly affects the complex magnetic properties of these nanoparticles necessitating a thorough examination of magnetism fundamentals related to these systems. One important aspect is the ability of magnetic nanoparticles to retain previous magnetic state configurations known as memory effect, primarily governed by domain structure and magnetic anisotropy. Despite its relevance to advanced applications, comprehensive studies on magnetic relaxation and memory effects remain limited. Here, the present review aims to bridge this gap by investigating relaxation mechanisms, synthesis strategies, and applications, fostering further innovation. It investigates the memory effects and their dependence on particle composition and morphology along with key synthesis techniques for large-scale production in industrial adoption. Structured into focused sections on magnetic properties and their influence on biomedical and technological applications, this review provides essential insights into memory effects, magneto-relaxation mechanisms, influencing factors, and both experimental and theoretical methodologies. It also delves into computational modelling and AI-driven design, which are revolutionizing the prediction, discovery, and optimization of materials with tailored properties.

36 MATERIALS SCIENCE↗

Beyond Binary: Automated PLC Memory Forensics through RGB Image Analysis and Deep Learning

The introduction of Industry 4.0 and the evolution of industrial control systems (ICS) to adopt Internet-based technologies enhanced productivity, but have inadvertently increased their vulnerability to cyber-based malicious attacks. When an ICS system is compromised, security analysts need to identify the root cause quickly to start the recovery process and develop mitigation strategies to safeguard against future instances. Memory forensics is critical in the analysis process to ascertain what occurred. To date, approaches to analyze the persistent memory in ICS devices are limited, and almost nonexistent for volatile memory. This paper proposes an automated methodology, COMA, for PLC memory dump analysis using computer vision and deep learning techniques. Specifically, COMA converts the sequences of bytes in a PLC memory dump to RGB pixels and creates a deep learning model that learns the underlying patterns and features of pre-labeled forensic artifacts in images and segments them into distinct regions. COMA then uses the trained model to automatically segment new memory images and extract forensic artifacts. We evaluate COMA on a Schneider Electric Modicon M221 PLC involving two cyber-based attack scenarios: (i) code injection and (ii) code modification. The empirical results show that COMA can successfully detect attack artifacts in memory dumps in both scenarios.

Asmar Awad, Rima↗

Improving scalability of electronic structure code for molecular simulations in the presence of environment

A scalable density functional electronic code with Gaussian basis set, called UTEP-NRLMOL, is developed to perform simulations of molecular systems in the presence of the environment with particular attention to the memory requirements. In the electronic structure calculations, the memory and computation time are proportional to the number of atoms. Memory requirements for density functional calculations scale as N*N, where N is the number of atoms. While the recent advances in HPC offer platforms with large numbers of cores, the limited amount of memory available on a given node and poor scalability of the electronic structure codes hinder their efficient usage of these platforms. We have introduced new scaling and parallelization paradigms using MPI-3 shared-memory functionality combined with usage of sparse algebra and storage of matrices in sparse format. This extends the range of applicability of the UTEP-NRLMOL code to large systems over 10,000 atoms, or using up to 67,000 basis functions, and making use of HPC architectures using over 6,000 processors utilizing all available cores. We have also interfaced code with effective fragment potential and polarizable continuum model libraries. The code was used in simulations of several applications which are published in reputed scientific journals.

74 ATOMIC AND MOLECULAR PHYSICS↗

A scalable superconducting nanowire memory array with row–column addressing

Scalable superconducting memory is required for the development of low-energy superconducting computers and fault-tolerant quantum computers. Conventional superconducting logic-based memory cells possess a large footprint that limits scaling; nanowire-based superconducting memory cells, although more compact, have high error rates, which hinders integration into large arrays. Here we report a 4 × 4 superconducting nanowire memory array that is designed for scalable row–column operations and has a functional density of 2.6 Mbit cm −2 . Each memory cell is based on a nanowire loop consisting of two temperature-dependent superconducting switches and a variable kinetic inductor. The arrays operate at 1.3 K, where we implement and characterize multiflux quanta state storage and destructive read-out. By optimizing the write- and read-pulse sequences, we minimize bit errors and maximize operating margins. We achieve a minimum bit error rate of 10 −5 . Here, we also use circuit-level simulations to understand the memory cell’s dynamics, performance limits and stability under varying pulse amplitudes.

Electrical and electronic engineering↗

Distributed Multi-GPU Community Detection on Exascale Computing Platforms

Community detection is a fundamental operation in graph mining, and by uncovering hidden structures and patterns within complex systems it helps solve fundamental problems pertaining to social networks, such as information diffusion, epidemics, and recommender systems. Scaling graph algorithms for massive networks becomes challenging on modern distributed-memory multi-GPU (Graphics Processing Unit) systems due to limitations such as irregular memory access patterns, load imbalances, higher communication-computation ratios, and cross-platform support. We present a novel algorithm HiPDPL-GPU (distributed parallel Louvain) to address these challenges. We conduct experiments involving different partitioning techniques to achieve optimized performance of HiPDPL-GPU on the two largest supercomputers: Frontier and Summit. Remarkably, HiPDPL-GPU processes a graph with 4.2 billion edges in less than 3 minutes using 1024 GPUs. Qualitatively performance of HiPDPL-GPU is similar or better compared to other state-of-the-art CPU- and GPU-based implementations. While prior GPU implementations have predominantly employed CUDA, our first-of-its-kind implementation for community detection is cross-platform, accommodating both AMD and NVIDIA GPUs.

graph algorithms, high performance comptuing↗

Distributed Multi-GPU Community Detection on Exascale Computing Platforms

Community detection is a fundamental operation in graph mining, and by uncovering hidden structures and patterns within complex systems it helps solve fundamental problems pertaining to social networks, such as information diffusion, epidemics, and recommender systems. Scaling graph algorithms for massive networks becomes challenging on modern distributed-memory multi-GPU (Graphics Processing Unit) systems due to limitations such as irregular memory access patterns, load imbalances, higher communication-computation ratios, and cross-platform support. We present a novel algorithm HiPDPL-GPU (Distributed Parallel Louvain) to address these challenges. We conduct experiments involving different partitioning techniques to achieve an optimized performance of HiPDPL-GPU on the two largest supercomputers: Frontier and Summit. Remarkably, HiPDPL-GPU processes a graph with 4.2 billion edges in less than 3 minutes using 1024 GPUs. Qualitatively, the performance of HiPDPL-GPU is similar or better compared to other state-of-the-art CPU- and GPU-based implementations. While prior GPU implementations have predominantly employed CUDA, our first-of-its-kind implementation for community detection is cross-platform, accommodating both AMD and NVIDIA GPUs.

Sattar, Naw Safrin↗

Automated Programmable Logic Controller Memory Forensics Using RGB Image Analysis and Deep Learning

The introduction of Industry 4.0 and Internet-based technologies has enhanced industrial control system operations but have inadvertently increased their vulnerabilities to cyber attacks. When an industrial control system is compromised, security analysts need to identify the root cause quickly to start the recovery process and develop mitigation strategies. Memory forensics is critical in the incident analysis process to ascertain what occurred. Approaches for analyzing the persistent memory in industrial control devices are limited and almost nonexistent for volatile memory. This chapter proposes an automated methodology for programmable logic controller memory dump analysis using computer vision and deep learning techniques. The methodology converts the sequences of bytes in a programmable logic controller memory dump to red-green-blue pixels and employs a deep learning model that learns the underlying patterns and features of pre-labeled forensic artifacts in images and segments them into distinct regions. The trained model is employed to automatically segment new memory images and identify forensic artifacts. Evaluation of the methodology on a Schneider Electric Modicon M221 programmable logic controller under code injection and code modification attacks demonstrates its ability to detect attack artifacts in memory dumps.

Asmar Awad, Rima [ORNL] (ORCID:0000000233407742)↗

Intrinsic optical bistability of photon avalanching nanocrystals

Optically bistable materials respond to a single input with two possible optical outputs, contingent on excitation history. Such materials would be ideal for optical switching and memory, but the limited understanding of intrinsic optical bistability (IOB) prevents the development of nanoscale IOB materials suitable for devices. Here we demonstrate IOB in Nd 3+ -doped KPb 2 Cl 5 avalanching nanoparticles, which switch with high contrast between luminescent and non-luminescent states, with hysteresis characteristic of bistability. Here we elucidate a non-thermal mechanism in which IOB originates from suppressed non-radiative relaxation in Nd 3+ ions and from the positive feedback of photon avalanching, resulting in extreme, >200th-order optical nonlinearities. The modulation of laser pulsing tunes the hysteresis widths, and dual-laser excitation enables transistor-like optical switching. This control over nanoscale IOB establishes avalanching nanoparticles for photonic devices in which light is used to manipulate light.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗