Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “memory spaces”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Logic in Memory Emulator

Logic in Memory Emulator (LiME) is a hardware/software tool specially designed for memory system evaluation and experiment. Emerging memories display a wide range of bandwidths, latencies, and capacities, making it challenging for the computer architect to navigate the design space of potential memory configurations, and for the application developer to assess performance implications of using such memories. With the LiME framework, architectural ideas can be prototyped in great detail yet with sufficient performance to support realistic evaluation on long running applications. LiME consists of two fundamental components: 1) the hardware and OS infrastructure for the emulator, and 2) a suite of benchmark applications to assist in characterizing the performance of current and future computer architectures. Some of the applications have been collected from other open source projects. Uses: Logging, replay and analysis of an application's memory behavior Evaluate impact of emerging memory technology on application performance. Emulate complex memory interactions in whole applications orders of magnitude faster than software simulation. Emulate acceleration hardware co-located with the memory subsystem. Features: Capture and log external memory accesses to a separate off-chip memory device without affecting application execution. Memory traces include the address, length, timestamp, and optionally the data for each transaction. Captured trace data can be saved to an SD card for off-line analysis. Configure a wide range of memory latencies in sub-nanosecond increments that encompass highbandwidth and storage class memories. Specify regions of interest (ROI) in applications to reduce the amount of trace data captured for analysis. Currently supports execution on Xilinx Zynq SoC which integrates an ARM processor with FPGA logic on a single device. Applications can be run under Linux or in bare metal mode on the ARM cores.

Jain, AbhishekK↗

GPU Direct I/O with HDF5

Exascale HPC systems are being designed with accelerators, such as GPUs, to accelerate parts of applications. In machine learning workloads as well as large-scale simulations that use GPUs as accelerators, the CPU (or host) memory is currently used as a buffer for data transfers between GPU (or device) memory and the file system. If the CPU does not need to operate on the data, then this is sub-optimal because it wastes host memory by reserving space for duplicated data. Furthermore, this “bounce buffer” approach wastes CPU cycles spent on transferring data. A new technique, NVIDIA GPUDirect Storage (GDS), can eliminate the need to use the host memory as a bounce buffer. Thereby, it becomes possible to transfer data directly between the device memory and the file system. This direct data path shortens latency by omitting the extra copy and enables higher-bandwidth. To take full advantage of GDS in existing applications, it is necessary to provide support with existing I/O libraries, such as HDF5 and MPI-IO, which are heavily used in applications. In this paper, we describe our effort of integrating GDS with HDF5, the top I/O library at NERSC and at DOE leadership computing facilities. We design and implement this integration using a HDF5 Virtual File Driver (VFD). The GDS VFD provides a file system abstraction to the application that allows HDF5 applications to perform I/O without the need to move data between CPUs and GPUs explicitly. We compare performance of the HDF5 GDS VFD with explicit data movement approaches and demonstrate superior performance with the GDS method.

Ravi, J↗

The Memory Scaling of Reverse-Mode Differentiation in Particle Accelerator Simulations with Space Charge

The recent development of differentiable simulation codes for particle accelerators has enabled gradient-based workflows that promise finer control and more realistic modeling of accelerator facilities. However, when using reverse-mode automatic differentiation, the memory usage continuously increases during the simulation, and can potentially exceed the available hardware memory - especially when costly space charge computation is included. To study the memory requirements for differentiable simulations, we have implemented space charge in Cheetah, a PyTorch-based beam tracking code that supports reverse-mode differentiation. We find that the memory usage for reverse-mode differentiation grows linearly with the number of macroparticles and cells, and that it is proportional to the number of space charge kicks involved in the simulation. This general scaling can be used to evaluate whether a given differentiable simulation is feasible given hardware memory constraints.

Dhamrait, Arjun↗

Berry curvature memory through electrically driven stacking transitions

In two-dimensional layered quantum materials, the interlayer stacking order determines both crystalline symmetry and quantum electronic properties such as Berry curvature, topology and electron correlation. Electrical stimuli can strongly influence quasi-particle interactions and the free energy landscape, thus making it possible to access hidden stacking orders with novel quantum properties and enabling dynamic engineering of these attributes. In this paper, we demonstrate electrically driven stacking transitions and a new type of nonvolatile memory based on Berry curvature in few-layer WTe 2 . The interplay of out-of-plane electric fields and electrostatic doping controls in-plane interlayer sliding and creates multiple polar and centrosymmetric stacking orders. In-situ nonlinear Hall transport reveals such stacking rearrangements result in a layer-parity-selective Berry curvature memory in momentum space, where the sign reversal of the Berry curvature only occurs in odd layer crystals. Our findings open an avenue towards exploring coupling between topology, electron correlations, and ferroelectricity in hidden stacking orders and demonstrate a new low energy cost, electrically-controlled topological memory in the atomically thin limit.

36 MATERIALS SCIENCE↗

Memory access optimization for particle operations in computational fluid dynamics-discrete element method simulations

Computational Fluid Dynamics - Discrete Element Method is used to model gas-solid systems in several applications in energy, pharmaceutical and petrochemical industries. Computational performance bottlenecks often limit the problem sizes that can be simulated at industrial scale. The data structures used to store several millions of particles in such large-scale simulations have a large memory footprint that does not fit into the processor cache hierarchies on current high-performance-computing platforms, leading to reduced computational performance. This paper specifically addresses this aspect of memory access bottlenecks in industrial scale simulations. The use of space-filling curves to improve memory access patterns is described and their impact on computational performance is quantified in both shared and distributed memory parallelization paradigms. The Morton space filling curve applied to uniform grids and k-dimensional tree partitions are used to reorder the particle data-structure thus improving spatial and temporal locality in memory. The performance impact of these techniques when applied to two benchmark problems, namely the homogeneous-cooling-system and a fluidized-bed, are presented. We report these optimization techniques lead to approximately two-fold performance improvement in particle focused operations such as neighbor-list creation and data-exchange, with ~ 1.5 times overall improvement in a fluidization simulation with 1.27 million particles.

97 MATHEMATICS AND COMPUTING↗

Ice sculpting: An artificial spin ice Tutorial on controlling microstate and geometry for magnonics and neuromorphic computing

Artificial spin ice, arrays of strongly interacting nanomagnets, are complex magnetic systems with many emergent properties, rich microstate spaces, intrinsic physical memory, high-frequency dynamics in the GHz range, and compatibility with a broad range of measurement approaches. This Tutorial article aims to provide the foundational knowledge needed to understand, design, develop, and improve the dynamic properties of artificial spin ice. Special emphasis is placed on introducing the theory of micromagnetics, which describes the complex dynamics within these systems, along with their design, fabrication methods, and standard measurement and control techniques. The article begins with a review of the historical background, introducing the underlying physical phenomena and interactions that govern artificial spin ice. We then explore the standard experimental techniques used to prepare the microstate space of the nanomagnetic array and to characterize magnetization dynamics, both in artificial spin ice and more broadly in ferromagnetic materials. Finally, we introduce the basics of neuromorphic computing applied to the case of artificial spin ice systems with a goal to help researchers new to the field grasp these exciting new developments.

Sultana, Rawnak [Univ. of Delaware, Newark, DE (Un↗

Design Space Exploration of Ferroelectric Tunnel Junction Toward Crossbar Memories

We perform a simulation-based analysis on the potential of emerging ferroelectric tunnel junctions (FTJs) as a memory device for crossbar arrays. Though FTJs are promising due to their low power switching characteristics compared to other emerging technologies, the greatest challenge for FTJs is the tradeoff between integration density and read performance. Our analysis highlights the need to co-optimize the ferroelectric thickness of the FTJ and read/write voltages to achieve proper functionality at large array sizes. Our analysis shows that FTJ-based crossbar achieves 93% higher sense margin at isoread power of 116 nW (per bit), but this FTJ design comes at a cost of 9.28× higher write power at isowrite time of 250 ns. In response, we study the potential tradeoffs of design points outside the feasible region to understand what device characteristics are desired to overcome such challenges.

Jao, Nicholas↗

Memory access statistics monitoring

Systems, apparatuses, and methods related to memory access statistics monitoring are described. A host is configured to map pages of memory for applications to a number of memory devices coupled thereto. A first memory device comprises a monitoring component configured to monitor access statistics of pages of memory mapped to the first memory device. A second memory device does not include a monitoring component capable of monitoring access statistics of pages of memory mapped thereto. The host is configured to map a portion of pages of memory for an application to the first memory device in order to obtain access statistics corresponding to the portion of pages of memory upon execution of the application despite there being space available on the second memory device and adjust mappings of the pages of memory for the application based on the obtained access statistics corresponding to the portion of pages.

Roberts, David A.↗

Evaluating the potential of disaggregated memory systems for HPC applications

Summary Disaggregated memory is a promising approach that addresses the limitations of traditional memory architectures by enabling memory to be decoupled from compute nodes and shared across a data center. Cloud platforms have deployed such systems to improve overall system memory utilization, but performance can vary across workloads. High‐performance computing (HPC) is crucial in scientific and engineering applications, where HPC machines also face the issue of underutilized memory. As a result, improving system memory utilization while understanding workload performance is essential for HPC operators. Therefore, learning the potential of a disaggregated memory system before deployment is a critical step. This paper proposes a methodology for exploring the design space of a disaggregated memory system. It incorporates key metrics that affect performance on disaggregated memory systems: memory capacity, local and remote memory access ratio, injection bandwidth, and bisection bandwidth, providing an intuitive approach to guide machine configurations based on technology trends and workload characteristics. We apply our methodology to analyze thirteen diverse workloads, including AI training, data analysis, genomics, protein, fusion, atomic nuclei, and traditional HPC bookends. Our methodology demonstrates the ability to comprehend the potential and pitfalls of a disaggregated memory system and provides motivation for machine configurations. Our results show that eleven of our thirteen applications can leverage injection bandwidth disaggregated memory without affecting performance, while one pays a rack bisection bandwidth penalty and two pay the system‐wide bisection bandwidth penalty. In addition, we also show that intra‐rack memory disaggregation would meet the application's memory requirement and provide enough remote memory bandwidth.

Ding, Nan↗

Efficient Treatment of Large Active Spaces through Multi-GPU Parallel Implementation of Direct Configuration Interaction

In this study, we have extended our graphical processing unit (GPU)-accelerated direct configuration interaction program to multiple devices, reducing iteration times for configuration spaces of 165 million determinants to only 3 s using NVIDIA P100 GPUs. Similar improvements in the one- and two-particle reduced density matrix formation allow for fast analytical energy gradients and electronic properties. Our parallel algorithm enables the calculation of arbitrarily large configuration spaces (limited only by available system memory), with iteration times of 13 min for an active space of 18 electrons in 18 orbitals (2.4 billion determinants) using six consumer grade NVIDIA 1080Ti GPUs. These advances enable routine molecular dynamics simulations, geometry optimizations, and absorption spectrum calculations for molecules with large configuration spaces, a task that has heretofore required massive computational effort. In this work, we demonstrate the utility of our program by generating the absorption spectrum for diphenyl acetylene at the floating occupation molecular orbital complete active space configuration interaction level of theory. Lastly, several active spaces were investigated to assess the dependence of spectral features on orbital space dimension.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

NESSi : The N on- E quilibrium S ystems S imulation package

The nonequilibrium dynamics of correlated many-particle systems is of interest in connection with pump–probe experiments on molecular systems and solids, as well as theoretical investigations of transport properties and relaxation processes. Nonequilibrium Green’s functions are a powerful tool to study interaction effects in quantum many-particle systems out of equilibrium, and to extract physically relevant information for the interpretation of experiments. Here, we present the open-source software package NESSi (The Non-Equilibrium Systems Simulation package) which allows to perform many-body dynamics simulations based on Green’s functions on the L-shaped Kadanoff–Baym contour. NESSi contains the library libcntr which implements tools for basic operations on these nonequilibrium Green’s functions, for constructing Feynman diagrams, and for the solution of integral and integro-differential equations involving contour Green’s functions. The library employs a discretization of the Kadanoff–Baym contour into time points and a high-order implementation of integration routines. The total integrated error scales up to $\mathcal{O}(N^{-7})$, which is important since the numerical effort increases at least cubically with the simulation time. A distributed-memory parallelization over reciprocal space allows large-scale simulations of lattice systems. We provide a collection of example programs ranging from dynamics in simple two-level systems to problems relevant in contemporary condensed matter physics, including Hubbard clusters and Hubbard or Holstein lattice models. The libcntr library is the basis of a follow-up software package for nonequilibrium dynamical mean-field theory calculations based on strong-coupling perturbative impurity solvers.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Collective neural network behavior in a dynamically driven disordered system of superconducting loops

Collective properties of complex systems composed of many interacting components such as neurons in our brain can be modeled by artificial networks based on disordered systems. We show that a disordered neural network of superconducting loops with Josephson junctions can exhibit computational properties like categorization and associative memory in the time evolution of its state in response to information from external excitations. Superconducting loops can trap multiples of fluxons in many discrete memory configurations defined by the local free energy minima in the configuration space of all possible states. A memory state can be updated by exciting the Josephson junctions to fire or allow the movement of fluxons through the network as the current through them surpasses their critical current thresholds. Simulations performed with a lumped element circuit model of a 4-loop network show that information written through excitations is translated into stable states of trapped flux and their time evolution. Experimental implementation on a high-Tc superconductor YBCO-based 4-loop network shows dynamically stable flux flow in each pathway characterized by the correlations between junction firing statistics. Neural network behavior is observed as energy barriers separating state categories in simulations in response to multiple excitations, and experimentally as junction responses characterizing different flux flow patterns in the network. The state categories that produce these patterns have different temporal stabilities relative to each other and the excitations. This provides strong evidence for time-dependent (short-to-long-term) memories, that are dependent on the geometrical and junction parameters of the loops, as described with a network model.

Josephson junctions↗

Profusion of symmetry-protected qubits from stable ergodicity breaking

We show how combining a discrete symmetry with topological Hilbert space fragmentation can give rise to exponentially many topologically stable qubits protected by a single discrete symmetry. We illustrate this explicitly with the example of the CZ𝑝 model, where the encoded qubits are prethermally stable to arbitrary symmetry-respecting perturbations for parametrically long times, substantially enhancing the robustness of a recently proposed construction based on nontopological fragmentation. In this model, the encoded qubits naturally come in pairs for which a universal set of transversal logical gates can be performed, ruling out (by the Eastin-Knill theorem) the possibility of using them for quantum error correction. We also comment on the combination of symmetry enrichment and topological fragmentation more generally, and the implications for use of systems exhibiting Hilbert space fragmentation as quantum memories.

kinetically constrained models↗

Distributed coherence directory subsystem with exclusive data regions

A processing system includes a first set of one or more processing units including a first processing unit, a second set of one or more processing units including a second processing unit, and a memory having an address space shared by the first and second sets. The processing system further includes a distributed coherence directory subsystem having a first coherence directory to support a first subset of one or more address regions of the address space and a second coherence directory to support a second subset of one or more address regions of the address space. In some implementations, the first coherence directory is implemented in the system so as to have a lower access latency for the first set, whereas the second coherence directory is implemented in the system so as to have a lower access latency for the second set.

97 MATHEMATICS AND COMPUTING↗

Distributed coherence directory subsystem with exclusive data regions

A processing system includes a first set of one or more processing units including a first processing unit, a second set of one or more processing units including a second processing unit, and a memory having an address space shared by the first and second sets. The processing system further includes a distributed coherence directory subsystem having a first coherence directory to support a first subset of one or more address regions of the address space and a second coherence directory to support a second subset of one or more address regions of the address space. In some implementations, the first coherence directory is implemented in the system so as to have a lower access latency for the first set, whereas the second coherence directory is implemented in the system so as to have a lower access latency for the second set.

Eckert, Yasuko↗

Utah FORGE: Well 56-32 Drilling Data and Logs

This dataset consists of drilling data (Pason data spreadsheets, daily reports, days v. depth, mud logs), Schlumberger logs (FMI, shear anisotropy analysis, memory, sonic, array induction/spectral density/dual spaced neutron/gamma ray/caliper, spectral GR/temperature, and Gardner density correlation), and an end of well report (EOWR) for Utah FORGE well 56-32. This is a vertical well that will be used for seismic monitoring. It was drilled between February 7th and February 21st 2021 to a depth of 9,145 feet. More information about this well can be found at: https://utahforge.com/2021/02/09/drilling-progress-of-well-56-32/ (linked below)

15 GEOTHERMAL ENERGY↗

Co-design of Advanced Architectures for Graph Analytics using Machine Learning

A graph is an excellent way of representing relationships among entities. We can use graph analytics to synthesize and analyze such relational data, and extract relevant features that are useful for various tasks such as machine learning. Considering the crucial role of graph analytics in various domains, it is important and timely to investigate the right hardware configurations that can achieve optimal performance for graph workloads on future high-performance computing systems. Design space exploration studies facilitate the selection of appropriate configurations (e.g. memory) to achieve a desired system performance. Recently, the approach of accelerating graph analytics using persistent non-volatile memory has gained a lot of attention. Traditional system simulators such as Gem5 and NVMain can be used to explore the design space of these advanced memory architectures for graph workloads. However, these simulators are slow in execution thus limiting the efficiency of design space exploration studies. To overcome this challenge, we proposed a machine learning based approach to co-design advanced memory architectures for graph workloads. We tested our approach with DRAM, non-volatile memory, and hybrid memory (DRAM+NVM) using a breadth first search benchmark algorithm. Our results showed the applicability of the proposed machine learning based approach to the co-design of the advanced memory architectures. In this paper, we provide recommendations on selecting advanced memory architectures to achieve desired performance for graph workloads. We also discuss the performances of different machine learning models that were considered in this study.

Kurte, Kuldeep↗