Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Memory systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

GASNet-EX Memory Kinds: Support for Device Memory in PGAS Programming Models

There is an emerging need for adaptive, lightweight communication in irregular HPC applications at exascale, where GPU accelerators provide the majority of available compute cycles. To address this need, Lawrence Berkeley National Lab is developing a programming system to support distributed-memory HPC application development using the Partitioned Global Address Space (PGAS) model. This work includes two major components: UPC++ and GASNet-EX. UPC++ is a C++ template library providing Remote Memory Access (RMA) and Remote Procedure Call (RPC) communication interfaces. GASNet-EX is a portable, high-performance communication middleware library, used by the implementations of UPC++ and many other PGAS programming models. We describe recent advances in GASNet-EX to efficiently implement zero-copy Remote Memory Access (RMA) communication to and from memory on accelerator devices such as GPUs. We demonstrate performance improvements via benchmark results from UPC++ (on Summit) and the Legion programming system (on DGX-1), both using GASNet-EX for communication.

Hargrove, Paul H↗

Synchronization between processes in a coordination namespace

A system and method of supporting point-to-point synchronization among processes/nodes implementing different hardware barriers in a tuple space/coordinated namespace (CNS) extended memory storage architecture. The system-wide CNS provides an efficient means for storing data, communications, and coordination within applications and workflows implementing barriers in a multi-tier, multi-nodal tree hierarchy. The system provides a hardware accelerated mechanism to support barriers between the participating processes. Also architected is a tree structure for a barrier processing method where processes are mapped to nodes of a tree, e.g., a tree of degree k to provide an efficient way of scaling the number of processes in a tuple space/coordination namespace.

Jacob, Philip↗

Synchronization between processes in a coordination namespace

A system and method of supporting point-to-point synchronization among processes/nodes implementing different hardware barriers in a tuple space/coordinated namespace (CNS) extended memory storage architecture. The system-wide CNS provides an efficient means for storing data, communications, and coordination within applications and workflows implementing barriers in a multi-tier, multi-nodal tree hierarchy. The system provides a hardware accelerated mechanism to support barriers between the participating processes. Also architected is a tree structure for a barrier processing method where processes are mapped to nodes of a tree, e.g., a tree of degree k, to provide an efficient way of scaling the number of processes in a tuple space/coordination namespace.

Jacob, Philip↗

Expandable implant and implant system

An embodiment of the invention includes an expandable implant to endovascularly embolize an anatomical void or malformation, such as an aneurysm. An embodiment is comprised of a chain or linked sequence of expandable polymer foam elements. Another embodiment includes an elongated length of expandable polymer foam coupled to a backbone. Another embodiment includes a system for endovascular delivery of an expandable implant (e.g., shape memory polymer) to embolize an aneurysm. The system may include a microcatheter, a lumen-reducing collar coupled to the distal tip of the microcatheter, a flexible pushing element detachably coupled to an expandable implant, and a flexible tubular sheath inside of which the compressed implant and pushing element are pre-loaded. Other embodiments are described herein.

Wilson, Thomas S.↗

Expandable implant and implant system

An embodiment of the invention includes an expandable implant to endovascularly embolize an anatomical void or malformation, such as an aneurysm. An embodiment is comprised of a chain or linked sequence of expandable polymer foam elements. Another embodiment includes an elongated length of expandable polymer foam coupled to a backbone. Another embodiment includes a system for endovascular delivery of an expandable implant (e.g., shape memory polymer) to embolize an aneurysm. The system may include a microcatheter, a lumen-reducing collar coupled to the distal tip of the microcatheter, a flexible pushing element detachably coupled to an expandable implant, and a flexible tubular sheath inside of which the compressed implant and pushing element are pre-loaded. Other embodiments are described herein.

Wilson, Thomas S.↗

Origins of magnetic memory and strong exchange bias bordering magnetic compensation in mixed-lanthanide systems

Here, the unexpected physical phenomena resulting from the seemingly inconsequential substitutions of chemically similar lanthanide elements in the Pr 1-x Gd x ScGe system are exploited to further the understanding of rare-earth magnetism and inform materials design. By directly probing magnetic moments of crystallographically indistinguishable Pr and Gd we solve the puzzles of how an unusual magnetic memory and strong exchange bias emerge at specific, easily predictable chemistries. Both effects are rooted in a robust antiparallel arrangement of large 4f magnetic moments of light and heavy lanthanides. This enables precise control of nearly zero net magnetization either opposed to, or aligned with, the external magnetic field that persists over a wide range of temperatures and fields. Further, spontaneous perturbations in the random distribution of lanthanide ions makes strong exchange bias possible in bulk single-phase compounds bordering magnetic compensation, consequently expanding the materials base beyond artificial magnetic multilayers and broadening the range of potential applications of the phenomenon.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Reconfigurable quantum phononic circuits via piezo-acoustomechanical interactions

Abstract We show that piezoelectric strain actuation of acoustomechanical interactions can produce large phase velocity changes in an existing quantum phononic platform: aluminum nitride on suspended silicon. Using finite element analysis, we demonstrate a piezo-acoustomechanical phase shifter waveguide capable of producing ± π phase shifts for GHz frequency phonons in 10s of μm with 10s of volts applied. Then, using the phase shifter as a building block, we demonstrate several phononic integrated circuit elements useful for quantum information processing. In particular, we show how to construct programmable multi-mode interferometers for linear phononic processing and a dynamically reconfigurable phononic memory that can switch between an ultra-long-lifetime state and a state strongly coupled to its bus waveguide. From the master equation for the full open quantum system of the reconfigurable phononic memory, we show that it is possible to perform read and write operations with over 90% quantum state transfer fidelity for an exponentially decaying pulse.

Taylor, Jeffrey C. (ORCID:0000000215982436)↗

Enabling Scalable and Extensible Memory-mapped Datastores in Userspace

Exascale workloads are expected to incorporate data-intensive processing in close coordination with traditional physics simulations. These emerging scientific, data-analytics and machine learning applications need to access a wide variety of datastores in flat files and structured databases. Programmer productivity is greatly enhanced by mapping datastores into the application process's virtual memory space to provide a unified “in-memory” interface. Currently, memory mapping is provided by system software primarily designed for generality and reliability. However, scalability at high concurrency is a formidable challenge on exascale systems. Also, there is a need for extensibility to support new datastores potentially requiring HPC data transfer services. In this article, we present UMap , a scalable and extensible userspace service for memory-mapping datastores. Furthermore, through decoupled queue management, concurrency aware adaptation, and dynamic load balancing, UMap enables application performance to scale even at high concurrency. We evaluate UMap in data-intensive applications, including sorting, graph traversal, database operations, and metagenomic analytics. Our results show that UMap as a userspace service outperforms an optimized kernel-based service across a wide range of intra-node concurrency by 1.22-1.9 × . We performed two case studies to demonstrate UMap 's extensibility. First, a new datastore residing in remote memory is incorporated into UMap as an application-specific plugin. Second, we present a persistent memory allocator Metall built atop UMap for unified storage/memory.

97 MATHEMATICS AND COMPUTING↗

Inductance meets memory in a quantum magnet

Orbital degrees of freedom offer a largely untapped route to emergent dynamical phenomena in correlated quantum materials. However, it remains unclear whether collective orbital states can intrinsically generate both reactive and memory functionalities in a bulk system. Here we show that in the ferrimagnet Mn₃Si₂Te₆, nonequilibrium reconfiguration of chiral orbital currents produces both emergent inductance and nonvolatile memristance as intrinsic properties of a single crystal. At low frequency and under a magnetic field along the c axis, coherent orbital-current domains generate robust clockwise inductive I-V loops. At higher frequency and low field, current-driven first-order reconfiguration leads to incomplete reversal and metastable trapping, producing an intrinsic electromotive force and a finite remanent voltage at zero current. These results establish orbital currents as a class of quantum state variables that encode both reactive and memory functionalities, opening routes toward intrinsically reconfigurable and energy-efficient electronic systems.

Cao, Tristan [University of Colorado, Boulder]↗

A stilbene–strontium iodide based radioxenon detection system for monitoring nuclear explosions

Atmospheric measurement of noble gases has been extensively used for monitoring clandestine nuclear weapon explosions for many years. The ratios of four xenon isotopes of interest ( 131 mXe, 133 mXe, 133 Xe, and 135 Xe) help in discriminating regular reactor operations from nuclear tests. A new coincidence-based detection system using stilbene and strontium iodide [SrI 2 (Eu)] for electron and photon detection respectively was developed at Oregon State University to address some of the challenges of the radioxenon systems deployed in the field such as memory effect, and poor energy resolution. Silicon photomultipliers (SiPMs) were used for sensing optical photons from all scintillation media. Real-time coincidence identification was achieved using the eight-channel digital pulse processor. The detection system was evaluated using lab check sources and Oregon State TRIGA reactor irradiated radioxenon samples. A 48-hour background coincidence spectrum was collected yielding a coincidence count rate and background rejection rate of 0.0174 ± 0.0003 counts per second (cps) and 98.9% respectively. The minimum detectable concentration (MDC) of the system was evaluated to be 0.11 ± 0.01, 0.13 ± 0.02, 0.20 ± 0.02, and 0.73 ± 0.08 for 131 mXe, 133 mXe, 133 Xe, and 135 Xe respectively. The memory effect of the detection system was found to be 0.069 ± 0.015%, which is almost a 70-fold reduction compared to traditional plastic scintillators. Here, the detection elements, custom-designed electronics, and the detector response to radioxenon are detailed in this work.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Profusion of symmetry-protected qubits from stable ergodicity breaking

We show how combining a discrete symmetry with topological Hilbert space fragmentation can give rise to exponentially many topologically stable qubits protected by a single discrete symmetry. We illustrate this explicitly with the example of the CZ𝑝 model, where the encoded qubits are prethermally stable to arbitrary symmetry-respecting perturbations for parametrically long times, substantially enhancing the robustness of a recently proposed construction based on nontopological fragmentation. In this model, the encoded qubits naturally come in pairs for which a universal set of transversal logical gates can be performed, ruling out (by the Eastin-Knill theorem) the possibility of using them for quantum error correction. We also comment on the combination of symmetry enrichment and topological fragmentation more generally, and the implications for use of systems exhibiting Hilbert space fragmentation as quantum memories.

kinetically constrained models↗

Design and Performance of Kokkos Staging Space toward Scalable Resilient Application Couplings

With the growing number of applications designed for heterogeneous HPC devices, application programmers and users are finding it challenging to compose scalable workflows as ensembles of these applications, that are portable, performant and resilient. The Kokkos C++ library has been designed to simplify this cumbersome procedure by providing an intra-application uniform programming model and portable performance. However, assembling multiple Kokkos-enabled applications into a complex workflow is still a challenge. Although Kokkos enables a uniform programming model, the inter-application data exchange still remains a challenge from both performance and software development cost perspectives. In order to address this issue, we propose Kokkos data staging memory space, an extension of Kokkos' data abstraction (memory space) for heterogeneous computing systems. This new abstraction allows to express data on a virtual shared-space for multiple Kokkos applications, thus extending Kokkos to support inter-application data exchange to build an efficient application workflow. Additionally, we study the effectiveness of asynchronous data layout conversions for applications requiring different memory access patterns for the shared data. Our preliminary evaluation with a synthetic benchmark indicate the effectiveness of this conversion adapted to three different scenarios representing access frequency and use patterns of the shared data.

97 MATHEMATICS AND COMPUTING↗

Active oscillatory associative memory

Traditionally, physical models of associative memory assume conditions of equilibrium. Here, we consider a prototypical oscillator model of associative memory and study how active noise sources that drive the system out of equilibrium, as well as nonlinearities in the interactions between the oscillators, affect the associative memory properties of the system. Our simulations show that pattern retrieval under active noise is more robust to the number of learned patterns and noise intensity than under passive noise. To understand this phenomenon, we analytically derive an effective energy correction due to the temporal correlations of active noise in the limit of short correlation decay time. We find that active noise deepens the energy wells corresponding to the patterns by strengthening the oscillator couplings, where the more nonlinear interactions are preferentially enhanced. Using replica theory, we demonstrate qualitative agreement between this effective picture and the retrieval simulations. Our work suggests that the nonlinearity in the oscillator couplings can improve memory under nonequilibrium conditions.

Chemistry↗

A two-level GPU-accelerated incomplete LU preconditioner for general sparse linear systems

This paper presents a parallel preconditioning approach based on incomplete LU (ILU) factorizations in the framework of Domain Decomposition (DD) for general sparse linear systems. We focus on distributed memory parallel architectures, specifically, those that are equipped with graphic processing units (GPUs). In addition to block-Jacobi, we present general purpose two-level ILU Schur complement-based approaches, where different strategies are presented to solve the coarse-level reduced system. These strategies are combined with modified ILU methods in the construction of the coarse-level operator, in order to effectively remove smooth errors by targeting an algebraically smooth vector. We leverage available GPU-based sparse matrix kernels to accelerate the setup and the solve phases of the proposed ILU preconditioner. We evaluate the efficiency of the proposed methods as a smoother for algebraic multigrid (AMG) and as a preconditioner for Krylov subspace methods on challenging anisotropic diffusion problems and a collection of general sparse matrices.

97 MATHEMATICS AND COMPUTING↗

UNITY: Unified Memory and Storage Space

UNITY is a 36-month project focused on providing design and evaluate a new distributed storage paradigm that unifies the traditionally distinct application views of memory- and file-based data storage into a single scalable and resilient environment. The project is a collaboration among Oak Ridge National Laboratory (ORNL, Lead institution), Los Alamos National Labs (LANL), and Georgia Tech (GT). The main contributions of the GT team have been around development of low level systems software for best leveraging the capabilities of new types of persistent memory technologies, and for development of methods for intelligent data management across memory/storage substrates with heterogeneous components. GT contributed the Phoenix library for optimized checkpoint/restart for HPC I/O for systems with non-volatile memory (NVM), the NVStream library for NVM-specialized streaming I/O for HPC workflows, the CoMerge, Mnemo and Kleio solutions for intelligent data management on NVM-based systems. These contributions result in significant improvements in both application performance and system efficiency.

97 MATHEMATICS AND COMPUTING↗

Nanoprobe Based Information Processing: Nanoprobe‐Electronics

Abstract With computational architectures becoming data‐centric and with the rise of in‐memory computing, the role of memory will be ever more crucial. In turn, storing massive amount of data demands for ultra‐low power and non‐volatile types necessitating new memory technologies. Here, a controllable nanoelectromechanical (NEM) spin memory is proposed for compact nonvolatile memory arrays. It combines the advantages of both nanomechanical and spin‐transfer torque (STT) magnetic memory devices to overcome fundamental scalability roadblocks of non‐hybrid alternatives. The hybrid system paves the way to a new memory technology in the regime of sub‐10‐nm lateral size with a sub‐1‐MA cm −2 switching energy.

Yi, Bao↗

Position Papers for the ASCR Workshop on the Management and Storage of Scientific Data

The purpose of this workshop is to identify priority research directions in the area of data management for high-performance and scientific computing above and beyond HPC’s traditional "the parallel file system is the data-management system" model. Supporting the breadth of the DOE mission, including the explosion of AI uses and the growing needs of experimental and observational science, motivates revisiting our assumptions about data management. There are many facets of this topic to explore including: (1) Interfaces for accessing data that resides on traditional persistent storage as well as memory devices; (2) Storage-system architecture design that supports scientific workflows on varied hierarchical storage and networking devices; (3) Devising metadata management infrastructure to support FAIR principles (Findability, Accessibility, Interoperability, and Reusability); (4) Capturing provenance information about scientific data; (5) Utilizing AI to learn I/O patterns of emerging workloads for efficient data management; (6) Providing data management support for AI and complex workflows; and (7) Understanding the overlap between traditional storage systems and I/O (SSIO) efforts and data management. While the program committee has identified these topics as important areas for discussion, we welcome position papers from the community that propose additional topics of interest for discussion at the workshop. The workshop agenda will include breakout sessions for discussing these and selected topic areas to inform priority research directions for data management for high-performance and scientific computing.

97 MATHEMATICS AND COMPUTING↗

Report for the ASCR Workshop on the Management and Storage of Scientific Data

The purpose of this workshop is to identify priority research directions in the area of data management for high-performance and scientific computing above and beyond HPC’s traditional "the parallel file system is the data-management system" model. Supporting the breadth of the DOE mission, including the explosion of AI uses and the growing needs of experimental and observational science, motivates revisiting our assumptions about data management. There are many facets of this topic to explore including: (1) Interfaces for accessing data that resides on traditional persistent storage as well as memory devices; (2) Storage-system architecture design that supports scientific workflows on varied hierarchical storage and networking devices; (3) Devising metadata management infrastructure to support FAIR principles (Findability, Accessibility, Interoperability, and Reusability); (4) Capturing provenance information about scientific data; (5) Utilizing AI to learn I/O patterns of emerging workloads for efficient data management; (6) Providing data management support for AI and complex workflows; and (7) Understanding the overlap between traditional storage systems and I/O (SSIO) efforts and data management. While the program committee has identified these topics as important areas for discussion, we welcome position papers from the community that propose additional topics of interest for discussion at the workshop. The workshop agenda will include breakout sessions for discussing these and selected topic areas to inform priority research directions for data management for high-performance and scientific computing.

97 MATHEMATICS AND COMPUTING↗