Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Memory systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Unveiling the nature of Ga-based chalcogenides for electrical switching selectors

Three-dimensional phase-change memory with stackable crossbar architecture is a promising technology to meet the urgent demands for high-density storage and rapid information processing in the era of explosive data growth. The performance depends strongly on the properties of ovonic threshold switching (OTS) selectors, which control the on/off states of memory units. Amorphous GaS serves as an outstanding OTS material, distinguished by its sizable mobility gap and high crystallization temperature, while the underlying mechanism continues to be inadequately comprehended. Here, in this work, we systematically studied the structural and electronic properties of amorphous Ga-X (X = S/Se/Te) using first-principles calculations. The results show that Ga atoms adopt tetrahedral motifs, while S/Se/Te atoms predominantly exhibit the structure of a distorted triangular pyramid. This structural arrangement is ascribed to the substantial dative bonds formed by the lone-pair electrons of the anions and the vacant sp3 orbitals around Ga atoms. Large mobility gaps (e.g., GaS: 2.43 eV, GaSe: 1.76 eV, GaTe: 1.26 eV) and distinct mid-gap states (e.g., ∼0.66 eV above valence band tail) ensure that these three chalcogenide glasses can be switched on under an external electric field while effectively suppressing leakage current without a bias, and the defect electronic states originate from short, robust Ga-Ga bonds due to the formation of distorted chain-like local structures. Our research elucidates the mechanisms of amorphous Ga-X as OTS materials, enriching the spectrum of electrical switching selectors by incorporating III-VI chalcogenides. This inclusion offers novel opportunities for the refinement and optimization of high-density integrated memory systems.

36 MATERIALS SCIENCE↗

Formal Definitions and Performance Comparison of Consistency Models for Parallel File Systems

The semantics of HPC storage systems are defined by the consistency models to which they abide. Storage consistency models have been less studied than their counterparts in memory systems, with the exception of the POSIX standard and its strict consistency model. The use of POSIX consistency imposes a performance penalty that becomes more significant as the scale of parallel file systems increases and the access time to storage devices, such as node-local solid storage devices, decreases. While some efforts have been made to adopt relaxed storage consistency models, these models are often defined informally and ambiguously as by-products of a particular implementation. Here in this work, we establish a connection between memory consistency models and storage consistency models and revisit the key design choices of storage consistency models from a high-level perspective. Further, we propose a formal and unified framework for defining storage consistency models and a layered implementation that can be used to easily evaluate their relative performance for different I/O workloads. Finally, we conduct a comprehensive performance comparison of two relaxed consistency models on a range of commonly seen parallel I/O workloads, such as checkpoint/restart of scientific applications and random reads of deep learning applications. We demonstrate that for certain I/O scenarios, a weaker consistency model can significantly improve the I/O performance. For instance, in small random reads that are typically found in deep learning applications, session consistency achieved a 5x improvement in I/O bandwidth compared to commit consistency, even at small scales.

97 MATHEMATICS AND COMPUTING↗

Kinetics of the xanthophyll cycle and its role in photoprotective memory and response

Efficiently balancing photochemistry and photoprotection is crucial for survival and productivity of photosynthetic organisms in the rapidly fluctuating light levels found in natural environments. The ability to respond quickly to sudden changes in light level is clearly advantageous. In the alga Nannochloropsis oceanica we observed an ability to respond rapidly to sudden increases in light level which occur soon after a previous high-light exposure. This ability implies a kind of memory. In this work, we explore the xanthophyll cycle in N. oceanica as a short-term photoprotective memory system. By combining snapshot fluorescence lifetime measurements with a biochemistry-based quantitative model, we show that short-term memory arises from the xanthophyll cycle. In addition, the model enables us to characterize the relative quenching abilities of the three xanthophyll cycle components. Given the ubiquity of the xanthophyll cycle in photosynthetic organisms the model described here will be of utility in improving our understanding of vascular plant and algal photoprotection with important implications for crop productivity.

59 BASIC BIOLOGICAL SCIENCES↗

Fixing Amdahl's Law within the Limits of Accelerated Systems: FALLACY

Closeout report for FALLACY project. The performance of Data Model Convergence Initiative (DMC) applications on parallel machines is far below the limit set by Amdahl’s law. Whether the machine is based on many-core, GPUs, FPGAs, or a heterogeneous combination, usually the most significant bottleneck is accessing data from the memory system. Aligning with DMC’s HW/architecture thrust, this project developed a set of memory-centric tools called ‘MemGaze’ that inform the HW/SW stack about an application’s memory behavior, including data access latency and diagnosing poor data layout and data composition. Our approach uses architectural modeling and analysis of workload data accesses.

97 MATHEMATICS AND COMPUTING↗

Bias control for a memory device

Methods, systems, and devices for bias control for a memory device are described. A memory system may store indication of whether data is coherent. In some examples, the indication may be stored as metadata, where a first value indicates that the data is not coherent and a second value or a third value indicate that the data is coherent. When a processing unit or other component of the memory system processes a command to access data, the memory system may operate according to a device bias mode when the indication is the first value, and according to a host bias mode when the indication is the second value or the third value.

97 MATHEMATICS AND COMPUTING↗

Bias control for a memory device

Methods, systems, and devices for bias control for a memory device are described. A memory system may store indication of whether data is coherent. In some examples, the indication may be stored as metadata, where a first value indicates that the data is not coherent and a second value or a third value indicate that the data is coherent. When a processing unit or other component of the memory system processes a command to access data, the memory system may operate according to a device bias mode when the indication is the first value, and according to a host bias mode when the indication is the second value or the third value.

Walker, Dean↗

Response of hypoxia to future climate change is sensitive to methodological assumptions

Climate-induced changes in hypoxia are among the most serious threats facing estuaries, which are among the most productive ecosystems on Earth. Future projections of estuarine hypoxia typically involve long-term multi-decadal continuous simulations or more computationally efficient time slice and delta methods that are restricted to short historical and future periods. We make a first comparison of these three methods by applying a linked terrestrial–estuarine model to the Chesapeake Bay, a large coastal-plain estuary in the eastern United States. Results show that the time slice approach accurately captures the behavior of the continuous approach, indicating a minimal impact of model memory. However, increases in mean annual hypoxic volume by the mid-twenty-first century simulated by the delta approach (+ 19%) are approximately twice as large as the time slice and continuous experiments (+ 9% and + 11%, respectively), indicating an important impact of changes in climate variability. Our findings suggest that system memory and projected changes in climate variability, as well as simulation length and natural variability of system hypoxia, should be considered when deciding to apply the more computationally efficient delta and time slice methods.

54 ENVIRONMENTAL SCIENCES↗

Embedding security into ferroelectric FET array via in situ memory operation

Non-volatile memories (NVMs) have the potential to reshape next-generation memory systems because of their promising properties of near-zero leakage power consumption, high density and non-volatility. However, NVMs also face critical security threats that exploit the non-volatile property. Compared to volatile memory, the capability of retaining data even after power down makes NVM more vulnerable. Existing solutions to address the security issues of NVMs are mainly based on Advanced Encryption Standard (AES), which incurs significant performance and power overhead. In this paper, we propose a lightweight memory encryption/decryption scheme by exploiting in-situ memory operations with negligible overhead. To validate the feasibility of the encryption/decryption scheme, device-level and array-level experiments are performed using ferroelectric field effect transistor (FeFET) as an example NVM without loss of generality. Besides, a comprehensive evaluation is performed on a 128 × 128 FeFET AND-type memory array in terms of area, latency, power and throughput. Compared with the AES-based scheme, our scheme shows ~22.6×/~14.1× increase in encryption/decryption throughput with negligible power penalty. Furthermore, we evaluate the performance of our scheme over the AES-based scheme when deploying different neural network workloads. Our scheme yields significant latency reduction by 90% on average for encryption and decryption processes.

47 OTHER INSTRUMENTATION↗

Early Performance Results on 4th Gen Intel(R) Xeon (R) Scalable Processors with DDR and Intel(R) Xeon(R) processors, codenamed Sapphire Rapids with HBM

The Crossroads supercomputer was designed to simulate some of the most complex physical devices in the world. These simulations routinely require 1/2 petabyte or more of system memory running on thousands of compute nodes for months at a time on the most powerful supercomputers. Improvements in time to solutions for these workloads can have major impact on our mission capabilities. In this paper we present early results of representative application workloads on 4th Gen Intel Xeon and Intel Xeon Processors codenamed Sapphire Rapids with HBM. These results demonstrate an extremely promising 8.57x improvement (node to node) over our prior generation Intel Broadwell (BDW) based HPC systems. No code modifications were required to achieve this speedup, providing a compelling path forward toward major reductions in time to solution and the complexity of physical systems that can be simulated in the future.

97 MATHEMATICS AND COMPUTING↗

Memory and rejuvenation in glassy systems

Here, the memory effect in a single crystal spin glass (Cu 0.92 Mn 0.08 ) has been measured using 1 Hz ac susceptibility measurements over a reduced temperature range of 0.4 - 0.7 T g and a model of the memory effect has been developed. A double-waiting-time protocol is carried out where the spin glass is first allowed to age at a temperature below T g , T w$_{1}$ , followed by a second aging 4 K lower, T w$_{2}$ . The 4 K separation is sufficient to ensure rejuvenation has occurred. The model is based on calculating overlaps between the growth of the correlation lengths at the two temperatures. It accounts for the absolute magnitude of the memory effect as a function of both waiting times and temperatures. The data can be explained by the memory loss being a function of the relative change in the correlated volume at the first waiting temperature due to growth in the correlations at the second waiting temperature.

36 MATERIALS SCIENCE↗

Computing material volume fractions on a superimposed mesh as applied to Monte Carlo particle transport simulations

Here, we present a newly implemented ray tracing algorithm in OpenMC for efficiently computing material volume fractions on superimposed meshes in complex geometries. By firing rays along each coordinate direction through the geometry, the approach accumulates track-length data in each mesh element, thereby determining the fractional composition of each material. Scaling studies on three different models—a random tetrahedra configuration, the Frascati Neutron Generator ITER dose rate benchmark, and a stellarator design—show excellent parallel performance, with nearly linear speedup on modern multi-threaded and distributed-memory systems. An analysis of the residual error relative to high-resolution reference solutions demonstrated that under optimal conditions it decreases as 1/R, where R is the number of rays fired, making it straightforward to achieve user-prescribed accuracy. This new functionality enables practical, mesh-based approaches for detailed nuclear analyses in production Monte Carlo workflows without resorting to expensive, fully conformal or unstructured meshing.

Monte Carlo↗

CG-Kit: Code Generation Toolkit for performant and maintainable variants of source code applied to Flash-X hydrodynamics simulations

CG-Kit is a new Code Generation tool-Kit that we have developed as a part of the solution for portability and maintainability for multiphysics computing applications. The development of CG-Kit is rooted in the urgent need created by the shifting landscape of high-performance computing platforms and the algorithmic complexities of a particular large-scale multiphysics application: Flash-X. To efficiently use computing resources on a heterogeneous node, an application must have a map of computation to resources and a mechanism to move the data and computation to the resources according to the map. Most existing performance portability solutions are focussed on abstracting the expression of computations so that a unified source code can be specialized to run on different resources. However, such an approach is insufficient for a code like Flash-X, which has a multitude of code components that can be assembled in various permutations and combinations to form different instances of applications. Similar challenges apply to any code that has composability, where a single specified way of apportioning work among devices may not be optimal. Additionally, use cases arise where the optimal control flow of computation may differ for different devices while the underlying numerics remain identical. This combination leads to unique challenges including handling an existing large code base in Fortran and/or C/C++, subdivision of code into a great variety of units supporting a wide range of physics and numerical methods, different parallelization techniques for distributed and shared memory systems and accelerator devices, and heterogeneity of computing platforms requiring coexisting variants of parallel algorithms. All of these challenges demand that scientific software developers apply existing knowledge about domain applications, algorithms, and computing platforms to determine custom abstractions and granularity for code generation. There is a critical lack of tools to tackle those problems. CG-Kit is designed to fill this gap by providing a user with the ability to express their desired control flow and computation-to-resource map in the form a pseudocode-like recipe. It consists of standalone tools that can be combined into highly specific and, we argue, highly effective portability and maintainability toolchains. Here we present the design of our new tools: parametrized source trees, control flow graphs, and recipes. The tools are implemented in Python. They are agnostic to the programming language of the source code targeted for code generation. In conclusion, we demonstrate the capabilities of the toolkit with two examples, first, multithreaded variants of the basic AXPY operation, and second, variants of parallel algorithms within a hydrodynamics solver, called Spark, from Flash-X that operates on block-structured adaptive meshes.

Algorithmic portability↗

Reinforcement Learning for Load-balanced Parallel Particle Tracing

We explore an online reinforcement learning (RL) paradigm to dynamically optimize parallel particle tracing performance in distributed-memory systems. Our method combines three novel components: (1) a work donation algorithm, (2) a high-order workload estimation model, and (3) a communication cost model. First, we design an RL-based work donation algorithm. Our algorithm monitors workloads of processes and creates RL agents to donate data blocks and particles from high-workload processes to low-workload processes to minimize program execution time. The agents learn the donation strategy on the fly based on reward and cost functions designed to consider processes' workload changes and data transfer costs of donation actions. Second, we propose a workload estimation model, helping RL agents estimate the workload distribution of processes in future computations. Third, we design a communication cost model that considers both block and particle data exchange costs, helping RL agents make effective decisions with minimized communication costs. We demonstrate that our algorithm adapts to different flow behaviors in large-scale fluid dynamics, ocean, and weather simulation data. Our algorithm improves parallel particle tracing performance in terms of parallel efficiency, load balance, and costs of I/O and communication for evaluations with up to 16,384 processors.

Distributed and parallel particle tracing↗

Synchronization for CXL Based Memory

Compute Express Link (CXL) is an important emerging standard for disaggregated memory. While this standard provisions coherency across numerous hosts and devices, implementing hardware support for type three devices is challenging. In this work, we look at the overhead of software synchronization and using software-based coherency. Moreover, we discuss the limits of software-based coherency in fully expressing modern synchronization techniques for a CXL-based disaggregate memory system. We demonstrate our approach using a CXL hardware prototype and running a version of the famous Peterson Lock (enhanced to run with more than two threads). We analyze its performance and share how more advanced synchronization techniques might interact with software-based coherence CXL hardware and program execution models.

High Performance Computing (HPC)↗

SpecSims: A Scalable Speculative Tree-based Simulation Cloning Framework for Finite Memory Machines

Simulation cloning is a technique in which cloned simulations whose state spaces differ partially from their parent simulation due to intervening events are spawned at runtime and concurrently advanced. It is a powerful method to carry out what-if analysis by speculatively exploring and evaluating the impact of various permutations of intervening cascade of events. Due to the exponential growth in the number of possible clones even for a small number of distinct intervening events, the practical efficacy of the approach is often severely limited by the maximum available memory of the computing host. In this paper, we introduce a novel speculative simulation cloning framework that executes a simulation cloning campaign capable of efficiently exploring an exponentially large space of clone simulations created by permutation of intervening events under a finite memory constraint. We provide a theoretical analysis of the runtime characteristics of our proposed approach and highlight its novel advantages such as memory-aware and as-long-as-needed execution. Furthermore, in support of our analytical findings and to demonstrate its practical feasibility, we implement a prototype of the cloning framework on a shared memory system and report its performance characteristics in the context of a heat diffusion simulation, and a power grid simulation subject to cascading disruptions from geomagnetic disturbances.

Simulation framework↗

R-Adaptivity to Enable Compression of Elementary Computations in Extreme-Scale Finite Element Simulators

Modern computing systems are capable of exascale calculations, which are revolutionizing the development and application of high-fidelity numerical models in computational science and engineering. While these systems continue to grow in processing power, the available system memory has not increased commensurately, and electrical power consumption continues to grow. A predominant approach to limit the memory usage in large-scale applications is to exploit the abundant processing power and continually recompute many low-level simulation quantities, rather than storing them. However, this approach can adversely impact the throughput of the simulation and diminish the benefits of modern computing architectures. We present three novel contributions to reduce the memory burden while maintaining, and sometimes improving, performance in simulations based on finite element discretizations. The first contribution develops dictionary-based data compression schemes that detect and exploit the structure of the discretization, due to redundancies across the finite element mesh. While these schemes are shown to reduce memory requirements by more than 99% on meshes with large numbers of identical mesh cells, there are applications where this structure does not exist. The second contribution leverages a recently developed augmented Lagrangian optimization algorithm to enable r-adaptivity for meshes with the goal of enhancing the redundancies in the mesh. The third contribution extends these methods to patch-based linear solvers and preconditioners by compressing local matrices. Numerical results demonstrate the effectiveness of the proposed methods to detect, enhance and exploit mesh structure on a suite of examples inspired by large-scale applications.

97 MATHEMATICS AND COMPUTING↗

Parallel Runtime Interface for Fortran (PRIF): A Multi-Image Solution for LLVM Flang

Fortran compilers that provide support for Fortran’s native parallel features often do so with a runtime library that depends on details of both the compiler implementation and the communication library, while others provide limited or no support at all. This paper introduces a new generalized interface that is both compiler- and runtime-library-agnostic, providing flexibility while fully supporting all of Fortran’s parallel features. The Parallel Runtime Interface for Fortran (PRIF) was developed to be portable across shared- and distributed-memory systems, with varying operating systems, toolchains and architectures. It achieves this by defining a set of Fortran procedures corresponding to each of the parallel features defined in the Fortran standard that may be invoked by a Fortran compiler and implemented by a runtime library. PRIF aims to be used as the solution for LLVM Flang to provide parallel Fortran support. This paper also briefly describes our PRIF prototype implementation: Caffeine.

Bonachea, Dan↗

Technical note: Optimizing the in situ cosmogenic 36 Cl extraction and measurement workflow for geologic applications

Abstract. In situ cosmogenic 36Cl analysis by accelerator mass spectrometry (AMS) is routinely employed to date Quaternary surfaces and assess rates of landscape evolution. However, standard laboratory preparation procedures for 36Cl dating require the addition of large amounts of isotopically enriched chlorine spike solution; these solutions are expensive and increasingly difficult to acquire from commercial sources. In addition, the typical workflow for 36Cl dating involves measuring both 35Cl/37Cl and 36Cl/Cl concurrently on the high-energy (post-accelerator) end of the AMS system, but 35Cl/37Cl determinations using this technique can be complicated by isotope fractionation and system memory during measurement. The traditional workflow also does not provide 36Cl extraction laboratories with the data needed to calculate native Cl concentrations in advance of 36Cl/Cl measurements. In light of these concerns, we present an improved workflow for extracting and measuring chlorine in geologic materials. Our initial step is to characterize 35Cl/37Cl on sample aliquots of up to ∼1 g prepared in Ag(Cl, Br) matrices, which greatly reduces the amount of isotopically enriched spike solution required to measure native Cl content in each sample. To avoid potential issues with isotope fractionation through the accelerator, 35Cl/37Cl is measured on the low-energy, pre-accelerator end of the AMS line. Then, for 36Cl/Cl measurements, we extract Cl as AgCl or Ag(Cl, Br) in analytical batches with a consistent total Cl load across all samples; this step is intended to minimize source memory effects during 36Cl/Cl measurements and allows the preparation of AMS standards that are customized to match known Cl contents in the samples. To assess the efficacy of this extraction and measurement workflow, we compare chlorine isotope ratio measurements on seven geologic samples prepared using standard procedures and the updated workflow. Measurements of 35Cl/37Cl and 36Cl/Cl are consistent between the two workflows, and 35Cl/37Cl values measured using our methods have considerably higher precision than those measured following standard protocols. The chemical preparation and measurement workflow presented here (1) reduces the amount of isotopically enriched chlorine spike used per rock sample by up to 95 %; (2) identifies rocks with high native Cl concentrations, which may be lower priority for 36Cl surface exposure dating, at an early stage of analysis; and (3) allows laboratory users to maintain control over the total chlorine content within and across analytical batches. These methods can be incorporated into existing laboratory and AMS protocols for 36Cl analyses and will increase the accessibility of 36Cl dating for geologic applications.

58 GEOSCIENCES↗