Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Memory systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

SpecSims: A Scalable Speculative Tree-based Simulation Cloning Framework for Finite Memory Machines

Simulation cloning is a technique in which cloned simulations whose state spaces differ partially from their parent simulation due to intervening events are spawned at runtime and concurrently advanced. It is a powerful method to carry out what-if analysis by speculatively exploring and evaluating the impact of various permutations of intervening cascade of events. Due to the exponential growth in the number of possible clones even for a small number of distinct intervening events, the practical efficacy of the approach is often severely limited by the maximum available memory of the computing host. In this paper, we introduce a novel speculative simulation cloning framework that executes a simulation cloning campaign capable of efficiently exploring an exponentially large space of clone simulations created by permutation of intervening events under a finite memory constraint. We provide a theoretical analysis of the runtime characteristics of our proposed approach and highlight its novel advantages such as memory-aware and as-long-as-needed execution. Furthermore, in support of our analytical findings and to demonstrate its practical feasibility, we implement a prototype of the cloning framework on a shared memory system and report its performance characteristics in the context of a heat diffusion simulation, and a power grid simulation subject to cascading disruptions from geomagnetic disturbances.

Simulation framework↗

R-Adaptivity to Enable Compression of Elementary Computations in Extreme-Scale Finite Element Simulators

Modern computing systems are capable of exascale calculations, which are revolutionizing the development and application of high-fidelity numerical models in computational science and engineering. While these systems continue to grow in processing power, the available system memory has not increased commensurately, and electrical power consumption continues to grow. A predominant approach to limit the memory usage in large-scale applications is to exploit the abundant processing power and continually recompute many low-level simulation quantities, rather than storing them. However, this approach can adversely impact the throughput of the simulation and diminish the benefits of modern computing architectures. We present three novel contributions to reduce the memory burden while maintaining, and sometimes improving, performance in simulations based on finite element discretizations. The first contribution develops dictionary-based data compression schemes that detect and exploit the structure of the discretization, due to redundancies across the finite element mesh. While these schemes are shown to reduce memory requirements by more than 99% on meshes with large numbers of identical mesh cells, there are applications where this structure does not exist. The second contribution leverages a recently developed augmented Lagrangian optimization algorithm to enable r-adaptivity for meshes with the goal of enhancing the redundancies in the mesh. The third contribution extends these methods to patch-based linear solvers and preconditioners by compressing local matrices. Numerical results demonstrate the effectiveness of the proposed methods to detect, enhance and exploit mesh structure on a suite of examples inspired by large-scale applications.

97 MATHEMATICS AND COMPUTING↗

Parallel Runtime Interface for Fortran (PRIF): A Multi-Image Solution for LLVM Flang

Fortran compilers that provide support for Fortran’s native parallel features often do so with a runtime library that depends on details of both the compiler implementation and the communication library, while others provide limited or no support at all. This paper introduces a new generalized interface that is both compiler- and runtime-library-agnostic, providing flexibility while fully supporting all of Fortran’s parallel features. The Parallel Runtime Interface for Fortran (PRIF) was developed to be portable across shared- and distributed-memory systems, with varying operating systems, toolchains and architectures. It achieves this by defining a set of Fortran procedures corresponding to each of the parallel features defined in the Fortran standard that may be invoked by a Fortran compiler and implemented by a runtime library. PRIF aims to be used as the solution for LLVM Flang to provide parallel Fortran support. This paper also briefly describes our PRIF prototype implementation: Caffeine.

Bonachea, Dan↗

Technical note: Optimizing the in situ cosmogenic 36 Cl extraction and measurement workflow for geologic applications

Abstract. In situ cosmogenic 36Cl analysis by accelerator mass spectrometry (AMS) is routinely employed to date Quaternary surfaces and assess rates of landscape evolution. However, standard laboratory preparation procedures for 36Cl dating require the addition of large amounts of isotopically enriched chlorine spike solution; these solutions are expensive and increasingly difficult to acquire from commercial sources. In addition, the typical workflow for 36Cl dating involves measuring both 35Cl/37Cl and 36Cl/Cl concurrently on the high-energy (post-accelerator) end of the AMS system, but 35Cl/37Cl determinations using this technique can be complicated by isotope fractionation and system memory during measurement. The traditional workflow also does not provide 36Cl extraction laboratories with the data needed to calculate native Cl concentrations in advance of 36Cl/Cl measurements. In light of these concerns, we present an improved workflow for extracting and measuring chlorine in geologic materials. Our initial step is to characterize 35Cl/37Cl on sample aliquots of up to ∼1 g prepared in Ag(Cl, Br) matrices, which greatly reduces the amount of isotopically enriched spike solution required to measure native Cl content in each sample. To avoid potential issues with isotope fractionation through the accelerator, 35Cl/37Cl is measured on the low-energy, pre-accelerator end of the AMS line. Then, for 36Cl/Cl measurements, we extract Cl as AgCl or Ag(Cl, Br) in analytical batches with a consistent total Cl load across all samples; this step is intended to minimize source memory effects during 36Cl/Cl measurements and allows the preparation of AMS standards that are customized to match known Cl contents in the samples. To assess the efficacy of this extraction and measurement workflow, we compare chlorine isotope ratio measurements on seven geologic samples prepared using standard procedures and the updated workflow. Measurements of 35Cl/37Cl and 36Cl/Cl are consistent between the two workflows, and 35Cl/37Cl values measured using our methods have considerably higher precision than those measured following standard protocols. The chemical preparation and measurement workflow presented here (1) reduces the amount of isotopically enriched chlorine spike used per rock sample by up to 95 %; (2) identifies rocks with high native Cl concentrations, which may be lower priority for 36Cl surface exposure dating, at an early stage of analysis; and (3) allows laboratory users to maintain control over the total chlorine content within and across analytical batches. These methods can be incorporated into existing laboratory and AMS protocols for 36Cl analyses and will increase the accessibility of 36Cl dating for geologic applications.

58 GEOSCIENCES↗

Method of efficiently identifying rollback requests

Disclosed in some examples are methods, systems, memory devices, and machine-readable mediums that allow a memory device to efficiently mark memory extents involved in an enhanced memory operation. An extent is marked if a meta state associated with the extent indicates that the extent is included in the enhanced memory operation. The largest memory extents of the operation are maintained in the memory device as a list of unmarked extents. When a primitive memory operation is received, the memory address is compared to the unmarked extents in the list to the meta state for that memory line. If the address is covered by the list of extents, or that line's meta state is marked, then the memory operation is performed including the enhanced memory operation.

Brewer, Tony M.↗

System and method for protecting GPU memory instructions against faults

A system and method for protecting memory instructions against faults are described. The system and method include converting the slave instructions to dummy operations, modifying memory arbiter to issue up to N master and N slave global/shared memory instructions per cycle, sending master memory requests to memory system, using slave requests for error checking, entering master requests to the GM/LM FIFO, storing slave requests in a register, and comparing the entered master requests with the stored slave requests.

Kalamatianos, John↗

A Synthesis Methodology for Intelligent Memory Interfaces in Accelerator Systems

Domain-specific systems improve the performance of a specific set of applications compared to general-purpose processing systems by deploying custom hardware accelerators. These hardware accelerators are generated using high-level synthesis (HLS) tools. The HLS tools enable a comprehensive design space exploration to optimize the compute performance of the generated accelerators. However, they often ignore the challenges of implementing the accelerators in a system-on-chip, particularly how the accelerators access memory. Our work introduces a buffering system design that improves accelerators' memory accesses by intelligently employing burst transactions to prefetch useful data from external memory to on-chip local buffers. Our design is dynamic, parametric, and transparent to the accelerators generated by HLS tools. We derive the buffering system parameters using appropriate compiler-based analysis passes and memory channel latency constraints. The proposed buffering system design results in, on average, 8.8x performance improvements while lowering memory channel utilization on average by 53.2% for a set of PolyBench kernels.

Limaye, Ankur M. (ORCID:0000000194062584)↗

Instructions for performing multi-line memory accesses

A system is described that performs memory access operations. The system includes a processor in a first node, a memory in a second node, a communication interconnect coupled to the processor and the memory, and an interconnect controller in the first node coupled between the processor and the communication interconnect. Upon executing a multi-line memory access instruction, the processor prepares a memory access operation for accessing, in the memory, a block of data including at least some of each of at least two lines of data. The processor then causes the interconnect controller to use a single remote direct memory access memory transfer to perform the memory access operation for the block of data via the communication interconnect.

Roberts, David A.↗

Memory object tagged memory monitoring method and system

Described are a method and processing apparatus to tag and track objects related to memory allocation calls. An application or software adds a tag to a memory allocation call to enable object level tracking. An entry is made into an object tracking table, which stores the tag and a variety of statistics related to the object and associated memory devices. The object statistics may be queried by the application to tune power/performance characteristics either by the application making runtime placement decisions, or by off-line code tuning based on a previous run. The application may add a tag to a memory allocation call to specify the type of memory characteristics requested based on the object statistics.

Roberts, David A.↗

Memory access response merging in a memory hierarchy

A system and method for efficiently processing memory requests are described. A computing system includes multiple compute units, multiple caches of a memory hierarchy and a communication fabric. A compute unit generates a memory access request that misses in a higher level cache, which sends a miss request to a lower level shared cache. During servicing of the miss request, the lower level cache merges identification information of multiple memory access requests targeting a same cache line from multiple compute units into a merged memory access response. The lower level shared cache continues to insert information into the merged memory access response until the lower level shared cache is ready to issue the merged memory access response. An intermediate router in the communication fabric broadcasts the merged memory access response into multiple memory access responses to send to corresponding compute units.

97 MATHEMATICS AND COMPUTING↗

Mapping entry invalidation

A memory access system may include a first memory address translator, a second memory address translator and a mapping entry invalidator. The first memory address translator translates a first virtual address in a first protocol of a memory access request to a second virtual address in a second protocol and tracks memory access request completions. The second memory address translator is to translate the second virtual address to a physical address of a memory. The mapping entry invalidator requests invalidation of a first mapping entry of the first mapping address translator requests invalidation of a second mapping entry of the second memory address translator corresponding to the first mapping entry following invalidation of the first mapping entry and based upon the tracked memory access request completions.

Walker, Shawn K.↗

Dynamic voltage and frequency scaling based on memory channel slack

A processing system scales power to memory and memory channels based on identifying causes of stalls of threads of a wavefront. If the cause is other than an outstanding memory request, the processing system throttles power to the memory to save power. If the stall is due to memory stalls for a subset of the memory channels servicing memory access requests for threads of a wavefront, the processing system adjusts power of the memory channels servicing memory access request for the wavefront based on the subset. By boosting power to the subset of channels, the processing system enables the wavefront to complete processing more quickly, resulting in increased processing speed. Conversely, by throttling power to the remainder of channels, the processing system saves power without affecting processing speed.

Das, Shomit N.↗

Shape memory embolectomy devices and systems

An embolectomy device comprised of an expansion unit and a support unit is disclosed. The expansion unit can be actuated in response to one or more external stimuli, and the support unit, located proximately to the expansion unit, provides a force to hold the expansion unit in place and to further induce the expansion unit's radial expansion. The radial expansion of the expansion unit causes the expansion unit to physically contact a blood clot, enabling the blood clot to be removed. In some embodiments, the expansion unit can be fabricated from a shape memory polymer foam. In some embodiments the support unit can be fabricated from any elastic material including, without limitation, shape memory alloys.

59 BASIC BIOLOGICAL SCIENCES↗

Dynamic Control of Sodium Cold Trap Purification Temperature Using LSTM System Identification

This study investigates the dynamic regulation of the sodium cold trap purification temperature at Argonne National Laboratory’s liquid sodium test facility, employing long short-term memory (LSTM) system identification techniques. The investigation introduces an innovative hybrid approach by integrating model predictive control (MPC) based on first principles dynamic models with a multi-step time–frequency LSTM model in predicting the temperature profiles of a sodium cold trap purification system. The long short-term memory–model predictive controller (LSTM-MPC) model employs a sliding window scheme to gather training samples for multi-step prediction, leveraging historical data to construct predictive models that capture the non-linearities of the complex system dynamics without explicitly modeling the underlying physical processes. The performance of the LSTM-MPC and MPC were evaluated through simulation experiments, where both models were assessed on their capacity to maintain the cold trap temperature within predefined set-points while minimizing deviations and overshoots. Results obtained show how the data-driven LSTM-MPC model demonstrates stability and adaptability. In contrast, the traditional MPC model exhibits irregularities, particularly evident as overshoots around set-point limits, which can potentially compromise its effectiveness over long prediction time intervals. The findings obtained offer valuable insights into integrating data-driven techniques for enhancing real-time monitoring systems.

LSTM-MPC↗

Memory instruction for memory tiers

Various embodiments provide for one or more processor instructions and memory instructions that enable a memory sub-system to copy, move, or swap data across (e.g., between) different memory tiers of the memory sub-system, where each of the memory tiers is associated with different memory locations (e.g., different physical memory locations) on one or more memory devices of the memory sub-system.

Roberts, David Andrew↗

UltraLiM: In-Memory Boolean Logic Architecture Using UltraRAM

Conventional computing architectures encounter ‘von Neumann’ and ‘memory wall’ bottlenecks which arise due to the back-and-forth data movement between the physically separate memory and processing units and the speed mismatch between them, respectively. These bottlenecks hurt both energy efficiency and the throughput of computing systems. To address these challenges, in-memory computing architectures have emerged as a promising alternative. They reduce the need for frequent data movement by executing different computing tasks inside the memory system. Here, we present UltraLiM, a logic-in-memory architecture using the UltraRAM-based memory system. UltraRAM holds the promise of developing a ‘universal memory’, overcoming the limitations of charge-based memories thanks to their non-volatile behavior with lower operating voltage. This work presents an in-memory computing architecture that integrates an UltraRAM-based memory array with a custom-designed peripheral circuitry. With this architecture, we can perform various in-memory Boolean logic operations (such as NOT, NAND, NOR, and XOR) in a single cycle. Leveraging the separate read-write paths in the UltraRAM-based memory array, we optimize read operations without encountering design conflicts. This optimization enhances the sense margin, enabling the use of simpler peripheral circuitry for in-memory logic operations.

Alam, Shamiul [University of Tennessee, Knoxville ↗

Register-Like Storage Block Used as Histograms, Cluster Buffers, and Hough Transform Accumulators for HEP Trigger Systems

In high energy physics experiment trigger systems, block memories are utilized for various purposes, especially in binned searching algorithms. In these algorithms, the storages are demanded to perform like a large set of registers. The writing and reading operation must be performed in single clock cycle and once an event is processed, the memory must be globally reset. These demands can be fulfilled with registers but the cost of using registers for large memory is unaffordable. Another common requirement is the boundary coverage feature during reading process. Additionally, when a memory bin is addressed, the stored contents in the addressed bin and its neighboring bin must be output simultaneously. In this paper, a register-like block storage design scheme is described, which allows updating memory locations in single clock cycle, reading two adjacent bins, and effectively refreshing entire memory within a single clock. The implementation and test results are presented.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗