Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “persistent memory”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Persistent Memory Object Storage and Indexing for Scientific Computing

This paper presents Mosiqs, a persistent memory object storage framework with metadata indexing and querying for scientific computing. We design Mosiqs based on the key idea that memory objects on shared PM pool can live beyond the application lifetime and can become the sharing currency for applications and scientists. Mosiqs provides an aggregate memory pool atop an array of persistent memory devices to store and access memory objects. Mosiqs uses a lightweight persistent memory key-value store to manage the metadata of memory objects such as persistent pointer mappings, which enables memory object sharing for effective scientific collaborations. Mosiqs is implemented atop PMDK. We evaluate the proposed approach on many-core server with an array of real PM devices. The preliminary evaluation confirms a 100% improvement for write and 30% in read performance against a PM-aware file system approach.

Khan, Awais↗

MOSIQS: Persistent Memory Object Storage With Metadata Indexing and Querying for Scientific Computing

Scientific applications often require high-bandwidth shared storage to perform joint simulations and collaborative data analytics. Shared memory pools provide a chance to satisfy such needs. Recently, a high-speed network such as Gen-Z utilizing persistent memory (PM) offers an opportunity to create a shared memory pool connected to compute nodes. However, there are several challenges to use scientific applications on the shared memory pool directly such as scalability, failure-atomicity, and lack of scientific metadata-based search and query. In this paper, we propose MOSIQS, a persistent memory object storage framework with metadata indexing and querying for scientific computing. We design MOSIQS based on the key idea that memory objects on PM pool can live beyond the application lifetime and can become the sharing currency for applications and scientists. MOSIQS provides an aggregate memory pool atop an array of persistent memory devices to store and access memory objects to accelerate scientific computing. MOSIQS uses a lightweight persistent memory key-value store to manage the metadata of memory objects, which enables memory object sharing. To facilitate metadata search and query over millions of memory objects resident on memory pool, we introduce Group Split and Merge (GSM), a novel persistent index data structure designed primarily for scientific datasets. GSM splits and merges dynamically to minimize the query search space and maintains low query processing time while overcoming the index storage overhead. MOSIQS is implemented on top of PMDK. We evaluate the proposed approach on many-core server with an array of real PM devices. Experimental results show that MOSIQS gains a 100% write performance improvement and executes multi-attribute queries efficiently with 2.7× less index storage overhead offering significant potential to speed up scientific computing applications.

97 MATHEMATICS AND COMPUTING↗

Metall: A persistent memory allocator for data-centric analytics

Data analytics applications transform raw input data into analytics-specific data structures before performing analytics. Unfortunately, such data ingestion steps are often more expensive than analytics. In addition, various types of NVRAM devices are already used in many HPC systems today. Such devices will be useful for storing and reusing data structures beyond a single process life cycle. We developed Metall, a persistent memory allocator built on top of the memory-mapped file mechanism. Metall enables applications to transparently allocate custom C++ data structures into various types of persistent memories. Metall incorporates a concise and high-performance memory management algorithm inspired by Supermalloc and the rich C++ interface developed by Boost.Interprocess library. On a dynamic graph construction workload, Metall achieved up to 11.7x and 48.3x performance improvements over Boost.Interprocess and memkind (PMEM kind), respectively. We also demonstrate Metall’s high adaptability by integrating Metall into a graph processing framework, GraphBLAS Template Library. Here this study’s outcomes indicate that Metall will be a strong tool for accelerating future large-scale data analytics by allowing applications to leverage persistent memory efficiently.

97 MATHEMATICS AND COMPUTING↗

Persistent memory as an effective alternative to random access memory in metagenome assembly

Abstract Background The assembly of metagenomes decomposes members of complex microbe communities and allows the characterization of these genomes without laborious cultivation or single-cell metagenomics. Metagenome assembly is a process that is memory intensive and time consuming. Multi-terabyte sequences can become too large to be assembled on a single computer node, and there is no reliable method to predict the memory requirement due to data-specific memory consumption pattern. Currently, out-of-memory (OOM) is one of the most prevalent factors that causes metagenome assembly failures. Results In this study, we explored the possibility of using Persistent Memory (PMem) as a less expensive substitute for dynamic random access memory (DRAM) to reduce OOM and increase the scalability of metagenome assemblers. We evaluated the execution time and memory usage of three popular metagenome assemblers (MetaSPAdes, MEGAHIT, and MetaHipMer2) in datasets up to one terabase. We found that PMem can enable metagenome assemblers on terabyte-sized datasets by partially or fully substituting DRAM. Depending on the configured DRAM/PMEM ratio, running metagenome assemblies with PMem can achieve a similar speed as DRAM, while in the worst case it showed a roughly two-fold slowdown. In addition, different assemblers displayed distinct memory/speed trade-offs in the same hardware/software environment. Conclusions We demonstrated that PMem is capable of expanding the capacity of DRAM to allow larger metagenome assembly with a potential tradeoff in speed. Because PMem can be used directly without any application-specific code modification, these findings are likely to be generalized to other memory-intensive bioinformatics applications.

59 BASIC BIOLOGICAL SCIENCES↗

FitCache: A Transparent Drop-In Framework for Multi-Tier Caching to Accelerate Distributed Deep Learning Workloads

Training in Deep learning (DL) remains highly compute- and data-intensive, with I/O becoming a critical bottleneck as models and datasets scale. Recent studies report that data loading can dominate training time, especially on large-scale HPC systems with shared parallel file systems (PFS). Existing caching approaches either rely on single-tier designs or require intrusive modifications to training pipelines, limiting their portability and effectiveness. In this work, we present FitCache, a transparent drop-in framework for multi-tier caching to accelerate distributed DL training by coordinating fast local memory (e.g., DRAM, Persistent Memory (PMem)) and NVMe as hierarchical caches atop PFS. Our design adapts to hardware diversity, i.e., if NVMe is missing, memory transparently acts as a caching tier, ensuring stable performance. FitCache transparently intercepts I/O requests and issues concurrent fetches across all tiers, returning data from the fastest responder without centralized metadata or static redirection paths. FitCache adapts to dynamic workloads and heterogeneous clusters while maintaining POSIX compatibility. Experiments on Frontier (2048 GPUs) and smaller research clusters show that FitCache reduces training time by up to 40% and per-batch I/O latency by up to 71.6% compared to Lustre Orion PFS, offering a drop-in solution for scalable DL training.

Hu, Guangxing [ORNL] (ORCID:0009000283203614)↗

Vaccination with mycobacterial lipid loaded nanoparticle leads to lipid antigen persistence and memory differentiation of antigen-specific T cells

Mycobacterium tuberculosis (Mtb) infection elicits both protein and lipid antigen-specific T cell responses. However, the incorporation of lipid antigens into subunit vaccine strategies and formulations has been underexplored, and the characteristics of vaccine-induced Mtb lipid-specific memory T cells have remained elusive. Mycolic acid (MA), a major lipid component of the Mtb cell wall, is presented by human CD1b molecules to unconventional T cell subsets. These MA-specific CD1b-restricted T cells have been detected in the blood and disease sites of Mtb-infected individuals, suggesting that MA is a promising lipid antigen for incorporation into multicomponent subunit vaccines. In this study, we utilized the enhanced stability of bicontinuous nanospheres (BCN) to efficiently encapsulate MA for in vivo delivery to MA-specific T cells, both alone and in combination with an immunodominant Mtb protein antigen (Ag85B). Pulmonary administration of MA-loaded BCN (MA-BCN) elicited MA-specific T cell responses in humanized CD1 transgenic mice. Simultaneous delivery of MA and Ag85B within BCN activated both MA- and Ag85B-specific T cells. Notably, pulmonary vaccination with MA-Ag85B-BCN resulted in the persistence of MA, but not Ag85B, within alveolar macrophages in the lung. Vaccination of MA-BCN through intravenous or subcutaneous route, or with attenuated Mtb likewise reproduced MA persistence. Moreover, MA-specific T cells in MA-BCN-vaccinated mice differentiated into a T follicular helper-like phenotype. Overall, the BCN platform allows for the dual encapsulation and in vivo activation of lipid and protein antigen-specific T cells and leads to persistent lipid depots that could offer long-lasting immune responses.

59 BASIC BIOLOGICAL SCIENCES↗

The persistence of memory in ionic conduction probed by nonlinear optics

Predicting practical rates of transport in condensed phases enables the rational design of materials, devices and processes. This is especially critical to developing low-carbon energy technologies such as rechargeable batteries. For ionic conduction, the collective mechanisms, variation of conductivity with timescales and confinement, and ambiguity in the phononic origin of translation, call for a direct probe of the fundamental steps of ionic diffusion: ion hops. However, such hops are rare-event large-amplitude translations, and are challenging to excite and detect. Here we use single-cycle terahertz pumps to impulsively trigger ionic hopping in battery solid electrolytes. This is visualized by an induced transient birefringence, enabling direct probing of anisotropy in ionic hopping on the picosecond timescale. The relaxation of the transient signal measures the decay of orientational memory, and the production of entropy in diffusion. We extend experimental results using in silico transient birefringence to identify vibrational attempt frequencies for ion hopping. Using nonlinear optical methods, we probe ion transport at its fastest limit, distinguish correlated conduction mechanisms from a true random walk at the atomic scale, and demonstrate the connection between activated transport and the thermodynamics of information.

25 ENERGY STORAGE↗

Single-node Partitioned-Memory for Huge Graph Analytics: Cost and Performance Trade-offs

Nonvolatile memory NVDIMMs, available as Intel Optane, are less expensive than DRAM and bring large byte-addressable storage within reach to many applications. Evaluations on graph analytics have shown promising performance only when DRAM is used as a hardware cache (Memory mode). An open question is whether graph applications can exploit Optane and DRAM directly (AppDirect mode) and achieving better-than-DRAM average bandwidth and run times. We evaluate Optane as a volatile pool on two large-scale graph applications with very different computational patterns, Grappolo and Ripples. We show that AppDirect mode can deliver better-than-DRAM performance, by allocating data structures to Optane and DRAM according to their access characteristics, resulting in higher average memory bandwidth and lower average latency. Memory mode provides DRAM-competitive performance with capacity equal to persistent memory. We demonstrate occasional 4x improvement using the latest AppDirect option and frequently observe competitive performance between Optane AppDirect Memory modes and DRAM.

Ghosh, Sayan↗

Enabling Scalable and Extensible Memory-mapped Datastores in Userspace

Exascale workloads are expected to incorporate data-intensive processing in close coordination with traditional physics simulations. These emerging scientific, data-analytics and machine learning applications need to access a wide variety of datastores in flat files and structured databases. Programmer productivity is greatly enhanced by mapping datastores into the application process's virtual memory space to provide a unified “in-memory” interface. Currently, memory mapping is provided by system software primarily designed for generality and reliability. However, scalability at high concurrency is a formidable challenge on exascale systems. Also, there is a need for extensibility to support new datastores potentially requiring HPC data transfer services. In this article, we present UMap , a scalable and extensible userspace service for memory-mapping datastores. Furthermore, through decoupled queue management, concurrency aware adaptation, and dynamic load balancing, UMap enables application performance to scale even at high concurrency. We evaluate UMap in data-intensive applications, including sorting, graph traversal, database operations, and metagenomic analytics. Our results show that UMap as a userspace service outperforms an optimized kernel-based service across a wide range of intra-node concurrency by 1.22-1.9 × . We performed two case studies to demonstrate UMap 's extensibility. First, a new datastore residing in remote memory is incorporated into UMap as an application-specific plugin. Second, we present a persistent memory allocator Metall built atop UMap for unified storage/memory.

97 MATHEMATICS AND COMPUTING↗

Automated Programmable Logic Controller Memory Forensics Using RGB Image Analysis and Deep Learning

The introduction of Industry 4.0 and Internet-based technologies has enhanced industrial control system operations but have inadvertently increased their vulnerabilities to cyber attacks. When an industrial control system is compromised, security analysts need to identify the root cause quickly to start the recovery process and develop mitigation strategies. Memory forensics is critical in the incident analysis process to ascertain what occurred. Approaches for analyzing the persistent memory in industrial control devices are limited and almost nonexistent for volatile memory. This chapter proposes an automated methodology for programmable logic controller memory dump analysis using computer vision and deep learning techniques. The methodology converts the sequences of bytes in a programmable logic controller memory dump to red-green-blue pixels and employs a deep learning model that learns the underlying patterns and features of pre-labeled forensic artifacts in images and segments them into distinct regions. The trained model is employed to automatically segment new memory images and identify forensic artifacts. Evaluation of the methodology on a Schneider Electric Modicon M221 programmable logic controller under code injection and code modification attacks demonstrates its ability to detect attack artifacts in memory dumps.

Asmar Awad, Rima [ORNL] (ORCID:0000000233407742)↗

HDF5 in the exascale era: Delivering efficient and scalable parallel I/O for exascale applications

Accurately modeling real-world systems requires scientific applications at exascale to generate massive amounts of data and manage data storage efficiently. However, parallel input and output (I/O) faces challenges due to new application workflows and the state-of-the-art memory, interconnect, and storage architectures considered in exascale designs. The storage hierarchy has expanded with node-local persistent memory, solid-state storage, and traditional disk and tape-based storage, thus requiring efficiency at each layer and much more efficient data movement among these layers. This paper discusses how the ExaHDF5 project improved the I/O performance and data management for exascale architectures by enhancing HDF5, a widely used parallel I/O library. The team developed an Asynchronous I/O Virtual Object Layer (VOL) connector that allowed overlapping I/O with computation. They also created a Cache VOL to complement asynchronous I/O by incorporating fast storage layers, such as burst buffer and node-local storage, into the parallel I/O workflow through caching and staging data. Additionally, the team enabled data aggregation and I/O at the node level by using a Subfiling Virtual File Driver (VFD). To demonstrate superior I/O performance with HDF5 at exascale, the ExaHDF5 team collaborated with several exascale applications. In this paper, we show I/O performance improvements for three applications: Cabana (a particle-based simulation library), EQSIM (a regional earthquake simulation software), and E3SM (a climate system modeling library).

Asynchronous I/Ol↗

Beyond Binary: Automated PLC Memory Forensics through RGB Image Analysis and Deep Learning

The introduction of Industry 4.0 and the evolution of industrial control systems (ICS) to adopt Internet-based technologies enhanced productivity, but have inadvertently increased their vulnerability to cyber-based malicious attacks. When an ICS system is compromised, security analysts need to identify the root cause quickly to start the recovery process and develop mitigation strategies to safeguard against future instances. Memory forensics is critical in the analysis process to ascertain what occurred. To date, approaches to analyze the persistent memory in ICS devices are limited, and almost nonexistent for volatile memory. This paper proposes an automated methodology, COMA, for PLC memory dump analysis using computer vision and deep learning techniques. Specifically, COMA converts the sequences of bytes in a PLC memory dump to RGB pixels and creates a deep learning model that learns the underlying patterns and features of pre-labeled forensic artifacts in images and segments them into distinct regions. COMA then uses the trained model to automatically segment new memory images and extract forensic artifacts. We evaluate COMA on a Schneider Electric Modicon M221 PLC involving two cyber-based attack scenarios: (i) code injection and (ii) code modification. The empirical results show that COMA can successfully detect attack artifacts in memory dumps in both scenarios.

Asmar Awad, Rima↗

Asynchronous aging and turnover of human circulating and tissue-resident memory T cells across sites

Memory T cells are maintained in tissues as circulating effector-memory (T EM ) and tissue-resident (T RM ) populations for protective immunity, though the role of site and subset in memory persistence remains undefined. Here, in this work, we investigated age-associated dynamics of human T cells in lymphoid organs, mucosal sites, and blood over 10 decades of life using retrospective radiocarbon ( 14 C) birth dating, along with cellular, transcriptome, and epigenetic profiling. Memory T cells across peripheral sites exhibited continuous turnover with mean lifespans of 1–2 years, while the spleen contained longer-lived T cells. Over age, T EM cells expressed senescent markers and a GZMK transcriptional signature, while T RM cells maintained site-specific resident phenotypes without exhibiting features of senescence. Both T EM and T RM cells showed age-associated DNA hypomethylation, though T RM cells exhibited more epigenetically regulated genes. Together, our findings reveal asynchronous aging of human memory T cells by subset and site, as well as persistence of T RM cells without immunosenescence.

T cells↗

A scale-wise analysis of intermittent momentum transport in dense canopy flows

We investigate the intermittent dynamics of momentum transport and its underlying time scales in the near-wall region of the neutrally stratified atmospheric boundary layer in the presence of a vegetation canopy. This is achieved through an empirical analysis of the persistence time scales (periods between successive zero-crossings) of momentum flux events, and their connection to the ejection–sweep cycle. Using high-frequency measurements from the GoAmazon campaign, spanning multiple heights within and above a dense canopy, the analysis suggests that, when the persistence time scales ( $t_p$ ) of momentum flux events from four different quadrants are separately normalized by $\varGamma _{w}$ (integral time scale of the vertical velocity), their distributions $P(t_p/\varGamma _{w})$ remain height-invariant. This result points to a persistent memory imposed by canopy-induced coherent structures, and to their role as an efficient momentum-transporting mechanism between the canopy airspace and the region immediately above. Moreover, $P(t_p/\varGamma _{w})$ exhibits a power-law scaling at times $t_{p}<\varGamma _{w}$ , with an exponential tail appearing for $t_{p} \geq \varGamma _{w}$ . By separating the flux events based on $t_p$ , we discover that around 80 % of the momentum is transported through the long-lived events ( $t_{p} \geq \varGamma _{w}$ ) at heights immediately above the canopy, while the short-lived ones ( $t_{p} < \varGamma _{w}$ ) only contribute marginally ( $\approx 20\,\%$ ). To explain the role of instantaneous flux amplitudes in momentum transport, we compare the measurements with newly developed surrogate data and establish that the range of time scales involved with amplitude variations in the fluxes tends to increase as one transitions from within to above the canopy.

Mechanics↗

Method and apparatus for providing thermal wear leveling

Exemplary embodiments provide thermal wear spreading among a plurality of thermal die regions in an integrated circuit or among dies by using die region wear-out data that represents a cumulative amount of time each of a number of thermal die regions in one or more dies has spent at a particular temperature level. In one example, die region wear-out data is stored in persistent memory and is accrued over a life of each respective thermal region so that a long term monitoring of temperature levels in the various die regions is used to spread thermal wear among the thermal die regions. In one example, spreading thermal wear is done by controlling task execution such as thread execution among one or more processing cores, dies and/or data access operations for a memory.

Roberts, David A.↗