Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “persistent memory”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Optimizing Management of Persistent Data Structures in High-Performance Analytics

Large-scale data analytics workflows ingest massive input data into various data structures, including graphs and key-value datastores. These data structures undergo multiple transformations and computations and are typically reused in incremental and iterative analytics workflows. Persisting in-memory views of these data structures enables reusing them beyond the scope of a single program run while avoiding repetitive raw data ingestion overheads. Memory-mapped I/O enables persisting in-memory data structures without data serialization and deserialization overheads. However, memory-mapped I/O lacks the key feature of persisting consistent snapshots of these data structures for incremental ingestion and processing. The obstacles to efficient virtual memory snapshots using memory-mapped I/O include background writebacks outside the application’s control, and the significantly high storage footprint of such snapshots. To address these limitations, we present Privateer, a memory and storage management tool that enables storage-efficient virtual memory snapshotting while also optimizing snapshot I/O performance. Here, we integrated Privateer into Metall, a state-of-the-art persistent memory allocator for C++, and the Lightning Memory-Mapped Database (LMDB), a widely-used key-value datastore in data analytics and machine learning. Privateer optimized application performance by 1.22× when storing data structure snapshots to node-local storage, and up to 16.7× when storing snapshots to a parallel file system. Privateer also optimizes storage efficiency of incremental data structure snapshots by up to 11× using data deduplication and compression.

Computer science↗

Mneme

A simple tool allowing recording the execution of a GPU (CUDA) kernel and replaying that kernel as an independent executable. The tool operates in 3 phases. During compile time the user needs to apply a provided LLVM pass to instrument the code. The pass detects all device global variables and device functions and stores this information with the respective LLVM-IR in the global device memory. The compilation generates a record-able executable. The second phase involves running the application executable with a desired input and using LD_PRELOAD to enable recording. When recording before invoking a device kernel the pre-loaded library stores device memory in persistent storage and associates the memory with the device kernel and an LLVM IR file. At the end of the recorded execution the pre-load library generates a database in the form of a JSON file containing information regarding the LLVM-IR files and the snapshots of device memory. During the third and last phase the user can replay the execution of an kernel as a separate independent executable. Besides executing it the user can modify the LLVM IR file and auto-tune parameters such as kernel launch-bounds or kernel runtime execution parameters (e.g. Kernel Block and Grid Dimensions). Is

Parasyris, Konstantinos↗

Global decline in ocean memory over the 21st century

Ocean memory, the persistence of ocean conditions, is a major source of predictability in the climate system beyond weather time scales. We show that ocean memory, as measured by the year-to-year persistence of sea surface temperature anomalies, is projected to steadily decline in the coming decades over much of the globe. This global decline in ocean memory is predominantly driven by shoaling of the upper-ocean mixed layer depth in response to global surface warming, while thermodynamic and dynamic feedbacks can contribute substantially regionally. As the mixed layer depth shoals, stochastic forcing becomes more effective in driving sea surface temperature anomalies, increasing high-frequency noise at the expense of persistent signals. Reduced ocean memory results in shorter lead times of skillful persistence-based predictions of sea surface thermal conditions, which may present previously unknown challenges for predicting climate extremes and managing marine biological resources under climate change.

54 ENVIRONMENTAL SCIENCES↗

The ECP SICM project: Managing complex memory hierarchies for exascale applications

The Exascale Computing Project (ECP)’s Simplified Interface to Complex Memories (SICM) effort focuses on developing universal interfaces for discovering, managing, and sharing data across complex memory hierarchies. These facilitate the exploitation of emerging memory technologies and support precise control over their various trade-offs such as high-bandwidth versus low-latency, persistent versus ephemeral, high-capacity versus low-capacity, and near-CPU versus near-GPU. SICM comprises three interrelated components: a low-level interface, a high-level interface, and a persistent-heap interface. The low-level SICM interface is intended for system and run-time developers as well as expert application developers who prefer full control of the memory objects used within their application. The high-level SICM interface builds upon the low-level interface, employing application-level profiling and analysis to optimize data management for complex memory hierarchies. The persistent-heap interface provides applications with a persistent memory allocator that can allocate custom C++ data structures in both block-storage and byte-addressable persistent memories.

97 MATHEMATICS AND COMPUTING↗

Hydrology Copilot: A Cloud-Native Ai System for Hydrological Data Analysis

The emergence of AI-driven Earth observation systems promises to broaden access to petabyte-scale geospatial data beyond domain specialists. However, translating this vision into operational scientific infrastructure requires addressing fundamental challenges in data virtualization, code transparency, and domain-specific reasoning. We present Hydrology Copilot, a cloud-native AI framework for natural-language-driven analysis of Earth observation data. To demonstrate operational capabilities at scale, we implement the system using NASA's North American Land Data Assimilation System version 3 (NLDAS-3), which provides surface meteorological forcing and land-surface model output across North and Central America at 1-km resolution, from which drought diagnostics are derived. The system integrates five core contributions: (1) scalable data virtualization using Kerchunk-based cloud optimized access, achieving a 1.5 to 4.6 times improvement in I/O latency across benchmark queries spanning regional single-day extractions (4.6 times speedup) to continental monthly aggregations (1.5 times speedup); (2) transparent code generation through Microsoft Azure AI Foundry agents that expose executable Python workflows for scientific verification; (3) persistent conversational memory enabling multi-turn analytical discourse across sessions; (4) intelligent query validation that enforces dataset boundaries and resolves ambiguous requests before execution; and (5) a multi-agent architecture coordinating query parsing, code generation, and visualization. We evaluate the system through drought-monitoring workflows, demonstrating reliable code generation, accurate results validated against reference computations and the operational U.S. Drought Monitor, and efficient operation across increasingly complex tasks. By bridging natural-language interfaces with rigorous hydrological analysis, Hydrology Copilot advances beyond proof-of-concept demonstrations to provide a deployable framework for operational Earth science applications.

Data virtualization↗

Two memories for geographical slant: separation and interdependence of action and awareness

The present study extended previous findings of geographical slant perception, in which verbal judgments of the incline of hills were greatly overestimated but motoric (haptic) adjustments were much more accurate. In judging slant from memory following a brief or extended time delay, subjects' verbal judgments were greater than those given when viewing hills. Motoric estimates differed depending on the length of the delay and place of response. With a short delay, motoric adjustments made in the proximity of the hill did not differ from those evoked during perception. When given a longer delay or when taken away from the hill, subjects' motoric responses increased along with the increase in verbal reports. These results suggest two different memorial influences on action. With a short delay at the hill, memory for visual guidance is separate from the explicit memory informing the conscious response. With short or long delays away from the hill, short-term visual guidance memory no longer persists, and both motor and verbal responses are driven by an explicit representation. These results support recent research involving visual guidance from memory, where actions become influenced by conscious awareness, and provide evidence for communication between the "what" and "how" visual processing systems.

NASA Discipline Space Human Factors↗

Alphanumeric character generator for oscilloscope

Compact portable alphanumeric display device can be used with any general-purpose externally-triggered oscilloscope without need for Z-axis modulation. Factors limiting size of display are: output line capacitance, read-only memory speed, and persistence of cathode-ray-tube.

Lockerson, D. C.↗

Conditions for Interference Versus Facilitation During Sequential Sensorimotor Adaptation

We investigated how sensorimotor adaptation acquired during one experimental session influenced the adaptation in a subsequent session. The subjects' task was to track a visual target using a joystick-controlled cursor, while the relationship between joystick and cursor position was manipulated to introduce a sensorimotor discordance. Each subject participated in two sessions, separated by a pause of 2 min to 1 month duration. We found that adaptation was achieved within minutes, and persisted in the memory for at least a month, with only a small decay (experiment A). When the discordances administered in the two sessions were in mutual conflict, we found evidence for task interference (experiment B). However, when the discordances were independent, we found facilitation rather than interference (experiment C); the latter finding could not be explained by the use of an "easier" discordance in the second session (experiment D). We conclude that interference is due to an incompatibility between task requirements, and not to a competition of tasks for short-term memory. We further conclude that the ability to adapt to a sensorimotor discordance.

Bock, Otmar↗

Predictability Associated with Land Surface Moisture States: Studies with the NSIPP System

Hydrologists have long speculated that soil moisture information can be used to increase skill in monthly to seasonal forecast systems. For this to be true, though, three conditions must be satisfied: (1) an imposed initial soil moisture anomaly in the forecast system must have some memory, so that it persists into the forecast period; (2) the modeled atmosphere must respond in a predictable way to the persisted anomaly; and (3) the forecast model must correctly represent both the soil moisture memory and the atmospheric response as they occur in nature. In this short paper, we review some recent work at NSIPP (NASA Seasonal-to-Interannual Prediction Project) that addresses all three conditions.

Koster, R. D.↗

Co-design of Advanced Architectures for Graph Analytics using Machine Learning

A graph is an excellent way of representing relationships among entities. We can use graph analytics to synthesize and analyze such relational data, and extract relevant features that are useful for various tasks such as machine learning. Considering the crucial role of graph analytics in various domains, it is important and timely to investigate the right hardware configurations that can achieve optimal performance for graph workloads on future high-performance computing systems. Design space exploration studies facilitate the selection of appropriate configurations (e.g. memory) to achieve a desired system performance. Recently, the approach of accelerating graph analytics using persistent non-volatile memory has gained a lot of attention. Traditional system simulators such as Gem5 and NVMain can be used to explore the design space of these advanced memory architectures for graph workloads. However, these simulators are slow in execution thus limiting the efficiency of design space exploration studies. To overcome this challenge, we proposed a machine learning based approach to co-design advanced memory architectures for graph workloads. We tested our approach with DRAM, non-volatile memory, and hybrid memory (DRAM+NVM) using a breadth first search benchmark algorithm. Our results showed the applicability of the proposed machine learning based approach to the co-design of the advanced memory architectures. In this paper, we provide recommendations on selecting advanced memory architectures to achieve desired performance for graph workloads. We also discuss the performances of different machine learning models that were considered in this study.

Kurte, Kuldeep↗

Oineus v1.0

A library for multi-threaded computation of persistence diagrams. The algorithm for computing persistence diagrams in a lock-free manner was published in the 'Towards Lock-free Persistent Homology' (D. Morozov, A. Nigmetov, Brief Announcement: SPAA 2020); it scales better than the only other shared-memory parallel implementation of persistent homology computation PHAT. Library includes python bindings to compute persistence diagrams and their vectorizations for lower-star filtrations on grid data. It is intended to be used by scientists working on Topological Data Analysis and its applications in different areas.

Nigmetov, Arnur↗

Ultrafast high-endurance memory based on sliding ferroelectrics

The persistence of voltage-switchable collective electronic phenomena down to the atomic scale has extensive implications for area- and energy-efficient electronics, especially in emerging nonvolatile memory technology. We investigate the performance of a ferroelectric field-effect transistor (FeFET) based on sliding ferroelectricity in bilayer boron nitride at room temperature. Sliding ferroelectricity represents a different form of atomically thin two-dimensional (2D) ferroelectrics, characterized by the switching of out-of-plane polarization through interlayer sliding motion. We examined the FeFET device employing monolayer graphene as the channel layer, which demonstrated ultrafast switching speeds on the nanosecond scale and high endurance exceeding 10 11 switching cycles, comparable to state-of-the-art FeFET devices. These characteristics highlight the potential of 2D sliding ferroelectrics for inspiring next-generation nonvolatile memory technology.

Science & Technology - Other Topics↗

Short-term memory and dual task performance

Two hypotheses concerning the way in which short-term memory interacts with another task in a dual task situation are considered. It is noted that when two tasks are combined, the activity of controlling and organizing performance on both tasks simultaneously may compete with either task for a resource; this resource may be space in a central mechanism or general processing capacity or it may be some task-specific resource. If a special relationship exists between short-term memory and control, especially if there is an identity relationship between short-term and a central controlling mechanism, then short-term memory performance should show a decrement in a dual task situation. Even if short-term memory does not have any particular identity with a controlling mechanism, but both tasks draw on some common resource or resources, then a tradeoff between the two tasks in allocating resources is possible and could be reflected in performance. The persistent concurrence cost in memory performance in these experiments suggests that short-term memory may have a unique status in the information processing system.

Regan, J. E.↗

PERKS: a Locality-Optimized Execution Model for Iterative Memory-bound GPU Applications

Iterative memory-bound solvers commonly occur in HPC codes. Typical GPU implementations have a loop on the host side that invokes the GPU kernel as much as time/algorithm steps there are. The termination of each kernel implicitly acts the barrier required after advancing the solution every time step. We propose an execution model for running memory-bound iterative GPU kernels: PERsistent KernelS (PERKS). In this model, the time loop is moved inside persistent kernel, and device-wide barriers are used for synchronization. We then reduce the traffic to device memory by caching subset of the output in each time step in the unused registers and shared memory. PERKS can be generalized to any iterative solver: they largely independent of the solver's implementation. We explain the design principle of PERKS and demonstrate effectiveness of PERKS for a wide range of iterative 2D/3D stencil benchmarks (geomean speedup of 2.12x for 2D stencils and 1.24x for 3D stencils over state-of-art libraries), and a Krylov subspace conjugate gradient solver (geomean speedup of 4.86x in smaller SpMV datasets from SuiteSparse and 1.43x in larger SpMV datasets over a state-of-art library). All PERKS-based implementations available at: https://github.com/neozhang307/PERKS.

Zhang, Lingqi↗

Representing the Sub-Grid Heterogeneity of Surface Precipitation in A General Circulation Model

Precipitation variability on spatial scales smaller than a typical general circulation model (GCM) grid box is often neglected, with the grid-mean precipitation rate being applied uniformly to underlying surface tiles. This reduces the extrema seen by the surface, with corresponding reductions in surface runoff and altered land-atmosphere fluxes. Here we present a novel approach to stochastically distribute precipitation across sub-grid surface tiles within a GCM. Based on 4 km Stage IV precipitation data, the scheme parameterizes the dry area fraction as a function of grid mean precipitation rate, and defines the relative distribution of intensities across non-dry surface tiles. To incorporate memory and mimic the persistence of precipitating storms, the relative intensity assigned to each sub-grid tile is determined by an autoregressive process. Using single column experiments, the scheme is shown to reproduce observed precipitation statistics at the scale of model surface tiles. We also document impacts on surface hydrology and energy partitioning, with notable increases in precipitation runoff, surface temperature variance, and the Bowen ratio.

GCM↗

Sptrace

Sptrace is a general-purpose space utilization tracing system that is conceptually similar to the commercial Purify product used to detect leaks and other memory usage errors. It is designed to monitor space utilization in any sort of heap, i.e., a region of data storage on some device (nominally memory; possibly shared and possibly persistent) with a flat address space. This software can trace usage of shared and/or non-volatile storage in addition to private RAM (random access memory). Sptrace is implemented as a set of C function calls that are invoked from within the software that is being examined. The function calls fall into two broad classes: (1) functions that are embedded within the heap management software [e.g., JPL's SDR (Simple Data Recorder) and PSM (Personal Space Management) systems] to enable heap usage analysis by populating a virtual time-sequenced log of usage activity, and (2) reporting functions that are embedded within the application program whose behavior is suspect. For ease of use, these functions may be wrapped privately inside public functions offered by the heap management software. Sptrace can be used for VxWorks or RTEMS realtime systems as easily as for Linux or OS/X systems.

Burleigh, Scott C.↗

Position Papers for the ASCR Workshop on the Management and Storage of Scientific Data

The purpose of this workshop is to identify priority research directions in the area of data management for high-performance and scientific computing above and beyond HPC’s traditional "the parallel file system is the data-management system" model. Supporting the breadth of the DOE mission, including the explosion of AI uses and the growing needs of experimental and observational science, motivates revisiting our assumptions about data management. There are many facets of this topic to explore including: (1) Interfaces for accessing data that resides on traditional persistent storage as well as memory devices; (2) Storage-system architecture design that supports scientific workflows on varied hierarchical storage and networking devices; (3) Devising metadata management infrastructure to support FAIR principles (Findability, Accessibility, Interoperability, and Reusability); (4) Capturing provenance information about scientific data; (5) Utilizing AI to learn I/O patterns of emerging workloads for efficient data management; (6) Providing data management support for AI and complex workflows; and (7) Understanding the overlap between traditional storage systems and I/O (SSIO) efforts and data management. While the program committee has identified these topics as important areas for discussion, we welcome position papers from the community that propose additional topics of interest for discussion at the workshop. The workshop agenda will include breakout sessions for discussing these and selected topic areas to inform priority research directions for data management for high-performance and scientific computing.

97 MATHEMATICS AND COMPUTING↗

Report for the ASCR Workshop on the Management and Storage of Scientific Data

The purpose of this workshop is to identify priority research directions in the area of data management for high-performance and scientific computing above and beyond HPC’s traditional "the parallel file system is the data-management system" model. Supporting the breadth of the DOE mission, including the explosion of AI uses and the growing needs of experimental and observational science, motivates revisiting our assumptions about data management. There are many facets of this topic to explore including: (1) Interfaces for accessing data that resides on traditional persistent storage as well as memory devices; (2) Storage-system architecture design that supports scientific workflows on varied hierarchical storage and networking devices; (3) Devising metadata management infrastructure to support FAIR principles (Findability, Accessibility, Interoperability, and Reusability); (4) Capturing provenance information about scientific data; (5) Utilizing AI to learn I/O patterns of emerging workloads for efficient data management; (6) Providing data management support for AI and complex workflows; and (7) Understanding the overlap between traditional storage systems and I/O (SSIO) efforts and data management. While the program committee has identified these topics as important areas for discussion, we welcome position papers from the community that propose additional topics of interest for discussion at the workshop. The workshop agenda will include breakout sessions for discussing these and selected topic areas to inform priority research directions for data management for high-performance and scientific computing.

97 MATHEMATICS AND COMPUTING↗