Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “persistent memory”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Method and apparatus for providing thermal wear leveling

Exemplary embodiments provide thermal wear spreading among a plurality of thermal die regions in an integrated circuit or among dies by using die region wear-out data that represents a cumulative amount of time each of a number of thermal die regions in one or more dies has spent at a particular temperature level. In one example, die region wear-out data is stored in persistent memory and is accrued over a life of each respective thermal region so that a long term monitoring of temperature levels in the various die regions is used to spread thermal wear among the thermal die regions. In one example, spreading thermal wear is done by controlling task execution such as thread execution among one or more processing cores, dies and/or data access operations for a memory.

Roberts, David A.↗

Optimizing Management of Persistent Data Structures in High-Performance Analytics

Large-scale data analytics workflows ingest massive input data into various data structures, including graphs and key-value datastores. These data structures undergo multiple transformations and computations and are typically reused in incremental and iterative analytics workflows. Persisting in-memory views of these data structures enables reusing them beyond the scope of a single program run while avoiding repetitive raw data ingestion overheads. Memory-mapped I/O enables persisting in-memory data structures without data serialization and deserialization overheads. However, memory-mapped I/O lacks the key feature of persisting consistent snapshots of these data structures for incremental ingestion and processing. The obstacles to efficient virtual memory snapshots using memory-mapped I/O include background writebacks outside the application’s control, and the significantly high storage footprint of such snapshots. To address these limitations, we present Privateer, a memory and storage management tool that enables storage-efficient virtual memory snapshotting while also optimizing snapshot I/O performance. Here, we integrated Privateer into Metall, a state-of-the-art persistent memory allocator for C++, and the Lightning Memory-Mapped Database (LMDB), a widely-used key-value datastore in data analytics and machine learning. Privateer optimized application performance by 1.22× when storing data structure snapshots to node-local storage, and up to 16.7× when storing snapshots to a parallel file system. Privateer also optimizes storage efficiency of incremental data structure snapshots by up to 11× using data deduplication and compression.

Computer science↗

Mneme

A simple tool allowing recording the execution of a GPU (CUDA) kernel and replaying that kernel as an independent executable. The tool operates in 3 phases. During compile time the user needs to apply a provided LLVM pass to instrument the code. The pass detects all device global variables and device functions and stores this information with the respective LLVM-IR in the global device memory. The compilation generates a record-able executable. The second phase involves running the application executable with a desired input and using LD_PRELOAD to enable recording. When recording before invoking a device kernel the pre-loaded library stores device memory in persistent storage and associates the memory with the device kernel and an LLVM IR file. At the end of the recorded execution the pre-load library generates a database in the form of a JSON file containing information regarding the LLVM-IR files and the snapshots of device memory. During the third and last phase the user can replay the execution of an kernel as a separate independent executable. Besides executing it the user can modify the LLVM IR file and auto-tune parameters such as kernel launch-bounds or kernel runtime execution parameters (e.g. Kernel Block and Grid Dimensions). Is

Parasyris, Konstantinos↗

Global decline in ocean memory over the 21st century

Ocean memory, the persistence of ocean conditions, is a major source of predictability in the climate system beyond weather time scales. We show that ocean memory, as measured by the year-to-year persistence of sea surface temperature anomalies, is projected to steadily decline in the coming decades over much of the globe. This global decline in ocean memory is predominantly driven by shoaling of the upper-ocean mixed layer depth in response to global surface warming, while thermodynamic and dynamic feedbacks can contribute substantially regionally. As the mixed layer depth shoals, stochastic forcing becomes more effective in driving sea surface temperature anomalies, increasing high-frequency noise at the expense of persistent signals. Reduced ocean memory results in shorter lead times of skillful persistence-based predictions of sea surface thermal conditions, which may present previously unknown challenges for predicting climate extremes and managing marine biological resources under climate change.

54 ENVIRONMENTAL SCIENCES↗

The ECP SICM project: Managing complex memory hierarchies for exascale applications

The Exascale Computing Project (ECP)’s Simplified Interface to Complex Memories (SICM) effort focuses on developing universal interfaces for discovering, managing, and sharing data across complex memory hierarchies. These facilitate the exploitation of emerging memory technologies and support precise control over their various trade-offs such as high-bandwidth versus low-latency, persistent versus ephemeral, high-capacity versus low-capacity, and near-CPU versus near-GPU. SICM comprises three interrelated components: a low-level interface, a high-level interface, and a persistent-heap interface. The low-level SICM interface is intended for system and run-time developers as well as expert application developers who prefer full control of the memory objects used within their application. The high-level SICM interface builds upon the low-level interface, employing application-level profiling and analysis to optimize data management for complex memory hierarchies. The persistent-heap interface provides applications with a persistent memory allocator that can allocate custom C++ data structures in both block-storage and byte-addressable persistent memories.

97 MATHEMATICS AND COMPUTING↗

Co-design of Advanced Architectures for Graph Analytics using Machine Learning

A graph is an excellent way of representing relationships among entities. We can use graph analytics to synthesize and analyze such relational data, and extract relevant features that are useful for various tasks such as machine learning. Considering the crucial role of graph analytics in various domains, it is important and timely to investigate the right hardware configurations that can achieve optimal performance for graph workloads on future high-performance computing systems. Design space exploration studies facilitate the selection of appropriate configurations (e.g. memory) to achieve a desired system performance. Recently, the approach of accelerating graph analytics using persistent non-volatile memory has gained a lot of attention. Traditional system simulators such as Gem5 and NVMain can be used to explore the design space of these advanced memory architectures for graph workloads. However, these simulators are slow in execution thus limiting the efficiency of design space exploration studies. To overcome this challenge, we proposed a machine learning based approach to co-design advanced memory architectures for graph workloads. We tested our approach with DRAM, non-volatile memory, and hybrid memory (DRAM+NVM) using a breadth first search benchmark algorithm. Our results showed the applicability of the proposed machine learning based approach to the co-design of the advanced memory architectures. In this paper, we provide recommendations on selecting advanced memory architectures to achieve desired performance for graph workloads. We also discuss the performances of different machine learning models that were considered in this study.

Kurte, Kuldeep↗

Oineus v1.0

A library for multi-threaded computation of persistence diagrams. The algorithm for computing persistence diagrams in a lock-free manner was published in the 'Towards Lock-free Persistent Homology' (D. Morozov, A. Nigmetov, Brief Announcement: SPAA 2020); it scales better than the only other shared-memory parallel implementation of persistent homology computation PHAT. Library includes python bindings to compute persistence diagrams and their vectorizations for lower-star filtrations on grid data. It is intended to be used by scientists working on Topological Data Analysis and its applications in different areas.

Nigmetov, Arnur↗

Ultrafast high-endurance memory based on sliding ferroelectrics

The persistence of voltage-switchable collective electronic phenomena down to the atomic scale has extensive implications for area- and energy-efficient electronics, especially in emerging nonvolatile memory technology. We investigate the performance of a ferroelectric field-effect transistor (FeFET) based on sliding ferroelectricity in bilayer boron nitride at room temperature. Sliding ferroelectricity represents a different form of atomically thin two-dimensional (2D) ferroelectrics, characterized by the switching of out-of-plane polarization through interlayer sliding motion. We examined the FeFET device employing monolayer graphene as the channel layer, which demonstrated ultrafast switching speeds on the nanosecond scale and high endurance exceeding 10 11 switching cycles, comparable to state-of-the-art FeFET devices. These characteristics highlight the potential of 2D sliding ferroelectrics for inspiring next-generation nonvolatile memory technology.

Science & Technology - Other Topics↗

PERKS: a Locality-Optimized Execution Model for Iterative Memory-bound GPU Applications

Iterative memory-bound solvers commonly occur in HPC codes. Typical GPU implementations have a loop on the host side that invokes the GPU kernel as much as time/algorithm steps there are. The termination of each kernel implicitly acts the barrier required after advancing the solution every time step. We propose an execution model for running memory-bound iterative GPU kernels: PERsistent KernelS (PERKS). In this model, the time loop is moved inside persistent kernel, and device-wide barriers are used for synchronization. We then reduce the traffic to device memory by caching subset of the output in each time step in the unused registers and shared memory. PERKS can be generalized to any iterative solver: they largely independent of the solver's implementation. We explain the design principle of PERKS and demonstrate effectiveness of PERKS for a wide range of iterative 2D/3D stencil benchmarks (geomean speedup of 2.12x for 2D stencils and 1.24x for 3D stencils over state-of-art libraries), and a Krylov subspace conjugate gradient solver (geomean speedup of 4.86x in smaller SpMV datasets from SuiteSparse and 1.43x in larger SpMV datasets over a state-of-art library). All PERKS-based implementations available at: https://github.com/neozhang307/PERKS.

Zhang, Lingqi↗

Position Papers for the ASCR Workshop on the Management and Storage of Scientific Data

The purpose of this workshop is to identify priority research directions in the area of data management for high-performance and scientific computing above and beyond HPC’s traditional "the parallel file system is the data-management system" model. Supporting the breadth of the DOE mission, including the explosion of AI uses and the growing needs of experimental and observational science, motivates revisiting our assumptions about data management. There are many facets of this topic to explore including: (1) Interfaces for accessing data that resides on traditional persistent storage as well as memory devices; (2) Storage-system architecture design that supports scientific workflows on varied hierarchical storage and networking devices; (3) Devising metadata management infrastructure to support FAIR principles (Findability, Accessibility, Interoperability, and Reusability); (4) Capturing provenance information about scientific data; (5) Utilizing AI to learn I/O patterns of emerging workloads for efficient data management; (6) Providing data management support for AI and complex workflows; and (7) Understanding the overlap between traditional storage systems and I/O (SSIO) efforts and data management. While the program committee has identified these topics as important areas for discussion, we welcome position papers from the community that propose additional topics of interest for discussion at the workshop. The workshop agenda will include breakout sessions for discussing these and selected topic areas to inform priority research directions for data management for high-performance and scientific computing.

97 MATHEMATICS AND COMPUTING↗

Report for the ASCR Workshop on the Management and Storage of Scientific Data

The purpose of this workshop is to identify priority research directions in the area of data management for high-performance and scientific computing above and beyond HPC’s traditional "the parallel file system is the data-management system" model. Supporting the breadth of the DOE mission, including the explosion of AI uses and the growing needs of experimental and observational science, motivates revisiting our assumptions about data management. There are many facets of this topic to explore including: (1) Interfaces for accessing data that resides on traditional persistent storage as well as memory devices; (2) Storage-system architecture design that supports scientific workflows on varied hierarchical storage and networking devices; (3) Devising metadata management infrastructure to support FAIR principles (Findability, Accessibility, Interoperability, and Reusability); (4) Capturing provenance information about scientific data; (5) Utilizing AI to learn I/O patterns of emerging workloads for efficient data management; (6) Providing data management support for AI and complex workflows; and (7) Understanding the overlap between traditional storage systems and I/O (SSIO) efforts and data management. While the program committee has identified these topics as important areas for discussion, we welcome position papers from the community that propose additional topics of interest for discussion at the workshop. The workshop agenda will include breakout sessions for discussing these and selected topic areas to inform priority research directions for data management for high-performance and scientific computing.

97 MATHEMATICS AND COMPUTING↗

Persistent optical phenomena in oxide semiconductors

The interaction of transparent oxide semiconductors with light is critically important for a range of applications. Persistent effects could be exploited for holographic memory or optically defined circuits. Conversely, they may also be detrimental to device operation. Large, room-temperature persistent photoconductivity (PPC) was discovered in strontium titanate (SrTiO 3 , STO) after annealing in a hydrogen-containing atmosphere. Barium titanate (BaTiO 3 , BTO), a ferroelectric material, was recently found to also exhibit PPC. Room-temperature photodarkening was observed in Cu-doped gallium oxide (β-Ga 2 O 3 ) after exposure to sub-bandgap light. Hydrogen is believed to play a central role in these persistent phenomena. In the proposed model, a photon excites substitutional hydrogen (a proton inside an oxygen vacancy), making the defect unstable. The proton leaves and binds to a host oxygen atom, forming an O-H bond that is observed with infrared spectroscopy. Here, an oxygen vacancy is left behind. Because oxygen vacancies in STO and BTO are shallow donors, this process results in PPC. In β-Ga 2 O 3 :Cu, however, the oxygen vacancy neighbors a Cu acceptor. In that case, photoexcitation results in the rare Cu 3+ state, which absorbs visible light. The effect can be “erased” by annealing at 300-400°C.

Annealing↗

Manipulating Ferroelectric Topological Polar Structures with Twisted Light

Abstract The dynamic control of non‐equilibrium states represents a central challenge in condensed matter physics. While intense terahertz fields drive metal‐insulator transitions and ferroelectricity via soft phonon modes, recent theory suggests that twisted light with orbital angular momentum (OAM) offers a distinct route to manipulate ferroelectric order and stabilize topological excitations including skyrmions, vortices, and Hopfions. Control of ferroelectric polarization in quasi‐2D CsBiNb 2 O 7 (CBNO) is demonstrated using non‐resonant twisted ultra‐violet (UV) light (375 nm, 800 THz). Combining in situ X‐ray Bragg coherent diffractive imaging (BCDI), twisted optical Raman spectroscopy, and density functional theory (DFT), three‐dimensional (3D) ionic displacements, strain fields, and polarization changes are resolved in single crystals. Operando measurements reveal light‐induced strain hysteresis under twisted light–a hallmark of nonlinear, history‐dependent ferroelastic switching driven by OAM. Discrete, irreversible domain transitions emerge as the topological charge ℓ is cycled, stabilizing non‐trivial domain textures including vortex‐antivortex pairs, Bloch/anti‐Bloch points, and merons. These persist after OAM removal, indicating a memory effect. Competing mechanisms are discussed, including multiphoton absorption, strain‐mediated polarization switching, and defect‐wall interactions. The findings establish structured light as a tool for deterministic, reversible control of ferroic states, enabling optically reconfigurable non‐volatile devices.

Chemistry↗

Phase transition amplification of proton number fluctuations in nuclear collisions from a transport model approach

The time evolution of particle number fluctuations in nuclear collisions at intermediate energies (E lab = 1.23–10⁢A GeV) is studied by means of the UrQMD-3.5 transport model. The transport description incorporates baryonic interactions through a density-dependent potential. This allows for an implementation of a first-order phase transition including a mechanically unstable region at large baryon density. The scaled variance of the baryon and proton number distributions is calculated in the central cubic spatial volume of the collisions at different times. A significant enhancement of fluctuations associated with the unstable region is observed. This enhancement persists to late times, reflecting a memory effect for the fluctuations. The presence of the phase transition has a much smaller influence on the observable event-by-event fluctuations of protons in momentum space.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Influence of plateau, slope, and valley on soil hydrology during the dry season in a Central Amazon old‐growth forest

Soil moisture regulates plant water supply and drought sensitivity in tropical forests, yet its vertical and topographic variation remains poorly characterized. We combined high-frequency time-domain reflectometry measurements from 5 to 100 cm across plateau, slope, and valley landforms at the Zona Florestal 2 research site north of Manaus, Central Amazonia, to quantify how soil moisture memory, timing of responses to rainfall, dry-down rates (τ), and soil–water depletion vary across these contrasting landforms. Landform-specific soil moisture calibration curves ensured accurate volumetric water content estimates in these highly weathered soils. During the 2023 dry-to-wet transition (August–November), soil moisture memory showed strong topographic contrasts, with valley profiles increasing from ∼47 h at 5 cm to ∼154 h at 100 cm, while plateaus exhibited higher near-surface persistence (∼124 h at 5 cm) but weaker memory at depth. Dry-down behavior reinforced these differences as valley soils exhibited τ values exceeding ∼200 h, more than double the characteristic τ of plateau soils (∼90 h). Rainfall–soil moisture correlations indicated immediate responses at shallow depths in valleys and progressively longer lags with depth on plateaus and slopes. These hydrologic patterns were mirrored in depletion profiles, which declined sharply below 30 cm on plateaus but remained high and sustained throughout the upper meter in slopes and valleys. Together, these findings provide the first depth-resolved field measurements of soil moisture memory, rainfall coupling, dry-down constants, and depletion dynamics across major upland landforms in Central Amazonia and offer clear observational benchmarks for improving land-surface and ecosystem model representations of soil–water processes.

Hillslope↗

Persistence length regulates emergent dynamics in active roller ensembles

Active colloidal fluids, biological and synthetic, often demonstrate complex self-organization and the emergence of collective behavior. Spontaneous formation of multiple vortices has been recently observed in a variety of active matter systems, however, the generation and tunability of the active vortices not controlled by geometrical confinement remain challenging. Here, we exploit the persistence length of individual particles in ensembles of active rollers to tune the formation of vortices and to orchestrate their characteristic sizes. We use two systems and employ two different approaches exploiting shape anisotropy or polarization memory of individual units for control of the persistence length. We characterize the dynamics of emergent multi-vortex states and reveal a direct link between the behavior of the persistence length and properties of the emergent vortices. We further demonstrate common features between the two systems including anti-ferromagnetic ordering of the neighboring vortices and active turbulent behavior with a characteristic energy cascade in the particles velocity field energy spectra. Finally, our findings provide insights into the onset of spatiotemporal coherence in active roller systems and suggest a control knob for manipulation of dynamic self-assembly in active colloidal ensembles.

42 ENGINEERING↗

Modeling Workloads of a Linear Electromagnetic Code for Load Balancing Matrix Assembly

This report presents our work to model the workloads of a linear electromagnetic application based on the method of moments in the frequency domain to effectively load balance the matrix assembly. This application is particularly challenging to load balance due to its lack of persistent iterative behavior, its operation under tight memory constraint (where the matrix may fill 80% of memory on each node), and the algorithmic complexity of the computational method. This report describes the first step in our work to apply an inspector-executor approach for load balancing workloads where key parameters are exposed during the inspector phase and a pre-trained model is applied to predict relative task weights for the load balancer.

97 MATHEMATICS AND COMPUTING↗