Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “memory spaces”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

SIMULATION CLONING FOR DIGITAL TWINS: A SCALABLE APPROACH

Digital Twin (DT) methods represent an important technology in which a simulated model of the operations of a physical system uses real-time sensor data to simulate, monitor, and consequently, improve its operations. One of the primary objectives of such a DT is to inform the physical system of measures to be taken in response to one or multiple intervening events that can change the state of the physical system. As such, a capability that is able to carry out multiple scenario assessments in real time in readiness for such events is a very effective tool in the use of simulations as DTs. However, continuous evaluation with highly probable event simulation scenarios are challenging due to the constraints of finite memory and a large exploration space. This paper reports a novel methodology for the continuous evaluation of $k$ probabilistic \textit{what-if} event scenarios under finite resource constraints and demonstrates its use as a digital-twin for a real-world application.

Yoginath, Srikanth↗

Memory Analysis Tool

This tool traces all memory accesses to stack (static allocation) and heap (dynamic allocation) on a trace run in a particular hardware and then estimates execution time on given arbitrary hardware configurations for hardware design space exploration. It can also compute memory access statistics, such as reuse distance.

Sato, Kento↗

Distributed-Memory Sparse Deep Neural Network Inference Using Global Arrays

Partitioned Global Address Space (PGAS) models exhibit tremendous promise in developing efficient and productive distributed-memory parallel applications. They have been used extensively in scientific computations due to conveniently offering a ``shared-memory''-like model and convenient interfaces that separate communication with synchronization. Traditionally, PGAS communication models have been applied to dense/contiguously distributed data, but most modern applications depict varied levels of sparsity. Existing PGAS models require certain adaptations to support distributed sparse computations, since associated computations often require matrix arithmetic, in addition to data movement. The Global Arrays toolkit from Pacific Northwest National Laboratory (PNNL) is one of the earliest PGAS models to combine one-sided data communication and distributed matrix operations and is still used in the popular NWChem quantum chemistry suite. Recently, we have expanded the Global Arrays toolkit to support common sparse operations, like sparse matrix-dense matrix multiplies (SpMM), sparse matrix-sparse matrix multiplication (SpGEMM) and Sampled Dense-Dense Matrix Multiplication (SDDMM). As it turns out, these operations are the bedrock of sparse Deep Learning (DL); sparse deep neural networks and Graph Neural Networks (GNNs) have gained increasing attention recently in achieving speedups on training and inference with reduced memory footprints. Unlike scientific applications in High Performance Computing (HPC), modern (distributed-memory capable) DL toolkits often rely on non-standardized and closed-source vendor software optimizations, creating challenges in software-hardware co-design at scale. Our goal is to support a variety of distributed-memory sparse matrix operations and helper functions in the newly created Sparse Global Arrays (SGA), such that it is possible to build portable and productive Machine Learning scenarios for algorithm/software and hardware codesign purposes. Contemporary data-parallel schemes for training/inference are undergoing a major overhaul since model replication limits scalability and causes resource inefficiencies. As such, we have adopted tensor parallelism in decomposing the model and inputs, to mitigate memory issues. Current implementation is built on top of MPI and uses CPUs to maximize the portability across the platforms.

Distributed computing, machine learning↗

UPC++ v1.0 Programmer’s Guide, Revision 2021.9.0

UPC++ is a C++ library that provides Partitioned Global Address Space (PGAS) programming. It is designed for writing parallel programs that run efficiently and scale well on distributed-memory parallel computers. The PGAS model is single program, multiple-data (SPMD), with each separate constituent process having access to local memory as it would in C++. PGAS additionally provides one-sided Remote Memory Access (RMA) to a global address space, which is allocated in shared segments that are distributed over the processes. UPC++ also features Remote Procedure Call (RPC) communication, making it easy to move computation to operate on data that resides on remote processes. In UPC++, all communication operations are explicit, which encourages programmers to be aware of the cost of communication and data movement. Moreover, all communication operations are asynchronous by default, to enable programmers to write code that scales well even on hundreds of thousands of cores.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Gravitational memory and compact extra dimensions

Here we develop a general formalism for treating radiative degrees of freedom near $\mathscr{I}^+$ in theories with an arbitrary Ricci-flat internal space. These radiative modes are encoded in a generalized news tensor which decomposes into gravitational, electromagnetic, and scalar components. We find a preferred gauge which simplifies the asymptotic analysis of the full nonlinear Einstein equations and makes the asymptotic symmetry group transparent. This asymptotic symmetry group extends the Bondi–Metzner–Sachs (BMS) group to include angle-dependent isometries of the internal space. We apply this formalism to study memory effects, which are expected to be observed in future experiments, that arise from bursts of higher-dimensional gravitational radiation. We outline how measurements made by gravitational wave observatories might probe properties of the compact extra dimensions.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

FLAMES─Fast, Low-Storage, Accurate, and Memory-Efficient Adaptive Sampling─Approach to Resolve Spatially Dependent Dynamics of Molecular Liquids

Many critical phenomena in soft matter occur at large length scales, necessitating the resolution of their structure and dynamics at low wavenumbers. However, resolving wavenumber-dependent dynamics computationally via molecular dynamics simulations presents significant challenges, as these phenomena span several orders of magnitude in both time and length scales, resulting in high computational costs and memory demands. Here, this work highlights the computational and memory challenges associated with analyzing molecular trajectories in reciprocal space and demonstrates a method to address them. We introduce FLAMESFast, Low-storage, Accurate, and Memory-Efficient adaptive Sampling, which is a direct method for calculation of structure factors, allowing us to select only the required number of wavevectors for binning. We also use wavenumber-dependent time steps to extract dynamics. Our FLAMES approach effectively mitigates computational and memory/storage bottlenecks. We demonstrate the method using simulations of a model system, liquid octane, at various temperatures. Comparisons with experimental data and real space computation show that the FLAMES technique achieves high accuracy in resolving temperature- and spatially dependent dynamics while being significantly more computationally efficient and requiring less memory and storage than methods based on a uniform wavevector grid and fixed temporal spacing.

Chen, Guang [Argonne National Laboratory (ANL), Ar↗

Effect of hatch spacing and laser power on microstructure, texture, and thermomechanical properties of laser powder bed fusion (L-PBF) additively manufactured NiTi

This study systematically evaluates the effects of laser powder bed fusion additive manufacturing (L-PBF-AM) parameters (hatch spacing and laser power) on the thermomechanical behavior and microstructure of Ni 50.8 Ti 49.2 shape memory alloy. The samples were fabricated with hatch spacings from 40 to 240 µm and laser powers of 50 and 100 W at a constant scanning speed of 125 mm/s, resulting in parts with volumetric energy density levels from 55 to 666 J/mm 3 and two sets of linear energy densities of 0.4 and 0.8 J/mm. The results showed a reduced melt pool size and discontinuity of scan tracks with decreased laser power. Additionally, the porosity level was increased with larger hatch spacing and lower laser power. More notably, the transformation temperatures increased, and the critical stress, recoverable strain, and functional stability of samples improved with lower hatch spacing, where the recovery ratio of up to 90% was observed, regardless of the employed laser power. This study also discussed the relationship between the fabrication process and texture formation in the L-PBF-AM process. In conclusion, the advantage of L-PBF-AM was revealed in tailoring the microstructure from highly textured samples in [1 1 1] or [0 0 1] direction when hatch spacing lower than laser beam focused was employed, to the appearance of equiaxed solidification front with island grains and random orientations.

36 MATERIALS SCIENCE↗

Unconventional Quantum Advantages for Computation (U-QuAC)

While quantum computing offers the promise of exponential advantages, limited quantum speedups are known, especially for practical applications. To open new avenues for quantum advantages, we propose Unconventional Quantum Advantages for Computation (U-QuACs), with respect to unconventional resources such as space (number of bits or quantum bits of memory required to solve a problem), accuracy of solution, communication, or energy consumption. We focus on space-efficient quantum algorithms, where we seek to design algorithms that solve a problem using much less space than the total size of the input. A natural setting in which space is critical is the streaming model of computation, where the input data arrives sequentially in pieces that must each be processed individually. Streaming is motivated by a variety of problems including analysis of internet traffic or social networks. We design the first exponential quantum space advantage for a natural streaming problem, which also constitutes the first quantum advantage for approximating a discrete optimization problem, albeit with respect to space.

97 MATHEMATICS AND COMPUTING↗

Non-DNA radiosensitive targets that initiate persistent behavioral deficits in rats exposed to space radiation

Predicting future CNS risks for astronauts during deep-space missions will rely substantially on ground-based rodent data with space-relevant ions and behaviors. For rats, the accumulated evidence indicates that less densely ionizing radiation, such as 4 He and 12 C ions, induce behavior deficits at lower doses than densely ionizing ions, such as 48 Ti and 56 Fe. However, this observation conflicts with standard somatic radiobiology, in which densely ionizing ions are generally more effective than less densely ionizing ions, and where the DNA/nucleus is the accepted target for radiation-induced tumorigenesis, cytogenetic aberrations, genetic mutations, and reproductive cell death. To gain deeper insight into the subcellular nature of the radiation targets for behavior risks, we compared the effects of dose, fluence, and linear energy transfer (LET) of 4 He and 56 Fe particles using existing datasets for four distinct behavioral outcomes in rats: elevated plus maze (EPM-anxiety), novel object recognition (NOR-memory), operant responding (OR-response to environmental stimuli), and attentional set-shifting (ATSET-cognitive flexibility). We confirmed that less densely ionizing particles (except protons) showed ~100-fold lower threshold doses than densely ionizing particles for behavioral deficits (0.1–1 cGy for 4 He vs. 15–100 cGy for 56 Fe). However, when analyzed by fluence the behavioral responses converged, indicating that 4 He and 56 Fe were equally effective on a per-track basis. When analyzed by LET, there were ~100-fold differences in the LET for maximum effectiveness for behavioral deficits and DNA endpoints (~1 vs ~100 keV/μm, respectively). These unique features of radiation-induced behavioral deficits (high sensitivity to particles in the 1-keV/μm range, insensitivity to protons in the 0.2 keV/μm range, and isofluence dependence for particles with LET>1 keV/μm) provide evidence in support of a new hypothesis of sub-micron sized radiosensitive targets for behavioral effects consistent with the thickness of plasma membranes and/or small subcellular structures, smaller than a whole synapse. Like our behavior findings, mouse immature oocyte killing which is known to have a plasma membrane target was also better explained by fluence, rather than dose. In contrast, fluence analyses for DNA/nuclear endpoints in somatic cells (e.g., tumor induction, chromosome aberrations) showed opposite results, suggesting that behavior targets are not DNA. Our findings raise questions regarding the identity of subcellular targets and the multi-cellular functional unit for behavior risks, low-dose susceptibility, and generalizability from rat to other species and astronauts.

63 RADIATION, THERMAL, AND OTHER ENVIRON. POLLUTAN↗

Reconfigurable training and reservoir computing in an artificial spin-vortex ice via spin-wave fingerprinting

Strongly interacting artificial spin systems are moving beyond mimicking naturally occurring materials to emerge as versatile functional platforms, from reconfigurable magnonics to neuromorphic computing. Typically, artificial spin systems comprise nanomagnets with a single magnetization texture: collinear macrospins or chiral vortices. Here, by tuning nanoarray dimensions we have achieved macrospin–vortex bistability and demonstrated a four-state metamaterial spin system, the ‘artificial spin-vortex ice’ (ASVI). ASVI can host Ising-like macrospins with strong ice-like vertex interactions and weakly coupled vortices with low stray dipolar field. Vortices and macrospins exhibit starkly differing spin-wave spectra with analogue mode amplitude control and mode frequency shifts of Δf = 3.8 GHz. The enhanced bitextural microstate space gives rise to emergent physical memory phenomena, with ratchet-like vortex injection and history-dependent non-linear fading memory when driven through global magnetic field cycles. We employed spin-wave microstate fingerprinting for rapid, scalable readout of vortex and macrospin populations, and leveraged this for spin-wave reservoir computation. ASVI performs non-linear mapping transformations of diverse input and target signals in addition to chaotic time-series forecasting.

97 MATHEMATICS AND COMPUTING↗

The neutrino gravitational memory from a core collapse supernova: phenomenology and physics potential

General Relativity predicts that the passage of matter or radiation from an asymmetrically-emitting source should cause a permanent change in the local space-time metric. This phenomenon, called the gravitational memory effect, has never been observed, however supernova neutrinos have long been considered a promising avenue for its detection in the future. With the advent of deci-Hertz gravitational wave interferometers, observing the supernova neutrino memory will be possible, with important implications for multimessenger astronomy and for tests of gravity. In this work, we develop a phenomenological (analytical) toy model for the supernova neutrino memory effect, which is overall consistent with the results of numerical simulations. This description is then generalized to several case studies of interest. We find that, for a galactic supernova, the dimensionless strain, h(t), is of order ~ 10 -22 - 10 -21 , and develops over a typical time scale that varies between ~ 0.1 - 10 s, depending on the time-evolution of the anisotropy of the neutrino emission. The characteristic strain, h c (f), has a maximum at a frequency f max ~ Script $\mathcal{O}$(10 -1 ) - Script $\mathcal{O}$(1) Hz. The detailed features of the time- and frequency-structure of the memory strain will inform us of the matter dynamics near the collapsed core, and allow to distinguish between different stellar collapse scenarios. Next generation gravitational wave detectors like DECIGO and BBO will be sensitive to the neutrino memory effect for supernovae at typical galactic distances and beyond; with Ultimate DECIGO exceeding a detectability distance of 10 Mpc

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Privateer

Privateer is a general-purpose data store that optimizes the tradeoff between storage space utilization and I/O performance. Privateer uses memory-mapped I/O with private mapping and an optimized writeback mechanism to maximize write parallelism and eliminate redundant writes; it also uses contentaddressable storage to optimize storage space via de-duplication.

Iwabuchi, Keita↗

GASNet-EX Memory Kinds: Support for Device Memory in PGAS Programming Models

There is an emerging need for adaptive, lightweight communication in irregular HPC applications at exascale, where GPU accelerators provide the majority of available compute cycles. To address this need, Lawrence Berkeley National Lab is developing a programming system to support distributed-memory HPC application development using the Partitioned Global Address Space (PGAS) model. This work includes two major components: UPC++ and GASNet-EX. UPC++ is a C++ template library providing Remote Memory Access (RMA) and Remote Procedure Call (RPC) communication interfaces. GASNet-EX is a portable, high-performance communication middleware library, used by the implementations of UPC++ and many other PGAS programming models. We describe recent advances in GASNet-EX to efficiently implement zero-copy Remote Memory Access (RMA) communication to and from memory on accelerator devices such as GPUs. We demonstrate performance improvements via benchmark results from UPC++ (on Summit) and the Legion programming system (on DGX-1), both using GASNet-EX for communication.

Hargrove, Paul H↗

pnnl/memgaze

MemGaze is a memory analysis toolset that combines high-resolution trace analysis and low overhead measurement, both with respect to time and space. It provides high-resolution by collecting world-level memory access traces, where the highest resolution supported is back-to-back sequences. It achieves low-overhead in space and time by leveraging sampling and various methods of hardware support for collecting traces. Its post-mortem trace processing provides multiresolution analysis for locations vs. operations; accesses vs. spatio-temporal reuse, and reuse (distance, rate, volume) vs. access patterns.

Central, PNNL Developer↗

Time-series learning of latent-space dynamics for reduced-order model closure

In this work, we study the performance of long short-term memory networks (LSTMs) and neural ordinary differential equations (NODEs) in learning latent-space representations of dynamical equations for an advection-dominated problem given by the viscous Burgers equation. Our formulation is devised in a nonintrusive manner with an equation-free evolution of dynamics in a reduced space with the latter being obtained through a proper orthogonal decomposition. In addition, we leverage the sequential nature of learning for both LSTMs and NODEs to demonstrate their capability for closure in systems that are not completely resolved in the reduced space. We assess our hypothesis for two advection-dominated problems given by the viscous Burgers equation. We observe that both LSTMs and NODEs are able to reproduce the effects of the absent scales for our test cases more effectively than does intrusive dynamics evolution through a Galerkin projection. This result empirically suggests that time-series learning techniques implicitly leverage a memory kernel for coarse-grained system closure as is suggested through the Mori–Zwanzig formalism.

97 MATHEMATICS AND COMPUTING↗

SpecSims: A Scalable Speculative Tree-based Simulation Cloning Framework for Finite Memory Machines

Simulation cloning is a technique in which cloned simulations whose state spaces differ partially from their parent simulation due to intervening events are spawned at runtime and concurrently advanced. It is a powerful method to carry out what-if analysis by speculatively exploring and evaluating the impact of various permutations of intervening cascade of events. Due to the exponential growth in the number of possible clones even for a small number of distinct intervening events, the practical efficacy of the approach is often severely limited by the maximum available memory of the computing host. In this paper, we introduce a novel speculative simulation cloning framework that executes a simulation cloning campaign capable of efficiently exploring an exponentially large space of clone simulations created by permutation of intervening events under a finite memory constraint. We provide a theoretical analysis of the runtime characteristics of our proposed approach and highlight its novel advantages such as memory-aware and as-long-as-needed execution. Furthermore, in support of our analytical findings and to demonstrate its practical feasibility, we implement a prototype of the cloning framework on a shared memory system and report its performance characteristics in the context of a heat diffusion simulation, and a power grid simulation subject to cascading disruptions from geomagnetic disturbances.

Simulation framework↗

UPC++ v1.0 Programmer’s Guide (Rev. 2023.9.0)

UPC++ is a C++ library that supports Partitioned Global Address Space (PGAS) programming. It is designed for writing efficient, scalable parallel programs on distributed-memory parallel computers. The key communication facilities in UPC++ are one-sided Remote Memory Access (RMA) and Remote Procedure Call (RPC). The UPC++ control model is single program, multiple-data (SPMD), with each separate constituent process having access to local memory as it would in C++. The PGAS memory model additionally provides one-sided RMA communication to a global address space, which is allocated in shared segments that are distributed over the processes. UPC++ also features Remote Procedure Call (RPC) communication, making it easy to move computation to operate on data that resides on remote processes. UPC++ was designed to support exascale high-performance computing, and the library interfaces and implementation are focused on maximizing scalability. In UPC++, all communication operations are syntactically explicit, which encourages programmers to consider the costs associated with communication and data movement. Moreover, all communication operations are asynchronous by default, encouraging programmers to seek opportunities for overlapping communication latencies with other useful work. UPC++ provides expressive and composable abstractions designed for efficiently managing aggressive use of asynchrony in programs. Together, these design principles are intended to enable programmers to write applications using UPC++ that perform well even on hundreds of thousands of cores.

97 MATHEMATICS AND COMPUTING↗

UPC++ v1.0 Programmer’s Guide (Revision 2022.3.0)

UPC++ is a C++ library that supports Partitioned Global Address Space (PGAS) programming. It is designed for writing efficient, scalable parallel programs on distributed-memory parallel computers. The key communication facilities in UPC++ are one-sided Remote Memory Access (RMA) and Remote Procedure Call (RPC). The UPC++ control model is single program, multiple-data (SPMD), with each separate constituent process having access to local memory as it would in C++. The PGAS memory model additionally provides one-sided RMA communication to a global address space, which is allocated in shared segments that are distributed over the processes. UPC++ also features Remote Procedure Call (RPC) communication, making it easy to move computation to operate on data that resides on remote processes. UPC++ was designed to support exascale high-performance computing, and the library interfaces and implementation are focused on maximizing scalability. In UPC++, all communication operations are syntactically explicit, which encourages programmers to consider the costs associated with communication and data movement. Moreover, all communication operations are asynchronous by default, encouraging programmers to seek opportunities for overlapping communication latencies with other useful work. UPC++ provides expressive and composable abstractions designed for efficiently managing aggressive use of asynchrony in programs. Together, these design principles are intended to enable programmers to write applications using UPC++ that perform well even on hundreds of thousands of cores.

97 MATHEMATICS AND COMPUTING↗