Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “memory spaces”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Picasso: Memory-Efficient Graph Coloring Using Palettes With Applications in Quantum Computing

A coloring of a graph is an assignment of colors to vertices such that no two neighboring vertices have the same color. The need for memory-efficient coloring algorithms is motivated by their application in computing clique partitions of graphs arising in quantum computations where the objective is to map a large set of Pauli strings into a compact set of unitaries. We present Picasso, a randomized memory-efficient iterative parallel graph coloring algorithm with theoretical sublinear space guarantees under practical assumptions. The parameters of our algorithm provide a trade-off between coloring quality and resource consumption. To assist the user, we also propose a machine learning model to predict the coloring algorithm’s parameters considering these trade-offs. We provide a sequential and a parallel implementation of the proposed algorithm. We perform an experimental evaluation on a 64-core AMD CPU equipped with 512 GB of memory and an Nvidia A100 GPU with 40GB of memory. For a small dataset where existing coloring algorithms can be executed within the 512 GB memory budget, we show up to 68× memory savings. On massive datasets we demonstrate that GPU-accelerated Picasso can process inputs with 49.5× more Pauli strings (vertex set in our graph) and 2,478× more edges than state-of-the-art parallel approaches.

artificial intelligence, quantum computing↗

Nonvolatile electrochemical memory at 600°C enabled by composition phase separation

Silicon-based microelectronics are limited to ~150°C and therefore not suitable for the extremely high temperatures in aerospace, energy, and space applications. While wide-band-gap semiconductors can provide high-temperature logic, nonvolatile memory devices at high temperatures have been challenging. In this work, we develop a nonvolatile electrochemical memory cell that stores and retains analog and digital information at temperatures as high as 600°C. Through correlative scanning transmission electron microscopy, we show that this high-temperature information retention is a result of composition phase separation between the oxidized and reduced forms of amorphous tantalum oxide. This result demonstrates a memory concept that is resilient at extreme temperatures and reveals phase separation as the principal mechanism that enables nonvolatile information storage in these electrochemical memory cells.

42 ENGINEERING↗

Quantum optical memory for entanglement distribution

Optical photons are powerful carriers of quantum information, which can be delivered in free space by satellites or in fibers on the ground over long distances. Entanglement of quantum states over long distances can empower quantum computing, quantum communications, and quantum sensing. Quantum optical memories are devices designed to store quantum information in the form of stationary excitations, such as atomic coherence, and are capable of coherently mapping these excitations to flying qubits. Quantum memories can effectively store and manipulate quantum states, making them indispensable elements in future long-distance quantum networks. Over the past two decades, quantum optical memories with high fidelities, high efficiencies, long storage times, and promising multiplexing capabilities have been developed, especially at the single-photon level. In this review, we introduce the working principles of commonly used quantum memory protocols and summarize the recent advances in quantum memory demonstrations. We also offer a vision for future quantum optical memory devices that may enable entanglement distribution over long distances.

Lei, Yisheng↗

BCSR on GPU: A Way Forward Extreme-scale Graph Processing on Accelerator-enabled Frontier Supercomputer

Handling large graphs in a distributed environment requires effective partitioning across processors and efficient management of local partitions. In 2D partitioning, local graphs often become too sparse, making memory-efficient data structures crucial. Using the Compressed Sparse Row (CSR) format wastes space, especially for > 83% of vertices with empty edges for the sparse graphs. This study explores bit-CSR (BCSR), a modified CSR representation, on GPUs to reduce memory usage in graph computations. We achieved 16.67% memory savings on a sparse rmat dataset with 268 million vertices and 357 million edges, without performance degradation, supported by both theoretical and experimental storage savings of 33%. However, we observed a 1.7× slowdown in degree lookup times due to bitwise operations on AMD CPUs. This analysis highlights the potential of BCSR on GPUs for improving Graph500 benchmark performance on GPU-accelerated systems, such as the Frontier supercomputer.

Sattar, Naw Safrin↗

A Decentralized Approach for Modeling Organized Convection Based on Thermal Populations on Microgrids

Abstract In this study, a spectral model for convective transport is coupled to a thermal population model on a two‐dimensional horizontal “microgrid,” covering the typical gridbox size of general circulation models. The goal is to explore new ways of representing impacts of spatial organization in cumulus cloud fields. The thermals are considered the smallest building block of convection, with thermal life cycle and movement represented through binomial functions. Thermals interact through two simple rules, reflecting pulsating growth and environmental deformation. Long‐lived thermal clusters thus form on the microgrid, exhibiting scale growth and spacing that represent simple forms of spatial organization and memory. Size distributions of cluster number are diagnosed from the microgrid through an online clustering algorithm, and provided as input to a spectral multiplume eddy‐diffusivity mass flux scheme. This yields a decentralized transport system, in that the thermal clusters acting as independent but interacting nodes that carry information about spatial structure. The main objectives of this study are (a) to seek proof of concept of this approach, and (b) to gain insight into impacts of spatial organization on convective transport. Single‐column model experiments demonstrate satisfactory skill in reproducing two observed cases of continental shallow convection. Metrics expressing self‐organization and spatial organization match well with large‐eddy simulation results. We find that in this coupled system, spatial organization impacts convective transport primarily through the scale break in the size distribution of cluster number. The rooting of saturated plumes in the subcloud mixed layer plays a key role in this process.

54 ENVIRONMENTAL SCIENCES↗

Fourier-based three-dimensional multistage transformer for aberration correction in multicellular specimens

High-resolution tissue imaging is often compromised by sample-induced optical aberrations that degrade resolution and contrast. Although wavefront sensor-based adaptive optics (AO) can measure these aberrations, such hardware solutions are typically complex, expensive to implement and slow when serially mapping spatially varying aberrations across large fields of view. Here we introduce AOViFT (adaptive optical vision Fourier transformer)—a machine learning-based aberration sensing framework built around a three-dimensional multistage vision transformer that operates on Fourier domain embeddings. AOViFT infers aberrations and restores diffraction-limited performance in puncta-labeled specimens with substantially reduced computational cost, training time and memory footprint compared to conventional architectures or real-space networks. We validated AOViFT on live gene-edited zebrafish embryos, demonstrating its ability to correct spatially varying aberrations using either a deformable mirror or postacquisition deconvolution. By eliminating the need for the guide star and wavefront sensing hardware and simplifying the experimental workflow, AOViFT lowers technical barriers for high-resolution volumetric microscopy across diverse biological samples.

Alshaabi, Thayer [Howard Hughes Medical Institute,↗

An infrared on-shell action and its implications for soft charge fluctuations in asymptotically flat spacetimes

We study the infrared on-shell action of Einstein gravity in asymptotically flat spacetimes (AFSs), obtaining an effective, gauge-invariant boundary action for memory and shockwave spacetimes. We show that the phase space is in both cases parameterized by the leading soft variables in AFSs, thereby extending the equivalence between shockwave and soft commutators to spacetimes with non-vanishing Bondi mass. We then demonstrate that our on-shell action is equal to three quantities studied separately in the literature: (i) the soft supertranslation charge; (ii) the shockwave effective action, or equivalently the modular Hamiltonian; and (iii) the soft effective action. Finally, we compute the quantum fluctuations in the soft supertranslation charge and, assuming the supertranslation parameter may be promoted to an operator, we obtain an area law, consistent with earlier results showing that the modular Hamiltonian has such fluctuations.

asymptotically flat spacetimes↗

Implementation of a Binary Neural Network on a Passive Array of Magnetic Tunnel Junctions

The increasing scale of neural networks and their growing application space have produced demand for more energy- and memory-efficient artificial-intelligence-specific hardware. Avenues to mitigate the main issue, the von Neumann bottleneck, include in-memory and near-memory architectures, as well as algorithmic approaches. In this report we leverage the low-power and the inherently binary operation of magnetic tunnel junctions (MTJs) to demonstrate neural network hardware inference based on passive arrays of MTJs. In general, transferring a trained network model to hardware for inference is confronted by degradation in performance due to device-to-device variations, write errors, parasitic resistance, and nonidealities in the substrate. To quantify the effect of these hardware realities, we benchmark 300 unique weight matrix solutions of a two-layer perceptron to classify the Wine dataset for both classification accuracy and write fidelity. Despite device imperfections, we achieve software-equivalent accuracy of up to 95.3% with proper tuning of network parameters in 15 x 15 MTJ arrays having a range of device sizes. The success of this tuning process shows that new metrics are needed to characterize the performance and quality of networks reproduced in mixed signal hardware.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

P38 heterogeneous multi-tiled system with support for message queues (MoSAIC) v0.1

The proposed system is written in the hardware description language (HDL) verilog targeting an FPGA board. It is intended as a testbed to explore architecture tradeoffs in multi-tiled heterogeneous architectures. Although we target FPGAs, the system can be implemented as a monolithic SoC or a package comprised of many chiplets that are interconnected in the same package using a NoC. The proposed NoC is lightweight and follows an axi-lite interface. The endpoints of the NoC are a heterogeneous mix of "tiles" as endpoints that are general purpose processors, fixed function accelerators, and programmable accelerators. We assume that the network interfaces for the NoC endpoints are all addressable in a global name-space in that they represent an address range (for memory addresses) or a range of unique identifiers that are associated with each individual tile. This makes the functionality abstract from the standpoint of the NoC design details. Message queues offer a direct inter-processor interface between peer general purpose cores and diverse accelerators that comprise an SoC. Although they share the same NoC infrastructure for inter-tile communication within an SoC or SiP, the hardware message queues bypass the memory hierarchy and thus do not pollute the memory state or invoke the cache coherence mechanism.

Gonzalez, LouisaPatricia↗

Performance Characteristics of the BlueField-2 SmartNIC

High-performance computing (HPC) researchers have long envisioned scenarios where application workflows could be improved through the use of programmable processing elements embedded in the network fabric. Recently, vendors have introduced programmable Smart Network Interface Cards (SmartNICs) that enable computations to be offloaded to the edge of the network. There is great interest in both the HPC and high-performance data analytics (HPDA) communities in understanding the roles these devices may play in the data paths of upcoming systems. This paper focuses on characterizing both the networking and computing aspects of NVIDIA’s new BlueField-2 SmartNIC when used in a 100Gb/s Ethernet environment. For the networking evaluation we conducted multiple transfer experiments between processors located at the host, the SmartNIC, and a remote host. These tests illuminate how much effort is required to saturate the network and help estimate the processing headroom available on the SmartNIC during transfers. For the computing evaluation we used the stress-ng benchmark to compare the BlueField-2 to other servers and place realistic bounds on the types of offload operations that are appropriate for the hardware. Our findings from this work indicate that while the BlueField-2 provides a flexible means of processing data at the network’s edge, great care must be taken to not overwhelm the hardware. While the host can easily saturate the network link, the SmartNIC’s embedded processors may not have enough computing resources to sustain more than half the expected bandwidth when using kernel-space packet processing. From a computational perspective, encryption operations, memory operations under contention, and on-card IPC operations on the SmartNIC perform significantly better than the general-purpose servers used for comparisons in our experiments. Therefore, applications that mainly focus on these operations may be good candidates for offloading to the SmartNIC.

97 MATHEMATICS AND COMPUTING↗

Finding MIDDLE Ground: Scalable and Secure Distributed Learning

Edge computing methods allow devices to efficiently train a high-performing, robust, and personalized model for predictive tasks. However, these methods succumb to privacy and scalability concerns such as adversarial data recovery and expensive model communication. Furthermore, edge computing methods unrealistically assume that all devices train an identical model. In practice, edge devices have varying computational and memory constraints which may not allow certain devices to have the space or speed to train a specific model. To overcome these issues, we propose MIDDLE: a model independent distributed learning algorithm which allows heterogeneous edge devices to assist each other’s training while communicating only non-sensitive information. MIDDLE unlocks the ability for edge devices, regardless of computational or memory constraints, to assist each other even with completely different model architectures. Furthermore, MIDDLE does not require model or gradient communication which greatly reduces communication size and time. We prove that MIDDLE attains the optimal convergence rate O(1/sqrt(TM)) of stochastic gradient descent for convex and non-convex smooth optimization (for total iterations T and batch size M). Finally, our experimental results demonstrate that MIDDLE (even in non-IID data settings) attains robust and high-performing models without model or gradient communication.

Bornstein, Marc I.↗

Spinodal enhancement of fluctuations in nucleus-nucleus collisions

Subensemble Acceptance Method (SAM) [1, 2] is an essential link between measured event-by-event fluctuations and their grand canonical theoretical predictions such as lattice QCD. The method allows quantifying the global conservation law effects in fluctuations. In its basic formulation, SAM requires a sufficiently large system such as created in central nucleus-nucleus collisions and sufficient space-momentum correlations. Directly in the spinodal region of the First Order Phase Transition (FOPT) different approximations should be used that account for finite size effects. Thus, we present the generalization of SAM applicable in both the pure phases, metastable and unstable regions of the phase diagram [3]. Obtained analytic formulas indicate the enhancement of fluctuations due to crossing the spinodal region of FOPT and are tested using molecular dynamics simulations. A rather good agreement is observed. Using transport model calculations with interaction potential we show that the spinodal enhancement of fluctuations survives till the later stages of collision via the memory effect [4]. However, at low collision energies the space-momentum correlation is not strong enough for this signal to be transferred to second and third order cumulants measured in momentum subspace. This result agrees well with recent HADES data on proton number fluctuations at $\sqrt{S_{NN}}$ = 2.4 GeV which are found to be consistent with the binomial momentum space acceptance [5].

Poberezhnyuk, Roman↗

LDRD Abbreviated report: High-Order General-Discrete-Ordinates Method Enabling Efficient Deterministic Transport in Hydrodynamic Simulations

Deterministic transport simulations for national-security and energy applications often operate in high-dimensional phase-space, where accuracy and cost both become major challenges. A common numerical artifact in such problems is the “ray-effect,” which appears as unphysical streaks. Beyond misinterpretation, these artifacts can contaminate tightly coupled physics, such as fluid dynamics, radiation-hydrodynamics, and laser-plasma interactions, eroding the predictive capability of entire multiphysics workflows. Our objective was to make high-dimension studies practical on modern hardware while mitigating the ray-effect without relying on prohibitively expensive sampling approaches such as Monte Carlo methods. We developed the Generic Discretization Library (GenDiL), a Graphics Processing Unit (GPU)-first framework that uses high-order Discontinuous Galerkin (DG) methods and matrix-free algorithms to reduce memory usage and improve computational efficiency, critical for phase-space simulations. GenDiL supports phase-space adaptivity in both mesh size and polynomial order (hp-adaptivity) to place resolution only where it is needed. A central capability is Local Dimensional Refinement (LDR), which couples lower-dimension continuum models to higher-dimension kinetic models through stable and conservative interfaces, so that high-fidelity physics is applied only in regions where it is essential. Building on the GenDiL framework, we developed the General SN (GSN) family of algorithms as a true generalization of the polar SN approach (discrete ordinates, often denoted SN). Rather than tying discrete ordinates to a specific polar change of coordinates, GSN formulates transport on an arbitrary change of coordinates chosen to reduce ray-effect. We studied two complementary variants: an analytic variant, where the coordinate map is prescribed in advance by a closed-form function; and a data-driven variant, where a quantity of interest, such as the net flux, guides the coordinate system. GenDiL provides the library infrastructure for efficient GPU execution, but the GSN concept is algorithmic and independent of any one library. Across representative high-dimension tests, including non-symmetric solutions, both variants delivered strong ray-effect mitigation at practical cost, moving four- to six-dimensional analysis toward repeatable, routine studies.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

$χ$-$MeRA$: Computationally efficient adaptive mesh refinement of Monte Carlo mesh based tallies

Here, the reactor physics community is always focused on reducing the computational time and memory required for simulations. $χ$-$MeRA$, which stands for flux-based-($χ$)-Mesh tally Refinement Adaptively, was built to reduce the computational time and memory required to solve the neutronics side of a multiphysics problem when compared to traditional methods for mesh based tallies in Monte Carlo (MC) simulations. $χ$-$MeRA$ couples a MC code with an adaptive mesh refinement (AMR) algorithm to take advantage of the accuracy of a MC code and the efficiency of an AMR algorithm. Also developed within $χ$-$MeRA$ was a set of metrics to assess the effects of the refinement on various parameters in the simulation space. For a plutonium sphere, $χ$-$MeRA$ shows a reduction in memory usage and computation time when compared to a fully refined mesh by a factor of 14.7 and 6.7, respectively. When compared to an unstructured mesh, improvement of 1.3 and 4.8 was achieved for memory usage and computation time. The development of $χ$-$MeRA$ helps solve the neutronics side of a multiphysics problem in a faster, more computationally efficient manner than traditional methods, and the final mesh created contains accurate results that can be passed onto the next physics code.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

DEVELOPMENT OF INEXPENSIVE HIGH TEMPERATURE NITI-BASED SHAPE MEMORY ALLOYS FOR POWDER BED ADDITIVE MANUFACTURING

NiTi and NiTi-based Shape Memory Alloys (SMA) exhibit a reversible solid-state phase transformation from martensite to austenite driven by thermal energy. High temperature (Mf>100°C) SMAs are martensite at room temperature and can be fabricated into solid-state actuators that return to a pre-programmed shape against a designed load after heating to transformation threshold. Reactive as-fabricated additively manufactured parts (4-D printing) is the current state of the art in manufacturing of SMAs but requires compositions compliant to rapid solidification. Existing actuator designs are developed from commercially available, highly investigated material compositions. However, existing high temperature high performance (high actuation strain, low thermal hysteresis) shape memory alloys contain significant (>10% at.) portions of high-cost Platinum Group Metals (PGMs). It is of significant scientific interest to investigate material compositions that are peer performing or superior to PGMs whose constituent elements represent a significant cost savings. Shape memory alloy properties vary significantly with small (0.1% at.) compositional changes making robust investigative sample sets very large. Computational material design can be deployed to shrink the compositional space of possible alloy combinations and reduce the experimental load in material discovery. Investigating shape memory effect (SME) and validating process additive process parameters for a single novel composition is cost intensive in both time and consumed materials. Additionally, sub-optimal processing, oxygen, or solidification rate sensitivity could render additively manufacturing specimens without micro, macro cracks, or significant chemical variance impossible. Unfortunately, such failure susceptibility cannot be simulated. Therefore, a research pathway to validate novel shape memory alloy compositions for powder bed fusion additive manufacturing without the need for powdered feedstock is also proposed. This research investigates novel high temperature shape memory alloys for actuators without platinum group alloying elements to discover one that could be commercially viable as an additive manufacturing feedstock.

Sundermann, Tayler↗

Quantum Time-Space Tradeoffs for Matrix Problems

We consider the time and space required for quantum computers to solve a wide variety of problems involving matrices, many of which have only been analyzed classically in prior work. Our main results show that for a range of linear algebra problems—including matrix-vector product, matrix inversion, matrix multiplication and powering—existing classical time-space tradeoffs, several of which are tight for every space bound, also apply to quantum algorithms with at most a constant factor loss. For example, for almost all fixed matrices 𝐴, including the discrete Fourier transform matrix, we prove that quantum circuits with at most 𝑇 input queries and 𝑆 qubits of memory require 𝑇 = Ω⁢(𝑛 2 /𝑆) to compute matrix-vector product 𝐴⁢𝑥 for 𝑥 ∈{0,1 𝑛 . We similarly prove that matrix multiplication for 𝑛 ×𝑛 binary matrices requires 𝑇 = Ω⁢(𝑛 3 /$\sqrt{𝑆}$). Because many of our lower bounds are matched by deterministic algorithms with the same time and space complexity, our results show that quantum computers cannot provide any asymptotic advantage for these problems with any space bound. We obtain matching lower bounds for the stronger notion of quantum cumulative memory complexity—the sum of the space per layer of a circuit. We also consider Boolean (i.e., AND-OR) matrix multiplication and matrix-vector products, improving the previous quantum time-space tradeoff lower bounds for 𝑛 × 𝑛 Boolean matrix multiplication to 𝑇 = Ω⁢(𝑛 2.5 /𝑆 1/4 ) from 𝑇 = Ω⁢(𝑛 2.5 /𝑆 1/2 ). Our improved lower bound for Boolean matrix multiplication is based on a new coloring argument that extracts more from the strong direct product theorem that was the basis for prior work. To obtain our tight lower bounds for linear algebra problems, we require much stronger bounds than strong direct product theorems. We obtain these bounds by adding a new bucketing method to the quantum recording-query technique of Zhandry that lets us apply classical arguments to upper bound the success probability of quantum circuits.

lower bounds↗

Data shuffling with hierarchical tuple spaces

Methods and systems for shuffling data to generate a dataset are described. A first map module may generate first pair data, and a second map module may generate second pair data, from source data. The first map module may insert the first pair data into a first local tuple space accessible to the first map module. The second map module may insert the second pair data into a second local tuple space accessible to the second map module. A shuffle module may request pair data that includes a particular key. The first and second pair data may be inserted into a global tuple space accessible by the first and second map modules. The shuffle module may identify the requested pair data in the global tuple space, and may fetch the identified pair data from a memory. The shuffle module may shuffle the fetched pair data to generate the dataset.

Kayi, Abdullah↗

Optimal control of large quantum systems: assessing memory and runtime performance of GRAPE

Abstract Gradient Ascent Pulse Engineering (GRAPE) is a popular technique in quantum optimal control, and can be combined with automatic differentiation (AD) to facilitate on-the-fly evaluation of cost-function gradients. We illustrate that the convenience of AD comes at a significant memory cost due to the cumulative storage of a large number of states and propagators. For quantum systems of increasing Hilbert space size, this imposes a significant bottleneck. We revisit the strategy of hard-coding gradients in a scheme that fully avoids propagator storage and significantly reduces memory requirements. Separately, we present improvements to numerical state propagation to enhance runtime performance. We benchmark runtime and memory usage and compare this approach to AD-based implementations, with a focus on pushing towards larger Hilbert space sizes. The results confirm that the AD-free approach facilitates the application of optimal control for large quantum systems which would otherwise be difficult to tackle.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗