Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “spatial memory”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

The Case for Strong Scaling in Deep Learning: Training Large 3D CNNs with Hybrid Parallelism

Here, we present scalable hybrid-parallel algorithms for training large-scale 3D convolutional neural networks. Deep learning-based emerging scientific workflows often require model training with large, high-dimensional samples, which can make training much more costly and even infeasible due to excessive memory usage. We solve these challenges by extensively applying hybrid parallelism throughout the end-to-end training pipeline, including both computations and I/O. Our hybrid-parallel algorithm extends the standard data parallelism with spatial parallelism, which partitions a single sample in the spatial domain, realizing strong scaling beyond the mini-batch dimension with a larger aggregated memory capacity. We evaluate our proposed training algorithms with two challenging 3D CNNs, CosmoFlow and 3D U-Net. Our comprehensive performance studies show that good weak and strong scaling can be achieved for both networks using up to 2K GPUs. More importantly, we enable training of CosmoFlow with much larger samples than previously possible, realizing an order-of-magnitude improvement in prediction accuracy.

97 MATHEMATICS AND COMPUTING↗

A review of materials used in tomographic volumetric additive manufacturing

Abstract Volumetric additive manufacturing is a novel fabrication method allowing rapid, freeform, layer-less 3D printing. Analogous to computer tomography (CT), the method projects dynamic light patterns into a rotating vat of photosensitive resin. These light patterns build up a three-dimensional energy dose within the photosensitive resin, solidifying the volume of the desired object within seconds. Departing from established sequential fabrication methods like stereolithography or digital light printing, volumetric additive manufacturing offers new opportunities for the materials that can be used for printing. These include viscous acrylates and elastomers, epoxies (and orthogonal epoxy-acrylate formulations with spatially controlled stiffness) formulations, tunable stiffness thiol-enes and shape memory foams, polymer derived ceramics, silica-nanocomposite based glass, and gelatin-based hydrogels for cell-laden biofabrication. Here we review these materials, highlight the challenges to adapt them to volumetric additive manufacturing, and discuss the perspectives they present. Graphical abstract

36 MATERIALS SCIENCE↗

Gravitational memory and soft theorems: The local perspective

In general relativity, gravitational memory describes the lasting change in the separation and relative velocity of freely falling detectors after the passage of gravitational waves (GWs). In this paper, we elucidate the relation between Bondi-Metzner-Sachs transformations at future null infinity and the description of gravitational memory in local synchronous coordinates, commonly used in GW detectors like LISA. We show that gravitational memory corresponds to large residual diffeomorphisms in this gauge, such as volume-preserving spatial rescalings. We reproduce the associated soft theorems for scattering amplitudes. Finally, we derive novel soft theorems for equal-time (in-in) correlation functions, which are recognized as the flat space analogues of inflationary consistency relations with a soft tensor mode. Furthermore, these relations provide a pathway toward uncovering deeper connections between gravitational memory and cosmological correlators.

General relativity↗

Understanding the Design Space of Sparse/Dense Multiphase Dataflows for Mapping Graph Neural Networks on Spatial Accelerators

Graph Neural Networks (GNNs) have garnered a lot of recent interest because of their success in learning representations from graph-structured data across several critical applications in cloud and HPC. Owing to their unique compute and memory characteristics that come from an interplay between dense and sparse phases of computations, the emergence of reconfigurable dataflow (aka spatial) accelerators offers promise for acceleration by mapping optimized dataflows (i.e., computation order and parallelism) for both phases. The goal of this work is to characterize and understand the design-space of dataflow choices for running GNNs on spatial accelerators in order for the compilers to optimize the dataflow based on the workload. Specifically, we propose a taxonomy to describe all possible choices for mapping the dense and sparse phases of GNNs spatially and temporally over a spatial accelerator, capturing both the intra-phase dataflow and the inter-phase (pipelined) dataflow. Using this taxonomy, we do deep-dives into the cost and benefits of several dataflows and perform case studies on implications of hardware parameters for dataflows and value of flexibility to support pipelined execution.

97 MATHEMATICS AND COMPUTING↗

Enhanced deep neural networks with transfer learning for distribution LMP considering load and PV uncertainties

As the flexibility of generation and demand increases in distribution systems, the residential loads are emerging as a promising means to participate in demand response and the transactive energy market. Market pricing is an instrumental mechanism for the distribution system operator to exploit the full potential of the flexible resources. The distribution locational marginal price (DLMP) can be used to guide the residential load consumption. This type of market signal helps the distribution system operator to optimize the scheduling of all resources while satisfying related network constraints through a day-ahead market. However, solving the optimization problem for large-scale systems can be computationally expensive. To address the scalability and practicability limitations of the DLMP framework, a learning-based approach is proposed in this paper to complement the day-ahead distribution market framework. Here, the proposed approach combines long short-term memory and transfer learning to develop deep neural network that can capture the spatial–temporal correlation of the input data. The model can determine the optimal DLMP for each node in a distribution system without the system parameters required to formulate the optimization problem. Testing results on IEEE 33-bus and 123-bus systems show that the proposed approach can generate a comparable DLMP against the optimization solutions.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Enhancing Lattice Kinetic Schemes for Fluid Dynamics with Lattice-Equivariant Neural Networks

A new class of equivariant neural networks is presented, hereby dubbed lattice-equivariant neural networks (LENNs), designed to satisfy local symmetries of a lattice structure. The approach develops within a recently introduced framework aimed at learning neural network-based surrogate models’ lattice Boltzmann collision operators. Whenever neural networks are employed to model physical systems, respecting symmetries and equivariance properties has been shown to be key for accuracy, numerical stability, and performance. Here, hinging on ideas from group representation theory, trainable layers are defined whose algebraic structure is equivariant with respect to the symmetries of the lattice cell. In this work, the presented method naturally allows for efficient implementations, in terms of both memory usage and computational costs, supporting scalable training/testing for lattices in two spatial dimensions and higher (in which the size of symmetry group grows). The approach is validated and tested considering 2D and 3D flowing dynamics, both in laminar and turbulent regimes. It is compared with group-averaged-based symmetric networks and with plain, nonsymmetric, networks, showing how the presented approach unlocks the (a posteriori) accuracy and training stability of the former models and the train/inference speed of the latter networks. (LENNs are about one order of magnitude faster than group-averaged networks in 3D.) The work in this paper opens toward practical use of machine learning-augmented lattice Boltzmann CFD in real-world simulations.

97 MATHEMATICS AND COMPUTING↗

Switching current reduction in magnetoresistive random access memories

We present an approach for minimizing the critical current for the magnetization switching in magnetic tunnel junctions by optimizing the spatial distribution of the current density. We show that such a minimization is possible because critical current is determined by the condition of making one of the magnetization eigenstates grow in time. The excitation of the eigenstates is enhanced when the spatial distributions of the eigenstates and current density overlap. Critical current can be viewed as a functional of the current density spatial distribution and it can be minimized by optimizing this distribution. Such an optimization results in a major reduction of the critical current and increase of the switching efficiency, viz. the ratio between the energy barrier and critical current. The minimized critical current increases approximately linearly with the magnetic tunnel junction size, which is much slower than critical current for the case of a uniform current density. The optimized efficiency can be approximately a constant with respect to the magnetic tunnel junction size, which is much higher than the efficiency for the uniform current density. Additional optimization can be achieved by spatially modulating the material parameters, e.g., the saturation magnetization. The presented approach and obtained scaling of the critical current and efficiency offers opportunities for the magnetic tunnel junction optimization.

36 MATERIALS SCIENCE↗

Accelerating geostatistical modeling using geostatistics-informed machine Learning

Ordinary Kriging (OK) is a popular geostatistical algorithm for spatial interpolation and estimation. The computational complexity of OK changes quadratically and cubically for memory and speed, respectively, given the number of data. Therefore, it is computationally intensive and also challenging to process a large set of data, especially in three-dimensional (3D) cases. This paper develops a geostatistics-informed machine learning (GIML) model to improve the efficiency of OK by reducing the number of points required to be estimated using OK. Specifically, only a very few of the unknown points are estimated by OK to get the weights and estimations, which are used as the training dataset. Moreover, the governing equations of OK are used to guide our proposed machine learning to better reproduce the spatial distributions. Our results show that the proposed GIML can reduce the computational time of OK by at least one order of magnitude. The effectiveness of the GIML is evaluated and compared using a 2D case. Furthermore, we demonstrate its efficiency and robustness by considering a different number of training samples on various 3D simulation grids.

58 GEOSCIENCES↗

Attribute-Aware RBFs: Interactive Visualization of Time Series Particle Volumes Using RT Core Range Queries

Smoothed-particle hydrodynamics (SPH) is a mesh-free method used to simulate volumetric media in fluids, astrophysics, and solid mechanics. Visualizing these simulations is problematic because these datasets often contain millions, if not billions of particles carrying physical attributes and moving over time. Radial basis functions (RBFs) are used to model particles, and overlapping particles are interpolated to reconstruct a high-quality volumetric field; however, this interpolation process is expensive and makes interactive visualization difficult. Existing RBF interpolation schemes do not account for color-mapped attributes and are instead constrained to visualizing just the density field. To address these challenges, we exploit ray tracing cores in modern GPU architectures to accelerate scalar field reconstruction. We use a novel RBF interpolation scheme to integrate per-particle colors and densities, and leverage GPU-parallel tree construction and refitting to quickly update the tree as the simulation animates over time or when the user manipulates particle radii. We also propose a Hilbert reordering scheme to cluster particles together at the leaves of the tree to reduce tree memory consumption. Finally, we reduce the noise of volumetric shadows by adopting a spatially temporal blue noise sampling scheme. Our method can provide a more detailed and interactive view of these large, volumetric, time-series particle datasets than traditional methods, leading to new insights into these physics simulations.

Particle Volumes↗

When ancient numerical demons meet physics-informed machine learning: adjoint-based gradients for implicit differentiable modeling

Recent advances in differentiable modeling, a genre of physics-informed machine learning that trains neural networks (NNs) together with process-based equations, have shown promise in enhancing hydrological models' accuracy, interpretability, and knowledge-discovery potential. Current differentiable models are efficient for NN-based parameter regionalization, but the simple explicit numerical schemes paired with sequential calculations (operator splitting) can incur numerical errors whose impacts on models' representation power and learned parameters are not clear. Implicit schemes, however, cannot rely on automatic differentiation to calculate gradients due to potential issues of gradient vanishing and memory demand. Here we propose a “discretize-then-optimize” adjoint method to enable differentiable implicit numerical schemes for the first time for large-scale hydrological modeling. The adjoint model demonstrates comprehensively improved performance, with Kling–Gupta efficiency coefficients, peak-flow and low-flow metrics, and evapotranspiration that moderately surpass the already-competitive explicit model. Therefore, the previous sequential-calculation approach had a detrimental impact on the model's ability to represent hydrological dynamics. Furthermore, with a structural update that describes capillary rise, the adjoint model can better describe baseflow in arid regions and also produce low flows that outperform even pure machine learning methods such as long short-term memory networks. The adjoint model rectified some parameter distortions but did not alter spatial parameter distributions, demonstrating the robustness of regionalized parameterization. Despite higher computational expenses and modest improvements, the adjoint model's success removes the barrier for complex implicit schemes to enrich differentiable modeling in hydrology.

58 GEOSCIENCES↗

Fast HARDI Uncertainty Quantification and Visualization with Spherical Sampling

In this paper, we study uncertainty quantification and visualization of orientation distribution functions (ODF), which corresponds to the diffusion profile of high angular resolution diffusion imaging (HARDI) data. The shape inclusion probability (SIP) function is the state‐of‐the‐art method for capturing the uncertainty of ODF ensembles. The current method of computing the SIP function with a volumetric basis exhibits high computational and memory costs, which can be a bottleneck to integrating uncertainty into HARDI visualization techniques and tools. We propose a novel spherical sampling framework for faster computation of the SIP function with lower memory usage and increased accuracy. In particular, we propose direct extraction of SIP isosurfaces, which represent confidence intervals indicating spatial uncertainty of HARDI glyphs, by performing spherical sampling of ODFs. Our spherical sampling approach requires much less sampling than the state‐of‐the‐art volume sampling method, thus providing significantly enhanced performance, scalability, and the ability to perform implicit ray tracing. Our experiments demonstrate that the SIP isosurfaces extracted with our spherical sampling approach can achieve up to 8164× speedup, 37282× memory reduction, and 50.2% less SIP isosurface error compared to the classical volume sampling approach. We demonstrate the efficacy of our methods through experiments on synthetic and human‐brain HARDI datasets.

97 MATHEMATICS AND COMPUTING↗

Programmable Cryogenic Memory in a Ge/GeSi Heterostructure

Programmable memory components that operate optimally at cryogenic temperatures are essential for cryogenic computing architectures that seek to implement computing-in-memory. In this work, we demonstrate highly programmable memory in a Ge/GeSi heterostructure field-effect transistor (HFET). To operate, the HFET is gated to introduce positive carriers within the Ge quantum well, creating a high-conductance state. We show that this device can be set to a low-conductance state by sweeping a negative bias on the device drain, and reset it to its high-conductance state by sweeping a more positive bias on the device gate, thereby creating memory. We then determine that the device can be programmed within a 103 range of conductances using either the SET or the RESET operation. We propose that memory is achieved through charge trapping as carriers tunnel out of the quantum well, and that altering the density and spatial distribution of carriers modulates the device conductance. This mechanism exhibits endurance over 1000 cycles at temperatures ≤ 25 K, suggesting that the carrier traps are located at the oxide-semiconductor interface. As a first demonstration of programmable conductance in a Ge/GeSi HFET, this work highlights the potential of group-IV HFETs to perform as analog cryogenic memory components.

cryogenic memory↗

Protonic nickelate device networks for spatiotemporal neuromorphic computing

Computation in biological neural circuits arises from the interplay of nonlinear temporal responses and spatially distributed dynamic network interactions. Replicating this richness in hardware has remained challenging, as most neuromorphic devices emulate only isolated neuron- or synapse-like functions. Here we introduce an integrated neuromorphic computing platform in which both nonlinear spatiotemporal processing and programmable memory are realized within a single perovskite nickelate material system. By engineering symmetric and asymmetric hydrogenated NdNiO 3 junction devices on the same wafer, we combine ultrafast, proton-mediated transient dynamics with stable multilevel resistance states. Networks of symmetric NdNiO 3 junctions exhibit emergent spatial interactions mediated by proton redistribution, while each node simultaneously provides short-term temporal memory, enabling nanosecond-scale operation with an energy cost of ~0.2 nJ per input. When interfaced with asymmetric output units serving as reconfigurable long-term weights, these networks allow both feature transformation and linear classification in the same material system. Leveraging these emergent interactions, the platform enables real-time pattern recognition and achieves high accuracy in spoken digit classification and early seizure detection, outperforming temporal-only or uncoupled architectures. These results position protonic nickelates as a compact, energy-efficient, CMOS-compatible platform that integrates processing and memory for scalable intelligent hardware.

Electrical and electronic engineering↗

Online thermal profile prediction for large format additive manufacturing: A hybrid CNN-LSTM based approach

Large format additive manufacturing (LFAM) is an advanced 3D printing technique that efficiently fabricates large-scale components through a layer-by-layer extrusion and deposition process. Accurate surface layer temperature monitoring is essential to prevent manufacturing failures and ensure final product quality. Traditional physics-based offline approaches for simulating thermal behavior are often inefficient and complex, posing challenges on real-time, in-situ monitoring. Here, to address this, we propose a data-driven hybrid CNN-LSTM model to predict sequential thermal images of arbitrary length using real-time infrared thermal imaging. In this approach, a Convolutional Neural Networks (CNN) is trained offline to capture spatial features, reduce dimensional complexity, and enhance time efficiency, while a stacked Long Short-Term Memory (LSTM) is applied online to capture temporal information for improved prediction of future thermal behavior in subsequent printing layers. Model performance is evaluated using MSE, SSIM, and PSNR metrics and is benchmarked against stacked LSTM and convolutional LSTM models, demonstrating superior accuracy and applicability. Additionally, to mitigate noise from moving extruders and gantry backgrounds in thermal images, a fine-tuned semantic segmentation model is implemented offline to extract printing geometry, enabling precise temperature tracking along the tool path for further thermal analysis. The frameworks developed in this study significantly advance temperature monitoring, thermal analysis, and in-situ manufacturing control for LFAM, bridging the gap between theoretical modeling and practical application.

Geometry extraction↗

Smart material based multilayered microbeam structures for spatial self-deployment and reconfiguration: A residual stress approach

Alleviation of the potentially damaging effects induced by residual stresses was comprehensively investigated in previous research. Here, this paper, however, presents a spatially self-deployable and reconfigurable multilayered microbeam which takes advantage of residual stresses and shape memory effects. Reconfigurable mechanism of a typical four-layered microbeam composed of Pt\Ni 50 Ti 50 \Ni 50 Ti 50 \Pt is introduced, followed by analytical modeling of the maximum distance of the self-deployed gap as functions of variable structural and material parameters, including compressive residual stress in Pt layers and tensile residual stress in Ni 50 Ti 50 layers. Analytical solutions given by the static model agree well with the results obtained via finite element models (FEMs). Fabrication, characterization, and in-situ experiments were carried out to validate the feasibility of deployment of the as-released four-layered microbeam. The maximum distance of the gap was measured to be 41.39 μm at 20 °C, which could be increased to 51.73 μm thanks to controllable reconfiguration driven by shape memory effects. Theoretical analysis of such self-deployment and reconfiguration suggested a tensile residual stress increase by 52 MPa in Ni 50 Ti 50 layers. The multilayered microbeam structure with capabilities of self-deployment and reconfiguration offers great potential for various emerging applications, such as micro robotics, medical drug delivery devices, and intelligent chip scale spacecraft.

36 MATERIALS SCIENCE↗

Gravitational memory and Ward identities in the local detector frame

Gravitational memory, which describes the permanent shift in the strain after the passage of gravitational waves, is directly related to Weinberg’s soft graviton theorems and the Bondi-Metzner-Sachs (BMS) symmetry group of asymptotically flat space-times. In this work, we provide an equivalent description of the phenomenon in local coordinates around gravitational wave detectors, such as transverse-traceless (TT) gauge. We show that gravitational memory is encoded in large residual diffeomorphisms in this gauge, which include time-dependent anisotropic spatial rescalings, and prove their equivalence to BMS transformations when translated to TT gauge. We then derive the associated Ward identities and associated soft theorems, for both scattering amplitudes and equal-time (in-in) correlation functions, and explicitly check their validity for planar gravitational waves. Furthermore, the in-in identities are recognized as the flat-space analog of the well-known inflationary consistency relations.

General relativity↗

Spectral quadrature for the first principles study of crystal defects: Application to magnesium

In this work, we present an accurate and efficient finite-difference formulation and parallel implementation of Kohn-Sham Density (Operator) Functional Theory (DFT) for non periodic systems embedded in a bulk environment. Specifically, employing non-local pseudopotentials, local reformulation of electrostatics, and truncation of the spatial Kohn-Sham Hamiltonian, and the Linear Scaling Spectral Quadrature method to solve for the pointwise electronic fields in real-space and the non-local component of the atomic force, we develop a parallel finite difference framework suitable for distributed memory computing architectures to simulate non-periodic systems embedded in a bulk environment. Choosing examples from magnesium-aluminum alloys, we first demonstrate the convergence of energies and forces with respect to spectral quadrature polynomial order, and the width of the spatially truncated Hamiltonian. Next, we demonstrate the parallel scaling of our framework, and show that the computation time and memory scale linearly with respect to the number of atoms. Next, we use the developed framework to simulate isolated point defects and their interactions in magnesium-aluminum alloys. Our findings conclude that the binding energies of divacancies, Al solute-vacancy and two Al solute atoms are anisotropic and are dependent on cell size. Furthermore, the binding is favorable in all three cases.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

A hybrid CNN-LSTM surrogate model for hyper-resolution spatiotemporal flood forecasting in Norfolk, Virginia

Study region: Norfolk, Virginia, United States Study focus: Accurate and timely flood forecasting is essential for enhancing resilience in coastal urban areas in the context of increasing frequency and intensity of rainfall, sea level rise and urbanization. This study presents a hybrid deep learning-based surrogate model that integrates Convolutional Neural Networks (CNN) and Long Short-Term Memory (LSTM) networks to enable real-time spatiotemporal flood forecasting. The model leverages CNN to capture spatial features from inputs such as elevation and Topographic Wetness Index (TWI), while LSTM processes time-series inputs of rainfall and tide data to capture temporal features. New hydrologic insights for the region: The hybrid CNN-LSTM model was trained using the physics-based hydrodynamic model simulations obtained from the Two-dimensional Unsteady FLOW (TUFLOW) model for Norfolk, Virginia, and achieved high predictive accuracy across diverse flood-prone areas. The reduced computational time from four to six hours using TUFLOW to 3.2 min per event using CNN-LSTM enables rapid flood inundation mapping and early warning applications. The model effectively captured both spatial flood extents and their temporal evolution across different flooding scenarios, providing forecasts at a 2.5-m spatial resolution and 15-min temporal resolution and a one-hour-ahead prediction horizon. While challenges remain in terms of transferability to new regions and real-time data assimilation, this approach demonstrates strong potential for supporting operational flood risk management in coastal urban environments.

Coastal urban flooding↗