Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “spatial memory”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Uncertainty-aware Continuous Implicit Neural Representations for Remote Sensing Object Counting

Many existing object counting methods rely on density map estimation (DME) of the discrete grid representation by decoding extracted image semantic features from designed convolutional neural networks (CNNs). Relying on discrete density maps not only leads to information loss dependent on the original image resolution, but also has a scalability issue when analyzing high-resolution images with cubically increasing memory complexity. Furthermore, none of the existing methods can offer reliable uncertainty quantification (UQ) for the derived count estimates. To overcome these limitations, we design UNcertainty-aware, hypernetwork-based Implicit neural representations for Counting (UNIC) to assign probabilities and the corresponding counting confidence over continuous spatial coordinates. We derive a sampling-based Bayesian counting loss function and develop the corresponding model training algorithm. UNIC outperforms existing methods on the Remote Sensing Object Counting (RSOC) dataset with reliable UQ and improved interpretability of the derived count estimates. Our code is available at https://github.com/SiyuanXu-tamu/UNIC.

97 MATHEMATICS AND COMPUTING↗

Domain decomposition in the GPU-accelerated Shift Monte Carlo code

The GPU solver within the Shift continuous-energy Monte Carlo neutron transport code has been extended to provide domain decomposition in addition to domain replication to enable the solution of problems with memory requirements exceeding the capacity of a single GPU. The strategy follows the Multiple Set, Overlapping Domain (MSOD) approach that is used in Shift’s CPU solver and integrates into the event-based algorithm used for Shift’s GPU solver. Furthermore, the ability to assign processors to spatial domains non-uniformly has been maintained. In this work, two different approaches for communicating particle data between domains are considered, and multiple criteria for load balancing problems have been investigated. Numerical results are presented for both fresh and depleted small modular nuclear reactor (SMR) cores. A parallel efficiency of approximately 80% was achieved with up to 16 spatial domains measured relative to full domain replication. A scaling study on the Summit supercomputer demonstrates a weak scaling parallel efficiency of over 90% on over 24000 GPUs.

97 MATHEMATICS AND COMPUTING↗

Selective area doping for Mott neuromorphic electronics

The cointegration of artificial neuronal and synaptic devices with homotypic materials and structures can greatly simplify the fabrication of neuromorphic hardware. We demonstrate experimental realization of vanadium dioxide (VO 2 ) artificial neurons and synapses on the same substrate through selective area carrier doping. By locally configuring pairs of catalytic and inert electrodes that enable nanoscale control over carrier density, volatility or nonvolatility can be appropriately assigned to each two-terminal Mott memory device per lithographic design, and both neuron- and synapse-like devices are successfully integrated on a single chip. Feedforward excitation and inhibition neural motifs are demonstrated at hardware level, followed by simulation of network-level handwritten digit and fashion product recognition tasks with experimental characteristics. Spatially selective electron doping opens up previously unidentified avenues for integration of emerging correlated semiconductors in electronic device technologies.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Impact of Random Spatial Fluctuation in Non-Uniform Crystalline Phases on the Device Variation of Ferroelectric FET

In this work, a comprehensive study of random spatial fluctuation of the ferroelectric (FE) phase and dielectric (DE) phase in FeFETs is conducted to understand its impact on device variation. It is found that: i) there exists a certain DE percentage threshold that below which the increase of the DE phase does not significantly impact the device memory window and variation and only above which evident device degradation can be observed; ii) increasing the DE phase increases the variation in the memory window and the coercive field distribution further exacerbates the variation, hence degrading the sensing margin; iii) decreasing the number of grains degrades the device variation, which calls for further grain size engineering for variation suppression.

42 ENGINEERING↗

Mechanisms enabling reconfigurability and long-term retention in vanadium oxide electrochemical memory

Phase coexistence in nanoscale electrochemical random-access memory (ECRAM) has recently been demonstrated to enable both information storage and extraordinary reconfigurability. These proof-of-principle demonstrations have left the mechanistic details of such a process unresolved. Particularly, the mechanisms that stabilize the multiple phases, and the underlying processes behind sustained memory retention, remain unclear, and are necessary to design such devices. Here we report microscale ECRAM devices composed of V⁢O𝑥, which enables us to directly probe the active region in an operando fashion using optical techniques. Using Raman mapping, we show the phase coexistence driven by the electrochemical injection of O vacancies to be spatially uniform (i.e., with no filaments). The stability was observed to be unusually long, with 1% loss over 14 years in ambient conditions. First-principles calculations of the oxygen vacancy formation energies in V⁢O 𝑥 further support the thermodynamic coexistence of multiple V⁢O 𝑥 phases and clarify the origin of the observed long-term retention in the ECRAM devices. Further, we demonstrate single devices that can be voltage programmed to exhibit synaptic, neuronal, and reconfigurable logic gate functionalities. Furthermore, we not only uncover the phase coexistence mechanism that may help device design, but also demonstrate the circuit-level applications of reconfigurability.

Electrical conductivity↗

Phase‐Change‐Memory Process at the Limit: A Proposal for Utilizing Monolayer Sb 2 Te 3

Abstract One central task of developing nonvolatile phase change memory (PCM) is to improve its scalability for high‐density data integration. In this work, by first‐principles molecular dynamics, to date the thinnest PCM material possible (0.8 nm), namely, a monolayer Sb 2 Te 3 , is proposed. Importantly, its SET (crystallization) process is a fast one‐step transition from amorphous to hexagonal phase without the usual intermediate cubic phase. An increased spatial localization of electrons due to geometrical confinement is found to be beneficial for keeping the data nonvolatile in the amorphous phase at the 2D limit. The substrate and superstrate can be utilized to control the phase change behavior: e.g., with passivated SiO 2 (001) surfaces or hexagonal Boron Nitride, the monolayer Sb 2 Te 3 can reach SET recrystallization in 0.54 ns or even as fast as 0.12 ns, but with unpassivated SiO 2 (001), this would not be possible. Besides, working with small volume PCM materials is also a natural way to lower power consumption. Therefore, the proposed PCM working process at the 2D limit will be an important potential strategy of scaling the current PCM materials for ultrahigh‐density data storage.

2D limit↗

The separate effect of halo mass and stellar mass on the evolution of massive disc galaxies

ABSTRACT We analyse a sample of massive disc galaxies selected from the fourth-generation Sloan Digital Sky Survey/Mapping Nearby Galaxies at Apache Point Observatory survey to investigate how the evolution of these galaxies depends on their stellar and halo masses. We applied a semi-analytic spectral fitting approach to the data from different regions in the galaxies to derive several of their key physical properties. From the best-fitting model results, together with direct observables such as morphology, colour, and the Mgb/〈Fe〉 index ratio measured within 1Re, we find that for central galaxies both their stellar and halo masses have a significant influence in their evolution. For a given halo mass, galaxies with higher stellar mass accumulate their stellar mass and become chemically enriched earlier than those with smaller stellar mass. Furthermore, at a given stellar mass, galaxies living in more massive haloes have longer star formation time-scales and are delayed in becoming chemically enriched. In contrast, the evolution of massive satellite galaxies is mostly determined by their stellar mass. The results indicate that both the assembled halo mass and the halo assembly history impact the evolution of central galaxies. Our spatially resolved analysis indicates that only the galaxy properties in the central region (0.0–0.5Re) show the dependencies described above. This fact supports a halo-driven formation scenario since the galaxies’ central regions are more likely to contain old stars formed along with the halo itself, keeping a memory of the halo formation process.

Zhou, Shuang (ORCID:0000000289996814)↗

A Predictor-Corrector Strategy for Adaptivity in Dynamical Low-Rank Approximations

Here, in this paper, we present a predictor-corrector strategy for constructing rank-adaptive, dynamical low-rank approximations (DLRAs) of matrix-valued ODE systems. The strategy is a compromise between (i) low-rank step-truncation approaches that alternately evolve and compress solutions and (ii) strict DLRA approaches that augment the low-rank manifold using subspaces generated locally in time by the DLRA integrator. The strategy is based on an analysis of the error between a forward temporal update into the ambient full-rank space, which is typically computed in a step-truncation approach before recompressing, and the standard DLRA update, which is forced to live in a low-rank manifold. We use this error, without requiring its full-rank representation, to correct the DLRA solution. A key ingredient for maintaining a low-rank representation of the error is a randomized SVD, which introduces some degree of stochastic variability into the implementation. The strategy is formulated and implemented in the context of discontinuous Galerkin spatial discretizations of PDEs and applied to several versions of DLRA methods found in the literature as well as a new variant. Numerical experiments comparing the predictor-corrector strategy to other methods demonstrate robustness to overcome shortcomings of step truncation or strict DLRA approaches: The former may require more memory than is strictly needed, while the latter may miss transients solution features that cannot be recovered. The effect of randomization, tolerances, and other implementation parameters is also explored.

97 MATHEMATICS AND COMPUTING↗

Automated Hybrid Variance Reduction on Advanced Architectures in the Shift Monte Carlo Code

Monte Carlo transport methods are the most accurate schemes for solving problems with complex energy and spatial features, but they come with a high computational cost. Although hybrid methods have enabled the use of Monte Carlo transport for a large class of problems, they still require significant computing resources. Modern multicore CPUs with large numbers of compute cores and graphical processing units (GPUs) provide opportunities to optimize the memory and run-time costs of hybrid Monte Carlo methods. This paper documents the development and analysis of three Monte Carlo transport algorithms that support hybrid transport using the consistent adjoint-driven importance sampling (CADIS) and forward-weighted CADIS methods in the Shift Monte Carlo code: history-based transport using static and dynamic threading on multicore CPUs and event-based transport enabling weight window tracking on GPUs. The results are shown for two challenging hybrid problems on the Frontier supercomputer at the Oak Ridge Leadership Computing Facility. The results show that all three methods yield good performance and enable solutions of difficult fixed-source transport problems in less than 2 min on 20 nodes of Frontier. Dynamic threading was observed to give up to 20% better scaling behavior than static threading. Moreover, the AMD Instinct 250X GPU was found to give 9 to 11 times greater throughput per graphics compute die than the best CPU performance. In conclusion, additional opportunities for optimization of hybrid transport on GPUs are discussed.

Denovo↗

Hardware-Based Emulator with Deep Learning Model for Building Energy Control and Prediction Based on Occupancy Sensors’ Data

Heating, ventilation, and air conditioning (HVAC) is the largest source of residential energy consumption. Occupancy sensors’ data can be used for HVAC control since it indicates the number of people in the building. HVAC and sensors form a typical cyber-physical system (CPS). In this paper, we aim to build a hardware-based emulation platform to study the occupancy data’s features, which can be further extracted by using machine learning models. In particular, we propose two hardware-based emulators to investigate the use of wired/wireless communication interfaces for occupancy sensor-based building CPS control, and the use of deep learning to predict the building energy consumption with the sensor data. We hypothesize is that the building energy consumption may be predicted by using the occupancy data collected by the sensors, and question what type of prediction model should be used to accurately predict the energy load. Another hypothesis is that an in-lab hardware/software platform could be built to emulate the occupancy sensing process. The machine learning algorithms can then be used to analyze the energy load based on the sensing data. To test the emulator, the occupancy data from the sensors is used to predict energy consumption. The synchronization scheme between sensors and the HVAC server will be discussed. We have built two hardware/software emulation platforms to investigate the sensor/HVAC integration strategies, and used an enhanced deep learning model—which has sequence-to-sequence long short-term memory (Seq2Seq LSTM)—with an attention model to predict the building energy consumption with the preservation of the intrinsic patterns. Because the long-range temporal dependencies are captured, the Seq2Seq models may provide a higher accuracy by using LSTM architectures with encoder and decoder. Meanwhile, LSTMs can capture the temporal and spatial patterns of time series data. The attention model can highlight the most relevant input information in the energy prediction by allocating the attention weights. The communication overhead between the sensors and the HVAC control server can also be alleviated via the attention mechanism, which can automatically ignore the irrelevant information and amplify the relevant information during CNN training. Our experiments and performance analysis show that, compared with the traditional LSTM neural network, the performance of the proposed method has a 30% higher prediction accuracy.

Ye, Zhijing↗

Predicting weather impacts on corn production in a data-limited region using a transfer learning approach

The stability of food supply and prices may depend more on annual changes in yields from year-to-year variability in weather than on longer-term average changes from changing climatic conditions. However, the absence of high-quality data on crop yields at fine spatial resolutions in many regions of the world makes it challenging to statistically model their response to interannual variability in weather patterns. Therefore, there is a need for empirical methods that can project annual crop yield changes even in limited data regions. Here, we propose a transfer learning algorithm that uses high spatial resolution data from one region to project yields in another region with more limited data. The goal of our work is to understand what data types can be beneficial for transferring learning from a source region to a very different target region with more limited data. We utilize Long Short-Term Memory to develop a transfer learning model that is trained on historical county-level corn yield in the United States and predicts district-level corn yield variations in India. Even using smaller amounts of data in India, simulating a data-scarce region, we achieve an average root mean square error of 0.48 bu acre−1 in predicting interannual yield variations. Using Shapley values to interpret results, we explore the contribution of the different weather parameters to interannual yield variability and find a larger influence of precipitation-related variables. Our study demonstrates the usefulness of this method for transferring models of weather impacts on crop yields trained on a data-rich country to one with more limited data. It suggests the potential of applying the transfer learning model to mitigate the need for extensive raw data globally.

Vishwakarma, Srishti [ORNL] (ORCID:000000031674419↗

Scalable multiscale modeling of platelets with 100 million particles

Here, we developed the core components of the AI-aided multiple time stepping algorithm for multiscale modeling of cell dynamics. This algorithm was implemented and analyzed on two supercomputer architectures with an application of simulating the aggregation of 250 platelets, or 102 million particles. To scale on these computers with complex memory and network architectures with GPUs, we devised a biomechanics-informed task mapping scheme to optimize load imbalance, communications, and memory utilization. Our simulations, scaling well up to 192 nodes on a Summit-like supercomputer with a peak speed of 11 petaflops, achieved a rate of 423 μs/day which is 500 times faster than the conventional algorithm using static time step and this has enabled studies of record size blood clots at record spatial–temporal resolutions. Additionally, we discovered the sensitive dependence of the scalability and execution time on the methods of decomposition, CPU–GPU coupling, and task mapping.

97 MATHEMATICS AND COMPUTING↗

Improving deep learning performance for predicting large-scale geological ${{CO}_{2}}$ sequestration modeling through feature coarsening

Physics-based reservoir simulation for fluid flow in porous media is a numerical simulation method to predict the temporal-spatial patterns of state variables (e.g. pressure p) in porous media, and usually requires prohibitively high computational expense due to its non-linearity and the large number of degrees of freedom (DoF). This work describes a deep learning (DL) workflow to predict the pressure evolution as fluid flows in large-scale 3-dimensional(3D) heterogeneous porous media. In particular, we develop an efficient feature coarsening technique to extract the most representative information and perform the training and prediction of DL at the coarse scale, and further recover the resolution at the fine scale by spatial interpolation. We validate the DL approach to predict pressure field against physics-based simulation data for a field-scale 3D geologic CO 2 sequestration reservoir model. We evaluate the impact of feature coarsening on DL performance, and observe that the feature coarsening not only decreases the training time by >74% and reduces the memory consumption by >75%, but also maintains temporal error 0.63% on average. Besides, the DL workflow provides predictive efficiency with 1406 times speedup compared to physics-based numerical simulation. The key findings from this research significantly improve the training and prediction efficiency of deep learning model to deal with large-scale heterogeneous reservoir models, and thus it can also be further applied to accelerate workflows of history matching and reservoir optimization for close-loop reservoir management.

58 GEOSCIENCES↗

An a priori evaluation of a principal component and artificial neural network based combustion model in diesel engine conditions

A principal component analysis (PCA) and artificial neural network (ANN) based chemistry tabulation approach is presented. ANNs are used to map the thermochemical state onto a low-dimensional manifold consisting of five control variables that have been identified using PCA. Three canonical configurations are considered to train the PCA-ANN model: a series of homogeneous reactors, a nonpremixed flamelet, and a two-dimensional lifted flame. The performance of the model in predicting the thermochemical manifold of a spatially-developing turbulent jet flame in diesel engine thermochemical conditions is a priori evaluated using direct numerical simulation (DNS) data. The PCA-ANN approach is compared with a conventional tabulation approach (tabulation using ad hoc defined control variables and linear interpolation). The PCA-ANN model provides higher accuracy and requires several orders of magnitude less memory. Here, these observations indicate that the PCA-ANN model is superior for chemistry tabulation, especially for modelling complex chemistries that present multiple combustion modes as observed in diesel combustion. The performance of the PCA-ANN model is then compared to the optimal estimator, i.e. the conditional mean from the DNS. The results indicate that the PCA-ANN model gives high prediction accuracy, comparable to the optimal estimator, especially for major species and the thermophysical properties. Higher errors are observed for the minor species and reaction rate predictions when compared to the optimal estimator. It is shown that the prediction of minor species and reaction rates can be improved by using training data that exhibits a variation of parameters as observed in the turbulent flame. The output of the ANN is analysed to assess mass conservation. It is observed that the ANN incurs a mean absolute error of 0.05% in mass conservation. Furthermore, it is demonstrated that this error can be reduced by modifying the cost function of the ANN to penalise for deviation from mass conservation.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Unified many-worlds browsing of arbitrary physics-based animations

Manually tuning physics-based animation parameters to explore a simulation outcome space or achieve desired motion outcomes can be notoriously tedious. This problem has motivated many sophisticated and specialized optimization-based methods for fine-grained (keyframe) control, each of which are typically limited to specific animation phenomena, usually complicated, and, unfortunately, not widely used. In this paper, we propose Unified Many-Worlds Browsing (UMWB), a practical method for sample-level control and exploration of physics-based animations. Our approach supports browsing of large simulation ensembles of arbitrary animation phenomena by using a unified volumetric WORLDPACK representation based on spatiotemporally compressed voxel data associated with geometric occupancy and other low-fidelity animation state. Beyond memory reduction, the WORLDPACK representation also enables unified query support for interactive browsing: it provides fast evaluation of approximate spatiotemporal queries, such as occupancy tests that find ensemble samples ("worlds") where material is either IN or NOT IN a user-specified spacetime region. WORLDPACKS also support real-time hardware-accelerated voxel rendering by exploiting the spatially hierarchical and temporal RLE raster data structure. Our UMWB implementation supports interactive browsing (and offline refinement) of ensembles containing thousands of simulation samples, and fast spatiotemporal queries and ranking. We show UMWB results using a wide variety of physics-based animation phenomena---not just JELL-O ® .

Computer Science↗

Uncertainty-Guided Prediction Horizon of Phase-Resolved Ocean Wave Forecasting Under Data Sparsity: Experimental and Numerical Evaluation

Accurate short-term wave forecasting is critical for the safe and efficient operation of marine structures that rely on real-time, phase-resolved ocean wave information for control and monitoring purposes (e.g., digital twins). These systems often depend on environmental sensors (e.g., waverider buoys, wave-sensing LIDAR). Challenges arise when upstream sensor data are missing, sparse, or phase-shifted due to drift. This study investigates the performance of two machine learning models, time-series dense encoder (TiDE) and long short-term memory (LSTM), for forecasting phase-resolved ocean surface elevations under varying degrees of data degradation. We introduce the τ-trimming algorithm, which adapts the prediction horizon based on uncertainty thresholds derived from historical forecasts. Numerical wave tank (NWT) and wave basin experiments are used to benchmark model performance under short- and long-term data masking, spatially coarse sensor grids, and upstream phase shifts. Results show under a 50% probability of upstream data loss, the τ-trimmed TiDE model achieves a 46% reduction in error at the most upstream target, compared to 22% for LSTM. Furthermore, phase misalignment in upstream data introduces a near-linear increase in forecast error. Under moderate model settings, a ±3 s misalignment increases the mean absolute error by approximately 0.5 m, while the same error is accumulated at ±4 s using the more conservative approach. These findings inform the design of resilient, uncertainty-aware wave forecasting systems suited for realistic offshore sensing environments.

42 ENGINEERING↗

Skyrmions in an oblique field

Here we study skyrmion states induced by the oblique field in a square magnetic island. To describe stable states and dynamics of the skyrmion, we employ the lattice model that accounts for all relevant interactions and uses parameters in the range of practical interest. Depending on the orientation of the field, two or four spatially separated energy minima emerge in the oblique field. The depth of the energy minima can be controlled by the strength and orientation of the magnetic field. We study the dynamics of the skyrmion and show that it can be moved between any two states by the field gradient. Possible applications for developing binary and quaternary skyrmion-based memory are discussed.

36 MATERIALS SCIENCE↗

Spatiotemporal Thermal Coupling in VO 2 Device Arrays

Correlated oxides such as VO 2 exhibit an electrically driven insulator–metal transition (IMT) that underlies their promise for neuromorphic and memory devices. Yet the IMT is not a uniform bulk process but a spatiotemporal phenomenon in which local heating nucleates filaments, contracts or dissolves them with the electric field, and couples to the environment. In this work, we directly image the VO 2 IMT dynamics by mid-wave infrared, thermography synchronized with electrical transport, resolving device temperature with micrometer spatial and microsecond temporal resolution. At the single-device level, we capture the full cycle of filament nucleation, contraction, and relaxation during current/voltage-driven resistive switching. At the array level, we show that heat propagates across etched gaps with an effective length scale of ∼131 µm, enabling cooperative behaviors among electrically isolated devices. Short-range distanced devices exhibit mutual filament attraction and sequential dissolution, while long-range distanced devices differentiate into distinct roles: drivers that initiate switching, cooperative responders that undergo assisted self-oscillations, and passive reporters that record the thermal field. Furthermore, these results reframe thermal crosstalk, long regarded as parasitic, as an intrinsic coupling channel and design principle for organizing collective switching behaviors, with direct implications for emergent circuit functionality in neuromorphic and unconventional computing architectures.

coupling↗