Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “spatial memory”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Autonomous Multistate Nanoencoding Using Combinatorial Ferroelectric Closure Domains in BiFeO 3

Recent advances in ferroic materials have identified topological defects as promising candidates for enabling additional functionalities in future electronic systems. The generation of stable and customizable polar topologies is needed to achieve multistates that enable beyond-binary device architectures. Here, in this study, we show how to autonomously pattern on-demand highly tunable striped closure domains in pristine rhombohedral-phase BiFeO 3 thin films through precise scanning of a biased atomic force microscopy tip along carefully designed paths. By employing this strategy, we generate and manipulate closed-loop structures with high spatial resolution in an automated manner, allowing the creation of highly tunable and intricate topological domain structures that exhibit distinct polarization configurations without the need for electrode deposition or complex heterostructure growth. As a proof-of-concept for ferroelectric beyond-binary memory devices, we use such topological domains as multistates, engineering an alphabet and automating the symbolic writing/reading process using autonomous microscopy. The resulting information density is compared with that of current commercially available memory devices, demonstrating the potential of ferroelectric topological domains for multistate information storage applications.

BiFeO3↗

Ferroelectric Fractals: Switching Mechanism of Wurtzite AlN

The advent of wurtzite ferroelectrics is enabling new ferroelectric devices for computer memory that have the potential to bypass the von Neumann bottleneck due to their robust polarization and silicon compatibility. However, the atomistic switching mechanism of wurtzites is still undetermined due to the limitations of density functional theory simulation size and experimental temporal and spatial resolution. Thus, physics-informed materials engineering to reduce coercive field and breakdown in these devices has been limited. In this work, the atomistic mechanism of domain wall migration and domain growth in aluminum nitride-based wurtzites is uncovered using molecular dynamics and Monte Carlo simulations. We reveal the anomalous switching mechanism of fast 1D single columns of atoms propagating from a slow-moving 2D fractallike domain wall. We find that the critical nucleus is a single aluminum ion that breaks its bond with one nitrogen and bonds to another nitrogen; this creates a cascade that flips atoms directly only in the same column, due to the extreme locality (sharpness) of the domain walls in wurtzites. We further show how the fractallike shape of the domain wall in the 2D plane breaks assumptions in the Kolmogorov, Avrami, and Ishibashi (KAI) model and leads to the anomalously fast switching in wurtzite structured ferroelectrics.

36 MATERIALS SCIENCE↗

Design of a Variational Multiscale Method for Turbulent Compressible Flows

A spectral-element framework is presented for the simulation of subsonic compressible high-Reynolds-number flows. The focus of the work is maximizing the efficiency of the computational schemes to enable unsteady simulations with a large number of spatial and temporal degrees of freedom. A collocation scheme is combined with optimized computational kernels to provide a residual evaluation with computational cost independent of order of accuracy up to 16th order. The optimized residual routines are used to develop a low-memory implicit scheme based on a matrix-free Newton-Krylov method. A preconditioner based on the finite-difference diagonalized ADI scheme is developed which maintains the low memory of the matrix-free implicit solver, while providing improved convergence properties. Emphasis on low memory usage throughout the solver development is leveraged to implement a coupled space-time DG solver which may offer further efficiency gains through adaptivity in both space and time.

Design↗

Electrode and Microstructure Dependence of Oxygen Diffusion in Ferroelectric Hafnium Zirconium Oxide Thin Films

Hafnia-based ferroelectrics hold promise to reduce energy demand for computing by enabling compute-in-memory and as non-volatile memories. The ferroelectric phase in this material system is, in part, stabilized by oxygen vacancies. While oxygen vacancies may be a necessity for phase stability, they limit device endurance through diffusion and accumulation into conducting channels. Herein, it is shown that oxygen diffusion is spatially variable within individual grains of ferroelectric hafnium zirconium oxide (HZO). Using 18 O tracers and finite difference modeling, it is shown that grain boundaries and regions near electrode interfaces allow for relatively rapid oxygen diffusion, with values as much as 10 4 larger than the grain cores. Further, the selection of electrode material affects the diffusion coefficients across all microstructural regions. HZO films in contact with TiN electrodes result in more oxygen-deficient HZO films and higher oxygen diffusion coefficients. Tungsten electrodes result in fewer vacancies and lower diffusion coefficients. Diffusion activation energy differences between the HZO with the two electrodes is reconciled by differing populations of charged and uncharged oxygen vacancies. This insight into the local vacancy populations and diffusion pathways provides a platform for designing hafnia-based films, deposition processes, and integration strategies to reduce vacancy gradients and improve performance.

36 MATERIALS SCIENCE↗

Models and error analyses in urban air quality estimation

Estimation theory has been applied to a wide range of aerospace problems. Application of this expertise outside the aerospace field has been extremely limited, however. This paper describes the use of covariance error analysis techniques in evaluating the accuracy of pollution estimates obtained from a variety of concentration measuring devices. It is shown how existing software developed for aerospace applications can be applied to the estimation of pollution through the processing of measurement types involving a range of spatial and temporal responses. The modeling of pollutant concentration by meandering Gaussian plumes is described in some detail. Time averaged measurements are associated with a model of the average plume, using some of the same state parameters and thus avoiding the problem of state memory. The covariance analysis has been implemented using existing batch estimation software. This usually involves problems in handling dynamic noise; however, the white dynamic noise has been replaced by a band-limited process which can be easily accommodated by the software.

Englar, T., Jr.↗

Accuracy Enhancement of Nuclear Power Plant Simulators Utilizing High Accuracy Simulation Predictions

More recently, reactor core simulators for core designs associated with commercial nuclear power plants that utilize what is believed to be higher fidelity models have been developed. Features such as neutronics models that utilize transport equation solvers with fine spatial meshes and many energy-groups, thermal-hydraulic models that utilize sub-channel solvers with fine spatial mesh and capable of treating a wide range of fluid conditions, and fuel-coolant chemistry interaction models capable of treating CRUD deposition are to be found in these higher fidelity core simulators. These reactor core simulators require access to higher performance computers, characterized by many processors, cores and large memory. So associated with utilization of these simulators is access to high performance computers and ability to accommodate in one’s workflow longer execution times. By contrast, currently used core simulators by the nuclear industry can execute on engineering workstations and have execution times of seconds to minutes. The desirability for having short execution times is not only desired for support of time critical tasks but supports the mental process of decision making by engineers. The goal of the work reported upon here has the objective of retaining the fidelity of higher fidelity models while retaining the ability to utilize engineering workstations. Beyond the core simulator goal, additional goals of this work include incorporating the just described core simulator capability into a Nuclear Steam Supply System (NSSS) simulator, and to incorporate the resulting capability into an environment supportive of design and operational decision making associated with nuclear power stations. The model selected for the core neutronics model is the NESTLE code, for the core thermal-hydraulic model is the CTF code utilizing coarse mesh, and for the NSSS model is the RELAP5-3D code. WSC’s proprietary 3KEYMASTERTM platform is being used to provide software coupling, user interface, visualization, and reporting. The NESTLE core neutronics simulator was first integrated with the CTF core thermal-hydraulic simulator using CTF developed communication commands which are also used for CTF to communicate with RELAP5-3D under WSC’s proprietary 3KEYMASTERTM platform. To assure NESTLE prediction consistency with higher fidelity core neutronic simulators, buffer codes have been created to automatically generate from output files written by the VERA core simulator the NESTLE nodal neutronic parameter’ library, geometry, and pin-power reconstruction input files, thereby avoiding a number of challenges associated with utilizing lattice physics codes and providing consistency with VERA predictions. To treat absorber rod effects a multi-set library is utilized, where a set refers to a specific absorber rod fully inserted pattern. A coarse spatial mesh CTF model was developed with features added that support using CTF as envisioned in the engineering quality simulator. A hybrid meshing approach was implemented to allow for automated construction of models with mixed levels of refinement. Specifically, a core model could resolve some assemblies at a nodal level (4 subchannels per assembly) and others at a pin-resolution (one subchannel per coolant subchannel in the assembly). The intention is that this will allow for better resolution of limiting conditions such as DNBR and PCT, which are based on local rod and subchannel conditions. Further development was done of features that enhance the capabilities for the envisioned engineering quality simulator that has been developed, but now for RELAP-3D. The RELAP5-3D code development includes ability to model more than 999 components and the addition of the cross-channels turbulence mixing model and the void drift model that are implemented in CTF, aiming to achieve closer prediction agreement of the two codes for transient simulations, specifically, more accurate matches of the overall mass, momentum, and energy exchanges of both the liquid and gas phases between the neighboring core assemblies. Graphics were also developed for the Instructor Station for this project under WSC’s proprietary 3KEYMASTERTM platform to facilitate design and operational decision making.

42 ENGINEERING↗

Outline for a theory of intelligence

Intelligence is defined as that which produces successful behavior. Intelligence is assumed to result from natural selection. A model is proposed that integrates knowledge from research in both natural and artificial systems. The model consists of a hierarchical system architecture wherein: (1) control bandwidth decreases about an order of magnitude at each higher level, (2) perceptual resolution of spatial and temporal patterns contracts about an order-of-magnitude at each higher level, (3) goals expand in scope and planning horizons expand in space and time about an order-of-magnitude at each higher level, and (4) models of the world and memories of events expand their range in space and time by about an order-of-magnitude at each higher level. At each level, functional modules perform behavior generation (task decomposition planning and execution), world modeling, sensory processing, and value judgment. Sensory feedback control loops are closed at every level.

Albus, James S.↗

Runge-Kutta Methods for Linear Ordinary Differential Equations

Three new Runge-Kutta methods are presented for numerical integration of systems of linear inhomogeneous ordinary differential equations (ODES) with constant coefficients. Such ODEs arise in the numerical solution of the partial differential equations governing linear wave phenomena. The restriction to linear ODEs with constant coefficients reduces the number of conditions which the coefficients of the Runge-Kutta method must satisfy. This freedom is used to develop methods which are more efficient than conventional Runge-Kutta methods. A fourth-order method is presented which uses only two memory locations per dependent variable, while the classical fourth-order Runge-Kutta method uses three. This method is an excellent choice for simulations of linear wave phenomena if memory is a primary concern. In addition, fifth- and sixth-order methods are presented which require five and six stages, respectively, one fewer than their conventional counterparts, and are therefore more efficient. These methods are an excellent option for use with high-order spatial discretizations.

Zingg, David W.↗

SHF: Symmetrical Hierarchical Forest with Pretrained Vision Transformer Encoder for High-Resolution Medical Segmentation

This paper presents a novel approach to addressing the long-sequence problem in high-resolution medical images for Vision Transformers (ViTs). Using smaller patches as tokens can enhance ViT performance, but quadratically increases computation and memory requirements. Therefore, the common practice for applying ViTs to high-resolution images is either to: (a) employ complex sub-quadratic attention schemes or (b) use large to medium-sized patches and rely on additional mechanisms within the model to capture the spatial hierarchy of details. We propose Symmetrical Hierarchical Forest (SHF), a lightweight approach that adaptively patches the input image to increase token information density and encode hierarchical spatial structures into the input embedding. We then apply a reverse depatching scheme to the output embeddings of the transformer encoder, eliminating the need for convolution-based decoders. Unlike previous methods that modify attention mechanisms or use a complex hierarchy of interacting models, SHF can be retrofitted to any ViT model to allow it to learn the hierarchical structure of details in high-resolution images without requiring architectural changes. Experimental results demonstrate significant gains in computational efficiency and performance: on the PAIP WSI dataset, we achieved a 3∼32×speedup or a 2.95%∼7.03% increase in accuracy (measured by Dice score) at a 64K2 resolution with the same computational budget, compared to state-of-the-art production models. On the 3D medical datasets BTCV and KiTS, training was 6×faster, with accuracy gains of 6.93% and 5.9%, respectively, compared to models without SHF.

Zhang, Enzhi [Hokkaido University, Japan]↗

A fast, matrix-based method to perform omnidirectional pressure integration

Abstract Experimentally-measured pressure fields play an important role in understanding many fluid dynamics problems. Unfortunately, pressure fields are difficult to measure directly with non-invasive, spatially resolved diagnostics, and calculations of pressure from velocity have proven sensitive to error in the data. Omnidirectional line integration methods are usually more accurate and robust to these effects as compared to implicit Poisson equations, but have seen slower uptake due to the higher computational and memory costs, particularly in 3D domains. This paper demonstrates how omnidirectional line integration approaches can be converted to a matrix inversion problem. This novel formulation uses an iterative approach so that the boundary conditions are updated each step, preserving the convergence behavior of omnidirectional schemes while also keeping the computational efficiency of Poisson solvers. This method is implemented in Matlab and also as a GPU-accelerated code in CUDA-C++. The behavior of the new method is demonstrated on 2D and 3D synthetic and experimental data. Three-dimensional grid sizes of up to 125 million grid points are tractable with this method, opening exciting opportunities to perform volumetric pressure field estimation from 3D PIV measurements.

42 ENGINEERING↗

Uncertainty-aware Continuous Implicit Neural Representations for Remote Sensing Object Counting

Many existing object counting methods rely on density map estimation (DME) of the discrete grid representation by decoding extracted image semantic features from designed convolutional neural networks (CNNs). Relying on discrete density maps not only leads to information loss dependent on the original image resolution, but also has a scalability issue when analyzing high-resolution images with cubically increasing memory complexity. Furthermore, none of the existing methods can offer reliable uncertainty quantification (UQ) for the derived count estimates. To overcome these limitations, we design UNcertainty-aware, hypernetwork-based Implicit neural representations for Counting (UNIC) to assign probabilities and the corresponding counting confidence over continuous spatial coordinates. We derive a sampling-based Bayesian counting loss function and develop the corresponding model training algorithm. UNIC outperforms existing methods on the Remote Sensing Object Counting (RSOC) dataset with reliable UQ and improved interpretability of the derived count estimates. Our code is available at https://github.com/SiyuanXu-tamu/UNIC.

97 MATHEMATICS AND COMPUTING↗

Domain decomposition in the GPU-accelerated Shift Monte Carlo code

The GPU solver within the Shift continuous-energy Monte Carlo neutron transport code has been extended to provide domain decomposition in addition to domain replication to enable the solution of problems with memory requirements exceeding the capacity of a single GPU. The strategy follows the Multiple Set, Overlapping Domain (MSOD) approach that is used in Shift’s CPU solver and integrates into the event-based algorithm used for Shift’s GPU solver. Furthermore, the ability to assign processors to spatial domains non-uniformly has been maintained. In this work, two different approaches for communicating particle data between domains are considered, and multiple criteria for load balancing problems have been investigated. Numerical results are presented for both fresh and depleted small modular nuclear reactor (SMR) cores. A parallel efficiency of approximately 80% was achieved with up to 16 spatial domains measured relative to full domain replication. A scaling study on the Summit supercomputer demonstrates a weak scaling parallel efficiency of over 90% on over 24000 GPUs.

97 MATHEMATICS AND COMPUTING↗

Selective area doping for Mott neuromorphic electronics

The cointegration of artificial neuronal and synaptic devices with homotypic materials and structures can greatly simplify the fabrication of neuromorphic hardware. We demonstrate experimental realization of vanadium dioxide (VO 2 ) artificial neurons and synapses on the same substrate through selective area carrier doping. By locally configuring pairs of catalytic and inert electrodes that enable nanoscale control over carrier density, volatility or nonvolatility can be appropriately assigned to each two-terminal Mott memory device per lithographic design, and both neuron- and synapse-like devices are successfully integrated on a single chip. Feedforward excitation and inhibition neural motifs are demonstrated at hardware level, followed by simulation of network-level handwritten digit and fashion product recognition tasks with experimental characteristics. Spatially selective electron doping opens up previously unidentified avenues for integration of emerging correlated semiconductors in electronic device technologies.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Parallel spatial direct numerical simulations on the Intel iPSC/860 hypercube

The implementation and performance of a parallel spatial direct numerical simulation (PSDNS) approach on the Intel iPSC/860 hypercube is documented. The direct numerical simulation approach is used to compute spatially evolving disturbances associated with the laminar-to-turbulent transition in boundary-layer flows. The feasibility of using the PSDNS on the hypercube to perform transition studies is examined. The results indicate that the direct numerical simulation approach can effectively be parallelized on a distributed-memory parallel machine. By increasing the number of processors nearly ideal linear speedups are achieved with nonoptimized routines; slower than linear speedups are achieved with optimized (machine dependent library) routines. This slower than linear speedup results because the Fast Fourier Transform (FFT) routine dominates the computational cost and because the routine indicates less than ideal speedups. However with the machine-dependent routines the total computational cost decreases by a factor of 4 to 5 compared with standard FORTRAN routines. The computational cost increases linearly with spanwise wall-normal and streamwise grid refinements. The hypercube with 32 processors was estimated to require approximately twice the amount of Cray supercomputer single processor time to complete a comparable simulation; however it is estimated that a subgrid-scale model which reduces the required number of grid points and becomes a large-eddy simulation (PSLES) would reduce the computational cost and memory requirements by a factor of 10 over the PSDNS. This PSLES implementation would enable transition simulations on the hypercube at a reasonable computational cost.

Joslin, Ronald D.↗

Message Passing vs. Shared Address Space on a Cluster of SMPs

The convergence of scalable computer architectures using clusters of PCs (or PC-SMPs) with commodity networking has become an attractive platform for high end scientific computing. Currently, message-passing and shared address space (SAS) are the two leading programming paradigms for these systems. Message-passing has been standardized with MPI, and is the most common and mature programming approach. However message-passing code development can be extremely difficult, especially for irregular structured computations. SAS offers substantial ease of programming, but may suffer from performance limitations due to poor spatial locality, and high protocol overhead. In this paper, we compare the performance of and programming effort, required for six applications under both programming models on a 32 CPU PC-SMP cluster. Our application suite consists of codes that typically do not exhibit high efficiency under shared memory programming. due to their high communication to computation ratios and complex communication patterns. Results indicate that SAS can achieve about half the parallel efficiency of MPI for most of our applications: however, on certain classes of problems SAS performance is competitive with MPI. We also present new algorithms for improving the PC cluster performance of MPI collective operations.

Shan, Hongzhang↗

Impact of Random Spatial Fluctuation in Non-Uniform Crystalline Phases on the Device Variation of Ferroelectric FET

In this work, a comprehensive study of random spatial fluctuation of the ferroelectric (FE) phase and dielectric (DE) phase in FeFETs is conducted to understand its impact on device variation. It is found that: i) there exists a certain DE percentage threshold that below which the increase of the DE phase does not significantly impact the device memory window and variation and only above which evident device degradation can be observed; ii) increasing the DE phase increases the variation in the memory window and the coercive field distribution further exacerbates the variation, hence degrading the sensing margin; iii) decreasing the number of grains degrades the device variation, which calls for further grain size engineering for variation suppression.

42 ENGINEERING↗

Mechanisms enabling reconfigurability and long-term retention in vanadium oxide electrochemical memory

Phase coexistence in nanoscale electrochemical random-access memory (ECRAM) has recently been demonstrated to enable both information storage and extraordinary reconfigurability. These proof-of-principle demonstrations have left the mechanistic details of such a process unresolved. Particularly, the mechanisms that stabilize the multiple phases, and the underlying processes behind sustained memory retention, remain unclear, and are necessary to design such devices. Here we report microscale ECRAM devices composed of V⁢O𝑥, which enables us to directly probe the active region in an operando fashion using optical techniques. Using Raman mapping, we show the phase coexistence driven by the electrochemical injection of O vacancies to be spatially uniform (i.e., with no filaments). The stability was observed to be unusually long, with 1% loss over 14 years in ambient conditions. First-principles calculations of the oxygen vacancy formation energies in V⁢O 𝑥 further support the thermodynamic coexistence of multiple V⁢O 𝑥 phases and clarify the origin of the observed long-term retention in the ECRAM devices. Further, we demonstrate single devices that can be voltage programmed to exhibit synaptic, neuronal, and reconfigurable logic gate functionalities. Furthermore, we not only uncover the phase coexistence mechanism that may help device design, but also demonstrate the circuit-level applications of reconfigurability.

Electrical conductivity↗

Phase‐Change‐Memory Process at the Limit: A Proposal for Utilizing Monolayer Sb 2 Te 3

Abstract One central task of developing nonvolatile phase change memory (PCM) is to improve its scalability for high‐density data integration. In this work, by first‐principles molecular dynamics, to date the thinnest PCM material possible (0.8 nm), namely, a monolayer Sb 2 Te 3 , is proposed. Importantly, its SET (crystallization) process is a fast one‐step transition from amorphous to hexagonal phase without the usual intermediate cubic phase. An increased spatial localization of electrons due to geometrical confinement is found to be beneficial for keeping the data nonvolatile in the amorphous phase at the 2D limit. The substrate and superstrate can be utilized to control the phase change behavior: e.g., with passivated SiO 2 (001) surfaces or hexagonal Boron Nitride, the monolayer Sb 2 Te 3 can reach SET recrystallization in 0.54 ns or even as fast as 0.12 ns, but with unpassivated SiO 2 (001), this would not be possible. Besides, working with small volume PCM materials is also a natural way to lower power consumption. Therefore, the proposed PCM working process at the 2D limit will be an important potential strategy of scaling the current PCM materials for ultrahigh‐density data storage.

2D limit↗