Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “spatial memory”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Spectral quadrature for the first principles study of crystal defects: Application to magnesium

In this work, we present an accurate and efficient finite-difference formulation and parallel implementation of Kohn-Sham Density (Operator) Functional Theory (DFT) for non periodic systems embedded in a bulk environment. Specifically, employing non-local pseudopotentials, local reformulation of electrostatics, and truncation of the spatial Kohn-Sham Hamiltonian, and the Linear Scaling Spectral Quadrature method to solve for the pointwise electronic fields in real-space and the non-local component of the atomic force, we develop a parallel finite difference framework suitable for distributed memory computing architectures to simulate non-periodic systems embedded in a bulk environment. Choosing examples from magnesium-aluminum alloys, we first demonstrate the convergence of energies and forces with respect to spectral quadrature polynomial order, and the width of the spatially truncated Hamiltonian. Next, we demonstrate the parallel scaling of our framework, and show that the computation time and memory scale linearly with respect to the number of atoms. Next, we use the developed framework to simulate isolated point defects and their interactions in magnesium-aluminum alloys. Our findings conclude that the binding energies of divacancies, Al solute-vacancy and two Al solute atoms are anisotropic and are dependent on cell size. Furthermore, the binding is favorable in all three cases.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

A hybrid CNN-LSTM surrogate model for hyper-resolution spatiotemporal flood forecasting in Norfolk, Virginia

Study region: Norfolk, Virginia, United States Study focus: Accurate and timely flood forecasting is essential for enhancing resilience in coastal urban areas in the context of increasing frequency and intensity of rainfall, sea level rise and urbanization. This study presents a hybrid deep learning-based surrogate model that integrates Convolutional Neural Networks (CNN) and Long Short-Term Memory (LSTM) networks to enable real-time spatiotemporal flood forecasting. The model leverages CNN to capture spatial features from inputs such as elevation and Topographic Wetness Index (TWI), while LSTM processes time-series inputs of rainfall and tide data to capture temporal features. New hydrologic insights for the region: The hybrid CNN-LSTM model was trained using the physics-based hydrodynamic model simulations obtained from the Two-dimensional Unsteady FLOW (TUFLOW) model for Norfolk, Virginia, and achieved high predictive accuracy across diverse flood-prone areas. The reduced computational time from four to six hours using TUFLOW to 3.2 min per event using CNN-LSTM enables rapid flood inundation mapping and early warning applications. The model effectively captured both spatial flood extents and their temporal evolution across different flooding scenarios, providing forecasts at a 2.5-m spatial resolution and 15-min temporal resolution and a one-hour-ahead prediction horizon. While challenges remain in terms of transferability to new regions and real-time data assimilation, this approach demonstrates strong potential for supporting operational flood risk management in coastal urban environments.

Coastal urban flooding↗

StressNet - Deep learning to predict stress with fracture propagation in brittle materials

Abstract Catastrophic failure in brittle materials is often due to the rapid growth and coalescence of cracks aided by high internal stresses. Hence, accurate prediction of maximum internal stress is critical to predicting time to failure and improving the fracture resistance and reliability of materials. Existing high-fidelity methods, such as the Finite-Discrete Element Model (FDEM), are limited by their high computational cost. Therefore, to reduce computational cost while preserving accuracy, a deep learning model, StressNet, is proposed to predict the entire sequence of maximum internal stress based on fracture propagation and the initial stress data. More specifically, the Temporal Independent Convolutional Neural Network (TI-CNN) is designed to capture the spatial features of fractures like fracture path and spall regions, and the Bidirectional Long Short-term Memory (Bi-LSTM) Network is adapted to capture the temporal features. By fusing these features, the evolution in time of the maximum internal stress can be accurately predicted. Moreover, an adaptive loss function is designed by dynamically integrating the Mean Squared Error (MSE) and the Mean Absolute Percentage Error (MAPE), to reflect the fluctuations in maximum internal stress. After training, the proposed model is able to compute accurate multi-step predictions of maximum internal stress in approximately 20 seconds, as compared to the FDEM run time of 4 h, with an average MAPE of 2% relative to test data.

36 MATERIALS SCIENCE↗

Scaling Resolution of Gigapixel Whole Slide Images Using Spatial Decomposition on Convolutional Neural Networks

Gigapixel images are prevalent in scientific domains ranging from remote sensing, and satellite imagery to microscopy, etc. However, training a deep learning model at the natural resolution of those images has been a challenge in terms of both, overcoming the resource limit (e.g. HBM memory constraints), as well as scaling up to a large number of GPUs. In this paper, we trained Residual neural Networks (ResNet) on 22,528 x 22,528-pixel size images using a distributed spatial decomposition method on 2,304 GPUs on the Summit Supercomputer. We applied our method on a Whole Slide Imaging (WSI) dataset from The Cancer Genome Atlas (TCGA) database. WSI images can be in the size of 100,000 x 100,000 pixels or even larger, and in this work we studied the effect of image resolution on a classification task, while achieving state-of-the-art AUC scores. Moreover, our approach doesn't need pixel-level labels, since we're avoiding patching from the WSI images completely, while adding the capability of training arbitrary large-size images. This is achieved through a distributed spatial decomposition method, by leveraging the non-block fat-tree interconnect network of the Summit architecture, which enabled GPU-to-GPU direct communication. Finally, detailed performance analysis results are shown, as well as a comparison with a data-parallel approach when possible.

Tsaris, Aristeidis (aris)↗

Technical note: Using long short-term memory models to fill data gaps in hydrological monitoring networks

Abstract. Quantifying the spatiotemporal dynamics in subsurface hydrological flows over a long time window usually employs a network of monitoring wells. However, such observations are often spatially sparse with potential temporal gaps due to poor quality or instrument failure. In this study, we explore the ability of recurrent neural networks to fill gaps in a spatially distributed time-series dataset. We use a well network that monitors the dynamic and heterogeneous hydrologic exchanges between the Columbia River and its adjacent groundwater aquifer at the U.S. Department of Energy's Hanford site. This 10-year-long dataset contains hourly temperature, specific conductance, and groundwater table elevation measurements from 42 wells with gaps of various lengths. We employ a long short-term memory (LSTM) model to capture the temporal variations in the observed system behaviors needed for gap filling. The performance of the LSTM-based gap-filling method was evaluated against a traditional autoregressive integrated moving average (ARIMA) method in terms of error statistics and accuracy in capturing the temporal patterns of river corridor wells with various dynamics signatures. Our study demonstrates that the ARIMA models yield better average error statistics, although they tend to have larger errors during time windows with abrupt changes or high-frequency (daily and subdaily) variations. The LSTM-based models excel in capturing both high-frequency and low-frequency (monthly and seasonal) dynamics. However, the inclusion of high-frequency fluctuations may also lead to overly dynamic predictions in time windows that lack such fluctuations. The LSTM can take advantage of the spatial information from neighboring wells to improve the gap-filling accuracy, especially for long gaps in system states that vary at subdaily scales. While LSTM models require substantial training data and have limited extrapolation power beyond the conditions represented in the training data, they afford great flexibility to account for the spatial correlations, temporal correlations, and nonlinearity in data without a priori assumptions. Thus, LSTMs provide effective alternatives to fill in data gaps in spatially distributed time-series observations characterized by multiple dominant frequencies of variability, which are essential for advancing our understanding of dynamic complex systems.

54 ENVIRONMENTAL SCIENCES↗

Accessing pluripotent materials through tempering of dynamic covalent polymer networks

Pluripotency, which is defined as a system not fixed as to its developmental potentialities, is typically associated with biology and stem cells. Here, inspired by this concept, we report synthetic polymers that act as a single “pluripotent” feedstock and can be differentiated into a range of materials that exhibit different mechanical properties, from hard and brittle to soft and extensible. To achieve this, we have exploited dynamic covalent networks that contain labile, dynamic thia-Michael bonds, whose extent of bonding can be thermally modulated and retained through tempering, akin to the process used in metallurgy. In addition, we show that the shape memory behavior of these materials can be tailored through tempering and that these materials can be patterned to spatially control mechanical properties.

36 MATERIALS SCIENCE↗

Fourier-based three-dimensional multistage transformer for aberration correction in multicellular specimens

High-resolution tissue imaging is often compromised by sample-induced optical aberrations that degrade resolution and contrast. Although wavefront sensor-based adaptive optics (AO) can measure these aberrations, such hardware solutions are typically complex, expensive to implement and slow when serially mapping spatially varying aberrations across large fields of view. Here we introduce AOViFT (adaptive optical vision Fourier transformer)—a machine learning-based aberration sensing framework built around a three-dimensional multistage vision transformer that operates on Fourier domain embeddings. AOViFT infers aberrations and restores diffraction-limited performance in puncta-labeled specimens with substantially reduced computational cost, training time and memory footprint compared to conventional architectures or real-space networks. We validated AOViFT on live gene-edited zebrafish embryos, demonstrating its ability to correct spatially varying aberrations using either a deformable mirror or postacquisition deconvolution. By eliminating the need for the guide star and wavefront sensing hardware and simplifying the experimental workflow, AOViFT lowers technical barriers for high-resolution volumetric microscopy across diverse biological samples.

Alshaabi, Thayer [Howard Hughes Medical Institute,↗

Performance Improvements for the Griffin Transport Solvers

Griffin is a Multiphysics Object-Oriented Simulation Environment based reactor multiphysics analysis application jointly developed by Idaho National Laboratory and Argonne National Laboratory. Griffin includes a variety of deterministic radiation transport solvers for fixed source, k-eigenvalue, adjoint, and subcritical multiplication, as well as transient solvers for point-kinetics, improved quasi-static, and spatial dynamics. A code assessment performed in FY-20 identified two significant issues with the transport solvers in Griffin: first, the primary heterogeneous SN (discrete ordinates) transport solver based on continuous finite element methods required significant mesh refinement and higher memory usage compared to solvers based on the method of characteristic for equivalent accuracy. Second, the homogeneous PN (spherical harmonics expansion) transport solver did not adequately support polynomial refinement, which is a feature usually required for problems with spatial homogenization and pronounced streaming, typical in fast or gas-cooled reactor systems. To address the first issue, the development effort focused on the more promising discontinuous finite element method (DFEM)-based SN transport solver in Griffin. The addition of an asynchronous parallel transport sweeper and coarse mesh finite difference (CMFD) acceleration have rendered a superior heterogeneous SN transport capability for multiphysics problems that requires far less computing resources in terms of both CPU time and memory usage. This is demonstrated with typical thermal- and fast-spectrum reactor benchmark problems, including 2D Transient Reactor Test, 3D Advanced Burner Test Reactor (ABTR), and 2D and 3D Empire microreactor. For the second issue, the development effort focused on a new transport solver based on the hybrid finite element PN method (HFEM-PN), equivalent to the variational nodal method, as well as a new diffusion solver based on HFEM-Diffusion. This solver is intended for homogenized domains with multiphysics coupling (i.e., supports mesh displacement, seamless temperature feedback, etc.). Initial calculations with the HFEM-Diffusion implementation show very good parallel efficiency for the residual evaluations with the 2D ABTR benchmark. A future development effort will be centered on further improvements to the CMFD, HFEM-PN, and DFEM diffusion solvers to ensure Griffin meets performance and software quality assurance requirements for advanced reactor design and analysis.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Upsampling Monte Carlo reactor simulation tallies in depleted LWR assemblies fueled with LEU and HALEU using a convolutional neural network

Simulating nuclear reactor cores at the highest achievable spatial and energy resolution is critical in modeling these systems accurately. Increasing the resolution, however, can dramatically increase the memory and central processing unit time required to run simulations. A convolutional neural network was shown previously to accurately upsample tally results of simulated light water reactor assemblies fueled with fresh, low enriched uranium. Here, we show that a convolutional neural network can be used to upsample tally results in assemblies containing fresh and depleted fuel enriched from 1.6 to 19.9 atom percent. The network was trained using neutron flux tallies from simulations of light water reactor assemblies with a range of fuel and coolant temperatures and a diverse selection of geometries. Accurate predictions of flux tallies are possible even on test assemblies with geometries and burnup levels well outside the range of those present in the training and validation data. The network improves the data density by a factor of 8 over a broad range of light water reactor assemblies while incurring insignificant additional computational cost to a Monte Carlo simulation.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

New modelling capabilities in IDT

This work concerns the enhanced modelling capabilities of the discrete ordinates transport solver IDT. The novelties introduced allow for modelling unstructured geometries composed by a collection of X/Y segments and circles, and the use of reciprocity and conservation relations reduce the memory imprint as well as the computational cost of the method. IDT decomposes geometries in modular Cartesian patterns, which are the so-called Heterogeneous Cartesian Cells (HCCs), containing a chunk of the original unstructured geometries. Each HCC can be then discretized by superimposing a XY grid to refine locally the HCC. Unlike the most popular MOC, IDT performs the spatial sweeping by directional collision probabilities instead of trajectories. The sources and interface angular fluxes are expanded up to linear order. The accuracy of ray-tracing, the memory imprints together with the novel mesh refinement capabilities have been verified. A first set of preliminary results on PWR lattice problems will be presented in this paper.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

GrainNN: A neighbor-aware long short-term memory network for predicting microstructure evolution during polycrystalline grain formation

High fidelity simulations of grain formation in alloys are an indispensable tool for process-to-mechanical-properties characterization. Such simulations, however, can be computationally expensive as they require fine spatial and temporal discretizations. Their cost becomes an obstacle to parametric studies and ensemble runs and ultimately makes downstream tasks like optimal control and uncertainty quantification challenging. To enable such downstream tasks, we introduce GrainNN, an efficient and accurate reduced-order model for epitaxial grain growth in additive manufacturing conditions. GrainNN is a sequence-to-sequence long-short-term-memory (LSTM) deep neural network that evolves the dynamics of manually crafted features. Its innovations are (1) an attention mechanism with grain-microstructure-specific transformer architecture; and (2) an overlapping combination of several clones of the network to generalize to grain configurations that are different from those used for training. This design enables GrainNN to predict grain formation for unseen physical parameters, grain number, domain size and geometry. Furthermore, GrainNN not only reconstructs the quantities of interest but also can be pointwise accurate. In our numerical experiments, we use a polycrystalline phase field method to both generate the training data and assess GrainNN. For multiparametric, ensemble simulations with many grains, GrainNN can be orders of magnitude faster than phase field simulations, while delivering 5%–15% pointwise error. Additionally, this speedup includes the cost of the phase field simulations for generating training data.

36 MATERIALS SCIENCE↗

The Stellar decomposition: A compact representation for simplicial complexes and beyond

Here, we introduce the Stellar decomposition, a model for efficient topological data structures over a broad range of simplicial and cell complexes. A Stellar decomposition of a complex is a collection of regions indexing the complex’s vertices and cells such that each region has sufficient information to locally reconstruct the star of its vertices, i.e., the cells incident in the region’s vertices. Stellar decompositions are general in that they can compactly represent and efficiently traverse arbitrary complexes with a manifold or non-manifold domain. They are scalable to complexes in high dimension and of large size, and they enable users to easily construct tailored application-dependent data structures using a fraction of the memory required by a corresponding global topological data structure on the complex. As a concrete realization of this model for spatially embedded complexes, we introduce the Stellar tree, which combines a nested spatial tree with a simple tuning parameter to control the number of vertices in a region. Stellar trees exploit the complex’s spatial locality by reordering vertex and cell indices according to the spatial decomposition and by compressing sequential ranges of indices. Stellar trees are competitive with state-of-the-art topological data structures for manifold simplicial complexes and offer significant improvements for cell complexes and non-manifold simplicial complexes. We conclude with a high-level description of several mesh processing and analysis applications that utilize Stellar trees to process large datasets.

97 MATHEMATICS AND COMPUTING↗

Mixed Delay/Nondelay Embeddings Based Neuromorphic Computing with Patterned Nanomagnet Arrays

Patterned nanomagnet arrays (PNAs) have been shown to exhibit a strong geometrically frustrated dipole interaction. Some PNAs have also shown emergent domain wall dynamics. Previous works have demonstrated methods to physically probe these magnetization dynamics of PNAs to realize neuromorphic reservoir systems that exhibit chaotic dynamical behavior and high-dimensional nonlinearity. These PNA reservoir systems from prior works leverage echo state properties and linear/nonlinear short-term memory of component reservoir nodes to map and preserve the dynamical information of the input time-series data into nondelay spatial embeddings. Such mappings enable these PNA reservoir systems to imitate and predict/forecast the input time series data. However, these prior PNA reservoir systems are based solely on the nondelay spatial embeddings obtained at component reservoir nodes. As a result, they require a massive number of component reservoir nodes, or a very large spatial embedding (i.e., high-dimensional spatial embedding) per reservoir node, or both, to achieve acceptable imitation and prediction accuracy. These requirements reduce the practical feasibility of such PNA reservoir systems. To address this shortcoming, we present a mixed delay/nondelay embeddings-based PNA reservoir system. Our system uses a single PNA reservoir node with the ability to obtain a mixture of delay/nondelay embeddings of the dynamical information of the time-series data applied at the input of a single PNA reservoir node. Our analysis shows that when these mixed delay/nondelay embeddings are used to train a perceptron at the output layer, our reservoir system outperforms existing PNA-based reservoir systems for the imitation of NARMA 2, NARMA 5, NARMA 7, and NARMA 10 time series data, and for the short-term and long-term prediction of the Mackey Glass time series data.

Ti, Changpeng↗

Energy Exascale Computational Fluid Dynamics Simulations With the Spectral Element Method

Development and application of the open-source GPU-based fluid-thermal simulation code, NekRS, are described. Time advancement is based on an efficient kth-order accurate timesplit formulation coupled with scalable iterative solvers. Spatial discretization is based on the high-order spectral element method (SEM), which affords the use of fast, low-memory, matrix-free operator evaluation. Further, recent developments include support for nonconforming meshes using overset grids and for GPU-based Lagrangian particle tracking. Results of large-eddy simulations of atmospheric boundary layers for wind-energy applications as well as extensive nuclear energy applications are presented.

42 ENGINEERING↗

Architecture-Aware Models of AI Engines for High-Performance Matrix Matrix Multiplication

The AI Engine (AIE) architecture, available in systems from mobile SoCs to server-class FPGAs, aims to efficiently execute AI/ML tasks through a two-dimensional array of compute tiles. Previous work on AIEs has explored different approaches to mapping computation across spatial arrays, but the compute kernel running on each tile has not been the focus. Additionally, the AIE-ML architecture introduces memory tiles and omits programmable logic, requiring new approaches to staging and moving data throughout the array. In this work we update analytical models developed for CPUs to produce the design of high performance kernels while introducing new model considerations such as memory structure, throughput, and latency as required by the AIE hardware. We evaluate our models by developing AIE-ML kernels for matrix multiplication in low-precision data types showing performance up to 95% of compute peak for the kernel when data resides in local memory and above 90% of compute peak when data resides in main memory.

Binder, Elliott D. [Carnegie Mellon University, Pi↗

Wireless Patch Antenna Characterization for Live Health Monitoring Using Machine Learning

Temperature monitoring in extreme environments, such as coal-fired power plants, was addressed by designing and testing wireless patch antennas for use in machine learning-aided temperature estimation. The sensors were designed to monitor the temperature and health of boiler systems. Wireless interrogation of the sensor was performed using a Vector Network Analyzer (VNA) and a pair of interrogation antennas to capture resonance behavior under varying thermal and spatial conditions with sensitivities ranging from 0.052 to 0.20 $\frac{𝑀𝐻𝑧}{°C}$. Sensor calibration was conducted using a Long Short-Term Memory (LSTM) model, which leveraged temporal patterns to account for hysteresis effects. The calibration method demonstrated improved performance when combined with an LSTM model, achieving up to a 76% improvement in temperature estimation error when compared with Linear Regression (LR). The experiments highlighted an innovative solution for patch antenna-based non-contact temperature measurement, which addresses limitations with conventional methods such as RFID-based systems, infrared, and thermocouples.

20 FOSSIL-FUELED POWER PLANTS↗

Quantum Imaging of Ferromagnetic van der Waals Magnetic Domain Structures at Ambient Conditions

Recently discovered 2D van der Waals magnetic materials, and specifically iron–germanium–telluride (Fe5GeTe2), have attracted significant attention both from a fundamental perspective and for potential applications. Key open questions concern their domain structure and magnetic phase transition temperature as a function of sample thickness and external field, as well as implications for integration into devices such as magnetic memories and logic. Here we address key questions using a nitrogen-vacancy center based quantum magnetic microscope, enabling direct imaging of the magnetization of Fe5GeTe2 at submicrometer spatial resolution as a function of temperature, magnetic field, and thickness. This quantum imaging technique provides noninvasive, high-sensitivity measurements with high spatial resolution under ambient conditions, making it particularly well suited for probing 2D magnets. We employ spatially resolved measures, including magnetization variance and cross-correlation, and find a significant spread in transition temperature yet with no clear dependence on thickness down to 15 nm. We also identify previously unknown stripe features in the optical as well as magnetic images, which we attribute to modulations of the constituting elements during crystal synthesis and subsequent oxidation. Our results suggest that the magnetic anisotropy in this material does not play a crucial role in their magnetic properties, leading to a magnetic phase transition of Fe5GeTe2 which is largely thickness-independent down to 15 nm. Our findings could be significant in designing future spintronic devices, magnetic memories, and logic with 2D van der Waals magnetic materials.

Bindu, Bindu [Hebrew University of Jerusalem, Isra↗

Kelvin Probe Force Microscopy Imaging of Plasticity in Hydrogenated Perovskite Nickelate Multilevel Neuromorphic Devices

Ion drift in nanoscale electronically inhomogeneous semiconductors is among the most important mechanisms being studied for designing neuromorphic computing hardware. However, nondestructive imaging of the ion drift in operando devices directly responsible for multiresistance states and synaptic memory represents a formidable challenge. Here, we present Kelvin probe force microscopy imaging of hydrogen-doped perovskite nickelate device channels subject to high-speed electric field pulses to directly visualize proton distribution by monitoring surface potential changes spatially, which is also supported with finite element-based electric field distribution studies. First-principles calculations provide mechanistic insights into the origin of surface potential changes as a function of hydrogen donor doping that serves as the contrast mechanism. We demonstrate 128 (7-bit) nonvolatile conductance levels in such devices relevant to in-memory computing applications. The synaptic plasticity measurements are implemented in spiking neural networks and show promising results for classification (SciKit Learn’s Iris and Wine data sets) and control (OpenAI’s CartPole-v1 and BipedalWalker-v3) simulation tasks.

Kelvin probe force microscopy↗