Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “memory mapping”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Machine learning models of intermittent operation of RO wellhead water treatment for salinity reduction and nitrate removal

Machine learning models were developed for intermittent multi-mode operation of a wellhead reverse osmosis water purification and desalination system to predict salt passage, nitrate passage, and permeate flux. The models, based on long short-term memory (LSTM) recurrent neural network (RNN) architecture, included an attention mechanism to increase model performance in proximity of the regulatory limit for nitrate. Training and testing of the models for the Startup, Production, Shutdown and Flushing operational modes were based on operational data (consisting of 22 process variables per data sample) acquired every 2–5 s over a six-month period. The significant sets of model input attributes for the different operational modes were assessed via Spearman ranking correlation, Self-Organizing Map (SOM) analysis and feed forward feature selection (FFFS). Although the variability of nitrate passage, salt passage and permeate flux was significant over the four operational modes, prediction performance for the three outcomes were with R2 and Average Absolute Relative Error (AARE) of 0.78–0.95 and 2.96–6.16 %, respectively. Model updates post membrane elements replacement demonstrated similar levels of prediction accuracy. The study results suggest that there is merit in exploring the utility of multi-mode models for sensor fault detection, data imputation, and for potential use in model-predictive control.

Intermittent RO operation↗

Ferroelectric HfO 2 and the importance of strain

Ferroelectric oxides based on HfO 2 show tremendous promise for the next generation of memory and logic devices. The ferroelectric polymorph is one of several that can be derived from the high symmetry cubic fluorite structure of HfO 2 . A single grain of HfO 2 may consist of a coherent mixture of multiple orientational and translational variants of different polymorphs. Here, we use symmetry-adapted strain-order parameters to elucidate the relationship between the different HfO 2 polymorphs and their symmetrically equivalent variants. We use first-principles electronic structure methods to identify minimum energy pathways and map them in subspaces of the symmetry-adapted strain order parameters. We next investigate the atomic structure of domain boundaries that separate coexisting variants of ferroelectric HfO 2 . Further, we rely on Gibbsian excess quantities and a precise specification of mechanical boundary conditions to describe the thermodynamic properties of domain boundaries. Our first-principles calculations show that the O and Hf shuffle arrangement within a domain boundary is closely related to the intermediate shuffle patterns of the homogeneous pathways between ferroelectric variants. Furthermore, the preferred structure within a boundary is very sensitive to local strain constraints imposed by the adjacent ferroelectric variants, leading to highly anisotropic domain boundary energies.

36 MATERIALS SCIENCE↗

Fast and Flexible Inference Framework for Continuum Reverberation Mapping Using Simulation-based Inference with Deep Learning

Continuum reverberation mapping (CRM) of active galactic nuclei (AGN) monitors multiwavelength variability signatures to constrain accretion disk structure and supermassive black hole (SMBH) properties. The upcoming Vera Rubin Observatory’s Legacy Survey of Space and Time will survey tens of millions of AGN over the next decade, with thousands of AGN monitored with almost daily cadence in the deep drilling fields. However, existing CRM methodologies often require long computation time and are not designed to handle such large amounts of data. In this paper, we present a fast and flexible inference framework for CRM using simulation-based inference (SBI) with deep learning to estimate SMBH properties from AGN light curves. We use a long short-term memory summary network to reduce the high dimensionality of the light curve data and then use a neural density estimator to estimate the posterior of SMBH parameters. Using simulated light curves, we find SBI can produce more accurate SMBH parameter estimation with 10 3 –10 5 times speed up in inference efficiency compared to traditional methods. The SBI framework is particularly suitable for wide-field CRM surveys as the light curves will have identical observing patterns, which can be incorporated into the SBI simulation. We explore the performance of our SBI model on light curves with irregular-sampled, realistic observing cadence and alternative variability characteristics to demonstrate the flexibility and limitation of the SBI framework.

79 ASTRONOMY AND ASTROPHYSICS↗

Uniform-in-phase-space data selection with iterative normalizing flows

Improvements in computational and experimental capabilities are rapidly increasing the amount of scientific data that are routinely generated. In applications that are constrained by memory and computational intensity, excessively large datasets may hinder scientific discovery, making data reduction a critical component of data-driven methods. Datasets are growing in two directions: the number of data points and their dimensionality. Whereas dimension reduction typically aims at describing each data sample on lower-dimensional space, the focus here is on reducing the number of data points. A strategy is proposed to select data points such that they uniformly span the phase-space of the data. The algorithm proposed relies on estimating the probability map of the data and using it to construct an acceptance probability. An iterative method is used to accurately estimate the probability of the rare data points when only a small subset of the dataset is used to construct the probability map. Instead of binning the phase-space to estimate the probability map, its functional form is approximated with a normalizing flow. Therefore, the method naturally extends to high-dimensional datasets. The proposed framework is demonstrated as a viable pathway to enable data-efficient machine learning when abundant data are available.

97 MATHEMATICS AND COMPUTING↗

Understanding Complex Magnetic Spin Textures with Simulation-Assisted Lorentz Transmission Electron Microscopy

There is an increased interest in topologically nontrivial magnetic spin textures such as skyrmions and chiral domain-wall solitons, both from a point of fundamental physics understanding as well as potential technological interest in low-power memory applications. In order to control their behavior, it is necessary to understand their complex spin texture at the nanoscale. Lorentz transmission electron microscopy (LTEM) is a suitable technique for studying these systems due to its high spatial resolution and capability to simultaneously characterize magnetic texture and microstructure. In this work, we present the application of PyLorentz, an open-source software suite that we have developed, for quantitative image analysis of Neel-type skyrmions in thin-film heterostructures. PyLorentz enhances LTEM capabilities by enabling reconstruction of magnetic induction maps from experimental images, as well as simulating LTEM images using micromagnetic simulation data. We demonstrate this for simulated Neel skyrmions as well as experimental data from [Pt/Co/W] multilayer heterostructures. Finally, we also show how simulation-assisted LTEM analysis is crucial for understanding these complex magnetic spin textures, in which the reconstructed magnetic induction map (seen in the LTEM images) differs significantly from the magnetization configuration.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Attribute-Aware RBFs: Interactive Visualization of Time Series Particle Volumes Using RT Core Range Queries

Smoothed-particle hydrodynamics (SPH) is a mesh-free method used to simulate volumetric media in fluids, astrophysics, and solid mechanics. Visualizing these simulations is problematic because these datasets often contain millions, if not billions of particles carrying physical attributes and moving over time. Radial basis functions (RBFs) are used to model particles, and overlapping particles are interpolated to reconstruct a high-quality volumetric field; however, this interpolation process is expensive and makes interactive visualization difficult. Existing RBF interpolation schemes do not account for color-mapped attributes and are instead constrained to visualizing just the density field. To address these challenges, we exploit ray tracing cores in modern GPU architectures to accelerate scalar field reconstruction. We use a novel RBF interpolation scheme to integrate per-particle colors and densities, and leverage GPU-parallel tree construction and refitting to quickly update the tree as the simulation animates over time or when the user manipulates particle radii. We also propose a Hilbert reordering scheme to cluster particles together at the leaves of the tree to reduce tree memory consumption. Finally, we reduce the noise of volumetric shadows by adopting a spatially temporal blue noise sampling scheme. Our method can provide a more detailed and interactive view of these large, volumetric, time-series particle datasets than traditional methods, leading to new insights into these physics simulations.

Particle Volumes↗

Diagnostics of Mixed-State Topological Order and Breakdown of Quantum Memory

Topological quantum memory can protect information against local errors up to finite error thresholds. Such thresholds are usually determined based on the success of decoding algorithms rather than the intrinsic properties of the mixed states describing corrupted memories. Here we provide an intrinsic characterization of the breakdown of topological quantum memory, which both gives a bound on the performance of decoding algorithms and provides examples of topologically distinct mixed states. We employ three information-theoretical quantities that can be regarded as generalizations of the diagnostics of ground-state topological order, and serve as a definition for topological order in error-corrupted mixed states. We consider the topological contribution to entanglement negativity and two other metrics based on quantum relative entropy and coherent information. In the concrete example of the two-dimensional (2D) Toric code with local bit-flip and phase errors, we map three quantities to observables in 2D classical spin models and analytically show they all undergo a transition at the same error threshold. This threshold is an upper bound on that achieved in any decoding algorithm and is indeed saturated by that in the optimal decoding algorithm for the Toric code. Published by the American Physical Society 2024

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Non-Intrusive Appliance Identification with Appliance-Specific Networks

The problem of noninstrusive load monitoring (NILM) is usually formulated as a single-channel blind source separation task, whose successful solution enable fast and convenient load identification and energy disaggregation. When applied at test time, NILM algorithms aim to identify the operating characteristics of individual appliances from an aggregate power measurement of the entire house. Recent advances in deep learning gave rise to many methods that mostly focus on learning a direct mapping from aggregate measurement to individual appliance power. However, these methods are not only computationally expensive, but they often suffer from overfitting and do not generalize very well. In this article, we propose a novel NILM method that leverages advances in statistical learning that have not been properly applied in this domain before. The proposed method consists of three stages: first, a Bayesian nonparametric learning-based approach for appliance state extraction; second, synthetic minority oversampling technique for data augmentation and mitigating the heavy imbalance in switching events; and third, appliance-specific lightweight long short-term memory networks for status classification for each appliance. Here, we adopt a “differential” input (the difference before and after the switching event) to reduce the complexity of network training and make the proposed method robust to multiappliance switching events. Experiments are conducted to demonstrate the effectiveness of the proposed method, achieving superior performance when compared to recent methods. An ablation study is conducted to demonstrate the effectiveness of each module of our method. Finally, we investigate the quality of generated synthetic samples.

42 ENGINEERING↗

Porting the WAVEWATCH III (v6.07) wave action source terms to GPU

Abstract. Surface gravity waves play a critical role in several processes, including mixing, coastal inundation, and surface fluxes. Despite the growing literature on the importance of ocean surface waves, wind–wave processes have traditionally been excluded from Earth system models (ESMs) due to the high computational costs of running spectral wave models. The development of the Next Generation Ocean Model for the DOE’s (Department of Energy) E3SM (Energy Exascale Earth System Model) Project partly focuses on the inclusion of a wave model, WAVEWATCH III (WW3), into E3SM. WW3, which was originally developed for operational wave forecasting, needs to be computationally less expensive before it can be integrated into ESMs. To accomplish this, we take advantage of heterogeneous architectures at DOE leadership computing facilities and the increasing computing power of general-purpose graphics processing units (GPUs). This paper identifies the wave action source terms, W3SRCEMD, as the most computationally intensive module in WW3 and then accelerates them via GPU. Our experiments on two computing platforms, Kodiak (P100 GPU and Intel(R) Xeon(R) central processing unit, CPU, E5-2695 v4) and Summit (V100 GPU and IBM POWER9 CPU) show respective average speedups of 2× and 4× when mapping one Message Passing Interface (MPI) per GPU. An average speedup of 1.4× was achieved using all 42 CPU cores and 6 GPUs on a Summit node (with 7 MPI ranks per GPU). However, the GPU speedup over the 42 CPU cores remains relatively unchanged (∼ 1.3×) even when using 4 MPI ranks per GPU (24 ranks in total) and 3 MPI ranks per GPU (18 ranks in total). This corresponds to a 35 %–40 % decrease in both simulation time and usage of resources. Due to too many local scalars and arrays in the W3SRCEMD subroutine and the huge WW3 memory requirement, GPU performance is currently limited by the data transfer bandwidth between the CPU and the GPU. Ideally, OpenACC routine directives could be used to further improve performance. However, W3SRCEMD would require significant code refactoring to make this possible. We also discuss how the trade-off between the occupancy, register, and latency affects the GPU performance of WW3.

58 GEOSCIENCES↗

Low-temperature grapho-epitaxial La-substituted BiFeO 3 on metallic perovskite

Bismuth ferrite has garnered considerable attention as a promising candidate for magnetoelectric spin-orbit coupled logic-in-memory. As model systems, epitaxial BiFeO 3 thin films have typically been deposited at relatively high temperatures (650–800 °C), higher than allowed for direct integration with silicon-CMOS platforms. Here, we circumvent this problem by growing lanthanum-substituted BiFeO 3 at 450 °C (which is reasonably compatible with silicon-CMOS integration) on epitaxial BaPb 0.75 Bi 0.25 O 3 electrodes. Notwithstanding the large lattice mismatch between the La-BiFeO 3 , BaPb 0.75 Bi 0.25 O 3 , and SrTiO 3 (001) substrates, all the layers in the heterostructures are well ordered with a [001] texture. Polarization mapping using atomic resolution STEM imaging and vector mapping established the short-range polarization ordering in the low temperature grown La-BiFeO 3 . Current-voltage, pulsed-switching, fatigue, and retention measurements follow the characteristic behavior of high-temperature grown La-BiFeO 3 , where SrRuO 3 typically serves as the metallic electrode. These results provide a possible route for realizing epitaxial multiferroics on complex-oxide buffer layers at low temperatures and opens the door for potential silicon-CMOS integration.

36 MATERIALS SCIENCE↗

Soft error-mitigating semiconductor design system and associated methods

A soft error-mitigating semiconductor design system and associated methods that tailor circuit design steps to mitigate corruption of data in storage elements (e.g., flip flops) due to Single Events Effects (SEEs). Required storage elements are automatically mapped to triplicated redundant nodes controlled by a voting element that enforces majority-voting logic for fault-free output (i.e., Triple Modular Redundancy (TMR)). Storage elements are also optimally positioned for placement in keeping with SEE-tolerant spacing constraints. Additionally, clock delay insertion (employing either a single global clock or clock triplication) in the TMR specification may introduce useful skew that protects against glitch propagation through the designed device. The resultant layout generated from the TMR configuration may relax constraints imposed on register transfer level (RTL) engineers to make rad-hard designs, as automation introduces TMR storage registers, memory element spacing, and clock delay/triplication with minimal designer input.

Miryala, Sandeep↗

Soft error-mitigating semiconductor design system and associated methods

A soft error-mitigating semiconductor design system and associated methods that tailor circuit design steps to mitigate corruption of data in storage elements (e.g., flip flops) due to Single Events Effects (SEEs). Required storage elements are automatically mapped to triplicated redundant nodes controlled by a voting element that enforces majority-voting logic for fault-free output (i.e., Triple Modular Redundancy (TMR)). Storage elements are also optimally positioned for placement in keeping with SEE-tolerant spacing constraints. Additionally, clock delay insertion (employing either a single global clock or clock triplication) in the TMR specification may introduce useful skew that protects against glitch propagation through the designed device. The resultant layout generated from the TMR configuration may relax constraints imposed on register transfer level (RTL) engineers to make rad-hard designs, as automation introduces TMR storage registers, memory element spacing, and clock delay/triplication with minimal designer input.

Miryala, Sandeep↗

Time-series machine-learning error models for approximate solutions to parameterized dynamical systems

This work proposes a machine-learning framework for modeling the error incurred by approximate solutions to parameterized dynamical systems. In particular, we extend the machine-learning error models (MLEM) framework proposed in Ref. Freno and Carlberg (2019) to dynamical systems. The proposed Time-Series Machine-Learning Error Modeling (T-MLEM) method constructs a regression model that maps features – which comprise error indicators that are derived from standard a posteriori error-quantification techniques – to a random variable for the approximate-solution error at each time instance. The proposed framework considers a wide range of candidate features, regression methods, and additive noise models. We consider primarily recursive regression techniques developed for time-series modeling, including both classical time-series models (e.g., autoregressive models) and recurrent neural networks (RNNs), but also analyze standard non-recursive regression techniques (e.g., feed-forward neural networks) for comparative purposes. Finally, numerical experiments conducted on multiple benchmark problems illustrate that the long short-term memory (LSTM) neural network, which is a type of RNN, outperforms other methods and yields substantial improvements in error predictions over traditional approaches.

42 ENGINEERING↗

Random insights into the complexity of two-dimensional tensor network calculations

Projected entangled pair states (PEPS) offer memory-efficient representations of some quantum many-body states that obey an entanglement area law and are the basis for classical simulations of ground states in two-dimensional (2d) condensed matter systems. However, rigorous results show that exactly computing observables from a 2d PEPS state is generically a computationally hard problem. Yet approximation schemes for computing properties of 2d PEPS are regularly used, and empirically seen to succeed, for a large subclass of (“not too entangled”) condensed matter ground states. Adopting the philosophy of random matrix theory, in this work, we analyze the complexity of approximately contracting a 2d random PEPS by exploiting an analytic mapping to an effective replicated statistical mechanics model that permits a controlled analysis at a large bond dimension. Through this statistical-mechanics lens, we argue that (i) although approximately sampling wave-function amplitudes of random PEPS faces a computational-complexity phase transition above a critical bond dimension, and (ii) one can generically efficiently estimate the norm and correlation functions for any finite bond dimension. Furthermore, these results are supported numerically for various bond-dimension regimes. It is an important open question whether the above results for random PEPS apply more generally also to PEPS representing physically relevant ground states.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Advancing attenuation estimation through integration of the Hessian in multiparameter viscoacoustic full-waveform inversion

Accurate seismic attenuation models of subsurface structures not only enhance subsequent migration processes by improving fidelity, resolution, and facilitating amplitude-compliant angle gather generation but also provide valuable constraints on subsurface physical properties. Leveraging full-wavefield information, multiparameter viscoacoustic full-waveform inversion ( Q-FWI) simultaneously estimates seismic velocity and attenuation ( Q) models. However, a major challenge in Q-FWI is the contamination of crosstalk artifacts, where inaccuracies in the velocity model are mistakenly mapped to the inverted attenuation model. While incorporating the Hessian is expected to mitigate these artifacts, the explicit implementation is prohibitively expensive due to its formidable computational cost. In this study, we formulate and develop a Q-FWI algorithm via the Newton-conjugate gradient (CG) framework, where the search direction at each iteration is determined through an internal CG loop. In particular, the Hessian is integrated into each CG step in a matrix-free fashion using the second-order adjoint-state method. We find through synthetic experiments that our Newton-CG Q-FWI significantly mitigates crosstalk artifacts compared with the limited-memory Broyden-Fletcher-Goldfarb-Shanno method and the CG method, albeit with a notable computational cost. In the discussion of several key implementation details, we also determine the significance of the approximate Gauss-Newton Hessian, the second-order adjoint-state method, and the two-stage inversion strategy.

Geochemistry & Geophysics↗

A hybrid CNN-LSTM surrogate model for hyper-resolution spatiotemporal flood forecasting in Norfolk, Virginia

Study region: Norfolk, Virginia, United States Study focus: Accurate and timely flood forecasting is essential for enhancing resilience in coastal urban areas in the context of increasing frequency and intensity of rainfall, sea level rise and urbanization. This study presents a hybrid deep learning-based surrogate model that integrates Convolutional Neural Networks (CNN) and Long Short-Term Memory (LSTM) networks to enable real-time spatiotemporal flood forecasting. The model leverages CNN to capture spatial features from inputs such as elevation and Topographic Wetness Index (TWI), while LSTM processes time-series inputs of rainfall and tide data to capture temporal features. New hydrologic insights for the region: The hybrid CNN-LSTM model was trained using the physics-based hydrodynamic model simulations obtained from the Two-dimensional Unsteady FLOW (TUFLOW) model for Norfolk, Virginia, and achieved high predictive accuracy across diverse flood-prone areas. The reduced computational time from four to six hours using TUFLOW to 3.2 min per event using CNN-LSTM enables rapid flood inundation mapping and early warning applications. The model effectively captured both spatial flood extents and their temporal evolution across different flooding scenarios, providing forecasts at a 2.5-m spatial resolution and 15-min temporal resolution and a one-hour-ahead prediction horizon. While challenges remain in terms of transferability to new regions and real-time data assimilation, this approach demonstrates strong potential for supporting operational flood risk management in coastal urban environments.

Coastal urban flooding↗

Polarization-Controlled Structural Modulation in the Single Atomic Layer at the PbZr 0.2 Ti 0.8 O 3 /LaNiO 3 Interface

Conductivity modulation via ferroelectric polarization coupling with LaNiO 3 (LNO) is demonstrated at an epitaxial ferroelectric–LNO interface. Conductivity measurements, varying the thickness of the LNO channel, show that this phenomenon is confined to a few atomic layers at the interface. Combining in situ biasing and off-axis holography, we mapped out electrostatic potentials at the PbZr 0.2 Ti 0.8 O 3 (PZT)/LNO/SrTiO 3 (STO) heterostructure upon polarization switching. Using aberration-corrected STEM, the interfacial atomic structures were investigated for the two different PZT polarization states. Polarization in PZT induces a significant change in the in-plane O–Ni–O bond angles, with a 37° modulation in the topmost 1 or 2 LNO unit cells, driven by strain in the oxygen sublattice for the two opposite polarization directions in PZT. Both oxygen and cation sublattices exhibit strain responses exceeding 10% upon switching. This atomic-layer structural modulation highlights a mechanism for functional oxide heterostructure development, offering pathways for advancements in nonvolatile memory, sensors, and energy-efficient transistors.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Localized strain profile in surface electrode array for programmable composite multiferroic devices

In this work, we investigate localized in-plane strains on the microscale, induced by arrays of biased surface electrodes patterned on piezoelectrics. Particular focus is given to the influence that adjacent electrode pairs have on one another to study the impact of densely packed electrode arrays. We present a series of X-ray microdiffraction studies to reveal the spatially resolved micrometer-scale strain distribution. The strain maps with micrometer-scale resolution highlight how the local strain profile in square regions up to 250 x 250 lm 2 in size is affected by the surface electrodes that are patterned on ferroelectric single-crystal [Pb(Mg 1/3 Nb 2/3 )O 3 ] x -[PbTiO 3 ] 1-x . The experimental measurements and simulation results show the influence of electrode pair distance, positioning of the electrode pair, including the angle of placement, and neighboring electrode pair arrangements on the strength and direction of the regional strain. Our findings are relevant to the development of microarchitected strain-mediated multiferroic devices. The electrode arrays could provide array-addressable localized strain control for applications including straintronic memory, probabilistic computing platforms, microwave devices, and magnetic-activated cell sorting platforms.

42 ENGINEERING↗