Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Memory device”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

A Graph Dynamical neural network approach for decoding dynamical states in ferroelectrics.

Ferroelectric materials such as BaTiO 3 show tremendous potential for emerging advances in memory devices, particular neuromorphic type devices. High density of memory can be obtained by stabilising polar domain walls at the nanoscale, regions of discontinuity between the well-defined polarization order parameter, but little is known about what controls their structure and dynamics in real nanoscale materials. Indeed, chiral polar domain walls have been observed in heterogeneous ferroelectrics, such as oxygen-deficient BaTiO 3 , but very little is known about how such polar-domains walls interact with defects. Indeed, a critical understanding of how dynamics of domain-walls depend on point-defects is crucial to create engineered ferroelectric memory devices. For this work, we perform large-scale simulations of nansocale domain-wall dynamics in pristine and defective BaTiO 3 using reactive force-field developed by us earlier (Phys. Chem. Chem. Phys., 2019, 21, 18240–18249), and capture their dynamical dependence on point defects using a graph dynamical neural-network approach, which we adapted to interrogate solids with well-defined order-parameters, and implemented using Pytorch based libraries. Our machine learning (ML) approach goes beyond the traditional post-processing methods to capture both spatial and temporal heterogeneities of large-scale molecular dynamics simulations of complex defective ferroelectric oxide materials. We crucially find that isolated oxygen vacancies introduce very localized spatial regions (~1–2 unit-cell in length) that show slow dipole relaxation due to formation of defect-dipoles, and that these defect-dipoles in turn slow the intrinsic dynamics of domain walls. Further, the roughness of domain walls, also influenced by vacancies, introduce dynamic heterogeneity along the domain-wall. As such we find a novel mechanism by which quenched disorder due to defects introduce dynamic heterogeneity thereby influencing response to external fields (particularly time varying fields) in a ferroelectric. Our study also emphasizes the need for creating digital twins of dynamical quantities to achieve autonomous in operando control of nanoscale switching.

42 ENGINEERING↗

Mneme

A simple tool allowing recording the execution of a GPU (CUDA) kernel and replaying that kernel as an independent executable. The tool operates in 3 phases. During compile time the user needs to apply a provided LLVM pass to instrument the code. The pass detects all device global variables and device functions and stores this information with the respective LLVM-IR in the global device memory. The compilation generates a record-able executable. The second phase involves running the application executable with a desired input and using LD_PRELOAD to enable recording. When recording before invoking a device kernel the pre-loaded library stores device memory in persistent storage and associates the memory with the device kernel and an LLVM IR file. At the end of the recorded execution the pre-load library generates a database in the form of a JSON file containing information regarding the LLVM-IR files and the snapshots of device memory. During the third and last phase the user can replay the execution of an kernel as a separate independent executable. Besides executing it the user can modify the LLVM IR file and auto-tune parameters such as kernel launch-bounds or kernel runtime execution parameters (e.g. Kernel Block and Grid Dimensions). Is

Parasyris, Konstantinos↗

FPGA-based computing system for processing data in size, weight, and power constrained environments

Technologies that are well-suited for use in size, weight, and power (SWAP)-constrained environments are described herein. A host controller dispatches data processing instructions to hardware acceleration engines (HAEs) of one or more field programmable gate arrays (FPGAs) and further dispatches data transfer instructions to a memory controller, such that the HAEs perform processing operations on data stored in local memory devices of the HAEs in parallel with other data being transferred from external memory devices coupled to the FPGA(s) to the local memory devices.

Napier, Matthew↗

Resistive Switching Memory Performance of Two-Dimensional Polyimide Covalent Organic Framework Films

Two-dimensional polyimide covalent organic framework (2D PI-NT COF) films were constructed on indium tin oxide-coated glass substrates to fabricate two-terminal sandwiched resistive memory devices. The 2D PI-NT COF films condensated from the reaction between 4,4',4"-triaminotriphenylamine and naphthalene-1,4,5,8-tetracarboxylic dianhydride under solvothermal conditions demonstrated high crystallinity, good orientation preference, tunable thickness, and low surface roughness. The well-aligned electron-donor (triphenylamine unit) and -acceptor (naphthalene diimide unit) arrays rendered the 2D PI-NT COF films a promising candidate for electronic applications. The memory devices based on 2D PI-NT COF films exhibited a typical write-once-read-many-time resistive switching behavior under an operating voltage of +2.30 V on the positive scan and -2.64 V on the negative scan. A high ON/OFF current ratio (>10 6 for the positive scan and 10 4 -10 6 for the negative scan) and long-term retention time indicated the high fidelity, low error, and high stability of the resistive memory devices. The memory behavior was attributed to an electric field-induced intramolecular charge transfer in an ordered donor-acceptor system, which provided the effective charge-transfer channels for injected charge carriers. Furthermore, this work represents the first example that explores the resistive memory properties of 2D PI-COF films, shedding light on the potential application of 2D COFs as information storage media.

2D covalent organic framework film↗

Feature Classification for Control System Devices

Control systems are used to automate industrial processes, smart grids, and smart cities. Unfortunately, cyber attacks on control systems are on the rise. Additionally, control systems lack the plethora of tools available for commodity systems for forensic investigation. An important step towards the proper forensic investigation is to analyze device memory. To assist in identifying features of device memory, we present a machine learning-based technique that integrates ontology information for feature classification in a control system device’s memory.

ahmed mithu, M Rayhan↗

HAM: Hotspot-Aware Manager for Improving Communications with 3D-Stacked Memory

merging High-Performance Computing (HPC) workloads, such as graph analytics, machine learning, and big data science, are data-intensive. Data-intensive workloads usually present fine-grained memory accesses with limited or no data locality, and thus incur frequent cache misses and low utilization of memory bandwidth. 3D-stacked memory devices such as Hybrid Memory Cube (HMC) and High Bandwidth Memory (HBM) can provide significantly higher bandwidth than conventional memory modules. However, the traditional interfaces and optimization methods for JEDEC DDR devices do not allow to fully exploit the potential performance of 3D-stacked memory with the massive amount of irregular memory accesses of data-intensive applications. In this paper, we propose a novel Hotspot-Aware Manager (HAM) infrastructure for 3D-stacked memory devices capable of optimizing memory access streams via request aggregation, hotspot detection, and in-memory prefetching. %and an associated hotspot-aware page policy. We present the HAM design and implementation, and simulate it on a system using RISC-V embedded cores with attached HMC devices. We extensively evaluate HAM with over 12 benchmarks and applications representing diverse irregular memory access patterns. The results show that, on average, HAM reduces redundant requests by 37.51\% and increases the prefetch buffer hit rate by 4.2 times, compared to a baseline streaming prefetcher. On the selected benchmark set, HAM provides performance gains of 21.81\% in average (up to 34.28\%) and power savings of 35.07\% over a standard 3D-stacked memory.

Wang, Xi↗

Cycle accurate and cycle reproducible memory for an FPGA based hardware accelerator

A method, system and computer program product are disclosed for using a Field Programmable Gate Array (FPGA) to simulate operations of a device under test (DUT). The DUT includes a device memory having a number of input ports, and the FPGA is associated with a target memory having a second number of input ports, the second number being less than the first number. In one embodiment, a given set of inputs is applied to the device memory at a frequency Fd and in a defined cycle of time, and the given set of inputs is applied to the target memory at a frequency Ft. Ft is greater than Fd and cycle accuracy is maintained between the device memory and the target memory. In an embodiment, a cycle accurate model of the DUT memory is created by separating the DUT memory interface protocol from the target memory storage array.

Asaad, Sameh W.↗

Si/SiGe QuBus for single electron information-processing devices with memory and micron-scale connectivity function

The connectivity within single carrier information-processing devices requires transport and storage of single charge quanta. Single electrons have been adiabatically transported while confined to a moving quantum dot in short, all-electrical Si/SiGe shuttle device, called quantum bus (QuBus). Here we show a QuBus spanning a length of 10 μm and operated by only six simply-tunable voltage pulses. We introduce a characterization method, called shuttle-tomography, to benchmark the potential imperfections and local shuttle-fidelity of the QuBus. The fidelity of the single-electron shuttle across the full device and back (a total distance of 19 μm) is (99.7 ± 0.3) %. Using the QuBus, we position and detect up to 34 electrons and initialize a register of 34 quantum dots with arbitrarily chosen patterns of zero and single-electrons. The simple operation signals, compatibility with industry fabrication and low spin-environment-interaction in 28 Si/SiGe, promises long-range spin-conserving transport of spin qubits for quantum connectivity in quantum computing architectures.

97 MATHEMATICS AND COMPUTING↗

GPU Direct I/O with HDF5

Exascale HPC systems are being designed with accelerators, such as GPUs, to accelerate parts of applications. In machine learning workloads as well as large-scale simulations that use GPUs as accelerators, the CPU (or host) memory is currently used as a buffer for data transfers between GPU (or device) memory and the file system. If the CPU does not need to operate on the data, then this is sub-optimal because it wastes host memory by reserving space for duplicated data. Furthermore, this “bounce buffer” approach wastes CPU cycles spent on transferring data. A new technique, NVIDIA GPUDirect Storage (GDS), can eliminate the need to use the host memory as a bounce buffer. Thereby, it becomes possible to transfer data directly between the device memory and the file system. This direct data path shortens latency by omitting the extra copy and enables higher-bandwidth. To take full advantage of GDS in existing applications, it is necessary to provide support with existing I/O libraries, such as HDF5 and MPI-IO, which are heavily used in applications. In this paper, we describe our effort of integrating GDS with HDF5, the top I/O library at NERSC and at DOE leadership computing facilities. We design and implement this integration using a HDF5 Virtual File Driver (VFD). The GDS VFD provides a file system abstraction to the application that allows HDF5 applications to perform I/O without the need to move data between CPUs and GPUs explicitly. We compare performance of the HDF5 GDS VFD with explicit data movement approaches and demonstrate superior performance with the GDS method.

Ravi, J↗

Computing rank‐revealing factorizations of matrices stored out‐of‐core

This paper describes efficient algorithms for computing rank-revealing factorizations of matrices that are too large to fit in main memory (RAM), and must instead be stored on slow external memory devices such as disks (out-of-core or out-of-memory). Traditional algorithms for computing rank-revealing factorizations (such as the column pivoted QR factorization and the singular value decomposition) are very communication intensive as they require many vector-vector and matrix-vector operations, which become prohibitively expensive when data is not in RAM. Randomization allows to reformulate new methods so that large contiguous blocks of the matrix are processed in bulk. The paper describes two distinct methods. The first is a blocked version of column pivoted Householder QR, organized as a “left-looking” method to minimize the number of the expensive write operations. The second method results employs a UTV factorization. It is organized as an algorithm-by-blocks to overlap computations and I/O operations. As it incorporates power iterations, it is much better at revealing the numerical rank. Numerical experiments on several computers demonstrate that the new algorithms are almost as fast when processing data stored on slow memory devices as traditional algorithms are for data stored in RAM.

97 MATHEMATICS AND COMPUTING↗

Memory instruction for memory tiers

Various embodiments provide for one or more processor instructions and memory instructions that enable a memory sub-system to copy, move, or swap data across (e.g., between) different memory tiers of the memory sub-system, where each of the memory tiers is associated with different memory locations (e.g., different physical memory locations) on one or more memory devices of the memory sub-system.

Roberts, David Andrew↗

Differential emissivity based evaporable particle measurement

A differential emissivity imaging device for measuring evaporable particle properties can include a heated plate, a thermal camera, a memory device, and an output interface. The heated plate can have an upper surface oriented to receive falling evaporable particles. The evaporable particles have a particle emissivity and the upper surface has a plate surface emissivity. The thermal camera can be oriented to produce a thermal image of the upper surface. A memory device can include instructions that cause the imaging device to calculate a mass of the individual evaporable particle via heat conduction using a calculated surface area and an evaporation time.

Garrett, Tim↗

Two-terminal electronic charge resistance switching device

A two-terminal memory device and methods for its use are provided. In the device, a bottom electrode is electrically continuous with a first operating terminal, and a control gate electrode is electrically continuous with a second operating terminal. A stack of insulator layers comprising a hopping conduction layer and a tunnel layer is contactingly interposed between the bottom electrode and the control gate electrode. The tunnel layer is thinner than the hopping conduction layer, and it has a wider bandgap than the hopping conduction layer. The hopping conduction layer consists of a material that supports electron hopping transport.

Marinella, Matthew↗

Autonomous Multistate Nanoencoding Using Combinatorial Ferroelectric Closure Domains in BiFeO 3

Recent advances in ferroic materials have identified topological defects as promising candidates for enabling additional functionalities in future electronic systems. The generation of stable and customizable polar topologies is needed to achieve multistates that enable beyond-binary device architectures. Here, in this study, we show how to autonomously pattern on-demand highly tunable striped closure domains in pristine rhombohedral-phase BiFeO 3 thin films through precise scanning of a biased atomic force microscopy tip along carefully designed paths. By employing this strategy, we generate and manipulate closed-loop structures with high spatial resolution in an automated manner, allowing the creation of highly tunable and intricate topological domain structures that exhibit distinct polarization configurations without the need for electrode deposition or complex heterostructure growth. As a proof-of-concept for ferroelectric beyond-binary memory devices, we use such topological domains as multistates, engineering an alphabet and automating the symbolic writing/reading process using autonomous microscopy. The resulting information density is compared with that of current commercially available memory devices, demonstrating the potential of ferroelectric topological domains for multistate information storage applications.

BiFeO3↗

A new era of ferroelectric thin films for nonvolatile memories

Ferroelectric films have potential applications in nonvolatile memory devices. In addition to the well-established perovskite-structure oxide ferroelectrics, hafnia- and wurtzite-based ferroelectrics have recently attracted considerable attention because of the improved scalability of the ferroelectric response (down to nanometer thicknesses), their compatibility with silicon fabrication processes, and the availability of deposition methods that realize three-dimensional structures. Furthermore, due to their high compatibility with silicon processes, these are also expected to be used in emerging energy-efficient applications, such as neuromorphic computing and reservoir computing. In this article, we review the fundamentals of hafnia- and wurtzite-based ferroelectrics and their advantages and issues for developing in nonvolatile memory devices. Then, possible emerging applications are discussed.

36 MATERIALS SCIENCE↗

Shape memory embolectomy devices and systems

An embolectomy device comprised of an expansion unit and a support unit is disclosed. The expansion unit can be actuated in response to one or more external stimuli, and the support unit, located proximately to the expansion unit, provides a force to hold the expansion unit in place and to further induce the expansion unit's radial expansion. The radial expansion of the expansion unit causes the expansion unit to physically contact a blood clot, enabling the blood clot to be removed. In some embodiments, the expansion unit can be fabricated from a shape memory polymer foam. In some embodiments the support unit can be fabricated from any elastic material including, without limitation, shape memory alloys.

59 BASIC BIOLOGICAL SCIENCES↗

Field-free switching of perpendicular magnetization in a ferrimagnetic insulator with spin reorientation transition

Writing magnetic bits through spin-orbit torque (SOT) switching is promising for fast and efficient magnetic random-access memory devices. While SOT switching of out-of-plane (OOP) magnetized states requires lateral symmetry breaking, in-plane (IP) magnetized states suffer from low storage density. Here, we demonstrate a field-free switching scheme using a 5-nanometer europium iron garnet film grown with a (110) orientation that shows a spin reorientation transition from OOP to IP above room temperature. This scheme combines the benefits of high-density storage in the OOP states at room temperature and the efficient field-free SOT switching in the IP states at elevated temperatures. While conventional switching of OOP bits faces the dilemma that high OOP anisotropy is required to improve bit stability and low OOP anisotropy is required to lower switching current density, this scheme disentangles this interdependence, allowing for low switching currents to be possible without sacrificing the bit stability, offering opportunities for future memory devices.

Science & Technology - Other Topics↗

Magnetic Switching in Monolayer 2D Diluted Magnetic Semiconductors via Spin‐to‐Spin Conversion

Abstract The integration of 2D van der Waals (vdW) magnets with topological insulators or heavy metals holds great potential for realizing next‐generation spintronic memory devices. However, achieving high‐efficiency spin–orbit torque (SOT) switching of monolayer vdW magnets at room temperature poses a significant challenge, particularly without an external magnetic field. Here, it is shown field‐free, deterministic, and nonvolatile SOT switching of perpendicular magnetization in the monolayer, diluted magnetic semiconductor (DMS), Fe‐doped MoS 2 (Fe:MoS 2 ) at up to 380 K with a current density of ≈7 × 10 4 A cm −2 . The in situ doping of Fe into monolayer MoS 2 via chemical vapor deposition and the geometry‐induced strain in the crystal break the rotational switching symmetry in Fe:MoS 2 , promoting field‐free SOT switching by generating out‐of‐plane spins via spin‐to‐spin conversion. An apparent anomalous Hall effect (AHE) loop shift at a zero in‐plane magnetic field verifies the existence of z spins in Fe:MoS 2 , inducing an antidamping‐like torque that facilitates field‐free SOT switching. This field‐free SOT application using a 2D ferromagnetic monolayer provides a new pathway for developing highly power‐efficient spintronic memory devices.

Chen, Siwei [Department of Mechanical Engineering ↗