Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “memory mapping”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Identifying location of data granules in global virtual address space

An approach is disclosed that identifies a home node of a data granule. The process is performed by an information handling system (a local node) that retrieves a global virtual address directory. The global virtual address directory maps shared virtual addresses to a number nodes that includes the local node with one of the nodes being the home node. The shared virtual addresses correspond to a plurality of memory addresses that are stored in a shared virtual memory that is shared amongst the plurality of nodes. The approach receives a selected shared virtual address, retrieves, from the global virtual address directory, the home node associated with the selected shared virtual address, and accesses the data granule corresponding to the selected shared virtual address from the home node.

Johns, Charles R.↗

Coordination and establishment of centralized facilities and services for the University of Alaska ERTS survey of the Alaskan environment

The author has identified the following significant results. Specifications have been prepared for the engineering design and construction of a digital color display unit which will be used for automatic processing of ERTS data. The color display unit is a disk refresh memory with computer interfaced input and a color cathode ray tube output display. The system features both analog and digital post disk data manipulation and a versatile color coding device suitable for displaying not only images, but also computer generated graphics such as diagrams, maps, and overlays. Input is from IBM compatible 9 track, 800 BPI tapes, as generated by an IBM 360 computer. ERTS digital tapes are read into the 360, where various analyses such as maximum likelihood classification are performed and the results are written on a magnetic tape which is the input to the color display unit. The greatest versatility in the data manipulation area is provided by the minicomputer built into the color display unit, which is off-line from the main 360 computer. The minicomputer is able to read any line from the refresh disk and place it in its 4K-16 bit memory. Considerable flexibility is available for post-processing enhancement of images by the investigator.

Belon, A. E.↗

ReSpike: A Co-Design Framework for Evaluating SNNs on ReRAM-Based Neuromorphic Processors

With Moore’s law approaching its end, traditional von Neumann architectures are struggling to keep up with the exceeding performance and memory requirements of artificial intelligence and machine learning algorithms. Unconventional computing approaches such as neuromorphic computing that leverage spiking neural networks (SNNs) to perform computation are gaining traction and seek the paradigm shift necessary to sustain the increasing demands of modern applications. Novel memory technologies, such as resistive RAM (ReRAM), employ a crossbar architecture that possesses the inherent capability of efficiently computing vector-matrix multiplication—a dominant operation in SNNs. The prospect of naturally mapping SNNs to the crossbar structures provides a unique opportunity for achieving a high-performance, power-efficient neuromorphic system. In this work, we present ReSpike, which is a new framework, behavioral simulator, and architectural design based on ReRAM crossbar architectures, enabling modeling and co-design to achieve efficient execution of SNNs. We drive this co-design forward by quantifying the impact that ReRAM cell nonidealities have on the corresponding accuracy of an SNN application.

Asifuzzaman, Kazi [ORNL] (ORCID:0000000240044791)↗

Probing boron vacancy defects in hBN via single spin relaxometry

Spin defects in solids offer promising platforms for quantum sensing and memory due to their long coherence times and optical addressability. Here, we integrate a single nitrogen-vacancy (NV) center in diamond with scanning probe microscopy to detect, read out, and spatially map spin-based quantum sensors at the nanoscale. Using the boron vacancy ($V$$^{–}_{B}$) center in hexagonal boron nitride—an emerging two-dimensional spin system—as a model, we detect its electron spin resonance indirectly via changes in the spin relaxation time (T 1 ) of a nearby NV center, eliminating the need for optical excitation or fluorescence detection of the $V$$^{–}_{B}$. Cross-relaxation between NV and $V$$^{–}_{B}$ ensembles significantly reduces NV T1, enabling quantitative nanoscale mapping of defect densities beyond the optical diffraction limit and clear resolution of hyperfine splitting in isotopically enriched h 10 B 15 N. Our method demonstrates interactions between spin sensors in 3D and 2D materials, establishing NV centers as versatile probes for characterizing otherwise inaccessible spin defects.

Quantum metrology↗

Identification of nonlinear system parameters in joints using the force-state mapping technique

A procedure is presented for identifying the potentially strong nonlinear properties of structural members, such as joints, by expressing the force transmitted by the member as a function of its mechanical state. By explicitly including position and rate dependent effects, the surface of transmitted force versus state, the force-state map, has distinct, unique, superposable features for common structural nonlinearities, even those which appear to indicate hysteresis on a force-stroke presentation. An analysis is performed on the influence of true memory effects, transient response, and uncertainty in the measurements and system mass on the precision of the procedure. The successful identification of simulated data verifies the accuracy of the identification algorithm. Tests are then conducted on three actual joint models, with incomplete state measurements typical of an actual testing environment. The ability of the procedure to estimate the complete state vector and to analyze and reconstruct the measured nonlinear characteristics is demonstrated.

Crawley, E. F.↗

3D-ReG: A 3D ReRAM-based Heterogeneous Architecture for Training Deep Neural Networks

Deep neural network (DNN) models are being expanded to a broader range of applications. The computational capability of traditional hardware platforms cannot accommodate the growth of model complexity. Among recent technologies to accelerate DNN, resistive memory (ReRAM)-based processing-in-memory (PIM) emerged as a promising solution for DNN inference due to its high efficiency for matrix-based computation. We face two major technical challenges in extending the use of ReRAM-based accelerators for training: (1) full-precision data is essential in back-propagation; (2) the need to support both feed-forward and back-propagation aggravates the data-movement burden. We propose a heterogeneous architecture named as 3D-ReG, which leverages full-precision GPU to ensure training accuracy and low-overhead 3D integration to provide low-cost data movements. Moreover, we introduce conservative and aggressive task-mapping schemes, which partition the computation phases in different ways to balance execution efficiency and training accuracy. We evaluate 3D-ReG implemented with two 3D integration technologies, through-silicon vias (TSVs) and monolithic inter-tier vias (MIVs), and compare them with GPU-only and PIM-only counterparts. Various GPU-only platforms using two main-memory technologies (DRAM, ReRAM) and three interconnect technologies (2D, TSV, MIV) are evaluated as well. Experimental results show that 3D-ReG can achieve on average 5.64× training speedup and 3.56× higher energy efficiency compared with the GPU with DRAM as main memory, at the cost of 0.05%–3.39% accuracy drop. We define a new metric, gain-loss ratio (GLR), which quantitatively evaluates the capability of a DNN training hardware in terms of the model accuracy and hardware efficiency. The results of our comparison show that the aggressive task-mapping scheme on MIV-based 3D-ReG outperforms the other methods.

Computer Science↗

Tracking algorithms using log-polar mapped image coordinates

The use of log-polar image sampling coordinates rather than conventional Cartesian coordinates offers a number of advantages for visual tracking and docking of space vehicles. Pixel count is reduced without decreasing the field of view, with commensurate reduction in peripheral resolution. Smaller memory requirements and reduced processing loads are the benefits in working environments where bulk and energy are at a premium. Rotational and zoom symmetries of log-polar coordinates accommodate range and orientation extremes without computational penalties. Separation of radial and rotational coordinates reduces the complexity of several target centering algorithms, described below.

Weiman, Carl F. R.↗

A Code-Agnostic Driver Application for Coupled Neutronics and Thermal-Hydraulic Simulations

While the literature has numerous examples of Monte Carlo and computational fluid dynamics (CFD) coupling, most are hard-wired codes intended primarily for research rather than as standalone, general-purpose applications. In this work, we describe an open source application, ENRICO, that enables coupled neutronic and thermal-hydraulic simulations between multiple codes that can be chosen at runtime (as opposed to a coupling between two specific codes). The application has been designed such that the control flow logic, domain mapping, nonlinear fixed-point iteration, solution transfers, and convergence checks are all agnostic to the underlying physics solvers used. Special emphasis has also been placed on enabling efficient execution on distributed-memory computing environments. The transfer of solution fields between solvers is performed in memory rather than through filesystem I/O. Additionally, solvers can be configured to run on overlapping or disjoint sets of processes. To date, coupling with the OpenMC and Shift Monte Carlo codes, the Nek5000 CFD code, and a simplified heat diffusion and subchannel solver has been implemented in ENRICO. We present results for coupled simulations of a single light-water reactor fuel assembly based on the NuScale reactor using various combinations of the physics solvers. For this problem, the coupled simulations are shown to converge in about four Picard iterations. A comparison of the heat source and temperature distributions computed by ENRICO using OpenMC coupled with Nek5000 and Shift coupled with Nek5000 illustrates remarkable agreement between the codes.

42 ENGINEERING↗

SMC 2021 Data Challenge: Analyzing Resource Utilization and User Behavior on Titan Supercomputer

Resource utilization statistics of submitted jobs on a supercomputer can help us understand how users from various scientific domains use HPC platforms and better design a job scheduler. We explore to generate insight regarding workload distribution and usage pattern domains from job scheduler trace, GPU failure information, and project-specific information collected from Titan supercomputer. Furthermore, we want to know how the scheduler performance varies over time and how the users' scheduling behavior changes following a system failure. These observations have the potential to provide valuable insight, which is helpful to prepare for system failures. These practices will help us develop and apply novel machine learning algorithms in understanding system behavior, requirement, and better scheduling of HPC systems. There are two datasets, RUR and GPU: RUR dataset is the job scheduler traces collected from the Titan supercomputer from 01/01/2015 to 07/31/2019 (2015.csv - 2019.csv). These were collected using resource Utilization Report (RUR), a Cray-developed resource-usage data collection and reporting system. It contains the usage information of its critical resources (CPU, Memory, GPU, and I/O) of each running job on Titan during that period (https://ieeexplore.ieee.org/abstract/document/8891001). It includes ProjectAreas as additional information, every job is associated with a project ID. TheProjectAreas.csv dataset provides a mapping of the project ID to its domain science. GPU dataset has information regarding GPU failure on Titan. There have been some hardware-related issues in the GPUs in Titan that caused some GPUs to fail, sometimes irrecoverably during some job runs. This dataset provides information regarding these failures during the execution of the submitted jobs. GPUs on Titan are uniquely identified by a serial number (SN), and they are installed in a location. A GPU can be installed in a location, then removed from that location following a failure, and then re-installed in a different location after fixing the problem. If the failure can't be recovered, the GPU might be removed entirely from Titan. There are two prominent types of failures that resulted in the removal of GPUs from Titan: Double Bit Error (DBE) and Out of the Bus (OTB). The dataset (gc_full.csv) has seven attributes, we provided a short description of these attributes in the ReadMe file. To learn more about this dataset, please refer to the git repository https://github.com/olcf/TitanGPULife and the related publication (https://ieeexplore.ieee.org/abstract/document/9355319).

42 ENGINEERING↗

Non-Reflecting Regions for Finite Difference Methods in Modeling of Elastic Wave Propagation in Plates

Solution of the wave equation using techniques such as finite difference or finite element methods can model elastic wave propagation in solids. This requires mapping the physical geometry into a computational domain whose size is governed by the size of the physical domain of interest and by the required resolution. This computational domain, in turn, dictates the computer memory requirements as well as the calculation time. Quite often, the physical region of interest is only a part of the whole physical body, and does not necessarily include all the physical boundaries. Reduction of the calculation domain requires positioning an artificial boundary or region where a physical boundary does not exist. It is important however that such a boundary, or region, will not affect the internal domain, i.e., it should not cause reflections that propagate back into the material. This paper concentrates on the issue of constructing such a boundary region.

Kishoni, Doron↗

Image applications for coastal resource planning: Elkhorn Slough Pilot Project

The purpose of this project has been to evaluate the utility of digital spectral imagery at two levels of resolution for large scale, accurate, auto-classification of land cover along the Central California Coast. Although remote sensing technology offers obvious advantages over on-the-ground mapping, there are substantial trade-offs that must be made between resolving power and costs. Higher resolution images can theoretically be used to identify smaller habitat patches, but they usually require more scenes to cover a given area and processing these images is computationally intense requiring much more computer time and memory. Lower resolution images can cover much larger areas, are less costly to store, process, and manipulate, but due to their larger pixel size can lack the resolving power of the denser images. This lack of resolving power can be critical in regions such as the Central California Coast where important habitat change often occurs on a scale of 10 meters. Our approach has been to compare vegetation and habitat classification results from two aircraft-based spectral scenes covering the same study area but at different levels of resolution with a previously produced ground-truthed land cover base map of the area. Both of the spectral images used for this project were of significantly higher resolution than the satellite-based LandSat scenes used in the C-CAP program. The lower reaches of the Elkhorn Slough watershed was chosen as an ideal study site because it encompasses a suite of important vegetation types and habitat loss processes characteristic of the central coast region. Dramatic habitat alterations have and are occurring within the Elkhorn Slough drainage area, including erosion and sedimentation, land use conversion, wetland loss, and incremental loss due to development and encroachnnent by agriculture. Additonally, much attention has already been focused on the Elkhorn Slough due to its status as a National Marine Education and Research Reserve and as part of the Monterey Bay National Marine Sanctuary. These destinations have resulted in a rich collection of prior spatial and temporal habitat data.

Kvitek, Rikk G.↗

Incremental Parallelization of Non-Data-Parallel Programs Using the Charon Message-Passing Library

Message passing is among the most popular techniques for parallelizing scientific programs on distributed-memory architectures. The reasons for its success are wide availability (MPI), efficiency, and full tuning control provided to the programmer. A major drawback, however, is that incremental parallelization, as offered by compiler directives, is not generally possible, because all data structures have to be changed throughout the program simultaneously. Charon remedies this situation through mappings between distributed and non-distributed data. It allows breaking up the parallelization into small steps, guaranteeing correctness at every stage. Several tools are available to help convert legacy codes into high-performance message-passing programs. They usually target data-parallel applications, whose loops carrying most of the work can be distributed among all processors without much dependency analysis. Others do a full dependency analysis and then convert the code virtually automatically. Even more toolkits are available that aid construction from scratch of message passing programs. None, however, allows piecemeal translation of codes with complex data dependencies (i.e. non-data-parallel programs) into message passing codes. The Charon library (available in both C and Fortran) provides incremental parallelization capabilities by linking legacy code arrays with distributed arrays. During the conversion process, non-distributed and distributed arrays exist side by side, and simple mapping functions allow the programmer to switch between the two in any location in the program. Charon also provides wrapper functions that leave the structure of the legacy code intact, but that allow execution on truly distributed data. Finally, the library provides a rich set of communication functions that support virtually all patterns of remote data demands in realistic structured grid scientific programs, including transposition, nearest-neighbor communication, pipelining, gather/scatter, and redistribution. At the end of the conversion process most intermediate Charon function calls will have been removed, the non-distributed arrays will have been deleted, and virtually the only remaining Charon functions calls are the high-level, highly optimized communications. Distribution of the data is under complete control of the programmer, although a wide range of useful distributions is easily available through predefined functions. A crucial aspect of the library is that it does not allocate space for distributed arrays, but accepts programmer-specified memory. This has two major consequences. First, codes parallelized using Charon do not suffer from encapsulation; user data is always directly accessible. This provides high efficiency, and also retains the possibility of using message passing directly for highly irregular communications. Second, non-distributed arrays can be interpreted as (trivial) distributions in the Charon sense, which allows them to be mapped to truly distributed arrays, and vice versa. This is the mechanism that enables incremental parallelization. In this paper we provide a brief introduction of the library and then focus on the actual steps in the parallelization process, using some representative examples from, among others, the NAS Parallel Benchmarks. We show how a complicated two-dimensional pipeline-the prototypical non-data-parallel algorithm- can be constructed with ease. To demonstrate the flexibility of the library, we give examples of the stepwise, efficient parallel implementation of nonlocal boundary conditions common in aircraft simulations, as well as the construction of the sequence of grids required for multigrid.

VanderWijngaart, Rob F.↗

Pressure Distribution Over Thick Tapered Airfoils, NACA 81, USA 27c Modified and USA 35

At the request of the United States Army Air Service, the tests reported herein were conducted in the 5-foot atmospheric wind tunnel of the Langley Memorial Aeronautical Laboratory. The object was the measurment of pressures over three representative thick, tapered airfoils which are being used on existing or forthcoming army airplanes. The results are presented in the form of pressure maps, cross-plan load and normal force coefficient curves and load contours. The pressure distribution along the chord was found very similar to that for thin wings, but with a tendency toward greater negative pressures. The characteristics of the loading across the span of the U. S. A. 27 C modified are inferior to those of the other two wings; in the latter the distribution is almost exactly elliptical throughout the usual range of flying angles. The form of tip incorporated in these models is not completely satisfactory and a modification is recommended. (author)

Reid, Elliott G↗

Distributed Macroscopic Traffic Simulation with Open Traffic Models

This paper presents OTM-MPI, an extension of the Open Traffic Models platform (OTM) for running macroscopic traffic simulations in high-performance computing environments. OTM-MPI represents the first open-source, distributed-memory, macroscopic simulation model developed for modern high performance parallel machines and large networks. Macroscopic simulations are appropriate for studying regional traffic scenarios when aggregate trends are of interest, rather than individual vehicle traces. They are also appropriate for studying the routing behavior of classes of vehicles, such as app-informed vehicles. The network partitioning was performed with METIS. Inter-process communication was done with MPI (message-passing interface). Results are provided for two networks: one realistic network which was obtained from Open Street Maps for Chattanooga, TN, and another larger synthetic grid network. The software recorded a speedups of 198x using 256 cores for Chattanooga, and 475x with 1,024 cores for the synthetic network.

macro-scopic traffic simulation↗

Mechanisms enabling reconfigurability and long-term retention in vanadium oxide electrochemical memory

Phase coexistence in nanoscale electrochemical random-access memory (ECRAM) has recently been demonstrated to enable both information storage and extraordinary reconfigurability. These proof-of-principle demonstrations have left the mechanistic details of such a process unresolved. Particularly, the mechanisms that stabilize the multiple phases, and the underlying processes behind sustained memory retention, remain unclear, and are necessary to design such devices. Here we report microscale ECRAM devices composed of V⁢O𝑥, which enables us to directly probe the active region in an operando fashion using optical techniques. Using Raman mapping, we show the phase coexistence driven by the electrochemical injection of O vacancies to be spatially uniform (i.e., with no filaments). The stability was observed to be unusually long, with 1% loss over 14 years in ambient conditions. First-principles calculations of the oxygen vacancy formation energies in V⁢O 𝑥 further support the thermodynamic coexistence of multiple V⁢O 𝑥 phases and clarify the origin of the observed long-term retention in the ECRAM devices. Further, we demonstrate single devices that can be voltage programmed to exhibit synaptic, neuronal, and reconfigurable logic gate functionalities. Furthermore, we not only uncover the phase coexistence mechanism that may help device design, but also demonstrate the circuit-level applications of reconfigurability.

Electrical conductivity↗

Reconfigurable training and reservoir computing in an artificial spin-vortex ice via spin-wave fingerprinting

Strongly interacting artificial spin systems are moving beyond mimicking naturally occurring materials to emerge as versatile functional platforms, from reconfigurable magnonics to neuromorphic computing. Typically, artificial spin systems comprise nanomagnets with a single magnetization texture: collinear macrospins or chiral vortices. Here, by tuning nanoarray dimensions we have achieved macrospin–vortex bistability and demonstrated a four-state metamaterial spin system, the ‘artificial spin-vortex ice’ (ASVI). ASVI can host Ising-like macrospins with strong ice-like vertex interactions and weakly coupled vortices with low stray dipolar field. Vortices and macrospins exhibit starkly differing spin-wave spectra with analogue mode amplitude control and mode frequency shifts of Δf = 3.8 GHz. The enhanced bitextural microstate space gives rise to emergent physical memory phenomena, with ratchet-like vortex injection and history-dependent non-linear fading memory when driven through global magnetic field cycles. We employed spin-wave microstate fingerprinting for rapid, scalable readout of vortex and macrospin populations, and leveraged this for spin-wave reservoir computation. ASVI performs non-linear mapping transformations of diverse input and target signals in addition to chaotic time-series forecasting.

97 MATHEMATICS AND COMPUTING↗

Particle simulation of plasmas on the massively parallel processor

Particle simulations, in which collective phenomena in plasmas are studied by following the self consistent motions of many discrete particles, involve several highly repetitive sets of calculations that are readily adaptable to SIMD parallel processing. A fully electromagnetic, relativistic plasma simulation for the massively parallel processor is described. The particle motions are followed in 2 1/2 dimensions on a 128 x 128 grid, with periodic boundary conditions. The two dimensional simulation space is mapped directly onto the processor network; a Fast Fourier Transform is used to solve the field equations. Particle data are stored according to an Eulerian scheme, i.e., the information associated with each particle is moved from one local memory to another as the particle moves across the spatial grid. The method is applied to the study of the nonlinear development of the whistler instability in a magnetospheric plasma model, with an anisotropic electron temperature. The wave distribution function is included as a new diagnostic to allow simulation results to be compared with satellite observations.

Gledhill, I. M. A.↗

Real-time processor for the Danish airborne SAR

A real-time processor for the Danish high-resolution SAR is presented in terms of its functional performance, algorithm, architecture, and implementation. The real-time processor is mainly intended to assist the operator in using the SAR system, but since the processor has been designed to produce high-quality images, it is expected to make off-line processing superfluous in many cases. The range-Doppler algorithm is adopted and supplemented with an extensive motion compensation, considering the special conditions related to the real-time strip mapping of large scenes. The processor is a pipeline of about 20 elements interconnected by a dedicated data path and a control bus. Only three different types of elements are involved: a programmable signal processing element, a multipurpose memory element, and a multipurpose interface element. The prototypes of these three elements have been tested with satisfactory results.

Dall, J.↗