Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “memory mapping”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Particle displacement tracking applied to air flows

Electronic Particle Image Velocimeter (PIV) techniques offer many advantages over conventional photographic PIV methods such as fast turn around times and simplified data reduction. A new all electronic PIV technique was developed which can measure high speed gas velocities. The Particle Displacement Tracking (PDT) technique employs a single cw laser, small seed particles (1 micron), and a single intensified, gated CCD array frame camera to provide a simple and fast method of obtaining two-dimensional velocity vector maps with unambiguous direction determination. Use of a single CCD camera eliminates registration difficulties encountered when multiple cameras are used to obtain velocity magnitude and direction information. An 80386 PC equipped with a large memory buffer frame-grabber board provides all of the data acquisition and data reduction operations. No array processors of other numerical processing hardware are required. Full video resolution (640x480 pixel) is maintained in the acquired images, providing high resolution video frames of the recorded particle images. The time between data acquisition to display of the velocity vector map is less than 40 sec. The new electronic PDT technique is demonstrated on an air nozzle flow with velocities less than 150 m/s.

Wernet, Mark P.↗

Particle displacement tracking applied to air flows

Electronic Particle Image Velocimetric (PIV) techniques offer many advantages over conventional photographic PIV methods such as fast turn around times and simplified data reduction. A new all electronic PIV technique was developed which can measure high speed gas velocities. The Particle Displacement Tracking (PDT) technique employs a single CW laser, small seed particles (1 micron), and a single intensified, gated CCD array frame camera to provide a simple and fast method of obtaining two-dimensional velocity vector maps with unambiguous direction determination. Use of a single CCD camera eliminates registration difficulties encountered when multiple cameras are used to obtain velocity magnitude and direction information. An 80386 PC equipped with a large memory buffer frame-grabber board provides all of the data acquisition and data reduction operations. No array processors of other numerical processing hardware are required. Full video resolution (640 x 480 pixel) is maintained in the acquired images, providing high resolution video frames of the recorded particle images. The time between data acquisition to display of the velocity vector map is less than 40 sec. The new electronic PDT technique is demonstrated on an air nozzle flow with velocities less than 150 m/s.

Wernet, Mark P.↗

Refining HPCToolkit for application performance analysis at exascale

As part of the US Department of Energy’s Exascale Computing Project (ECP), Rice University has been refining its HPCToolkit performance tools to better support measurement and analysis of applications executing on exascale supercomputers. To efficiently collect performance measurements of GPU-accelerated applications, HPCToolkit employs novel non-blocking data structures to communicate performance measurements between tool threads and application threads. To attribute performance information in detail to source lines, loop nests, and inlined call chains, HPCToolkit performs parallel analysis of large CPU and GPU binaries involved in the execution of an exascale application to rapidly recover mappings between machine instructions and source code. To analyze terabytes of performance measurements gathered during executions at exascale, HPCToolkit employs distributed-memory parallelism, multithreading, sparse data structures, and out-of-core streaming analysis algorithms. To support interactive exploration of profiles up to terabytes in size, HPCToolkit’s hpcviewer graphical user interface uses out-of-core methods to visualize performance data. The result of these efforts is that HPCToolkit now supports collection, analysis, and presentation of profiles and traces of GPU-accelerated applications at exascale. These improvements have enabled HPCToolkit to efficiently measure, analyze and explore terabytes of performance data for executions using as many as 64K MPI ranks and 64K GPU tiles on ORNL’s Frontier supercomputer. HPCToolkit’s support for measurement and analysis of GPU-accelerated applications has been employed to study a collection of open-science applications developed as part of ECP. This paper reports on these experiences, which provided insight into opportunities for tuning applications, strengths and weaknesses of HPCToolkit itself, as well as unexpected behaviors in executions at exascale.

Adhianto, Laksono↗

Analog In-Memory Computing for the Synthetic Aperture Radar Polar Format Algorithm

As the utility of synthetic aperture radar (SAR) systems increases in autonomous vehicles, satellites, and other power- and space-constrained edge applications, there is a growing need for processors that can form SAR images at low power. In recent years, analog in-memory compute (AIMC) has shown immense promise for accelerating neural networks and other matrix-vector multiplication (MVM) heavy workloads at the edge. Here, in this work, we examine how the polar format algorithm (PFA), a popular SAR image formation algorithm, can be mapped to these AIMC systems. The PFA maps readily onto analog MVMs because it primarily consists of two linear operations: interpolation of frequency-domain data to a Cartesian grid, followed by a 2-D Fourier transform. This work presents two approaches to map the interpolation operation onto MVMs in analog hardware: a chirp transform and a modified form of sinc interpolation. These mappings introduce algorithmic errors, and their effect on the quality of SAR image formation is examined, both quantitatively and qualitatively. In addition, the impact of errors introduced by the analog hardware is explored to determine which approach is optimal under varying assumptions about the underlying analog memory devices and circuits.

Analog computing↗

Cricket: A Mapped, Persistent Object Store

This paper describes Cricket, a new database storage system that is intended to be used as a platform for design environments and persistent programming languages. Cricket uses the memory management primitives of the Mach operating system to provide the abstraction of a shared, transactional single-level store that can be directly accessed by user applications. In this paper, we present the design and motivation for Cricket. We also present some initial performance results which show that, for its intended applications, Cricket can provide better performance than a general-purpose database storage system.

Shekita, Eugene↗

Exploring Domain-Wall Pinning in Ferroelectrics via Automated High-Throughput Atomic Force Microscopy

Domain-wall dynamics in ferroelectric materials are strongly position-dependent, since each polar interface is locked into a unique local microstructure. This necessitates spatially resolved studies of wall pinning using scanning-probe microscopy techniques. The pinning centers and pre-existing domain walls are usually sparse within the image plane, precluding the use of dense hyperspectral imaging modes and requiring time-consuming human experimentation. Here, a large-area epitaxial PbTiO 3 film on cubic KTaO 3 was investigated to quantify the electric-field-driven dynamics of the polar–strain domain structures using ML-controlled automated piezoresponse force microscopy. Analysis of 1500 switching events reveals that domain-wall displacement depends not only on field parameters but also on the local ferroelectric–ferroelastic configuration. For example, twin boundaries in polydomains regions, like a 1 – /c+ ∥ a 2 – /c – , stay pinned up to a certain level of bias magnitude and change only marginally as the bias increases from 20 to 30 V, whereas single-variant boundaries, like the a 2 + /c + ∥ a 2 – /c – stack, are already activated at 20 V. These statistics on the possible ferroelectric and ferroelastic wall orientations, together with the automated high-throughput AFM workflow, can be distilled into a predictive map that links domain configurations to pulse parameters. Here, this microstructure-specific rule set forms the foundation for the design of ferroelectric memories.

automated scanning probe microscopy↗

Orbit Operations at 433 Eros: Navigation for the NEAR Shoemaker Mission

NASA's Near Earth Asteroid Rendezvous Mission began its record-setting exploration of the asteroid 433 Eros by inserting the spacecraft into orbit about Eros on February 14, 2000. This is the first spacecraft from any country to orbit an asteroid. The mission has overcome a failed insertion burn attempt on December 20, 1998, an event that would have ended most planetary missions, to return to the same target and successfully begin its science mapping a little more than a year later. Shortly after the successful insertion into orbit, the mission was renamed NEAR Shoemaker (NEAR) in memory of the late astronomer and geologist Eugene Shoemaker. NEAR will gather science data at Eros until February 14, 2001, which is the nominal end of mission. The NEAR mission is managed by the Johns Hopkins University, Applied Physics Laboratory in Laurel, Maryland. Since the initial mission concept in 1992, the design and implementation of the NEAR navigation system have been the responsibility of the Jet Propulsion Laboratory, California Institute of Technology. This presentation will show some of the unique features of navigation and mission design related to orbiting an asteroid and to designing a robust navigation system for the NEAR spacecraft. The problem of navigating a spacecraft about an asteroid is made difficult by the relative uncertainty in the asteroid physical properties which perturb the orbit: i.e., the mass, gravity field, and spin state. To help solve this problem, the navigation system for NEAR uses traditional DSN radio metric Doppler and range tracking, along with new technologies of optical landmark tracking and laser ranging to the asteroid surface. The experiences to date for each of these data types in the navigation solutions will be presented. Plans for the remainder of the NEAR mission will be presented, which include low orbits (down to 35 km radius circular orbits), and close flybys that may pass within 1 km of the surface. In addition, at the end of mission, NASA has approved a controlled descent and hovering phase that will culminate with the spacecraft impacting the surface. The maneuver planning for this final phase will also be presented.

Williams, B. G.↗

On the Development of an Efficient Parallel Hybrid Solver with Application to Acoustically Treated Aero-Engine Nacelles

A finite element solution to the convected Helmholtz equation in a nonuniform flow is used to model the noise field within 3-D acoustically treated aero-engine nacelles. Options to select linear or cubic Hermite polynomial basis functions and isoparametric elements are included. However, the key feature of the method is a domain decomposition procedure that is based upon the inter-mixing of an iterative and a direct solve strategy for solving the discrete finite element equations. This procedure is optimized to take full advantage of sparsity and exploit the increased memory and parallel processing capability of modern computer architectures. Example computations are presented for the Langley Flow Impedance Test facility and a rectangular mapping of a full scale, generic aero-engine nacelle. The accuracy and parallel performance of this new solver are tested on both model problems using a supercomputer that contains hundreds of central processing units. Results show that the method gives extremely accurate attenuation predictions, achieves super-linear speedup over hundreds of CPUs, and solves upward of 25 million complex equations in a quarter of an hour.

Watson, Willie R.↗

Beam controlled nano-robotic device

A system and method (referred to as a method) to fabricate nanorobots. The method generates a pixel map of an atomic object and identifies portions of the atomic object that form a nanorobot. The method stores those identifications in a memory. The method adjusts an electron beam to a noninvasive operating level and images the portions of the atomic object that form the nanorobot. The method executes a plurality of scanning profiles by the electron beam to form the nanorobot and detects nanorobot characteristics and their surroundings via the electron beam in response to executing the plurality of scanning profiles.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

Deep learning simulations of the microwave sky

Here we present 500 high-resolution, full-sky millimeter-wave deep learning (DL) simulations that include lensed CMB maps and correlated foreground components. We find that these MillimeterDL simulations can reproduce a wide range of non-Gaussian summary statistics matching the input training simulations, while only being optimized to match the power spectra. The procedure we develop in this work enables the capability to mass produce independent full-sky realizations from a single expensive full-sky simulation, when ordinarily the latter would not provide enough training data. We also circumvent a common limitation of high-resolution DL simulations that they be confined to small sky areas, often due to memory or GPU issues; we do this by developing a “stitching” procedure that can recover the large-scale, high-order statistics and avoid discontinuities or repeated features in the maps. In addition, since our network takes as input a full-sky lensing convergence map, it can in principle take a full-sky lensing convergence map from any large-scale structure (LSS) simulation and generate the corresponding lensed CMB and correlated foreground components at millimeter wavelengths; this is especially useful in the current era of combining results from both CMB and LSS surveys, which require a common set of simulations.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

A Passive fuel-Cell Surface power System (PaCeSS)

This effort produces a fully-passive surface power generation capability for fuel cell technology. The concept replaces the actively pumped thermal management of a state of the art (SOA) fuel cell/electrolyzer stack with a passive 2-phase thermosyphon for heat transport and a passive shape memory actuating radiator for temperature management. The concept addresses NASA fuel cell technology roadmap needs for greater reliability and longer operating life by eliminating life limiting elements of SOA fuel cell system design that drive failure modes associated with moving parts and control complexity. Applications map to surface power (primary fuel cell), surface energy storage (regenerable fuel cell) and in-situ resource utilization (electrolyzer/unitized reversible fuel cell). The first year of this effort concluded with independent demonstrations of 2-phase thermosyphon heat transport from a simulated fuel cell and temperature driven actuation of a shape memory actuating radiator. Results are in-character with underlying physics and demonstrate the potential to realize the objectives of this concept. his effort produces a fully-passive surface power generation capability for fuel cell technology. The concept replaces the actively pumped thermal management of a state of the art (SOA) fuel cell/electrolyzer stack with a passive 2-phase thermosyphon for heat transport and a passive shape memory actuating radiator for temperature management.

deionized water↗

Design and analysis of CXL performance models for tightly-coupled heterogeneous computing

Truly heterogeneous systems enable partitioned workloads to be mapped to the hardware that nets the best performance. However, current practice requires that inter-device communication between different vendors' hardware use host memory as an intermediary step. To date, there are no widely adopted solutions that allow accelerators to directly transfer data. A new cache-coherent protocol, CXL, aims to facilitate easier, fine-grained sharing between accelerators. In this work we analyze existing methods for designing heterogeneous applications that target GPUs and FPGAs working collaboratively, followed by an exploration to show the benefits of a CXL-enabled system. Specifically, we develop a test application that utilizes both an NVIDIA P100 GPU and a Xilinx U250 FPGA to show current communication limitations. From this application, we capture overall execution time and throughput measurements on the FPGA and GPU. We use these measurements as inputs to novel CXL performance models to show that using CXL caching instead of host memory results in a 1.31X speedup, while a more tightly-coupled pipelined implementation using CXL-enabled hardware would result in a speedup of 1.45X.

Cabrera, Anthony↗

SGI Implementation of MPI-2

MPI-2 is the natural successor to MPI-1 in many ways, and includes 3 major features: parallel I/O (which has a NAS legacy) , remote memory operations (alias one-sided I/O), and dynamic process management. SGI supports a subset of MPI-2. It is useful to have a map of what one can and cannot do with MPI-2 on SGI Origin and Altix systems, to be aware of known problems, and to contrast performance on different platforms.

Nelson, Terry↗

Design of Hopfield Networks Based on Superconducting Coupled Oscillators

The global energy shortage has driven the development of many energy-efficient computational platforms beyond Moore's law, among which brain-inspired neuromorphic computing is one of the promising solutions. Associative memory and pattern recognition are important computations solved by brain-inspired Hopfield networks. Classical Hopfield networks store memories via fixed point attractors of their dynamics. In oscillatory Hopfield networks, these attractors are replaced by periodic orbits. Here, we design an oscillatory Hopfield network based on coupled superconducting oscillators. We first employ a mathematical phase reduction approach to map networks of coupled superconducting rapid single flux quantum (RSFQ) ring oscillators to coupled Kuramoto phase-oscillator networks. We use this theory to numerically optimize the hardware's mutual inductances in order to directly match the phase-reduced superconducting oscillators to a model of phase-oscillator-based Hopfield networks. The resulting network can store multiple oscillatory phase-locked memory patterns and recover the patterns based on the initial phase conditions. As different pattern recognition tasks, or learning, require tunable connectivity strengths between the oscillatory nodes, we further employ a coupler circuit that enables tuning the coupling strength between two oscillators by applying an external flux. We demonstrate the functionality of our design through numerical simulations of a small example network with oscillators operating at 86 GHz and recognizing patterns within 10 ns. Our approach enables the learning and retrieval of dynamical memory patterns with a wide range of applications where rhythmic dynamic output is beneficial.

Cheng, Ran↗

Data-driven based coordinated smart inverter control for distributed energy resources

Smart inverters (SI) for distributed energy resources (DER) are becoming popular since they have the ability to stabilize as well as restore the voltage and frequency of power systems. Aiming at establishing the mathematical models combined with SI control methods, multiple optimization methods are developed. However, the computational complexity of solving such a mathematical model with various uncertainties limits the real-time application of the SI control. To conquer this challenge, a data-driven-based SI control approach is developed to achieve coordinated control in the high penetration DER system. First, an optimization problem for maximizing the active power generation and minimizing the power loss is designed using the Volt/VAR control. To reduce the time consumption, the recurrent neural network (RNN) is proposed to model the relationship between the uncertainties and control actions during the offline site. The RNN with different sub-structures such as the long short-term memory cell and gated recurrent unit cell are included to enrich the diversity of features. In the last stage, different experiment comparisons, including multiple uncertainties maps and stateof- art machine learning methods, are conducted to verify the effectiveness of the proposed method based on the IEEE 123 bus power system. The results demonstrate that the proposed method can effectively achieve a rapid and coordinated control with a lower error rate.

Qiu, Wei↗

Architecture-Aware Models of AI Engines for High-Performance Matrix Matrix Multiplication

The AI Engine (AIE) architecture, available in systems from mobile SoCs to server-class FPGAs, aims to efficiently execute AI/ML tasks through a two-dimensional array of compute tiles. Previous work on AIEs has explored different approaches to mapping computation across spatial arrays, but the compute kernel running on each tile has not been the focus. Additionally, the AIE-ML architecture introduces memory tiles and omits programmable logic, requiring new approaches to staging and moving data throughout the array. In this work we update analytical models developed for CPUs to produce the design of high performance kernels while introducing new model considerations such as memory structure, throughput, and latency as required by the AIE hardware. We evaluate our models by developing AIE-ML kernels for matrix multiplication in low-precision data types showing performance up to 95% of compute peak for the kernel when data resides in local memory and above 90% of compute peak when data resides in main memory.

Binder, Elliott D. [Carnegie Mellon University, Pi↗

SMC 2021 : Analyzing Resource Utilization and User Behavior on Titan Supercomputer

Resource utilization statistics of submitted jobs on a supercomputer can help us understand how users from various scientific domains use HPC platforms and better design a job scheduler. We explore to generate insight regarding workload distribution and usage pattern domains from job scheduler trace, GPU failure information, and project-specific information collected from Titan supercomputer. Furthermore, we want to know how the scheduler performance varies over time and how the users' scheduling behavior changes following a system failure. These observations have the potential to provide valuable insight, which is helpful to prepare for system failures. These practices will help us develop and apply novel machine learning algorithms in understanding system behavior, requirement, and better scheduling of HPC systems. There are two datasets, RUR and GPU. RUR: This dataset is the job scheduler traces collected from the Titan supercomputerfrom 01/01/2015 to 07/31/2019 (2015.csv - 2019.csv). These were collected usingResource Utilization Report (RUR), a Cray-developed resource-usage data collectionand reporting system. It contains the usage information of its critical resources (CPU,Memory, GPU, and I/O) of each running job on Titan during that period [2]. ProjectAreas: Every job is associated with a project ID. TheProjectAreas.csvdatasetprovides a mapping of the project ID to its domain science. GPU: There have been some hardware-related issues in the GPUs in Titan that caused some GPUs to fail, sometimes irrecoverably during some job runs. This dataset provides information regarding these failures during the execution of the submitted jobs. GPUs on Titan are uniquely identified by a serial number (SN), and they are installed in a location. A GPU can be installed in a location, then removed from that location following a failure, and then re-installed in a different location after fixing the problem. If the failure can't be recovered, the GPU might be removed entirely from Titan. There are two prominent types of failures that resulted in the removal of GPUs from Titan: Double Bit Error (DBE) and Out of the Bus (OTB). The dataset (gc_full.csv) has the following fields: 1. SN : Serial number of a GPU 2. location : The location where it is installed 3. insert : The time when it was inserted into that location 4. remove : The time when it was removed from that location 5. duration : Amount of time the GPU spent in this location 6. out : If the device was taken out entirely w/o a re-installment into a new location. 7. event : If the GPU was taken out entirely, the reason for its removal. To learn more about this dataset, please refer to the git repositoryhttps://github.com/olcf/TitanGPULifeand the related publication [1]. References [1] George Ostrouchov, Don Maxwell, Rizwan A Ashraf, Christian Engelmann, MallikarjunShankar, and James H Rogers. Gpu lifetimes on titan supercomputer: Survival analysisand reliability. InSC20: International Conference for High Performance Computing,Networking, Storage and Analysis, pages 1-14. IEEE, 2020. [2] Feiyi Wang, Sarp Oral, Satyabrata Sen, and Neena Imam. Learning from five-yearresource-utilization data of titan system. In2019 IEEE International Conference onCluster Computing (CLUSTER), pages 1-6. IEEE, 2019.

42 ENGINEERING↗

Multiple degree of freedom force-state component identification

Force-state mapping is extended to allow the quasistatic characterization of realistic multiple degree of freedom systems whose constitutive relations are nonlinear, dissipative, coupled and depend on memory of past states. A general constitutive force-state model is formulated, which includes dynamic hysteresis phenomena typical of that found in deployable structures. A two step parameter identification approach is developed, which employs a linear extended Kalman filter and a nonlinear least squares fit. A six degree of freedom force-state testing device was constructed, and used to obtain data on one bay of a simple truss and on one bay of the MODE deployable truss.

Masters, Brett P.↗