Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “limited memory”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

HPF Implementation of NPB2.3

We present an HPF implementation of BT, SP, LU, FT, CG and MG of NP132.3-serial benchmark set. The implementation is based on HPF performance model of the benchmark specific primitive operations with distributed arrays. We present profiling and performance data on SGI Origin 2000 and compare the results with NPB2.3. We discuss an HPF limitation in organizing pipelined computations which results in memory and communication overhead in BT and SP.

Frumkin, Michael↗

Continuous Test and Transition Infrastructure for Quantum Networking for Science Complex

We propose a network architecture with a separate quantum dataplane and a conventional control plane: (a) Quantum Data Plane: The Quantum Data plane consists of links of dark fibers connecting quantum switches and repeaters, which in turn, connect to quantum computers, memory and sensors. Since the current reach is limited to local areas, it will begin as a collection of site networks at laboratories, (b) Conventional Control Plane: The Control plane provides management access to quantum devices for configuration and provisioning via control nodes with firewall and encryption capabilities. Continuous Test and Transition Infrastructure: We propose an infrastructure with a control plane connecting multiple sites via encrypted tunnels over conventional networks consisting of the following: (a) Site Quantum Networks: Individual site networks supported by their fiber plants connect laboratories that house and connect to their quantum devices for testing and interoperability.(b) Site Control Planes: Sites are connected over individual control planes with control hosts with Software Defined Networking capabilities, which can be peered with other networks via ESnet.(c) Progressive Expansion: Initially, sites will develop individual data planes with their specific quantum devices and fiber connections, and mechanisms to interface with conventional networks. They progressively expand and interconnect, under a wide-area ecosystem of peered control planes for device testing, interoperability development and roll off into production environments.

Rao, Nageswara S.↗

Design, processing, and testing of lsi arrays for space station

The design of a MOS 256-bit Random Access Memory (RAM) is discussed. Technological achievements comprise computer simulations that accurately predict performance; aluminum-gate COS/MOS devices including a 256-bit RAM with current sensing; and a silicon-gate process that is being used in the construction of a 256-bit RAM with voltage sensing. The Si-gate process increases speed by reducing the overlap capacitance between gate and source-drain, thus reducing the crossover capacitance and allowing shorter interconnections. The design of a Si-gate RAM, which is pin-for-pin compatible with an RCA bulk silicon COS/MOS memory (type TA 5974), is discussed in full. The Integrated Circuit Tester (ICT) is limited to dc evaluation, but the diagnostics and data collecting are under computer control. The Silicon-on-Sapphire Memory Evaluator (SOS-ME, previously called SOS Memory Exerciser) measures power supply drain and performs a minimum number of tests to establish operation of the memory devices. The Macrodata MD-100 is a microprogrammable tester which has capabilities of extensive testing at speeds up to 5 MHz. Beam-lead technology was successfully integrated with SOS technology to make a simple device with beam leads. This device and the scribing are discussed.

Lile, W. R.↗

tobac v1.5: introducing fast 3D tracking, splits and mergers, and other enhancements for identifying and analysing meteorological phenomena

There is a continuously increasing need for reliable feature detection and tracking tools based on objective analysis principles for use with meteorological data. Many tools have been developed over the previous 2 decades that attempt to address this need but most have limitations on the type of data they can be used with, feature computational and/or memory expenses that make them unwieldy with larger datasets, or require some form of data reduction prior to use that limits the tool's utility. The Tracking and Object-Based Analysis of Clouds (tobac) Python package is a modular, open-source tool that improves on the overall generality and utility of past tools. A number of scientific improvements (three spatial dimensions, splits and mergers of features, an internal spectral filtering tool) and procedural enhancements (increased computational efficiency, internal regridding of data, and treatments for periodic boundary conditions) have been included in tobac as a part of the tobac v1.5 update. These improvements have made tobac one of the most robust, powerful, and flexible identification and tracking tools in our field to date and expand its potential use in other fields. Future plans for tobac v2 are also discussed.

54 ENVIRONMENTAL SCIENCES↗

Co-design of Advanced Architectures for Graph Analytics using Machine Learning

A graph is an excellent way of representing relationships among entities. We can use graph analytics to synthesize and analyze such relational data, and extract relevant features that are useful for various tasks such as machine learning. Considering the crucial role of graph analytics in various domains, it is important and timely to investigate the right hardware configurations that can achieve optimal performance for graph workloads on future high-performance computing systems. Design space exploration studies facilitate the selection of appropriate configurations (e.g. memory) to achieve a desired system performance. Recently, the approach of accelerating graph analytics using persistent non-volatile memory has gained a lot of attention. Traditional system simulators such as Gem5 and NVMain can be used to explore the design space of these advanced memory architectures for graph workloads. However, these simulators are slow in execution thus limiting the efficiency of design space exploration studies. To overcome this challenge, we proposed a machine learning based approach to co-design advanced memory architectures for graph workloads. We tested our approach with DRAM, non-volatile memory, and hybrid memory (DRAM+NVM) using a breadth first search benchmark algorithm. Our results showed the applicability of the proposed machine learning based approach to the co-design of the advanced memory architectures. In this paper, we provide recommendations on selecting advanced memory architectures to achieve desired performance for graph workloads. We also discuss the performances of different machine learning models that were considered in this study.

Kurte, Kuldeep↗

Compute in‐Memory with Non‐Volatile Elements for Neural Networks: A Review from a Co‐Design Perspective

Abstract Deep learning has become ubiquitous, touching daily lives across the globe. Today, traditional computer architectures are stressed to their limits in efficiently executing the growing complexity of data and models. Compute‐in‐memory (CIM) can potentially play an important role in developing efficient hardware solutions that reduce data movement from compute‐unit to memory, known as the von Neumann bottleneck. At its heart is a cross‐bar architecture with nodal non‐volatile‐memory elements that performs an analog multiply‐and‐accumulate operation, enabling the matrix‐vector‐multiplications repeatedly used in all neural network workloads. The memory materials can significantly influence final system‐level characteristics and chip performance, including speed, power, and classification accuracy. With an over‐arching co‐design viewpoint, this review assesses the use of cross‐bar based CIM for neural networks, connecting the material properties and the associated design constraints and demands to application, architecture, and performance. Both digital and analog memory are considered, assessing the status for training and inference, and providing metrics for the collective set of properties non‐volatile memory materials will need to demonstrate for a successful CIM technology.

36 MATERIALS SCIENCE↗

Implementation of Ferroelectric Memories for Space Applications

Ferroelectric random access semiconductor memories (FeRAMs) are an ideal nonvolatile solution for space applications. These memories have low power performance, high endurance and fast write times. By combining commercial ferroelectric memory technology with radiation hardened CMOS technology, nonvolatile semiconductor memories for space applications can be attained. Of the few radiation hardened semiconductor manufacturers, none have embraced the development of radiation hardened FeRAMs, due a limited commercial space market and funding limitations. Government funding may be necessary to assure the development of radiation hardened ferroelectric memories for space applications.

Philpy, Stephen C.↗

Dynamically Rendering Rough Terrain with Minimal Memory Overhead

Rendering highly detailed terrain is a process with the potential to consume a great deal of a computer’s random access memory (RAM). In a browser-based application, this resource is limited even further, leading to the necessity to use alternative methods of rendering the large amount of data needed for high detail. This report describes one such method that places the onus of rendering on the speed of the graphics processing unit (GPU) rather than on the computer’s memory. By removing attribute buffers, which contribute greatly to memory costs, from the rendering pipeline and generating the requisite attributes on the fly using a heightmap texture instead, it is estimated that memory usage can be cut down to one-sixth that of the previous method.

Visualization↗

Block encoding of the three-dimensional heterogeneous Poisson equation with application to fracture flow

Quantum linear system (QLS) algorithms offer the potential to solve large-scale linear systems exponentially faster than classical methods. However, applying QLS algorithms to real-world problems remains challenging due to issues such as state preparation, data loading, and efficient information extraction. In this work, we study the feasibility of applying QLS algorithms to solve discretized three-dimensional (3D) heterogeneous Poisson equations, with specific examples relating to groundwater flow through geologic fracture networks. We explicitly construct a block encoding for the 3D heterogeneous Poisson matrix by leveraging the sparse local structure of the discretized operator. While classical solvers benefit from preconditioning, we show that block encoding the system matrix and preconditioner separately does not improve the effective condition number that dominates the QLS run-time. This differs from classical approaches where the preconditioner and the system matrix can often be implemented independently. Nevertheless, due to the structure of the problem in three dimensions, the quantum algorithm achieves a run-time of 𝑂⁡(𝑁 2/3 polylog 𝑁 ⋅log (1/𝜖)), outperforming the best classical methods (with run times of 𝑂⁡(𝑁⁢log 𝑁 ⋅log (1/𝜖))) and offering exponential memory savings. These results highlight both the promise and limitations of QLS algorithms for practical scientific computing, and point to effective condition-number reduction as a key barrier in achieving quantum advantages.

58 GEOSCIENCES↗

Diabatic Eddy Forcing Increases Persistence and Opposes Propagation of the Southern Annular Mode in MERRA-2

Abstract As a dominant mode of jet variability on subseasonal time scales, the Southern Annular Mode (SAM) provides a window into how the atmosphere can produce internal oscillations on longer-than-synoptic time scales. While SAM’s existence can be explained by dry, purely barotropic theories, the time scale for its persistence and propagation is set by a lagged interaction between barotropic and baroclinic mechanisms, making the exact physical mechanisms challenging to identify and to simulate, even in latest generation models. By partitioning the eddy momentum flux convergence in MERRA-2 using an eddy–mean flow interaction framework, we demonstrate that diabatic processes (condensation and radiative heating) are the main contributors to SAM’s persistence in its stationary regime, as well as the key for preventing propagation in this regime. In SAM’s propagating regime, baroclinic and diabatic feedbacks also dominate the eddy–jet feedback. However, propagation is initiated by barotropic shifts in upper-level wave breaking and then sustained by a baroclinic response, leading to a roughly 60-day oscillation period. This barotropic propagation mechanism has been identified in dry, idealized models, but here we show evidence of this mechanism for the first time in reanalysis. The diabatic feedbacks on SAM are consistent with modulation of the storm-track latitude by SAM, altering the emission temperature and cloud cover over individual waves. Therefore, future attempts to improve the SAM time scale in models should focus on the storm-track location, as well as the roles of the cloud and moisture parameterizations. Significance Statement As they circumnavigate the planet, the tropospheric jet streams slowly drift north and south over about 30 days, longer than the normal limit of weather prediction. Understanding the source of this “memory” could improve our knowledge of how the atmosphere organizes itself and our ability to make long-term forecasts. Current theories have identified several possible internal atmospheric interactions responsible for this memory. Yet most of the theories for understanding the jets’ behavior assume that this behavior is only weakly influenced by atmospheric water vapor. We show that this assumption is not enough to understand jet persistence. Instead, clouds and precipitation are more important contributors in reanalysis data than internal “dry” mechanisms to this memory of the Southern Hemisphere jet.

54 ENVIRONMENTAL SCIENCES↗

Towards parallel I/O in finite element simulations

I/O issues in finite element analysis on parallel processors are addressed. Viable solutions for both local and shared memory multiprocessors are presented. The approach is simple but limited by currently available hardware and software systems. Implementation is carried out on a CRAY-2 system. Performance results are reported.

Farhat, Charbel↗

Modeling of supersonic combustor flows using parallel computing

While current 3D CFD codes and modeling techniques have been shown capable of furnishing engineering data for complex scramjet flowfields, the usefulness of such efforts is primarily limited by solutions' CPU time requirements, and secondarily by memory requirements. Attention is presently given to the use of parallel computing capabilities for engineering CFD tools for the analysis of supersonic reacting flows, and to an illustrative incompressible CFD problem using up to 16 iPSC/2 processors with single-domain decomposition.

Riggins, D.↗

R-Hope: Development Approach to Extreme Non-volatile Memory Reuse Onboard the Curiosity Rover

The MSL Curiosity rover landed on Mars on August~5, 2012. Over time, one of its two computers experienced critical hardware memory failure. This non-volatile NAND flash memory held file system partitions and tunable parameters needed for running rover flight software. The project assembled a design and development team to re-purpose a NOR flash memory hardware chip, only 1.5\% of the size of the NAND, to hold the file systems and parameters. The usable NOR memory required major software changes to accommodate the new limitations of slower access speeds, vastly different physical layout, and smaller size. This presentation discusses the approach, challenges, and outcomes of restoring function to the computer so it can act as a ``lifeboat'' in event of problems with the primary computer.

Peper, Nick↗

Best Merge Region Growing Segmentation with Integrated Non-Adjacent Region Object Aggregation

Best merge region growing normally produces segmentations with closed connected region objects. Recognizing that spectrally similar objects often appear in spatially separate locations, we present an approach for tightly integrating best merge region growing with non-adjacent region object aggregation, which we call Hierarchical Segmentation or HSeg. However, the original implementation of non-adjacent region object aggregation in HSeg required excessive computing time even for moderately sized images because of the required intercomparison of each region with all other regions. This problem was previously addressed by a recursive approximation of HSeg, called RHSeg. In this paper we introduce a refined implementation of non-adjacent region object aggregation in HSeg that reduces the computational requirements of HSeg without resorting to the recursive approximation. In this refinement, HSeg s region inter-comparisons among non-adjacent regions are limited to regions of a dynamically determined minimum size. We show that this refined version of HSeg can process moderately sized images in about the same amount of time as RHSeg incorporating the original HSeg. Nonetheless, RHSeg is still required for processing very large images due to its lower computer memory requirements and amenability to parallel processing. We then note a limitation of RHSeg with the original HSeg for high spatial resolution images, and show how incorporating the refined HSeg into RHSeg overcomes this limitation. The quality of the image segmentations produced by the refined HSeg is then compared with other available best merge segmentation approaches. Finally, we comment on the unique nature of the hierarchical segmentations produced by HSeg.

Tilton, James C.↗

Arithmetic Data Cube as a Data Intensive Benchmark

Data movement across computational grids and across memory hierarchy of individual grid machines is known to be a limiting factor for application involving large data sets. In this paper we introduce the Data Cube Operator on an Arithmetic Data Set which we call Arithmetic Data Cube (ADC). We propose to use the ADC to benchmark grid capabilities to handle large distributed data sets. The ADC stresses all levels of grid memory by producing 2d views of an Arithmetic Data Set of d-tuples described by a small number of parameters. We control data intensity of the ADC by controlling the sizes of the views through choice of the tuple parameters.

Frumkin, Michael A.↗

Nanoengineered Shape-Memory Hemostat

Uncontrolled hemorrhage is the predominant cause of preventable combat deaths. Various biomaterials serve as hemostatic agents due to their procoagulant or absorptive activity. However, these biomaterials often lack expansion capabilities, which severely limits use in noncompressible wounds. This study combines a hemostatic nanocomposite with a shape-memory polymer foam to design a composite material with both hemostatic and physical expansion properties. This composite is fabricated in two formulations: a foam externally coated in a highly concentrated nanocomposite (“coated composite”) and a foam containing a diluted nanocomposite infused throughout its pores (“infused composite”). Both formulations retain the shape-memory foam's expansion property. Further, the coated composite shows improved fluid uptake (>2-fold) versus infused composites or foam. The nanocomposite component dissociates from the foam under degradative conditions, with the foam remaining stable for 30 days. Hemostatic studies illustrate that the coated composite reduces the clotting time by ≈20%. Alternatively, the infused composite improves clotting over a larger distance (up to ≈2× distance from the composite). These results signify a modular hemostatic ability: the coated composite reduces clotting and improves fluid uptake, while the infused composite achieves diffuse clotting and maintains mechanical properties. Thus, these materials pose a strong potential for use in noncompressible wounds.

60 APPLIED LIFE SCIENCES↗

Status and prospects of computational fluid dynamics for unsteady transonic viscous flows

Applications of computational aerodynamics to aeronautical research, design, and analysis have increased rapidly over the past decade, and these applications offer significant benefits to aeroelasticians. The past developments are traced by means of a number of specific examples, and the trends are projected over the next several years. The crucial factors that limit the present capabilities for unsteady analyses are identified; they include computer speed and memory, algorithm and solution methods, grid generation, turbulence modeling, vortex modeling, data processing, and coupling of the aerodynamic and structural dynamic analyses. The prospects for overcoming these limitations are presented, and many improvements appear to be readily attainable. If so, a complete and reliable numerical simulation of the unsteady, transonic viscous flow around a realistic fighter aircraft configuration could become possible within the next decade. The possibilities of using artificial intelligence concepts to hasten the achievement of this goal are also discussed.

Mccroskey, W. J.↗

Status and prospects of computational fluid dynamics for unsteady transonic flow

Applications of computational aerodynamics to aeronautical research, design, and analysis have increased rapidly over the past decade, and these applications offer significant benefits to aeroelasticians. The past developments are traced by means of a number of specific examples, and the trends are projected over the next several years. The crucial factors that limit the present capabilities for unsteady analyses are identified; they include computer speed and memory, algorithm and solution methods, grid generation, turbulence modeling, vortex modeling, data processing, and coupling of the aerodynamic and structural dynamic analyses. The prospects for overcoming these limitations are presented, and many improvements appear to be readily attainable. If so, a complete and reliable numerical simulation of the unsteady, transonic viscous flow around a realistic fighter aircraft configuration could become possible within the next decade. The possibilities of using artificial intelligence concepts to hasten the achievement of this goal are also discussed.

Mccroskey, W. J.↗