Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “limited memory”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Efficient packing of patterns in sparse distributed memory by selective weighting of input bits

When a set of patterns is stored in a distributed memory, any given storage location participates in the storage of many patterns. From the perspective of any one stored pattern, the other patterns act as noise, and such noise limits the memory's storage capacity. The more similar the retrieval cues for two patterns are, the more the patterns interfere with each other in memory, and the harder it is to separate them on retrieval. A method is described of weighting the retrieval cues to reduce such interference and thus to improve the separability of patterns that have similar cues.

Kanerva, Pentti↗

UltraLiM: In-Memory Boolean Logic Architecture Using UltraRAM

Conventional computing architectures encounter ‘von Neumann’ and ‘memory wall’ bottlenecks which arise due to the back-and-forth data movement between the physically separate memory and processing units and the speed mismatch between them, respectively. These bottlenecks hurt both energy efficiency and the throughput of computing systems. To address these challenges, in-memory computing architectures have emerged as a promising alternative. They reduce the need for frequent data movement by executing different computing tasks inside the memory system. Here, we present UltraLiM, a logic-in-memory architecture using the UltraRAM-based memory system. UltraRAM holds the promise of developing a ‘universal memory’, overcoming the limitations of charge-based memories thanks to their non-volatile behavior with lower operating voltage. This work presents an in-memory computing architecture that integrates an UltraRAM-based memory array with a custom-designed peripheral circuitry. With this architecture, we can perform various in-memory Boolean logic operations (such as NOT, NAND, NOR, and XOR) in a single cycle. Leveraging the separate read-write paths in the UltraRAM-based memory array, we optimize read operations without encountering design conflicts. This optimization enhances the sense margin, enabling the use of simpler peripheral circuitry for in-memory logic operations.

Alam, Shamiul [University of Tennessee, Knoxville ↗

The persistence of a visual dominance effect in a telemanipulator task: A comparison between visual and electrotactile feedback

The possibility to use an electrotactile stimulation in teleoperation and to observe the interpretation of such information as a feedback to the operator was investigated. It is proposed that visual feedback is more informative than an electrotactile one; and that complex electrotactile feedback slows down both the motor decision and motor response processes, is processed as an all or nothing signal, and bypasses the receptive structure and accesses directly in a working memory where information is sequentially processed and where memory is limited in treatment capacity. The electrotactile stimulation is used as an alerting signal. It is suggested that the visual dominance effect is the result of the advantage of both a transfer function and a sensory memory register where information is pretreated and memorized for a short time. It is found that dividing attention has an effect on the acquisition of the information but not on the subsequent decision processes.

Gaillard, J. P.↗

Very Large Scale Optimization

The purpose of this research under the NASA Small Business Innovative Research program was to develop algorithms and associated software to solve very large nonlinear, constrained optimization tasks. Key issues included efficiency, reliability, memory, and gradient calculation requirements. This report describes the general optimization problem, ten candidate methods, and detailed evaluations of four candidates. The algorithm chosen for final development is a modern recreation of a 1960s external penalty function method that uses very limited computer memory and computational time. Although of lower efficiency, the new method can solve problems orders of magnitude larger than current methods. The resulting BIGDOT software has been demonstrated on problems with 50,000 variables and about 50,000 active constraints. For unconstrained optimization, it has solved a problem in excess of 135,000 variables. The method includes a technique for solving discrete variable problems that finds a "good" design, although a theoretical optimum cannot be guaranteed. It is very scalable in that the number of function and gradient evaluations does not change significantly with increased problem size. Test cases are provided to demonstrate the efficiency and reliability of the methods and software.

Vanderplaats, Garrett↗

Memory-Efficient Onboard Rock Segmentation

Rockster-MER is an autonomous perception capability that was uploaded to the Mars Exploration Rover Opportunity in December 2009. This software provides the vision front end for a larger software system known as AEGIS (Autonomous Exploration for Gathering Increased Science), which was recently named 2011 NASA Software of the Year. As the first step in AEGIS, Rockster-MER analyzes an image captured by the rover, and detects and automatically identifies the boundary contours of rocks and regions of outcrop present in the scene. This initial segmentation step reduces the data volume from millions of pixels into hundreds (or fewer) of rock contours. Subsequent stages of AEGIS then prioritize the best rocks according to scientist- defined preferences and take high-resolution, follow-up observations. Rockster-MER has performed robustly from the outset on the Mars surface under challenging conditions. Rockster-MER is a specially adapted, embedded version of the original Rockster algorithm ("Rock Segmentation Through Edge Regrouping," (NPO- 44417) Software Tech Briefs, September 2008, p. 25). Although the new version performs the same basic task as the original code, the software has been (1) significantly upgraded to overcome the severe onboard re source limitations (CPU, memory, power, time) and (2) "bulletproofed" through code reviews and extensive testing and profiling to avoid the occurrence of faults. Because of the limited computational power of the RAD6000 flight processor on Opportunity (roughly two orders of magnitude slower than a modern workstation), the algorithm was heavily tuned to improve its speed. Several functional elements of the original algorithm were removed as a result of an extensive cost/benefit analysis conducted on a large set of archived rover images. The algorithm was also required to operate below a stringent 4MB high-water memory ceiling; hence, numerous tricks and strategies were introduced to reduce the memory footprint. Local filtering operations were re-coded to operate on horizontal data stripes across the image. Data types were reduced to smaller sizes where possible. Binary- valued intermediate results were squeezed into a more compact, one-bit-per-pixel representation through bit packing and bit manipulation macros. An estimated 16-fold reduction in memory footprint relative to the original Rockster algorithm was achieved. The resulting memory footprint is less than four times the base image size. Also, memory allocation calls were modified to draw from a static pool and consolidated to reduce memory management overhead and fragmentation. Rockster-MER has now been run onboard Opportunity numerous times as part of AEGIS with exceptional performance. Sample results are available on the AEGIS website at http://aegis.jpl.nasa.gov.

Burl, Michael C.↗

Analog Nonvolatile Computer Memory Circuits

In nonvolatile random-access memory (RAM) circuits of a proposed type, digital data would be stored in analog form in ferroelectric field-effect transistors (FFETs). This type of memory circuit would offer advantages over prior volatile and nonvolatile types: In a conventional complementary metal oxide/semiconductor static RAM, six transistors must be used to store one bit, and storage is volatile in that data are lost when power is turned off. In a conventional dynamic RAM, three transistors must be used to store one bit, and the stored bit must be refreshed every few milliseconds. In contrast, in a RAM according to the proposal, data would be retained when power was turned off, each memory cell would contain only two FFETs, and the cell could store multiple bits (the exact number of bits depending on the specific design). Conventional flash memory circuits afford nonvolatile storage, but they operate at reading and writing times of the order of thousands of conventional computer memory reading and writing times and, hence, are suitable for use only as off-line storage devices. In addition, flash memories cease to function after limited numbers of writing cycles. The proposed memory circuits would not be subject to either of these limitations. Prior developmental nonvolatile ferroelectric memories are limited to one bit per cell, whereas, as stated above, the proposed memories would not be so limited. The design of a memory circuit according to the proposal must reflect the fact that FFET storage is only partly nonvolatile, in that the signal stored in an FFET decays gradually over time. (Retention times of some advanced FFETs exceed ten years.) Instead of storing a single bit of data as either a positively or negatively saturated state in a ferroelectric device, each memory cell according to the proposal would store two values. The two FFETs in each cell would be denoted the storage FFET and the control FFET. The storage FFET would store an analog signal value, between the positive and negative FFET saturation values. This signal value would represent a numerical value of interest corresponding to multiple bits: for example, if the memory circuit were designed to distinguish among 16 different analog values, then each cell could store 4 bits. Simultaneously with writing the signal value in the storage FFET, a negative saturation signal value would be stored in the control FFET. The decay of this control-FFET signal from the saturation value would serve as a model of the decay, for use in regenerating the numerical value of interest from its decaying analog signal value. The memory circuit would include addressing, reading, and writing circuitry that would have features in common with the corresponding parts of other memory circuits, but would also have several distinctive features. The writing circuitry would include a digital-to-analog converter (DAC); the reading circuitry would include an analog-to-digital converter (ADC). For writing a numerical value of interest in a given cell, that cell would be addressed, the saturation value would be written in the control FFET in that cell, and the non-saturation analog value representing the numerical value of interest would be generated by use of the DAC and stored in the storage FFET in that cell. For reading the numerical value of interest stored in a given cell, the cell would be addressed, the ADC would convert the decaying control and storage analog signal values to digital values, and an associated fast digital processing circuit would regenerate the numerical value from digital values.

MacLeod, Todd↗

Linear complexity

We present factorization and solution phases for a new linear complexity direct solver designed for concurrent batch operations on fine-grained parallel architectures, for matrices amenable to hierarchical representation. We focus on the strong-admissibility-based $\mathscr{H}^{2}$ format, where strong recursive skeletonization factorization compresses remote interactions. We build upon previous implementations of $\mathscr{H}^{2}$ matrix construction for efficient factorization and solution algorithm design, which are illustrated graphically in stepwise detail. The algorithms are ‘blackbox’ in the sense that the only inputs are the matrix and right-hand side, without analytical or geometrical information about the origin of the system. We demonstrate linear complexity scaling in both time and memory on four representative families of dense matrices up to one million in size. Parallel scaling up to 16 threads is enabled by a multi-level matrix graph coloring and avoidance of dynamic memory allocations thanks to prefix-sum memory management. An experimental backward error analysis is included. We break down the timings of different phases, identify phases that are memory-bandwidth limited, and discuss alternatives for phases that may be sensitive to the trend to employ lower precisions for performance.

Boukaram, Wajih↗

Scaling the memory wall using mixed-precision - HPG-MxP on an exascale-class machine

Mixed-precision algorithms have been proposed as a way for scientific computing to benefit from some of the gains seen for AI on recent high performance computing (HPC) platforms. A few applications dominated by dense matrix operations have seen substantial speedups by utilizing low precision formats such as FP16. However, a majority of scientific simulation applications are memory bandwidth limited. Beyond preliminary studies, the practical gain from using mixed-precision algorithms on a given high-performance computing (HPC) system is largely unclear. The High Performance GMRES Mixed Precision (HPG-MxP) benchmark has been proposed to measure the useful performance of a HPC system on sparse matrix-based mixed-precision applications. In this work, we present an implementation of the HPG-MxP benchmark for an exascale system and describe our algorithm enhancements. We show for the first time a speedup of 1.6x using a combination of double- and single-precision keeping the same residual level on modern GPU-based supercomputers.

Kashi, Aditya [ORNL] (ORCID:0000000325893792)↗

A GPU-based compressible combustion solver for applications exhibiting disparate space and time scales

High-speed chemically active flows pose significant computational challenges due to their disparate space and time scales, with stiff chemistry often dominating simulation time. While modern scientific computing programs achieve exascale performance by leveraging graphics processing units (GPUs), existing GPU-based compressible combustion solvers face critical limitations in memory management, load balancing, and handling the highly localized nature of chemical reactions. To this end, we present a high-performance compressible reacting flow solver built on the AMReX framework and optimized for multi-GPU settings. Here, our approach addresses three GPU performance bottlenecks: memory access patterns through column-major storage optimization, computational workload variability via a bulk-sparse integration strategy for chemical kinetics, and multi-GPU load distribution for adaptive mesh refinement applications. The solver adapts existing matrix-based chemical kinetics formulations to multi-grid contexts. Using representative combustion applications, including 2D and 3D detonations and a 3D jet-in-crossflow configuration, we demonstrate 1.4–5× performance improvements over initial implementations on an in-house cluster of NVIDIA H100 GPUs, and near-ideal weak scaling on the Frontier supercomputer (Oak Ridge Leadership Computing Facility) with up to 1024 AMD Instinct MI250X GPUs. Roofline analysis reveals substantial improvements in arithmetic intensity for both convection (∼ 10 ×) and chemistry (∼ 4 ×) routines, confirming efficient utilization of GPU memory bandwidth and computational resources.

42 ENGINEERING↗

The recovery of the HEAO-2 observatory

In August 1980, the Second High Energy Astronomy Observatory suffered the simultaneous double failure of two gyros, leaving the observatory with one gyro less than what is required for operation. The function of the missing gyro was replaced to allow restoration of spacecraft operation by reprogramming the on-board computer to derive rate from star and sun sensors, and by utilizing complex procedures. Large sun sensor quantization, jump characteristics of the star tracker, and extremely limited computer memory capacity further complicated the problem. Recovery attempts continued for over three months, during which time a suitable rate algorithm was derived. Final verification of the technique was never completed, however, due to the unexpected revival of one of the failed gyros.

Rose, R. E.↗

The computation of steady 3-D separated flows over aerodynamic bodies at incidence and yaw

This paper describes the implementation of a general purpose 3-D NS code and its application to simulated 3-D separated vortical flows over aerodynamic bodies. The thin-layer Reynolds-averaged NS equations are solved by an implicit approximate factorization scheme. The pencil data structure enables the code to run on very fine grids using only limited incore memories. Solutions of a low subsonic flow over an inclined ellipsoid are compared with experimental data to validate the code. Transonic flows over a yawed elliptical wing at incidence are computed and separations occurred at different yaw angles are discussed.

Pulliam, T. H.↗

The use of the QR factorization in the partial realization problem

The use of the QR factorization of the Hankel matrix in solving the partial realization problem is analyzed. Straightforward use of the QR factorization results in a realization scheme that possesses all of the computational advantages of Rissanen's realization scheme. These latter properties are computational efficiency, recursiveness, use of limited computer memory, and the realization of a system triplet having a condensed structure. Moreover, this scheme is robust when the order of the system corresponds to the rank of the Hankel matrix. When this latter condition is violated, an approximate realization could be determined via the QR factorization. In this second scheme, the given Hankel matrix is approximated by a low-rank non-Hankel matrix. Furthermore, it is demonstrated that column pivoting might be incorporated in this second scheme. The results presented are derived for a single input/single output system, but this does not seem to be a restriction.

Verhaegen, M. H.↗

A survey of fault diagnosis technology

Existing techniques and methodologies for fault diagnosis are surveyed. The techniques run the gamut from theoretical artificial intelligence work to conventional software engineering applications. They are shown to define a spectrum of implementation alternatives where tradeoffs determine their position on the spectrum. Various tradeoffs include execution time limitations and memory requirements of the algorithms as well as their effectiveness in addressing the fault diagnosis problem.

Riedesel, Joel↗

Transitioning Autonomous Systems Technology Research to a Flight Software Environment

NASA has developed methods and algorithms for autonomous spacecraft operations,including automated planning and scheduling, fault diagnostics and impact determination,procedure management and display. Making the transition from technology research tooperational flight software requires overcoming significant technical, programmatic andcultural challenges. Technology research is aimed at developing methods that performspecific functions correctly, but the resulting software may not be designed for flightprocessors with limited CPU, memory and network resources, and may not be easilyintegrated into spacecraft flight software. Our objective in the Autonomous Systems andOperations Project is to make significant strides toward the transformation from technologyto operational use. Our focus was twofold: maturing research grade autonomy software intoa flight software environment using broadly accepted languages and tools; and integratingautonomy applications with each other and with representative systems and their data andcommand interfaces. For a target flight software environment, we chose Core FlightSoftware, developed by Goddard Space Flight Center as a common operating systemindependent framework. Our hardware integration environment was provided by theIntegrated Power and Avionics Systems (iPAS) Lab at Johnson Space Center, in whichvarious subsystem development has been conducted to address engineering challenges forthe vehicles and systems required for long-duration missions into the solar system. The iPASand its network of connected facilities provides realistic subsystem hardware or simulationsof spacecraft power, life support, guidance, navigation and control, and command and datahandling subsystems. Interfaces between autonomy applications and the subsystems beingassessed and controlled were developed, assessed and refined. The hardware and softwareenvironment using CFS and the iPAS facility has proven to be a highly flexible and realisticenvironment in which to rapidly integrate applications in an iterative, low cost setting. Usingthe integration environment we have developed, we will turn our focus to performance andsizing analysis to determine the computational requirements for full-scale deployment ofautonomy technology. Scalability of reasoners and the spacecraft models upon which theyoperate, and robustness across the full range of spacecraft conditions and environments willbe explored and improved. We are making significant contributions to the future programsthat will build the spacecraft that will take humans beyond the Earth-Moon system, in whichprogram Systems Engineers will be able to accurately and confidently design in accurate,robust and mature autonomous operations systems.

Flight Software↗

Preserving and Curating the Moon: Adventures in Lunar Core Processing

The lunar crust is the most easily accessible part of the Moon to both remote sensing and sample analyses and provides an archive of information about planetary formation, crustal evolution, and contains a wealth of information about the origin of the Earth-Moon system [e.g., 1-5]. The Apollo mission returned 382 kg of rocks, soil and core samples. Studies of these lunar samples are crucial for our understanding of the Moon’s formation and geological evolution, and for the past 50 years these returned samples have provided the foundation for lunar science [5]. The returned samples are stored and cared for in the lunar curation facility at NASA’s Johnson Space Center. This facility is comprised of a large suite of clean rooms, sample vaults for pristine and return samples, thin section labs, core and saw rooms, storage and working areas, and ancillary labs all designed to minimize contamination from the environment and other samples. Some of the returned samples were intentionally set aside and left unopened. Recently, the Apollo Next Generation Sample Analysis (ANGSA) initiative was designed to examine these pristine samples so the next generation of lunar scientists can further our insight into the Moon’s history. Here, we present the meticulous process that involves preparing for, and ultimately opening, one of the unopened core samples: Apollo 17 drive tube 73002,0,which was collected on the Moon from a landslide deposit near Lara Crater by astronauts Gene Cernan and Jack Schmitt. In order to open, examine, and curate 73002,0withminimalpotential contamination, great care had to be taken prior to opening its container. Beginning18 months before extrusion of the sample, all core processing equipment was pulled out of storage, identified, sorted, cleaned, and purged with nitrogen gas. However, limited institutional memory has made this step challenging as most of the former core processors from the Apollo area have retired or passed away. Twelvemonths prior to extrusion, table-top rehearsals were initiated to identify equipment and learn how it fits together and operates. Five months before extruding the real core, preparations further evolved to include the extrusion and dissection of a lunar core simulant. In addition, a mock-up glovebox was designed and built to allow for a more realistic practice environment. One month prior to extrusion, the actual core cabinet was prepared for use, which included fitting it with lights, a webcam, and power. The tool and equipment cleaning procedure was also modified to include increased cleanliness and sterility requirements. While still sealed, the core was CT scanned at the University of Texas at Austin to maximize its scientific return. Days before the extrusion, witness plates and foil were deployed inside the core cabinet to monitor potential particle and organic contamination within the cabinet. On Nov. 5th, 2019, core sample 73002,0 was successfully opened and extruded(Fig.1). Dissection of 73002,0 began immediately afterwards and is still under way. Processing this sample will help us prepare for future sampling missions and core extrusions and will enable new scientific discoveries about the Moon.

C H Krysher↗

Progress in Computational Aeroelasticity Using High Fidelity Flow and Structural Equations on Parallel Computers

Aeroelasticity which involves strong coupling of fluids, structures and controls is an important element in designing an aircraft. Computational aeroelasticity using low fidelity methods such as the linear aerodynamic flow equations coupled with the modal structural equations are well advanced. Though these low fidelity approaches are computationally less intensive, they are not adequate for the analysis of modern aircraft such as High Speed Civil Transport (HSCT) and Advanced Subsonic Transport (AST) which can experience complex flow/structure interactions. HSCT can experience vortex induced aeroelastic oscillations whereas AST can experience transonic buffet associated structural oscillations. Both aircraft may experience a dip in the flutter speed at the transonic regime. For accurate aeroelastic computations at these complex fluid/structure interaction situations, high fidelity equations such as the Navier-Stokes for fluids and the finite-elements for structures are needed. Computations using these high fidelity equations require large computational resources both in memory and speed. Current conventional supercomputers have reached their limitations both in memory and speed. As a result, parallel computers have evolved to overcome the limitations of conventional computers. This paper will address the transition that is taking place in computational aeroelasticity from conventional computers to parallel computers. The paper will address special techniques needed to take advantage of the architecture of new parallel computers. Results will be illustrated from computations made on iPSC/860 and IBM SP2 computer by using ENASERO code that directly couples the Euler/Navier-Stokes flow equations with high resolution finite-element structural equations.

Guruswamy, Guru P.↗

MFLOP to GFLOP: The Impact on High Fidelity Based Computational Aeroelasticity

Aeroelasticity which involves strong coupling of fluids, structures and controls is an important element in designing an aircraft. Computational aeroelasticity using low fidelity methods such as the linear aerodynamic flow equations coupled with the modal structural equations are well advanced. Though these low fidelity approaches are computationally less intensive, they are not adequate for the analysis of modern aircraft which can experience complex flow/structure interactions. Even at moderate angles of attack supersonic aircraft can experience vortex induced aeroelastic oscillations. Near transonic speeds buffet associated structural oscillations are possible. Aircraft flying in transonic regime may experience a dip in the flutter speed. For accurate aeroelastic computations at these complex fluid/structure interaction situations, high fidelity equations such as the Navier-Stokes for fluids and the finite-elements for structures are needed. Computations using these high fidelity equations require large computational resources both in memory and speed. Current conventional supercomputers have reached their limitations both in memory and speed. As a result, parallel computers have evolved to overco me the limitations of conventional computers. This paper will address the transition that is taking place in computational aeroelasticity from conventional computers to parallel computers. The paper will address special techniques needed to take advantage of the architecture of new parallel computers. Results will be illustrated from computations made on iPSC/860 and IBM SP2 computer by using ENSAERO code that directly couples the Euler/Navier-Stokes flow equations with high resolution finite-element structural equations. Modifications required in both fluids and structural solvers in order to run efficiently on parallel computers will be discussed. Implementation of moving grids and fluid/structural interface on parallel computers will be discussed.

Guruswamy, Guru P.↗

Evaluating the potential of disaggregated memory systems for HPC applications

Summary Disaggregated memory is a promising approach that addresses the limitations of traditional memory architectures by enabling memory to be decoupled from compute nodes and shared across a data center. Cloud platforms have deployed such systems to improve overall system memory utilization, but performance can vary across workloads. High‐performance computing (HPC) is crucial in scientific and engineering applications, where HPC machines also face the issue of underutilized memory. As a result, improving system memory utilization while understanding workload performance is essential for HPC operators. Therefore, learning the potential of a disaggregated memory system before deployment is a critical step. This paper proposes a methodology for exploring the design space of a disaggregated memory system. It incorporates key metrics that affect performance on disaggregated memory systems: memory capacity, local and remote memory access ratio, injection bandwidth, and bisection bandwidth, providing an intuitive approach to guide machine configurations based on technology trends and workload characteristics. We apply our methodology to analyze thirteen diverse workloads, including AI training, data analysis, genomics, protein, fusion, atomic nuclei, and traditional HPC bookends. Our methodology demonstrates the ability to comprehend the potential and pitfalls of a disaggregated memory system and provides motivation for machine configurations. Our results show that eleven of our thirteen applications can leverage injection bandwidth disaggregated memory without affecting performance, while one pays a rack bisection bandwidth penalty and two pay the system‐wide bisection bandwidth penalty. In addition, we also show that intra‐rack memory disaggregation would meet the application's memory requirement and provide enough remote memory bandwidth.

Ding, Nan↗