Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “near memory computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Voyager at Uranus

The engineering changes that had to be made in the Voyager 2 spacecraft in order to enable it to fly beyond the originally planned encounters with Jupiter and Saturn and the underlying engineering strategy leading to the encounter with Uranus are discussed. Fixes of the azimuth actuator failure, receiver failure, and memory failure, and capability upgrades of Image Data Compression, aperture augmentation protective coding, smear reduction, Image Motion Compensation, power management, and contingency planning are summarized. The use of the Computer Command Subsystem in the extension of the voyage is described. The ways in which the strategic priorities, including spacecraft preservation, protection of the near-encounter load, development of new ground and spacecraft capabilities, and repair or circumvention of existing spacecraft faults, were accomplished are reviewed, and the accomplishment of additional tasks is also discussed.

Mclaughlin, W. I.↗

Recent Improvements in Aerodynamic Design Optimization on Unstructured Meshes

Recent improvements in an unstructured-grid method for large-scale aerodynamic design are presented. Previous work had shown such computations to be prohibitively long in a sequential processing environment. Also, robust adjoint solutions and mesh movement procedures were difficult to realize, particularly for viscous flows. To overcome these limiting factors, a set of design codes based on a discrete adjoint method is extended to a multiprocessor environment using a shared memory approach. A nearly linear speedup is demonstrated, and the consistency of the linearizations is shown to remain valid. The full linearization of the residual is used to precondition the adjoint system, and a significantly improved convergence rate is obtained. A new mesh movement algorithm is implemented and several advantages over an existing technique are presented. Several design cases are shown for turbulent flows in two and three dimensions.

Nielsen, Eric J.↗

Counter tube window and X-ray fluorescence analyzer study

A study was performed to determine the best design tube window and X-ray fluorescence analyzer for quantitative analysis of Venusian dust and condensates. The principal objective of the project was to develop the best counter tube window geometry for the sensing element of the instrument. This included formulation of a mathematical model of the window and optimization of its parameters. The proposed detector and instrument has several important features. The instrument will perform a near real-time analysis of dust in the Venusian atmosphere, and is capable of measuring dust layers less than 1 micron thick. In addition, wide dynamic measurement range will be provided to compensate for extreme variations in count rates. An integral pulse-height analyzer and memory accumulate data and read out spectra for detail computer analysis on the ground.

Hertel, R.↗

Enhancement and Extension of Porosity Model in the FDNS-500 Code to Provide Enhanced Simulations of Rocket Engine Components

In the past, the design of rocket engines has primarily relied on the cold flow/hot fire test, and the empirical correlations developed based on the database from previous designs. However, it is very costly to fabricate and test various hardware designs during the design cycle, whereas the empirical model becomes unreliable in designing the advanced rocket engine where its operating conditions exceed the range of the database. The main goal of the 2nd Generation Reusable Launching Vehicle (GEN-II RLV) is to reduce the cost per payload and to extend the life of the hardware, which poses a great challenge to the rocket engine design. Hence, understanding the flow characteristics in each engine components is thus critical to the engine design. In the last few decades, the methodology of computational fluid dynamics (CFD) has been advanced to be a mature tool of analyzing various engine components. Therefore, it is important for the CFD design tool to be able to properly simulate the hot flow environment near the liquid injector, and thus to accurately predict the heat load to the injector faceplate. However, to date it is still not feasible to conduct CFD simulations of the detailed flowfield with very complicated geometries such as fluid flow and heat transfer in an injector assembly and through a porous plate, which requires gigantic computer memories and power to resolve the detailed geometry. The rigimesh (a sintered metal material), utilized to reduce the heat load to the faceplate, is one of the design concepts for the injector faceplate of the GEN-II RLV. In addition, the injector assembly is designed to distribute propellants into the combustion chamber of the liquid rocket engine. A porosity mode thus becomes a necessity for the CFD code in order to efficiently simulate the flow and heat transfer in these porous media, and maintain good accuracy in describing the flow fields. Currently, the FDNS (Finite Difference Navier-Stakes) code is one of the CFD codes which are most widely used by research engineers at NASA Marshall Space Flight Center (MSFC) to simulate various flow problems related to rocket engines. The objective of this research work during the 10-week summer faculty fellowship program was to 1) debug the framework of the porosity model in the current FDNS code, and 2) validate the porosity model by simulating flows through various porous media such as tube banks and porous plate.

Cheng, Gary↗

Scalable parallel communications

Coarse-grain parallelism in networking (that is, the use of multiple protocol processors running replicated software sending over several physical channels) can be used to provide gigabit communications for a single application. Since parallel network performance is highly dependent on real issues such as hardware properties (e.g., memory speeds and cache hit rates), operating system overhead (e.g., interrupt handling), and protocol performance (e.g., effect of timeouts), we have performed detailed simulations studies of both a bus-based multiprocessor workstation node (based on the Sun Galaxy MP multiprocessor) and a distributed-memory parallel computer node (based on the Touchstone DELTA) to evaluate the behavior of coarse-grain parallelism. Our results indicate: (1) coarse-grain parallelism can deliver multiple 100 Mbps with currently available hardware platforms and existing networking protocols (such as Transmission Control Protocol/Internet Protocol (TCP/IP) and parallel Fiber Distributed Data Interface (FDDI) rings); (2) scale-up is near linear in n, the number of protocol processors, and channels (for small n and up to a few hundred Mbps); and (3) since these results are based on existing hardware without specialized devices (except perhaps for some simple modifications of the FDDI boards), this is a low cost solution to providing multiple 100 Mbps on current machines. In addition, from both the performance analysis and the properties of these architectures, we conclude: (1) multiple processors providing identical services and the use of space division multiplexing for the physical channels can provide better reliability than monolithic approaches (it also provides graceful degradation and low-cost load balancing); (2) coarse-grain parallelism supports running several transport protocols in parallel to provide different types of service (for example, one TCP handles small messages for many users, other TCP's running in parallel provide high bandwidth service to a single application); and (3) coarse grain parallelism will be able to incorporate many future improvements from related work (e.g., reduced data movement, fast TCP, fine-grain parallelism) also with near linear speed-ups.

Maly, K.↗

Neural networks for data compression and invariant image recognition

An approach to invariant image recognition (I2R), based upon a model of biological vision in the mammalian visual system (MVS), is described. The complete I2R model incorporates several biologically inspired features: exponential mapping of retinal images, Gabor spatial filtering, and a neural network associative memory. In the I2R model, exponentially mapped retinal images are filtered by a hierarchical set of Gabor spatial filters (GSF) which provide compression of the information contained within a pixel-based image. A neural network associative memory (AM) is used to process the GSF coded images. We describe a 1-D shape function method for coding of scale and rotationally invariant shape information. This method reduces image shape information to a periodic waveform suitable for coding as an input vector to a neural network AM. The shape function method is suitable for near term applications on conventional computing architectures equipped with VLSI FFT chips to provide a rapid image search capability.

Gardner, Sheldon↗

MFLOP to GFLOP: The Impact on High Fidelity Based Computational Aeroelasticity

Aeroelasticity which involves strong coupling of fluids, structures and controls is an important element in designing an aircraft. Computational aeroelasticity using low fidelity methods such as the linear aerodynamic flow equations coupled with the modal structural equations are well advanced. Though these low fidelity approaches are computationally less intensive, they are not adequate for the analysis of modern aircraft which can experience complex flow/structure interactions. Even at moderate angles of attack supersonic aircraft can experience vortex induced aeroelastic oscillations. Near transonic speeds buffet associated structural oscillations are possible. Aircraft flying in transonic regime may experience a dip in the flutter speed. For accurate aeroelastic computations at these complex fluid/structure interaction situations, high fidelity equations such as the Navier-Stokes for fluids and the finite-elements for structures are needed. Computations using these high fidelity equations require large computational resources both in memory and speed. Current conventional supercomputers have reached their limitations both in memory and speed. As a result, parallel computers have evolved to overco me the limitations of conventional computers. This paper will address the transition that is taking place in computational aeroelasticity from conventional computers to parallel computers. The paper will address special techniques needed to take advantage of the architecture of new parallel computers. Results will be illustrated from computations made on iPSC/860 and IBM SP2 computer by using ENSAERO code that directly couples the Euler/Navier-Stokes flow equations with high resolution finite-element structural equations. Modifications required in both fluids and structural solvers in order to run efficiently on parallel computers will be discussed. Implementation of moving grids and fluid/structural interface on parallel computers will be discussed.

Guruswamy, Guru P.↗

FFTs in external or hierarchical memory

A description is given of advanced techniques for computing an ordered FFT on a computer with external or hierarchical memory. These algorithms (1) require as few as two passes through the external data set, (2) use strictly unit stride, long vector transfers between main memory and external storage, (3) require only a modest amount of scratch space in main memory, and (4) are well suited for vector and parallel computation. Performance figures are included for implementations of some of these algorithms on Cray supercomputers. Of interest is the fact that a main memory version outperforms the current Cray library FFT routines on the Cray-2, the Cray X-MP, and the Cray Y-MP systems. Using all eight processors on the Cray Y-MP, this main memory routine runs at nearly 2 Gflops.

Bailey, David H.↗

A Least-Squares Finite Element Method for Electromagnetic Scattering Problems

The least-squares finite element method (LSFEM) is applied to electromagnetic scattering and radar cross section (RCS) calculations. In contrast to most existing numerical approaches, in which divergence-free constraints are omitted, the LSFF-M directly incorporates two divergence equations in the discretization process. The importance of including the divergence equations is demonstrated by showing that otherwise spurious solutions with large divergence occur near the scatterers. The LSFEM is based on unstructured grids and possesses full flexibility in handling complex geometry and local refinement Moreover, the LSFEM does not require any special handling, such as upwinding, staggered grids, artificial dissipation, flux-differencing, etc. Implicit time discretization is used and the scheme is unconditionally stable. By using a matrix-free iterative method, the computational cost and memory requirement for the present scheme is competitive with other approaches. The accuracy of the LSFEM is verified by several benchmark test problems.

Wu, Jie↗

TPSAS-NF1676L-17800-DND

Currently, there are two national challenge problems that guide much of the research in durability and damage tolerance at NASA Langley. The first, Airframe Digital Twin, is a concept that combines as-built vehicle components, as-experienced loads and environments, and other vehicle-specific characteristics to enable ultrahigh fidelity modeling of aircraft and spacecraft throughout their service lives. The second, Materials Genome Initiative, is an analog to the Human Genome Project, and is intended to improve the rate at which materials scientists can discover, understand fundamental physics, and improve material systems. This presentation will highlight several research projects ongoing at NASA Langley that are in support of the above challenge problems. Two of those topics will be the subject of detailed discussion. First, investigations of microstructurally-small fatigue cracking (MSFC) in Al-2Cu and Al-4Cu, fabricated in-house, will be presented. Single- and oligo-crystals of Al-Cu specimens were loaded in uniaxial fatigue, while high-resolution in-situ measurements of deformation were made using image correlation (IC) in a scanning-electron microscope (SEM). The Al-Cu specimens were then replicated as crystal plasticity finite element models (CPFEM), where evolution of slip localization near grain boundaries was computed. Comparison among experiment and CPFEM is made. In addition, XRay diffraction measurements of the as-fabricated specimens were carried out, where direct measurements of the embedded copper precipitates were made, and their influence on growing MSFCs were directly observed. The second main topic will illustrate ongoing work in the area of so-called damage-sensing particles. In this work, shape-memory alloys are embedded in an aluminum alloy matrix. Upon the propagation of a fatigue crack, these particles undergo a strain-induced phase transformation which is detected using an acoustic sensor, providing real-time information on propagating cracks. Experiments and simulations regarding the development of this system will also be detailed.

Jacob Hochhalter↗

Computational Modeling and Experimental Characterization of Martensitic Transformations in Nicoal for Self-Sensing Materials

Fundamental changes to aero-vehicle management require the utilization of automated health monitoring of vehicle structural components. A novel method is the use of self-sensing materials, which contain embedded sensory particles (SP). SPs are micron-sized pieces of shape-memory alloy that undergo transformation when the local strain reaches a prescribed threshold. The transformation is a result of a spontaneous rearrangement of the atoms in the crystal lattice under intensified stress near damaged locations, generating acoustic waves of a specific spectrum that can be detected by a suitably placed sensor. The sensitivity of the method depends on the strength of the emitted signal and its propagation through the material. To study the transition behavior of the sensory particle inside a metal matrix under load, a simulation approach based on a coupled atomistic-continuum model is used. The simulation results indicate a strong dependence of the particle's pseudoelastic response on its crystallographic orientation with respect to the loading direction and suggest possible ways of optimizing particle sensitivity. The technology of embedded sensory particles will serve as the key element in an autonomous structural health monitoring system that will constantly monitor for damage initiation in service, which will enable quick detection of unforeseen damage initiation in real-time and during onground inspections.

Wallace, T. A.↗

Spacecraft computer technology at Southwest Research Institute

Southwest Research Institute (SwRI) has developed and delivered spacecraft computers for a number of different near-Earth-orbit spacecraft including shuttle experiments and SDIO free-flyer experiments. We describe the evolution of the basic SwRI spacecraft computer design from those weighing in at 20 to 25 lb and using 20 to 30 W to newer models weighing less than 5 lb and using only about 5 W, yet delivering twice the processing throughput. Because of their reduced size, weight, and power, these newer designs are especially applicable to planetary instrument requirements. The basis of our design evolution has been the availability of more powerful processor chip sets and the development of higher density packaging technology, coupled with more aggressive design strategies in incorporating high-density FPGA technology and use of high-density memory chips. In addition to reductions in size, weight, and power, the newer designs also address the necessity of survival in the harsh radiation environment of space. Spurred by participation in such programs as MSTI, LACE, RME, Delta 181, Delta Star, and RADARSAT, our designs have evolved in response to program demands to be small, low-powered units, radiation tolerant enough to be suitable for both Earth-orbit microsats and for planetary instruments. Present designs already include MIL-STD-1750 and Multi-Chip Module (MCM) technology with near-term plans to include RISC processors and higher-density MCM's. Long term plans include development of whole-core processors on one or two MCM's.

Shirley, D. J.↗

Achieving High Efficiency in Reduced Order Modeling for Large Scale Polycrystal Plasticity Simulations

Reduced order models for the nonlinear response of heterogeneous microstructures typically require a construction (or training) stage to build the reduced order basis. In this manuscript, an efficient model construction strategy for the eigenstrain homogenization method (EHM) is presented. The proposed strategy relies on a parallel, element-by-element, conjugate gradient solver. Near linear scaling has been achieved with respect to the number of degrees of freedom used to resolve the microstructure. Linear scaling with respect to the number of pre-analyses required to construct the reduced order model (ROM) follows from the EHM formulation. Furthermore, a parallel implementation for fast evaluation of the constructed ROM has been developed using shared memory parallelization. It has been shown that for large microstructures with ≈ 10,000 grains, the total computational cost of evaluating the nonlinear response of a polycrystal could be reduced by approximately an order of magnitude using 32 cores with respect to serial ROM simulation. The present methodology has been verified using an additively manufactured polycrystalline microstructure of a nickel-based superalloy, Inconel 625. The capability of the developed framework to construct a ROM for such large microstructures, as well as the ability of the ROM to predict average and local quantities of interest has been demonstrated.

microscale↗

Developments and Validations of Fully Coupled CFD and Practical Vortex Transport Method for High-Fidelity Wake Modeling in Fixed and Rotary Wing Applications

A novel Computational Fluid Dynamics (CFD) coupling framework using a conventional Reynolds-Averaged Navier-Stokes (BANS) solver to resolve the near-body flow field and a Particle-based Vorticity Transport Method (PVTM) to predict the evolution of the far field wake is developed, refined, and evaluated for fixed and rotary wing cases. For the rotary wing case, the RANS/PVTM modules are loosely coupled to a Computational Structural Dynamics (CSD) module that provides blade motion and vehicle trim information. The PVTM module is refined by the addition of vortex diffusion, stretching, and reorientation models as well as an efficient memory model. Results from the coupled framework are compared with several experimental data sets (a fixed-wing wind tunnel test and a rotary-wing hover test).

Anusonti-Inthra, Phuriwat↗

Fast Machine Learning Lidar Surrogate Simulator: Pristine Clear Sky

The simulations of lidar signals and retrievals rely on a range of optic-physical models, such as radiative transfer models, particle scattering and absorption models, along with the output data from atmospheric physical models. Integrating these different models to represent signals of a lidar system is computationally expensive, and performing backward retrievals can be complex and ambiguous. However, with the advantages of Machine Learning, there is a new potential for building effective lidar signal database linked to corresponding atmospheric profiles. For this project, we are developing a fast pre-trained neural network as the lidar surrogate simulator using simulated data for a CALIPSO-like lidar (355 nm, 532nm, and 1064nm), and a CO2 differential absorption lidar (DIAL) near 1571nm. Specifically, we utilize a long short-term memory (LSTM) model to map the relationships between atmospheric profiles (pressure, temperature, air density and CO2 mixing ratio) and lidar signals. This approach allows us to build machine learning based simulators that can reconstruct lidar signals at specific bands from MERRA reanalysis data, and perform retrievals of atmospheric profiles using lidar signals at various wavelengths. As a first step, the results show the potential of this method to establish a foundational model for sensor signals. This model offers the promise of enabling both accurate predictions and rapid retrievals, providing a more efficient approach to signal processing and analysis.

Shan Zeng↗

A High-Order Finite Spectral Volume Method for Conservation Laws on Unstructured Grids

A time accurate, high-order, conservative, yet efficient method named Finite Spectral Volume (FSV) is developed for conservation laws on unstructured grids. The concept of a 'spectral volume' is introduced to achieve high-order accuracy in an efficient manner similar to spectral element and multi-domain spectral methods. In addition, each spectral volume is further sub-divided into control volumes (CVs), and cell-averaged data from these control volumes is used to reconstruct a high-order approximation in the spectral volume. Riemann solvers are used to compute the fluxes at spectral volume boundaries. Then cell-averaged state variables in the control volumes are updated independently. Furthermore, TVD (Total Variation Diminishing) and TVB (Total Variation Bounded) limiters are introduced in the FSV method to remove/reduce spurious oscillations near discontinuities. A very desirable feature of the FSV method is that the reconstruction is carried out only once, and analytically, and is the same for all cells of the same type, and that the reconstruction stencil is always non-singular, in contrast to the memory and CPU-intensive reconstruction in a high-order finite volume (FV) method. Discussions are made concerning why the FSV method is significantly more efficient than high-order finite volume and the Discontinuous Galerkin (DG) methods. Fundamental properties of the FSV method are studied and high-order accuracy is demonstrated for several model problems with and without discontinuities.

Wang, Z. J.↗

A PC-based hardware implementation of the maximum-likelihood classifier for the Shuttle Ice Detection System

A PC-based near-real time implementation of a two-channel maximum-likelihood classifier is described. The statistical distribution of the reflectance characteristics of ice, frost, and water formation on spray-on-foam-insulation, which covers the External Tank surface of the Space Shuttle, is acquired. The classification technique is based on these statistics. The computer, set in either a training or a classifying mode, learns the statistics of the various classes, or produces a color-coded image denoting the respective categories of classification. The classified results are memory-mapped for efficiency. The speed of the classification process is only limited by the speed of the digital frame grabber and the software that interfaces the frame grabber to the monitor. The process took 4 seconds for a 512 x 480 pixel image.

Jaggi, S.↗

Spacecraft automated operations

Trends in automation of planetary spacecraft are examined using data from missions as far back as Mariner '67 and up to the highly sophisticated Galileo. Nine design considerations which influence the degree of automation such as protection against catastrophic failures, highly repetitive functions, loss of spacecraft communications, and the need for near-real-time adaptivity are discussed. Rapid growth of automation is shown in terms of on-board hardware by plots of number of processors on board, the average speed of processors, and total core memory. The number of commands transmitted from the ground has grown to 5 million bits in Voyager, so that increases in mission complexity have increased both in spacecraft automation and ground operations. Achieving greater automation by transferring ground operations to the spacecraft with the current means of controlling missions, are considered noting proposed changes. For the future, improved computer technology, more microprocessors and increased core storage will be used, and the number of automated functions and their complexity will grow. It is concluded that using the growing computational capability of spacecraft will achieve more autonomy thus reversing the trend of increased mission complexity and cost.

Bird, T. H.↗