Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parallel Matrix Multiplication”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Microsystem Cooler Development

A patented microsystem Stirling cooler is under development with potential application to electronics, sensors, optical and radio frequency (RF) systems, microarrays, and other microsystems. The microsystem Stirling cooler is most suited to volume-limited applications that require cooling below the ambient or sink temperature. Primary components of the planar device include: two diaphragm actuators that replace the pistons found in traditional-scale Stirling machines; and a micro-regenerator that stores and releases thermal energy to the working gas during the Stirling cycle. The use of diaphragms eliminates frictional losses and bypass leakage concerns associated with pistons, while permitting reversal of the hot and cold sides of the device during operation to allow precise temperature control. Three candidate microregenerators were custom fabricated for initial evaluation: two constructed of porous ceramic, and one made of multiple layers of nickel and photoresist in an offset grating pattern. An additional regenerator was prepared with a random stainless steel fiber matrix commonly used in existing Stirling machines for comparison to the custom fabricated regenerators. The candidate regenerators were tested in a piezoelectric-actuated test apparatus designed to simulate the Stirling refrigeration cycle. In parallel with the regenerator testing, electrostatically-driven comb-drive diaphragm actuators for the prototype device have been designed for deep reactive ion etching (DRIE) fabrication.

Moran, Matthew E.↗

Numerically stable algorithm for discrete-ordinate-method radiative transfer in multiple scattering and emitting layered media

The transfer of monochromatic radiation in a scattering, absorbing, and emitting plane-parallel medium with a specified bidirectional reflectivity at the lower boundary is considered. The equations and boundary conditions are summarized. The numerical implementation of the theory is discussed with attention given to the reliable and efficient computation of eigenvalues and eigenvectors. Ways of avoiding fatal overflows and ill-conditioning in the matrix inversion needed to determine the integration constants are also presented.

Stamnes, Knut↗

Sensitivity analysis for large-deflection and postbuckling responses on distributed-memory computers

A computational strategy is presented for calculating sensitivity coefficients for the nonlinear large-deflection and postbuckling responses of laminated composite structures on distributed-memory parallel computers. The strategy is applicable to any message-passing distributed computational environment. The key elements of the proposed strategy are: (1) a multiple-parameter reduced basis technique; (2) a parallel sparse equation solver based on a nested dissection (or multilevel substructuring) node ordering scheme; and (3) a multilevel parallel procedure for evaluating hierarchical sensitivity coefficients. The hierarchical sensitivity coefficients measure the sensitivity of the composite structure response to variations in three sets of interrelated parameters; namely, laminate, layer and micromechanical (fiber, matrix, and interface/interphase) parameters. The effectiveness of the strategy is assessed by performing hierarchical sensitivity analysis for the large-deflection and postbuckling responses of stiffened composite panels with cutouts on three distributed-memory computers. The panels are subjected to combined mechanical and thermal loads. The numerical studies presented demonstrate the advantages of the reduced basis technique for hierarchical sensitivity analysis on distributed-memory machines.

Watson, Brian C.↗

Accelerated Adaptive MGS Phase Retrieval

The Modified Gerchberg-Saxton (MGS) algorithm is an image-based wavefront-sensing method that can turn any science instrument focal plane into a wavefront sensor. MGS characterizes optical systems by estimating the wavefront errors in the exit pupil using only intensity images of a star or other point source of light. This innovative implementation of MGS significantly accelerates the MGS phase retrieval algorithm by using stream-processing hardware on conventional graphics cards. Stream processing is a relatively new, yet powerful, paradigm to allow parallel processing of certain applications that apply single instructions to multiple data (SIMD). These stream processors are designed specifically to support large-scale parallel computing on a single graphics chip. Computationally intensive algorithms, such as the Fast Fourier Transform (FFT), are particularly well suited for this computing environment. This high-speed version of MGS exploits commercially available hardware to accomplish the same objective in a fraction of the original time. The exploit involves performing matrix calculations in nVidia graphic cards. The graphical processor unit (GPU) is hardware that is specialized for computationally intensive, highly parallel computation. From the software perspective, a parallel programming model is used, called CUDA, to transparently scale multicore parallelism in hardware. This technology gives computationally intensive applications access to the processing power of the nVidia GPUs through a C/C++ programming interface. The AAMGS (Accelerated Adaptive MGS) software takes advantage of these advanced technologies, to accelerate the optical phase error characterization. With a single PC that contains four nVidia GTX-280 graphic cards, the new implementation can process four images simultaneously to produce a JWST (James Webb Space Telescope) wavefront measurement 60 times faster than the previous code.

Lam, Raymond K.↗

NAS Experiences of Porting CM Fortran Codes to HPF on IBM SP2 and SGI Power Challenge

Current Connection Machine (CM) Fortran codes developed for the CM-2 and the CM-5 represent an important class of parallel applications. Several users have employed CM Fortran codes in production mode on the CM-2 and the CM-5 for the last five to six years, constituting a heavy investment in terms of cost and time. With Thinking Machines Corporation's decision to withdraw from the hardware business and with the decommissioning of many CM-2 and CM-5 machines, the best way to protect the substantial investment in CM Fortran codes is to port the codes to High Performance Fortran (HPF) on highly parallel systems. HPF is very similar to CM Fortran and thus represents a natural transition. Conversion issues involved in porting CM Fortran codes on the CM-5 to HPF are presented. In particular, the differences between data distribution directives and the CM Fortran Utility Routines Library, as well as the equivalent functionality in the HPF Library are discussed. Several CM Fortran codes (Cannon algorithm for matrix-matrix multiplication, Linear solver Ax=b, 1-D convolution for 2-D datasets, Laplace's Equation solver, and Direct Simulation Monte Carlo (DSMC) codes have been ported to Subset HPF on the IBM SP2 and the SGI Power Challenge. Speedup ratios versus number of processors for the Linear solver and DSMC code are presented.

Saini, Subhash↗

Efficient dynamic simulation for multiple chain robotic mechanisms

An efficient O(mN) algorithm for dynamic simulation of simple closed-chain robotic mechanisms is presented, where m is the number of chains, and N is the number of degrees of freedom for each chain. It is based on computation of the operational space inertia matrix (6 x 6) for each chain as seen by the body, load, or object. Also, computation of the chain dynamics, when opened at one end, is required, and the most efficient algorithm is used for this purpose. Parallel implementation of the dynamics for each chain results in an O(N) + O(log sub 2 m+1) algorithm.

Lilly, Kathryn W.↗

Fuzzy Modeling and Parallel Distributed Compensation for Aircraft Flight Control from Simulated Flight Data

A method is described that combines fuzzy system identification techniques with Parallel Distributed Compensation (PDC) to develop nonlinear control methods for aircraft using minimal a priori knowledge, as part of NASA’s Learn-to-Fly initiative. A fuzzy model was generated with simulated flight data, and consisted of a weighted average of multiple linear time invariant state-space cells having parameters estimated using the equation-error approach and a least-squares estimator. A compensator was designed for each subsystem using Linear Matrix Inequalities (LMI) to guarantee closed-loop stability and performance requirements. This approach is demonstrated using simulated flight data to automatically develop a fuzzy model and design control laws for a simplified longitudinal approximation of the F-16 nonlinear flight dynamics simulation. Results include a comparison of flight data with the estimated fuzzy models and simulations that illustrate the feasibility and utility of the combined fuzzy modeling and control approach.

Weinstein, Rose↗

Recent Advances in Radar Polarimetry and Polarimetric SAR Interferometry

The development of Radar Polarimetry and Radar Interferometry is advancing rapidly, and these novel radar technologies are revamping Synthetic Aperture Radar Imaging decisively. In this exposition the successive advancements are sketched; beginning with the fundamental formulations and high-lighting the salient points of these diverse remote sensing techniques. Whereas with radar polarimetry the textural fine-structure, target-orientation and shape, symmetries and material constituents can be recovered with considerable improvements above that of standard amplitude-only Polarization Radar ; with radar interferometry the spatial (in depth) structure can be explored. In Polarimetric-Interferometric Synthetic Aperture Radar (POL-IN-SAR) Imaging it is possible to recover such co-registered textural plus spatial properties simultaneously. This includes the extraction of Digital Elevation Maps (DEM) from either fully Polarimetric (scattering matrix) or Interferometric (dual antenna) SAR image data takes with the additional benefit of obtaining co-registered three-dimensional POL-IN-DEM information. Extra-Wide-Band POL-IN-SAR Imaging - when applied to Repeat-Pass Image Overlay Interferometry - provides differential background validation and measurement, stress assessment, and environmental stress-change monitoring capabilities with hitherto unattained accuracy, which are essential tools for improved global biomass estimation. More recently, by applying multiple parallel repeat-pass EWB-POL-D(RP)-IN-SAR imaging along stacked (altitudinal) or displaced (horizontal) flight-lines will result in Tomographic (Multi- Interferometric) Polarimetric SAR Stereo-Imaging , including foliage and ground penetrating capabilities. It is shown that the accelerated advancement of these modern EWB-POL-D(RP)-IN-SAR imaging techniques is of direct relevance and of paramount priority to wide-area dynamic homeland security surveillance and local-to-global environmental ground-truth measurement and validation, stress assessment, and stress-change monitoring of the terrestrial and planetary covers. In addition, various closely related topics of (i) acquiring additional and protecting existing spectral windows of the Natural Electromagnetic Spectrum (NES) pertinent to Remote Sensing; (ii) mitigating against common "Radio Frequency Interference (RFI)" and intentional Directive Jamming of Airborne & Space borne POL-IN-SAR Imaging Platforms are appraised.

Boerner, Wolfgang-Martin↗

Layout optimization with algebraic multigrid methods

Finding the optimal position for the individual cells (also called functional modules) on the chip surface is an important and difficult step in the design of integrated circuits. This paper deals with the problem of relative placement, that is the minimization of a quadratic functional with a large, sparse, positive definite system matrix. The basic optimization problem must be augmented by constraints to inhibit solutions where cells overlap. Besides classical iterative methods, based on conjugate gradients (CG), we show that algebraic multigrid methods (AMG) provide an interesting alternative. For moderately sized examples with about 10000 cells, AMG is already competitive with CG and is expected to be superior for larger problems. Besides the classical 'multiplicative' AMG algorithm where the levels are visited sequentially, we propose an 'additive' variant of AMG where levels may be treated in parallel and that is suitable as a preconditioner in the CG algorithm.

Regler, Hans↗

Parallelization of a Six Degree of Freedom Entry Vehicle Trajectory Simulation Using OpenMP and OpenACC

The art and science of writing parallelized software, using methods such as Open Multi-Processing (OpenMP) and Open Accelerators (OpenACC), is dominated by computer scientists. Engineers and non-computer scientists looking to apply these techniques to their project applications face a steep learning curve, especially when looking to adapt their original single threaded software to run multi-threaded on graphics processing units (GPUs). There are significant changes in mindset that must occur; such as how to manage memory, the organization of instructions, and the use of if statements (also known as branching). The purpose of this work is twofold: 1) to demonstrate the applicability of parallelized coding methodologies, OpenMP and OpenACC, to tasks outside of the typical large scale matrix mathematics; and 2) to discuss, from an engineer’s perspective, the lessons learned from parallelizing software using these computer science techniques. This work applies OpenMP, on both multi-core central processing units (CPUs) and Intel® Xeon Phi™ 7210, and OpenACC on GPUs. These parallelization techniques are used to tackle the simulation of thousands of entry vehicle trajectories through the integration of six degree of freedom (DoF) equations of motion (EoM). The forces and moments acting on the entry vehicle, and used by the EoM, are estimated using multiple models of varying levels of complexity. Several benchmark comparisons are made on the execution of six DoF trajectory simulation: single thread Intel® Xeon® E5-2670 CPU, multi-thread CPU using OpenMP, multi-thread Xeon Phi™ 7210 using OpenMP, and multi-thread NVIDIA® Tesla® K40 GPU using OpenACC. These benchmarks are run on the Pleiades Supercomputer Cluster at the National Aeronautics and Space Administration (NASA) Ames Research Center (ARC), and a Xeon Phi™ 7210 node at NASA Langley Research Center (LaRC).

Green, Justin S.↗

NASA Tech Briefs, December 2010

Topics include: Coherent Frequency Reference System for the NASA Deep Space Network; Diamond Heat-Spreader for Submillimeter-Wave Frequency Multipliers; 180-GHz I-Q Second Harmonic Resistive Mixer MMIC; Ultra-Low-Noise W-Band MMIC Detector Modules; 338-GHz Semiconductor Amplifier Module; Power Amplifier Module with 734-mW Continuous Wave Output Power; Multiple Differential-Amplifier MMICs Embedded in Waveguides; Rapid Corner Detection Using FPGAs; Special Component Designs for Differential-Amplifier MMICs; Multi-Stage System for Automatic Target Recognition; Single-Receiver GPS Phase Bias Resolution; Ultra-Wideband Angle-of-Arrival Tracking Systems; Update on Waveguide-Embedded Differential MMIC Amplifiers; Automation Framework for Flight Dynamics Products Generation; Product Operations Status Summary Metrics; Mars Terrain Generation; Application-Controlled Parallel Asynchronous Input/Output Utility; Planetary Image Geometry Library; Propulsion Design With Freeform Fabrication (PDFF); Economical Fabrication of Thick-Section Ceramic Matrix Composites; Process for Making a Noble Metal on Tin Oxide Catalyst; Stacked Corrugated Horn Rings; Refinements in an Mg/MgH2/H2O-Based Hydrogen Generator; Continuous/Batch Mg/MgH2/H2O-Based Hydrogen Generator; Strain System for the Motion Base Shuttle Mission Simulator; Ko Displacement Theory for Structural Shape Predictions; Pyrotechnic Actuator for Retracting Tubes Between MSL Subsystems; Surface-Enhanced X-Ray Fluorescence; Infrared Sensor on Unmanned Aircraft Transmits Time-Critical Wildfire Data; and Slopes To Prevent Trapping of Bubbles in Microfluidic Channels.

Source record↗

Performance Optimization Methods for a Memory-Bound, Unstructured-Grid CFD Application on Massively Parallel GPU Platforms

Computational performance of the FUN3D unstructured-grid computational fluid dynamics (CFD) application on massively parallel GPU environments is memory-bound and highly dependent upon efficient reads from and atomic updates to the irregular cell-, edge-, and node-based data structures. In this talk, we present recent efforts into optimizing select performance-critical kernels on NVIDIA Tesla V100 and A100 GPUs and AMD CDNA MI100 GPUs. A novel use of L2 cache residency controls and asynchronous loads into on-chip shared memory are explored on the A100 GPU for the sparse iterative solver, which is dominated by mixed-precision, sparse matrix vector multiplication. Demonstrations show that these methods improve global memory bandwidth utilization by 13.5% on the A100 GPU. Several techniques are also presented that use registers and/or shared memory to facilitate array transposition and aggregation which combine to reduce the frequency and increase the cache efficiency of floating-point atomic updates to the irregular data structures. These methods are demonstrated to improve the kernel throughput by nearly 500% on select kernels on the AMD MI100 over atomic updates directly to global memory. Overall, both V100 and A100 GPUs outperformed the MI100 GPU on kernels dominated by double-precision atomic updates; however, the techniques demonstrated here reduced the performance gap and improved the MI100 performance.

GPU CPU unstructured CFD memory↗

NASA Tech Briefs, June 2011

Topics covered include: Wind and Temperature Spectrometry of the Upper Atmosphere in Low-Earth Orbit; Health Monitor for Multitasking, Safety-Critical, Real-Time Software; Stereo Imaging Miniature Endoscope; Early Oscillation Detection Technique for Hybrid DC/DC Converters; Parallel Wavefront Analysis for a 4D Interferometer; Schottky Heterodyne Receivers With Full Waveguide Bandwidth; Carbon Nanofiber-Based, High-Frequency, High-Q, Miniaturized Mechanical Resonators; Ultracapacitor-Based Uninterrupted Power Supply System; Coaxial Cables for Martian Extreme Temperature Environments; Using Spare Logic Resources To Create Dynamic Test Points; Autonomous Coordination of Science Observations Using Multiple Spacecraft; Autonomous Phase Retrieval Calibration; EOS MLS Level 1B Data Processing Software, Version 3; Cassini Tour Atlas Automated Generation; Software Development Standard Processes (SDSP); Graphite Composite Panel Polishing Fixture; Material Gradients in Oxygen System Components Improve Safety; Ridge Waveguide Structures in Magnesium-Doped Lithium Niobate; Modifying Matrix Materials to Increase Wetting and Adhesion; Lightweight Magnetic Cooler With a Reversible Circulator; The Invasive Species Forecasting System; Method for Cleanly and Precisely Breaking Off a Rock Core Using a Radial Compressive Force; Praying Mantis Bending Core Breakoff and Retention Mechanism; Scoring Dawg Core Breakoff and Retention Mechanism; Rolling-Tooth Core Breakoff and Retention Mechanism; Vibration Isolation and Stabilization System for Spacecraft Exercise Treadmill Devices; Microgravity-Enhanced Stem Cell Selection; Diagnosis and Treatment of Neurological Disorders by Millimeter-Wave Stimulation; Passive Vaporizing Heat Sink; Remote Sensing and Quantization of Analog Sensors; Phase Retrieval for Radio Telescope and Antenna Control; Helium-Cooled Black Shroud for Subscale Cryogenic Testing; Receive Mode Analysis and Design of Microstrip Reflectarrays; and Chance-Constrained Guidance With Non-Convex Constraints.

Source record↗

Multiscale and Multifidelity Modeling of a 3D Woven Composite Thermal Protection System

Complex three-dimensional (3D) woven composites have been considered by multiple NASA projects in recent years as a means of offering improved mechanical and thermal performance over traditional laminated composite systems. Parallel efforts have focused on developing simulation capabilities for these systems, which have traditionally and heavily relied on experimental testing to evaluate composite performance. One system is the Heatshield for Extreme Entry Environment Technology (HEEET), which is being considered for the thermal protection system on reentry spacecraft. Optical microscopy was used to characterize the blended carbon and phenolic fiber tows. A section of HEEET insulation layer was imaged with high-resolution micro-computed tomography (microCT) and segmented to separate individual tows, porous matrix, and voids. These data were used to develop multiscale thermomechanical computational models within the NASA Multiscale Analysis Tool (NASMAT). Two NASMAT modeling approaches were considered to capture the details of the 3D woven architecture: a coarse model appropriate for inclusion in multiscale structural analyses and a high-fidelity model created by downsampling the microCT data. Both elastic and thermal properties were computed and compared. The feasibility and challenges associated with modeling complex, hybrid 3D woven composites were also addressed.

NASMAT↗

Exploring the Connection Between Sampling Problems in Bayesian Inference and Statistical Mechanics

The Bayesian and statistical mechanical communities often share the same objective in their work - estimating and integrating probability distribution functions (pdfs) describing stochastic systems, models or processes. Frequently, these pdfs are complex functions of random variables exhibiting multiple, well separated local minima. Conventional strategies for sampling such pdfs are inefficient, sometimes leading to an apparent non-ergodic behavior. Several recently developed techniques for handling this problem have been successfully applied in statistical mechanics. In the multicanonical and Wang-Landau Monte Carlo (MC) methods, the correct pdfs are recovered from uniform sampling of the parameter space by iteratively establishing proper weighting factors connecting these distributions. Trivial generalizations allow for sampling from any chosen pdf. The closely related transition matrix method relies on estimating transition probabilities between different states. All these methods proved to generate estimates of pdfs with high statistical accuracy. In another MC technique, parallel tempering, several random walks, each corresponding to a different value of a parameter (e.g. "temperature"), are generated and occasionally exchanged using the Metropolis criterion. This method can be considered as a statistically correct version of simulated annealing. An alternative approach is to represent the set of independent variables as a Hamiltonian system. Considerab!e progress has been made in understanding how to ensure that the system obeys the equipartition theorem or, equivalently, that coupling between the variables is correctly described. Then a host of techniques developed for dynamical systems can be used. Among them, probably the most powerful is the Adaptive Biasing Force method, in which thermodynamic integration and biased sampling are combined to yield very efficient estimates of pdfs. The third class of methods deals with transitions between states described by rate constants. These problems are isomorphic with chemical kinetics problems. Recently, several efficient techniques for this purpose have been developed based on the approach originally proposed by Gillespie. Although the utility of the techniques mentioned above for Bayesian problems has not been determined, further research along these lines is warranted

Pohorille, Andrew↗

Programmable remapper with single flow architecture

The invention relates to image processing systems and methods and in particular to a machine which accepts a real time video image in the form of a matrix of picture elements (pixels) and remaps such image according to a selectable one of a plurality of mapping functions to create an output matrix of pixels. Such mapping functions, or transformations, may be any one of a number of different transformations depending on the objective of the user of the system. The system remaps input images from one coordinate system to another using a set of look-up tables for the data necessary for the transform. The transforms, which are operator selectable, are precomputed and loaded into massive look-up tables. Input pixels, via the look-up tables of any particular transform selected, are mapped into output pixels with the radiance information of the input pixels being appropriately weighted. An earlier embodiment of the system included two parallel processors: a collective processor which mapped multiple input pixels into a single output pixel and an interpolative processor. The interpolative processor performed an interpolation among pixels in the input image where a given input pixel may affect the value of many output pixels. Several advantages are provided over previous embodiments in that the two distinct processors are replaced by a single processor capable of performing both types of operations (collective and interpolative) with no more complexity. Previously, there has existed no image processor or 'remapper' that can operate with sufficient speed and flexibility to permit investigating different transformation patterns in real time.

Fisher, Timothy E.↗

Implementation and Assessment of Advanced Analog Vector-Matrix Processor

This paper discusses the design and implementation of an analog optical vecto-rmatrix coprocessor with a throughput of 128 Mops for a personal computer. Vector matrix calculations are inherently parallel, providing a promising domain for the use of optical calculators. However, to date, digital optical systems have proven too cumbersome to replace electronics, and analog processors have not demonstrated sufficient accuracy in large scale systems. The goal of the work described in this paper is to demonstrate a viable optical coprocessor for linear operations. The analog optical processor presented has been integrated with a personal computer to provide full functionality and is the first demonstration of an optical linear algebra processor with a throughput greater than 100 Mops. The optical vector matrix processor consists of a laser diode source, an acoustooptical modulator array to input the vector information, a liquid crystal spatial light modulator to input the matrix information, an avalanche photodiode array to read out the result vector of the vector matrix multiplication, as well as transport optics and the electronics necessary to drive the optical modulators and interface to the computer. The intent of this research is to provide a low cost, highly energy efficient coprocessor for linear operations. Measurements of the analog accuracy of the processor performing 128 Mops are presented along with an assessment of the implications for future systems. A range of noise sources, including cross-talk, source amplitude fluctuations, shot noise at the detector, and non-linearities of the optoelectronic components are measured and compared to determine the most significant source of error. The possibilities for reducing these sources of error are discussed. Also, the total error is compared with that expected from a statistical analysis of the individual components and their relation to the vector-matrix operation. The sufficiency of the measured accuracy of the processor is compared with that required for a range of typical problems. Calculations resolving alloy concentrations from spectral plume data of rocket engines are implemented on the optical processor, demonstrating its sufficiency for this problem. We also show how this technology can be easily extended to a 100 x 100 10 MHz (200 Cops) processor.

Gary, Charles K.↗

Design of a LQR Controller of Reduced Inputs for Multiple Spacecraft Formation Flying

Regarding multiple spacecraft formation flying, the observation is made that control thrust need only be applied coplanar to the local horizon to achieve complete controllability of a two-satellite formation. Without the need for zenith-nadir (radial) thrust, simplifications and reduction of the weight of the propulsion system may be accomplished. This work focuses on the validation of this radial-excluding control system on its own merits, and in comparison to a related system which does provide thrust parallel to the orbital radius. Simulations are performed using commercial ODE solvers to propagate the Keplerian dynamics of a controlled satellite relative to an uncontrolled, leader satellite. The conclusion is drawn that, despite the exclusion of the radial thrust axis, the remaining control thrust available still provides enough control to design a gain matrix of adequate performance using linear-quadratic regulator (LQR) techniques.

Starin, Scott R.↗