Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “matrix multiplication”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Optical implementation of systolic array processing

Algorithms for matrix vector multiplication are implemented using acousto-optic cells for multiplication and input data transfer and using charge coupled devices detector arrays for accumulation and output of the results. No two dimensional matrix mask is required; matrix changes are implemented electronically. A system for multiplying a 50 component nonnegative real vector by a 50 by 50 nonnegative real matrix is described. Modifications for bipolar real and complex valued processing are possible, as are extensions to matrix-matrix multiplication and multiplication of a vector by multiple matrices.

Caulfield, H. J.↗

Parallelization of the Physical-Space Statistical Analysis System (PSAS)

Atmospheric data assimilation is a method of combining observations with model forecasts to produce a more accurate description of the atmosphere than the observations or forecast alone can provide. Data assimilation plays an increasingly important role in the study of climate and atmospheric chemistry. The NASA Data Assimilation Office (DAO) has developed the Goddard Earth Observing System Data Assimilation System (GEOS DAS) to create assimilated datasets. The core computational components of the GEOS DAS include the GEOS General Circulation Model (GCM) and the Physical-space Statistical Analysis System (PSAS). The need for timely validation of scientific enhancements to the data assimilation system poses computational demands that are best met by distributed parallel software. PSAS is implemented in Fortran 90 using object-based design principles. The analysis portions of the code solve two equations. The first of these is the "innovation" equation, which is solved on the unstructured observation grid using a preconditioned conjugate gradient (CG) method. The "analysis" equation is a transformation from the observation grid back to a structured grid, and is solved by a direct matrix-vector multiplication. Use of a factored-operator formulation reduces the computational complexity of both the CG solver and the matrix-vector multiplication, rendering the matrix-vector multiplications as a successive product of operators on a vector. Sparsity is introduced to these operators by partitioning the observations using an icosahedral decomposition scheme. PSAS builds a large (approx. 128MB) run-time database of parameters used in the calculation of these operators. Implementing a message passing parallel computing paradigm into an existing yet developing computational system as complex as PSAS is nontrivial. One of the technical challenges is balancing the requirements for computational reproducibility with the need for high performance. The problem of computational reproducibility is well known in the parallel computing community. It is a requirement that the parallel code perform calculations in a fashion that will yield identical results on different configurations of processing elements on the same platform. In some cases this problem can be solved by sacrificing performance. Meeting this requirement and still achieving high performance is very difficult. Topics to be discussed include: current PSAS design and parallelization strategy; reproducibility issues; load balance vs. database memory demands, possible solutions to these problems.

Larson, J. W.↗

Accuracy and speed in computing the Chebyshev collocation derivative

We studied several algorithms for computing the Chebyshev spectral derivative and compare their roundoff error. For a large number of collocation points, the elements of the Chebyshev differentiation matrix, if constructed in the usual way, are not computed accurately. A subtle cause is is found to account for the poor accuracy when computing the derivative by the matrix-vector multiplication method. Methods for accurately computing the elements of the matrix are presented, and we find that if the entities of the matrix are computed accurately, the roundoff error of the matrix-vector multiplication is as small as that of the transform-recursion algorithm. Results of CPU time usage are shown for several different algorithms for computing the derivative by the Chebyshev collocation method for a wide variety of two-dimensional grid sizes on both an IBM and a Cray 2 computer. We found that which algorithm is fastest on a particular machine depends not only on the grid size, but also on small details of the computer hardware as well. For most practical grid sizes used in computation, the even-odd decomposition algorithm is found to be faster than the transform-recursion method.

Don, Wai-Sun↗

Parallel Preconditioning for CFD Problems on the CM-5

Up to today, preconditioning methods on massively parallel systems have faced a major difficulty. The most successful preconditioning methods in terms of accelerating the convergence of the iterative solver such as incomplete LU factorizations are notoriously difficult to implement on parallel machines for two reasons: (1) the actual computation of the preconditioner is not very floating-point intensive, but requires a large amount of unstructured communication, and (2) the application of the preconditioning matrix in the iteration phase (i.e. triangular solves) are difficult to parallelize because of the recursive nature of the computation. Here we present a new approach to preconditioning for very large, sparse, unsymmetric, linear systems, which avoids both difficulties. We explicitly compute an approximate inverse to our original matrix. This new preconditioning matrix can be applied most efficiently for iterative methods on massively parallel machines, since the preconditioning phase involves only a matrix-vector multiplication, with possibly a dense matrix. Furthermore the actual computation of the preconditioning matrix has natural parallelism. For a problem of size n, the preconditioning matrix can be computed by solving n independent small least squares problems. The algorithm and its implementation on the Connection Machine CM-5 are discussed in detail and supported by extensive timings obtained from real problem data.

Simon, Horst D.↗

A Multiple Sphere T-Matrix Fortran Code for Use on Parallel Computer Clusters

A general-purpose Fortran-90 code for calculation of the electromagnetic scattering and absorption properties of multiple sphere clusters is described. The code can calculate the efficiency factors and scattering matrix elements of the cluster for either fixed or random orientation with respect to the incident beam and for plane wave or localized- approximation Gaussian incident fields. In addition, the code can calculate maps of the electric field both interior and exterior to the spheres.The code is written with message passing interface instructions to enable the use on distributed memory compute clusters, and for such platforms the code can make feasible the calculation of absorption, scattering, and general EM characteristics of systems containing several thousand spheres.

Mackowski, D. W.↗

Communication Lower Bounds and Optimal Algorithms for Symmetric Matrix Computations

In this article, we focus on the communication costs of three symmetric matrix computations: (i) multiplying a matrix with its transpose, known as a symmetric rank-k update (SYRK) (ii) adding the result of the multiplication of a matrix with the transpose of another matrix and the transpose of that result, known as a symmetric rank-2k update (SYR2K) (iii) performing matrix multiplication with a symmetric input matrix (SYMM). All three computations appear in the Level 3 Basic Linear Algebra Subroutines (BLAS) and have wide use in applications involving symmetric matrices. We establish communication lower bounds for these kernels using sequential and distributed-memory parallel computational models, and we show that our bounds are tight by presenting communication-optimal algorithms for each setting. Our lower bound proofs rely on applying a geometric inequality for symmetric computations and analytically solving constrained nonlinear optimization problems. As a result, the symmetric matrix and its corresponding computations are accessed and performed according to a triangular block partitioning scheme in the optimal algorithms.

Al Daas, Hussam [Rutherford Appleton Laboratory, D↗

Mapping unstructured grid computations to massively parallel computers

Investigated here is this mapping problem: assign the tasks of a parallel program to the processors of a parallel computer such that the execution time is minimized. First, a taxonomy of objective functions and heuristics used to solve the mapping problem is presented. Next, we develop a highly parallel heuristic mapping algorithm, called Cyclic Pairwise Exchange (CPE), and discuss its place in the taxonomy. CPE uses local pairwise exchanges of processor assignments to iteratively improve an initial mapping. A variety of initial mapping schemes are tested and recursive spectral bipartitioning (RSB) followed by CPE is shown to result in the best mappings. For the test cases studied here, problems arising in computational fluid dynamics and structural mechanics on unstructured triangular and tetrahedral meshes, RSB and CPE outperform methods based on simulated annealing. Much less time is required to do the mapping and the results obtained are better. Compared with random and naive mappings, RSB and CPE reduce the communication time two fold for the test problems used. Finally, we use CPE in two applications on a CM-2. The first application is a data parallel mesh-vertex upwind finite volume scheme for solving the Euler equations on 2-D triangular unstructured meshes. CPE is used to map grid points to processors. The performance of this code is compared with a similar code on a Cray-YMP and an Intel iPSC/860. The second application is parallel sparse matrix-vector multiplication used in the iterative solution of large sparse linear systems of equations. We map rows of the matrix to processors and use an inner-product based matrix-vector multiplication. We demonstrate that this method is an order of magnitude faster than methods based on scan operations for our test cases.

Hammond, Steven Warren↗

New pole placement algorithm - Polynomial matrix approach

A simple and direct pole-placement algorithm is introduced for dynamical systems having a block companion matrix A. The algorithm utilizes well-established properties of matrix polynomials. Pole placement is achieved by appropriately assigning coefficient matrices of the corresponding matrix polynomial. This involves only matrix additions and multiplications without requiring matrix inversion. A numerical example is given for the purpose of illustration.

Shafai, B.↗

Optical computing and image processing using photorefractive gallium arsenide

Recent experimental results on matrix-vector multiplication and multiple four-wave mixing using GaAs are presented. Attention is given to a simple concept of using two overlapping holograms in GaAs to do two matrix-vector multiplication processes operating in parallel with a common input vector. This concept can be used to construct high-speed, high-capacity, reconfigurable interconnection and multiplexing modules, important for optical computing and neural-network applications.

Cheng, Li-Jen↗

239 Pu R -matrix Analysis and Neutron Multiplicities in the Neutron Energy Region up to a few keVs [Abstract]

The evaluation of 239 Pu neutron resonance parameters coupled to neutron multiplicities $\overline{v}_p$ is of particular importance to investigate the ($\mathcal{n, γf}$) reaction in which a $\mathcal{γ}$-ray emission occurs before the scission of the compound nuclear. This reaction has offered one explanation for the fluctuations in the measured values of $\overline{v}_p$ In this regard, the competition between ($\mathcal{n, γf}$) reaction and the direct fission process can be also included in the R matrix analysis of fission and capture measured data. The goal of this work is the coupled evaluation of the $\mathcal{n}$+ 239 Pu resonance parameters and related neutron multiplicities by ensuring the adoption of thermal neutron constants recently evaluated at the International Atomic Nuclear Energy as well as the recommended (thermal-neutron) induced prompt neutron fission spectrum (PFNS). Moreover, this new set of physical evaluated quantities should also guarantee the agreement for high-leakage solution benchmarks while keeping the good performance of large thermal solution assemblies.

07 ISOTOPE AND RADIATION SOURCES↗

Synthesis of magnesiowüstite nanocrystallites embedded in an amorphous silicate matrix via low energy multiple ion implantations

The synthesis process is presented for experimentally simulating modifications in cosmic dust grains using sequential ion implantations or irradiations followed by thermal annealing. Cosmic silicate dust analogues were prepared via implantation of 20–80 keV Fe - , Mg - , and O - ions into commercially available p-type silicon (100)wafers. The as-implanted analogues are amorphous with a Mg/(Fe+Mg) ratio of 0.5 tailored to match theoretical abundances in circumstellar dusts. Before the ion implantations were performed, Monte-Carlo-based ion-solid interaction codes were used to model the dynamic redistribution of the implanted atoms in the silicon substrate. 600 keV helium ion irradiation was performed on one of the samples before thermal annealing. Two samples were thermally annealed at a temperature appropriate for an M-class stellar wind, 1000 K, for 8.3 h in a vacuum chamber with a pressure of 1 x 10 -7 torr. The elemental depth profiles were extracted utilizing Rutherford Backscattering Spectrometry (RBS) in the samples before and after thermal annealing. X-ray diffraction (XRD)analysis was employed for the identification of various phases in crystalline minerals in the annealed analogues. Transmission electron microscopy (TEM) analysis was utilized to identify specific crystal structures. RBS analysis shows redistribution of the implanted Fe, Mg, and O after thermal annealing due to incorporation into the crystal structures for each sample type. XRD patterns along with TEM analysis showed nanocrystalline Mg and Fe oxides with possible incorporation of additional silicate minerals.

79 ASTRONOMY AND ASTROPHYSICS↗

Satellite-matrix-switched, time-division-multiple-access network simulator

A versatile experimental Ka-band network simulator has been implemented at the NASA Lewis Research Center to demonstrate and evaluate a satellite-matrix-switched, time-division-multiple-access (SMS-TDMA) network and to evaluate future digital ground terminals and radiofrequency (RF) components. The simulator was implemented by using proof-of-concept RF components developed under NASA contracts and digital ground terminal and link simulation hardware developed at Lewis. This simulator provides many unique capabilities such as satellite range delay and variation simulation and rain fade simulation. All network parameters (e.g., signal-to-noise ratio, satellite range variation rate, burst density, and rain fade) are controlled and monitored by a central computer. The simulator is presently configured as a three-ground-terminal SMS-TDMA network.

Ivancic, William D.↗

Satellite-matrix-switched, time-division-multiple-access network simulator

A versatile experimental Ka-band network simulator has been implemented at the NASA Lewis Research Center to demonstrate and evaluate a satellite-matrix-switched, time-division-multiple-access (SMS-TDMA) network and to evaluate future digital ground terminals and radiofrequency (RF) components. The simulator was implemented by using proof-of-concept RF components developed under NASA contracts and digital ground terminal and link simulation hardware developed at Lewis. This simulator provides many unique capabilities such as satellite range delay and variation simulation and rain fade simulation. All network parameters (e.g., signal-to-noise ratio, satellite range variation rate, burst density, and rain fade) are controlled and monitored by a central computer. The simulator is presently configured as a three-ground-terminal SMS-TDMA network.

Ivancic, William D.↗

Visualization of newt aragonitic otoconial matrices using transmission electron microscopy

Otoconia are calcified protein matrices within the gravity-sensing organs of the vertebrate vestibular system. These protein matrices are thought to originate from the supporting or hair cells in the macula during development. Previous studies of mammalian calcitic, barrel-shaped otoconia revealed an organized protein matrix consisting of a thin peripheral layer, a well-defined organic core and a flocculent matrix inbetween. No studies have reported the microscopic organization of the aragonitic otoconial matrix, despite its protein characterization. Pote et al. (1993b) used densitometric methods and inferred that prismatic (aragonitic) otoconia have a peripheral protein distribution, compared to that described for the barrel-shaped, calcitic otoconia of birds, mammals, and the amphibian utricle. By using tannic acid as a negative stain, we observed three kinds of organic matrices in preparations of fixed, decalcified saccular otoconia from the adult newt: (1) fusiform shapes with a homogenous electron-dense matrix; (2) singular and multiple strands of matrix; and (3) more significantly, prismatic shapes outlined by a peripheral organic matrix. These prismatic shapes remain following removal of the gelatinous matrix, revealing an internal array of organic matter. We conclude that prismatic otoconia have a largely peripheral otoconial matrix, as inferred by densitometry.

NASA Discipline Neuroscience↗

Design and Implementation of the PALM-3000 Real-Time Control System

This paper reflects, from a computational perspective, on the experience gathered in designing and implementing realtime control of the PALM-3000 adaptive optics system currently in operation at the Palomar Observatory. We review the algorithms that serve as functional requirements driving the architecture developed, and describe key design issues and solutions that contributed to the system's low compute-latency. Additionally, we describe an implementation of dense matrix-vector-multiplication for wavefront reconstruction that exceeds 95% of the maximum sustained achievable bandwidth on NVIDIA Geforce 8800GTX GPU.

PALM-3000↗

Characterization of Delaminations and Transverse Matrix Cracks in Composite Laminates Using Multiple-Angle Ultrasonic Inspection

Delaminations and transverse matrix cracks often appear concurrently in composite laminates. Normal-incidence ultrasound is excellent at detecting delaminations, but is not optimum for matrix cracks. Non-normal incidence, or polar backscattering, has been shown to optimally detect matrix cracks oriented perpendicular to the ultrasonic plane of incidence. In this work, a series of six composite laminates containing slots were loaded in tension to achieve various levels of delamination and ply cracking. Ultrasonic backscattering was measured over a range of incident polar and azimuthal angles, in order to characterize the relative degree of damage of the two types. Sweptpolar- angle measurements were taken with a curved phased array, as a step toward an array-based approach to simultaneous measurement of combined flaws.

Johnston, Patrick H.↗

Stress Intensity Factor Solutions for Multiple Edge Cracks in Ceramic Matrix Composites

NASA Lewis Research Center conducted a study to determine the stress intensity factor solutions for periodic arrays of bridged cracks for various crack spacings and crack lengths. Initially, the stress intensity factor of an array of unbridged multiple edge cracks was determined under constant global displacement as well as at a point load along the crack wake. These solutions are expected to contribute toward the development of a damage-based life-prediction methodology for CMC engine components.

Ghosn, Louis↗