Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Matrix factorization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Factorization Of Positive Definite, Banded Hermitian Matrices

Report discusses application of Cholesky factorization algorithm to positive definite, banded Hermitian matrices. Begins by extending Cholesky factorization algorithm to cover uniformly-partitioned, banded, positive definite matrices of rank n that is real symmetric or Hermitian. Then two stratagems given for use of algorithm in concurrent-processing system in which N less than it has to be to enable factorization of matrix in as few serial steps as possible and where uniformly high efficiency expected from all processing elements. One of major purposes of this and related studies to maximize speedup and efficiency in system of concurrent-data-processing elements.

Salama, Moktar A.↗

Progress on a Taylor weak statement finite element algorithm for high-speed aerodynamic flows

A new finite element numerical Computational Fluid Dynamics (CFD) algorithm has matured to the point of efficiently solving two-dimensional high speed real-gas compressible flow problems in generalized coordinates on modern vector computer systems. The algorithm employs a Taylor Weak Statement classical Galerkin formulation, a variably implicit Newton iteration, and a tensor matrix product factorization of the linear algebra Jacobian under a generalized coordinate transformation. Allowing for a general two-dimensional conservation law system, the algorithm has been exercised on the Euler and laminar forms of the Navier-Stokes equations. Real-gas fluid properties are admitted, and numerical results verify solution accuracy, efficiency, and stability over a range of test problem parameters.

Baker, A. J.↗

An intrinsically n-dimensional generalized flux vector splitting implicit finite element Euler algorithm

A generalized flux-vector splitting implicit Galerkin finite-element algorithm for the Euler equations in curvilinear coordinates for ideal and reacting gases is derived. For an arbitrary equation of state, the curvilinear-coordinate flux vector is split in kinematic and kinetic components, and the associated jacobian matrix eigenvalues explicitly depend on the metric data. After directional semidiscretization, the terminal ordinary differential-equation system is solved via a nonlinearly stable implicit Runge-Kutta scheme in concert with an accurate tensor matrix product factorization. The results for selected two-dimensional supersonic and axisymmetric hypersonic flows validate the algorithm and verify its robustness for curvilinear-coordinate computations. The evolution towards a steady state is achieved for large Courant numbers without indication of numerical instabilities.

Iannelli, G. S.↗

Automatic Management of Parallel and Distributed System Resources

Viewgraphs on automatic management of parallel and distributed system resources are presented. Topics covered include: parallel applications; intelligent management of multiprocessing systems; performance evaluation of parallel architecture; dynamic concurrent programs; compiler-directed system approach; lattice gaseous cellular automata; and sparse matrix Cholesky factorization.

Yan, Jerry↗

Multivariable frequency domain identification via 2-norm minimization

The author develops a computational approach to multivariable frequency domain identification, based on 2-norm minimization. In particular, a Gauss-Newton (GN) iteration is developed to minimize the 2-norm of the error between frequency domain data and a matrix fraction transfer function estimate. To improve the global performance of the optimization algorithm, the GN iteration is initialized using the solution to a particular sequentially reweighted least squares problem, denoted as the SK iteration. The least squares problems which arise from both the SK and GN iterations are shown to involve sparse matrices with identical block structure. A sparse matrix QR factorization method is developed to exploit the special block structure, and to efficiently compute the least squares solution. A numerical example involving the identification of a multiple-input multiple-output (MIMO) plant having 286 unknown parameters is given to illustrate the effectiveness of the algorithm.

Bayard, David S.↗

Parallel 3D Multi-Stage Simulation of a Turbofan Engine

A 3D multistage simulation of each component of a modern GE Turbofan engine has been made. An axisymmetric view of this engine is presented in the document. This includes a fan, booster rig, high pressure compressor rig, high pressure turbine rig and a low pressure turbine rig. In the near future, all components will be run in a single calculation for a solution of 49 blade rows. The simulation exploits the use of parallel computations by using two levels of parallelism. Each blade row is run in parallel and each blade row grid is decomposed into several domains and run in parallel. 20 processors are used for the 4 blade row analysis. The average passage approach developed by John Adamczyk at NASA Lewis Research Center has been further developed and parallelized. This is APNASA Version A. It is a Navier-Stokes solver using a 4-stage explicit Runge-Kutta time marching scheme with variable time steps and residual smoothing for convergence acceleration. It has an implicit K-E turbulence model which uses an ADI solver to factor the matrix. Between 50 and 100 explicit time steps are solved before a blade row body force is calculated and exchanged with the other blade rows. This outer iteration has been coined a "flip." Efforts have been made to make the solver linearly scaleable with the number of blade rows. Enough flips are run (between 50 and 200) so the solution in the entire machine is not changing. The K-E equations are generally solved every other explicit time step. One of the key requirements in the development of the parallel code was to make the parallel solution exactly (bit for bit) match the serial solution. This has helped isolate many small parallel bugs and guarantee the parallelization was done correctly. The domain decomposition is done only in the axial direction since the number of points axially is much larger than the other two directions. This code uses MPI for message passing. The parallel speed up of the solver portion (no 1/0 or body force calculation) for a grid which has 227 points axially.

Turner, Mark G.↗

An Analytical Thermal Model for Autonomous Soaring Research

A viewgraph presentation describing an analytical thermal model used to enable research on autonomous soaring for a small UAV aircraft is given. The topics include: 1) Purpose; 2) Approach; 3) SURFRAD Data; 4) Convective Layer Thickness; 5) Surface Heat Budget; 6) Surface Virtual Potential Temperature Flux; 7) Convective Scaling Velocity; 8) Other Calculations; 9) Yearly trends; 10) Scale Factors; 11) Scale Factor Test Matrix; 12) Statistical Model; 13) Updraft Strength Calculation; 14) Updraft Diameter; 15) Updraft Shape; 16) Smoothed Updraft Shape; 17) Updraft Spacing; 18) Environment Sink; 19) Updraft Lifespan; 20) Autonomous Soaring Research; 21) Planned Flight Test; and 22) Mixing Ratio.

Allen, Michael↗

Noninteracting Control of Robotic Space Vehicles

This paper develops methods for noninteracting control of articulated, possibly flexible, multibody space vehicles based on the diagonalized equation of motion.

robotic space vehicles noninteracting control mass↗

Fast and accurate calculation of EXAFS Debye-Waller factors in U⁢O2 using the dynamical matrix method

Theoretical modeling of bonding dynamics in metal oxides is required for predicting their thermal conductivity, catalytic activity, and mechanical properties. A primary challenge is the scarcity of experimental methods for validating theoretical predictions of these atomic-scale dynamics. This work presents a workflow that uses experimental extended x-ray absorption fine structure (EXAFS) data collected at high temperatures to validate an interatomic force field for uranium dioxide (UO2), an important model material. The validated force field is then used to drive computationally intensive molecular dynamics (MD) simulations and as input for the much faster dynamical matrix Debye-Waller (DMDW) method. The predicted values of the Debye-Waller factors from the DMDW calculations are in good agreement with those obtained from the MD simulations, with residual pair-specific differences attributable to quantum zero-point motion at low temperatures and lattice anharmonicity at high temperatures. We further show that theoretical EXAFS spectra constructed directly from DMDW-derived Debye-Waller factors reproduce the experimental data (at relatively low temperatures) with accuracy comparable to full MD-EXAFS, providing an additional validation of the choice of the potential. This study establishes a validated, rapid computational pathway for modeling bond dynamics, naturally incorporating quantum nuclear\\\\r\\\\nstatistics absent in classical simulations, which are essential for the mechanistic understanding of complex oxide materials.

58 GEOSCIENCES↗

GSoFa: Scalable Sparse Symbolic LU Factorization on GPUs

Decomposing a matrix $\mathbf {A}$ into a lower matrix $\mathbf {L}$ and an upper matrix $\mathbf {U}$, which is also known as LU decomposition, is an essential operation in numerical linear algebra. For a sparse matrix, LU decomposition often introduces more nonzero entries in the $\mathbf {L}$ and $\mathbf {U}$ factors than in the original matrix. A symbolic factorization step is needed to identify the nonzero structures of $\mathbf {L}$ and $\mathbf {U}$ matrices. Attracted by the enormous potentials of the Graphics Processing Units (GPUs), an array of efforts have surged to deploy various LU factorization steps except for the symbolic factorization, to the best of our knowledge, on GPUs. This article introduces gSoFa, the first GPU-based symbolic factorization design with the following three optimizations to enable scalable LU symbolic factorization for nonsymmetric pattern sparse matrices on GPUs. First, here we introduce a novel fine-grained parallel symbolic factorization algorithm that is well suited for the Single Instruction Multiple Thread (SIMT) architecture of GPUs. Second, we tailor supernode detection into a SIMT friendly process and strive to balance the workload, minimize the communication and saturate the GPU computing resources during supernode detection. Third, we introduce a three-pronged optimization to reduce the excessive space consumption problem faced by multi-source concurrent symbolic factorization. Taken together, gSoFa achieves up to 31× speedup from 1 to 44 Summit nodes (6 to 264 GPUs) and outperforms the state-of-the-art CPU project, on average, by 5×. Notably, gSoFa also achieves up to 47 percent of the peak memory throughput of a V100 GPU in the Summit Supercomputer.

97 MATHEMATICS AND COMPUTING↗

Recursive flexible multibody system dynamics using spatial operators

This paper uses spatial operators to develop new spatially recursive dynamics algorithms for flexible multibody systems. The operator description of the dynamics is identical to that for rigid multibody systems. Assumed-mode models are used for the deformation of each individual body. The algorithms are based on two spatial operator factorizations of the system mass matrix. The first (Newton-Euler) factorization of the mass matrix leads to recursive algorithms for the inverse dynamics, mass matrix evaluation, and composite-body forward dynamics for the systems. The second (innovations) factorization of the mass matrix, leads to an operator expression for the mass matrix inverse and to a recursive articulated-body forward dynamics algorithm. The primary focus is on serial chains, but extensions to general topologies are also described. A comparison of computational costs shows that the articulated-body, forward dynamics algorithm is much more efficient than the composite-body algorithm for most flexible multibody systems.

Jain, A.↗

HYMPS: Numerical Techniques

The digital calculations that drive the classification portion of the Hybrid Pattern Recognition System (HYMPS) are considered. A revision of these calculations is suggested that involves three items: matrix inversion, det calculations and singularity, and covariance factorization. It is shown that it is more economical to first factor the covariance matrix, and by so doing, delete the matrix inversion routine. In addition, the necessary det calculations can be more easily realized by use of simple theoretical facts about the factorization.

Decell, H. P., Jr.↗

Application of fiber bridging models to fatigue crack growth in unidirectional titanium matrix composites

Several fiber bridging models were reviewed and applied to study the matrix fatigue crack growth behavior in center notched (0)(sub 8) SCS-6/Ti-15-3 and (0)(sub 4) SCS-6/Ti-6Al-4V laminates. Observations revealed that fatigue damage consisted primarily of matrix cracks and fiber matrix interfacial failure in the (0)(sub 8) SCS-6/Ti-15-3 laminates. Fiber-matrix interface failure included fracture of the brittle reaction zone and cracking between the two carbon rich fiber coatings. Intact fibers in the wake of the matrix cracks reduce the stress intensity factor range. Thus, an applied stress intensity factor range is inappropriate to characterize matrix crack growth behavior. Fiber bridging models were used to determine the matrix stress intensity factor range in titanium metal matrix composites. In these models, the fibers in the wake of the crack are idealized as a closure pressure. An unknown constant frictional shear stress is assumed to act along the debond or slip length of the bridging fibers. The frictional shear stress was used as a curve fitting parameter to available data (crack growth data, crack opening displacement data, and debond length data). Large variations in the frictional shear stress required to fit the experimental data indicate that the fiber bridging models in their present form lack predictive capabilities. However, these models provide an efficient and relatively simple engineering method for conducting parametric studies of the matrix growth behavior based on constituent properties.

Bakuckas, J. G., Jr.↗

Recursive dynamics for flexible multibody systems using spatial operators

Due to their structural flexibility, spacecraft and space manipulators are multibody systems with complex dynamics and possess a large number of degrees of freedom. Here the spatial operator algebra methodology is used to develop a new dynamics formulation and spatially recursive algorithms for such flexible multibody systems. A key feature of the formulation is that the operator description of the flexible system dynamics is identical in form to the corresponding operator description of the dynamics of rigid multibody systems. A significant advantage of this unifying approach is that it allows ideas and techniques for rigid multibody systems to be easily applied to flexible multibody systems. The algorithms use standard finite-element and assumed modes models for the individual body deformation. A Newton-Euler Operator Factorization of the mass matrix of the multibody system is first developed. It forms the basis for recursive algorithms such as for the inverse dynamics, the computation of the mass matrix, and the composite body forward dynamics for the system. Subsequently, an alternative Innovations Operator Factorization of the mass matrix, each of whose factors is invertible, is developed. It leads to an operator expression for the inverse of the mass matrix, and forms the basis for the recursive articulated body forward dynamics algorithm for the flexible multibody system. For simplicity, most of the development here focuses on serial chain multibody systems. However, extensions of the algorithms to general topology flexible multibody systems are described. While the computational cost of the algorithms depends on factors such as the topology and the amount of flexibility in the multibody system, in general, it appears that in contrast to the rigid multibody case, the articulated body forward dynamics algorithm is the more efficient algorithm for flexible multibody systems containing even a small number of flexible bodies. The variety of algorithms described here permits a user to choose the algorithm which is optimal for the multibody system at hand. The availability of a number of algorithms is even more important for real-time applications, where implementation on parallel processors or custom computing hardware is often necessary to maximize speed.

Jain, A.↗

A Parallel Non-Overlapping Domain-Decomposition Algorithm for Compressible Fluid Flow Problems on Triangulated Domains

This paper considers an algebraic preconditioning algorithm for hyperbolic-elliptic fluid flow problems. The algorithm is based on a parallel non-overlapping Schur complement domain-decomposition technique for triangulated domains. In the Schur complement technique, the triangulation is first partitioned into a number of non-overlapping subdomains and interfaces. This suggests a reordering of triangulation vertices which separates subdomain and interface solution unknowns. The reordering induces a natural 2 x 2 block partitioning of the discretization matrix. Exact LU factorization of this block system yields a Schur complement matrix which couples subdomains and the interface together. The remaining sections of this paper present a family of approximate techniques for both constructing and applying the Schur complement as a domain-decomposition preconditioner. The approximate Schur complement serves as an algebraic coarse space operator, thus avoiding the known difficulties associated with the direct formation of a coarse space discretization. In developing Schur complement approximations, particular attention has been given to improving sequential and parallel efficiency of implementations without significantly degrading the quality of the preconditioner. A computer code based on these developments has been tested on the IBM SP2 using MPI message passing protocol. A number of 2-D calculations are presented for both scalar advection-diffusion equations as well as the Euler equations governing compressible fluid flow to demonstrate performance of the preconditioning algorithm.

Barth, Timothy J.↗

Ionizing radiation induces heritable disruption of epithelial cell interactions

Ionizing radiation (IR) is a known human breast carcinogen. Although the mutagenic capacity of IR is widely acknowledged as the basis for its action as a carcinogen, we and others have shown that IR can also induce growth factors and extracellular matrix remodeling. As a consequence, we have proposed that an additional factor contributing to IR carcinogenesis is the potential disruption of critical constraints that are imposed by normal cell interactions. To test this hypothesis, we asked whether IR affected the ability of nonmalignant human mammary epithelial cells (HMEC) to undergo tissue-specific morphogenesis in culture by using confocal microscopy and imaging bioinformatics. We found that irradiated single HMEC gave rise to colonies exhibiting decreased localization of E-cadherin, beta-catenin, and connexin-43, proteins necessary for the establishment of polarity and communication. Severely compromised acinar organization was manifested by the majority of irradiated HMEC progeny as quantified by image analysis. Disrupted cell-cell communication, aberrant cell-extracellular matrix interactions, and loss of tissue-specific architecture observed in the daughters of irradiated HMEC are characteristic of neoplastic progression. These data point to a heritable, nonmutational mechanism whereby IR compromises cell polarity and multicellular organization.

Non-NASA Center↗