Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,729 records · Page 96

A compact high-speed parallel multiplication scheme

This paper discusses a compact, fast, parallel multiplication scheme of the generation-reduction type using generalized Dadda-type pseudoadders for reduction and m x m multipliers for generation. The implications of present and future LSI are considered, a partitioning algorithm is presented, and the results obtained for a 24 x 24-bit implementation are discussed.

Stenzel, W. J.↗

New cellular automaton model for magnetohydrodynamics

A new type of two-dimensional cellular automation method is introduced for computation of magnetohydrodynamic fluid systems. Particle population is described by a 36-component tensor referred to a hexagonal lattice. By appropriate choice of the coefficients that control the modified streaming algorithm and the definition of the macroscopic fields, it is possible to compute both Lorentz-force and magnetic-induction effects. The method is local in the microscopic space and therefore suited to massively parallel computations.

Chen, Hudong↗

Fully-coupled analysis of jet mixing problems. Three-dimensional PNS model, SCIP3D

Numerical procedures formulated for the analysis of 3D jet mixing problems, as incorporated in the computer model, SCIP3D, are described. The overall methodology closely parallels that developed in the earlier 2D axisymmetric jet mixing model, SCIPVIS. SCIP3D integrates the 3D parabolized Navier-Stokes (PNS) jet mixing equations, cast in mapped cartesian or cylindrical coordinates, employing the explicit MacCormack Algorithm. A pressure split variant of this algorithm is employed in subsonic regions with a sublayer approximation utilized for treating the streamwise pressure component. SCIP3D contains both the ks and kW turbulence models, and employs a two component mixture approach to treat jet exhausts of arbitrary composition. Specialized grid procedures are used to adjust the grid growth in accordance with the growth of the jet, including a hybrid cartesian/cylindrical grid procedure for rectangular jets which moves the hybrid coordinate origin towards the flow origin as the jet transitions from a rectangular to circular shape. Numerous calculations are presented for rectangular mixing problems, as well as for a variety of basic unit problems exhibiting overall capabilities of SCIP3D.

Wolf, D. E.↗

Image reconstruction from multiple 1-D scans using filtered localized projection

The spatial resolution that can be attained using scanning linear arrays (consisting of discrete IR solid-state detectors) for image acquisition is considered, and a filtered local projection (FLP) method is described which efficiently combines all available scan information into one rectangular grid without the need for explicit interpolation. Mathematically, the FLP algorithm consists of a localized summation followed by an inverse-filter operation, and it has application to nonlinear restoring techniques. The present method is applied, using a linear array, to simulated data for staggered parallel scans and to multiple scan directions. Noise effects and limitations of the technique are also considered.

Frieden, B. Roy↗

Memory management in traceback Viterbi decoders

The new Viterbi decoder for long constraint length codes, under development for the Deep Space Network, stores path information according to an algorithm called traceback. The details of a particular implementation of this algorithm, based on three memory buffers, are described. The penalties in increased storage requirement and longer decoding delay are offset by the reduced amount of data that needs to be exchanged between processors, in a parallel architecture decoder.

Collins, O.↗

Cumulative reports and publications through December 31, 1989

A complete list of reports from the Institute for Computer Applications in Science and Engineering (ICASE) is presented. The major categories of the current ICASE research program are: numerical methods, with particular emphasis on the development and analysis of basic numerical algorithms; control and parameter identification problems, with emphasis on effectual numerical methods; computational problems in engineering and the physical sciences, particularly fluid dynamics, acoustics, structural analysis, and chemistry; computer systems and software, especially vector and parallel computers, microcomputers, and data management. Since ICASE reports are intended to be preprints of articles that will appear in journals or conference proceedings, the published reference is included when it is available.

Source record↗

GPU Implementation of the OVERFLOW CFD Code

The high-performance computing (HPC) landscape is quickly changing to systems where most of the performance comes from specialized chips, specifically graphics processing units (GPUs). Such GPU systems are throughput machines, where efficient use of the GPU often requires code refactoring to expose a few orders of magnitude more fine grain parallelism than was previously used on the CPU. Recent modifications to OVERFLOW, an overset, structured grid, computational fluid dynamics flow solver, written in Fortran will be presented. These modifications include both code modernization efforts and algorithmic changes to enable OVERFLOW to efficiently utilize GPUs. Many of these algorithmic changes would likely also be applicable for other structured grid, stencil-based codes wanting to utilize GPUs. The capabilities that have been ported to run on the GPUs are presented, along with the performance gains of the GPU version relative the CPU version of OVERFLOW.

GPU Programming↗

GPU Implementation of the OVERFLOW CFD Code

The high-performance computing (HPC) landscape is quickly changing to systems where most of the performance comes from specialized chips, specifically graphics processing units (GPUs). Such GPU systems are throughput machines, where efficient use of the GPU often requires code refactoring to expose a few orders of magnitude more fine grain parallelism than was previously used on the CPU. Recent modifications to OVERFLOW, an overset, structured grid, computational fluid dynamics flow solver, written in Fortran will be presented. These modifications include both code modernization efforts and algorithmic changes to enable OVERFLOW to efficiently utilize GPUs. Many of these algorithmic changes would likely also be applicable for other structured grid, stencil-based codes wanting to utilize GPUs. The capabilities that have been ported to run on the GPUs are presented, along with the performance gains of the GPU version relative the CPU version of OVERFLOW.

GPU Programming↗

Second Generation Readout For Large Format Photon Counting Microwave Kinetic Inductance Detectors

We present the development of a second generation digital readout system for photon counting microwave kinetic inductance detector (MKID) arrays operating in the optical and near-infrared wavelength bands. Our system retains much of the core signal processing architecture from the first generation system but with a significantly higher bandwidth, enabling the readout of kilopixel MKID arrays. Each set of readout boards is capable of reading out 1024 MKID pixels multiplexed over 2 GHz of bandwidth; two such units can be placed in parallel to read out a full 2048 pixel microwave feedline over a 4 GHz–8 GHz band. As in the first generation readout, our system is capable of identifying, analyzing, and recording photon detection events in real time with a time resolution of order a few microseconds. Here, we describe the hardware and firmware, and present an analysis of the noise properties of the system. We also present a novel algorithm for efficiently suppressing IQ mixer sidebands to below −30 dBc.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Empirical validation and comparison of methodologies to simulate micro and macro-encapsulated PCMs in the building envelope

Thermal Energy Storage (TES) has the potential to shift peak electricity demand. Passive TES is usually implemented in building envelope as micro and macro encapsulated phase change materials (PCM) to shift electric energy demand and therefore requires careful heat transfer analysis. Whole building energy modelling with simplified heat transfer analysis has become extremely important for designers, architects, engineers, and researchers to predict energy performance of buildings. It is important to validate PCM modelling algorithms used in building energy programs to quantify their error and prove their capacity to model different PCM encapsulation types. This study uses data from a microencapsulated PCM and two macroencapsulated PCMs (Bio based PCM and hydrate salts) tested in full-scale using the Advanced Multiscale Building Energy Research (AMBER) Lab located at the Colorado School of Mines and is used to validate a numerical algorithm written in MATLAB language. To approximate the heat transfer through a wall assembly with macroencapsulated PCM pouches, several modelling techniques that can reduce 3D heat transfer characteristics to 1D are explored in this research. A parallel path heat transfer modelling approach is found to give the closest agreement with the experimental data for the pouched PCMs in building envelope applications.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

A New Class of AMG Interpolation Methods Based on Matrix-Matrix Multiplications

A new class of distance-two interpolation methods for algebraic multigrid (AMG) that can be formulated in terms of sparse matrix-matrix multiplications is presented and analyzed. Compared with similar distance-two prolongation operators, the proposed algorithms exhibit improved efficiency and portability to various computing platforms, since they allow one to easily exploit existing high-performance sparse matrix kernels. The new interpolation methods have been implemented in hypre, a widely used parallel multigrid solver library. With the proposed interpolations, the overall time of hypre's BoomerAMG setup can be considerably reduced, while sustaining equivalent, sometimes improved, convergence rates. Numerical results for a variety of test problems on parallel machines are presented that support the superiority of the proposed interpolation operators over the existing ones in hypre.

97 MATHEMATICS AND COMPUTING↗

PyJMAK: An Open-Source Python Toolkit for Modeling Solid-State Metallurgical Phase Transformations

Accurate prediction of metallurgical phase transformations is an essential basis for autonomous optimization and rapid part qualification. Several methods can be used to estimate the evolution of phase fractions such as JMAK kinetics-based models, phase-field models, thermodynamic models, and data-driven machine learning models. Thermodynamic and phase-field-based methodologies solve multiphysics equations requiring numerous calibration parameters and significant computational resources. As a result, the computation domain is limited to a point or on order of micron-meters. The data-driven models rely on large datasets from experiments and simulations. While the JMAK model only provides information about phase fraction evolution, it can predict this evolution in near real-time using thermal history and thermodynamic data without restriction on the domain. JMAK models have been popularly used by researchers to model phase transformations occuring during additive manufacturing or over arbitrary temperature profiles. Commercial proprietary software such as Abaqus and Ansys or closed-source in-house implementations offer the ability to model JMAK based kinetics to predict phase transformation. However, these software packages are not open-source or freely available for use and development in conjunction with manufacturing machines, sensors, and machine learning algorithms. In addition, the use of the model is restricted by a license token. In contrast, given temperature profiles at multiple points in the domain, this Python-based PyJMAK model can compute phase evolution in parallel due to its stand-alone modular, voxel-based structure, and it can be executed on high-performance computing resources without any license restrictions.

Prabhune, Bhagya [Oak Ridge National Laboratory (O↗

On the inversion of block tridiagonals without storage constraints

A strategy was developed to permit trade-offs between the number of floating point operations required and the storage requirements for the solution of certain difference problems, such as block tridiagonal systems of equations. This is done by recomputing some intermediate results instead of storing them. Reducing the storage to the square root of the current requirement roughly doubles the number of computations. Reducing the storage more than this tends to make the number of computations prohibitively large. In theory, if m is the order of each sub-matrix in the block tridiagonal matrix, one can solve any linear system with only 5m(2) + 1 temporary storage cells. In many cases m is a constant and quite small. For example, in solving a factored form of the three-dimensional Navier-Stokes equations, the size m of the block tridiagonals is 5. This method lends itself to efficient use on computers with parallel processing or vector processing architectures. On these computers the larger number of floating point operations is more than offset by the decrease in I/O and the increased percentage of vector operations made possible by this algorithm.

Merriam, M. L.↗

On the factorization of block-tridiagonals without storage constraints

In many programs solving difference equations, problem size is restricted by the number of available memory cells. A strategy has been developed to permit trade-offs between the number of floating point operations required and storage requirements for the solution of certain problems such as block tridiagonal systems of equations. This is done by recomputing some intermediate results instead of storing them. Reducing the storage to the square root of the current requirement will roughly double the number of computations. In theory, if m is the order of each sub-matrix in the block tridiagonal matrix, one can solve any linear system with only 5 sq m + 1 temporary storage cells. This method lends itself to efficient use on computers with parallel processing or vector processing architectures. On these computers the larger number of floating point operations is more than offset by the decrease in I/O and the increased percentage of vector operations made possible by this algorithm.

Merriam, M. L.↗

Electromagnetic scattering analysis on a hypercube parallel architecture

The applicability of a parallel architecture to the solution of large electromagnetic scattering problems is demonstrated. Two techniques, finite difference and the method of moments, are used to provide insight into the comparative speedups which can be attained for several algorithms. The flexibility of the hypercube architecture for different analysis algorithms is illustrated.

Patterson, Jean E.↗

Semiannual report, 1 April - 30 September 1991

The major categories of the current Institute for Computer Applications in Science and Engineering (ICASE) research program are: (1) numerical methods, with particular emphasis on the development and analysis of basic numerical algorithms; (2) control and parameter identification problems, with emphasis on effective numerical methods; (3) computational problems in engineering and the physical sciences, particularly fluid dynamics, acoustics, and structural analysis; and (4) computer systems and software for parallel computers. Research in these areas is discussed.

Source record↗

Toward intelligent flight control

Flight control systems can benefit by being designed to emulate functions of natural intelligence. Intelligent control functions fall in three categories: declarative, procedural, and reflexive. Declarative actions involve decision-making, providing models for system monitoring, goal planning, and system/scenario identification. Procedural actions concern skilled behavior and have parallels in guidance, navigation, and adaptation. Reflexive actions are more-or-less spontaneous and are similar to inner-loop control and estimation. Intelligent flight control systems will contain a hierarchy of expert systems, procedural algorithms, and computational neural networks, each expanding on prior functions to improve mission capability to increase the reliability and safety of flight and to ease pilot workload.

Stengel, Robert F.↗

Full-Potential Modeling of Blade-Vortex Interactions

A study of the full-potential modeling of a blade-vortex interaction was made. A primary goal of this study was to investigate the effectiveness of the various methods of modeling the vortex. The model problem restricts the interaction to that of an infinite wing with an infinite line vortex moving parallel to its leading edge. This problem provides a convenient testing ground for the various methods of modeling the vortex while retaining the essential physics of the full three-dimensional interaction. A full-potential algorithm specifically tailored to solve the blade-vortex interaction (BVI) was developed to solve this problem. The basic algorithm was modified to include the effect of a vortex passing near the airfoil. Four different methods of modeling the vortex were used: (1) the angle-of-attack method, (2) the lifting-surface method, (3) the branch-cut method, and (4) the split-potential method. A side-by-side comparison of the four models was conducted. These comparisons included comparing generated velocity fields, a subcritical interaction, and a critical interaction. The subcritical and critical interactions are compared with experimentally generated results. The split-potential model was used to make a survey of some of the more critical parameters which affect the BVI.

Jones, Henry E.↗