Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,747 records · Page 97

Using Agent Base Models to Optimize Large Scale Network for Large System Inventories

The aim of this paper is to use Agent Base Models (ABM) to optimize large scale network handling capabilities for large system inventories and to implement strategies for the purpose of reducing capital expenses. The models used in this paper either use computational algorithms or procedure implementations developed by Matlab to simulate agent based models in a principal programming language and mathematical theory using clusters, these clusters work as a high performance computational performance to run the program in parallel computational. In both cases, a model is defined as compilation of a set of structures and processes assumed to underlie the behavior of a network system.

Shameldin, Ramez Ahmed↗

Fiber Bragg Grating Sensor System for Monitoring Smart Composite Aerospace Structures

Lightweight, electromagnetic interference (EMI) immune, fiber-optic, sensor- based structural health monitoring (SHM) will play an increasing role in aerospace structures ranging from aircraft wings to jet engine vanes. Fiber Bragg Grating (FBG) sensors for SHM include advanced signal processing, system and damage identification, and location and quantification algorithms. Potentially, the solution could be developed into an autonomous onboard system to inspect and perform non-destructive evaluation and SHM. A novel method has been developed to massively multiplex FBG sensors, supported by a parallel processing interrogator, which enables high sampling rates combined with highly distributed sensing (up to 96 sensors per system). The interrogation system comprises several subsystems. A broadband optical source subsystem (BOSS) and routing and interface module (RIM) send light from the interrogation system to a composite embedded FBG sensor matrix, which returns measurand-dependent wavelengths back to the interrogation system for measurement with subpicometer resolution. In particular, the returned wavelengths are channeled by the RIM to a photonic signal processing subsystem based on powerful optical chips, then passed through an optoelectronic interface to an analog post-detection electronics subsystem, digital post-detection electronics subsystem, and finally via a data interface to a computer. A range of composite structures has been fabricated with FBGs embedded. Stress tensile, bending, and dynamic strain tests were performed. The experimental work proved that the FBG sensors have a good level of accuracy in measuring the static response of the tested composite coupons (down to submicrostrain levels), the capability to detect and monitor dynamic loads, and the ability to detect defects in composites by a variety of methods including monitoring the decay time under different dynamic loading conditions. In addition to quasi-static and dynamic load monitoring, the system can capture acoustic emission events that can be a prelude to structural failure, as well as piezoactuator-induced ultrasonic Lamb-waves-based techniques as a basis for damage detection.

Moslehi, Behzad↗

Compiling global name-space parallel loops for distributed execution

Distributed memory machines do not provide hardware support for a global address space. Thus programmers are forced to partition the data across the memories of the architecture and use explicit message passing to communicate data between processors. The compiler support required to allow programmers to express their algorithms using a global name-space is examined. A general method is presented for analysis of a high level source program and its translation into a set of independently executing tasks communicating via messages. If the compiler has enough information, this translation can be carried out at compile time. Otherwise, run-time code is generated to implement the required data movement. The analysis required in both situations is described and the performance of the generated code on the Intel iPSC/2 is presented.

Koelbel, Charles↗

PLUM: Parallel Load Balancing for Unstructured Adaptive Meshes

Dynamic mesh adaption on unstructured grids is a powerful tool for computing large-scale problems that require grid modifications to efficiently resolve solution features. Unfortunately, an efficient parallel implementation is difficult to achieve, primarily due to the load imbalance created by the dynamically-changing nonuniform grid. To address this problem, we have developed PLUM, an automatic portable framework for performing adaptive large-scale numerical computations in a message-passing environment. First, we present an efficient parallel implementation of a tetrahedral mesh adaption scheme. Extremely promising parallel performance is achieved for various refinement and coarsening strategies on a realistic-sized domain. Next we describe PLUM, a novel method for dynamically balancing the processor workloads in adaptive grid computations. This research includes interfacing the parallel mesh adaption procedure based on actual flow solutions to a data remapping module, and incorporating an efficient parallel mesh repartitioner. A significant runtime improvement is achieved by observing that data movement for a refinement step should be performed after the edge-marking phase but before the actual subdivision. We also present optimal and heuristic remapping cost metrics that can accurately predict the total overhead for data redistribution. Several experiments are performed to verify the effectiveness of PLUM on sequences of dynamically adapted unstructured grids. Portability is demonstrated by presenting results on the two vastly different architectures of the SP2 and the Origin2OOO. Additionally, we evaluate the performance of five state-of-the-art partitioning algorithms that can be used within PLUM. It is shown that for certain classes of unsteady adaption, globally repartitioning the computational mesh produces higher quality results than diffusive repartitioning schemes. We also demonstrate that a coarse starting mesh produces high quality load balancing, at a fraction of the cost required a fine initial mesh. Results indicate that our parallel load balancing strategy will remain viable on large numbers of processors.

Oliker, Leonid↗

A scoping study of far-SOL main-wall protection limiters for steady-state operation of compact pilot plant tokamaks

We present a novel method for handling steady-state heat fluxes incident on the main wall of pilot plant-scale magnetic fusion devices, based on the utilization of protection limiters in the far scrape-off layer (SOL). This method helps avoid large plasma-wall gaps, without excessively compromising blanket performance. We present an optimization algorithm for determining the appropriate size and scale of these protection limiters given (1) probability distributions of SOL plasma parameters and (2) assumed risk tolerance. As part of this optimization, we have developed an analytic description of parallel heat fluxes across limiter shadows, and an objective cost function (the ‘Far-SOL Marginal Cost’) to quantify the impact that different main-wall thermal management design choices have on reactor capital cost. Applying the model to a midscale fusion pilot plant concept shows that making use of far-SOL protection limiters can reduce capital costs on the order of $500 M, relative to naively increasing the plasma-wall gap. Our analysis demonstrates that the far-SOL power decay length is the highest-leverage plasma assumption for thermal loading of the first wall, and the primary cost driver for main wall thermal management. The relative cost efficiency of protection limiters increases as assumptions on the far-SOL heat flux become more pessimistic. The concepts described in this paper motivate the further development of far-SOL protection limiters as part of larger efforts to design economical core-edge-wall compatible solutions for a fusion pilot plant.

Design under uncertainty↗

Emu v1.1

Emu is a particle-in-cell code for solving the neutrino quantum kinetic equations in 1, 2, or 3 spatial dimensions with arbitrary angular resolution in order to simulate neutrino flavor transformation in neutron star merger and core collapse supernova environments. Emu represents the neutrino distribution function as a set of particles, each of which represent a collection of neutrinos and antineutrinos with unique position and momentum. Each particle carries two density matrices to define the flavor state of the neutrinos and antineutrinos it represents. Emu includes the vacuum, matter, and neutrino self-interaction potentials. Emu calculates the self-interaction potential using PIC deposition and interpolation algorithms that efficiently compute the local number density and flux at particle locations. Emu is implemented in C++ and is based on the AMReX library for high-performance, block-structured adaptive mesh refinement. Emu is parallelized with MPI + OpenMP for CPUs and MPI + CUDA for GPUs.

Willcox, Donald↗

Phase Field Dislocation Dynamics (PFDD) version 2.x

This disclosure is for version 2.x of a mesoscale model called Phase Field Dislocation Dynamics (PFDD). PFDD is used for investigating deformation in nanoscale (grain sizes of ~300 nm and less) materials, such as metals and alloys. This approach models the motion and interaction of individual defects, namely dislocations, in the material using scalar-valued phase field variables, also called order parameters. The system is evolved through energy minimization thus the model calculates the total energy density in terms of the phase field variables. The energy minimization is completed using the Ginzburg-Landau equation, and is implemented with explicit time integration. The total system energy can be comprised of several terms, including the strain energy (which describes dislocation-dislocation interactions), the energy due to an applied stress (dislocation interactions with the applied stress), and a core/lattice (perfect dislocations) or generalized stacking fault (partial dislocations) energy (described the dislocation core structure). The latter term in particular may vary based on the crystal structure being modeled and is typically informed using lower length scale (e.g., atomistic) approaches, although no such (atomistic) calculations are completed within the PFDD algorithm. This basic formulation was previously reviewed by Los Alamos National Laboratory and released under license number C17113. This previously reviewed version we will henceforth refer to as PFDD v1.0. PFDD v1.0 consisted of 2 codes (one parallel and one serial) plus input files, all written in the C language. This new disclosure is addressing the next versions of the PFDD, versions 2.x. There have been several enhancements of PFDD v1.0, which are described here and included in the attached code, which we will refer to as PFDD v2.0. There are also several new features described here that are either planned or already in process and are expected to be subsequent releases, i.e., v2.1, v2.2, ...v2.x.

Hunter, Abigail↗

Performant Optimization Strategies for Multifidelity Stochastic Power Grid Models

This talk goes into the algorithmic work done under the Forest project in order to solve expensive power grid models. We explore multiple fidelities of models that balance accuracy and computational expense. We use bundling strategies and progressive hedging in order to parallelize large stochastic programs.

Alfant, Rachael May [Sandia National Laboratories ↗

Developing new architectures for the Block 2 VLBI correlator system

The overall LSI (large-scale integrated circuits) architecture design and current status of the VBLI (very long baseline interferometry) block 2 correlator is addressed. The VBLI correlator algorithms demand a computing system that provides a throughput of hundreds of millions of instructions per second to perform cross-correlation detection for six baselines. The LSI technology lights the way for the computation of complex parallel process and is raising the upper bound of computerization.

Peterson, J. C.↗

Automated matching of pairs of SIR-B images for elevation mapping

During the SIR-B mission in October 1984, a significant number of overlapping synthetic aperture radar (SAR) images of various ground areas was collected. This has offered the first opportunity to perform stereo analyses on images from space that cover large ground areas to determine elevation information. This paper presents the preliminary results of an investigation to obtain elevation data from stereo pairs of SIR-B images. First, the accuracy with which elevation information can be derived from SIR-B image pairs is evaluated theoretically. It is shown that elevation accuracy is a function of the slant range resolution, the incidence angles with which the stereo pair is obtained, the accuracies in spacecraft state estimation, and determination of corresponding pixels in the stereo pair. Next, a hierarchical method is developed to match the corresponding pixels. This method involves iterative removal of local distortions and correlations of pairs of local neighborhoods in the two images. Since it is necessary to perform the matching at every pixel in the image, it is very computationally intensive. Therefore, it has been implemented on the Massively Parallel Processor (MPP) at the Goddard Space Flight Center (GSFC). The MPP's speed permits two iterations of this technique to operate on a pair of 512 x 512 images within 7 s. Results of applying this algorithm of SIR-B images of Mount Shasta, CA, are shown. The matching algorithm performs well in regions of the image with significant features. An approximate elevation image derived from the matching process corresponds to published topographic map data, except for certain obvious discontinuities.

Ramapriyan, H. K.↗

Performance of an electromagnetic bearing for the vibration control of a supercritical shaft

The flexural vibrations of a rotating shaft, running through one or more critical speeds, can be reduced to an acceptably low level by applying suitable control forces at an intermediate span position. If electromagnets are used to produce the control forces then it is possible to implement a wide variety of control strategies. A test rig is described which includes a microprocessor-based controller, in which such strategies can be realised in terms of software-based algorithms. The electromagnet configuration and the method of stabilising the electromagnet force-gap characteristic are discussed. The bounds on the performance of the system are defined. A simple control algorithm is outlined, where the control forces are proportional to the measured displacement and velocity at a single point on the shaft span; in this case the electromagnet behaves in a similar manner to that of a parallel combination of a linear spring and damper. Experimental and predicted performance of the system are compared, for this type of control, where various programmable rates of damping are applied.

Bradfield, C. D.↗

Modeling of outgassing and matrix decomposition in carbon-phenolic composites

Work done in the period Jan. - June 1994 is summarized. Two threads of research have been followed. First, the thermodynamics approach was used to model the chemical and mechanical responses of composites exposed to high temperatures. The thermodynamics approach lends itself easily to the usage of variational principles. This thermodynamic-variational approach has been applied to the transpiration cooling problem. The second thread is the development of a better algorithm to solve the governing equations resulting from the modeling. Explicit finite difference method is explored for solving the governing nonlinear, partial differential equations. The method allows detailed material models to be included and solution on massively parallel supercomputers. To demonstrate the feasibility of the explicit scheme in solving nonlinear partial differential equations, a transpiration cooling problem was solved. Some interesting transient behaviors were captured such as stress waves and small spatial oscillations of transient pressure distribution.

Mcmanus, Hugh L.↗

Optimizations of a Hardware Decoder for Deep-Space Optical Communications

The National Aeronautics and Space Administration has developed a capacity approaching modulation and coding scheme that comprises a serial concatenation of an inner accumulate pulse-position modulation (PPM) and an outer convolutional code [or serially concatenated PPM (SCPPM)] for deep-space optical communications. Decoding of this code uses the turbo principle. However, due to the nonbinary property of SCPPM, a straightforward application of classical turbo decoding is very inefficient. Here, we present various optimizations applicable in hardware implementation of the SCPPM decoder. More specifically, we feature a Super Gamma computation to efficiently handle parallel trellis edges, a pipeline-friendly 'maxstar top-2' circuit that reduces the max-only approximation penalty, a low-latency cyclic redundancy check circuit for window-based decoders, and a high-speed algorithmic polynomial interleaver that leads to memory savings. Using the featured optimizations, we implement a 6.72 megabits-per-second (Mbps) SCPPM decoder on a single field-programmable gate array (FPGA). Compared to the current data rate of 256 kilobits per second from Mars, the SCPPM coded scheme represents a throughput increase of more than twenty-six fold. Extension to a 50-Mbps decoder on a board with multiple FPGAs follows naturally. We show through hardware simulations that the SCPPM coded system can operate within 1 dB of the Shannon capacity at nominal operating conditions.

quadratic polynomial interleaver↗

Integration of a Decentralized Linear-Quadratic-Gaussian Control into GSFC's Universal 3-D Autonomous Formation Flying Algorithm

A decentralized control is investigated for applicability to the autonomous formation flying control algorithm developed by GSFC for the New Millenium Program Earth Observer-1 (EO-1) mission. This decentralized framework has the following characteristics: The approach is non-hierarchical, and coordination by a central supervisor is not required; Detected failures degrade the system performance gracefully; Each node in the decentralized network processes only its own measurement data, in parallel with the other nodes; Although the total computational burden over the entire network is greater than it would be for a single, centralized controller, fewer computations are required locally at each node; Requirements for data transmission between nodes are limited to only the dimension of the control vector, at the cost of maintaining a local additional data vector. The data vector compresses all past measurement history from all the nodes into a single vector of the dimension of the state; and The approach is optimal with respect to standard cost functions. The current approach is valid for linear time-invariant systems only. Similar to the GSFC formation flying algorithm, the extension to linear LQG time-varying systems requires that each node propagate its filter covariance forward (navigation) and controller Riccati matrix backward (guidance) at each time step. Extension of the GSFC algorithm to non-linear systems can also be accomplished via linearization about a reference trajectory in the standard fashion, or linearization about the current state estimate as with the extended Kalman filter. To investigate the feasibility of the decentralized integration with the GSFC algorithm, an existing centralized LQG design for a single spacecraft orbit control problem is adapted to the decentralized framework while using the GSFC algorithm's state transition matrices and framework. The existing GSFC design uses both reference trajectories of each spacecraft in formation and by appropriate choice of coordinates and simplified measurement modeling is formulated as a linear time-invariant system. Results for improvements to the GSFC algorithm and a multiple satellite formation will be addressed. The goal of this investigation is to progressively relax the assumptions that result in linear time-invariance, ultimately to the point of linearization of the non-linear dynamics about the current state estimate as in the extended Kalman filter. An assessment will then be made about the feasibility of the decentralized approach to the realistic formation flying application of the EO-1/Landsat 7 formation flying experiment.

Folta, David C.↗

Two criteria for the selection of assembly plans - Maximizing the flexibility of sequencing the assembly tasks and minimizing the assembly time through parallel execution of assembly tasks

The authors introduce two criteria for the evaluation and selection of assembly plans. The first criterion is to maximize the number of different sequences in which the assembly tasks can be executed. The second criterion is to minimize the total assembly time through simultaneous execution of assembly tasks. An algorithm that performs a heuristic search for the best assembly plan over the AND/OR graph representation of assembly plans is discussed. Admissible heuristics for each of the two criteria introduced are presented. Some implementation issues that affect the computational efficiency are addressed.

Homem De Mello, Luiz S.↗

NAS Applications and Advanced Algorithms

This paper examines the applications most commonly run on the supercomputers at the Numerical Aerospace Simulation (NAS) facility. It analyzes the extent to which such applications are fundamentally oriented to vector computers, and whether or not they can be efficiently implemented on hierarchical memory machines, such as systems with cache memories and highly parallel, distributed memory systems.

Bailey, David H.↗

A Non-perturbative Approach to Computing Seismic Normal Modes in Rotating Planets

In this work, a continuous Galerkin method based approach is presented to compute the seismic normal modes of rotating planets. Special care is taken to separate out the essential spectrum in the presence of a fluid outer core using a polynomial filtering eigensolver. The relevant elastic-gravitational system of equations, including the Coriolis force, is subjected to a mixed finite-element method, while self-gravitation is accounted for with the fast multipole method. Our discretization utilizes fully unstructured tetrahedral meshes for both solid and fluid regions. The relevant eigenvalue problem is solved by a combination of several highly parallel and computationally efficient methods. We validate our three-dimensional results in the non-rotating case using analytical results for constant elastic balls, as well as numerical results for an isotropic Earth model from standard “radial” algorithms. We also validate the computations in the rotating case, but only in the slowly-rotating regime where perturbation theory applies, because no other independent algorithms are available in the general case. The algorithm and code are used to compute the point spectra of eigenfrequencies in several Earth and Mars models studying the effects of heterogeneity on a large range of scales.

58 GEOSCIENCES↗