Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,369 records · Page 76

Path Planning: Differential Dynamic Programming and Model Predictive Path Integral Control on VTOL Aircraft

This paper explores two optimal control approaches, widely used in robotics, to establish their viability as real-time trajectory planners for vehicle configurations envisioned for the emerging aviation sector of Urban Air Mobility (UAM). Differential Dynamic Programming (DDP) enables planning over highly nonlinear dynamics using second-order approximations along a nominal trajectory, and displays quadratic convergence to a local solution. Model Predictive Path Integral (MPPI) is a stochastic sampling-based algorithm that can optimize for general cost criteria, including potentially highly nonlinear formulations, and supports parallel computation through the use of modern GPU hardware. In this work, DDP and MPPI were implemented using model predictive control (MPC), and the results indicate they are able to successfully transition the aircraft over different flight envelopes and generate trajectories unique to UAM vehicles.

Differential Dynamic Programming↗

Development and Airborne Demonstration of the Concurrent Artificially-Intelligent Spectrometry and Adaptive Lidar System: Advancing Lidar Capabilities for the STV Observing System

We report on the design, build and planned airborne demonstration of a spaceflight-prototype Concurrent Artificially-intelligent Spectrometry and Adaptive Lidar System (CASALS). The CASALS lidar is an Adaptive Wavelength Scanning Lidar (AWSL) operating in push broom mode. The demonstration has three major goals: advance the Technical Readiness Level of the AWSL hardware, validate its measurement performance and mature algorithms and methods needed for the Surface Topography and Vegetation (STV) observing system. AWSL acquires parallel tracks of surface heights by rapidly steering a laser beam across a swath. A 1040nm-centered laser is tuned across 30nm and carved into 2-ns pulses, the pulse energy is fiber amplified and the pulses are dispersed cross-track using a non-mechanical wavelength-to-angle grating. For the spaceflight system the beam will be pointable to 1200 10m footprints across a 7km swath. For the airborne demonstration there will be 256 0.7m footprints across a 110m swath. In both cases the footprints overlap across- and along-track for uniform target illumination. For the airborne demonstration a steering mirror will increase the accessible swath width to 4km. At the receiver, solar radiation is filtered with a narrow-slit grating-spectrometer and the footprints are imaged onto a linear-mode, photon-sensitive HgCdTe APD-array. The received pulses are time-division-multiplexed to a few high-speed analog-to-digital converters to record waveforms. Spaceflight and airborne CASALS are designed to nominally detect 20 photons per pulse and, by averaging 27 overlapping footprints, achieve 2cm flat target range precision and high-quality vegetation structure waveforms The AWSL will be flown in the summer of 2024, along with a Headwall VNIR-SWIR hyperspectral sensor imaging a 4km wide swath, at NEON eddy covariance flux towers in the U.S. mid-Atlantic where high resolution hyperspectral and lidar data, acquired annually, are available for validation.

Guangning Yang↗

A computationally efficient statistically downscaled 100 m resolution Greenland product from the regional climate model MAR

The Greenland Ice Sheet (GrIS) has been contributing directly to sea level rise, and this contribution is projected to accelerate over the next decades. A crucial tool for studying the evolution of surface mass loss (e.g., surface mass balance, SMB) consists of regional climate models (RCMs), which can provide current estimates and future projections of sea level rise associated with such losses. However, one of the main limitations of RCMs is the relatively coarse horizontal spatial resolution at which outputs are currently generated. Here, we report results concerning the statistical downscaling of the SMB modeled by the Modèle Atmosphérique Régional (MAR) RCM from the original spatial resolution of 6 km to 100 m building on the relationship between elevation and mass losses in Greenland. To this goal, we developed a geospatial framework that allows the parallelization of the downscaling process, a crucial aspect to increase the computational efficiency of the algorithm. Using the results obtained in the case of the SMB, surface and air temperature are assessed through the comparison of the modeled outputs with in situ and satellite measurement. The downscaled products show a considerable improvement in the case of the downscaled product with respect to the original coarse output, with the coefficient of determination (R 2 ) increasing from 0.868 for the original MAR output to 0.935 for the SMB downscaled product. Moreover, the value of the slope and intercept of the linear regression fitting modeled and measured SMB values shifts from 0.865 for the original MAR to 1.015 for the downscaled product in the case of the slope and from the value −235 mm w.e. yr -1 (original) to −57 mm w.e. yr -1 (downscaled) in the case of the intercept, considerably improving upon results previously published in the literature.

Greenland Ice Sheet↗

Scheduling message processing for reducing rollback propagation

Traditional checkpointing and rollback recovery techniques for parallel systems have typically assumed the communication pattern is specified by program behavior. In this paper we exploit the property that the communication pattern can often be changed at run-time without affecting program correctness. A scheduling algorithm for message processing and its implementation for reducing rollback propagation are described. The algorithm incorporates a user-transparent prioritized scheme based upon the run-time communication and checkpointing history. Communication trace-driven simulation for several parallel programs written in the Chare Kernel language demonstrates that the probability of rollback propagation can be reduced at the cost of slight additional performance degradation.

Wang, Yi-Min↗

Enhancements to Program LAURA for computation of three-dimensional hypersonic flow

Changes to Program Laura (Langley Aerothermodynamic Upwind Relaxation Algorithm) are presented which enhance both stability and accuracy of the algorithm. A discussion of iteration/sweeping strategies and their relation to computer architectures is included to best exploit the capabilities of serial, vector, and parallel processor machines. Test cases for Mach 10 perfect gas flow and Mach 32 real gas flow in chemical nonequilibrium over a blunt, raked elliptic cone using the thin-layer Navier-Stokes equations are presented in order to demonstrate the current improved capabilities. Algorithm changes include the use of volume averaging, application of a symmetric total variation diminishing (TVD) scheme, and stronger interaction between the grid/shock alignment routine and the relaxation algorithm. Good comparisons with heat transfer and pitching moment data at three different angles of attack for the Mach 10 tests serve to further validate the present algorithm. Parameters are defined which control the coupling of the specie continuity equations with the solution of the mixture conservation equations. A discussion of the consequences involved in the choice of strong versus weak coupling is presented, and a sample nonequilibrium calculation on a fine grid over a full scale model of the Aeroassist Flight Experiment (AFE) demonstrates current capabilities.

Gnoffo, Peter A.↗

Algorithms for Efficient Reproducible Floating Point Summation

We define “reproducibility” as getting bitwise identical results from multiple runs of the same program, perhaps with different hardware resources or other changes that should not affect the answer. Many users depend on reproducibility for debugging or correctness. However, dynamic scheduling of parallel computing resources, combined with nonassociative floating point addition, makes reproducibility challenging even for summation, or operations like the BLAS. We describe a “reproducible accumulator” data structure (the “binned number”) and associated algorithms to reproducibly sum binary floating point numbers, independent of summation order. We use a subset of the IEEE Floating Point Standard 754-2008 and bitwise operations on the standard representations in memory. Our approach requires only one read-only pass over the data, and one reduction in parallel, using a 6-word reproducible accumulator (more words can be used for higher accuracy), enabling standard tiling optimization techniques. Summing n words with a 6-word reproducible accumulator requires approximately 9 n floating point operations (arithmetic, comparison, and absolute value) and approximately 3 n bitwise operations. The final error bound with a 6-word reproducible accumulator and our default settings can be up to 2 29 times smaller than the error bound for conventional (recursive) summation on ill-conditioned double-precision inputs.

Computer Science↗

Maximizing TDRS Command Load Lifetime

The GNC software onboard ISS utilizes TORS command loads, and a simplistic model of TORS orbital motion to generate onboard TORS state vectors. Each TORS command load contains five "invariant" orbital elements which serve as inputs to the onboard propagation algorithm. These elements include semi-major axis, inclination, time of last ascending node crossing, right ascension of ascending node, and mean motion. Running parallel to the onboard software is the TORS Command Builder Tool application, located in the JSC Mission Control Center. The TORS Command Builder Tool is responsible for building the TORS command loads using a ground TORS state vector, mirroring the onboard propagation algorithm, and assessing the fidelity of current TORS command loads onboard ISS. The tool works by extracting a ground state vector at a given time from a current TORS ephemeris, and then calculating the corresponding "onboard" TORS state vector at the same time using the current onboard TORS command load. The tool then performs a comparison between these two vectors and displays the relative differences in the command builder tool GUI. If the RSS position difference between these two vectors exceeds the tolerable lim its, a new command load is built using the ground state vector and uplinked to ISS. A command load's lifetime is therefore defined as the time from when a command load is built to the time the RSS position difference exceeds the tolerable limit. From the outset of TORS command load operations (STS-98), command load lifetime was limited to approximately one week due to the simplicity of both the onboard propagation algorithm, and the algorithm used by the command builder tool to generate the invariant orbital elements. It was soon desired to extend command load lifetime in order to minimize potential risk due to frequent ISS commanding. Initial studies indicated that command load lifetime was most sensitive to changes in mean motion. Finding a suitable value for mean motion was therefore the key to achieving this goal. This goal was eventually realized through development of an Excel spreadsheet tool called EMMIE (Excel Mean Motion Interactive Estimation). EMMIE utilizes ground ephemeris nodal data to perform a least-squares fit to inferred mean anomaly as a function of time, thus generating an initial estimate for mean motion. This mean motion in turn drives a plot of estimated downtrack position difference versus time. The user can then manually iterate the mean motion, and determine an optimal value that will maximize command load lifetime. Once this optimal value is determined, the mean motion initially calculated by the command builder tool is overwritten with the new optimal value, and the command load is built for uplink to ISS. EMMIE also provides the capability for command load lifetime to be tracked through multiple TORS ephemeris updates. Using EMMIE, TORS command load lifetimes of approximately 30 days have been achieved.

Brown, Aaron J.↗

A Length Adaptive Algorithm-Hardware Co-design of Transformer on FPGA Through Sparse Attention and Dynamic Pipelining

Transformers are considered one of the most important deep learning models since 2018, in part because it establishes state-of-the-art (SOTA) records and could potentially replace existing Deep Neural Networks (DNNs). Despite the remarkable triumphs, the prolonged turnaround time of Transformer models is a widely recognized roadblock. The variety of sequence lengths imposes additional computing overhead where inputs need to be zero-padded to the maximum sentence length in the batch to accommodate the parallel computing platforms. This paper targets the field-programmable gate array (FPGA) and proposes a coherent sequence length adaptive algorithm–hardware co-design for Transformer acceleration. Particularly, we develop a hardware-friendly sparse attention operator and a length-aware hardware resource scheduling algorithm. The proposed sparse attention operator brings the complexity of attention-based models down to linear complexity and alleviates the off-chip memory traffic. The proposed length-aware resource hardware scheduling algorithm dynamically allocates the hardware resources to fill up the pipeline slots and eliminates bubbles for NLP tasks. Experiments show that our design has very small accuracy loss and has 80.2 × and 2.6 × speedup compared to CPU and GPU implementation, and 4 × higher energy efficiency than state-of-the-art GPU accelerator optimized via CUBLAS GEMM.

Peng, Hongwu↗

Xyce Parallel Electronic Simulator Users' Guide (V.7.1)

This manual describes the use of the Xyce Parallel Electronic Simulator. Xyce has been designed as a SPICE-compatible, high-performance analog circuit simulator, and has been written to support the simulation needs of the Sandia National Laboratories electrical designers. This development has focused on improving capability over the current state-of-the-art in the following areas: 1) Capability to solve extremely large circuit problems by supporting large-scale parallel computing platforms (up to thousands of processors). This includes support for most popular parallel and serial computers. 2) A differential-algebraic-equation (DAE) formulation, which better isolates the device model package from solver algorithms. This allows one to develop new types of analysis without requiring the implementation of analysis-specific device models. 3) Device models that are specifically tailored to meet Sandia's needs, including some radiation-aware devices (for Sandia users only). 4) Object-oriented code design and implementation using modern coding practices. Xyce is a parallel code in the most general sense of the phrase a message passing parallel implementation which allows it to run efficiently a wide range of computing platforms. These include serial, shared-memory and distributed-memory parallel platforms. Attention has been paid to the specific nature of circuit-simulation problems to ensure that optimal parallel efficiency is achieved as the number of processors grows.

42 ENGINEERING↗

Xyce™ Parallel Electronic Simulator Users' Guide, Version 7.5.

This manual describes the use of the Xyce Parallel Electronic Simulator. Xyce has been designed as a SPICE-compatible, high-performance analog circuit simulator, and has been written to support the simulation needs of the Sandia National Laboratories electrical designers. This development has focused on improving capability over the current state-of-the-art in the following areas: (1) Capability to solve extremely large circuit problems by supporting large-scale parallel computing platforms (up to thousands of processors). This includes support for most popular parallel and serial computers. (2) A differential-algebraic-equation (DAE) formulation, which better isolates the device model package from solver algorithms. This allows one to develop new types of analysis without requiring the implementation of analysis-specific device models. (3) Device models that are specifically tailored to meet Sandia's needs, including some radiation-aware devices (for Sandia users only). (4) Object-oriented code design and implementation using modern coding practices. Xyce is a parallel code in the most general sense of the phrase — a message passing parallel implementation — which allows it to run efficiently a wide range of computing platforms. These include serial, shared-memory and distributed-memory parallel platforms. Attention has been paid to the specific nature of circuit-simulation problems to ensure that optimal parallel efficiency is achieved as the number of processors grows.

97 MATHEMATICS AND COMPUTING↗

Xyce™ Parallel Electronic Simulator Users' Guide (V.7.6)

This manual describes the use of the Xyce™ Parallel Electronic Simulator. Xyce™ has been designed as a SPICE-compatible, high-performance analog circuit simulator, and has been written to support the simulation needs of the Sandia National Laboratories electrical designers. This development has focused on improving capability over the current state-of-the-art in the following areas: (1) Capability to solve extremely large circuit problems by supporting large-scale parallel computing platforms (up to thousands of processors). This includes support for most popular parallel and serial computers. (2) A differential-algebraic-equation (DAE) formulation, which better isolates the device model package from solver algorithms. This allows one to develop new types of analysis without requiring the implementation of analysis-specific device models. (3) Device models that are specifically tailored to meet Sandia's needs, including some radiation-aware devices (for Sandia users only). (4) Object-oriented code design and implementation using modern coding practices. Xyce™ is a parallel code in the most general sense of the phrase—a message passing parallel implementation—which allows it to run efficiently a wide range of computing platforms. These include serial, shared-memory and distributed-memory parallel platforms. Attention has been paid to the specific nature of circuit-simulation problems to ensure that optimal parallel eficiency is achieved as the number of processors grows.

97 MATHEMATICS AND COMPUTING↗

Xyce™ Parallel Electronic Simulator Users’ Guide, Version 7.8

This manual describes the use of the Xyce Parallel Electronic Simulator. Xyce has been designed as a SPICE-compatible, high-performance analog circuit simulator, and has been written to support the simulation needs of the Sandia National Laboratories electrical designers. This development has focused on improving capability over the current state-of-the-art in the following areas: (1) Capability to solve extremely large circuit problems by supporting large-scale parallel computing platforms (up to thousands of processors). This includes support for most popular parallel and serial computers. (2) A differential-algebraic-equation (DAE) formulation, which better isolates the device model package from solver algorithms. This allows one to develop new types of analysis without requiring the implementation of analysis-specific device models. (3) Device models that are specifically tailored to meet Sandia’s needs, including some radiation-aware devices (for Sandia users only). (4) Object-oriented code design and implementation using modern coding practices. Xyce is a parallel code in the most general sense of the phrase — a message passing parallel implementation — which allows it to run efficiently a wide range of computing platforms. These include serial, shared-memory and distributed-memory parallel platforms. Attention has been paid to the specific nature of circuit-simulation problems to ensure that optimal parallel efficiency is achieved as the number of processors grows.

42 ENGINEERING↗

Real time identification of large space structures

Identification of frequencies, damping ratios, and mode shapes of large space structures (LSSs) are examined in real time. Real time processing allows for quick updates of model processing after a reconfiguration of structural failure. Recursive lattice least squares (RLLS) was selected as the baseline algorithm for the identification. Simulation results on a one dimensional LSS demonstrated that it provides good estimates, was not ill-conditioned in the presence of under-excited modes, allowed activity by a supervisory control system which prevented damage to the LSS or excessive drift, and was capable of real-time processing for typical LSS models. A suboptimal version of RLLS, which is equivalent to simulated parallel processing, was derived. A NASTRAN model of the dual keel U.S. space station was used to demonstrate the input/identification algorithm package in a more realistic simulation. Because the first eight flexible modes were very close together, the identification was much more difficult than in the simple examples. Even so, the model was accurately identified in real time.

Voss, Janice E.↗

Modal identification using single-mode projection filters and comparison with ERA and MLE results

The Single-Mode Projection Filter (SPF) is a newly developed algorithm for eigensystem parameter identification from both analytical results and test data. The SPF is formulated with a single mode only and practical for parallel processing implementation. Explicit formulations of SPF are derived for the multi-input multi-output (MIMO) system by using the orthogonal matrices of the controllability and observability matrices in the general sense. The modal parameters of SPF are initially obtained from an analytical model in modal space. The experimental data are then processed through SPF to update its modal parameters and to minimize a cost function defined by the norm of an error matrix. The updated modal parameters represent the characteristics of the test data. A two-dimensional global minimum optimization algorithm is developed and applied for the filter update by using the interval analysis method. The SPF is developed based on a single-mode subsystem and identifies only one modal frequency and one modal damping within a specified region. For an n-modes structure, n SPF can be implemented for parallel processing to reduce the computational burden. The SPF is applied to analyze the simulated data for the MAST beam structure. The estimated modal parameters are comparable to those from the Eigensystem Realization Algorithm (ERA) and repeated modal frequencies are identified. The modal analysis of the Spacecraft Control Laboratory Experiment (SCOLE) data is also performed by using the ERA and the Maximum Likelihood Estimate (MLE). The result shows that the first five modal frequencies are very close from ERA and MLE. However, there are slight disparities in the damping rates and the computational burdens are quite different among these two algorithms.

Huang, Jen-Kuang↗

Evaluation of a dual processor implementation for a fault inferring nonlinear detection system

The design of a modified fault inferring nonlinear detection system (FINDS) algorithm for a dual-processor configured flight computer is described. The algorithm was changed in order to divide it into its translational dynamics and rotational kinematics and to use it for parallel execution on the flight computer. The FINDS consists of: (1) a no-fail filter (NFF), (2) a set of test-of-mean detection tests, (3) a bank of first order filters to estimate failure levels in individual sensors, and (4) a decision function. NFF filter performance using flight recorded sensor data is analyzed using a filter autoinitialization routine. The failure detection and isolation capability of the partitioned algorithm is evaluated. A multirate implementation for the bias-free and bias filter gain and covariance matrices is discussed.

Godiwala, P. M.↗

Novel strategies for modal-based structural material identification

Here, we present modal-based methods for model calibration in structural dynamics, and address several key challenges in the solution of gradient-based optimization problems with eigenvalues and eigenvectors, including the solution of singular Helmholtz problems encountered in sensitivity calculations, non-differentiable objective functions caused by mode swapping during optimization, and cases with repeated eigenvalues. Unlike previous literature that relied on direct solution of the eigenvector adjoint equations, we present a parallel iterative domain decomposition strategy (Adjoint Computation via Modal Superposition with Truncation Augmentation) for the solution of the singular Helmholtz problems. For problems with repeated eigenvalues we present a novel Mode Separation via Projection algorithm, and in order to address mode swapping between inverse iterations we present a novel Injective mode ordering metric. We present the implementation of these methods in a massively parallel finite element framework with the ability to use measured modal data to extract unknown structural model parameters from large complex problems. A series of increasingly complex numerical examples are presented that demonstrate the implementation and performance of the methods in a massively parallel finite element framework [7], [5], using gradient-based optimization techniques in the Rapid Optimization Library (ROL) [21].

36 MATERIALS SCIENCE↗

Multifrontal Non-negative Matrix Factorization

Non-negative matrix factorization (Nmf) is an important tool in high-performance large scale data analytics with applications ranging from community detection, recommender system, feature detection and linear and non-linear unmixing. While traditional Nmf works well when the data set is relatively dense, however, it may not extract sufficient structure when the data is extremely sparse. Specifically, traditional Nmf fails to exploit the structured sparsity of the large and sparse data sets resulting in dense factors. We propose a new algorithm for performing Nmf on sparse data that we call multifrontal Nmf (Mf-Nmf) since it borrows several ideas from the multifrontal method for unconstrained factorization (e.g. LU and QR). We also present an efficient shared memory parallel implementation of Mf-Nmf and discuss its performance and scalability. We conduct several experiments on synthetic and realworld datasets and demonstrate the usefulness of the algorithm by comparing it against standard baselines. We obtain a speedup of 1.2x to 19.5x on 24 cores with an average speed up of 10.3x across all the real world datasets.

Sao, Piyush↗

Implementation and Characterization of Three-Dimensional Particle-in-Cell Codes on Multiple-Instruction-Multiple-Data Massively Parallel Supercomputers

A three-dimensional electrostatic particle-in-cell (PIC) plasma simulation code has been developed on coarse-grain distributed-memory massively parallel computers with message passing communications. Our implementation is the generalization to three-dimensions of the general concurrent particle-in-cell (GCPIC) algorithm. In the GCPIC algorithm, the particle computation is divided among the processors using a domain decomposition of the simulation domain. In a three-dimensional simulation, the domain can be partitioned into one-, two-, or three-dimensional subdomains ("slabs," "rods," or "cubes") and we investigate the efficiency of the parallel implementation of the push for all three choices. The present implementation runs on the Intel Touchstone Delta machine at Caltech; a multiple-instruction-multiple-data (MIMD) parallel computer with 512 nodes. We find that the parallel efficiency of the push is very high, with the ratio of communication to computation time in the range 0.3%-10.0%. The highest efficiency (> 99%) occurs for a large, scaled problem with 64(sup 3) particles per processing node (approximately 134 million particles of 512 nodes) which has a push time of about 250 ns per particle per time step. We have also developed expressions for the timing of the code which are a function of both code parameters (number of grid points, particles, etc.) and machine-dependent parameters (effective FLOP rate, and the effective interprocessor bandwidths for the communication of particles and grid points). These expressions can be used to estimate the performance of scaled problems--including those with inhomogeneous plasmas--to other parallel machines once the machine-dependent parameters are known.

Lyster, P. M.↗