Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,315 records · Page 73

High Impact Weather Forecasts and Warnings with the GOES-R Geostationary Lightning Mapper (GLM)

The Geostationary Operational Environmental Satellite (GOES-R) is the next series to follow the existing GOES system currently operating over the Western Hemisphere. A major advancement over the current GOES include a new capability for total lightning detection (cloud and cloud-to-ground flashes) from the Geostationary Lightning Mapper (GLM). The GLM will operate continuously day and night with near-uniform spatial resolution of 8 km with a product refresh rate of less than 20 sec over the Americas and adjacent oceanic regions. This will aid in forecasting severe storms and tornado activity, and convective weather impacts on aviation safety and efficiency. In parallel with the instrument development, a GOES-R Risk Reduction Science Team and Algorithm Working Group Lightning Applications Team have begun to develop cal/val performance monitoring tools and new applications using the GLM alone, in conjunction with other instruments, and merged or blended integrated observing system products combining satellite, radar, in-situ and numerical models. Proxy total lightning data from the NASA Lightning Imaging Sensor (LIS) on the Tropical Rainfall Measuring Mission (TRMM) satellite and regional ground-based lightning networks are being used to develop the pre-launch algorithms, test data sets, and applications, as well as improve our knowledge of thunderstorm initiation and evolution. In this presentation we review the planned implementation of the instrument and suite of operational algorithms.

Goodman, Steven J.↗

Low-Thrust Trajectory Optimization with Simplified SQP Algorithm

The problem of low-thrust trajectory optimization in highly perturbed dynamics is a stressing case for many optimization tools. Highly nonlinear dynamics and continuous thrust are each, separately, non-trivial problems in the field of optimal control, and when combined, the problem is even more difficult. This paper de-scribes a fast, robust method to design a trajectory in the CRTBP (circular restricted three body problem), beginning with no or very little knowledge of the system. The approach is inspired by the SQP (sequential quadratic programming) algorithm, in which a general nonlinear programming problem is solved via a sequence of quadratic problems. A few key simplifications make the algorithm presented fast and robust to initial guess: a quadratic cost function, neglecting the line search step when the solution is known to be far away, judicious use of end-point constraints, and mesh refinement on multiple shooting with fixed-step integration.In comparison to the traditional approach of plugging the problem into a “black-box” NLP solver, the methods shown converge even when given no knowledge of the solution at all. It was found that the only piece of information that the user needs to provide is a rough guess for the time of flight, as the transfer time guess will dictate which set of local solutions the algorithm could converge on. This robustness to initial guess is a compelling feature, as three-body orbit transfers are challenging to design with intuition alone. Of course, if a high-quality initial guess is available, the methods shown are still valid.We have shown that endpoints can be efficiently constrained to lie on 3-body repeating orbits, and that time of flight can be optimized as well. When optimizing the endpoints, we must make a trade between converging quickly on sub-optimal endpoints or converging more slowly on end-points that are arbitrarily close to optimal. It is easy for the mission design engineer to adjust this trade based on the problem at hand.The biggest limitation to the algorithm at this point is that multi-revolution transfers (greater than 2 revolutions) do not work nearly as well. This restriction comes in because the relationship between node 1 and node N becomes increasingly nonlinear as the angular distance grows. Trans-fers with more than about 1.5 complete revolutions generally require the line search to improve convergence. Future work includes: Comparison of this algorithm with other established tools; improvements to how multiple-revolution transfers are handled; parallelization of the Jacobian computation; in-creased efficiency for the line search; and optimization of many more trajectories between a variety of 3-body orbits.

Parrish, Nathan L.↗

INSPiRE – An Approach to Mission Quality Management using Network Slicing for Space Applications

Managing traffic between the Earth-Moon and Earth-Mars is a complex process requiring significant investment in resources and expertise at NASA. INSPiRE improves the performance of space networks by enabling a dynamic re-configuration process that works for any mixed topology over a heterogeneous and multi-vendor network. To achieve the desired functionality, INSPiRE incorporates a set of algorithms, machine learning processes, and policy inference to handle unpredictable, disruptive events. INSPiRE draws parallels from the current notion of the 3GPP (5G and beyond) Network Slicing approach, where the same physical network divides into several virtual networks, and for each of these virtual networks, there is a guaranteed Quality of Service for the missions that they serve.

cognitive communications↗

Hardware Implementation of Lossless Adaptive and Scalable Hyperspectral Data Compression for Space

On-board lossless hyperspectral data compression reduces data volume in order to meet NASA and DoD limited downlink capabilities. The technique also improves signature extraction, object recognition and feature classification capabilities by providing exact reconstructed data on constrained downlink resources. At JPL a novel, adaptive and predictive technique for lossless compression of hyperspectral data was recently developed. This technique uses an adaptive filtering method and achieves a combination of low complexity and compression effectiveness that far exceeds state-of-the-art techniques currently in use. The JPL-developed 'Fast Lossless' algorithm requires no training data or other specific information about the nature of the spectral bands for a fixed instrument dynamic range. It is of low computational complexity and thus well-suited for implementation in hardware. A modified form of the algorithm that is better suited for data from pushbroom instruments is generally appropriate for flight implementation. A scalable field programmable gate array (FPGA) hardware implementation was developed. The FPGA implementation achieves a throughput performance of 58 Msamples/sec, which can be increased to over 100 Msamples/sec in a parallel implementation that uses twice the hardware resources This paper describes the hardware implementation of the 'Modified Fast Lossless' compression algorithm on an FPGA. The FPGA implementation targets the current state-of-the-art FPGAs (Xilinx Virtex IV and V families) and compresses one sample every clock cycle to provide a fast and practical real-time solution for space applications.

FPGA implementation↗

High Performance Parallel Computational Nanotechnology

At a recent press conference, NASA Administrator Dan Goldin encouraged NASA Ames Research Center to take a lead role in promoting research and development of advanced, high-performance computer technology, including nanotechnology. Manufacturers of leading-edge microprocessors currently perform large-scale simulations in the design and verification of semiconductor devices and microprocessors. Recently, the need for this intensive simulation and modeling analysis has greatly increased, due in part to the ever-increasing complexity of these devices, as well as the lessons of experiences such as the Pentium fiasco. Simulation, modeling, testing, and validation will be even more important for designing molecular computers because of the complex specification of millions of atoms, thousands of assembly steps, as well as the simulation and modeling needed to ensure reliable, robust and efficient fabrication of the molecular devices. The software for this capacity does not exist today, but it can be extrapolated from the software currently used in molecular modeling for other applications: semi-empirical methods, ab initio methods, self-consistent field methods, Hartree-Fock methods, molecular mechanics; and simulation methods for diamondoid structures. In as much as it seems clear that the application of such methods in nanotechnology will require powerful, highly powerful systems, this talk will discuss techniques and issues for performing these types of computations on parallel systems. We will describe system design issues (memory, I/O, mass storage, operating system requirements, special user interface issues, interconnects, bandwidths, and programming languages) involved in parallel methods for scalable classical, semiclassical, quantum, molecular mechanics, and continuum models; molecular nanotechnology computer-aided designs (NanoCAD) techniques; visualization using virtual reality techniques of structural models and assembly sequences; software required to control mini robotic manipulators for positional control; scalable numerical algorithms for reliability, verifications and testability. There appears no fundamental obstacle to simulating molecular compilers and molecular computers on high performance parallel computers, just as the Boeing 777 was simulated on a computer before manufacturing it.

Saini, Subhash↗

Genarris 2.0: A Random Structure Generator for Molecular Crystals

Genarris is an open source Python package for generating random molecular crystal structures with physical constraints for seeding crystal structure prediction algorithms and training machine learning models. Here we present a new version of the code, containing several major improvements. A MPI-based parallelization scheme has been implemented, which facilitates the seamless sequential execution of user-defined workflows. A new method for estimating the unit cell volume based on the single molecule structure has been developed using a machine-learned model trained on experimental structures. A new algorithm has been implemented for generating crystal structures with molecules occupying special Wyckoff positions. A new hierarchical structure check procedure has been developed to detect unphysical close contacts efficiently and accurately. New intermolecular distance settings have been implemented for strong hydrogen bonds. To demonstrate these new features, we study two specific cases: benzene and glycine. Genarris finds the experimental structures of the two polymorphs of benzene and the three polymorphs of glycine. Program summary Program Title: Genarris 2.0 Program Files doi: http://dx.doi.org/10.17632/grx6mz4pjn.1 Licensing provisions: BSD-3 Clause Programming language: Python, C External routines/libraries: Spglib, ASE, pymatgen, SciPy, mpi4py, scikit-learn, PyTorch, FHI-aims. Nature of problem: Molecular crystal structure prediction. Solution method: Genarris 2.0 generates molecular crystal structures over the 230 space groups, on general and special Wyckoff positions, using physical constraints. Down-sampling of the generated structures may be performed subsequently, based on molecular crystal packing descriptors and an unsupervised machine learning algorithm. Lastly, ab initio structure relaxation may be performed for the final pool. Depending on the user-defined workflow implemented, Genarris may be used to generate diverse molecular crystal datasets to seed evolutionary algorithms or to train machine learning algorithms or as a standalone crystal structure prediction method. Restrictions: For crystal structure generation, the molecule of interest must be semi-rigid with no bond rotational degrees of freedom. Unusual features: Genarris 2.0 is a highly distributed program, making use of MPI for Python parallelization. The user has the ability to design and implement workflows by executing a user-defined list of procedures. Genarris 2.0 offers new features including a machine learning model for estimating the molecular volume in the solid state from the single molecule structure, structure generation in special Wyckoff positions of space groups, hierarchical structure checks including rigorous treatment of non-orthogonal structures, and clustering and down-selection workflows combining first principles simulations with machine learning. (C) 2020 Elsevier B.V. All rights reserved.

Crystal structure prediction↗

PISCES: An environment for parallel scientific computation

The parallel implementation of scientific computing environment (PISCES) is a project to provide high-level programming environments for parallel MIMD computers. Pisces 1, the first of these environments, is a FORTRAN 77 based environment which runs under the UNIX operating system. The Pisces 1 user programs in Pisces FORTRAN, an extension of FORTRAN 77 for parallel processing. The major emphasis in the Pisces 1 design is in providing a carefully specified virtual machine that defines the run-time environment within which Pisces FORTRAN programs are executed. Each implementation then provides the same virtual machine, regardless of differences in the underlying architecture. The design is intended to be portable to a variety of architectures. Currently Pisces 1 is implemented on a network of Apollo workstations and on a DEC VAX uniprocessor via simulation of the task level parallelism. An implementation for the Flexible Computing Corp. FLEX/32 is under construction. An introduction to the Pisces 1 virtual computer and the FORTRAN 77 extensions is presented. An example of an algorithm for the iterative solution of a system of equations is given. The most notable features of the design are the provision for several granularities of parallelism in programs and the provision of a window mechanism for distributed access to large arrays of data.

Pratt, T. W.↗

A Hierarchical and Distributed Approach for Mapping Large Applications to Heterogeneous Grids using Genetic Algorithms

In this paper, we propose a distributed approach for mapping a single large application to a heterogeneous grid environment. To minimize the execution time of the parallel application, we distribute the mapping overhead to the available nodes of the grid. This approach not only provides a fast mapping of tasks to resources but is also scalable. We adopt a hierarchical grid model and accomplish the job of mapping tasks to this topology using a scheduler tree. Results show that our three-phase algorithm provides high quality mappings, and is fast and scalable.

Sanyal, Soumya↗

GPU Lossless Hyperspectral Data Compression System for Space Applications

On-board lossless hyperspectral data compression reduces data volume in order to meet NASA and DoD limited downlink capabilities. At JPL, a novel, adaptive and predictive technique for lossless compression of hyperspectral data, named the Fast Lossless (FL) algorithm, was recently developed. This technique uses an adaptive filtering method and achieves state-of-the-art performance in both compression effectiveness and low complexity. Because of its outstanding performance and suitability for real-time onboard hardware implementation, the FL compressor is being formalized as the emerging CCSDS Standard for Lossless Multispectral & Hyperspectral image compression. The FL compressor is well-suited for parallel hardware implementation. A GPU hardware implementation was developed for FL targeting the current state-of-the-art GPUs from NVIDIA(Trademark). The GPU implementation on a NVIDIA(Trademark) GeForce(Trademark) GTX 580 achieves a throughput performance of 583.08 Mbits/sec (44.85 MSamples/sec) and an acceleration of at least 6 times a software implementation running on a 3.47 GHz single core Intel(Trademark) Xeon(Trademark) processor. This paper describes the design and implementation of the FL algorithm on the GPU. The massively parallel implementation will provide in the future a fast and practical real-time solution for airborne and space applications.

Graphic Processor Units↗

FLEET: Flexible Efficient Ensemble Training for Heterogeneous Deep Neural Networks

Parallel training of an ensemble of Deep Neural Networks (DNN) on a cluster of nodes is an effective approach to shorten the process of neural network architecture search and hyper-parameter tuning for a given learning task. Prior efforts have shown that data sharing, where the common preprocessing operation is shared across the DNN training pipelines, saves computational resources and improves pipeline efficiency. Data sharing strategy, however, performs poorly for a heterogeneous set of DNNs where each DNN has varying computational needs and thus different training rate and convergence speed. This paper proposes FLEET, a flexible ensemble DNN training framework for efficiently training a heterogeneous set of DNNs. We build FLEET via several technical innovations. We theoretically prove that an optimal resource allocation is NP-hard and propose a greedy algorithm to efficiently allocate resources for training each DNN with data sharing. We integrate data-parallel DNN training into ensemble training to mitigate the differences in training rates and introduce checkpointing into this context to address the issue of different convergence speeds. Experiments show that FLEET significantly improves the training efficiency of DNN ensembles without compromising the quality of the result.

Guan, Hui↗

On-board landmark navigation and attitude reference parallel processor system

An approach to autonomous navigation and attitude reference for earth observing spacecraft is described along with the landmark identification technique based on a sequential similarity detection algorithm (SSDA). Laboratory experiments undertaken to determine if better than one pixel accuracy in registration can be achieved consistent with onboard processor timing and capacity constraints are included. The SSDA is implemented using a multi-microprocessor system including synchronization logic and chip library. The data is processed in parallel stages, effectively reducing the time to match the small known image within a larger image as seen by the onboard image system. Shared memory is incorporated in the system to help communicate intermediate results among microprocessors. The functions include finding mean values and summation of absolute differences over the image search area. The hardware is a low power, compact unit suitable to onboard application with the flexibility to provide for different parameters depending upon the environment.

Gilbert, L. E.↗

A fully implicit, asymptotic-preserving, semi-Lagrangian algorithm for the time dependent anisotropic heat transport equation

In this paper, we extend the operator-split asymptotic-preserving, semi-Lagrangian algorithm for time dependent anisotropic heat transport equation proposed in Chacón et al. (2014) [18] to use a fully implicit time integration with backward differentiation formulas. The proposed implicit method can deal with arbitrary heat-transport anisotropy ratios $\mathcal{X}$∥ /$ \mathcal{X}$⟂ $\ggg$ 1 (with $\mathcal{X}$∥, $ \mathcal{X}$⟂ the parallel and perpendicular heat diffusivities, respectively) in complicated magnetic field topologies in an accurate and efficient manner. Further, the implicit algorithm is second-order accurate temporally and demonstrates an accurate treatment at boundary layers (e.g., island separatrices), which was not ensured by the operator-split implementation. The condition number of the resulting algebraic system is independent of the anisotropy ratio, and is inverted with preconditioned GMRES. We propose a simple preconditioner that renders the finite-dimensional linear operator compact, resulting in mesh-independent convergence rates for topologically simple magnetic fields, and convergence rates scaling as ~ (NΔt) 1/4 (with N the total mesh size and Δt the timestep) in topologically complex magnetic-field configurations. We demonstrate the accuracy and performance of the approach with test problems of varying complexity, including an analytically tractable boundary-layer problem in a straight magnetic field, and a topologically complex magnetic field featuring magnetic islands with extreme anisotropy ratios $\mathcal{X}$∥ /$ \mathcal{X}$⟂ = 10 10 ) .

97 MATHEMATICS AND COMPUTING↗

Physical Retrievals of Over-Ocean Rain Rate from Multichannel Microwave Imagery: Theoretical Characteristics of Normalized Polarization and Scattering Indices - Part 1

Microwave rain rate retrieval algorithms have most often been formulated in terms of the raw brightness temperatures observed by one or more channels of a satellite radiometer. Taken individually, single-channel brightness temperatures generally represent a near-arbitrary combination of positive contributions due to liquid water emission and negative contributions due to scattering by ice and/or visibility of the radiometrically cold ocean surface. Unfortunately, for a given rain rate, emission by liquid water below the freezing level and scattering by ice particles above the freezing level are rather loosely coupled in both a physical and statistical sense. Furthermore, microwave brightness temperatures may vary significantly (approx. 30-70 K) in response to geophysical parameters other than liquid water and precipitation. Because of these complications, physical algorithms which attempt to directly invert observed brightness temperatures have typically relied on the iterative adjustment of detailed micro-physical profiles or cloud models, guided by explicit forward microwave radiative transfer calculations. In support of an effort to develop a significantly simpler and more efficient inversion-type rain rate algorithm, the physical information content of two linear transformations of single-frequency, dual-polarization brightness temperatures is studied: the normalized polarization difference P of Petty and Katsaros (1990, 1992), which is intended as a measure of footprint-averaged rain cloud transmittance for a given frequency; and a scattering index S (similar to the polarization corrected temperature of Spencer et al.,1989) which is sensitive almost exclusively to ice. A reverse Monte Carlo radiative transfer model is used to elucidate the qualitative response of these physically distinct single-frequency indices to idealized 3-dimensional rain clouds and to demonstrate their advantages over raw brightness temperatures both as stand-alone indices of precipitation activity and as primary variables in physical, multichannel rain rate retrieval schemes. As a byproduct of the present analysis, it is shown that conventional plane-parallel analyses of the well-known foot-print-filling problem for emission-based algorithms may in some cases give seriously misleading results.

Petty, G. W.↗

Adaptive Control and Scaling Approach for the Emulation of Dynamic Sub-scale Torque Loads

Previous research by the authors proposed a control and scaling approach for emulating dynamic sub-scale torque loads. This approach produced an emulation controller that successfully regulated the dynamic behavior of a sub-scale electro-mechanical system intended to be a dynamical representation of the sub-scale turboelectric powertrain of a single-aisle commercial aircraft. The sub-scale system provides an environment without turbomachinery or rotors for the initial testing of electrified aircraft propulsion (EAP) control algorithms as they would be applied to a full-scale EAP system. The sub-scale turbomachinery/rotor torque loads were produced by electric machines (EMs) driven by the control and scaling approach, a full-scale turboelectric powertrain model, and an advanced EAP control algorithm. Although successfully tested, this approach produces an emulation controller that does not guarantee asymptotic stability. Modifying the original control law and integrating adaptive control techniques into the emulation controller allows the designer to guarantee asymptotic stability. This paper introduces the idea behind the emulation controller modifications, derives the controller, proves asymptotic stability, describes the implementation of the controller on a sub-scale electro-mechanical system intended to represent a parallel hybrid-electric turbofan engine, and describes the testing of a full-scale advanced EAP control algorithm. The turbofan engine and emulation controller performance are compared to results obtained using the previous, non-adaptive control and scaling approach.

adaptive↗

Adaptive Control and Scaling Approach for the Emulation of Dynamic Subscale Torque Loads

Previous research by the authors proposed a control and scaling approach for emulating dynamic sub-scale torque loads. This approach produced an emulation controller that successfully regulated the dynamic behavior of a sub-scale electro-mechanical system intended to be a dynamical representation of the sub-scale turboelectric powertrain of a single-aisle commercial aircraft. The sub-scale system provides an environment without turbomachinery or rotors for the initial testing of electrified aircraft propulsion (EAP) control algorithms as they would be applied to a full-scale EAP system. The sub-scale turbomachinery/rotor torque loads were produced by electric machines (EMs) driven by the control and scaling approach, a full-scale turboelectric powertrain model, and an advanced EAP control algorithm. Although successfully tested, this approach produces an emulation controller that does not guarantee asymptotic stability. Modifying the original control law and integrating adaptive control techniques into the emulation controller allows the designer to guarantee asymptotic stability. This paper introduces the idea behind the emulation controller modifications, derives the controller, proves asymptotic stability, describes the implementation of the controller on a sub-scale electro-mechanical system intended to represent a parallel hybrid-electric turbofan engine, and describes the testing of a full-scale advanced EAP control algorithm. The turbofan engine and emulation controller performance are compared to results obtained using the previous, non-adaptive control and scaling approach.

adaptive↗

Fast and Accurate Intersections on a Sphere

We introduce a fast, high-precision algorithm for calculating intersections between great circle arcs and lines of constant latitude on the unit sphere. We first propose a simplified intersection point formula with improved speed and numerical robustness over the ones traditionally implemented in geoscience software. We then show how algorithms based on the concept of error-free transformations (EFT) can be applied to evaluate this formula within a relative error bound that is on the order of machine precision. Here, we demonstrate that, with a vectorized and parallelized implementation, this enhanced accuracy is achieved with no compute time overhead compared to a direct calculation in hardware floating point, making our algorithm suitable for performance-sensitive applications like regridding of high-resolution climate data. In contrast, evaluating our formula using high-precision data types like quadruple precision and arbitrary precision, or using the robust intersection computation routines from the Computational Geometry Algorithms Library, leads to significant computational overhead, especially since these alternatives inhibit vectorization. More generally, our work demonstrates how EFT techniques can be combined and extended to implement nontrivial geometric calculations with high accuracy and speed.

Environmental sciences↗

Motion Cueing Algorithm Modification for Improved Turbulence Simulation

Atmospheric turbulence cueing produced by flight simulator motion systems has been less than satisfactory because the turbulence profiles have been attenuated by the motion cueing algorithms. Cardullo and Ellor initially addressed this problem by directly porting the turbulence model output to the motion system. Reid and Robinson addressed the problem by employing a parallel aircraft model, which is only stimulated by the turbulence inputs and adding a filter specially designed to pass the higher turbulence frequencies. There have been advances in motion cueing algorithm development at the Man-Machine Systems Laboratory, at SUNY Binghamton. In particular, the system used to generate turbulence cues has been studied. The Reid approach, implemented by Telban and Cardullo, was employed to augment the optimal motion cueing algorithm installed at the NASA LaRC Simulation Laboratory, driving the Visual Motion Simulator. In this implementation, the output of the primary flight channel was added to the output of the turbulence channel and then sent through a non-linear cueing filter. The cueing filter is an adaptive filter; therefore, it is not desirable for the output of the turbulence channel to be augmented by this type of filter. The likelihood of the signal becoming divergent was also an issue in this design. After testing on-site it became apparent that the architecture of the turbulence algorithm was generating unacceptable cues. As mentioned above, this cueing algorithm comprised a filter that was designed to operate at low bandwidth. Therefore, the turbulence was also filtered, augmenting the cues generated by the model. If any filtering is to be done to the turbulence, it will utilize a filter with a much higher bandwidth, above the frequencies produced by the aircraft response to turbulence. The authors have developed an implementation wherein only the signal from the primary flight channel passes through the nonlinear cueing filter. This paper discusses three new algorithms. Testing shows that the new methods provide the pilot with a more realistic sensation of turbulence; the cues are not attenuated by algorithm. Results of offline testing show the credibility of the models. Offline test verification was based primarily on the evaluation of the power spectral density of the outputs and the time response.

Ercole, Anthony V.↗

Research in Computational Aeroscience Applications Implemented on Advanced Parallel Computing Systems

Improving the numerical linear algebra routines for use in new Navier-Stokes codes, specifically Tim Barth's unstructured grid code, with spin-offs to TRANAIR is reported. A fast distance calculation routine for Navier-Stokes codes using the new one-equation turbulence models is written. The primary focus of this work was devoted to improving matrix-iterative methods. New algorithms have been developed which activate the full potential of classical Cray-class computers as well as distributed-memory parallel computers.

Wigton, Larry↗