Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,621 records · Page 90

Efficient Mosaicking of Spitzer Space Telescope Images

A parallel version of the MOPEX software, which generates mosaics of infrared astronomical images acquired by the Spitzer Space Telescope, extends the capabilities of the prior serial version. In the parallel version, both the input image space and the output mosaic space are divided among the available parallel processors. This is the only software that performs the point-source detection and the rejection of spurious imaging effects of cosmic rays required by Spitzer scientists. This software includes components that implement outlier-detection algorithms that can be fine-tuned for a particular set of image data by use of a number of adjustable parameters. This software has been used to construct a mosaic of the Spitzer Infrared Array Camera Shallow Survey, which comprises more than 17,000 exposures in four wavelength bands from 3.6 to 8 m and spans a solid angle of about 9 square degrees. When this software was executed on 32 nodes of the 1,024-processor Cosmos cluster computer at NASA s Jet Propulsion Laboratory, a speedup of 8.3 was achieved over the serial version of MOPEX. The performance is expected to improve dramatically once a true parallel file system is installed on Cosmos.

Jacob, Joseph↗

Image Gradient Decomposition for Parallel and Memory-Efficient Ptychographic Reconstruction

Ptychography is a popular microscopic imaging modality for many scientific discoveries and sets the record for highest image resolution. Unfortunately, the high image resolution for ptychographic reconstruction requires significant amount of memory and computations, forcing many applications to compromise their image resolution in exchange for a smaller memory footprint and a shorter reconstruction time. In this paper, we propose a novel image gradient decomposition method that significantly reduces the memory footprint for ptychographic reconstruction by tessellating image gradients and diffraction measurements into tiles. In addition, we propose a parallel image gradient decomposition method that enables asynchronous point-to-point communications and parallel pipelining with minimal overhead on a large number of GPUs. Our experiments on a Titanate material dataset (PbTiO3) with 16632 probe locations show that our Gradient Decomposition algorithm reduces memory footprint by 51 times. In addition, it achieves time-to-solution within 2.2 minutes by scaling to 4158 GPUs with a super-linear strong scaling efficiency at 364% compared to runtimes at 6 GPUs. This performance is 2.7 times more memory efficient, 9 times more scalable and 86 times faster than the state-of-the-art algorithm.

Wang, Xiao↗

Deep Reinforcement Learning Based Control of Wind Turbines for Fast Frequency Response

In order to fulfill vital auxiliary grid services, such as load regulation, spin and non-spin reserve provision, and frequency support during emergencies, there is often a requirement for certain wind farms to operate in de-loaded modes. Leveraging the swift response capabilities of wind farms, this study demonstrates that reserving power in de-loaded modes can significantly enhance power grid stability and reliability during system contingencies. Controlling wind farms optimally for frequency support is intricate due to the nonlinearity of models and controllers and the complexity of wind farm interactions with power systems. Here, to address this challenge, this paper introduces a novel approach that integrates wind turbines into reinforcement learning-based solutions for frequency response. This innovative methodology utilizes the state-of-the-art reinforcement learning algorithm known as the surrogate-gradient-based evolutionary strategy. The proposed learning-based algorithm provides continuous control of wind farm output to rapidly stabilize system frequency and prevent unnecessary trips of under-frequency load shedding relays. To facilitate efficient training, parallel computing techniques are employed. The proposed methodology is evaluated on a modified IEEE-39 bus system, and simulation results reveal its efficacy in reliably supporting power system frequency and preventing the need for unnecessary load shedding.

Gao, Wei [Argonne National Laboratory (ANL), Argon↗

Quandary

Quandary numerically simulates and optimizes the time-evolution of open quantum systems. The underlying dynamics are modelled by Lindblad's master equation, a linear ordinary differential equation (ODE) describing quantum systems interacting with the environment. Quandary solves this ODE numerically by applying a time-stepping integration scheme, and utilizes a gradient-based optimization approach to determine optimal control pulses that drive the quantum system to a desired target state. Two optimization objectives are considered: (a) Unitary gate optimization that finds controls to realize a unitary gate transformation, and (b) optimal reset that aims to drive the quantum system to the ground states. Gradient-based optimization schemes utilizing Petsc's Tao optimization package are applied to generate control pulses that minimize the respective measure. To evaluate the gradient of the objective function, the discrete adjoint method is used while leveraging techniques from Algorithmic Differentiation to produce exact and consistent gradients. To mitigate excessive execution run times, the software can be build together with the XBraid software library which provides a parallelization strategy to distribute the time-evolution of the underlying dynamics onto multiple processor.

Petersson, NilsA.↗

Designing a Framework for Solving Multiobjective Simulation Optimization Problems

Multiobjective simulation optimization (MOSO) problems are optimization problems with multiple conflicting objectives, where evaluation of at least one of the objectives depends on a black-box numerical code or real-world experiment, which we refer to as a simulation. Whereas an extensive body of research is dedicated to developing new algorithms and methods for solving these and related problems, it is challenging and time-consuming to integrate these techniques into real-world production-ready solvers. This is partly because of the diversity and complexity of modern state-of-the-art MOSO algorithms and methods and partly because of the complexity and specificity of many real-world problems and their corresponding computing environments. The complexity of this problem is only compounded when introducing potentially complex and/or domain-specific surrogate-modeling techniques, problem formulations, design spaces, and data acquisition functions. Here, this paper carefully surveys the current state of the art in MOSO algorithms, techniques, and solvers, as well as problem types and computational environments where MOSO is commonly applied. We then present several key challenges in the design of a parallel multiobjective simulation optimization framework (ParMOO) and how they have been addressed. Finally, we provide two case studies demonstrating how customized ParMOO solvers can be quickly built and deployed to solve real-world MOSO problems.

engineering design optimization↗

CSRI Summer Proceedings 2020

The Computer Science Research Institute (CSRI) brings university faculty and students to Sandia for focused collaborative research on Department of Energy (DOE) computer and computational science problems. The institute provides an opportunity for university researchers to learn about problems in computer and computational science at DOE laboratories. Participants conduct leading-edge research, interact with scientists and engineers at the laboratories, and help transfer results of their research to programs at the labs. Some specific CSRI research interest areas are: scalable solvers, optimization, adaptivity and mesh refinement, graph-based, discrete, and combinatorial algorithms, uncertainty estimation, mesh generation, dynamic load-balancing, virus and other malicious-code defense, visualization, scalable cluster computers, data-intensive computing, environments for scalable computing, parallel input/output, advanced architectures, and theoretical computer science. The CSRI Summer Program is organized by CSRI and typically includes the organization of a weekly seminar series and the publication of a summer proceedings. In 2020, the CSRI summer program was executed completely virtually; all student interns worked from home, due to the COVID-19 pandemic.

97 MATHEMATICS AND COMPUTING↗

Multiprocessor architecture: Synthesis and evaluation

Multiprocessor computed architecture evaluation for structural computations is the focus of the research effort described. Results obtained are expected to lead to more efficient use of existing architectures and to suggest designs for new, application specific, architectures. The brief descriptions given outline a number of related efforts directed toward this purpose. The difficulty is analyzing an existing architecture or in designing a new computer architecture lies in the fact that the performance of a particular architecture, within the context of a given application, is determined by a number of factors. These include, but are not limited to, the efficiency of the computation algorithm, the programming language and support environment, the quality of the program written in the programming language, the multiplicity of the processing elements, the characteristics of the individual processing elements, the interconnection network connecting processors and non-local memories, and the shared memory organization covering the spectrum from no shared memory (all local memory) to one global access memory. These performance determiners may be loosely classified as being software or hardware related. This distinction is not clear or even appropriate in many cases. The effect of the choice of algorithm is ignored by assuming that the algorithm is specified as given. Effort directed toward the removal of the effect of the programming language and program resulted in the design of a high-level parallel programming language. Two characteristics of the fundamental structure of the architecture (memory organization and interconnection network) are examined.

Standley, Hilda M.↗

Local parallel models for integration of stereo matching constraints and intrinsic image combination

Parallel relaxation computations such as those of connectionist networks offer a useful model for constraint integration and intrinsic image combination in developing a general-purpose stereo matching algorithm. This paper describes such a stereo algorithm that incorporates hierarchical, surface-structure, and edge-appearance constraints that are redefined and integrated at the level of individual candidate matches. The algorithm produces a high percentage of correct decisions on a wide variety of stereo pairs. Its few errors arise when the correlation measures defined by the constraints are either weakened or ambiguous, as in the case of periodic patterns in the images. Two additional mechanisms are discussed for overcoming the remaining errors.

Stewart, Charles V.↗

A spectral collocation method for compressible, non-similar boundary layers

An efficient and highly accurate algorithm based on a spectral collocation method is developed for numerical solution of the compressible, two-dimensional and axisymmetric boundary layer equations. The numerical method incorporates a fifth-order, fully implicit marching scheme in the streamwise (timelike) dimension and a spectral collocation method based on Chebyshev polynomial expansions in the wall-normal (spacelike) dimension. The spectral collocation algorithm is used to derive the nonsimilar mean velocity and temperature profiles in the boundary layer of a 'fuselage' (cylinder) in a high-speed (Mach 5) flow parallel to its axis. The stability of the flow is shown to be sensitive to the gradual streamwise evolution of the mean flow and it is concluded that the effects of transverse curvature on stability should not be ignored routinely.

Pruett, C. D.↗

Implicit transient finite element structural computations on MIMD systems - FETI vs. direct solvers

A domain decomposition method for implicit schemes that require significantly less storage and is several times faster than factorization algorithms is proposed. The transient domain decomposition method is an extension of the finite element tearing and interconnecting (FETI) method for the solution of static problems. Serial and parallel performance results obtained using the CRAY Y-MP/8 and the iPSC-860/128 systems demonstrate that the FETI method is superior to both serial and parallel direct methods.

Crivelli, Luis↗

Formation Flying With Decentralized Control in Libration Point Orbits

A decentralized control framework is investigated for applicability of formation flying control in libration orbits. The decentralized approach, being non-hierarchical, processes only direct measurement data, in parallel with the other spacecraft. Control is accomplished via linearization about a reference libration orbit with standard control using a Linear Quadratic Regulator (LQR) or the GSFC control algorithm. Both are linearized about the current state estimate as with the extended Kalman filter. Based on this preliminary work, the decentralized approach appears to be feasible for upcoming libration missions using distributed spacecraft.

Folta, David↗

NASA's Robotic Lunar Lander Development Project

Since early 2005, NASA's Robotic Lunar Lander Development (RLLD) office at NASA MSFC, in partnership with the Applied Physics Laboratory (APL), has developed mission concepts and preformed risk-reduction activities to address planetary science and exploration objectives uniquely met with landed missions. The RLLD team developed several concepts for lunar human-exploration precursor missions to demonstrate precision landing and in-situ resource utilization, a multi-node lunar geophysical network mission, either as a stand-alone mission, or as part of the International Lunar Network (ILN), a Lunar Polar Volatiles Explorer and a Mercury lander mission for the Planetary Science decadal survey, and an asteroid rendezvous and landing mission for the Exploration Precursor Robotics Mission (xPRM) office. The RLLD team has conducted an extensive number of risk-reduction activities in areas common to all lander concepts, including thruster testing, propulsion thermal control demonstration, composite deck design and fabrication, and landing leg stability and vibration. In parallel, the team has developed two robotic lander testbeds providing closed-loop, autonomous hover and descent activities for integration and testing of flight-like components and algorithms. A compressed-air test article had its first flight in September 2009 and completed over 150 successful flights. This small test article (107 kg dry/146 kg wet) uses a central throttleable thruster to offset gravity, plus 3 descent thrusters (~37lbf ea) and 6 attitude-control thrusters (~12lbf ea) to emulate the flight system with pulsed operation over approximately 10s of flight time. The test article uses carbon composite honeycomb decks, custom avionics (COTS components assembled in-house), and custom flight and ground software. A larger (206 kg dry/322 kg wet), hydrogen peroxide-propelled vehicle began flight tests in spring 2011 and fly over 30 successful flights to a maximum altitude of 30m. The monoprop testbed also uses a central gravity-canceling thruster and 3 descent thrusters, but has 12 attitude-control thrusters and a maximum flight time of over a minute. The testbed uses aluminum ortho-grid decks, an LN200-1 IMU, Roke Manor Radar Altimeter, Illunis optical cameras, Novatel Pro-Pak GPS truth data system, Pressure transducers & thermocouples for housekeeping, "In-Control" ground system software, and the core Flight Executive (cFE) modular software environment. The peroxide lander testbed is able to accept other sensors and algorithms for testing, both from within NASA and from other customers. Through these activities, the RLLD team has significantly reduced technical risks for all small and medium class robotic landers for the Moon and other airless planetary bodies.

Cohen, Barbara A.↗

CALIPSO Lidar Calibration at 532 nm: Version 4 Nighttime Algorithm

Data products from the Cloud-Aerosol Lidar with Orthogonal Polarization (CALIOP) on board Cloud-Aerosol Lidar and Infrared Pathfinder Satellite Observations (CALIPSO) were recently updated following the implementation of new (version 4) calibration algorithms for all of the level 1 attenuated backscatter measurements. In this work we present the motivation for and the implementation of the version 4 nighttime 532 nm parallel channel calibration. The nighttime 532 nm calibration is the most fundamental calibration of CALIOP data, since all of CALIOP’s other radiometric calibration procedures – i.e., the 532 nm daytime calibration and the 1064 nm calibrations during both nighttime and daytime – depend either directly or indirectly on the 532 nm nighttime calibration. The accuracy of the 532 nm nighttime calibration has been significantly improved by raising the molecular normalization altitude from 30-34 km to 36-39 km to substantially reduce stratospheric aerosol contamination. Due to the greatly reduced molecular number density and consequently reduced signal-to-noise ratio (SNR) at these higher altitudes, the signal is now averaged over a larger number of samples using data from multiple adjacent granules. As well, an enhanced strategy for filtering the radiation-induced noise from high energy particles was adopted. Further, the meteorological model used in the earlier versions has been replaced by the improved MERRA-2 model. An aerosol scattering ratio of 1.01 ± 0.01 is now explicitly used for the calibration altitude. These modifications lead to globally revised calibration coefficients which are, on average, 2-3% lower than in previous data releases. Further, the new calibration procedure is shown to eliminate biases at high altitudes that were present in earlier versions and consequently leads to an improved representation of stratospheric aerosols. Validation results using airborne lidar measurements are also presented. Biases relative to collocated measurements acquired by the Langley Research Center (LaRC) airborne high spectral resolution lidar (HSRL) are reduced from 3.6% ± 2.2% in the version 3 data set to 1.6% ± 2.4 % in the version 4 release.

Jayanta Kar↗

Characterization of a Pixelated Cadmium Telluride Detector System Using a Polychromatic X-Ray Source and Gold Nanoparticle-Loaded Phantoms for Benchtop X-Ray Fluorescence Imaging

In this paper, the imaging dose and scan time have been considered as the two major constraints for routine benchtop x-ray fluorescence computed tomography (XFCT) imaging. One way to address this issue is to acquire x-ray fluorescence (XRF) signals in parallel through a 2D array of single-crystal detectors or a pixelated detector along with the cone-beam x-ray source. To identify a detector system suitable for this purpose, a commercially available, fully spectroscopic cadmium telluride (CdTe) pixelated detector, HEXITEC (High-Energy X-ray Imaging Technology), was tested under the experimental conditions optimized for benchtop XFCT imaging of gold nanoparticles (GNPs). Specifically, two different parallel-hole stainless steel collimators were fabricated and coupled with the detector for seamless integration into our existing benchtop cone-beam XFCT system. After the detector deployment, this benchtop XFCT system was used to detect XRF photons from GNP-loaded phantoms. A pixel-merging algorithm was introduced to enhance the sensitivity of XRF photon detection thereby minimizing the scan time. The effect of pixel-level charge sharing correction algorithms was investigated within the context of benchtop XFCT imaging. The detector energy resolution, in terms of the full width at half maximum (FWHM) values at different gold K-shell XRF energies, was also determined. Of the two charge sharing correction algorithms examined, the charge sharing addition gave better sensitivity than the charge sharing discrimination (csd). On the other hand, under the current experimental conditions, the energy resolution of the HEXITEC detector was the best with the csd and estimated to be 1.56 keV FWHM at 66-69 keV photon energy. Overall, despite some degradation of the detector energy resolution (compared with typical single crystal CdTe detectors), the HEXITEC detector enabled parallel data acquisition under the experimental conditions typical of benchtop XFCT imaging and operated well within our benchtop XFCT setup.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Krylov subspace methods on supercomputers

A short survey of recent research on Krylov subspace methods with emphasis on implementation on vector and parallel computers is presented. Conjugate gradient methods have proven very useful on traditional scalar computers, and their popularity is likely to increase as three-dimensional models gain importance. A conservative approach to derive effective iterative techniques for supercomputers has been to find efficient parallel/vector implementations of the standard algorithms. The main source of difficulty in the incomplete factorization preconditionings is in the solution of the triangular systems at each step. A few approaches consisting of implementing efficient forward and backward triangular solutions are described in detail. Polynomial preconditioning as an alternative to standard incomplete factorization techniques is also discussed. Another efficient approach is to reorder the equations so as to improve the structure of the matrix to achieve better parallelism or vectorization. An overview of these and other ideas and their effectiveness or potential for different types of architectures is given.

Saad, Youcef↗

PUMIPic: A mesh-based approach to unstructured mesh Particle-In-Cell on GPUs

Unstructured mesh particle-in-cell, PIC, simulations executing on the current and next generation of massively parallel systems require new methods for both the mesh and particles to achieve performance and scalability on GPUs. The traditional approach to implementing PIC simulations defines data structures and algorithms in terms of particles with a full copy of the unstructured mesh on every process. To effectively scale the unstructured mesh and particles, mesh-based PIC uses the unstructured mesh as the predominant data structure with the particles stored in terms of the mesh entities. Here, this paper details the PUMIPic library, a framework for developing efficient and performance-portable mesh-based PIC simulations on GPU systems. A pseudo physics simulation based on a five-dimensional gyro-kinetic code for modeling plasma physics is used to examine the performance of PUMIPic. Scaling studies of the unstructured mesh partition and number of particles are performed up to 4096 nodes of the Summit system at Oak Ridge National Laboratory. The studies show that mesh-based PIC can utilize a partitioned mesh and maintain scaling up to system limitations.

97 MATHEMATICS AND COMPUTING↗

A robot conditioned reflex system modeled after the cerebellum.

Reduction of a theory of cerebellar function to computer software for the control of a mechanical manipulator. This reduction is achieved by considering the cerebellum, along with the higher-level brain centers which control it, as a type of finite-state machine with input entering the cerebellum via mossy fibers from the periphery and output from the cerebellum occurring via Purkinje cells. It is hypothesized that the cerebellum learns by an error-correction system similar to Perceptron training algorithms. An electromechanical model of the cerebellum is then developed for the control of a mechanical arm. The problem of modeling the granular layer which selects the set of parallel fibers which are active at any instant of time is considered, and a relevance matrix is constructed to model the relative degree of influence which mossy fibers from the various joints have on the sets of granule cells unique to each joint.

Albus, J. S.↗

Multispectral imaging and analysis system

Arrays of charge coupled devices or linear detector arrays simultaneously obtain spectral reflectance data of different wavelengths for a target area. Several accommodating a particular bandwidth, are individually associated with each array. Data from the arrays are read out in parallel and applied to a computer or microprocessor for processing. The microprocessor serves to analyze the data in real time and if possible, in accordance with hard-wired algorithms. The data are then displayed as an image on an appropriate display unit and also recorded for further use. The display system may be operationally connected to receive a terrain image such that the target area and the analyzed spectral reflectance data are superimposed and simultaneously displayed.

Goetz, A. F. H.↗