Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,639 records · Page 91

Parallel discrete event simulation: A shared memory approach

With traditional event list techniques, evaluating a detailed discrete event simulation model can often require hours or even days of computation time. Parallel simulation mimics the interacting servers and queues of a real system by assigning each simulated entity to a processor. By eliminating the event list and maintaining only sufficient synchronization to insure causality, parallel simulation can potentially provide speedups that are linear in the number of processors. A set of shared memory experiments is presented using the Chandy-Misra distributed simulation algorithm to simulate networks of queues. Parameters include queueing network topology and routing probabilities, number of processors, and assignment of network nodes to processors. These experiments show that Chandy-Misra distributed simulation is a questionable alternative to sequential simulation of most queueing network models.

Reed, Daniel A.↗

Parallel discrete event simulation using shared memory

With traditional event-list techniques, evaluating a detailed discrete-event simulation-model can often require hours or even days of computation time. By eliminating the event list and maintaining only sufficient synchronization to ensure causality, parallel simulation can potentially provide speedups that are linear in the numbers of processors. A set of shared-memory experiments, using the Chandy-Misra distributed-simulation algorithm, to simulate networks of queues is presented. Parameters of the study include queueing network topology and routing probabilities, number of processors, and assignment of network nodes to processors. These experiments show that Chandy-Misra distributed simulation is a questionable alternative to sequential-simulation of most queueing network models.

Reed, Daniel A.↗

Digital and optical shape representation and pattern recognition; Proceedings of the Meeting, Orlando, FL, Apr. 4-6, 1988

The present conference discusses topics in pattern-recognition correlator architectures, digital stereo systems, geometric image transformations and their applications, topics in pattern recognition, filter algorithms, object detection and classification, shape representation techniques, and model-based object recognition methods. Attention is given to edge-enhancement preprocessing using liquid crystal TVs, massively-parallel optical data base management, three-dimensional sensing with polar exponential sensor arrays, the optical processing of imaging spectrometer data, hybrid associative memories and metric data models, the representation of shape primitives in neural networks, and the Monte Carlo estimation of moment invariants for pattern recognition.

Juday, Richard D.↗

Rapid Corner Detection Using FPGAs

In order to perform precision landings for space missions, a control system must be accurate to within ten meters. Feature detection applied against images taken during descent and correlated against the provided base image is computationally expensive and requires tens of seconds of processing time to do just one image while the goal is to process multiple images per second. To solve this problem, this algorithm takes that processing load from the central processing unit (CPU) and gives it to a reconfigurable field programmable gate array (FPGA), which is able to compute data in parallel at very high clock speeds. The workload of the processor then becomes simpler; to read an image from a camera, it is transferred into the FPGA, and the results are read back from the FPGA. The Harris Corner Detector uses the determinant and trace to find a corner score, with each step of the computation occurring on independent clock cycles. Essentially, the image is converted into an x and y derivative map. Once three lines of pixel information have been queued up, valid pixel derivatives are clocked into the product and averaging phase of the pipeline. Each x and y derivative is squared against itself, as well as the product of the ix and iy derivative, and each value is stored in a WxN size buffer, where W represents the size of the integration window and N is the width of the image. In this particular case, a window size of 5 was chosen, and the image is 640 480. Over a WxN size window, an equidistance Gaussian is applied (to bring out the stronger corners), and then each value in the entire window is summed and stored. The required components of the equation are in place, and it is just a matter of taking the determinant and trace. It should be noted that the trace is being weighted by a constant k, a value that is found empirically to be within 0.04 to 0.15 (and in this implementation is 0.05). The constant k determines the number of corners available to be compared against a threshold sigma to mark a valid corner. After a fixed delay from when the first pixel is clocked in (to fill the pipeline), a score is achieved after each successive clock. This score corresponds with an (x,y) location within the image. If the score is higher than the predetermined threshold sigma, then a flag is set high and the location is recorded.

Morfopoulos, Arin C.↗

Improved Finite-Volume Method for Radiative Hydrodynamics

Fully coupled simulations of hydrodynamics and radiative transfer are essential to a number of fields ranging from astrophysics to engineering applications. Of particular interest in this work are hypersonic atmospheric entries and associated experimental apparatus, e.g., shock tubes and high enthalpy testing facilities. The radiative transfer calculations must supply to the CFD a heating term in the energy equation in the form of the divergence of the radiative heat flux and the radiative heat fluxes to bounding surfaces. It is most efficient to solve the radiative transfer equation on the same grid as the CFD solution, and this work presents an algorithm with improved accuracy for such simulations on structured and unstructured grids compared to more conventional approaches. Results will be shown for shock radiation during hypersonic reentry. Issues of parallelization within a radiation sweep will also be discussed.

Wray, Alan↗

The Use of Field Programmable Gate Arrays (FPGA) in Small Satellite Communication Systems

This paper will describe the use of digital Field Programmable Gate Arrays (FPGA) to contribute to advancing the state-of-the-art in software defined radio (SDR) transponder design for the emerging SmallSat and CubeSat industry and to provide advances for NASA as described in the TAO5 Communication and Navigation Roadmap (Ref 4). The use of software defined radios (SDR) has been around for a long time. A typical implementation of the SDR is to use a processor and write software to implement all the functions of filtering, carrier recovery, error correction, framing etc. Even with modern high speed and low power digital signal processors, high speed memories, and efficient coding, the compute intensive nature of digital filters, error correcting and other algorithms is too much for modern processors to get efficient use of the available bandwidth to the ground. By using FPGAs, these compute intensive tasks can be done in parallel, pipelined fashion and more efficiently use every clock cycle to significantly increase throughput while maintaining low power. These methods will implement digital radios with significant data rates in the X and Ka bands. Using these state-of-the-art technologies, unprecedented uplink and downlink capabilities can be achieved in a 1/2 U sized telemetry system. Additionally, modern FPGAs have embedded processing systems, such as ARM cores, integrated inside the FPGA allowing mundane tasks such as parameter commanding to occur easily and flexibly. Potential partners include other NASA centers, industry and the DOD. These assets are associated with small satellite demonstration flights, LEO and deep space applications. MSFC currently has an SDR transponder test-bed using Hardware-in-the-Loop techniques to evaluate and improve SDR technologies.

Varnavas, Kosta↗

Development of the Science Data System for the International Space Station Cold Atom Lab

Cold Atom Laboratory (CAL) is a facility that will enable scientists to study ultra-cold quantum gases in a microgravity environment on the International Space Station (ISS) beginning in 2016. The primary science data for each experiment consists of two images taken in quick succession. The first image is of the trapped cold atoms and the second image is of the background. The two images are subtracted to obtain optical density. These raw Level 0 atom and background images are processed into the Level 1 optical density data product, and then into the Level 2 data products: atom number, Magneto-Optical Trap (MOT) lifetime, magnetic chip-trap atom lifetime, and condensate fraction. These products can also be used as diagnostics of the instrument health. With experiments being conducted for 8 hours every day, the amount of data being generated poses many technical challenges, such as downlinking and managing the required data volume. A parallel processing design is described, implemented, and benchmarked. In addition to optimizing the data pipeline, accuracy and speed in producing the Level 1 and 2 data products is key. Algorithms for feature recognition are explored, facilitating image cropping and accurate atom number calculations.

bose einstein condensate↗

Simulations of particle acceleration in parallel shocks: Direct comparison between Monte Carlo and one-dimensional hybrid codes

We have made a direct comparison between two different computer simulations of a plane, parallel, collisionless shock including particle acceleration to energies typical of those of diffuse ions observed at the earth bow shock. Despite the fact that the one-dimensional hybrid and Monte Carlo techniques employ entirely different algorithms, they give surprisingly close agreement in the overall shapes of the complete distribution functions for protons as well as heavier ions. Both methods show that energetic ions emerge smoothly from the background thermal plasma with approximately the same relative injection rate and that the fraction of the incoming plasma's energy flux that is converted into downstream enthalpy flux of the accelerated population (i.e., the acceleration efficiency) is similar in the two cases. The fraction of the downstream proton distribution made up of superthermal particles is quite large, with at least 10% of the energy flux going into protons with energies above 10 keV. In addition, an upstream precursor, produced by backstreaming energetic particles, is present in both shocks, although the Monte Carlo precursor is considerably longer than that produced in the hybrid shock. These results offer convincing evidence that, at least in these ways, the two simulations are consistent in their description of parallel shock structure and particle acceleration, and they lay the groundwork for development of shock models employing a combination of both methods.

Ellison, Donald C.↗

Improved Characterization of PSC Processes Derived from a Third-Generation CALIOP and MLS Detection and Composition Classification Algorithm

The new 3-year CloudSat and CALIPSO Science Team project described in this poster will use a unique combination of data from the Cloud-Aerosol LIdar with Orthogonal Polarization (CALIOP) instrument on CALIPSO and the Microwave Limb Sounder (MLS) on Aura, in conjunction with supporting meteorological information and detailed modeling studies, to advance our understanding of polar stratospheric cloud (PSC) processes and their role in ozone depletion. We will develop a third-generation (Gen3) PSC detection and composition algorithm that incorporates a new, more robust two-dimensional, multi-channel CALIOP feature detection scheme (2D-McDA). We will also devise and implement an improved two-dimensional PSC composition classification scheme that utilizes multiple parameters (e.g., CALIOP 532-nm parallel and perpendicular scattering ratios, CALIOP 1064-nm total scattering ratio, MLS HNO3 and H2O, ambient temperature, and temperature histories) in a Bayesian approach to determine the most likely PSC composition and help constrain solid PSC particle number density and size/shape. The combined CALIOP/MLS analyses will allow us to study in detail the full life cycle of PSCs and their resulting impact on gas-phase HNO3 and H2O, which should lead to improved parameterizations of PSC microphysics in global CCMs where detailed particle information is not available. The Gen3 CALIOP PSC algorithm will be a natural stepping-stone toward the analysis of data collected during future spaceborne lidar missions, such as NASA’s Atmosphere Observation System (AtmOS) mission currently scheduled for launch late in this decade. We will also investigate possible trends in PSC occurrence and composition over the entire CALIOP data record and through further comparisons with the Stratospheric Aerosol Measurement (SAM) II solar occultation PSC record from 1979-1989. Finally, we will validate the mountain-wave parameterization and PSC schemes used in the UM-UKCA (Unified Model coupled to the United Kingdom Chemistry and Aerosol module) chemistry-climate model through detailed comparisons with earlier CALIOP PSC data products and those developed under this proposal.

CALIPSO↗

Improved Characterization of PSC Processes Derived from a Third-Generation CALIOP and MLS Detection and Composition Classification Algorithm

The new 3-year CloudSat and CALIPSO Science Team project described in this poster will use a unique combination of data from the Cloud-Aerosol LIdar with Orthogonal Polarization (CALIOP) instrument on CALIPSO and the Microwave Limb Sounder (MLS) on Aura, in conjunction with supporting meteorological information and detailed modeling studies, to advance our understanding of polar stratospheric cloud (PSC) processes and their role in ozone depletion. We will develop a third-generation (Gen3) PSC detection and composition algorithm that incorporates a new, more robust two-dimensional, multi-channel CALIOP feature detection scheme (2D-McDA). We will also devise and implement an improved two-dimensional PSC composition classification scheme that utilizes multiple parameters (e.g., CALIOP 532-nm parallel and perpendicular scattering ratios, CALIOP 1064-nm total scattering ratio, MLS HNO3 and H2O, ambient temperature, and temperature histories) in a Bayesian approach to determine the most likely PSC composition and help constrain solid PSC particle number density and size/shape. The combined CALIOP/MLS analyses will allow us to study in detail the full life cycle of PSCs and their resulting impact on gas-phase HNO3 and H2O, which should lead to improved parameterizations of PSC microphysics in global chemistry-climate models (CCMs) where detailed particle information is not available. The Gen3 CALIOP PSC algorithm will be a natural stepping-stone toward the analysis of data collected during future spaceborne lidar missions, such as NASA’s Atmosphere Observation System (AOS) mission currently scheduled for launch late in this decade. We will also investigate possible trends in PSC occurrence and composition over the entire CALIOP data record and through further comparisons with the Stratospheric Aerosol Measurement (SAM) II solar occultation PSC record from 1979-1989. Finally, we will validate the mountain-wave parameterization and PSC schemes used in the UM-UKCA (Unified Model coupled to the United Kingdom Chemistry and Aerosol module) CCM through detailed comparisons with earlier CALIOP PSC data products and those developed under this proposal.

CALIPSO↗

Monte Carlo Explicitly Correlated Second-Order Many-Body Green’s Function Calculations of Semiconductor Band Gaps

A systematically converging series of ab initio, post-density-functional, size-consistent, electron-correlated approximations is desired for predictive computing of felectronic band structures of insulating, semiconducting, and metallic solids. A series that meets all of these desiderata (except the applicability to metals) is ab initio many-body Green's function theory based on Gaussian-type-orbital (GTO) basis sets. Here, its leading-order approximation, the second-order Green's function (GF2) method in the diagonal and frequency-independent approximations with the aug-cc-pVDZ basis set, is applied to the fundamental band gaps of three semiconductors (diamond, silicon, and silicon carbide in the zincblende structure) using cluster models. Corrections are made to the basis-set-incompleteness errors by the explicit-correlation (F12) ansatz (GF2-F12) for the valence band edges. The crystals are modeled as surface-passivated clusters of increasing sizes, whose wave functions are expanded by up to 2709 GTO basis functions. Immense computational costs of these calculations are overcome by the highly scalable stochastic algorithm of the Monte Carlo GF2-F12 method, whose operation cost per state increases only as a cubic power of system size, which has a tiny memory footprint and easily achieves near-perfect parallel efficiency on thousands of CPUs or on hundreds of GPUs. The correlated, F12-corrected highest-occupied and lowest-unoccupied molecular-orbital energy (HOMO-LUMO) gap is 5.78 ± 0.07 eV for C 87 H 76 as compared with the experimental value of the fundamental (indirect) band gap of bulk diamond at 5.48 eV. The correlated, F12-corrected HOMO-LUMO gaps for Si 75 H 76 and Si 32 C 43 H 76 are 2.56 ± 0.15 eV and 3.50 ± 0.12 eV, respectively, which are expected to decrease further with increasing cluster sizes. As a result, the experimental fundamental (indirect) band gaps of bulk silicon and silicon carbide are 1.17 eV and 2.42 eV, respectively.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Pass-efficient methods for compression of high-dimensional turbulent flow data

The future of high-performance computing, specifically on future Exascale computers, will presumably see memory capacity and bandwidth fail to keep pace with data generated, for instance, from massively parallel partial differential equation (PDE) systems. Current strategies proposed to address this bottleneck entail the omission of large fractions of data, as well as the incorporation of in situ compression algorithms to avoid overuse of memory. To ensure that post-processing operations are successful, this must be done in a way that a sufficiently accurate representation of the solution is stored. Moreover, in situations where the input/output system becomes a bottleneck in analysis, visualization, etc., or the execution of the PDE solver is expensive, the number of passes made over the data must be minimized. In the interest of addressing this problem, this work focuses on the utility of pass-efficient, parallelizable, low-rank, matrix decomposition methods in compressing high-dimensional simulation data from turbulent flows. Additionally, a particular emphasis is placed on using coarse representation of the data – compatible with the PDE discretization grid – to accelerate the construction of the low-rank factorization. This includes the presentation of a novel single-pass matrix decomposition algorithm for computing the so-called interpolative decomposition. The methods are described extensively and numerical experiments on two turbulent channel flow data are performed. In the first (unladen) channel flow case, compression factors exceeding 400 are achieved while maintaining accuracy with respect to first- and second-order flow statistics. In the particle-laden case, compression factors of 100 are achieved and the compressed data is used to recover particle velocities. These results show that these compression methods can enable efficient computation of various quantities of interest in both the carrier and disperse phases.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

CP2K: An Electronic Structure and Molecular Dynamics Software Package - Quickstep: Efficient and Accurate Electronic Structure Calculations

CP2K is an open source electronic structure and molecular dynamics software package to perform atomistic simulations of solid-state, liquid, molecular and biological systems. It is especially aimed at massively-parallel and linear-scaling electronic structure methods and state-of-the-art ab-initio molecular dynamics simulations. Excellent performance for electronic structure calculations is achieved using novel algorithms implemented for modern high-performance computing systems. This review revisits the main capabilities of CP2K to perform efficient and accurate electronic structure simulations. The emphasis is put on density functional theory and multiple post-Hartree-Fock methods using the Gaussian and plane wave approach and its augmented all-electron extension. TDK has received funding from the European Research Council (ERC) under the European Union's Horizon 2020 research and innovation programme (grant agreement No. 716142). VRR has been supported by the Swiss National Science Foundation in the form of Ambizione grant No. PZ00P2 174227 and RZK by the Natural Sciences and Engineering Research Council of Canada (NSERC) through Discovery Grants (RGPIN-2016-0505). GKS and CJM are supported by the US Department of Energy, Office of Science, Office of Basic Energy Sciences, Division of Chemical Sciences, Geosciences, and Biosciences. UK based work was funded under the embedded CSE programme of the ARCHER UK National Supercomputing Service (http://www.archer.ac.uk), grants eCSE03-011, eCSE06-6, eCSE08-9, eCSE13-17 and the EPSRC (EP/P022235/1) grant “Surface and Interface Toolkit for the Materials Chemistry Community". Computational resources were provided by the Swiss National Supercomputing Centre (CSCS) and Compute Canada. The generous allocation of computing time on the FPGA-based supercomputer “Noctua" at PC2 is kindly acknowledged.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

mesoflow [SWR-22-56]

Mesoflow is a continuum scale simulation tool developed specifically for modeling transport and chemistry at the mesoscale. Our solver utilizes Cartesian block-structured adaptive mesh refinement to resolve complex surface morphologies (of catalysts/biomass particles among others) directly obtained from X-ray tomography data. An immersed boundary based formulation enables rapid representation of complex geometries prevalent in most mesoporous interfaces. The solver is developed on top of open-source performance portable library, AMReX, providing parallel execution capabilities on current and upcoming high-performance-computing (HPC) architectures. Our flexible software framework enables integration of complex chemical mechanisms at heterogenous interfaces and time-split algorithms for circumventing highly disparate reaction and flow time-scales. Our current studies indicate a ten-fold performance gain by using graphics-processing-units (GPU) compared to a single processor for representative problem sizes (2 million cell mesh).

Sitaraman, Hariswaran↗

Unprecedented cloud resolution in a GPU-enabled full-physics atmospheric climate simulation on OLCF’s summit supercomputer

Clouds represent a key uncertainty in future climate projection. While explicit cloud resolution remains beyond our computational grasp for global climate, we can incorporate important cloud effects through a computational middle ground called the Multi-scale Modeling Framework (MMF), also known as Super Parameterization. This algorithmic approach embeds high-resolution Cloud Resolving Models (CRMs) to represent moist convective processes within each grid column in a Global Climate Model (GCM). The MMF code requires no parallel data transfers and provides a self-contained target for acceleration. This study investigates the performance of the Energy Exascale Earth System Model-MMF (E3SM-MMF) code on the OLCF Summit supercomputer at an unprecedented scale of simulation. Hundreds of kernels in the roughly 10K lines of code in the E3SM-MMF CRM were ported to GPUs with OpenACC directives. A high-resolution benchmark using 4600 nodes on Summit demonstrates the computational capability of the GPU-enabled E3SM-MMF code in a full physics climate simulation.

58 GEOSCIENCES↗

Optimization of the moderators in the STS preliminary design

This report details the results for an optimization of the dimensions of the moderators in the preliminary design of the Spallation Neutron Source Second Target Station (STS). This study uses the optimization algorithms of Dakota and an unstructured mesh model for the moderators in MCNP. More details on the unstructured mesh model and the automated mesh generation can be found in [3]. Parallel to this effort, the same moderator geometries have been optimized using a constructive solid geometry (CSG) MCNP model. More details on this model and its results can be found in [4]. Three optimal designs are selected for each moderator: one that is optimized for maximum peak brightness, one for maximum time-integrated brightness, and one for a combination of peak and time-integrated brightness. The backbone of the optimization work flow is provided by Dakota. For each set of design parameters requested by Dakota, a new solid geometry is automatically built in Creo and SpaceClaim, and subsequently exported to Attila4MC to generate an unstructured mesh geometry for MCNP. After the MCNP calculation is finished, the objective function (e.g., brightness metric) is returned to Dakota. After the new design has been evaluated, a result-file is written, and Dakota proposes the next set of design parameters to be evaluated. The loop continues until a specified convergence criterion has been met. The design parameters of the cylindrical (upper) moderator include the hydrogen radius, the premoderator thickness (top, bottom, radial), the beryllium radius and the horizontal position of the moderator. The crucial design choice is the hydrogen radius. A radius of 62 mm is shown to provide the maximum time-integrated brightness. The maximum peak brightness occurs with a radius of 40 mm. A combined (middle) design, which balances peak and time-integrated brightnesses, is obtained with a hydrogen radius of 50 mm. The premoderator thicknesses and the beryllium radius are slightly larger in the design optimized for time-integrated brightness than in the design optimized for peak brightness. The sensitivity to these two parameters is relatively small close to the optimal configurations. The hydrogen vessel and vacuum vessel wall thicknesses are dependent on the radius of the liquid hydrogen due to structural integrity requirements. The increased wall thicknesses for larger vessels significantly penalize the time-integrated brightness, with the maximum obtainable value reduced by more than 10% relative to earlier studies which used fixed vessel wall thicknesses. The impact of the variable wall thicknesses is much less for the peak brightness and combined brightness designs. The design parameters of the tube (lower) moderator selected for the optimization are the tube length, the annular premoderator thickness, the beryllium radius and the horizontal position of the moderator. The tube length is the crucial parameter and is chosen large (210 mm) and small (125 mm) in the designs optimized for time-integrated and peak brightness respectively. A combined optimal design has a tube length of 170 mm. The premoderator thickness and the beryllium radius are chosen larger in the design optimized for time-integrated brightness.

42 ENGINEERING↗

Portable HCAL reconstruction in the CMS detector using the Alpaka library

CMS has deployed a number of different GPU algorithms at the High-Level Trigger (HLT) in Run 3. As the code base for GPU algorithms continues to grow, the burden for developing and maintaining separate implementations for GPU and CPU becomes increasingly challenging. To mitigate this, CMS has adopted the Alpaka (Abstraction Library for Parallel Kernel Acceleration) library as the performance portability solution to provide a single-code base for parallel execution on both GPUs and CPUs in CMS software (CMSSW). A direct CUDA version of HCAL energy reconstruction, called Minimization At Hcal, Iteratively (MAHI), has been deployed at the HLT in the 2022-2023 data taking period. This contribution will describe how the CUDA version is converted into a portable implementation using the Alpaka library. We will discuss the porting experience from CUDA to Alpaka, the validation process and the performance of the Alpaka version in CPU and GPU.

Kwok, Martin↗

Solution of partial differential equations on vector and parallel computers

The present status of numerical methods for partial differential equations on vector and parallel computers was reviewed. The relevant aspects of these computers are discussed and a brief review of their development is included, with particular attention paid to those characteristics that influence algorithm selection. Both direct and iterative methods are given for elliptic equations as well as explicit and implicit methods for initial boundary value problems. The intent is to point out attractive methods as well as areas where this class of computer architecture cannot be fully utilized because of either hardware restrictions or the lack of adequate algorithms. Application areas utilizing these computers are briefly discussed.

Ortega, J. M.↗