Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 757 records · Page 42

Multigrid solvers on parallel computers

Massively parallel computers, as considered in this investigation, are not yet available. However, a large-scale parallel computer cannot usefully be designed before the hypothetical algorithms which will employ it are studied. Most of the studies of parallel partial differential equations (PDE) solvers are based on solution techniques much slower (on sequential machines) than multigrid methods. Multigrid methods are highly parallelizable. Each of their processes can simultaneously be performed at all grid points. The present investigation is concerned with a preliminary exploration of the potential of multigrid, or, more generally, Multi-Level Adaptive Techniques (MLAT) on computers with many processors. Basic processes are considered, taking into account coarse-grid approximation, relaxation, coarse-grid corrections, full multigrid algorithms, nonlinear problems and eigenvalue problems, fine-to-coarse correction, and chains of problems. Details of parallel multigrid processing are also examined.

Brandt, A.↗

A space-time tracking algorithm for high occupancy events at future colliders

We propose to explore the potential advantages of a newclass of tracking algorithms loosely inspired by the Hough transformconcept and where we include the time of arrival of each hit as anadditional coordinate to be treated in the same way as a spatialcoordinate. A remarkable property of this algorithm is that theexecution time is proportional to the total number of hits to beprocessed, making it particularly attractive for high occupancysituations expected at future colliders. The particular structureof the algorithm also lends itself naturally to parallel hardwareimplementations which, combined to its intrinsic flexibility, shouldprovide a powerful tool for triggering at future colliders. To probethe effectiveness of the algorithm, we apply it to a quasi-realisticsimulated environment of a possible future muon collider experimentand report the performance.

Casarsa, Massimo [INFN, Trieste; Royal Inst. Tech.↗

A generic fine-grained parallel C

With the present availability of parallel processors of vastly different architectures, there is a need for a common language interface to multiple types of machines. The parallel C compiler, currently under development, is intended to be such a language. This language is based on the belief that an algorithm designed around fine-grained parallelism can be mapped relatively easily to different parallel architectures, since a large percentage of the parallelism has been identified. The compiler generates a FORTH-like machine-independent intermediate code. A machine-dependent translator will reside on each machine to generate the appropriate executable code, taking advantage of the particular architectures. The goal of this project is to allow a user to run the same program on such machines as the Massively Parallel Processor, the CRAY, the Connection Machine, and the CYBER 205 as well as serial machines such as VAXes, Macintoshes and Sun workstations.

Hamet, L.↗

A comparison using APPL and PVM for a parallel implementation of an unstructured grid generation program

Efforts to parallelize the VGRIDSG unstructured surface grid generation program are described. The inherent parallel nature of the grid generation algorithm used in VGRIDSG was exploited on a cluster of Silicon Graphics IRIS 4D workstations using the message passing libraries Application Portable Parallel Library (APPL) and Parallel Virtual Machine (PVM). Comparisons of speed up are presented for generating the surface grid of a unit cube and a Mach 3.0 High Speed Civil Transport. It was concluded that for this application, both APPL and PVM give approximately the same performance, however, APPL is easier to use.

Arthur, Trey↗

GPU Accelerated Prognostics

Prognostic methods enable operators and maintainers to predict the future performance for critical systems. However, these methods can be computationally expensive and may need to be performed each time new information about the system becomes available. In light of these computational requirements, we have investigated the application of graphics processing units (GPUs) as a computational platform for real-time prognostics. Recent advances in GPU technology have reduced cost and increased the computational capability of these highly parallel processing units, making them more attractive for the deployment of prognostic software. We present a survey of model-based prognostic algorithms with considerations for leveraging the parallel architecture of the GPU and a case study of GPU-accelerated battery prognostics with computational performance results.

Prognostics↗

An Overview of a Trajectory-Based Solution for En Route and Terminal Area Self-Spacing to Include Parallel Runway Operations

This paper presents an overview of an algorithm specifically designed to support NASA's Airborne Precision Spacing concept. This airborne self-spacing concept is trajectory-based, allowing for spacing operations prior to the aircraft being on a common path. This implementation provides the ability to manage spacing against two traffic aircraft, with one of these aircraft operating to a parallel dependent runway. Because this algorithm is trajectory-based, it also has the inherent ability to support required-time-of-arrival (RTA) operations

Abbott, Terence S.↗

Advanced Shuttle Strategies for Parallel QCCD Architectures

Trapped ions (TIs) are at the forefront of quantum computing implementation, offering unparalleled coherence, fidelity, and connectivity. However, the scalability of TI systems is hampered by the limited capacity of individual ion traps, necessitating intricate ion shuttling for advanced computational tasks. The quantum charge-coupled device (QCCD) framework has emerged as a promising solution, facilitating ion mobility for universal quantum computation. Current QCCD architectures predominantly feature a linear topology, which is increasingly recognized as inefficient for complex quantum operations. Anticipating the shift toward more efficacious designs, this article introduces an innovative quantum scheduling strategy optimized for parallel QCCD topologies. Our strategy proposes a probabilistic formula for ion movement, alongside ingenious methods for local layer generation and layer compression, yielding a significant reduction in ion shuttle times. Through simulations, we demonstrate that our strategy not only substantially outstrips the linear model but also exhibits better performance over other parallel strategies that employ greedy algorithms. This is achieved through our nuanced resolution of complexities, such as traffic blocks and trap capacity limitations. The consequent reduction in shuttle operations leads to lower energy consumption and an enhancement in the quantum computer's fidelity, ultimately accelerating program execution times.

43 PARTICLE ACCELERATORS↗

Memory-Aware External Facelist Calculation: A Data-Parallel Atomic Hash Counting Approach

Unstructured volumetric meshes serve as fundamental data representations in various scientific simulations and analyses. They play a crucial role in representing complex computational domains and are essential for important numerical techniques, such as finite element analysis. Whenever such a mesh is read from a file, streamed in-situ, or generated by algorithms, scientific visualization libraries rely on calculating the external surface of a geometry, named “external facelist”, to produce a polygonal mesh for rendering. Consequently, external facelist calculation has become one of the most widely used algorithms in the scientific visualization domain, necessitating optimal performance. In this paper, we explore relevant work on external facelist calculation algorithms in two common visualization libraries, VTK and Viskores, assess their performance and memory constraints, and introduce a novel memory-aware external facelist calculation algorithm employing an atomic hash counting approach. This algorithm fully leverages Viskores' data-parallel primitive operations, facilitating its execution across diverse many-core architectures. Our algorithm features the lowest memory footprint on the GPU and the second-lowest on the CPU among all evaluated methods, and it also delivers the fastest performance on both CPU and GPU. It has been made available under an open-source license in the VTK and Viskores visualization systems.

Tsalikis, Spiros [Kitware] (ORCID:0000000151137195↗

Unorthodox parallelization for Bayesian quantum state estimation

Quantum state tomography (QST) allows for the reconstruction of quantum states through measurements and some inference technique under the assumption of repeated state preparations. Bayesian inference provides a promising platform to achieve both efficient QST and accurate uncertainty quantification, yet is generally plagued by the computational limitations associated with long Markov chains. In this work, we present a novel Bayesian QST approach that leverages modern distributed parallel computer architectures to efficiently sample a D-dimensional Hilbert space. Using a parallelized preconditioned Crank–Nicholson Metropolis–Hastings algorithm, we demonstrate our approach on simulated data and experimental results from IBM Quantum systems up to four qubits, showing significant speedups through parallelization. Although highly unorthodox in pooling independent Markov chains, our method proves remarkably practical, with validation ex post facto via diagnostics like the intrachain autocorrelation time. We conclude by discussing scalability to higher-dimensional systems, offering a path toward efficient and accurate Bayesian characterization of large quantum systems.

Bayesian inference↗

VTK-m User's' Guide (V.1.7)

High-performance computing relies on ever finer threading. Advances in processor technology include ever greater numbers of cores, hyperthreading, accelerators with integrated blocks of cores, and special vectorized instructions, all of which require more software parallelism to achieve peak performance. Traditional visualization solutions cannot support this extreme level of concurrency. Extreme scale systems require a new programming model and a fundamental change in how we design algorithms. To address these issues we created VTK-m: the visualization toolkit for multi-/many-core architectures. VTK-m supports a number of algorithms and the ability to design further algorithms through a top-down design with an emphasis on extreme parallelism. VTK-m also provides support for finding and building links across topologies, making it possible to perform operations that determine manifold surfaces, interpolate generated values, and find adjacencies. Although VTK-m provides a simplified high-level interface for programming, its template-based code removes the overhead of abstraction. VTK-m simplifies the development of parallel scientific visualization algorithms by providing a framework of supporting functionality that allows developers to focus on visualization operations. Consider the listings in Figure 1.1 that compares the size of the implementation for the Marching Cubes algorithm in VTK-m with the equivalent reference implementation in the CUDA software development kit. Because VTK-m internally manages the parallel distribution of work and data, the VTK-m implementation is shorter and easier to maintain. Additionally, VTK-m provides data abstractions not provided by other libraries that make code written in VTK-m more versatile.This book includes contributions from the VTK-m community including the VTK-m development team and the user community.

97 MATHEMATICS AND COMPUTING↗

The VTK-m Users' Guide (V.1.9)

High-performance computing relies on ever finer threading. Advances in processor technology include ever greater numbers of cores, hyperthreading, accelerators with integrated blocks of cores, and special vectorized instructions, all of which require more software parallelism to achieve peak performance. Traditional visualization solutions cannot support this extreme level of concurrency. Extreme scale systems require a new programming model and a fundamental change in how we design algorithms. To address these issues we created VTK-m: the visualization toolkit for multi-/many-core architectures. VTK-m supports a number of algorithms and the ability to design further algorithms through a top-down design with an emphasis on extreme parallelism. VTK-m also provides support for finding and building links across topologies, making it possible to perform operations that determine manifold surfaces, interpolate generated values, and find adjacencies. Although VTK-m provides a simplified high-level interface for programming, its template-based code removes the overhead of abstraction. VTK-m simplifies the development of parallel scientific visualization algorithms by providing a framework of supporting functionality that allows developers to focus on visualization operations. Consider the listings in Figure 1.1 that compares the size of the implementation for the Marching Cubes algorithm in VTK-m with the equivalent reference implementation in the CUDA software development kit. Because VTK-m internally manages the parallel distribution of work and data, the VTK-m implementation is shorter and easier to maintain. Additionally, VTK-m provides data abstractions not provided by other libraries that make code written in VTK-m more versatile.

97 MATHEMATICS AND COMPUTING↗

The VTK-m Users' Guide (V.2.0)

High-performance computing relies on ever finer threading. Advances in processor technology include ever greater numbers of cores, hyperthreading, accelerators with integrated blocks of cores, and special vectorized instructions, all of which require more software parallelism to achieve peak performance. Traditional visualization solutions cannot support this extreme level of concurrency. Extreme scale systems require a new programming model and a fundamental change in how we design algorithms. To address these issues we created VTK-m: the visualization toolkit for multi-/many-core architectures. VTK-m supports a number of algorithms and the ability to design further algorithms through a top-down design with an emphasis on extreme parallelism. VTK-m also provides support for finding and building links across topologies, making it possible to perform operations that determine manifold surfaces, interpolate generated values, and find adjacencies. Although VTK-m provides a simplified high-level interface for programming, its template-based code removes the overhead of abstraction. VTK-m simplifies the development of parallel scientific visualization algorithms by providing a framework of supporting functionality that allows developers to focus on visualization operations. Consider the listings in Figure 1.1 that compares the size of the implementation for the Marching Cubes algorithm in VTK-m with the equivalent reference implementation in the CUDA software development kit. Because VTK-m internally manages the parallel distribution of work and data, the VTK-m implementation is shorter and easier to maintain. Additionally, VTK-m provides data abstractions not provided by other libraries that make code written in VTK-m more versatile.

97 MATHEMATICS AND COMPUTING↗

ORNL_AISD_NiPt

This dataset describes the nickel-platinum (NiPt) solid solution binary alloy, where the two constituent elements nickel (Ni) and platinum (Pt) are randomly placed on the face centered cubic (FCC) crystal structure, with the lattice constant of 3.840 angstroms. The dataset comprises data for three different sizes of the crystal structure: 256 atoms, 864 atoms, and 2,048 atoms, each of which contains 1900 configurations. For each size of the crystal structure, the data set was generated for concentrations ranging from 0at% of Pt to 100at% of Pt in the NiPt binary system, with increasing the concentration of Pt in the system every 5at%. For each one of the chemical compositions, 100 random configurations were generated, each with a different random seed. Each of the output files contains the mass, type, atomic coordinates, energy per atom, and forces in x, y, and z directions respectively. For each atomic configuration, the output was collected every 150 steps during the minimization stage and every 1000 steps during the replica exchange stage. Large-scale Atomic/Molecular Massively Parallel Simulator (LAMMPS) [1], which is a molecular dynamics code, was used to generate data for NiPt alloy. The simulation used the interatomic potential for NiPt binary system MEAM_LAMMPS_KimSeolJi_2017_PtNi__MO_020840179467_001 [3] from the OpenKIM library (Open Knowledgebase of Interatomic Models) [2]. This potential was developed based on the second nearest-neighbor modified embedded-atom method (2NN MEAM). The simulation process begins with the generation of the random NiPt structure and follows with the short minimization and replica exchange simulation. The minimization procedure adjusts atomic coordinates and performs energy minimization, which typically leads to a local potential energy minimum. The method used for the minimization was the conjugate gradient algorithm. A short replica exchange (parallel tempering) simulation involves four replicas (ensembles) of a system and follows the minimization stage. Multiple snapshots of the configuration were collected during the minimization and replica exchange stages. NiPt alloy is interesting due to its magnetic and charge transfer properties [4]. The data is provided in three compressed zipped folders: atoms256.zip, atoms864.zip, atoms2048.zip Each zipped folder contains the data that describes crystals of size 256 atoms, 864 atoms, and 2,048 atoms respectively. Each one of the three zipped folders contains the data structured in the following way: -Ni_ground_state.cfg --> atomic configuration for the pure nickel -Pt_ground_state.cfg --> atomic configuration for the pure platinum -Pt#_filtered --> folders containing atomic configurations for #at% concentration of platinum. The folder contains 100 atomic configurations, each saved in a subfolder. Each subfolder named config* is associated with a specific atomic configuration. Each of these subfolders contains files with .cfg format, corresponding to outputs for each atomic configuration The total number of atomic configurations contained in atoms256.zip is 65,046. The total number of atomic configurations contained in atoms864.zip is 63,936. The total number of atomic configurations contained in atoms2048.zip is 61,997. The total number of atomic configurations spanned by the entire dataset is 190,979. References [1] https://www.lammps.org/ [2] https://openkim.org/ [3] https://openkim.org/id/MEAM_LAMMPS_KimSeolJi_2017_PtNi__MO_020840179467_001 [4] El-Gendy, Ahmed A. and Hampel, Silke and Büccchner, Bernd and Klingeler, Rüdiger, Tuneable magnetic properties of carbon-shielded NiPt-nanoalloys, RSC Adv., volume 6, issue 57, pages 52427-52433, 2016, The Royal Society of Chemistry, doi:10.1039/C6RA05910D

36 MATERIALS SCIENCE↗

ORNL_AISD_NiPt_108atoms

This dataset describes the nickel-platinum (NiPt) solid solution binary alloy, where the two constituent elements nickel (Ni) and platinum (Pt) are randomly placed on the face centered cubic (FCC) crystal structure, with the lattice constant of 3.840 angstroms. The dataset comprises data for crystal structures with 108 atoms with 1,900 configurations. The data set was generated for concentrations ranging from 0at% of Pt to 100at% of Pt in the NiPt binary system, with increasing the concentration of Pt in the system every 5at%. For each one of the chemical compositions, 100 random configurations were generated, each with a different random seed. Each of the output files contains the mass, type, atomic coordinates, energy per atom, and forces in x, y, and z directions respectively. For each atomic configuration, the output was collected every 150 steps during the minimization stage and every 1000 steps during the replica exchange stage. Large-scale Atomic/Molecular Massively Parallel Simulator (LAMMPS) [1], which is a molecular dynamics code, was used to generate data for NiPt alloy. The simulation used the interatomic potential for NiPt binary system 'MEAM_LAMMPS_KimSeolJi_2017_PtNi__MO_020840179467_001' [3] from the OpenKIM library (Open Knowledgebase of Interatomic Models) [2]. This potential was developed based on the second nearest-neighbor modified embedded-atom method (2NN MEAM). The simulation process begins with the generation of the random NiPt structure and follows with the short minimization and replica exchange simulation. The minimization procedure adjusts atomic coordinates and performs energy minimization, which typically leads to a local potential energy minimum. The method used for the minimization was the conjugate gradient algorithm. A short replica exchange (parallel tempering) simulation involves four replicas (ensembles) of a system and follows the minimization stage. Multiple snapshots of the configuration were collected during the minimization and replica exchange stages. NiPt alloy is interesting due to its magnetic and charge transfer properties [4]. The data is provided in a compressed zipped folders atoms108.zip. The zipped folder contains the data structured in the following way: - Ni_ground_state.cfg --> atomic configuration for the pure nickel - Pt_ground_state.cfg --> atomic configuration for the pure platinum - Pt#_filtered --> folders containing atomic configurations for #at% concentration of platinum. The folder contains 100 atomic configurations, each saved in a subfolder - Each subfolder named config* is associated with a specific atomic configuration. Each of these subfolders contains files with .cfg format, corresponding to outputs for each atomic configuration The total number of atomic configurations contained in atoms108.zip is 66,132. This dataset is an extension to the dataset ORNL_AISD_NiPt [5] that has been previously released with crystal structures of 256 atoms, 864 atoms, and 2,048 atoms, with the same methodology for data collection. References [1] https://www.lammps.org/ [2] https://openkim.org/ [3] https://openkim.org/id/MEAM_LAMMPS_KimSeolJi_2017_PtNi__MO_020840179467_001 [4] El-Gendy, Ahmed A. and Hampel, Silke and Büchner, Bernd and Klingeler, Rüdiger, Tuneable magnetic properties of carbon-shielded NiPt-nanoalloys, RSC Adv., volume 6, issue 57, pages 52427-52433, 2016, The Royal Society of Chemistry, doi:10.1039/C6RA05910D [5] M. Karabin, M. Lupo Pasini, and M. Eisenbach. ORNL_AISD_NiPt. United States: N. p., 2023. Web. doi:10.13139/OLCF/1958172.

36 MATERIALS SCIENCE↗

Segmentation of remotely sensed data using parallel region growing

The improved spatial resolution of the new earth resources satellites will increase the need for effective utilization of spatial information in machine processing of remotely sensed data. One promising technique is scene segmentation by region growing. Region growing can use spatial information in two ways: only spatially adjacent regions merge together, and merging criteria can be based on region-wide spatial features. A simple region growing approach is described in which the similarity criterion is based on region mean and variance (a simple spatial feature). An effective way to implement region growing for remote sensing is as an iterative parallel process on a large parallel processor. A straightforward parallel pixel-based implementation of the algorithm is explored and its efficiency is compared with sequential pixel-based, sequential region-based, and parallel region-based implementations. Experimental results from on aircraft scanner data set are presented, as is a discussioon of proposed improvements to the segmentation algorithm.

Tilton, J. C.↗

Soft-output decoding algorithms in iterative decoding of turbo codes

In this article, we present two versions of a simplified maximum a posteriori decoding algorithm. The algorithms work in a sliding window form, like the Viterbi algorithm, and can thus be used to decode continuously transmitted sequences obtained by parallel concatenated codes, without requiring code trellis termination. A heuristic explanation is also given of how to embed the maximum a posteriori algorithms into the iterative decoding of parallel concatenated codes (turbo codes). The performances of the two algorithms are compared on the basis of a powerful rate 1/3 parallel concatenated code. Basic circuits to implement the simplified a posteriori decoding algorithm using lookup tables, and two further approximations (linear and threshold), with a very small penalty, to eliminate the need for lookup tables are proposed.

Benedetto, S.↗

Soft-Output Decoding Algorithms in Iterative Decoding of Turbo Codes

In this article, we present two versions of a simplified maximum a posteriori decoding algorithm. The algorithms work in a sliding window form, like the Viterbi algorithm, and can thus be used to decode continuously transmitted sequences obtained by parallel concatenated codes, without requiring code trellis termination. A heuristic explanation is also given of how to embed the maximum a posteriori algorithms into the iterative decoding of parallel concatenated codes (turbo codes). The performances of the two algorithms are compared on the basis of a powerful rate 1/3 parallel concatenated code. Basic circuits to implement the simplified a posteriori decoding algorithm using lookup tables, and two further approximations (linear and threshold), with a very small penalty, to eliminate the need for lookup tables are proposed.

Benedetto, S.↗

Quantum multi-programming for Grover’s search

Quantum multi-programming is a method utilizing contemporary noisy intermediate-scale quantum computers by executing multiple quantum circuits concurrently. Despite early research on it, the research remains on quantum gates or small-size quantum algorithms without correlation. In this paper, we propose a quantum multi-programming (QMP algorithm for Grover's search. Our algorithm decomposes Grover's algorithm by the partial diffusion operator and executes the decomposed circuits in parallel by QMP. We proved that this new algorithm increases the rotation angle of the Grover operator which, as a result, increases the success probability. The new algorithm is implemented on IBM quantum computers and compared with the canonical Grover's algorithm and other variations of Grover's algorithms. So, the empirical tests validate that our new algorithm outperforms other variations of Grover's algorithms as well as the canonical Grover's algorithm.

97 MATHEMATICS AND COMPUTING↗