Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallelization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 775 records · Page 43

Aperture size distribution, length, and preferential location of bed-parallel veins in shale

Bed-parallel, calcite-filled veins (BPVs) are common in shale formations, and although they have been widely described in other studies, little is known about their population aperture size distribution. To address this knowledge gap, we analyzed BPV sizes in outcrops and cores from the Vaca Muerta Formation, Neuquén Basin, Argentina; in two cores from the Marcellus Formation, Appalachian Basin, northeast Pennsylvania; and in one core from the Wolfcamp Shale, Delaware Basin, West Texas. Nine out of ten aperture size populations follow a negative exponential distribution, with one following a weak power law. Bed-parallel vein size distribution and intensity vary among formations and within the same shale. We define three groups of distributions: (1) Vaca Muerta outcrops, with the highest BPV intensity and the largest BPVs (cumulative frequency of 4.9 BPVs per meter [BPVs/m] for apertures 0.265 mm to 8.7 cm); (2) Vaca Muerta cores with a similar BPV intensity overall but with no apertures wider than 1.2 cm; and (3) Vaca Muerta, Wolfcamp, and Marcellus cores with the fewest BPVs (cumulative frequency up to 0.63 BPVs/m) and very few wider than 1 cm. Aperture and length in two outcrop data sets are weakly positively correlated and follow power laws with exponents of 0.44 and 0.49. Mechanical interfaces at boundaries between different lithologies exert a strong control on BPV location, with 65–75% of observed interfaces having BPVs along them. Only 25–30% of the BPVs occur at observed material interfaces, however, and unless subtle, unobserved mechanical layering is present, other factors must also control location. BPV intensity and organic richness (TOC) from Vaca Muerta well logs are correlated in some instances but not in others, indicating TOC is not always a good proxy for BPV location or intensity. Furthermore, these findings provide useful information for modeling of hydraulic fracture treatments where BPVs may influence development of the stimulated fracture network, for example by limiting height growth.

58 GEOSCIENCES↗

Scalable line and plane relaxation in a parallel structured multigrid solver

The efficient solution of sparse, linear systems that arise through the discretization of partial differential equations remains a key challenge for a range of high performance scientific simulations. One approach for reducing data movement and improving performance is by exposing and exploiting structure in a problem through the use of robust structured multilevel solvers. By choosing coarsening that preserves the structure of the problem, these methods maintain efficient structured computation and communication throughout the multigrid hierarchy. However, when coarsening is not permitted to be dependent on the operator, anisotropy must be addressed by the smoother — producing error compatible for coarse-grid correction with structured coarsening. Here, the components required in a scalable parallel structured solver are described with a focus on memory and communication efficiency of robust smoothers. While the implementation of communication and memory reduction techniques in smoothers integrated in a complete 3D solver present a significant engineering challenge, a novel approach is proposed that addresses these challenges systematically through a change to the solver’s execution model. Enabled by user-level threading paired with a set of data and communication abstractions, this approach permits seamless aggregation of communication in plane smoothers — directly reusing code for a 2D distributed multilevel cycle. Results show an effective reduction in communication costs for coarse-grid problems, and result in a speedup of 8.7x in smoothing routines shown in Fig. 12 using this approach. This produces a significant improvement to strong scalability while maintaining favorable weak scaling behavior. Finally, a parallel scaling study using a series of refined meshes is included that demonstrates the effectiveness of this approach in an application of interest.

97 MATHEMATICS AND COMPUTING↗

A case study on parallel HDF5 dataset concatenation for high energy physics data analysis

In High Energy Physics (HEP), experimentalists generate large volumes of data that, when analyzed, helps us better understand the fundamental particles and their interactions. This data is often captured in many files of small size, creating a data management challenge for scientists. In order to better facilitate data management, transfer, and analysis on large scale platforms, it is advantageous to aggregate data further into a smaller number of larger files. However, this translation process can consume significant time and resources, and if performed incorrectly the resulting aggregated files can be inefficient for highly parallel access during analysis on large scale platforms. In this paper, we present our case study on parallel I/O strategies and HDF5 features for reducing data aggregation time, making effective use of compression, and ensuring efficient access to the resulting data during analysis at scale. We focus on NOvA detector data in this case study, a large-scale HEP experiment generating many terabytes of data. Here, the lessons learned from our case study inform the handling of similar datasets, thus expanding community knowledge related to this common data management task.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Parallel Quantum-Enhanced Sensing

Quantum metrology takes advantage of quantum correlations to enhance the sensitivity of sensors and measurement techniques beyond their fundamental classical limit, given by the shot-noise limit. The use of both temporal and spatial correlations present in quantum states of light can extend quantum-enhanced sensing to a parallel configuration that can simultaneously probe an array of sensors or independently measure multiple parameters. To this end, we use multispatial-mode bright twin beams of light, which are characterized by independent quantum-correlated spatial subregions in addition to quantum temporal correlations, to probe a four-sensor quadrant plasmonic array. We show that it is possible to independently and simultaneously measure local changes in refractive index for all four sensors with a quantum enhancement in sensitivity in the range of 22% to 24% over the corresponding classical configuration. Finally, these results provide a first step toward highly parallel spatially resolved quantum-enhanced sensing techniques and pave the way toward more complex quantum sensing and quantum imaging platforms.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Parallel algorithms for hyperdynamics and local hyperdynamics

Hyperdynamics (HD) is a method for accelerating the timescale of standard molecular dynamics (MD). It can be used for simulations of systems with an energy potential landscape that is a collection of basins, separated by barriers, where transitions between basins are infrequent. HD enables the system to escape from a basin more quickly while enabling a statistically accurate renormalization of the simulation time, thus effectively boosting the timescale of the simulation. In [Kim, Perez, Voter, J Chem Phys, 139:144110, 2013)1, a local version of HD was formulated, which exploits the intrinsic locality characteristic typical of most systems to mitigate the poor scaling properties of standard HD as the system size is increased. In this paper, we discuss how both HD and local HD can be formulated to run efficiently in parallel. We have implemented these ideas in the LAMMPS MD code, which means HD can be used with any interatomic potential LAMMPS supports. Together, these parallel methods allow simulations of any size to achieve the time acceleration offered by HD (which can be orders of magnitude), at a cost 3-5x that of standard MD. As examples, we performed two simulations of a million-atom system to model the diffusion and clustering of Pt adatoms on a large patch of Pt(100) surface for 80 and 160 μs.

74 ATOMIC AND MOLECULAR PHYSICS↗

Cross-field and parallel dynamics of SOL filaments in TCV

Using recently installed scrape-off layer diagnostics on the tokamak à configuration variable, we characterise the poloidal and parallel properties of turbulent filaments. We access both attached and detached divertor conditions across a wide range of core densities (f G ϵ [0.09, 0.66]) in diverted L-mode plasma configurations. With a gas puff imaging (GPI) diagnostic at the outer midplane we observed filaments with a monotonic increase in radial velocity (from 390 m s –1 to 800 m s –1 ) and cross-field radii (from 8.5 mm to 13.4 mm) with increasing core density. Interpreting the filament behaviour in the context of the two-region model, we find that they populate the ideal-interchange regime (C i ) in discharges at very low densities, and the resistive X (RX)-point regime for all other discharges. The scaling of filament velocity versus size shows good agreement with this interpretation. These results are discussed and compared with previous probe-based measurements for similar conditions, which mostly placed filaments in TCV in the resistive ballooning (RB) regime. In addition, for the first time in TCV, the parallel filament extension is studied by magnetically aligning the GPI measurements at the outboard midplane with a reciprocating probe in the divertor. In agreement with the filaments being in the ideal-interchange and the RX-point regimes, they are found to extend beyond the X-point into the outer divertor leg.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Contingency Analysis Based on Partitioned and Parallel Holomorphic Embedding

In the steady-state contingency analysis, the traditional Newton-Raphson method suffers from non-convergence issues when solving post-outage power flow problems, which hinders the integrity and accuracy of security assessment. In this paper, we propose a novel robust contingency analysis approach based on holomorphic embedding (HE). Here, the HE-based simulator provides theoretical convergence guarantee, which is desirable because it avoids the influence of numerical issues and provides a credible security assessment conclusion. In addition, based on the multi-area characteristics of real-world power systems, a partitioned HE (PHE) method is proposed with an interfacebased partitioning of HE formulation. The PHE method does not undermine the numerical robustness of HE and significantly reduces the computation burden in large-scale contingency analysis. The PHE method is further enhanced by parallel or distributed computation to become parallel PHE (P2HE). Tests on a 458-bus system, a synthetic 419-bus system and a large-scale 21447-bus system demonstrate the advantages of the proposed methods in robustness and efficiency.

42 ENGINEERING↗

Advanced Shuttle Strategies for Parallel QCCD Architectures

Trapped ions (TIs) are at the forefront of quantum computing implementation, offering unparalleled coherence, fidelity, and connectivity. However, the scalability of TI systems is hampered by the limited capacity of individual ion traps, necessitating intricate ion shuttling for advanced computational tasks. The quantum charge-coupled device (QCCD) framework has emerged as a promising solution, facilitating ion mobility for universal quantum computation. Current QCCD architectures predominantly feature a linear topology, which is increasingly recognized as inefficient for complex quantum operations. Anticipating the shift toward more efficacious designs, this article introduces an innovative quantum scheduling strategy optimized for parallel QCCD topologies. Our strategy proposes a probabilistic formula for ion movement, alongside ingenious methods for local layer generation and layer compression, yielding a significant reduction in ion shuttle times. Through simulations, we demonstrate that our strategy not only substantially outstrips the linear model but also exhibits better performance over other parallel strategies that employ greedy algorithms. This is achieved through our nuanced resolution of complexities, such as traffic blocks and trap capacity limitations. The consequent reduction in shuttle operations leads to lower energy consumption and an enhancement in the quantum computer's fidelity, ultimately accelerating program execution times.

43 PARTICLE ACCELERATORS↗

Asynchronous and Load-Balanced Union-Find for Distributed and Parallel Scientific Data Visualization and Analysis

We present a novel distributed union-find algorithm that features asynchronous parallelism and k-d tree based load balancing for scalable visualization and analysis of scientific data. Applications of union-find include level set extraction and critical point tracking, but distributed union-find can suffer from high synchronization costs and imbalanced workloads across parallel processes. In this study, we prove that global synchronizations in existing distributed union-find can be eliminated without changing final results, allowing overlapped communications and computations for scalable processing. We also use a k-d tree decomposition to redistribute inputs, in order to improve workload balancing. We benchmark the scalability of our algorithm with up to 1,024 processes using both synthetic and application data. Here, we demonstrate the use of our algorithm in critical point tracking and super-level set extraction with high-speed imaging experiments and fusion plasma simulations, respectively.

97 MATHEMATICS AND COMPUTING↗

Parallel Algorithms for Computing the Tensor-Train Decomposition

The tensor-train (TT) decomposition expresses a tensor in a data-sparse format used in molecular simulations, high-order correlation functions, and optimization. In this paper, we propose four parallelizable algorithms that compute the TT format from various tensor inputs: (1) Parallel-TTSVD for traditional format, (2) PSTT and its variants for streaming data, (3) Tucker2TT for Tucker format, and (4) TT-fADI for solutions of Sylvester tensor equations. We provide theoretical guarantees of accuracy, parallelization methods, scaling analysis, and numerical results. For example, for a d-dimension tensor in $\mathbb{R}$ $n\times∙∙∙$$\times$$n$ a two-sided sketching algorithm PSTT2 is shown to have a memory complexity of $O(n^{[d/2]})$, improving upon $O(n^{d—1})$ from previous algorithms.

97 MATHEMATICS AND COMPUTING↗

SPARTA: High-Level Synthesis of Parallel Multi-Threaded Accelerators

This article presents a methodology for the Synthesis of PARallel multi-Threaded Accelerators (SPARTA) from OpenMP annotated C/C++ specifications. SPARTA extends an open-source HLS tool, enabling the generation of accelerators that provide latency tolerance for irregular memory accesses through multithreading, support fine-grained memory-level parallelism through a hot-potato deflection-based network-on-chip (NoC), support synchronization constructs, and can instantiate memory-side caches. Our approach is based on a custom runtime OpenMP library, providing flexibility and extensibility. Experimental results show high scalability when synthesizing irregular graph kernels. The accelerators generated with our approach are, on average, 2.29x faster than state-of-the-art HLS methodologies.

Design automation↗

Designing a parallel Feel-the-Way clustering algorithm on HPC systems

This paper introduces a new parallel clustering algorithm, named Feel-the-Way clustering algorithm, that provides better or equivalent convergence rate than the traditional clustering methods by optimizing the synchronization and communication costs. Our algorithm design centers on how to optimize three factors simultaneously: reduced synchronizations, improved convergence rate, and retained same or comparable optimization cost. To compare the optimization cost, we use the Sum of Square Error (SSE) cost as the metric, which is the sum of the square distance between each data point and its assigned clusters. Compared with the traditional MPI k-means algorithm, the new Feel-the-Way algorithm requires less communications among participating processes. As for the convergence rate, the new algorithm requires fewer number of iterations to converge. As for the optimization cost, it obtains the SSE costs that are close to the k-means algorithm. In the paper, we first design the full-step Feel-the-Way k-means clustering algorithm that can significantly reduce the number of iterations that are required by the original k-means clustering method. Next, we improve the performance of the full-step algorithm by adopting an optimized sampling-based approach, named reassignment-history-aware sampling. Our experimental results show that the optimized sampling-based Feel-the-Way method is significantly faster than the widely used k-means clustering method, and can provide comparable optimization costs. More extensive experiments with several synthetic datasets and real-world datasets (e.g., MNIST, CIFAR-10, ENRON, and PLACES-2) show that the new parallel algorithm can outperform the open source MPI k-means library by up to 110% on a high-performance computing system using 4,096 CPU cores. In addition, the new algorithm can take up to 51% fewer iterations to converge than the k-means clustering algorithm.

97 MATHEMATICS AND COMPUTING↗

Real-Time Adaptive Lossless Hyperspectral Image Compression using CCSDS on Parallel GPGPU and Multicore Processor Systems

The proposed CCSDS (Consultative Committee for Space Data Systems) Lossless Hyperspectral Image Compression Algorithm was designed to facilitate a fast hardware implementation. This paper analyses that algorithm with regard to available parallelism and describes fast parallel implementations in software for GPGPU and Multicore CPU architectures. We show that careful software implementation, using hardware acceleration in the form of GPGPUs or even just multicore processors, can exceed the performance of existing hardware and software implementations by up to 11x and break the real-time barrier for the first time for a typical test application.

realtime↗

Stage-by-Stage and Parallel Flow Path Compressor Modeling for a Variable Cycle Engine

This paper covers the development of stage-by-stage and parallel flow path compressor modeling approaches for a Variable Cycle Engine. The stage-by-stage compressor modeling approach is an extension of a technique for lumped volume dynamics and performance characteristic modeling. It was developed to improve the accuracy of axial compressor dynamics over lumped volume dynamics modeling. The stage-by-stage compressor model presented here is formulated into a parallel flow path model that includes both axial and rotational dynamics. This is done to enable the study of compressor and propulsion system dynamic performance under flow distortion conditions. The approaches utilized here are generic and should be applicable for the modeling of any axial flow compressor design.

Parallel Flow Path modeling↗

Stage-by-Stage and Parallel Flow Path Compressor Modeling for a Variable Cycle Engine, NASA Advanced Air Vehicles Program - Commercial Supersonic Technology Project - AeroServoElasticity

This paper covers the development of stage-by-stage and parallel flow path compressor modeling approaches for a Variable Cycle Engine. The stage-by-stage compressor modeling approach is an extension of a technique for lumped volume dynamics and performance characteristic modeling. It was developed to improve the accuracy of axial compressor dynamics over lumped volume dynamics modeling. The stage-by-stage compressor model presented here is formulated into a parallel flow path model that includes both axial and rotational dynamics. This is done to enable the study of compressor and propulsion system dynamic performance under flow distortion conditions. The approaches utilized here are generic and should be applicable for the modeling of any axial flow compressor design accurate time domain simulations. The objective of this work is as follows. Given the parameters describing the conditions of atmospheric disturbances, and utilizing the derived formulations, directly compute the transfer function poles and zeros describing these disturbances for acoustic velocity, temperature, pressure, and density. Time domain simulations of representative atmospheric turbulence can then be developed by utilizing these computed transfer functions together with the disturbance frequencies of interest.

Compressor Modeling Modeling↗

An Open-Source Parallel EMT Simulation Framework: Preprint

As the integration level of inverter-based resources (IBR) increases, ensuring the reliable operation of the bulk power systems requires the use of electromagnetic transient (EMT) simulation tools to identify and mitigate system-wide stability risks. Conducting EMT studies for large-scale, IBR-rich grids, however, is challenging due to the inherent computational bottleneck caused by the underlying high-fidelity models and required small time steps. This paper introduces ParaEMT: an open-source, generic EMT simulation framework designed to accelerate simulations by leveraging advanced parallel computational technologies, such as high-performance computers. This paper presents a comprehensive exposition of ParaEMT, covering its modeling library, simulation strategy, framework structure, operational procedures, and auxiliary features, alongside its extensible parallel computational architecture. Notably, ParaEMT is a publicly accessible and modularized framework written in Python, thereby facilitating future development and the integration of new models and algorithms. The accuracy and efficiency of ParaEMT are demonstrated by rigorous validations via multiple case studies.

electromagnetic transient simulation↗

OpenGraphGym: A Parallel Reinforcement Learning Framework for Graph Optimization Problems

This paper presents an open-source, parallel AI environment (named OpenGraphGym) to facilitate the application of reinforcement learning (RL) algorithms to address combinatorial graph optimization problems. This environment incorporates a basic deep reinforcement learning method, and several graph embeddings to capture graph features, it also allows users to rapidly plug in and test new RL algorithms and graph embeddings for graph optimization problems. This new open-source RL framework is targeted at achieving both high performance and high quality of the computed graph solutions. This RL framework forms the foundation of several ongoing research directions, including 1) benchmark works on different RL algorithms and embedding methods for classic graph problems; 2) advanced parallel strategies for extreme-scale graph computations, as well as 3) performance evaluation on real-world graph solutions.

Zheng, Weijian↗