Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallelization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,477 records · Page 82

Enhanced PDV waveform search and analysis method using parallel circular-convolution / cross-correlation for improved dynamic surface velocity extraction [Poster]

Previous work on exhaustive search methodologies for extracting best-match parameters pertaining to dynamic surface quantities from PDV was done by cross-correlating synthetically generated PDV waveforms with observed counterparts using the circular-convolution theorem. This work was further developed into an open-source PDV analysis toolkit called CCPDVANALYSIS which expands upon and enhances the previously tested methods by parallelizing serial algorithmic components and incorporating a comprehensive script library for different flavors of instantaneous frequency functions utilized in generating synthetic PDV waveforms. Results of these enhancements have been shown to markedly decrease execution times of exhaustive search and extraction algorithms and produce improved velocity recoveries for low-velocity and dynamically varying velocity signals. The CCPDVANALYSIS script library demonstrates an advanced method for extracting velocities from low-velocity and non-constant velocity signals further extending and improving the methods beyond capabilities of traditional frequency domain tools.

97 MATHEMATICS AND COMPUTING↗

Data-flow parallelism for high-energy and nuclear physics computing frameworks

The processing tasks of a scientific workflow in high-energy and nuclear physics (HENP) can typically be represented as a directed acyclic graph formed according to the data flow—i.e. the data dependencies among algorithms executed as part of the workflow. With this representation, an HENP computing framework can optimally execute a workflow, exploiting the parallelism inherent among independent tasks. Despite such a natural description of a workflow, most HENP frameworks do not make use of technologies that provide concurrent execution of graph-based tasking structures. In this session, we describe Fermilab efforts to adopt a graph-based technology (specifically Intel’s oneTBB flow graph) for meeting the framework needs of its experiments, notably DUNE. After introducing the physics DUNE intends to explore, we will show that all common processing idioms supported by current HENP frameworks can naturally be supported by oneTBB’s data-flow technology, optimally leveraging the concurrent capabilities of the machine. In addition, we discuss collaborative efforts between Fermilab and the Intel oneTBB development team, who is considering improvements to the flow-graph technology to better support HENP use cases.

43 PARTICLE ACCELERATORS↗

Data-flow parallelism for high-energy and nuclear physics frameworks

The processing tasks of an event-processing workflow in high-energy and nuclear physics (HENP) can typically be represented as a directed acyclic graph formed according to the data flow—i.e. the data dependencies among algorithms executed as part of the workflow. With this representation, an HENP framework can optimally execute a workflow, exploiting the parallelism inherent among independent tasks. Despite such a natural description of a workflow, most HENP frameworks do not make use of technologies that provide concurrent execution of graph-based tasking structures. In this talk, we describe Fermilab efforts to adopt a graph-based technology (specifically Intel’s oneTBB flow graph) for meeting the framework needs of its experiments, notably DUNE. Building on the Meld project as presented at CHEP2023, we demonstrate that all common processing idioms supported by current frameworks can naturally be supported by oneTBB’s data-flow technology, optimally leveraging the concurrent capabilities of the machine. In addition, we discuss collaborative efforts between Fermilab and the Intel oneTBB development team, who is considering improvements to the flow-graph technology to better support HENP use cases.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Cable-Driven Parallel Robot (CDPR) for Panelized Envelope Retrofits: Feasible Workspace Analysis

Recent decades have seen remarkable progress in the field of robotic-assisted construction. Cable-driven parallel robots (CDPRs) emerge as promising tools for automating construction processes, due to their advantageous features such as scalability, reconfigurability, compact design, and high payload-to-weight ratio. This paper uses a simple static model to determine the feasibility of a CDPR for overclad panel installation in building envelope retrofits. Given that the building facade needs to be a subset of the CDPR’s wrench-feasible workspace, we focus on the sensitivity of the workspace concerning various cable arrangements and CDPR frame sizes (e.g., height and width extensions). Our analysis indicates that no cable arrangement satisfies the requirement of complete facade coverage and avoids cable-to-panel collisions. Thus, frame extension is needed to enhance coverage. However, in densely populated areas where width extension is limited by space constraints, height extension alone is insufficient to guarantee full facade coverage. This paper pioneers the investigation of CDPRs for panelized envelope retrofits, showcasing their advantages and limitations and paving the way for further research and development.

Liu, Yifang↗

Dynamic Analysis of a Six-Cable Parallel Robot for Automated Panelized Building Retrofits

Cable-Driven Parallel Robots (CDPRs) are highly suitable for automated panelized building retrofits, thanks to their compact footprint and high payload-to-weight ratio. A common CDPR configuration featuring eight cables, where the anchors form a rectangular prism in front of the building facade, offers a large wrench feasible workspace and good control versatility. However, installing upper anchor points requires additional support structures, such as towers or beams, increasing setup complexity and posing logistical challenges in construction settings. To mitigate these challenges, we propose a six-cable CDPR model specifically designed for automated panelized building retrofits. Although the feasible workspace is limited, our analysis shows that the proposed CDPR adequately covers the critical areas required for panel installation. To validate that the six-cable system can effectively transport the end effector to the desired installation pose, we calculated optimal trajectories based on a constrained dynamic model. The simulation results of the six-cable CDPR demonstrate promising potential for automated panelized building retrofits, effectively balancing simplicity, cost-effectiveness, and functionality.

Liu, Yifang [ORNL] (ORCID:0000000190817417)↗

Viskores: Integrating Parallel Scientific Visualization Research into Applications

Viskores is a scientific visualization library that is the primary deployment of such algorithms to the parallel accelerated processors of modern DOE supercomputers. In this paper, we review the capabilities provided by Viskores and how these capabilities are leveraged by other software in the high-performance computing ecosystem. We discuss the Viskores data representation and pay particular attention to array management. Through this array management we describe how data is adapted between Viskores and other software along with strategies for converting dynamic, polymorphic objects to static representations better suited to GPU processing. We conclude with several examples of Viskores integrating with high-performance software that is used in production today.

Moreland, Ken [ORNL] (ORCID:0000000270513288)↗

CLUE: A Fast Parallel Clustering Algorithm for High Granularity Calorimeters in High-Energy Physics

One of the challenges of high granularity calorimeters, such as that to be built to cover the endcap region in the CMS Phase-2 Upgrade for HL-LHC, is that the large number of channels causes a surge in the computing load when clustering numerous digitized energy deposits (hits) in the reconstruction stage. In this article, we propose a fast and fully parallelizable density-based clustering algorithm, optimized for high-occupancy scenarios, where the number of clusters is much larger than the average number of hits in a cluster. The algorithm uses a grid spatial index for fast querying of neighbors and its timing scales linearly with the number of hits within the range considered. We also show a comparison of the performance on CPU and GPU implementations, demonstrating the power of algorithmic parallelization in the coming era of heterogeneous computing in high-energy physics.

Rovere, Marco↗

Parallel Broadband Femtosecond Reflection Spectroscopy at a Soft X-Ray Free-Electron Laser

X-ray absorption spectroscopy (XAS) and the directly linked X-ray reflectivity near absorption edges yield a wealth of specific information on the electronic structure around the resonantly addressed element. Observing the dynamic response of complex materials to optical excitations in pump–probe experiments requires high sensitivity to small changes in the spectra which in turn necessitates the brilliance of free electron laser (FEL) pulses. However, due to the fluctuating spectral content of pulses generated by self-amplified spontaneous emission (SASE), FEL experiments often struggle to reach the full sensitivity and time-resolution that FELs can in principle enable. Here, we implement a setup which solves two common challenges in this type of spectroscopy using FELs: First, we achieve a high spectral resolution by using a spectrometer downstream of the sample instead of a monochromator upstream of the sample. Thus, the full FEL bandwidth contributes to the measurement at the same time, and the FEL pulse duration is not elongated by a monochromator. Second, the FEL beam is divided into identical copies by a transmission grating beam splitter so that two spectra from separate spots on the sample (or from the sample and known reference) can be recorded in-parallel with the same spectrometer, enabling a spectrally resolved intensity normalization of pulse fluctuations in pump–probe scenarios. We analyze the capabilities of this setup around the oxygen K- and nickel L-edges recorded with third harmonic radiation of the free electron laser in Hamburg (FLASH), demonstrating the capability for pump–probe measurements with sensitivity to reflectivity changes on the per mill level.

36 MATERIALS SCIENCE↗

Dynamic Modeling, Trajectory Optimization, and Linear Control of Cable-Driven Parallel Robots for Automated Panelized Building Retrofits

The construction industry faces a growing need for automation to reduce costs, improve accuracy and productivity, and address labor shortages. One area that stands to benefit significantly from automation is panelized prefabricated building envelope retrofits, which can improve a building’s energy efficiency in heating and cooling interior spaces. In this paper, we propose using cable-driven parallel robots (CDPRs), which can effectively lift and handle large objects, to install these panels. However, implementing CDPRs presents significant challenges because of their nonlinear dynamics, complex trajectory planning, and precise control requirements. To tackle these challenges, this work focuses on a new application of established control and trajectory optimization theories in a CDPR simulation of a building envelope retrofit under real-world conditions. We first model the dynamics of CDPRs, highlighting the critical role of damping in system behavior. Building on this dynamic model, we formulate a trajectory optimization problem to generate feasible and efficient motion plans for the robot under operational and environmental constraints. Given the high precision required in the construction industry, accurately tracking the optimized trajectory is essential. However, challenges such as partial observability and external vibrations complicate this task. To address these issues, a Linear Quadratic Gaussian control framework is applied, enabling the robot to track the optimized trajectories with precision. Simulation results show that the proposed controller enables precise end effector positioning with errors under 4 mm, even in the presence of external wind disturbances. Through comprehensive simulations, our approach allows for an in-depth exploration of the system’s nonlinear dynamics, trajectory optimization, and control strategies under controlled yet highly realistic conditions. The results demonstrate the feasibility of CDPRs for automating panel installation and provide insights into their practical deployment.

CDPR↗

Experimental Investigation of Low-Frequency Distributed Acoustic Sensor Responses to Two Parallel Propagating Fractures

Low-frequency distributed acoustic sensing (LF-DAS) is a diagnostic tool for hydraulic fracture propagation with far-field monitoring using fiber optic sensors. LF-DAS senses strain rate variation caused by stress field change due to fracture propagation. Fiber optic sensors are installed in the monitoring wells in the vicinity of a fractured well. From the strain responses, fracture propagation can be evaluated. To understand subsurface conditions with multiple propagating fractures, a laboratory-scale hydraulic fracture experiment was performed simulating the LF-DAS response to fracture propagation with embedded distributed optical fiber strain sensors under these conditions. The experiment was performed using a transparent cube of epoxy with two parallel radial initial flaws centered in the cube. Fluid was injected into the sample to generate fractures along the initial flaws. The experiment used distributed high-definition fiber optic strain sensors with tight spatial resolutions. The sensors were embedded at two different locations on opposite sides of the initial flaws, serving as observation/monitoring locations. We also employed finite element modeling to numerically solve the linear elastic equations of equilibrium continuity and stress–strain relationships. The measured strains from the experiment were compared to simulation results from the finite element model. The experimentally derived strain and strain-rate waterfall plots from this study show the responses to both fractures propagating, while the fracture at the lower position took most of the fluid during the experiment. Interestingly, a fracture first began propagating from the upper flaw of the two flaws, but once the lower fracture was initiated, it grew much faster than the upper fracture. Both fibers were intercepted by the lower fracture, further verifying the strain signature as a fracture is approaching and intersecting an offset fiber.

Chemistry↗

Role of Parallel Solenoidal Electric Field on Energy Conversion in 2.5D Decaying Turbulence with a Guide Magnetic Field

We perform 2.5D particle-in-cell simulations of decaying turbulence in the presence of a guide (out-of-plane) background magnetic field. The fluctuating magnetic field initially consists of Fourier modes at low wavenumbers (long wavelengths). With time, the electromagnetic energy is converted to plasma kinetic energy (bulk flow+thermal energy) at the rate per unit volume of J· E for current density J and electric field E. Such decaying turbulence is well known to evolve toward a state with strongly intermittent plasma current. Here we decompose the electric field into components that are irrotational, E ir , and solenoidal (divergence-free), E so . E ir is associated with charge separation, and J · E ir is a rate of energy transfer between ions and electrons with little net change in plasma kinetic energy. Therefore, the net rate of conversion of electromagnetic energy to plasma kinetic energy is strongly dominated by J · E so , and for a strong guide magnetic field, this mainly involves the component E so,∥ parallel to the total magnetic field B. We examine various indicators of the spatial distribution of the energy transfer rate J ∥ · E so,∥ , which relates to magnetic reconnection, the best of which are (1) the ratio of the out-of-plane electric field to the in-plane magnetic field, (2) the out-of-plane component of the nonideal electric field, and (3) the magnitude of the estimate of current helicity.

79 ASTRONOMY AND ASTROPHYSICS↗

Measurement of Energy Reduction of Inertial Alfvén Waves Propagating through Parallel Gradients in the Alfvén Speed

We have studied the propagation of inertial Alfvén waves through parallel gradients in the Alfvén speed using the Large Plasma Device at the University of California, Los Angeles. The reflection and transmission of Alfvén waves through inhomogeneities in the background plasma are important for understanding wave propagation, turbulence, and heating in space, laboratory, and astrophysical plasmas. Here we present inertial Alfvén waves under conditions relevant to solar flares and the solar corona. We find that the transmission of the inertial Alfvén waves is reduced as the sharpness of the gradient is increased. Any reflected waves were below the detection limit of our experiment, and reflection cannot account for all of the energy not transmitted through the gradient. Our findings indicate that, for both kinetic and inertial Alfvén waves, the controlling parameter for the transmission of the waves through an Alfvén speed gradient is the ratio of the Alfvén wavelength along the gradient divided by the scale length of the gradient. Furthermore, our results suggest that an as-yet-unidentified damping process occurs in the gradient.

79 ASTRONOMY AND ASTROPHYSICS↗

Hubble Frontier Field Clusters and Their Parallel Fields: Photometric and Photometric Redshift Catalogs

We present a multiband analysis of the six Hubble Frontier Field clusters and their parallel fields, producing catalogs with measurements of source photometry and photometric redshifts. We release these catalogs to the public along with maps of intracluster light and models for the brightest galaxies in each field. This rich data set covers a wavelength range from 0.2 to 8 μm, utilizing data from the Hubble Space Telescope, Keck Observatories, Very Large Telescope array, and Spitzer Space Telescope. We validate our products by injecting into our fields and recovering a population of synthetic objects with similar characteristics to those in real extragalactic surveys. The photometric catalogs contain a total of over 32,000 entries, with 50% completeness at a threshold of mag AB ~ 29.1 for unblended sources and magAB ~ 29 for blended ones, in the IR-weighted detection band. Photometric redshifts were obtained by means of template fitting and have an average outlier fraction of 10.3% and scatter σ = 0.067 when compared to spectroscopic estimates. The software we devised, after being tested in the present work, will be applied to new data sets from ongoing and future surveys.

79 ASTRONOMY AND ASTROPHYSICS↗

PCMS: Parallel Coupler For Multimodel Simulations

This paper presents the Parallel Coupler for Multimodel Simulations (PCMS), a new GPU accelerated generalized coupling framework for coupling simulation codes on leadership class supercomputers. PCMS includes distributed control and field mapping methods for up to five dimensions. For field mapping PCMS can utilize discretization and field information to accommodate physics constraints. PCMS is demonstrated with a coupling of the gyrokinetic microturbulence code XGC with a Monte Carlo neutral transport code DEGAS2 and with a 5D distribution function coupling of an energetic particle transport code (GNET) to a gyrokinetic microturbulence code (GTC). Weak scaling is also demonstrated on up to 2,080 GPUs of Frontier with a weak scaling efficiency of 85%.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

SERGHEI (SERGHEI-SWE) v1.0: a performance-portable high-performance parallel-computing shallow-water solver for hydrology and environmental hydraulics

The Simulation EnviRonment for Geomorphology, Hydrodynamics, and Ecohydrology in Integrated form (SERGHEI) is a multi-dimensional, multi-domain, and multi-physics model framework for environmental and landscape simulation, designed with an outlook towards Earth system modelling. At the core of SERGHEI's innovation is its performance-portable high-performance parallel-computing (HPC) implementation, built from scratch on the Kokkos portability layer, allowing SERGHEI to be deployed, in a performance-portable fashion, in graphics processing unit (GPU)-based heterogeneous systems. In this work, we explore combinations of MPI and Kokkos using OpenMP and CUDA backends. In this contribution, we introduce the SERGHEI model framework and present with detail its first operational module for solving shallow-water equations (SERGHEI-SWE) and its HPC implementation. This module is designed to be applicable to hydrological and environmental problems including flooding and runoff generation, with an outlook towards Earth system modelling. Its applicability is demonstrated by testing several well-known benchmarks and large-scale problems, for which SERGHEI-SWE achieves excellent results for the different types of shallow-water problems. Finally, SERGHEI-SWE scalability and performance portability is demonstrated and evaluated on several TOP500 HPC systems, with very good scaling in the range of over 20 000 CPUs and up to 256 state-of-the art GPUs.

58 GEOSCIENCES↗

Diagnostic development for parallel wave-number measurement of lower hybrid waves in EAST

An eight-channel magnetic probe diagnostic system has been designed and installed adjacent to the 4.6 GHz lower hybrid (LH) grill antenna in the low-field side of the EAST tokamak in order to study the n|| evolution of lower hybrid waves in the first pass from the launcher to the core plasma. The magnetic probes are separated by 6.6 mm, which allows measurement of the dominant parallel refractive index n|| up to n|| =5 for 4.6GHz LH waves. The magnetic probes are designed to be sensitive to the magnetic field component perpendicular to the background magnetic field with a slit on the casing that encloses the probe. The intermediate frequency (IF) stage, which consists of two mixing stages, down-coverts the frequency of the measured wave signals at 4.6 GHz to 20 MHz. A bench test demonstrates the phase stability of the magnetic probe diagnostic system. By evaluating the phase variation of the measured signals along the background magnetic field, the dominant n|| of the LH wave in the scrape-off layer (SOL) has been deduced during the 2019 experimental campaign. In the low density plasma, the measured dominant n|| of the lower hybrid waves is about 2.1, corresponding to the main peak 2.04 of the launched n|| spectrum. The n|| deduced by the least square linear fit method remains near this value in the low density plasma with a high spatial correlation magnitude of 0.9. With an eight-channel probe system a wave-number spectrum has also been deduced, which has a peak near to the measured dominant n||.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗