Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallelization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24

Parallel measurement of transcriptomes and proteomes from same single cells using nanodroplet splitting

Single-cell multiomics provides comprehensive insights into gene regulatory networks, cellular diversity, and temporal dynamics. While tools for co-profiling single-cell genomes, transcriptomes, and epigenomes are available, accessing proteomes in parallel is more challenging. We developed nanoSPLITS (nanodroplet SPlitting for Linked-multimodal Investigations of Trace Samples), an integrated platform that enables global profiling of the transcriptome and proteome from same single cells using RNA sequencing and mass spectrometry-based proteomics, respectively. nanoSPLITS can precisely quantify over 5000 genes, 2000 proteins, and 140 phosphopeptides per single cell and identify candidate cell markers from these modalities. By exploring Cdk1-mediated cell cycle arrest, we demonstrate how nanoSPLITS single-cell multiomics can provide comprehensive cellular characterization with insights into covarying protein/gene clusters, unique phosphorylation events, and mitotic pathways.

59 BASIC BIOLOGICAL SCIENCES↗

Parallel Algebraic Multigrid for Fusion and Higher-Order PDEs

Multigrid methods play a key role in large-scale scientific simulation because they are among the fastest and most scalable approaches for solving the underlying sparse linear systems of equations that arise from a wide array of Partial Differential Equation (PDE) discretizations. Algebraic multigrid (AMG) is a special type of multigrid method that depends only on the description of the linear system, giving it better portability and broader applicability than geometric multigrid, as it requires no explicit knowledge of the problem geometry. Even though these methods are widely used today, there are still applications where further development is needed. In this report, we focus on PDEs with higher-order terms (e.g., fourth order), concentrating on a PDE that arises in tokamak edge plasma simulations (a tokamak is a machine that confines a plasma using magnetic fields and is believed to be the leading plasma confinement concept for future fusion power plants). General multigrid relaxes a linear system on coarser grids and reverses this process with interpolation, but standard AMG methods struggle with the aforementioned higher-order PDEs. We investigate cyclic coarsening and interpolation heuristics, as well as new iterative approximation methods of refining the solution at each grid to improve the existing multigrid approach. To this end, we ensure that these techniques are transferable to a parallelized setting with LLNL’s supercomputers.

97 MATHEMATICS AND COMPUTING↗

Intelligent Partitioning based Fully Parallel AC Security-Constrained Optimal Power Flow

Today’s power grid is becoming more diverse and integrated with high-level distributed energy resources and smart control technologies that is creating a new set of grid management challenges in terms of large-scale, nonlinear, and non-convex problem modeling, complex and time-consuming computation, as well as difficult uncertainty handling. This project focused on solving a challenging multi-period security-constrained generation scheduling problem, which is of great importance for maximizing the social welfare of real-time dispatch, day-ahead market, as well as weekly planning of power systems. Our developed software explored parallel optimization algorithms for complex and realistic power system models, and develop fast, efficient, and robust grid optimization solutions on the high-performance computing platform that will enable increased grid economics, flexibility, resilience, as well as energy security in the United States.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Xyce™ Parallel Electronic Simulator Users’ Guide (V.7.7)

This manual describes the use of the Xyce Parallel Electronic Simulator. Xyce has been designed as a SPICE-compatible, high-performance analog circuit simulator, and has been written to support the simulation needs of the Sandia National Laboratories electrical designers.

42 ENGINEERING↗

Enhanced PDV waveform search and analysis method using parallel circular-convolution / cross-correlation for improved dynamic surface velocity extraction [Poster]

Previous work on exhaustive search methodologies for extracting best-match parameters pertaining to dynamic surface quantities from PDV was done by cross-correlating synthetically generated PDV waveforms with observed counterparts using the circular-convolution theorem. This work was further developed into an open-source PDV analysis toolkit called CCPDVANALYSIS which expands upon and enhances the previously tested methods by parallelizing serial algorithmic components and incorporating a comprehensive script library for different flavors of instantaneous frequency functions utilized in generating synthetic PDV waveforms. Results of these enhancements have been shown to markedly decrease execution times of exhaustive search and extraction algorithms and produce improved velocity recoveries for low-velocity and dynamically varying velocity signals. The CCPDVANALYSIS script library demonstrates an advanced method for extracting velocities from low-velocity and non-constant velocity signals further extending and improving the methods beyond capabilities of traditional frequency domain tools.

97 MATHEMATICS AND COMPUTING↗

Data-flow parallelism for high-energy and nuclear physics computing frameworks

The processing tasks of a scientific workflow in high-energy and nuclear physics (HENP) can typically be represented as a directed acyclic graph formed according to the data flow—i.e. the data dependencies among algorithms executed as part of the workflow. With this representation, an HENP computing framework can optimally execute a workflow, exploiting the parallelism inherent among independent tasks. Despite such a natural description of a workflow, most HENP frameworks do not make use of technologies that provide concurrent execution of graph-based tasking structures. In this session, we describe Fermilab efforts to adopt a graph-based technology (specifically Intel’s oneTBB flow graph) for meeting the framework needs of its experiments, notably DUNE. After introducing the physics DUNE intends to explore, we will show that all common processing idioms supported by current HENP frameworks can naturally be supported by oneTBB’s data-flow technology, optimally leveraging the concurrent capabilities of the machine. In addition, we discuss collaborative efforts between Fermilab and the Intel oneTBB development team, who is considering improvements to the flow-graph technology to better support HENP use cases.

43 PARTICLE ACCELERATORS↗

Data-flow parallelism for high-energy and nuclear physics frameworks

The processing tasks of an event-processing workflow in high-energy and nuclear physics (HENP) can typically be represented as a directed acyclic graph formed according to the data flow—i.e. the data dependencies among algorithms executed as part of the workflow. With this representation, an HENP framework can optimally execute a workflow, exploiting the parallelism inherent among independent tasks. Despite such a natural description of a workflow, most HENP frameworks do not make use of technologies that provide concurrent execution of graph-based tasking structures. In this talk, we describe Fermilab efforts to adopt a graph-based technology (specifically Intel’s oneTBB flow graph) for meeting the framework needs of its experiments, notably DUNE. Building on the Meld project as presented at CHEP2023, we demonstrate that all common processing idioms supported by current frameworks can naturally be supported by oneTBB’s data-flow technology, optimally leveraging the concurrent capabilities of the machine. In addition, we discuss collaborative efforts between Fermilab and the Intel oneTBB development team, who is considering improvements to the flow-graph technology to better support HENP use cases.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Cable-Driven Parallel Robot (CDPR) for Panelized Envelope Retrofits: Feasible Workspace Analysis

Recent decades have seen remarkable progress in the field of robotic-assisted construction. Cable-driven parallel robots (CDPRs) emerge as promising tools for automating construction processes, due to their advantageous features such as scalability, reconfigurability, compact design, and high payload-to-weight ratio. This paper uses a simple static model to determine the feasibility of a CDPR for overclad panel installation in building envelope retrofits. Given that the building facade needs to be a subset of the CDPR’s wrench-feasible workspace, we focus on the sensitivity of the workspace concerning various cable arrangements and CDPR frame sizes (e.g., height and width extensions). Our analysis indicates that no cable arrangement satisfies the requirement of complete facade coverage and avoids cable-to-panel collisions. Thus, frame extension is needed to enhance coverage. However, in densely populated areas where width extension is limited by space constraints, height extension alone is insufficient to guarantee full facade coverage. This paper pioneers the investigation of CDPRs for panelized envelope retrofits, showcasing their advantages and limitations and paving the way for further research and development.

Liu, Yifang↗

Dynamic Analysis of a Six-Cable Parallel Robot for Automated Panelized Building Retrofits

Cable-Driven Parallel Robots (CDPRs) are highly suitable for automated panelized building retrofits, thanks to their compact footprint and high payload-to-weight ratio. A common CDPR configuration featuring eight cables, where the anchors form a rectangular prism in front of the building facade, offers a large wrench feasible workspace and good control versatility. However, installing upper anchor points requires additional support structures, such as towers or beams, increasing setup complexity and posing logistical challenges in construction settings. To mitigate these challenges, we propose a six-cable CDPR model specifically designed for automated panelized building retrofits. Although the feasible workspace is limited, our analysis shows that the proposed CDPR adequately covers the critical areas required for panel installation. To validate that the six-cable system can effectively transport the end effector to the desired installation pose, we calculated optimal trajectories based on a constrained dynamic model. The simulation results of the six-cable CDPR demonstrate promising potential for automated panelized building retrofits, effectively balancing simplicity, cost-effectiveness, and functionality.

Liu, Yifang [ORNL] (ORCID:0000000190817417)↗

Viskores: Integrating Parallel Scientific Visualization Research into Applications

Viskores is a scientific visualization library that is the primary deployment of such algorithms to the parallel accelerated processors of modern DOE supercomputers. In this paper, we review the capabilities provided by Viskores and how these capabilities are leveraged by other software in the high-performance computing ecosystem. We discuss the Viskores data representation and pay particular attention to array management. Through this array management we describe how data is adapted between Viskores and other software along with strategies for converting dynamic, polymorphic objects to static representations better suited to GPU processing. We conclude with several examples of Viskores integrating with high-performance software that is used in production today.

Moreland, Ken [ORNL] (ORCID:0000000270513288)↗

CLUE: A Fast Parallel Clustering Algorithm for High Granularity Calorimeters in High-Energy Physics

One of the challenges of high granularity calorimeters, such as that to be built to cover the endcap region in the CMS Phase-2 Upgrade for HL-LHC, is that the large number of channels causes a surge in the computing load when clustering numerous digitized energy deposits (hits) in the reconstruction stage. In this article, we propose a fast and fully parallelizable density-based clustering algorithm, optimized for high-occupancy scenarios, where the number of clusters is much larger than the average number of hits in a cluster. The algorithm uses a grid spatial index for fast querying of neighbors and its timing scales linearly with the number of hits within the range considered. We also show a comparison of the performance on CPU and GPU implementations, demonstrating the power of algorithmic parallelization in the coming era of heterogeneous computing in high-energy physics.

Rovere, Marco↗

Parallel Broadband Femtosecond Reflection Spectroscopy at a Soft X-Ray Free-Electron Laser

X-ray absorption spectroscopy (XAS) and the directly linked X-ray reflectivity near absorption edges yield a wealth of specific information on the electronic structure around the resonantly addressed element. Observing the dynamic response of complex materials to optical excitations in pump–probe experiments requires high sensitivity to small changes in the spectra which in turn necessitates the brilliance of free electron laser (FEL) pulses. However, due to the fluctuating spectral content of pulses generated by self-amplified spontaneous emission (SASE), FEL experiments often struggle to reach the full sensitivity and time-resolution that FELs can in principle enable. Here, we implement a setup which solves two common challenges in this type of spectroscopy using FELs: First, we achieve a high spectral resolution by using a spectrometer downstream of the sample instead of a monochromator upstream of the sample. Thus, the full FEL bandwidth contributes to the measurement at the same time, and the FEL pulse duration is not elongated by a monochromator. Second, the FEL beam is divided into identical copies by a transmission grating beam splitter so that two spectra from separate spots on the sample (or from the sample and known reference) can be recorded in-parallel with the same spectrometer, enabling a spectrally resolved intensity normalization of pulse fluctuations in pump–probe scenarios. We analyze the capabilities of this setup around the oxygen K- and nickel L-edges recorded with third harmonic radiation of the free electron laser in Hamburg (FLASH), demonstrating the capability for pump–probe measurements with sensitivity to reflectivity changes on the per mill level.

36 MATERIALS SCIENCE↗

Dynamic Modeling, Trajectory Optimization, and Linear Control of Cable-Driven Parallel Robots for Automated Panelized Building Retrofits

The construction industry faces a growing need for automation to reduce costs, improve accuracy and productivity, and address labor shortages. One area that stands to benefit significantly from automation is panelized prefabricated building envelope retrofits, which can improve a building’s energy efficiency in heating and cooling interior spaces. In this paper, we propose using cable-driven parallel robots (CDPRs), which can effectively lift and handle large objects, to install these panels. However, implementing CDPRs presents significant challenges because of their nonlinear dynamics, complex trajectory planning, and precise control requirements. To tackle these challenges, this work focuses on a new application of established control and trajectory optimization theories in a CDPR simulation of a building envelope retrofit under real-world conditions. We first model the dynamics of CDPRs, highlighting the critical role of damping in system behavior. Building on this dynamic model, we formulate a trajectory optimization problem to generate feasible and efficient motion plans for the robot under operational and environmental constraints. Given the high precision required in the construction industry, accurately tracking the optimized trajectory is essential. However, challenges such as partial observability and external vibrations complicate this task. To address these issues, a Linear Quadratic Gaussian control framework is applied, enabling the robot to track the optimized trajectories with precision. Simulation results show that the proposed controller enables precise end effector positioning with errors under 4 mm, even in the presence of external wind disturbances. Through comprehensive simulations, our approach allows for an in-depth exploration of the system’s nonlinear dynamics, trajectory optimization, and control strategies under controlled yet highly realistic conditions. The results demonstrate the feasibility of CDPRs for automating panel installation and provide insights into their practical deployment.

CDPR↗

Experimental Investigation of Low-Frequency Distributed Acoustic Sensor Responses to Two Parallel Propagating Fractures

Low-frequency distributed acoustic sensing (LF-DAS) is a diagnostic tool for hydraulic fracture propagation with far-field monitoring using fiber optic sensors. LF-DAS senses strain rate variation caused by stress field change due to fracture propagation. Fiber optic sensors are installed in the monitoring wells in the vicinity of a fractured well. From the strain responses, fracture propagation can be evaluated. To understand subsurface conditions with multiple propagating fractures, a laboratory-scale hydraulic fracture experiment was performed simulating the LF-DAS response to fracture propagation with embedded distributed optical fiber strain sensors under these conditions. The experiment was performed using a transparent cube of epoxy with two parallel radial initial flaws centered in the cube. Fluid was injected into the sample to generate fractures along the initial flaws. The experiment used distributed high-definition fiber optic strain sensors with tight spatial resolutions. The sensors were embedded at two different locations on opposite sides of the initial flaws, serving as observation/monitoring locations. We also employed finite element modeling to numerically solve the linear elastic equations of equilibrium continuity and stress–strain relationships. The measured strains from the experiment were compared to simulation results from the finite element model. The experimentally derived strain and strain-rate waterfall plots from this study show the responses to both fractures propagating, while the fracture at the lower position took most of the fluid during the experiment. Interestingly, a fracture first began propagating from the upper flaw of the two flaws, but once the lower fracture was initiated, it grew much faster than the upper fracture. Both fibers were intercepted by the lower fracture, further verifying the strain signature as a fracture is approaching and intersecting an offset fiber.

Chemistry↗

Role of Parallel Solenoidal Electric Field on Energy Conversion in 2.5D Decaying Turbulence with a Guide Magnetic Field

We perform 2.5D particle-in-cell simulations of decaying turbulence in the presence of a guide (out-of-plane) background magnetic field. The fluctuating magnetic field initially consists of Fourier modes at low wavenumbers (long wavelengths). With time, the electromagnetic energy is converted to plasma kinetic energy (bulk flow+thermal energy) at the rate per unit volume of J· E for current density J and electric field E. Such decaying turbulence is well known to evolve toward a state with strongly intermittent plasma current. Here we decompose the electric field into components that are irrotational, E ir , and solenoidal (divergence-free), E so . E ir is associated with charge separation, and J · E ir is a rate of energy transfer between ions and electrons with little net change in plasma kinetic energy. Therefore, the net rate of conversion of electromagnetic energy to plasma kinetic energy is strongly dominated by J · E so , and for a strong guide magnetic field, this mainly involves the component E so,∥ parallel to the total magnetic field B. We examine various indicators of the spatial distribution of the energy transfer rate J ∥ · E so,∥ , which relates to magnetic reconnection, the best of which are (1) the ratio of the out-of-plane electric field to the in-plane magnetic field, (2) the out-of-plane component of the nonideal electric field, and (3) the magnitude of the estimate of current helicity.

79 ASTRONOMY AND ASTROPHYSICS↗

Measurement of Energy Reduction of Inertial Alfvén Waves Propagating through Parallel Gradients in the Alfvén Speed

We have studied the propagation of inertial Alfvén waves through parallel gradients in the Alfvén speed using the Large Plasma Device at the University of California, Los Angeles. The reflection and transmission of Alfvén waves through inhomogeneities in the background plasma are important for understanding wave propagation, turbulence, and heating in space, laboratory, and astrophysical plasmas. Here we present inertial Alfvén waves under conditions relevant to solar flares and the solar corona. We find that the transmission of the inertial Alfvén waves is reduced as the sharpness of the gradient is increased. Any reflected waves were below the detection limit of our experiment, and reflection cannot account for all of the energy not transmitted through the gradient. Our findings indicate that, for both kinetic and inertial Alfvén waves, the controlling parameter for the transmission of the waves through an Alfvén speed gradient is the ratio of the Alfvén wavelength along the gradient divided by the scale length of the gradient. Furthermore, our results suggest that an as-yet-unidentified damping process occurs in the gradient.

79 ASTRONOMY AND ASTROPHYSICS↗