Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19

Rasterization with Data-Parallel Primitives

Parallel rasterization can suffer from race conditions during fragment generation, which is traditionally addressed by using specialized hardware accessible via vendor graphics APIs. Unfortunately, graphics APIs are increasingly problematic on high-performance computers, either because they are not provided or because of concerns about dependencies with in situ visualization. In response, we present a hardware-agnostic rasterization algorithm that handles race conditions using only data-parallel primitives (DPPs), enabling efficient rendering on HPC systems without graphics API dependencies and aligning with recent efforts to deliver visualization software with DPPs. Our evaluation consists of three phases: (1) evaluating portability across different CPU and GPU architectures, (2) evaluating competitiveness with a community standard, and (3) evaluating performance across varying workloads and available parallelism. The supporting experiments run on both AMD and NVIDIA GPUs, considering data sets as large as 460 million triangles and 160 million pixels. While performance generally falls short of graphics API baselines, it achieves interactive frame rates on most workloads. As a result, we conclude our approach is a viable solution for rasterization on high-performance computers since our approach is portably performant across different architectures without the need for specialized vendor support.

Buckley, Makani [University of Oregon] (ORCID:0009↗

High-performance data management for whole slide image analysis in digital pathology

When dealing with giga-pixel digital pathology in whole-slide imaging, a notable proportion of data records holds relevance during each analysis operation. For instance, when deploying an image analysis algorithm on whole-slide images (WSI), the computational bottleneck often lies in the input-output (I/O) system. This is particularly notable as patch-level processing introduces a considerable I/O load onto the computer system. However, this data management process could be further paralleled, given the typical independence of patch-level image processes across different patches. This paper details our endeavors in tackling this data access challenge by implementing the Adaptable IO System version 2 (ADIOS2). Our focus has been constructing and releasing a digital pathology-centric pipeline using ADIOS2, which facilitates streamlined data management across WSIs. Additionally, we’ve developed strategies aimed at curtailing data retrieval times. The performance evaluation encompasses two key scenarios: (1) a pure CPU-based image analysis scenario (“CPU scenario”), and (2) a GPU-based deep learning framework scenario (“GPU scenario”). Our findings reveal noteworthy outcomes. Under the CPU scenario, ADIOS2 showcases an impressive two-fold speed-up compared to the brute-force approach. In the GPU scenario, its performance stands on par with the cutting-edge GPU I/O acceleration framework, NVIDIA Magnum IO GPU Direct Storage (GDS). From what we know, this appears to be among the initial instances, if any, of utilizing ADIOS2 within the field of digital pathology. The source code has been made publicly available at https://github.com/hrlblab/adios.

Wang, Xiao↗

Performance Evaluation of Multi-Vendor Grid-Forming Inverters for Grid-Connected Operation Through Hardware Experimentation: Preprint

Existing real-world projects of GFM inverters that operate in parallel to power grids typically are sized between dozens and a few hundreds Megawatt (MW) scale according to a recent NERC GFM inverter white paper. These large systems are often difficult to evaluate prior to deployments because of their large size. The performance of smaller GFM inverters (dozens to a few hundreds MW) that operate parallel with power grids (distribution systems) is even less understood. There is an opportunity to better understand these systems through hardware testing under controlled laboratory conditions. Therefore, this paper presents the functional performance evaluation tests of multiple (three) commercial GFM inverters when they operate in parallel with the grid through hardware experiments. The goal of these tests is to explore the GFM inverters' functionalities and dynamic response when in parallel with power grids to eventually develop universal specifications for GFM inverters. Both steady state (changing the inverter's frequency and voltage droop) and transient (adding step change in grid's frequency/voltage) tests are performed for each GFM inverter with the same testing circuit and testing protocol. The experimental results indicate the bench-marked performance that: 1) the GFM inverters can be dispatched through frequency and voltage droop intercepts to output the target power when paralleled to the grid; 2) the GFM inverters automatically respond to system frequency and voltage events to output the needed power, however, the GFM inverters all show stability issues when absorbing reactive power from the grid.

grid-forming inverters↗

Modified Andronov-Hopf Oscillator-Based Grid-Forming Converter with Emulated Virtual Cable for Enhanced Power Sharing Performance

Nonlinear oscillator-based grid-forming converters offer superior dynamic and steady-state performance, making them an attractive solution for interconnecting renewable resources. This paper proposes a novel modified Andronov-Hopf oscillator to enhance the operating spectrum and facilitate the integration of renewable energy sources. An inner loop controller based on the Lyapunov energy function is implemented to achieve robust stability and performance, while a virtual cable emulation strategy enables seamless parallel operation. Comprehensive modeling and simulation studies validate the effectiveness of the proposed system, demonstrating its capabilities in addressing diverse operating scenarios, including grid faults, renewable energy fluctuations, and parallel operation. The proposed solution exhibits fast transient response, robust stability, and flexible operation, making it a valuable contribution to the field of renewable energy integration. The results of this study can be used to inform the design and implementation of next-generation grid-forming converters, enabling a more sustainable and reliable energy future. Additionally, the proposed system's ability to operate in both grid-connected and islanded modes makes it an ideal candidate for remote and off-grid renewable energy applications. The proposed solution's scalability and modularity also make it suitable for large-scale renewable energy integration. The proposed system is verified through MATLAB/Simulink and PLECS simulations, demonstrating its effectiveness in ensuring robust and efficient operation.

Andronov-Hopf Oscillator (AHO)↗

Hybrid Simulation of Proton Cyclotron Waves Upstream of Mars Generated by Pickup Ion Beam Distribution

Linear instability analysis as well as a corresponding two‐dimensional hybrid simulation are performed to examine the excitation of the proton cyclotron waves observed upstream of Mars. The waves are believed to be excited by the pickup ions produced from the ionization of the Martian hydrogen exosphere. And previous statistical analysis of wave observations suggested that the waves are mostly related to pickup ion beam velocity distributions. While earlier linear instability analysis of pickup ion beam distributions has mainly been focused on the parallel unstable modes, our analysis reveals that the maximum growth rate occurs at very oblique propagation. The corresponding hybrid simulation confirms the linear analysis results and further demonstrates that the pickup ions are scattered toward an isotropic shell velocity distribution by the waves excited. Interestingly, the waves at oblique propagation gradually damp out and the system is eventually dominated by waves of quasi‐parallel propagation.

79 ASTRONOMY AND ASTROPHYSICS↗

New scaling and nuclear structure aspects in heavy-ion fusion reactions

Three new behaviors have been found in comparisons of fusion cross sections for different collision systems. root (1) Replacing the energy E with a scaling one, E scal = (E-V g )/($\sqrt{2}$W g ), is successful for washing out the Coulomb interaction in the spectra of fusion cross sections, where V g and W g are barrier height and width of the single-Gaussian barrier distribution model. (2) In a representation of σE vs the scaling energy, E scal , all data sets display in parallel. Here, the ratio for sigma E from any two fusion systems over the whole range is a constant value. That behavior is also studied in another representation, in which the data sets display as parallel horizontal lines for any heavy-ion fusion system. (3) The constant ratio value is the ratio of parameter products, $R^2_gW_g$, of the two systems; where R g is the barrier radius obtained in the single-Gaussian barrier distribution model. Moreover, when comparing neighboring collision systems at the same E scal , the ratio of sigma is near a constant value within a few percent over the whole range. Thus a quantitative comparison for the fusion enhancement for neighboring systems is developed. The present finding could be beneficial for predicting unmeasured fusion cross sections.

Jiang, C. L. [Argonne National Laboratory (ANL), A↗

Monitoring pipeline integrity of underground gas storage facilities using membrane-based electrochemical sensors

Effective monitoring of internal corrosion risk is crucial to ensuring the safety and longevity of natural gas pipeline infrastructure. While electrochemical sensors are commonly used to assess corrosion rates and corrosion indicators in aqueous fluids, they are rarely used in gas pipelines as these fluids lack the ionic conductivity needed for electrochemical measurements. The inclusion of ion-conductive membranes into electrochemical sensors can extend their functionality into humidified gas streams, providing critical information about emerging corrosion events that are common during withdrawal season in pipeline systems downstream from underground storage facilities. In parallel, new protective films, like those obtained through cold spray coating, are being developed to protect oil and gas pipelines and recover losses in structural integrity due to corrosion damage. Herein, we demonstrate how membrane-based electrochemical sensors (MBES) can be used to monitor fluid corrosivity by examining their response to changes in water content for a wide range of fluid compositions. It was found that MBES readings were highly sensitive to water content changes with membrane conductivity measurements varying from 10 –6 to 10 –1 S cm -1 , and corrosion rate measurements which varied from 10 –7 to 1 mm y -1 . Electron microscopy confirmed that the self-healing characteristics of metal coating films were still active despite their inclusion into an MBES probe. In conclusion, these findings indicate that membrane-based corrosion monitoring can be expanded to monitor coated-pipeline materials and provide early detection of emerging corrosion upsets relevant to underground gas storage facilities.

Electrochemical sensor↗

CHICOX, a heavy-ion detection system for GRETA

The CHICOX detection system is an array of position sensitive parallel plate avalanche counters used for heavy ion detection in in-beam particle-gamma coincidence experiments. CHICOX, an upgrade of the CHICO2 array, is geometrically compatible with the GRETINA and GRETA gamma-ray arrays and provides an angular resolution sufficient to utilize the excellent position resolution of GRETINA/GRETA. CHICOX was successfully commissioned and fielded as an auxiliary detector to GRETINA in a campaign of stable-beam Coulomb excitation experiments at the ATLAS facility of Argonne National Laboratory. We report here on the design and in-beam performance of the CHICOX array.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Load‐Balancing Intense Physics Calculations to Embed Regionalized High‐Resolution Cloud Resolving Models in the E3SM and CESM Climate Models

Abstract We design a new strategy to load‐balance high‐intensity sub‐grid atmospheric physics calculations restricted to a small fraction of a global climate simulation's domain. We show why the current parallel load balancing infrastructure of Community Earth System Model (CESM) and Energy Exascale Earth Model (E3SM) cannot efficiently handle this scenario at large core counts. As an example, we study an unusual configuration of the E3SM Multiscale Modeling Framework (MMF) that embeds a binary mixture of two separate cloud‐resolving model grid structures that is attractive for low cloud feedback studies. Less than a third of the planet uses high‐resolution (MMF‐HR; sub‐km horizontal grid spacing) relative to standard low‐resolution (MMF‐LR) cloud superparameterization elsewhere. To enable MMF runs with Multi‐Domain cloud resolving models (CRMs), our load balancing theory predicts the most efficient computational scale as a function of the high‐intensity work's relative overhead and its fractional coverage. The scheme successfully maximizes model throughput and minimizes model cost relative to precursor infrastructure, effectively by devoting the vast majority of the processor pool to operate on the few high‐intensity (and rate‐limiting) high‐resolution (HR) grid columns. Two examples prove the concept, showing that minor artifacts can be introduced near the HR/low‐resolution CRM grid transition boundary on idealized aquaplanets, but are minimal in operationally relevant real‐geography settings. As intended, within the high (low) resolution area, our Multi‐Domain CRM simulations exhibit cloud fraction and shortwave reflection convergent to standard baseline tests that use globally homogenous MMF‐LR and MMF‐HR. We suggest this approach can open up a range of creative multi‐resolution climate experiments without requiring unduly large allocations of computational resources.

54 ENVIRONMENTAL SCIENCES↗

Multi-dimensional incoherent Thomson scattering system in PHAse Space MApping (PHASMA) facility

A multi-dimensional incoherent Thomson scattering diagnostic system capable of measuring electron temperature anisotropies at the level of the electron velocity distribution function (EVDF) is implemented on the PHAse Space MApping facility to investigate electron energization mechanisms during magnetic reconnection. This system incorporates two injection paths (perpendicular and parallel to the axial magnetic field) and two collection paths, providing four independent EVDF measurements along four velocity space directions. For strongly magnetized electrons, a 3D EVDF comprised of two characteristic electron temperatures perpendicular and parallel to the local magnetic field line is reconstructed from the four measured EVDFs. As a result, validation of isotropic electrons in a single magnetic flux rope and a steady-state helicon plasma is presented.

47 OTHER INSTRUMENTATION↗

Modeling of the ECCD injection effect on the Heliotron J and LHD plasma stability

The aim of the study is to analyze the stability of the energetic particle modes (EPM) and Alfven Eigenmodes (AE) in Helitron J and LHD plasma if the electron cyclotron current drive (ECCD) is applied. Additionally, the analysis is performed using the code FAR3d that solves the reduced MHD equations describing the linear evolution of the poloidal flux and the toroidal component of the vorticity in a full 3D system, coupled with equations of density and parallel velocity moments for the energetic particle (EP) species, including the effect of the acoustic modes. The Landau damping and resonant destabilization effects are added via the closure relation. The simulation results show that the n = 1 EPM and n = 2 global AE (GAE) in Heliotron J plasma can be stabilized if the magnetic shear is enhanced at the plasma periphery by an increase (co-ECCD injection) or decrease (ctr-ECCD injection) of the rotational transform at the magnetic axis ($\rlap{-} \iota_{0}$). In the ctr-ECCD simulations, the EPM/AE growth rate decreases only below a given $\rlap{-} \iota_{0}$, similar to the ECCD intensity threshold observed in the experiments. In addition, ctr-ECCD simulations show an enhancement of the continuum damping. The simulations of the LHD discharges with ctr-ECCD injection indicate the stabilization of the n = 1 EPM, n = 2 toroidal AE (TAE) and n = 3 TAE, caused by an enhancement of the continuum damping in the inner plasma leading to a higher EP β threshold with respect to the co- and no-ECCD simulations.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Study of the Alfven eigenmodes stability in CFQS plasma using a Landau closure model

The aim of this study is to analyze the stability of the Alfven eigenmodes (AE) in the Chinese First Quasi-axisymmetric Stellarator (CFQS). The AE stability is calculated using the code FAR3d that solves the reduced MHD equations to describe the linear evolution of the poloidal flux and the toroidal component of the vorticity in a full 3D system, coupled with equations of density and parallel velocity moment for the energetic particles (EP) species including the effect of the helical couplings and acoustic modes. The Landau damping and resonant destabilization effects are added in the model by a given closure relation. The simulation results indicate the destabilization of n = 1 to 4 AEs by EP during the slowing down process, particularly n = 1 and n = 2 toroidal AEs (TAE), n = 3 elliptical AE (EAE) and n = 4 non circular AE (NAE). If the resonance is caused by EPs with an energy above 17 keV (weakly thermalized EP), n = 2 EAEs and n = 3 NAEs are unstable. On the other hand, EPs with an energy below 17 keV (late thermalization stage) lead to the destabilization of n = 3 and n = 4 TAEs. The simulations for an off-axis NBI injection indicate the further destabilization of n = 2 to 4 AEs although the growth rate of the n = 1 AEs slightly decreases, so no clear optimization trend with respect to the NBI deposition region is identified. In addition, n = 2, 4 helical AE (HAE) are unstable above an EP β threshold. Also, if the thermal β of the simulation increases (higher thermal plasma density) the AE stability of the plasma improves. Overall, the simulations including the effect of the finite Larmor radius and electron-ion Landau damping show the stabilization of the n = 1 to 4 EAE/NAEs as well as a decrease of the growth rate and frequency of the n = 1 to 4 BAE/TAEs.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

High-Fidelity Ion State Detection Using Trap-Integrated Avalanche Photodiodes

Integrated technologies greatly enhance the prospects for practical quantum information processing and sensing devices based on trapped ions. High-speed and high-fidelity ion state readout is critical for any such application. Integrated detectors offer significant advantages for system portability and can also greatly facilitate parallel operations if a separate detector can be incorporated at each ion-trapping location. Here, we demonstrate ion quantum state detection at room temperature utilizing single-photon avalanche diodes (SPADs) integrated directly into the substrate of silicon ion trapping chips. Furthermore, we detect the state of a trapped Sr + ion via fluorescence collection with the SPAD, achieving 99.92(1)% average fidelity in 450 μs, opening the door to the application of integrated state detection to quantum computing and sensing utilizing arrays of trapped ions.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Decentralized Carrier Phase Shifting for Optimal Harmonic Minimization in Asymmetric Parallel-Connected Inverters

This paper presents a carrier phase shifting technique for minimizing the aggregate harmonics in networks of asymmetric parallel-connected inverters for distributed power generation system applications. The proposed technique is: 1) implemented in a decentralized manner, relying only on local voltage and current measurements, and 2) optimal in the sense that it minimizes a cost function representing the carrier-frequency current harmonics. The analysis indicates that the proposed optimal carrier phase shifting technique can enable order-of-magnitude reductions in harmonic power, and also universal improvements compared to symmetric carrier interleaving for asymmetric inverter networks. Moreover, compared to existing methods that require either centralized communication or information exchange between inverters to coordinate carriers, the proposed technique is completely decentralized, which provides important practical benefits for implementation, including improved robustness and reduced cost. The technique is experimentally validated on a network of three single-phase 2-kW inverters and demonstrates a 36.5% reduction in the weighted total harmonic distortion factor of the aggregate inverter current, and the ability to converge to the optimal carrier phase spacing dynamically in less than one line frequency cycle (16.7 ms) in steady state and transient operating conditions.

42 ENGINEERING↗

RingX: Scalable Parallel Attention for Long-Context Learning on HPC

The attention mechanism has become foundational for remarkable AI breakthroughs since the introduction of the Transformer, driving the demand for increasingly longer context to power frontier models such as large-scale reasoning language models and high-resolution image/video generators. However, its quadratic computational and memory complexities present substantial challenges. Current state-of-the-art parallel attention methods, such as ring attention, are widely adopted for long-context training but utilize a point-to-point communication strategy that fails to fully exploit the capabilities of modern HPC network architectures. In this work, we propose ringX, a scalable family of parallel attention methods optimized explicitly for HPC systems. By enhancing workload partitioning, refining communication patterns, and improving load balancing, ringX achieves up to 3.4 × speedup compared to conventional ring attention on the Frontier supercomputer. Optimized for both bi-directional and causal attention mechanisms, ringX demonstrates its effectiveness through training benchmarks of a Vision Transformer (ViT) on a climate dataset and a Generative Pre-Trained Transformer (GPT) model, Llama3 8B. Our method attains an end-to-end training speedup of approximately 1.5 × in both scenarios. To our knowledge, the achieved 38% model FLOPs utilization (MFU) for training Llama3 8B with a 1M-token sequence length on 4,096 GPUs represents one of the highest training efficiencies reported for long-context learning on HPC systems. Our code implementation is available at https://github.com/jqyin/ringX-attention.

Yin, Junqi [ORNL] (ORCID:0000000338435520)↗

Simulating microgalvanic corrosion in alloys using the PRISMS phase-field framework

In this prospective paper, we first review the existing simulation tools to simulate microgalvanic corrosion during free immersion. Then, we describe a recently developed application that employs PRISMS-PF, an open-source, high-performance phase-field modeling framework. The model employed in the application accounts for the electrochemical reaction at the metal/electrolyte interface and ionic migration in the electrolyte to determine the evolution of the corrosion front. We present the implementation details for the application and discuss its features such as super-linear parallel scaling performance for a sufficiently large system. Finally, we demonstrate the capability of the application by simulating corrosion of the matrix phase of an alloy near a secondary phase particle in two and three dimensions.

36 MATERIALS SCIENCE↗

Green's function methods for multiphysics simulations (Final Report)

Green's functions are important tools for analyzing mathematical properties of partial differential equations (PDEs), and for numerically solving PDEs, especially when equations for the same operator but with multiple right hand sides need to be solved simultaneously. This proposal aims at developing efficient and accurate numerical methods for computing Green's functions, which can be used to tackle a challenging question in multiphysics simulation of DOE-mission science problems: how to couple quantum physics with classical physics. The key mathematical difficulty is properly formulate a "boundary condition" for the region described by quantum physics, and conventional approaches often model such boundary conditions in an empirical way. The proposed Green's function methods use the Dirichlet-to-Neumann map to formulate a boundary condition that is non-empirical and can couple the quantum and classical regions in an in principle exact way. The key ingredient of the new methods is to construct the Dirichlet-to-Neumann map in an efficient, accurate and versatile manner. The new methods have provably low complexity and are ideally suited for massively parallel and emerging many-core computational systems.

97 MATHEMATICS AND COMPUTING↗