Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel communication”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Digital Assurance for Grid Reliability in the Era of Large Load Growth

The rapid expansion of large electric loads is reshaping the operational and regulatory landscape of the U.S. electric grid. These facilities are reaching new scales of expansion, now exceeding a gigawatt per site, and their highly sensitive, digitally driven behaviors introduce new reliability risks. Recent grid events, including large load losses following routine transmission disturbances, highlight the consequences of limited ride-through capability, inconsistent protection settings, inadequate modeling, and lack of behind-the-meter visibility. Parallels to earlier integration challenges of new grid technologies suggest that the grid’s existing processes, standards, and interconnection frameworks are no longer adequate for emerging large loads. This brief synthesizes lessons from the evolution of inverter-based resource regulation and applies them to large-load integration. It identifies critical gaps in modeling accuracy, interconnection processes, performance standards, and compliance mechanisms. Technical recommendations emphasize advanced monitoring, improved modeling, coordinated communication protocols, modernized substations, and structured behind-the-meter control schemes. Collectively, these measures provide a roadmap to maintain bulk power system reliability while enabling the continued growth of large, electrified digital infrastructure.

24 - POWER TRANSMISSION AND DISTRIBUTION↗

Scalable FBP decomposition for cone-beam CT reconstruction

Filtered Back-Projection (FBP) is a fundamental compute intense algorithm used in tomographic image reconstruction. Cone-Beam Computed Tomography (CBCT) devices use a cone-shaped X-ray beam, in comparison to the parallel beam used in older CT generations. Distributed image reconstruction of cone-beam datasets typically relies on dividing batches of images into different nodes. This simple input decomposition, however, introduces limits on input/output sizes and scalability.We propose a novel decomposition scheme and reconstruction algorithm for distributed FPB. This scheme enables arbitrarily large input/output sizes, eliminates the redundancy arising in the end-to-end pipeline and improves the scalability by replacing two communication collectives with only one segmented reduction. Finally, we implement the proposed decomposition scheme in a framework that is useful for all current-generation CT devices (7th gen). In our experiments using up to 1024 GPUs, our framework can construct 40963 volumes, for real-world datasets, in under 16 seconds (including I/O).

Chen, Peng↗

Large-Signal Stability Improvement of Parallel Grid-Forming Inverter-Driven Black Start: Preprint

With the rapid increase of inverter-based resources in modern grids, advanced grid-forming (GFM) inverter capa- bilities, such as system restoration and operation under faults, are urgently needed to realize power electronics-dominant grids at scale. One such capability is inverter-driven black start using GFM inverters. This paper analyzes the ability of two recently proposed advanced GFM controls to help GFM inverters sustain system-wide, off-nominal conditions and remain synchronized until they can overcome the momentary overloading as more GFMs join the process without generator sequence coordination or communications and finally stabilize the grid. Through an extensive set of 1,200 full-order electromagnetic transient simu- lations, we evaluate the black-start process while employing the various GFM inverter controls. The results show that the GFM current limiter and primary control have a significant impact on the stability of the system during dynamic operating conditions and thus impact the success of inverter-driven system restoration.

current limit↗

Assessment methods for determining small changes in hearing performance over time

Although the behavioral pure-tone threshold audiogram is considered the gold standard for quantifying hearing loss, assessment of speech understanding, especially in noise, is more relevant to quality of life but is only partly related to the audiogram. Metrics of speech understanding in noise are therefore an attractive target for assessing hearing over time. However, speech-in-noise assessments have more potential sources of variability than pure-tone threshold measures, making it a challenge to obtain results reliable enough to detect small changes in performance. Here, this review examines the benefits and limitations of speech-understanding metrics and their application to longitudinal hearing assessment, and identifies potential sources of variability, including learning effects, differences in item difficulty, and between- and within-individual variations in effort and motivation. We conclude by recommending the integration of non-speech auditory tests, which provide information about aspects of auditory health that have reduced variability and fewer central influences than speech tests, in parallel with the traditional audiogram and speech-based assessments.

60 APPLIED LIFE SCIENCES↗

Medium Voltage Solid State Transformer for Extreme Fast Charging Applications

A modular and scalable solid state transformer (SST) with direct medium voltage (MV) AC connectivity is proposed to enable DC extreme fast charging (XFC) of electric vehicles. Single-phase-modules (SPMs), each consisting of an active-front-end (AFE) stage and an isolated DC-DC stage, are connected in input-series-output-parallel (ISOP) configuration. The modular hardware is co-designed with decentralized control of the DC-DC stages where voltage and power balancing are achieved by each SPM using only its local sensor feedback; a centralized controller (CC) regulates the low voltage (LV) DC bus through the AFE stages without any sensor feedback form the SPMs. The controller architecture contrasts sharply with the prior art for MV AC to LV DC SSTs where high-speed bidirectional communication among SPMs and a CC are required for module-level voltage and power balancing, which severely limits the scalability and practical realization of higher voltage and higher power units. Detailed small-signal analysis and controller design guidelines are developed. Furthermore, a soft start-up strategy is presented. The proposed converter and control structure are validated through simulation and experimental results.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Efficient Routing of Quantum LDPC Codes on Programmable 2D Toric Architectures

Quantum low-density parity-check codes are promising candidates towards scalable fault-tolerant quantum computation. Among these, bivariate bicycle (BB) codes offer superior encoding rates and large code distance compared to surface codes. However, their requirement on long-range stabilizer measurements poses significant challenges for implementation on realistic hardware with limited connectivity, such as superconducting circuit platforms. In this work, we introduce a novel hardware-software co-design that leverages a programmable communication network architecture to address these limitations. Our approach utilizes a 2D toric network of oscillators as a flexible communication fabric linking qubits at each site. Such architecture significantly reduces the number of long-range couplers required from O ( n ) to O (√ n ). Dual-rail qubits, along with native gates including Swap-Wait-Swap gates and beamsplitter SWAPs, ensure that long-range two-qubit gates can be executed with high fidelity and low latency. To further enhance performance, our qubit layout and routing algorithm utilize symmetries of the codes and enable maximum parallelism for long-range two-qubit gates, maintaining a low syndrome extraction cycle duration and scalability over the code length. We perform circuit-level simulation with realistic noise modeling based on experimental hardware parameters, observing an logical error rate per logical qubit per cycle of 3.06% for [[18,4,4]] BB code, 2.6× less than the existing experimental result. These findings provide a practical roadmap and identify key technological advancements needed to achieve low-overhead fault-tolerant quantum computing at scale.

Liu, Kun [Yale Univ., New Haven, CT (United States↗

UPC++ v1.0 Programmer’s Guide, Revision 2020.10.0

UPC++ is a C++11 library that provides Partitioned Global Address Space (PGAS) programming. It is designed for writing parallel programs that run efficiently and scale well on distributed-memory parallel computers. The PGAS model is single program, multiple-data (SPMD), with each separate constituent process having access to local memory as it would in C++. However, PGAS also provides access to a global address space, which is allocated in shared segments that are distributed over the processes. UPC++ provides numerous methods for accessing and using global memory. In UPC++, all operations that access remote memory are explicit, which encourages programmers to be aware of the cost of communication and data movement. Moreover, all remote-memory access operations are by default asynchronous, to enable programmers to write code that scales well even on hundreds of thousands of cores.

97 MATHEMATICS AND COMPUTING↗

UPC++ v1.0 Programmer’s Guide, Revision 2020.3.0

UPC++ is a C++11 library that provides Partitioned Global Address Space (PGAS) programming. It is designed for writing parallel programs that run efficiently and scale well on distributed-memory parallel computers. The PGAS model is single program, multiple-data (SPMD), with each separate constituent process having access to local memory as it would in C++. However, PGAS also provides access to a global address space, which is allocated in shared segments that are distributed over the processes. UPC++ provides numerous methods for accessing and using global memory. In UPC++, all operations that access remote memory are explicit, which encourages programmers to be aware of the cost of communication and data movement. Moreover, all remote-memory access operations are by default asynchronous, to enable programmers to write code that scales well even on hundreds of thousands of cores.

97 MATHEMATICS AND COMPUTING↗

Interactive Web Application for Traffic Simulation Data Management and Visualization

As traffic simulation software becomes more effective for realistically simulating and analyzing traffic dynamics and vehicle interactions on the mesoscopic and microscopic level, the management, dissemination, and collaborative visualization of traffic simulation results produced by individual transportation planners presents a significant challenge. Existing online content management systems have a very limited capability in allowing users to query specific traffic simulation scenarios and geospatially visualize simulation results through shareable and interactive web interfaces. This paper presents a web-based application for promoting the archiving, sharing, and visualization of large-scale traffic simulation outputs. The application is developed to enhance cyber-physical controls, communications, and public education for collaborative transportation planning. Unique features of the web application include: (a) allowing users to upload their new traffic simulation scenarios (parameters and outputs), as well as search existing scenarios using easily accessible interfaces; (b) optimizing simulation output files with heterogeneous data formats and projected coordinate systems for web-based storage and management using a scalable and searchable data/metadata standard; (c) standardizing user-uploaded simulation outputs using web interfaces and data processing libraries with parallel computing capacity; and (d) providing shareable web visual interfaces for visualizing the traffic flow and signal information stored in simulation outputs (e.g., regional traffic patterns and individual vehicle interactions) and visually comparing multiple simulation outputs both spatially and temporally. Furthermore, the paper presents the conceptual design and implementation of this application, and demonstrates the application’s performance for sharing, comparing, and visualizing simulation outputs from VISSIM and SUMO, two commonly used traffic simulation software programs.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Flexible User-Defined Domain Decomposition in Kilometer-Scale E3SM Land Model Simulation

The Energy Exascale Earth System Model (E3SM) Land Model (ELM) has been extended to kilometer-scale (km-ELM) resolutions, enabling high-fidelity simulations of terrestrial processes at 1 km x 1 km grid spacing. In ELM, domain decomposition partitions the computational domain across processors, ensuring efficient parallel execution. Currently, round-robin decomposition is applied, providing a straightforward way to distribute computational workload. As ELM continues evolving at the kilometer-scale (km-scale), particularly with integrating lateral flow modeling, decomposition strategies must also account for the increased workload and data movement. This paper introduces a flexible user-defined domain decomposition framework, allowing users to customize domain partitioning based on application requirements. The impact of different decomposition strategies is evaluated across various applications concerning computation, communication, and I/O. Results demonstrate that while 1D partitioning yields superior I/O performance, k-nearest neighbors (KNN) clustering effectively reduces inter-process communication overhead. This study lays the groundwork for scalable partitioning in large-scale land surface simulations, enhancing next-generation Earth system modeling.

Wang, Dali [ORNL] (ORCID:0000000168065108)↗

ExaSGD: 2022 Kernel Thrust Activities

The Kernel Thrust milestone ADSE22-407 covers the development of device-capable optimization algorithms and solvers technologies required by the ExaSGD project’s software stack in order to solve security-constrained alternating current optimal power flow (SC-ACOPF) problems on emerging exascale architectures. To this extent, in FY22 the main objective of the Kernel Thrust was (i) provide sparse optimization solver that runs efficiently on hardware accelerator devices (i.e., NVIDIA and AMD GPUs) to perform intra-node computations, (ii) strengthen the reliability and increase the performance of the mixed-dense sparse (MDS) solver of HiOp for deployment on the FY22 target architectures, Summit and Crusher, and (iii) increase performance by improving the mathematical algorithm and refining the parallel MPI-based implementation of the coarse-grain parallel solver HiOp-PriDec for capabilities deployment on the FY22 target architectures, Summit and Crusher. This document presents the developments and contributions done by the Kernels Thrust Team in FY22 toward completion of the above-mentioned objectives. These contributions progressed along four main development (sub)thrusts: (1) Design and implementation of a sparse optimization solver for use on hardware accelerators; (2) Improvement of the mathematical algorithm and of the parallel implementation of HiOp-PriDec to ensure readiness and efficient coarse-grain parallelism for FY23 target exascale machine; and (3) Support Software and Application Development Thrusts of the exaSGD project in their deployment of the project’s software stack on AMD- and NVIDIA-based architectures. The development of the sparse optimization solver (thrust 1 above) was new in FY22 and resulted in a new sparse solver in HiOp (available as of version 0.6). The second development thrust was a continuation of the efforts from FY21 and improved the mathematical algorithm and the communication strategy of the HiOp-PriDec solver. The last developement thrust is a large collaborative effort. Namely, the project’s teams from multiple labs (LLNL, PNNL, ORNL, and NREL) performed large-scale demonstration of the ExaSGD software stack, namely the optimization solvers of HiOp interfaced with the modeling front-end ExaGO and the stochastic sampler PowerScenarios. These demonstration efforts solved large-scale instances of the SC-ACOPF challenge problem of medium network sizes (10, 000-bus system) and large number of contingencies on Summit (NVIDIA accelerators) and Crusher (AMD accelerators) systems at ORNL.

97 MATHEMATICS AND COMPUTING↗

HipBone: A performance-portable graphics processing unit-accelerated C++ version of the NekBone benchmark

We present hipBone, an open-source performance-portable proxy application for the Nek5000 (and NekRS) computational fluid dynamics applications. HipBone is a fully GPU-accelerated C++ implementation of the original NekBone CPU proxy application with several novel algorithmic and implementation improvements which optimize its performance on modern fine-grain parallel GPU accelerators. Our optimizations include a conversion to store the degrees of freedom of the problem in assembled form in order to reduce the amount of data moved during the main iteration and a portable implementation of the main Poisson operator kernel. We demonstrate near-roofline performance of the operator kernel on three different modern GPU accelerators from two different vendors. We present a novel algorithm for splitting the application of the Poisson operator on GPUs which aggressively hides MPI communication required for both halo exchange and assembly. Our implementation of nearest-neighbor MPI communication then leverages several different routing algorithms and GPU-Direct RDMA capabilities, when available, which improves scalability of the benchmark. We demonstrate the performance of hipBone on three different clusters housed at Oak Ridge National Laboratory, namely, the Summit supercomputer and the Frontier early-access clusters, Spock and Crusher. Our tests demonstrate both portability across different clusters and very good scaling efficiency, especially on large problems.

Computer Science↗

Scaling Out a Combinatorial Algorithm for Discovering Carcinogenic Gene Combinations to Thousands of GPUs

Cancer is a leading cause of death in the US, second only to heart disease. It is primarily a result of a combination of an estimated two-nine genetic mutations (multi-hit combinations). Although a body of research has identified hundreds of cancer-causing genetic mutations, we don’t know the specific combination of mutations responsible for specific instances of cancer for most cancer types. An approximate algorithm for solving the weighted set cover problem was previously adapted to identify combinations of genes with mutations that may be responsible for individual instances of cancer. However, the algorithm’s computational requirement scales exponentially with the number of genes, making it impractical for identifying more than three-hit combinations, even after the algorithm was parallelized and scaled up to a V100 GPU. Since most cancers have been estimated to require more than three hits, we scaled out the algorithm to identify combinations of four or more hits using 1000 nodes (6000 V100 GPUs with ≈48×106 processing cores) on the Summit supercomputer at Oak Ridge National Laboratory. Efficiently scaling out the algorithm required a series of algorithmic innovations and optimizations for balancing an exponentially divergent workload across processors and for minimizing memory latency and inter-node communication. We achieved an average strong scaling efficiency of 90.14% (80.96%–97.96% for 200 to 1000 nodes), compared to a 100 node run, with 84.18% scaling efficiency for 1000 nodes. With experimental validation, the multi-hit combinations identified here could provide further insight into the etiology of different cancer subtypes and provide a rational basis for targeted combination therapy.

Dash, Sajal↗

Review of hybrid HVDC systems combining line communicated converter and voltage source converter

Hybrid HVDC system, which consists of the advantages of line commutated converter (LCC) and voltage source converter (VSC), is an emerging power transmission system. This paper presents a review of four LCC-VSC hybrid HVDC topologies. The first topology is the pole-hybrid HVDC system, in which the LCC and VSC form the positive- and negative-pole respectively. The second one is the terminal-hybrid HVDC system; in this topology one terminal adopts LCC and the other terminal adopts VSC. The series converter-hybrid HVDC system is the third topology wherein each terminal is formed by LCC and VSC in series. The fourth hybrid topology under consideration is a parallel converter-hybrid HVDC system with LCC and VSC connected in parallel in each terminal. Here, the main contribution of this paper is a comprehensive analysis and comparison of the four mentioned hybrid topologies in terms of PQ operating zone, power flow reversal method, and DC fault ride-through strategy.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

FLYING SERVING: On-the-Fly Parallelism Switching for Large Language Model Serving

Production LLM serving must simultaneously deliver high throughput, low latency, and sufficient context capacity under non-stationary traffic and mixed request requirements. Data parallelism (DP) maximizes throughput by running independent replicas, while tensor parallelism (TP) reduces per-request latency and pools memory for long-context inference. However, existing serving stacks typically commit to a static parallelism configuration at deployment; adapting to bursts, priorities, or long-context requests is often disruptive and slow. We present Flying Serving, a vLLM-based system that enables online DP-TP switching without restarting engine workers. Flying Serving makes reconfiguration practical by virtualizing the state that would otherwise force data movement: (i) a zero-copy Model Weights Manager that exposes TP shard views on demand, (ii) a KV Cache Adaptor that preserves request KV state across DP/TP layouts, (iii) an eagerly initialized Communicator Pool to amortize collective setup, and (iv) a deadlock-free scheduler that coordinates safe transitions under execution skew. Across three popular LLMs and realistic serving scenarios, Flying Serving improves performance by up to 4.79 × under high load and 3.47 × under low load while supporting latency- and memory-driven requests.

Gao, Shouwei [ORNL]↗

Massively scalable Kerr comb-driven silicon photonic link

Abstract The growth of computing needs for artificial intelligence and machine learning is critically challenging data communications in today’s data-centre systems. Data movement, dominated by energy costs and limited ‘chip-escape’ bandwidth densities, is perhaps the singular factor determining the scalability of future systems. Using light to send information between compute nodes in such systems can dramatically increase the available bandwidth while simultaneously decreasing energy consumption. Through wavelength-division multiplexing with chip-based microresonator Kerr frequency combs, independent information channels can be encoded onto many distinct colours of light in the same optical fibre for massively parallel data transmission with low energy. Although previous high-bandwidth demonstrations have relied on benchtop equipment for filtering and modulating Kerr comb wavelength channels, data-centre interconnects require a compact on-chip form factor for these operations. Here we demonstrate a massively scalable chip-based silicon photonic data link using a Kerr comb source enabled by a new link architecture and experimentally show aggregate single-fibre data transmission of 512 Gb s −1 across 32 independent wavelength channels. The demonstrated architecture is fundamentally scalable to hundreds of wavelength channels, enabling massively parallel terabit-scale optical interconnects for future green hyperscale data centres.

Rizzo, Anthony (ORCID:000000034752797X)↗

Design Methodologies for Integrated Quantum Frequency Processors

We report frequency-encoded quantum information offers intriguing opportunities for quantum communications and networking, with the quantum frequency processor paradigm—based on electro-optic phase modulators and Fourier-transform pulse shapers—providing a path for scalable construction of quantum gates. Yet all experimental demonstrations to date have relied on discrete fiber-optic components that occupy significant physical space and impart appreciable loss. In this article, we introduce a model for the design of quantum frequency processors comprising microring resonator-based pulse shapers and integrated phase modulators. We estimate the performance of single and parallel frequency-bin Hadamard gates, finding high fidelity values that extend to frequency bins with relatively wide bandwidths. By incorporating multi-order filter designs as well, we explore the limits of tight frequency spacings, a regime extremely difficult to obtain in bulk optics. Overall, our model is general, simple to use, and extendable to other material platforms, providing a much-needed design tool for future frequency processors in integrated photonics.

97 MATHEMATICS AND COMPUTING↗

A Two-Stage Decomposition Approach for AC Optimal Power Flow

The alternating current optimal power flow (AC-OPF) problem is critical to power system operations and planning, but it is generally hard to solve due to its nonconvex and large-scale nature. Furthermore, this paper proposes a scalable decomposition approach in which the power network is decomposed into a master network and a number of subnetworks, where each network has its own AC-OPF subproblem. This formulates a two-stage optimization problem and requires only a small amount of communication between the master and subnetworks. The key contribution is a smoothing technique that renders the response of a subnetwork differentiable with respect to the input from the master problem, utilizing properties of the barrier problem formulation that naturally arises when subproblems are solved by a primal-dual interior-point algorithm. Consequently, existing efficient nonlinear programming solvers can be used for both the master problem and the subproblems. The advantage of this framework is that speedup can be obtained by processing the subnetworks in parallel, and it has convergence guarantees under reasonable assumptions. The formulation is readily extended to instances with stochastic subnetwork loads. Numerical results show favorable performance and illustrate the scalability of the algorithm which is able to solve instances with more than 11 million buses.

24 POWER TRANSMISSION AND DISTRIBUTION↗