Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Communication architecture”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Framework for Large-scale Implementation of Wholesale-Retail Transactive Control Mechanism

Transactive energy is a control technique that uses market mechanisms to achieve desired control objectives. Several simulation studies and field demonstrations were carried out in recent years, but all focused on small scale systems and purpose-built simplified models which are not always capable of pointing out all the advantages and shortcomings of transac- tive energy methods. This work describes a co-simulation framework built on a hierarchical control architecture that allows for conducting studies of the impacts of a very large-scale deployment of transactive energy. The hierarchical transactive control architecture adopted in this work helps with alleviating computation and communication burden to facilitate a more effective large scale real-time market operation among device level resources and the system level operators. The co-simulation framework is evaluated an integrated power sys- tem model of unprecedented scale composed of the Western Electricity Coordination Council (WECC) transmission system with tens of thousands of distribution systems deployed with flexible device-level distributed energy resources (DERs) using off-the-shelf simulators.

Transactive energy, market-based controls, DER int↗

Cloud-based implementation of white-box model predictive control for a GEOTABS office building: A field test demonstration

Model predictive control (MPC) has been proven in simulations and pilot case studies to be a superior control strategy for large buildings. MPC can utilize the weather and occupancy schedule forecasts, together with the system model, to predict the future thermal behavior of the building and minimize the overall energy use and maximize thermal comfort. However, these advantages come with the cost of increased modeling effort, computational demands, communication infrastructure, and commissioning efforts. Thus a typical approach is to, often rapidly, simplify the building modeling and MPC optimization problem while paying a price of not reaching the full performance potential. It has been shown that by employing accurate physics-based models, MPC performance can be notably increased closer to its theoretical performance bound. However, implementation of such high-fidelity MPC in real buildings remains a challenge, resulting in a lack of successful field test studies. This work presents the methodology and field test demonstration of a computationally efficient implementation of the white-box MPC in an office building in Belgium. The detailed model of the building is based on first-principle physical equations. The deployment and supervision of MPC operation in a practical setting are supported by an automated cloud-based communication infrastructure. The motivating factor behind the cloud-based architecture is its compatibility with a commercially appealing control as a service concept. The building is equipped with a ground source heat pump (GSHP) and thermally activated building structures (TABS), where the combination of both is also known as GEOTABS. From a control perspective, GEOTABS buildings are particularly challenging systems due to large scale, complex heating, ventilation and air conditioning (HVAC) system, and slow dynamics with time delays. On the other hand, there is an increased potential for energy savings due to the high thermal mass, which acts as thermal storage. The MPC operation is demonstrated during the challenging transient seasons (switching between heating and cooling), and its performance is compared to a traditional rule-based controller (RBC). We provide a proof of concept of real MPC operation for the most difficult seasons with notable GSHP energy use savings equal to 53.5% and thermal comfort improvement by 36.9%. Other MPC applications found in the literature describe tests for only cooling or only heating, and up to now only for a black-box or a grey-box approach.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

VISION: a modular AI assistant for natural human-instrument interaction at scientific user facilities

Scientific user facilities, such as synchrotron beamlines, are equipped with a wide array of hardware and software tools that require a codebase for human-computer-interaction. This often necessitates developers to be involved to establish connection between users/researchers and the complex instrumentation. The advent of generative AI presents an opportunity to bridge this knowledge gap, enabling seamless communication and efficient experimental workflows. Here we present a modular architecture for the Virtual Scientific Companion by assembling multiple AI-enabled cognitive blocks that each scaffolds large language models (LLMs) for a specialized task. With VISION, we performed LLM-based operation on the beamline workstation with low latency and demonstrated the first voice-controlled experiment at an x-ray scattering beamline. The modular and scalable architecture allows for easy adaptation to new instruments and capabilities. Development on natural language-based scientific experimentation is a building block for an impending future where a science exocortex—a synthetic extension to the cognition of scientists—may radically transform scientific practice and discovery.

36 MATERIALS SCIENCE↗

Fundamental Studies of the Vibrational, Electronic, and Photophysical Properties of Tetrapyrrolic Architectures

The ability to capture and utilize light in the near-ultraviolet (NUV), visible and near-infrared (NIR-I and NIR-II) spectral regions (i.e., 320–400, 400–700, 700–1000, 1000–1700 nm) is essential for any solar-energy conversion scheme. Nature employs chlorophylls and bacteriochlorophylls in light-harvesting architectures to absorb light in the blue and red/NIR regions. Accessory pigments (carotenoids, bilins) augment absorption of the (bacterio)chlorophylls in the green region. The harvested energy is funneled to a reaction center protein, where charge separation occurs. Subsequent migration of the electron and the hole stabilizes and stores the energy from light via redox chemistry. The long-term objective of the Bocian/Holten&Kirmaier/Lindsey research program under this DOE grant has been to develop tetrapyrrole-based molecular architectures that absorb sunlight, funnel energy and separate charge with high efficiency. Integral to the program has been iterative cycles of design, synthesis and characterization that provided deep insights into the relationships between chemical composition, electronic structure, and key static and dynamic properties (vibrational, redox, photophysical, energy/charge transfer) of tetrapyrrolic systems. Such architectures included monomers, dyads, larger arrays, and complexes with accessory components. The objective was to develop molecular designs and guiding principles to enhance current and future energy-conversion schemes. Molecular arrays targeted to address one or more fundamental questions concerning light harvesting and energy/charge transfer were constructed from analogues of the naturally occurring hemes, chlorophylls and bacteriochlorophylls. Diverse, tunable synthetic building blocks were prepared that spanned the three respective tetrapyrrole families, which are the porphyrins, chlorins and bacteriochlorins. Thus, the research focused on porphyrins as well as synthetic surrogates for chlorophylls (chlorins, 13 1 -oxophorbines and chlorin-imides) and bacteriochlorophylls (bacteriochlorins, bacterio-13 1 -oxophorbines and bacteriochlorin-imides), generically termed hydroporphyrins. Although the three tetrapyrrole classes (porphyrins, chlorins and bacteriochlorins) absorb light strongly in the violet-blue spectral region, the long-wavelength absorption band typically lies in the green-orange, red, and NIR regions, respectively, with increasing intensity. Understanding the spectra, electronic structure, and energy/charge-transfer properties of such tetrapyrrolic macrocycles is of central importance for the rational design of molecular architectures for solar-energy conversion. Our integrated program of molecular design and synthesis coupled with a variety of spectroscopic, electrochemical, and computational studies have probed from first principles how structural and electronic properties of tetrapyrrolic macrocycles dictate spectral properties as well as the rates of ground-state hole/electron transfer and excited-state energy flow in multicomponent architectures. Individual molecules and multicomponent architectures were designed to test ideas of fundamental importance, often requiring the development of new synthetic methodology. The members of the collaborative team had almost daily discussions by phone and/or e-mail concerning design of molecules, flow of compounds between the labs, planning of physical characterization studies, discussing results and analysis and integrating into design of next generation architectures, and the preparation of manuscripts. Furthermore, students and postdocs in the different labs routinely communicated with one another to facilitate the advancement of the research activities. In short, a highly integrated and collaborative research program was well established among the groups. The research effort involved molecular design and synthesis of synthetic molecular architectures by the Lindsey group integrated with physicochemical and photophysical characterization by the Bocian group and the Holten&Kirmaier group (Figure 2). The Bocian group carried out electrochemical, electron paramagnetic resonance (EPR), resonance Raman (RR), and Fourier-transform infrared (FT-IR) studies, as well as density functional theory (DFT) calculations and the time-dependent extension (TDDFT) to gain insight into excited-state properties. The Holten&Kirmaier group carried out static and time-resolved absorption and fluorescence spectroscopy studies and simulated absorption spectra using molecular orbital (MO) energies from DFT as input to the four-orbital model to complement the TDDFT calculations. The combined measurements provided understanding of the vibrational/electronic properties of the individual molecules and the changes that occur upon incorporation into multicomponent architectures. This information underpinned elucidating the mechanisms and timescales of ground-state hole/electron transfer and excited-state energy and charge transfer.

14 SOLAR ENERGY↗

DoCeph: DPU-Offloaded Messaging in Ceph for Reduced Host CPU Utilization

Ceph is a widely used distributed object store, but its messenger layer imposes substantial CPU overhead on the host. To address this limitation, we propose DoCeph, a DPU-offloaded storage architecture for Ceph that disaggregates the system by offloading the communication-intensive messaging component to the DPU while retaining the storage backend on the host. The DPU efficiently manages communication, using lightweight RPC for metadata operations and DMA for data transfer. Moreover, DoCeph introduces a pipelining technique that overlaps data transmission with buffer preparation, mitigating hardware-imposed transfer size limitations. We implemented DoCeph on a Ceph cluster with NVIDIA BlueField-3 DPUs. Evaluation results indicate that DoCeph cuts host CPU usage by up to 92% while sustaining stable throughput and providing larger performance benefits for object writes over 1 MB.

Park, Kuri [Sogang University]↗

Distributed Grid Control of Flexible Loads and DERs for Optimized Provision of Synthetic Regulating Reserves

Over the course of this project, we have successfully de-risked our distributed microgrid control architecture by tightly integrating its associated control algorithms into a unified software library, installing the software library on several industrial-grade target hardware platforms, and validating the performance of the resulting microgrid controller in a real-life microgrid. Upon completion of the project, we demonstrated that our distributed control architecture is resilient against (i) failures in control devices, (ii) unreliable communication links, (iii) delays in transmitted data, and (iv) imperfect knowledge of the number of (and state of) generation and load assets in the microgrid. In this final report, we present results from all the milestones that were accomplished over the course of this project.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Reconfigurable Network Slicing Orchestration in Network Function Virtualization Compatible Operational Technology Environment

The ongoing transition to Industry 4.0, which is characterized by increased inter-connectivity of cyber-physical systems, requires having time-sensitive, high throughput, and secure transfer of critical data in industrial sites. In this context, network slicing emerges as a critical tool to ensure timely data delivery by provisioning the network resources to cater to specific applications’ requirements and mitigating potential cyber attacks. To address these challenges, this paper aims to tackle two key questions essential for the successful implementation of network slicing in industrial environments. First, it investigates architectural considerations for developing a network infrastructure capable of supporting network slicing functionalities effectively. The proposed approach significantly improves deployment efficiency over traditional manual configurations. Second, it delves into the automated orchestration process, elucidating the steps and components involved in transitioning from a static network management approach to dynamically leverage network function virtualization schemes for creating network slices in ad-hoc manner. The system demonstrates high throughput suitable for production-level solutions and maintains exceptionally low latency, making it ideal for ultra-reliable low-latency communications. Even with increased network demands, the system remains stable, with effective Quality of Service (QoS) management, ensuring reliable performance under varying conditions. The proposed architecture outlines the necessary components, services, and communication protocols required for a production-level orchestrator for network segmentation in SCADA environments.

Rodiles Delgado, Brian G.↗

CommBench: Micro-Benchmarking Hierarchical Networks with Multi-GPU, Multi-NIC Nodes

Modern high-performance computing systems have multiple GPUs and network interface cards (NICs) per node. The resulting network architectures have multilevel hierarchies of subnetworks with different interconnect and software technologies. These systems offer multiple vendor-provided communication capabilities and library implementations (IPC, MPI, NCCL, RCCL, OneCCL) with APIs providing varying levels of performance across the different levels. Understanding this performance is currently difficult because of the wide range of architectures and programming models (CUDA, HIP, OneAPI). We present CommBench, a library with cross-system portability and a high-level API that enables developers to easily build microbenchmarks relevant to their use cases and gain insight into the performance (bandwidth & latency) of multiple implementation libraries on different networks. We demonstrate CommBench with three sets of microbenchmarks that profile the performance of six systems. Our experimental results reveal the effect of multiple NICs on optimizing the bandwidth across nodes and also present the performance characteristics of four available communication libraries within and across nodes of NVIDIA, AMD, and Intel GPU networks.

Hidayetoglu, Mert↗

Toward a persistent event-streaming system for high-performance computing applications

High-performance computing (HPC) applications have traditionally relied on parallel file systems and file transfer services to manage data movement and storage. Alternative approaches have been proposed that use direct communications between application components, trading persistence and fault tolerance for speed. Event-driven architectures, as popularized in enterprise contexts, present a compelling middle ground, avoiding the performance cost and API constraints of parallel file systems while retaining persistence and offering impedance matching between application components. However, adapting streaming frameworks to HPC workloads requires addressing challenges unique to HPC systems. This paper investigates the potential for a streaming framework designed for HPC infrastructures and use cases. We introduce Mofka, a persistent event-streaming framework designed specifically for HPC environments. Mofka combines the capabilities of a traditional streaming service with optimizations tailored to the HPC context, such as support for massively multicore nodes, efficient scaling for large producer-consumer workflows, RDMA-enabled high-performance network communications, specialized network fabrics with multiple links per node, and efficient handling of large scientific data payloads. Built using the Mochi suite of HPC data service components, Mofka provides a lightweight, modular, and high-performance solution for persistent streaming in HPC systems. We present the architecture of Mofka and evaluate its performance against Kafka and Redpanda using benchmarks on diverse platforms, including Argonne's Polaris and Oak Ridge's Frontier supercomputers, showing up to 8× improvement in throughput in some scenarios. We then demonstrate its utility in several real-world applications: a tomographic reconstruction pipeline, a workflow for the discovery of metal-organic frameworks for carbon capture, and the instrumentation of Dask workflows for provenance tracking and performance analysis.

HPC↗

Technical Characterization and Benefit Evaluation of 5G-Enabled Grid Data Transport and Applications

This report summarizes the Year 1 work of Pacific Northwest National Laboratory’s (PNNL’s) 5G Fabricated Resource and Asset Management Encompassment for energy infrastructure (Energy FRAME) project funded by the Department of Energy Office of Science’s Advanced Scientific Computing Research Program. 5G is a breakthrough technology that enables a fully mobile and connected society, and a 5G-enabled digital continuum will be one of the critical foundations for a clean energy economy and grid modernization. In collaboration with PNNL’s Advanced Wireless Communication team and Center for Advanced Technology Evaluation team, the project team has been evaluating the system performance of 5G testbeds in the PNNL 5G Innovation Studio, and has formulated a co-simulation test case of power system transmission, distribution, and communication (T&D&C) networks considering 5G technology and high penetration of distributed energy resources. The methodology developed in the 5G Energy FRAME project can be customized to fit different future grid scenarios to evaluate multiple (dynamic) configurations (computing, sensing, communication, environment) for different stakeholders. In summary, our main technical highlights in project Year 1 are as follows: 1) Technical characterization of 5G standalone architectures, 2) Formulation of co-simulation test case of T&D&C networks embedded with 5G, 3) Initial benefit evaluation of 5G communication platform for grid use cases, and 4) Additional extended discussions on edge computing, artificial intelligence and machine learning, and high-performance computing and cloud computing adoptions. In addition, a collection of system performance data is shared through the publicly available weblink, https://www.pnnl.gov/projects/5g-energy-frame/publications

24 POWER TRANSMISSION AND DISTRIBUTION↗

A distributed voltage inference framework for cyber-physical attacks detection and localization in active distribution grids

The transition to active distribution grids with real-time monitoring and control depends on the proliferation of advanced communication networks and devices. This paradigm shift towards a cyber-physical architecture also introduces new vulnerabilities for adversaries to exploit and launch sophisticated cyber-physical attacks targeting grid observability. Current research highlights the challenges in distinguishing attacks on voltage phasor or nodal injection measurements and isolating multi-source attack locations in a multiphase distribution grid. The attack detection and localization methods in literature face accuracy issues, applications across diverse attack scenarios, or scalability limits. Here, to bridge these gaps, this paper proposes a distributed Voltage Inference framework for real-time detection and localization of cyber-physical attacks, addressing scalability, adaptability, and accuracy challenges in state-of-the-art methods. The proposed methodology leverages the distributed nature of the Voltage Inference framework through a two-step process of prediction and correction, together with a tractable graph partitioning approach, providing a reliable solution to identify compromised measurement sources and facilitate isolation. Extensive testing on IEEE 13 and 123-node distribution feeders underscores the algorithm’s efficacy, enhancing the security and resilience of active distribution grids against evolving cyber threats. Additionally, Hardware-in-the-Loop (HIL) implementation validates the proposed strategy’s practical applicability in real-world scenarios.

active distribution grids↗

MassiveGNN: Efficient Training via Prefetching for Massively Connected Distributed Graphs

Graph Neural Networks (GNN) are indispensable in learning from graph-structured data, yet their rising computational costs, especially on massively connected graphs, pose significant challenges in terms of execution performance. To tackle this, distributed-memory solutions such as partitioning the graph to concurrently train multiple replicas of GNNs are in practice. However, approaches requiring a partitioned graph usually suffer from communication overhead and load imbalance, even under optimal partitioning and communication strategies due to irregularities in the neighborhood minibatch sampling. This paper proposes practical trade-offs for improving the sampling and communication overheads for representation learn- ing on distributed graphs (using popular GraphSAGE architecture) by developing a parameterized prefetch and eviction scheme on top of the state-of-the-art Amazon DistDGL distributed GNN framework, demonstrating about 15–40% improvement in end-to-end training performance on the NERSC Perlmutter supercomputer for various OGB datasets.

Machine Leanring, high performance comptuing, grap↗

A Shift Selection Strategy for Parallel Shift-invert Spectrum Slicing in Symmetric Self-consistent Eigenvalue Computation

The central importance of large-scale eigenvalue problems in scientific computation necessitates the development of massively parallel algorithms for their solution. Recent advances in dense numerical linear algebra have enabled the routine treatment of eigenvalue problems with dimensions on the order of hundreds of thousands on the world’s largest supercomputers. In cases where dense treatments are not feasible, Krylov subspace methods offer an attractive alternative due to the fact that they do not require storage of the problem matrices. However, demonstration of scalability of either of these classes of eigenvalue algorithms on computing architectures capable of expressing massive parallelism is non-trivial due to communication requirements and serial bottlenecks, respectively. In this work, we introduce the SISLICE method: a parallel shift-invert algorithm for the solution of the symmetric self-consistent field (SCF) eigenvalue problem. The SISLICE method drastically reduces the communication requirement of current parallel shift-invert eigenvalue algorithms through various shift selection and migration techniques based on density of states estimation and k-means clustering, respectively. This work demonstrates the robustness and parallel performance of the SISLICE method on a representative set of SCF eigenvalue problems and outlines research directions that will be explored in future work.

97 MATHEMATICS AND COMPUTING↗

Blockchain for Fault-Tolerant Grid Operations Version 2.0

This report explores the potential of distributed ledger technology (DLT) as a transformative tool to enhance fault-tolerant operations in electrical distribution systems. Leveraging DLT's core attributes, including an immutable decentralized ledger, distributed consensus mechanisms, and state replication capabilities, this study focuses on three critical use cases. A central aspect of this research centers on the utilization of a consensus-driven ledger, providing actors within the system, such as distributed resources, with access to a reliable data repository. This empowers these actors to collaborate effectively and make informed decisions, all securely recorded on the blockchain. The first use case concentrates on data configuration, utilizing mathematical criteria---particularly, the chi-squared test for gross error detection---to identify trustworthy sensors for advanced decision-making. Building upon this foundation of trust, the second use case, topology identification, accurately determines circuit breaker states, unveiling the distribution network's topology. Ultimately, the third use case leverages this trust to execute switching actions, reconfiguring feeders and restoring power to disconnected customers after fault events. The concept of trust serves as a cornerstone in this approach, marking a departure from traditional fault location, isolation, and service restoration (FLISR) methods. Additionally, the blockchain-based architecture introduces decentralization, empowering disconnected areas to make autonomous decisions, even when communication with a central control center is disrupted. The primary contributions of this report are twofold: (1) a novel approach for evaluating distribution system voltage areas while preserving data ownership and (2) the implementation of interactions between distribution network areas using the actor model. Unlike the previous sequential approach for evaluating the area connection voltages, which required a radial network topology, this study's area model reduction enables a more versatile approach. The area model reduction addresses issues of prolonged data waiting times and multiple points of failure within the previous approach. Notably, the presented evaluation for the reduced network model area connection reveals a significant increase in the differences in voltage magnitudes. Simulation and evaluation of area agents across four distinct cases elucidate the area-level interaction behavior during a fault event. Simulations demonstrate that the proposed distributed FLISR (DFLISR) approach can successfully restore service to an affected area. Varying message delays and message loss probabilities in each simulation case underscore their impacts on restoration times, ranging from 3 min and 32 s to 6 min and 19 s. In contrast, power is not restored in an area in one of our simulation cases.

24 POWER TRANSMISSION AND DISTRIBUTION↗

A Framework for Neural Network Inference on FPGA-Centric SmartNICs

FPGA-based SmartNICs offer great potential to significantly improve the performance of high-performance computing and warehouse data processing by tightly coupling support for reconfigurable data-intensive computation with cross-node communication, thereby mitigating the von Neumann bottleneck. Existing work, however, has been generally been limited in that it assumes an accelerator model where kernels are offloaded to SmartNICs, but most control tasks are left to the CPUs. This leads to frequent waiting, inferior performance, and scaling challenges. In this work, we propose a new distributive data-centric computing framework, named FCsN, for reconfigurable SmartNIC-based systems. Through a lightweight task circulation execution model and its implementation architecture, FCsN allows the complete detaching of kernel execution, control logic, system scheduling, and network communication to the SmarNICs. This boosts performance by: (i) avoiding the control dependency with CPUs and (ii) supporting streaming kernel execution and network communication at line rate and in a very fine-grained manner. We demonstrate the efficiency and flexibility of FCsN using various types of neural network applications including graph neural networks; as these last are both irregular and data intensive they offer an especially robust demonstration. Evaluations using commonly-used neural network models and graph datasets show that a system with the support of FCsN can achieve, on average, 144 speedups over the MPI-based standard CPU baselines.

Guo, Anqi↗

Topology-Dependent Performance of Free-Space Photonic Quantum Networks Under Noise

Photonic quantum communication enables secure and high-fidelity information transfer beyond classical limits, with direct relevance to emerging quantum networks operating in free-space environments. While physical-layer models of depolarizing noise, Gamma–Gamma turbulence statistics, entanglement swapping, and decoy-state QKD security bounds are individually well established, prior work typically treats these components in isolation or under fixed network assumptions. In this work, we develop a unified topology-aware analytical framework that simultaneously integrates free-space optical link budgets, turbulence-induced visibility degradation, depolarizing qubit noise, multi-hop entanglement cascade dynamics, teleportation fidelity thresholds, CHSH nonlocality certification, and asymptotic decoy-state secret key rate bounds across star, mesh, and ring graph structures. Rather than introducing new physical channel models, we demonstrate that identical physical links exhibit fundamentally different end-to-end performance once embedded within different network topologies. Mesh architectures minimize visibility cascade through hop-count reduction but incur quadratic hardware scaling. Star topologies minimize link count but concentrate noise and synchronization overhead at the hub. Ring configurations offer linear hardware scaling with multiplicative fidelity degradation. The results establish topology as a first-order design parameter in near-term free-space quantum networks operating without full quantum repeater infrastructures. While motivated by distributed multi-agent architectures, the framework applies broadly to terrestrial, airborne, and satellite-assisted photonic quantum communication systems.

QKD↗

xSDK: Building an ecosystem of highly efficient math libraries for exascale

Current efforts to build increasingly powerful computer architectures are opening up new avenues for more complex and higher fidelity simulations coupled with data analytics and learning, leading to new scientific insights and deeper understanding. At one extreme, exascale computers will be much faster than previous computer generations (performing 10 18 operations per second—that is, 1,000 times faster than petascale). To achieve these performance improvements, computer architectures are becoming increasingly complex, with deep memory hierarchies, very high node and core counts, and heterogeneous features such as graphics processing units (GPUs). Such architectural changes impact the full breadth of computing scales, as heterogeneity pervades even current-generation laptops, workstations, and moderate-sized clusters. While emerging advanced architectures provide unprecedented opportunities, they also present significant challenges for developers of scientific applications, such as multiphysics and multiscale codes, who must adapt their software to handle disruptive changes in architectures and new programming models that have not yet stabilized. Developers must consider increasing concurrency while reducing communication and synchronization, and other complexities such as the potential for using mixed precision to leverage the compute power available in low-precision tensor cores. On one hand, developers must implement new scientific capabilities, which in turn increase code complexity. On the other hand, the codes must be ported to new architectures, requiring the inclusion of new programming models and the restructuring of code to achieve good performance. Addressing these issues is beyond the capability of any single person or team—leading to the need for collaboration among many teams, who encapsulate their expertise in reusable software and work together to create sustainable software ecosystems.

97 MATHEMATICS AND COMPUTING↗

SmartFuse: Reconfigurable Smart Switches to Accelerate Fused Collectives in HPC Applications

Communication switches have sometimes been augmented to process collectives (e.g., the IBM BlueGene project and the Mellanox SHArP switch). In this work, we find that there is a great acceleration opportunity through the further augmentation of switches to accelerate more complex functions that combine communication with computation. We consider three types of such functions. The first is fully-fused collectives built by fusing multiple existing collectives like Allreduce with Alltoall. The second is semi-fused collectives built by combining a collective with another computation. The third we refer to as higher-order collectives built by combining multiple computations and communications, such as to perform matrix-matrix multiply (PGEMM). In this work, we propose a framework called SmartFuse to accelerate fused collective functions. The core of SmartFuse is a reconfigurable smart switch to support these operations. The semi/fully fused collectives are implemented with a CGRAlike architecture, while higher-order collectives are implemented with a more specialized computational unit that can also schedule communication. Supporting our framework is software to evaluate and translate relevant parts of the input program, compile them into a control data flow graph, and then map this graph to the switch hardware. The proposed framework, once deployed, has the strong potential to accelerate existing HPC applications transparently by encapsulation within an MPI implementation. Experimental results show that this approach improves the performance of the PGEMM kernel, MINIFE, and AMG by, on average, 94%, 15%, and 13%, respectively.

Haghi, Pouya↗