Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Communication architecture”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Modeling Data Movement Performance on Heterogeneous Architectures

The cost of data movement on parallel systems varies greatly with machine architecture, job partition, and nearby jobs. Performance models that accurately capture the cost of data movement provide a tool for analysis, allowing for communication bottlenecks to be pinpointed. Modern heterogeneous architectures yield increased variance in data movement as there are a number of viable paths for inter-GPU communication. In this paper, we present performance models for the various paths of inter-node communication on modern heterogeneous architectures, including the trade-off between GPUDirect communication and copying to CPUs. Furthermore, we present a novel optimization for inter-node communication based on these models, utilizing all available CPU cores per node. Finally, we show associated performance improvements for MPI collective operations.

97 MATHEMATICS AND COMPUTING↗

Compressed basis GMRES on high-performance graphics processing units

Krylov methods provide a fast and highly parallel numerical tool for the iterative solution of many large-scale sparse linear systems. To a large extent, the performance of practical realizations of these methods is constrained by the communication bandwidth in current computer architectures, motivating the investigation of sophisticated techniques to avoid, reduce, and/or hide the message-passing costs (in distributed platforms) and the memory accesses (in all architectures). This article leverages Ginkgo’s memory accessor in order to integrate a communication-reduction strategy into the (Krylov) GMRES solver that decouples the storage format (i.e., the data representation in memory) of the orthogonal basis from the arithmetic precision that is employed during the operations with that basis. Given that the execution time of the GMRES solver is largely determined by the memory accesses, the cost of the datatype transforms can be mostly hidden, resulting in the acceleration of the iterative step via a decrease in the volume of bits being retrieved from memory. Together with the special properties of the orthonormal basis (whose elements are all bounded by 1), this paves the road toward the aggressive customization of the storage format, which includes some floating-point as well as fixed-point formats with mild impact on the convergence of the iterative process. We develop a high-performance implementation of the “compressed basis GMRES” solver in the Ginkgo sparse linear algebra library using a large set of test problems from the SuiteSparse Matrix Collection. We demonstrate robustness and performance advantages on a modern NVIDIA V100 graphics processing unit (GPU) of up to 50% over the standard GMRES solver that stores all data in IEEE double-precision.

97 MATHEMATICS AND COMPUTING↗

Architectural scaling tradeoffs in modular 3D bosonic quantum processors

We propose a modular three-dimensional bosonic quantum processor built from repeatable coupled-cavity modules linked by configurable interconnect networks. Using hardware-motivated graph-theoretic measures, we compare nearest-neighbor, hub-based, and hybrid architectures in terms of interconnect count, communication distance, resource concentration, and implementation complexity. Rather than identifying a universally optimal topology, our analysis shows how these architectures redistribute the costs of scaling, including wiring and port requirements, nonlocal communication distance, exposure to shared resources, routing bottlenecks, and scheduling overhead. Case studies of a \(3\times3\) processor and a larger hierarchical architecture further distinguish finite-size performance from asymptotic scaling. The resulting framework provides a systematic basis for evaluating modular three-dimensional bosonic processors and for identifying the device-level parameters required for quantitative hardware design.

Zhu, Shaojiang [Fermilab] (ORCID:0000000293180092)↗

Programmable photonic integrated meshes for modular generation of optical entanglement links

Abstract Large-scale generation of quantum entanglement between individually controllable qubits is at the core of quantum computing, communications, and sensing. Modular architectures of remotely-connected quantum technologies have been proposed for a variety of physical qubits, with demonstrations reported in atomic and all-photonic systems. However, an open challenge in these architectures lies in constructing high-speed and high-fidelity reconfigurable photonic networks for optically-heralded entanglement among target qubits. Here we introduce a programmable photonic integrated circuit (PIC), realized in a piezo-actuated silicon nitride (SiN)-in-oxide CMOS-compatible process, that implements an N × N Mach–Zehnder mesh (MZM) capable of high-speed execution of linear optical transformations. The visible-spectrum photonic integrated mesh is programmed to generate optical connectivity on up to N = 8 inputs for a range of optically-heralded entanglement protocols. In particular, we experimentally demonstrated optical connections between 16 independent pairwise mode couplings through the MZM, with optical transformation fidelities averaging 0.991 ± 0.0063. The PIC’s reconfigurable optical connectivity suffices for the production of 8-qubit resource states as building blocks of larger topological cluster states for quantum computing. Our programmable PIC platform enables the fast and scalable optical switching technology necessary for network-based quantum information processors.

47 OTHER INSTRUMENTATION↗

Distributed Energy Resource Cybersecurity Standards Development [Final Report]

Currently, the solar industry is operating with little application-specific guidance on how to protect and defend their systems from cyberattacks. This 3-year Department of Energy (DOE) Solar Energy Technologies Office-funded project helped advance the distributed energy resource (DER) cybersecurity state-of-the-art by (a) bolstering industry awareness of cybersecurity concepts, risks, and solutions through a webinar series and (b) developing recommendations for DER cybersecurity standards to improve the security performance of DER products and networks. Drafting DER standards is a lengthy, consensus-based process requiring effective leadership and stakeholder participation. This project was designed to reduce standard and guide writing times by creating well-researched recommendations that could act as a starting place for national and international standards development organizations. Working within the SunSpec/Sandia DER Cybersecurity Workgroup, the team produced guidance for DER cybersecurity certification, communication protocol standards, network architecture s, access control, and patching. The team also led subgroups within the IEEE P 1547.3 Guide for Cybersecurity of Distributed Energy Resources Interconnected with Electric Power Systems committee and pushed a draft to ballot in October 2021.

14 SOLAR ENERGY↗

Space-based quantum networks are an essential component of future architecture for distributed quantum computers and quantum-enhanced secure communication

Space-based quantum links show great promise for connecting and communicating between quantum computers over ultra-long distances without the high loss incurred through fiber. A successful US quantum satellite would require large investment, national priority, and a diverse set of expertise. But, it would deliver US-owned quantum links that would allow for quantum-enhanced secure communications and the ability to connect quantum computers over long distances.

97 MATHEMATICS AND COMPUTING↗

OmniFed: A Modular Framework for Configurable Federated Learning from Edge to HPC

Federated Learning (FL) is critical for edge and High Performance Computing (HPC) where data is not centralized and privacy is crucial. We present OmniFed, a modular framework designed around decoupling and clear separation of concerns for configuration, orchestration, communication, and training logic. Its architecture supports configuration-driven prototyping and code-level override-what-you-need customization. We also support different topologies, mixed communication protocols within a single deployment, and popular training algorithms. It also offers optional privacy mechanisms including Differential Privacy (DP), Homomorphic Encryption (HE), and Secure Aggregation (SA), as well as compression strategies. These capabilities are exposed through well-defined extension points, allowing users to customize topology and orchestration, learning logic, and privacy/compression plugins, all while preserving the integrity of the core system. We evaluate multiple models and algorithms to measure various performance metrics. By unifying topology configuration, mixed-protocol communication, and pluggable modules in one stack, OmniFed streamlines FL deployment across heterogeneous environments. Github repository is available at https://github.com/at-aaims/OmniFed.

Tyagi, Sahil [ORNL] (ORCID:0009000783144745)↗

Incremental Threshold Scheme Enabled IoT Group Key Management

Cyber landscape evolves rapidly. Internet of Things (IoT) and Edge Computing (EC) have rapidly become an integral part of the modern computing infrastructure. It is expected that there will be more than 50 billion active and connected IoT devices by 2025 [1]. Pervasive IoT/EC creates unprecedented opportunities bridging the gap between previously segregated cyber and physical spaces. However, this progress also brings along new security challenges. IoT devices typically have limited computation, communication, and storage resources. This leads to security architecture designs such as using symmetric keys for group communication. While secure and efficient in stable network settings, symmetric key solutions are ill-adapted for IoT's highly dynamic device mobility behavior and frequent group membership turnover. Whenever IoT members leave a group, the known symmetric keys cannot be made forgotten, posing a serious vulnerability. This leads to frequent re-groupings that require expensive re-authentication, key regeneration, and key redistribution in order to maintain IoT/EC security. We present a novel symmetric key management framework that integrate an Incremental Threshold Scheme (ITS) cryptographical function into communication protocol's key rotation mechanism to allow for secure and efficient symmetric key communication group member node revocation. This ITS-enabled key management framework alleviates the need of frequent and expensive re-grouping and re-keying needed by today's large and dynamic IoT/EC operations. We further applied this ITS-enabled key management framework to a distributed IoT/EC-integrated publish and subscribe framework for applicability validation.

Li, Mingyan↗

Standard Modular Architecture for Consumer End Plug and Play Interfaces

The growth in load and generation sources at the edge of the grid is driving the innovation in power electronic (PE) grid interfaces at the consumer end to improve the grid resiliency, reliability, security, and cost of the future infrastructure. Novel PE technologies are key enablers for the future grid infrastructure to handle the issues that arise because of the projected growth. This paper introduces a novel fundamental building block (FBB) architecture with plug-and play features. An example framework that enables co-ordination of multiple FBBs hierarchically to provide various grid functions is also presented. The FBBs are designed to be equipped with advanced features like online health monitoring, embedded intelligence, and decision-making capability for enhancing the metrics of consumer end plug and play interfaces. The paper elaborates on the proposed architecture and the developed framework i.e., the controls, communication, protection, and the corresponding timing requirements for the building blocks. The architecture and the framework have been validated through simulations and the communication framework has been validated with a control hardware in the loop platform.

Chinthavali, Madhu Sudhan↗

A Review of Edge Computing Technology and Its Applications in Power Systems

Recent advancements in network-connected devices have led to a rapid increase in the deployment of smart devices and enhanced grid connectivity, resulting in a surge in data generation and expanded deployment to the edge of systems. Classic cloud computing infrastructures are increasingly challenged by the demands for large bandwidth, low latency, fast response speed, and strong security. Therefore, edge computing has emerged as a critical technology to address these challenges, gaining widespread adoption across various sectors. This paper introduces the advent and capabilities of edge computing, reviews its state-of-the-art architectural advancements, and explores its communication techniques. A comprehensive analysis of edge computing technologies is also presented. Furthermore, this paper highlights the transformative role of edge computing in various areas, particularly emphasizing its role in power systems. It summarizes edge computing applications in power systems that are oriented from the architectures, such as power system monitoring, smart meter management, data collection and analysis, resource management, etc. Additionally, the paper discusses the future opportunities of edge computing in enhancing power system applications.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Mosaics , The Best of Both Worlds: Analog devices with Digital Spiking Communication to build a Hybrid Neural Network Accelerator

Neuromorphic architectures have seen a resurgence of interest in the past decade owing to 100x-1000x efficiency gain over conventional Von Neumann architectures. Digital neuromorphic chips like Intel's Loihi have shown efficiency gains compared to GPUs and CPUs and can be scaled to build larger systems. Analog neuromorphic architectures promise even further savings in energy efficiency, area, and latency than their digital counterparts. Neuromorphic analog and digital technologies provide both low-power and configurable acceleration of challenging artificial intelligence (AI) algorithms. We present a hybrid analog-digital neuromorphic architecture that can amplify the advantages of both high-density analog memory and spike-based digital communication while mitigating each of the other approaches' limitations.

97 MATHEMATICS AND COMPUTING↗

Grid-Connected Modular Soft-Switching Solid State Transformers (M-S4T)

The objective of this project is to develop and verify the concept of a flexible and modular soft-switching solid-state transformer (M-S4T) for direct grid-connected applications. The ability to directly connect power electronics converters to the medium voltage grid (4 kV – 13 kV), and to potentially replace the passive and bulky, but ubiquitous 60 hertz service transformer in the 25 kVA to 100 kVA range, with a more flexible and controllable device, has been regarded as the ‘holy grail’ in grid control. However, this has proven to be extremely difficult. This project has developed the solutions to several key challenges of the direct grid-connected power electronics and realized a 7.2 kV M-S4T prototype. First, a protection method to protect the M-S4T from the high voltages (110 kV for the 13 kV system) that occur on the grid due to transients and lightning strikes have been developed and experimentally verified. Second, the realization and the operation of the M-S4T based on high-voltage SiC devices (>3.3 kV) and a medium-frequency medium-voltage low-leakage transformer in a single-stage solid-state transformer with zero-voltage switching, low dv/dt, and low electromagnetic interference has been successfully demonstrated up to 7.5 kV peak. Third, an oil-cooling system and stable communication and distributed control system for converter module voltage sharing have been developed and experimentally verified. The developed M-S4T has realized a modular universal high-performance power conversion system. This conversion system is scalable to different voltage and power levels and adaptable to four-quadrant bidirectional operation. Moreover, the use of passive cooling techniques meets the equipment life requirements, and the lightning protection scheme fulfills the basic insulation level specifications for direct grid connection. Such power conversion system opens up near-term opportunities, including energy storage, solar PV, or electric vehicle charging with significant cost and footprint savings. In the longer term, the possibility of replacing the utility distribution transformer with an M-S4T will be transformative for future distribution grids with a compact footprint and full controllability to enable high renewable energy and storage penetration. In addition to the main project, this report expands on the Plus-Up projected including as part of the main award. This project developed and demonstrated the technology for autonomous collaborative inverters that can be connected in an ad hoc manner to the grid. The aim of the project was to: (1) evaluate the existing techniques for grid-connected inverters and find their limitations; (2) develop detailed requirements for grid-connected inverters in the modern grid with millions of active nodes; (3) design a unified control strategy that brings more autonomy and intelligence to grid-connected inverters, and addresses parts of the issues with the existing techniques. The proposed technique, called UniCon, enables inverters to 1) connect/disconnect to/from the grid in an ad hoc manner; (2) work based on local sensing. Slow communication could be used for a more optimized behavior; (3) work automatically in both grid-forming/grid-following mode; (4) handle large disturbances, e.g., big load step and fault, in an oscillation-free manner; (5) work collaboratively with other inverters in steady-state and during transients. UniCon can be implemented in the middle-level control; hence it is agnostic to the vendor and to the implementation of the inner voltage/current and protection loops. Furthermore, a new synchronization scheme, based on deep learning, was developed that can extract the grid voltage phase and amplitude in a stable manner. The method is cheap to implement can improve the dynamic performance of the grid-connected inverters during fast transients, e.g., fault. The proposed control scheme was validated by (1) MATLAB/Simulink; (2) hardware-in-the-loop results, and; (3) experimental results using three inverters that form a microgrid in a down-scaled feeder. Lastly, both the M-S4T and UniCon have achieved promising tangible paths to markets. In the case of the M-S4T, the underlying technology — the Soft Switching Solid State Transformer (S4T) developed at the Georgia Tech Center for Distributed Energy (GT-CDE) has been licensed by GridBlock from the Georgia Tech Research Corporation, and GridBlock has been working with manufacturing partner Jabil (one of the largest US-based contract manufacturers) and system integrator Power Secure (largest deployer of microgrids in the US with 4.7 GW under management), to meet the strong initial demand. Similarly, GridBlock has an exclusive license to the UniCon technology, developed under this award by GT-CDE. The UniCon provides an intermediate control layer that enables the implementation of the higher-level ‘transactive’ control commands for the system. The architecture of the system - slow communications with the cloud for system optimization and setpoints, and the use of locally measured quantities for real-time control, provide a very robust and secure way of implementing a real-time must-run grid that is also secure and stable. This is a brand-new functionality that is critical for the future grid and key to GridBlock’s business model.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Integrating 5G Technology for Improved Process Monitoring and Network Slicing in ICS

Industrial Control Systems (ICS) are crucial for monitoring physical processes that support essential cyber-enabled services like power generation. The use of proprietary communication and lack of effective intrusion detection mechanisms pose constraints for efficient operation. Therefore, there is a need to modernize these systems with decentralized technologies like Edge Computing and 5G. However, integrating 5G and Edge Computing into large-scale ICS networks presents implementation and performance challenges. To address these challenges, this paper proposes an integrated ICS architecture that combines 5G and Edge Computing technologies with traditional ICS protocols. The objective is to minimize implementation and operational difficulties while improving the monitoring of physical processes and enabling robust intrusion detection. The proposed architecture outlines the necessary components, services, and communication protocols required for the integration of 5G and Edge Computing.

Aguayo, Jared M.↗

Optimizing High Performance Markov Clustering for Pre-Exascale Architectures

HipMCL is a high-performance distributed memory implementation of the popular Markov Cluster Algorithm (MCL) and can cluster large-scale networks within hours using a few thousand CPU-equipped nodes. It relies on sparse matrix computations and heavily makes use of the sparse matrix-sparse matrix multiplication kernel (SpGEMM). The existing parallel algorithms in HipMCL are not scalable to Exascale architectures, both due to their communication costs dominating the runtime at large concurrencies and also due to their inability to take advantage of accelerators that are increasingly popular. In this work, we systematically remove scalability and performance bottlenecks of HipMCL. We enable GPUs by performing the expensive expansion phase of the MCL algorithm on GPU. Additionally, we propose a CPU-GPU joint distributed SpGEMM algorithm called pipelined Sparse SUMMA and integrate a probabilistic memory requirement estimator that is fast and accurate. Furthermore, we develop a new merging algorithm for the incremental processing of partial results produced by the GPUs, which improves the overlap efficiency and the peak memory usage. We also integrate a recent and faster algorithm for performing SpGEMM on CPUs. We validate our new algorithms and optimizations with extensive evaluations. With the enabling of the GPUs and integration of new algorithms, HipMCL is up to 12.4x faster, being able to cluster a network with 70 million proteins and 68 billion connections just under 15 minutes using 1024 nodes of ORNL's Summit supercomputer.

97 MATHEMATICS AND COMPUTING↗

Automated Controller Hardware-In-The-Loop Testbed for EV Charger Resilience Analysis

This paper focuses on the development of a tool that includes an automated testbed with controls, protection, and communications integrated into a real-time system to provide a platform to generate data sets for failure modes and effects analysis. This tool establishes a value for automation of data generation for different scenarios and addresses the gap of nonexistent field data for different applications and use cases. The features of this tool can further be expanded to include multiple power electronics models, communication protocols, and scaled system architectures. This general framework was evaluated for a DC fast charger system use case to provide quantitative solution for resiliency.

Starke, Michael↗

Scalable/Secure Cooperative Algorithms and Framework for Extremely-high Penetration Solar Integration (SolarExPert) (Final Technical Report)

This SolarExPert project has developed a Sustainable Grid Platform (SGP) with scalable architecture of distributed control and optimization. The SGP consists of the following major functions: 1) an advanced grid architecture with hierarchical and distributed communication and control, combined with the OpenFMB standard and implemented on the Multi-Agent OpenDSS (MA-OpenDSS) platform; 2) an online distributed stochastic optimal power flow; 3) an online distributed system state estimation algorithm; 4) the distributed Volt/VAR optimization and frequency control algorithms; 5) the distributed distribution system restoration strategy. The developed SGP together with advanced functions are tested in 1 million (1M)-node distribution system on the MA-OpenDSS platform. Furthermore, the models and algorithms are tested in 100,000-node system HiL simulation, and also in P-HiL implementation with 100 physical devices. The developed functions haven been validated and tested on the selected actual distribution feeder with the data collected from the field of Maui Meadows in Hawaii. The distributed PV hosting capacities with cooperative Volt/VAR and Volt/VAR/Watt control are estimated and compared to provide recommendations for customers and the utility company.

14 SOLAR ENERGY↗

Control Oriented Models for Co-Design: Technical Overview of MT HVDC, MVDC, and Solid State Transformer Building Blocks

The electric power system is shifting toward a power electronics–enabled grid, where converter based “building blocks” (e.g., high voltage direct current (HVDC) links, multi terminal HVDC (MT HVDC) networks, medium voltage DC (MVDC) links, and solid state transformers (SSTs)) provide fast, precise control of power flows, voltage, and frequency. This report develops and applies publicly shareable electromagnetic transient (EMT) and phasor models to examine how such building blocks can be composed and coordinated to support offshore wind integration, inter area transfers, feeder support, and resilience. Section 2 documents a modular multilevel converter (MMC)–based MT HVDC modeling framework and two use cases: a compact WSCC/IEEE 9 bus test system and a 240 bus “mini WECC” case with five offshore wind plants (OWFs). Phasor to EMT transfer, initialization, and sanity checks are summarized, and neutral demonstrations of normal and contingency operation are reported. Section 3 frames the problem of wind plant inertial frequency response (IFR): shaping energy release and recovery to improve nadir while avoiding aerodynamic stall; representative simulations illustrate the issues without disclosing proprietary control. Section 4 develops MVDC concepts through an IEEE 16 bus loop and an Olympic Peninsula case study that compares AC vs. MVDC corridors and shows how feeder headroom can be pooled via DC couplers. Section 5 surveys SST architectures and identifies a gap: scalable, communication free coordination of multiple SSTs for islanded feeder networks. Across the report, novel methods and configurations under separate publication and IP review are not disclosed; only topic oriented, replicable setups and non proprietary results are shown. These models and use cases are intended as foundations for future publications and co design studies on architecture, control, and coordination of PE enabled grids.

24 POWER TRANSMISSION AND DISTRIBUTION↗