Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Communication architecture”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Fast and Scalable Sparse Triangular Solver for Multi-GPU Based HPC Architectures

Designing efficient and scalable sparse linear algebra kernels on modern multi-GPU based HPC systems is a daunting task due to significant irregular memory references and workload imbalance across the GPUs. This is particularly the case for \textit{Sparse Triangular Solver (SpTRSV)} which introduces additional two-dimensional computation dependencies among subsequent computation steps. Dependency information is exchanged and shared among GPUs, thus warrant for efficient memory allocation, data partitioning, and workload distribution as well as fine-grained communication and synchronization support. In this work, we demonstrate that directly adopting unified memory can adversely affect the performance of SpTRSV on multi-GPU architectures, despite linking via fast interconnect like NVLinks and NVSwitches. Alternatively, we employ the latest NVSHMEM technology based on Partitioned Global Address Space programming model to enable efficient fine-grained communication and drastic synchronization overhead reduction. Furthermore, to handle workload imbalance, we propose a malleable task-pool execution model which can further enhance the utilization of GPUs. By applying these techniques, our experiments on the NVIDIA multi-GPU supernode V100-DGX-1 and DGX-2 systems demonstrate that our design can achieve on average 3.53x (up to 9.86x) speedup on a DGX-1 system and 3.66x (up to 9.64x) speedup on a DGX-2 system with 4-GPUs over the Unified-Memory design. The comprehensive sensitivity and scalability studies also show that the proposed zero-copy SpTRSV is able to fully utilize the computing and communication resources of the multi-GPU system.

Xie, Chenhao↗

Efficient Anomaly Detection Driven By Different Machine Learning Architectures And Models

The rapid growth and ubiquitous adoption of the internet and cyber-physical systems (CPS) have fundamentally transformed modern communication, work, and human-system interactions. While networks now form the backbone of critical digital ecosystems, enabling seamless data transmission across diverse, interconnected systems, this increased connectivity also expands the attack surface, making real-time detection of network intrusions and anomalies a pressing challenge. Detecting unusual activities within network infrastructure requires advanced data traffic analysis to differentiate between legitimate and malicious interactions. Traditional approaches to network anomaly detectionâ??such as rule-based and signature-based systemsâ??often depend on predefined patterns to identify known anomalies, limiting their effectiveness against emerging, stealthy, or previously unseen threats. These conventional methods suffer from high false alarm rates and fail to adapt to the ever-evolving nature of network traffic, particularly in large-scale, decentralized environments where data volume, velocity, and variety are constantly increasing. This dissertation presents artificial intelligence (AI)-driven approaches to anomaly detection that leverage graphics processing unit (GPU)-enabled high-performance computing (HPC) platforms for processing massive network traffic data and monitoring the components of cyber-physical systems (CPS) for potentially hazardous conditions. The research advances several key contributions: (1) Designing efficient machine learning techniques for CPS condition monitoring and anomaly detection; (2) enabling federated learning (FL) frameworks that enable distributed detection while preserving data privacy and system resilience; (3) exploring graph-based methodologies combining graph neural networks (GNN) and graph machine learning (ML) approaches for the Internet of Things (IoT) and automotive network security, and (4) performing distributed edge computing optimizations that integrate FL with scalable technologies for reduced communication overhead. Through extensive experiments, these methodologies demonstrate that complex anomaly detection and condition monitoring tasks can be achieved while balancing computational efficiency and detection accuracy through fine-grained network information processing. The frameworks developed in this research establish a robust foundation for network anomaly detection, providing scalable, adaptive, and privacy-preserving solutions for safeguarding CPS and IoT networks in an increasingly interconnected digital landscape. The practical implications of these research findings are significant, as they can inform the development of next-generation network security systems and contribute to the protection of critical infrastructure against sophisticated cyber attacks.

Marfo, William↗

Designing a Magnetic Measurement Data Acquisition and Control System With Reuse in Mind: A Rotating Coil System Example

Accelerator magnet test facilities frequently need to measure different magnets on differently equipped test stands and with different instrumentation. Designing a modular and highly reusable system that combines flexibility built-in at the architectural level as well as on the component level addresses this need. Specification of the backbone of the system, with the interfaces and dataflow for software components and core hardware modules, serves as a basis for building such a system. The design process and implementation of an extensible magnetic measurement data acquisition and control system are described, including techniques for maximizing the reuse of software. The discussion is supported by showing the application of this methodology to constructing two dissimilar systems for rotating coil measurements, both based on the same architecture and sharing core hardware modules and many software components. The first system is for production testing 10 m long cryo-assemblies containing two MQXFA quadrupole magnets for the high-luminosity upgrade of the Large Hadron Collider and the second for testing IQC conventional quadrupole magnets in support of the accelerator system at Fermilab.

43 PARTICLE ACCELERATORS↗

The Design and Evaluation of Zero Trust Architecture for Electric Vehicle Charging Infrastructure: EVs @ Scale Series on EV Charging Station Cybersecurity

Implementing a zero trust architecture can significantly bolster the security of electric vehicle (EV) charging infrastructure. EV charging infrastructure includes numerous networked interfaces, each of which can present potential vulnerabilities. When these vulnerabilities are exploited, they can compromise the entire system, leading to severe operational and security risks. Zero trust is a security model that operates on the principle of "never trust, always verify," which helps manage the attack surface and limit the scope of any potential compromises. Fundamentally, this model ensures that no entity, whether inside or outside the network, is trusted by default. The design principles of zero trust include continuous verification, strict deny-by-default access controls, and micro-segmentation. Continuous verification ensures that every request is thoroughly checked, regardless of its origin. Strict access controls enforce the principle of least privilege, allowing users and devices only the minimum necessary access to perform their functions. Micro-segmentation involves dividing the network into smaller, isolated segments to prevent lateral movement in case of a breach. In the context of EV charging infrastructure, zero trust can be implemented through various strategies. For example, multi-factor authentication (MFA) can be required for engineers to access the management interfaces and control systems of charging stations. Real-time monitoring and analysis of network traffic can help detect and respond to anomalies. Systems that do not need to communicate with each other can be micro-segmented to enhance security. All communications should adhere to predefined policies to be permitted. Additionally, encrypting communications can protect sensitive information exchanged between chargers and management systems. This paper presents a zero trust architecture specifically designed for EV charging infrastructure. Implementing zero trust not only mitigates risks but also builds a resilient infrastructure capable of withstanding and quickly recovering from cyber threats. The architecture addresses six defined security objectives. A comprehensive test plan is developed to assess the architecture against these objectives, and the results of the evaluation are reported. This approach is essential for maintaining the reliability and integrity of EV charging services in an increasingly interconnected and vulnerable digital landscape. This is the first in a planned series of papers exploring the implementation of zero trust in EV charging infrastructure. Each paper will delve into different aspects and applications of zero trust, highlighting how various work processes and requirements can lead to distinct architectural designs. These architectures will be tailored to address specific security challenges and operational needs within the EV charging ecosystem, ensuring a robust and adaptable security framework.

33 ADVANCED PROPULSION SYSTEMS↗

A Novel Architecture for Attack-Resilient Wide-Area Protection and Control System in Smart Grid

Wide-area protection and control (WAPAC) systems are widely applied in the energy management system (EMS) that rely on a wide-area communication network to maintain system stability, security, and reliability. As technology and grid infrastructure evolve to develop more advanced WAPAC applications, however, so do the attack surfaces in the grid infrastructure. This paper presents an attack-resilient system (ARS) for the WAPAC cybersecurity by seamlessly integrating the network intrusion detection system (NIDS) with intrusion mitigation and prevention system (IMPS). In particular, the proposed NIDS utilizes signature and behavior-based rules to detect attack reconnaissance, communication failure, and data integrity attacks. Further, the proposed IMPS applies state transition-based mitigation and prevention strategies to quickly restore the normal grid operation after cyberattacks. As a proof of concept, we validate the proposed generic architecture of ARS by performing experimental case study for wide-area protection scheme (WAPS), one of the critical WAPAC applications, and evaluate the proposed NIDS and IMPS components of ARS in a cyber-physical testbed environment. Our experimental results reveal a promising performance in detecting and mitigating different classes of cyberattacks while supporting an alert visualization dashboard to provide an accurate situational awareness in real-time.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Grid Architecture Mapping to Understand Transformation (GAMUT): Concept Definition

The electric power system is undergoing significant transformation driven by changes in the generation mix, increasing reliance on information and communication technologies, evolving customer expectations, and a dynamic cyber-physical environment. To address these challenges, substantial investments will be made over the next decade to create a flexible, affordable, secure, reliable, and resilient electric system. In response, the U.S. Department of Energy's Office of Electricity is developing knowledge management tools aimed at systematically documenting technology pilots and demonstration projects. This initiative seeks to improve understanding of operating contexts, planning steps, and integration requirements for innovative grid technologies. The Department of Energy project leverages Grid Architecture principles to collect and organize insights into the interdependencies, requirements, and capabilities of various grid solutions. By employing a systematic approach, the project aims to perform a feasibility study for the development of a software tool that will support stakeholders, including regulators, transmission and distribution system operators, distributed energy resources aggregators, and technology providers. By offering a systematic, software-based framework to document technologies, understand their benefits and impacts, and inform deployment strategies, the proposed Grid Architecture Mapping to Understand Transformation (GAMUT) Tool is intended to reduce costs, enhance decision-making, and facilitate scaling and effective deployment of new technologies. The feasibility study will evaluate the technical and financial viability of the GAMUT Tool, identifying stakeholder needs and assessing risks and mitigation strategies. Ultimately, the GAMUT team aims to develop an initial minimal viable product, followed by staged improvements, to aid in the modernization of the electric grid, contributing to a more sustainable and resilient energy future.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Scalable multiscale modeling of platelets with 100 million particles

Here, we developed the core components of the AI-aided multiple time stepping algorithm for multiscale modeling of cell dynamics. This algorithm was implemented and analyzed on two supercomputer architectures with an application of simulating the aggregation of 250 platelets, or 102 million particles. To scale on these computers with complex memory and network architectures with GPUs, we devised a biomechanics-informed task mapping scheme to optimize load imbalance, communications, and memory utilization. Our simulations, scaling well up to 192 nodes on a Summit-like supercomputer with a peak speed of 11 petaflops, achieved a rate of 423 μs/day which is 500 times faster than the conventional algorithm using static time step and this has enabled studies of record size blood clots at record spatial–temporal resolutions. Additionally, we discovered the sensitive dependence of the scalability and execution time on the methods of decomposition, CPU–GPU coupling, and task mapping.

97 MATHEMATICS AND COMPUTING↗

Allosteric control of dynamin-related protein 1 through a disordered C-terminal Short Linear Motif

Abstract The mechanochemical GTPase dynamin-related protein 1 (Drp1) catalyzes mitochondrial and peroxisomal fission, but the regulatory mechanisms remain ambiguous. Here we find that a conserved, intrinsically disordered, six-residue Short Linear Motif at the extreme Drp1 C-terminus, named CT-SLiM, constitutes a critical allosteric site that controls Drp1 structure and function in vitro and in vivo. Extension of the CT-SLiM by non-native residues, or its interaction with the protein partner GIPC-1, constrains Drp1 subunit conformational dynamics, alters self-assembly properties, and limits cooperative GTP hydrolysis, surprisingly leading to the fission of model membranes in vitro. In vivo, the involvement of the native CT-SLiM is critical for productive mitochondrial and peroxisomal fission, as both deletion and non-native extension of the CT-SLiM severely impair their progression. Thus, contrary to prevailing models, Drp1-catalyzed membrane fission relies on allosteric communication mediated by the CT-SLiM, deceleration of GTPase activity, and coupled changes in subunit architecture and assembly-disassembly dynamics.

Science & Technology - Other Topics↗

Enabling Efficient Sparse Computations using Linear Algebra Aware Compilers

This project developed the LAPIS compiler framework, built on the Multilevel Intermediate Representation (MLIR), to optimize sparse linear algebra operations and support performance portability across diverse architectures. The main innovation of LAPIS is the Kokkos dialect, which allows for lowering codes from a high productivity language to different architectures in an elegant way. The dialect also allows the conversion of lower-level MLIR code to C++ Kokkos code, facilitating the integration of scientific machine learning (SciML) models into applications. To extend LAPIS for distributed memory architectures, a new partition dialect was created to manage the distribution of sparse tensors and express communication patterns for sparse linear algebra operations. This dialect also supports the distributed execution of operators and includes algorithmic optimizations to minimize communication to improve performance. The project also demonstrates that MLIR can enable effective linear algebra-level optimizations, improving performance on different GPUs for both sparse and dense linear algebra kernels. Key applications of LAPIS include sparse linear algebra and graph kernels, TenSQL, a relational database management solution built on GraphBLAS, and the development of subgraph isomorphism and monomorphism kernels, showcasing performance portability. In summary, the LAPIS framework supports productivity, performance, portability, and distributed memory execution, while also enabling linear algebra-level optimizations that are challenging in traditional programming languages, with successful applications ranging from simple sparse linear algebra to complex graph kernels.

97 MATHEMATICS AND COMPUTING↗

Distributed, Intelligent Edge-Sensing for a Smarter Grid

The electric grid is undergoing major transformations and developments resulting in unprecedented levels of volatility, uncertainty, and stress on grid infrastructure. Smart sensors and methods aiding in advanced visibility and situational awareness are key for tackling these issues. In this work, a decentralized architecture is proposed, where sensing, local computation and control capability are embedded in the edge devices, communicating with a set of trusted 'data mules' in a 'delay-tolerant' manner, while functioning autonomously. This system has been designed and implemented as an overall platform – called Global Asset Monitoring, Management and Analytics (GAMMA) Platform intended to provide the backbone for a global array of sensors and actuators. Further, as a building block for advanced current sensing solutions, a smart, low-cost ‘clip-on’ current sensor based on PCB-embedded Rogowski coil has been developed. The sensor hosts a novel signal conditioning stage allowing an 'auto-tuning' feature, resulting in a universal current sensor design for measuring a wide range of currents, including faults for smart grid applications. Finally, the research proposes a method to instrument and monitor key parameters for the most common electric utility asset – the pole-top distribution transformer. The work done in this research enables scalable, edge-intelligent sensing solutions for monitoring grid infrastructure, allowing utility operators to gain advanced visibility in an economical way.

Kulkarni, Shreyas Bhalchandra↗

EUREICA: Efficient UltRa Endpoint IoT-enabled Coordinated Architecture

The electricity grid has evolved from a physical system to a cyber-physical system with digital devices that perform measurement, control, communication, computation, and actuation. The increased penetration of distributed energy resources (DERs) that include renewable generation, flexible loads, and storage provides extraordinary opportunities for improvements in efficiency and sustainability. However, they can introduce new vulnerabilities in the form of cyberattacks, which can cause significant challenges in ensuring grid resilience. The purpose of this project was to develop a framework ((Efficient, Ultra-REsilient, IoT-Coordinated Assets, or EUREICA)for achieving grid resilience through suitably coordinated assets including a network of Internet of Things (IoT) devices, and a local electricity market (LEM) to identify trustable assets and carry out this coordination. Situational Awareness (SA) of locally available DERs with the ability to inject power or reduce consumption is enabled by the market, together with a monitoring procedure for their trustability and commitment. Experiments conducted during this project demonstrated that, with this SA, a variety of cyberattacks can be mitigated using local trustable resources without stressing the bulk grid. The demonstrations were carried out using a variety of high-fidelity co-simulation platforms, real-time hardware-in-the-loop validation, and a utility-friendly simulator.

14 SOLAR ENERGY↗

XaaS: Acceleration as a Service to Enable Productive High-Performance Cloud Computing

High-performance computing (HPC) and the cloud have evolved independently, specializing their innovations into performance or productivity. Acceleration as a Service (XaaS) is a recipe to empower both fields with a shared execution platform that provides transparent access to computing resources, regardless of the underlying cloud or HPC service provider. Bridging HPC and cloud advancements, XaaS presents a unified architecture built on performance-portable containers. Here, our converged model concentrates on low-overhead, high-performance communication and computing, targeting resource-intensive workloads from climate simulations to machine learning. XaaS lifts the restricted allocation model of Function as a Service (FaaS), allowing users to benefit from the flexibility and efficient resource utilization of serverless computing while supporting long-running and performance-sensitive workloads from HPC.

97 MATHEMATICS AND COMPUTING↗

Parthenon—a performance portable block-structured adaptive mesh refinement framework

On the path to exascale the landscape of computer device architectures and corresponding programming models has become much more diverse. While various low-level performance portable programming models are available, support at the application level lacks behind. To address this issue, we present the performance portable block-structured adaptive mesh refinement (AMR) framework Parthenon, derived from the well-tested and widely used Athena++ astrophysical magnetohydrodynamics code, but generalized to serve as the foundation for a variety of downstream multi-physics codes. Parthenon adopts the Kokkos programming model, and provides various levels of abstractions from multidimensional variables, to packages defining and separating components, to launching of parallel compute kernels. Parthenon allocates all data in device memory to reduce data movement, supports the logical packing of variables and mesh blocks to reduce kernel launch overhead, and employs one-sided, asynchronous MPI calls to reduce communication overhead in multi-node simulations. Using a hydrodynamics miniapp, we demonstrate weak and strong scaling on various architectures including AMD and NVIDIA GPUs, Intel and AMD x86 CPUs, IBM Power9 CPUs, as well as Fujitsu A64FX CPUs. At the largest scale on Frontier (the first TOP500 exascale machine), the miniapp reaches a total of 1.7 × 10 13 zone-cycles/s on 9216 nodes (73,728 logical GPUs) at [Formula: see text] weak scaling parallel efficiency (starting from a single node). In combination with being an open, collaborative project, this makes Parthenon an ideal framework to target exascale simulations in which the downstream developers can focus on their specific application rather than on the complexity of handling massively-parallel, device-accelerated AMR.

97 MATHEMATICS AND COMPUTING↗

FENATE

FENATE: Fast Evaluation of Network Architecture -- Toolchain and Environment. Our framework is comprised of scalable tools to both skeletonize and simulate MPI communication patterns providing a functional view of the network under investigation (trading off accuracy for speed)

Young, Stephen↗

Architecture Of A Multi-channel Data Streaming Device With An Fpga As A Coprocessor

The design of a data acquisition system often involves the integration of a Field Programmable Gate Array (FPGA) with analog front-end components to achieve precise timing and control. Reuse of these hardware systems can be difficult since they need to be tightly coupled to the communications interface and timing requirements of the specific ADC used. A hybrid design exploring the use of FPGA as a coprocessor to a traditional CPU in a dataflow architecture is presented. Reduction in the volume of data and gradual transitioning of data processing away from a hard real-time environment are both discussed. Chief design concerns, including data throughput and precise synchronization with external stimuli, are addressed. The discussion is illustrated by the implementation of a multi-channel digital integrator, a device based entirely on commercial off-the-shelf (COTS) equipment.

Nogiec, Jerzy M.↗

Development of a Practical Secondary Control for Hardware Microgrids

Practical, vendor-agnostic interoperability guidelines for the secondary control architecture of microgrids (MGs) with multiple grid-forming (GFM) inverter-based resources (IBRs) have not yet been developed. Therefore, this paper proposes a generic and vendor-agnostic secondary control architecture that operates with all GFM IBRs and synchronous generators. This secondary control does not require the use of additional measurement devices in the MG and utilizes the inherent communication systems of the GFM units, such as Modbus TCP/IP. The practical challenges of Modbus registers, such as packet loss and quantization error, and their detrimental impacts on secondary control actions are investigated. The proposed three-stage modification for any secondary control architecture to mitigate these impacts includes: 1) averaging the data read, 2) situational event-triggering of the controller, and 3) finite iteration of the controlling action. The proposed method is validated using a three-phase electric power, 480 V, 60 Hz, 500 kVA laboratory hardware microgrid with commercial two GFM IBRs and one diesel generator. The experimental results corroborates the fact the proposed modification in the secondary control architecture is advantageous for practical usage under erroneous measurements.

grid-forming inverter↗

Distributed Load Shedding Application Architecture and Bi-Level Predictive Estimator Algorithm

Increasing penetrations of distributed renewables are decreasing the effectiveness of traditional decentralized under-frequency load shedding (UFLS) schemes. As more distribution circuits begin to back-feed the transmission system, operation of traditional UFLS may exacerbate frequency instability. This paper presents the conceptual framework for a data-rich environment to coordinate UFLS across multiple distribution providers based on the laminar coordination framework in order to ensure optimal adaptive setting of UFLS relays. Communication and control are enabled through a distributed implementation of the IEC 61968-1 Common Information Model message bus structure. In addition to the proposed architecture, a novel adaptive UFLS scheme informed by a bi-level state estimator to create optimal relay setpoints is introduced. Initial simulation results are presented for the IEEE 14-bus test system on scenarios leading to mis-operation of traditional UFLS.

Anderson, Alexander A.↗

Development of a Practical Secondary Control for Hardware Microgrids: Preprint

Practical vendor-agnostic interoperability guidelines for the secondary control architecture of microgrids (MGs) with multiple grid-forming (GFM) inverter-based resources (IBRs) have not yet been developed. Therefore, this paper proposes a generic and vendor-agnostic secondary control architecture that works with all GFM IBRs and synchronous generators. This secondary control does not need to employ any additional measurement devices in the MG and uses the inherent communication systems of the GFM units, such as Modbus TCP/IP. The practical challenges of Modbus holding registers such as packet loss, quantization error, etc. and their deteriorating impacts on the secondary control action are investigated. The proposed three-stage modification of any secondary control architecture to eliminate these impacts includes 1) averaging the data read, 2) situational event-triggering of the controller, and 3) the finite iteration of the controlling action. The proposed method is validated using a laboratory hardware MG with commercial GFM units.

inverter based resources↗