Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Low latency”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Autonomy Loops for Monitoring, Operational Data Analytics, Feedback, and Response in HPC Operations

Many High Performance Computing (HPC) facilities have developed and deployed frameworks in support of continuous monitoring and operational data analytics (MODA) to help improve efficiency and throughput. Because of the complexity and scale of systems and workflows and the need for low-latency response to address dynamic circumstances, automated feedback and response have the potential to be more effective than current human-in-the-loop approaches which are laborious and error prone. Progress has been limited, however, by factors such as the lack of infrastructure and feedback hooks, and successful deployment is often site- and case-specific. In this position paper we report on the outcomes and plans from a recent Dagstuhl Seminar, seeking to carve a path for community progress in the development of autonomous feedback loops for MODA, based on the established formalism of similar (MAPE-K) loops in autonomous computing and self-adaptive systems. By defining and developing such loops for significant cases experienced across HPC sites, we seek to extract commonalities and develop conventions that will facilitate interoperability and interchangeability with system hardware, software, and applications across different sites, and will motivate vendors and others to provide telemetry interfaces and feedback hooks to enable community development and pervasive deployment of MODA autonomy loops.

autonomy loops↗

Integration of 5G and Time Sensitive Networks in Fossil Energy Generation Systems: A Case Study

Precise timing and data transmission within stringent time constraints are critical for numerous applications, such as robotics, virtual and augmented reality, industrial automation, energy and medical business, and various other sectors. Time-sensitive networks (TSN) and fifth-generation wireless communications (5G) are crucial for industrial communications, enabling convergent communication for various services using a common network core. Applications that are time-sensitive and require deterministic communications with low latency fall into this category, such as generation plant systems. This paper presents a simulated model of 5G-TSN for a fossil power plant based on the wired network parameters implemented in situ. Metrics analysis, comparison and future work are presented. © 2024 IEEE.

01 COAL, LIGNITE, AND PEAT↗

A Brief Survey of Data Streaming Technologies

Streaming data is data that is emitted at variable volumes in a continuous, incremental manner with the goal of low-latency processing often at a different physical location. Network infrastructure is used to facilitate the connection between data sources and sinks, and must be robust to handle the requirements of the workflow. The U.S. Department of Energy Office of Science (DOE SC) a federal agency supporting fundamental scientific research for energy and the Nation’s largest supporter of basic research in the physical sciences. DOE SC has the responsibility for operating $\mathbf{1 0}$ National Laboratories, and 28 scientific user facilities supporting advanced supercomputers, particle accelerators, large x-ray light sources, neutron scattering sources, and other specialized facilities for nanoscience and genomics. This paper investigates the state of streaming data workfows, and details some of the approaches to this challenging problem.

Kissel, Ezra↗

Scaled-up Neuromorphic Array Communications Controller (SNACC) for Large-scale Neural Networks

Neuromorphic computing is one promising post-Moore’s law era technology, which takes inspiration from biological brains to perform computing tasks. The human brain contains billions of neurons with trillions of synapses and as neuromorphic hardware systems scale to larger and larger sizes, the communication system used to transfer information between neuromorphic elements and traditional computers must scale to keep up. In prior work, we describe the use of a separate neuromorphic array communications controller to support low-latency, high-throughput communication between our neuromorphic systems and a traditional computer. In this work, the neuromorphic array communications controller is used to support the scaling of a neuromorphic development system which uses multiple neuromorphic processors arranged in a two-dimensional array. The neuromorphic array communications controller, along with scalable local connections, is used to create a scalable neuromorphic platform to enable the development and testing of large neuromorphic network arrays.

Young, Aaron↗

Adaptive Spatially Aware I/O for Multiresolution Particle Data Layouts

Large-scale simulations on nonuniform particle distributions that evolve over time are widely used in cosmology, molecular dynamics, and engineering. Such data are often saved in an unstructured format that neither preserves spatial locality nor provides metadata for accelerating spatial or attribute subset queries, leading to poor performance of visualization tasks. Furthermore, the parallel I/O strategy used typically writes a file per process or a single shared file, neither of which is portable or scalable across different HPC systems. We present a portable technique for scalable, spatially aware adaptive aggregation that preserves spatial locality in the output. We evaluate our approach on two supercomputers, Stampede2 and Summit, and demonstrate that it outperforms prior approaches at scale, achieving up to 2.5× faster writes and reads for nonuniform distributions. Furthermore, the layout written by our method is directly suitable for visual analytics, supporting low-latency reads and attribute-based filtering with little overhead.

Usher, Will↗

A 5G Enabled Adaptive Computing Workflow for Greener Power Grid

5G wireless technology can deliver higher data speeds, ultra low latency, more reliability, massive network capacity, increased availability, and a more uniform user experience to users. It brings additional power to help address the challenges brought by renewable integration and decarbonization. In this paper, a 5G enabled adaptive computing workflow tool has been presented that consists of various computing resources, such as 5G equipment, edge computing, cluster, Graphics processing unit (GPU) and cloud computing, with two examples showing technical feasibility for edge-grid-cloud interaction for real-time monitoring, security assessment, and forecasting. Benefiting from the high data transmission speed and massive connection capability of 5G, the workflow shows its potential to seamlessly integrate various applications at distributed and/or centralized locations to build more complex and powerful functions, with better flexibility.

5G technology, computational workflow, edge comput↗

Optimizing Non-Terrestrial Hybrid RF/FSO Links With Reinforcement Learning: Navigating Through Clouds

In the pursuit of ubiquitous broadband connectivity, there has been a significant shift towards the vertical expansion of communication networks into space, particularly through the exploitation of low Earth orbit (LEO) satellite constellations, which are favored for their relatively low latency. However, this approach faces many challenges that need to be addressed, including atmospheric turbulence, high path loss, and dynamic cloud formations. High-altitude pseudo-satellites (HAPS) have emerged as promising relaying layers between LEO satellites and ground stations, enhancing coverage, latency, and direct terrestrial user connectivity. While radio frequency (RF) bands suffer from congestion and limited bandwidth, free space optical (FSO) communications offer higher data rates, but are susceptible to misalignment and weather-induced signal degradation. To address these challenges, a hybrid RF/FSO approach has been proposed to take advantage of both technologies by dynamic switching between RF and FSO based on propagation channel conditions. This paper introduces a reinforcement learning-based algorithm designed to optimize the trajectory of HAPS, maneuver around cloudy areas, and seamlessly switch between the RF and FSO communication modes to maximize the achievable capacity. The proposed approach aims to maximize system performance by intelligently adapting to environmental conditions and offering a promising solution for next-generation space communication networks.

actor-critic algorithm↗

Alerga: Alert Aggregation and Reasoning in GOOSE Simulation Pipeline

IEC 61850 specifies the Generic Object Oriented Substation Event (GOOSE) protocol as one option for low latency communication of substation-related events. Due to its strict timing requirements, GOOSE lacks any form of encryption or authentication and has only minimal integrity guarantees. These absences render the protocol vulnerable to a variety of communication anomalies, including adversarial action. In particular, an adversary with access to the substation network can launch man in the middle (MITM) attacks. We propose Alerga, a set of tools to allow operators to mitigate some of the risks of the protocol while retaining its strengths. To that end, we have developed first a GOOSE simulation pipeline including data generation, anomaly detection, alert handling, causal reasoning and data visualization components. The simulator is designed to be modular, allowing operators to swap components to better fit their network capabilities. The volume of alert traffic on a substation network threatens operators with alert fatigue. In order to combat this, we secondly present a novel form of alert aggregation and processing, offering operators a condensed view of any threats to the system. Thirdly, to facilitate the handling of these threats, our causal reasoning system traces the alerts back to their most likely cause, generating an initial hypothesis for operators to investigate.

alert aggregation↗

Real-time Electromagnetic Transient Simulation of Multi-Terminal HVDC-AC Grids based on GPU

High-fidelity electromagnetic transient (EMT) simulation plays a critical role in understanding the dynamic behavior and fast transients involved in operation, control, and protection of Multi-Terminal dc (MTdc) grids. Here, this paper proposes a cost-effective high-performance real-time EMT simulation platform for large-scale cross-continental MTdc grids based on graphics processing unit (GPU). The proposed simulation platform: i) assembles detailed EMT models of all components within an MTdc-ac grid into a single platform. This setup provides a complete simulation solution to capture fast transient signals required for high-bandwidth controller design and protection studies without any compromise; ii) implements the first GPU-based simulation architecture and corresponding algorithms for MTdc-ac grids with real-time performance at scales of 1s; iii) is highly-efficient and balances the high utilization of GPU resources and low latency required for the simulation; and iv) outperforms the existing central processing unit (CPU)- or digital signal processor (DSP)/field-programmable gate array (FPGA)-based simulators in terms of its higher scalability on large-scale MTdc-ac grids and superior price-performance ratio on the hardware. Accuracy and performance of the proposed platform are evaluated with respect to the reference results from PSCAD/EMTDC environment.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Spike-and-Slab Shrinkage Priors for Structurally Sparse Bayesian Neural Networks

Network complexity and computational efficiency have become increasingly significant aspects of deep learning. Sparse deep learning addresses these challenges by recovering a sparse representation of the underlying target function by reducing heavily overparameterized deep neural networks. Specifically, deep neural architectures compressed via structured sparsity (e.g., node sparsity) provide low-latency inference, higher data throughput, and reduced energy consumption. In this article, we explore two well-established shrinkage techniques, Lasso and Horseshoe, for model compression in Bayesian neural networks (BNNs). To this end, we propose structurally sparse BNNs, which systematically prune excessive nodes with the following: 1) spike-and-slab group Lasso (SS-GL) and 2) SS group Horseshoe (SS-GHS) priors, and develop computationally tractable variational inference, including continuous relaxation of Bernoulli variables. We establish the contraction rates of the variational posterior of our proposed models as a function of the network topology, layerwise node cardinalities, and bounds on the network weights. Furthermore, we empirically demonstrate the competitive performance of our models compared with the baseline models in prediction accuracy, model compression, and inference latency.

97 MATHEMATICS AND COMPUTING↗

Cooling and Timing Tests of the ATLAS Fast TracKer VME Boards

The Fast TracKer (FTK) is an ATLAS trigger upgrade built for full-event, low-latency, high-rate tracking. The FTK core, made of 9U VME boards, performs the most demanding computational task. The associative memory board (AMB) serial link processor and the auxiliary card (AUX), plugged on the front and back sides of the same VME slot, constitute the processing unit (PU), which finds tracks using hits from eight layers of the inner detector. The PU works in pipeline with the second stage board (SSB), which finds 12-layer tracks by adding extra hits to the identified tracks. In the designed configuration, 16 PUs and four SSBs are installed in a VME crate. The high power consumption of the AMB, AUX, and SSB (respectively, of about 250, 70, and 160 W per board) required the development of a custom cooling system. Even though the expected power consumption for each VME crate of the FTK system is high compared with a common VME setup, the 8 FTK core crates will use ≈60 kW, which is just a fraction of the power and the space needed for a CPU farm performing the same task. We report on the integration of 32 PUs and eight SSBs inside the FTK system, on the infrastructures needed to run and cool them, and on the tests performed to verify the system processing rate and the temperature stability at a safe value.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

MIC-DP: A Scalable Correlation-Aware Differential Privacy Framework for High-Dimensional Data

Conventional differential privacy (DP) assumes record independence, limiting effectiveness on real-world datasets with temporal, spatial, or structural correlations. These dependencies undermine privacy guarantees and degrade utility in domains like healthcare, IoT, and smart city analytics. We propose Maximum Information Correlated Differential Privacy (MIC-DP), a novel framework that dynamically calibrates noise based on statistical dependencies. MIC-DP uses the Maximum Information Coefficient (MIC) to capture both linear and nonlinear correlations without explicit modeling, enabling adaptive sensitivity adjustment and improved privacy–utility trade-offs. Evaluations on healthcare (MIMIC), demographic (ACI), and synthetic datasets show that MIC-DP reduces mean absolute error (MAE) by up to 5.2% under strict privacy budgets (ϵ≤1), with aggregate utility improvements reaching 18% across datasets and evaluation metrics. MIC-DP provides formal (ϵ,δ)-privacy guarantees, scales efficiently with feature count, and supports deployment in moderate-scale, privacy-sensitive applications. Its tunable performance and runtime efficiency make MIC-DP suitable for privacy-sensitive applications where low-latency analytics and strong privacy guarantees must coexist. These results demonstrate MIC-DP’s effectiveness as a correlation-aware solution for practical DP.

Yang, Wenjun [Univ. of Washington, Tacoma, WA (Uni↗

Scalable Asynchronous Domain Decomposition Solvers

We discuss how parallel implementations of linear iterative solvers generally alternate between phases of data exchange and phases of local computation. Increasingly large problem sizes and more heterogeneous compute architectures make load balancing and the design of low latency network interconnects that are able to satisfy the communication requirements of linear solvers very challenging tasks. In particular, global communication patterns such as inner products become increasingly limiting at scale. We explore the use of asynchronous communication based on one-sided Message Passing Interface primitives in the context of domain decomposition solvers. In particular, a scalable asynchronous two-level Schwarz method is presented. We discuss practical issues encountered in the development of a scalable solver and show experimental results obtained on a state-of-the-art supercomputer system that illustrate the benefits of asynchronous solvers in load balanced as well as load imbalanced scenarios. Using the novel method, we can observe speedups of up to four times over its classical synchronous equivalent.

97 MATHEMATICS AND COMPUTING↗

DDStore: Distributed Data Store for Scalable Training of Graph Neural Networks on Large Atomistic Modeling Datasets

Graph neural networks (GNNs) are a class of Deep Learning models used in designing atomistic materials for effective screening of large chemical spaces. To ensure robust prediction, GNN models must be trained on large volumes of atomistic data on leadership class supercomputers. Even with the advent of modern architectures that consist of multiple storage layers that include node-local NVMe devices in addition to device memory for caching large datasets, extreme-scale model training faces I/O challenges at scale.We present DDStore, an in-memory distributed data store designed for GNN training on large-scale graph data. DDStore provides a hierarchical, distributed, data caching technique that combines data chunking, replication, low-latency random access, and high throughput communication. DDStore achieves near-linear scaling for training a GNN model using up to 1000 GPUs on the Summit and Perlmutter supercomputers, and reaches up to a 6.15x reduction in GNN training time compared to state-of-the-art methodologies.

Choi, Jong Youl↗

Cookie-Jar: An Adaptive Re-configurable Framework for Wireless Network Infrastructures

5G advancements like Massive Multiple Input Multiple Output (MIMO) bring high capacity and low latency, but also intensify interference challenges. Static and dynamic coordination techniques address this, often at the cost of increased power draw. We introduce Cookie-Jar (CJ), an interference coordination (IC) framework using reinforcement learning for multi-goal optimization. By dynamically adjusting network, power, and topology parameters based on real-time conditions, CJ improves Signal to Noise and Interference Ratio (SINR) while minimizing power consumption. Simulated 5G experiments showcase CJ's potential, achieving a 15% SINR improvement with near-identical power draw compared to existing methods.

Network↗

An Intelligent Garbage Sorting System Based on Edge Computing and Visual Understanding of Social Internet of Vehicles

In order to enable Social Internet of Vehicles devices to achieve the purpose of intelligent and autonomous garbage classification in a public environment, while avoiding network congestion caused by a large amount of data accessing the cloud at the same time, it is therefore considered to combine mobile edge computing with Social Internet of Vehicles to give full play to mobile edge computing features of high bandwidth and low latency. At the same time, based on cutting-edge technologies such as deep learning, knowledge graph, and 5G transmission, the paper builds an intelligent garbage sorting system based on edge computing and visual understanding of Social Internet of Vehicles. First of all, for the massive multisource heterogeneous Social Internet of Vehicles big data in the public environment, different item modal data adopts different processing methods, aiming to obtain a visual understanding model. Secondly, using the 5G network, the model is deployed on the edge device and the cloud for cloud-side collaborative management, aiming to avoid the waste of edge node resources, while ensuring the data privacy of the edge node. Finally, the Social Internet of Vehicles devices is used to make intelligent decision-making on the big data of the items. First, the items are judged as garbage, and then the category is judged, and finally the task of grabbing and sorting is realized. The experimental results show that the system proposed in this paper can efficiently process the big data of Social Internet of Vehicles and make valuable intelligent decisions. At the same time, it also has a certain role in promoting the promotion of Social Internet of Vehicles devices.

Shen, Xuehao↗

Preparing MPICH for exascale

The advent of exascale supercomputers heralds a new era of scientific discovery, yet it introduces significant architectural challenges that must be overcome for MPI applications to fully exploit its potential. Among these challenges is the adoption of heterogeneous architectures, particularly the integration of GPUs to accelerate computation. Additionally, the complexity of multithreaded programming models has also become a critical factor in achieving performance at scale. The efficient utilization of hardware acceleration for communication, provided by modern NICs, is also essential for achieving low latency and high throughput communication in such complex systems. In response to these challenges, the MPICH library, a high-performance and widely used Message Passing Interface (MPI) implementation, has undergone significant enhancements. Here, this paper presents four major contributions that prepare MPICH for the exascale transition. First, we describe a lightweight communication stack that leverages the advanced features of modern NICs to maximize hardware acceleration. Second, our work showcases a highly scalable multithreaded communication model that addresses the complexities of concurrent environments. Third, we introduce GPU-aware communication capabilities that optimize data movement in GPU-integrated systems. Finally, we present a new datatype engine aimed at accelerating the use of MPI derived datatypes on GPUs. These improvements in the MPICH library not only address the immediate needs of exascale computing architectures but also set a foundation for exploiting future innovations in high-performance computing. By embracing these new designs and approaches, MPICH-derived libraries from HPE Cray and Intel were able to achieve real exascale performance on OLCF Frontier and ALCF Aurora respectively.

Guo, Yanfei [Argonne National Laboratory (ANL), Ar↗

Laboratory demonstration of the prediction of wind-blown turbulence by adaptive optics at 8 kHz with use of LQG control

The low-latency adaptive optical mirror system (LLAMAS) is designed to push the limits on achievable latencies and frame rates. It has 21 subapertures across its pupil. Here, a reformulated version of the linear quadratic Gaussian (LQG) method predictive Fourier control is implemented in LLAMAS; for all modes, it takes just 30 µs to compute. In the testbed, a turbulator mixes hot and ambient air to produce wind-blown turbulence. Wind prediction clearly improves correction when compared to an integral controller. Closed-loop telemetry shows that wind-predictive LQG removes the characteristic “butterfly” and reduces temporal error power by up to a factor of three for mid-spatial frequency modes. Strehl changes seen in focal plane images are consistent with telemetry and the system error budget.

47 OTHER INSTRUMENTATION↗