Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Adaptive Fault Tolerance”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Peer-to-Peer Communication Trade-Offs for Smart Grid Applications: Preprint

Peer-to-peer energy management systems for smart grids require developers to consider the trade-offs between the amount of communication traffic generated and the quality and speed of convergence of the control algorithms that are deployed. Employing a fully connected communication causes messages to scale exponentially with the number of nodes, while using a sparse connectivity causes less information dissemination leading to degradation of the algorithm performance. The best communication topology for a particular application lies somewhere in between and often requires empirical evaluation by application designers. Existing methods do not put focus on the needs for smart grid applications, which is information dissemination throughout the network and they do not provide a flexible solution for application developers to prototype and deploy different topologies without modifying the application code. This paper introduces a configurable virtual communication topology framework TopLinkMgr, allowing users to specify any chosen communication topology and deploy peer-to-peer applications using it. It also introduces a self-adaptive, fault-tolerant topology management algorithm, Bounded Path Dissemination that can ensure the dissemination of information to all peers within a specified threshold for a sparsely connected topology. Experiments show that the algorithm improves on convergence speed and accuracy over state-of-the-art methods and is also robust against node failures. The results indicate the possibility of achieving a close-to optimal convergence without overloading the network allowing the realization of peer-to-peer control platforms covering larger and more complex power systems.

Bounded Path Dissemination↗

Model Predictive Fault-Tolerant Tracking Control for PDF Control Systems With Packet Losses

In this article, a fault-tolerant tracking control strategy is investigated for nonlinear probability density function (PDF) control systems with the actuator fault, uncertainties, unknown disturbance, and random packet losses. The control input signal dropout and measurement signal dropouts are described as the independent Bernoulli distribution. An adaptive fault diagnosis (FD) observer based on the Lyapunov function is given to simultaneously estimate the fault, disturbance, and state with packet losses. Furthermore, different from the traditional robust fault-tolerant control (FTC), a new active fault-tolerant tracking controller is designed based on the model predictive control framework, which has better adaptive fault-tolerant performance. Finally, the validity of the proposed FTC method has been proved by a simulation study of a papermaking process.

42 ENGINEERING↗

Feedback-Based Fault-Tolerant and Health-Adaptive Optimal Charging of Batteries

The key technology barriers that hinder the growth of Electric Vehicles (EVs) are long charging time, the shorter life-time of EV batteries, and battery safety. Specifically, EV charging protocols have significant effects on battery lifetime and safety. If not charged properly, the battery could end up with shorter life, and more importantly, improper charging can cause battery faults leading to catastrophic failures. To overcome these barriers, we propose a closed-loop feedback based approach, that enables real-time optimal fast charging protocol adaptation to battery health and possess active diagnostic capabilities in the sense that, during charging, it detects real-time faults and takes corrective action to mitigate such fault effects. We utilize battery electrical-thermal model, explicit battery capacity and power fade aging models, and thermal fault model to capture battery behavior. In conjunction with the models, we adopt linear quadratic optimal control techniques to realize the feedback-based control algorithm. Simulation studies are presented to illustrate the effectiveness of the proposed scheme.

batteries↗

Efficient Client Selection in Federated Learning

Federated Learning (FL) enables decentralized machine learning while preserving data privacy. This paper proposes a novel client selection framework that integrates differential privacy and fault tolerance. The adaptive client selection adjusts the number of clients based on performance and system constraints, with noise added to protect privacy. Evaluated on the UNSW-NB15 and ROAD datasets for network anomaly detection, the method improves accuracy by 7% and reduces training time by 25 % compared to baselines. Fault tolerance enhances robustness with minimal performance trade-offs.

Marfo, William [University of Texas at El Paso,Dep↗

Adaptive Client Selection in Federated Learning: A Network Anomaly Detection Use Case

Federated Learning (FL) has become a ubiquitous approach for training machine learning models on decentralized data, addressing the myriad privacy concerns inherent in traditional centralized methods. However, the efficiency of FL depends on effective client selection and robust privacy preservation mechanisms. Inadequate client selection may lead to suboptimal model performance, while insufficient privacy measures risk exposing sensitive data. This paper proposes a client selection framework for FL that integrates differential privacy and fault tolerance. Our adaptive approach dynamically adjusts the number of selected clients based on model performance and system constraints, ensuring privacy through calibrated noise addition. We evaluate our method on a network anomaly detection use case using the UNSW-NB15 and ROAD datasets. Results show up to a 7% increase in accuracy and a 25% reduction in training time compared to FedL2P. Moreover, we highlight the trade-offs between privacy budgets and model performance, with higher privacy budgets reducing noise and improving accuracy. Our fault tolerance mechanism, while causing a slight performance drop, enhances robustness to client failures. Statistical validation using Mann-Whitney U tests confirms the significance of these improvements (p < 0.05).

Marfo, William [University of Texas at El Paso,Dep↗

Current possibilities and future opportunities for erasure coded computations

The key capability established through the research funded by this award are erasure coded computations for linear systems, in serial and in parallel. This capability enables powerful efficient and scalable alternatives to existing linear system solvers in fault-prone computational systems.

97 MATHEMATICS AND COMPUTING↗

JANUS: Resilient and Adaptive Data Transmission for Enabling Timely and Efficient Cross-Facility Scientific Workflows

In modern science, the growing complexity of large-scale scientific projects has led to an increasing reliance on cross-facility scientific workflows, where resources and expertise from multiple institutions and geographic locations are leveraged to accelerate scientific discovery. These workflows often require transmitting huge amounts of scientific data through wide-area networks. Although high-speed networks like ESnet and transfer services such as Globus have improved data mobility, several challenges remain. The sheer volume of data can overwhelm network bandwidth, widely used transport protocols such as TCP suffer from inefficiencies due to retransmissions triggered by packet loss, and existing fault-tolerance mechanisms like erasure coding introduce substantial overhead. In this paper, we propose Janus, a resilient and adaptable data transmission approach designed for cross-facility scientific workflows. Unlike traditional TCP-based methods, Janus leverages UDP, integrates erasure coding for fault tolerance, and combines it with error-bounded lossy compression to reduce overhead. This novel design allows users to balance data transmission time and accuracy, optimizing transfer performance based on specific scientific requirements. Additionally, Janus dynamically adjusts erasure coding parameters in response to real-time network conditions, ensuring efficient data transfers even in fluctuating environments. We develop optimization models for determining ideal configurations and implement adaptive data transfer protocols to enhance reliability. Through extensive simulations and real-network experiments, we demonstrate that Janus significantly improves transfer efficiency while maintaining data fidelity.

Esaulov, Vladislav [Georgia State University, Atla↗

Operational Evolution of FTS3: A DevOps Driven Approach to Elastic Operations

The File Transfer Service (FTS3) is a distributed data movement service developed at CERN and widely used to transfer data across the Worldwide LHC Computing Grid (WLCG). At Fermilab, FTS3 supports data transfers for multiple experiments, including Intensity Frontier experiments such as DUNE, enabling reliable data movement between WebDAV endpoints in Europe and the Americas.​ At CHEP 2021, we reported on the initial containerized deployment of FTS3 on OKD, the community Kubernetes distribution of Red Hat OpenShift. In this work, we present the subsequent evolution of this deployment, focusing on new operational capabilities introduced to improve scalability, robustness, and long-term maintainability.​ We describe the adoption of more secure and reproducible container build workflows, the integration of DevOps-driven operational practices, and enhancements in monitoring and automation. A key new result is the introduction of horizontal scaling and elastic resource management, allowing FTS3 components to dynamically adapt to workload variations while maintaining service reliability. We also discuss improvements in fault tolerance and operational procedures derived from production experience.​ Finally, we summarize lessons learned from operating FTS3 as a Kubernetes-native service and outline how these developments have improved the resilience and efficiency of data movement operations at Fermilab.

Munoz Flores, Victor Leopoldo [Fermilab]↗

Fault-Tolerant Operation of Bosonic Qubits with Discrete-Variable Ancillae

Fault-tolerant quantum computation with bosonic qubits often necessitates the use of noisy discrete-variable ancillae. In this work, we establish a comprehensive and practical fault-tolerance framework for such a hybrid system and synthesize it with fault-tolerant protocols by combining bosonic quantum error correction (QEC) and advanced quantum control techniques. We introduce essential building blocks of error-corrected gadgets by leveraging ancilla-assisted bosonic operations using a generalized variant of path-independent quantum control. Using these building blocks, we construct a universal set of error-corrected gadgets that tolerate a single-photon loss and an arbitrary ancilla fault for four-legged cat qubits. Notably, our construction requires only dispersive coupling between bosonic modes and ancillae, as well as beam-splitter coupling between bosonic modes, both of which have been experimentally demonstrated with strong strengths and high accuracy. Moreover, each error-corrected bosonic qubit is comprised of only a single bosonic mode and a three-level ancilla, featuring the hardware efficiency of bosonic QEC in the full fault-tolerant setting. We numerically demonstrate the feasibility of our schemes using current experimental parameters in the circuit-QED platform. Finally, we present a hardware-efficient architecture for fault-tolerant quantum computing by concatenating the four-legged cat qubits with an outer qubit code utilizing only beam-splitter couplings. Our estimates suggest that the overall noise threshold can be reached using existing hardware. These developed fault-tolerant schemes extend beyond their applicability to four-legged cat qubits and can be adapted for other rotation-symmetrical codes, offering a promising avenue toward scalable and robust quantum computation with bosonic qubits. Published by the American Physical Society 2024

Physics↗

Fault-tolerant grid frequency measurement algorithm during transients

Many critical electric grid operations rely on accurate grid frequency measurements. Unfortunately, the measurement accuracy can be easily undermined by power system transient faults. During a power system transient fault, the power grid voltages and currents are usually highly distorted by high-frequency components. What is worse, the power grid signals could have discontinuity during some system transient faults such as phase angle jump, and the discontinuity could result in large measurement errors to state-of-the-art grid measurement algorithms. In this study, a fault-tolerant grid frequency measurement algorithm during transients is proposed. The new algorithm consists of two stages. The first stage is a transient detector, and it can detect the occurrence of system transient faults instantaneously. The second stage is the intelligent frequency estimator, and it will adapt its measurements according to the transient detector. The performance of the algorithm is evaluated under different steady-state and transient conditions. Both dependability and security of the fault-tolerant algorithm are assessed by using PSCAD simulation data and IEEE Standard test data.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Federated Learning for Efficient Condition Monitoring and Anomaly Detection in Industrial Cyber-Physical Systems

Detecting and localizing anomalies in cyber-physical systems (CPS) has become increasingly challenging as systems grow in complexity, particularly due to varying sensor reliability and node failures in distributed environments. While federated learning (FL) offers a foundation for distributed model training, existing approaches lack mechanisms to handle these CPS-specific challenges. This paper presents an enhanced FL framework that introduces three key innovations: adaptive model aggregation based on sensor reliability, dynamic node selection for resource optimization, and Weibull-based checkpointing for fault tolerance. Our framework enables reliable condition monitoring while addressing the computational and reliability challenges of industrial CPS deployments. Experiments on NASA Bearing and Hydraulic System Datasets demonstrate superior performance over state-of-the-art FL methods, achieving 99.5% AUC-ROC in anomaly detection and maintaining accuracy under node failures. Statistical validation using Mann-Whitney (U) test confirms significant improvements (p < 0.05) in both detection accuracy and computational efficiency across diverse operational scenarios.1

Marfo, William [University of Texas at El Paso,Dep↗

Fault diagnosis and fault tolerant control for T-S fuzzy stochastic distribution systems subject to sensor and actuator faults

The problem of fault diagnosis (FD) and fault tolerant control (FTC) for a class of Takagi-Sugeno (T-S) fuzzy stochastic distribution control (SDC) systems subject to sensor and actuator faults is discussed in this paper. First, fuzzy logic models are used to approximate the output probability density function (PDF). Next, an adaptive augmented state/fault diagnosis observer is proposed to estimate the system state, sensor and the actuator faults simultaneously. New expected weights based on the sensor fault estimation information and a PI-type fuzzy feedback fault tolerant (FT) controller are designed to compensate the effect of sensor fault and actuator fault simultaneously. When the sensor fault occurs, the expected objective is redesigned to compensate the sensor fault. Meanwhile, the PI controller can compensate the effect of actuator fault, and the output PDF of the system can still track the desired PDF after the fault occurs. Finally, an example of quality distribution control in chemical reaction process is given to confirm the effectiveness of the algorithm.

42 ENGINEERING↗

5G integrated edge computing platform for efficient component monitoring in coal-fired power plants

This project developed a cutting-edge 5G-integrated edge computing framework to enhance operational efficiency and reliability in coal-fired power plants through real-time component monitoring and anomaly detection. The initiative focused on leveraging distributed machine learning, federated learning, and 5G-based dynamic network slicing to support scalable, fault-tolerant monitoring environments to meet the operational requirements in industrial control systems. With a Distributed Edge Computing Service (DECS) orchestration, this project enabled federated learning at edge for condition monitoring and introduced adaptive client selection strategies to minimize communication overhead. Scalable distributed training was achieved using the Horovod framework, thus enhancing performance across edge nodes. In the realm of 5G networking, the project designed and deployed reconfigurable, QoS-aware network slicing tailored for operational technology (OT) environments, integrating software-defined networks to bolster cyber-resilience and enabling dynamic slicing for federated learning workloads. A significant milestone was the development of a virtualized ICS environment with 5G core integration—which allowed elastic and fault tolerant distributed training on real-world datasets such as NASA Bearings, Hydraulic Systems, and TEP. To broaden the impact of the project, a TRL-3 virtualized ICS testbed for research and education was designed. This project engaged several graduate and undergraduate students to conduct research on the cutting-edge technology, and it resulted in one PhD dissertation, one MS thesis, and over 14 peer-reviewed publications. With the support of this project students also participated in national cybersecurity competitions to improve their professional development skills.

20 FOSSIL-FUELED POWER PLANTS↗

Towards Precision-Aware Fault Tolerance Approaches for Mixed-Precision Applications

Graphics Processing Units (GPUs), the dominantly adopted accelerators in HPC systems, are susceptible to transient hardware fault. New generation of GPUs feature mixed-precision architectures such as NVIDIA Tensor Cores to accelerate matrix multiplications. While being widely adapted, how would they behave under transient hardware faults remain unclear. In this study, we conduct a large-scale fault injection experiments on GEMM kernels implemented with different floating-point data types on the V100 and A100 Tensor Cores, and show distinct error resilience characteristics for the GEMMS with different formats. In the future, we plan to explore this space by building precision-aware floating-point fault tolerance techniques for applications such as DNNs that exercise low-precision computations.

Fang, Bo↗

ATTNChecker: Highly-Optimized Fault Tolerant Attention for Large Language Model Training

Large Language Models (LLMs) have demonstrated remarkable performance in various natural language processing tasks. However, the training of these models is computationally intensive and susceptible to faults, particularly in the attention mechanism, which is a critical component of transformer-based LLMs. In this paper, we investigate the impact of faults on LLM training, focusing on INF, NaN, and near-INF values in the computation results with systematic fault injection experiments. We observe the propagation patterns of these errors, which can trigger non-trainable states in the model and disrupt training, forcing the procedure to load from checkpoints. To mitigate the impact of these faults, we propose ATTNChecker, the first Algorithm-Based Fault Tolerance (ABFT) technique tailored for the attention mechanism in LLMs. ATTNChecker is designed based on fault propagation patterns of LLM and incorporates performance optimization to adapt to both system reliability and model vulnerability while providing lightweight protection for fast LLM training. Evaluations on four LLMs show that ATTNChecker on average incurs on average 7% overhead on training while detecting and correcting all extreme errors. Compared with the state-of-the-art checkpoint/restore approach, ATTNChecker reduces recovery overhead by up to 49×.

Liang, Yuhang [University of Alabama - Birmingham]↗

Orchestrating Fault Prediction with Live Migration and Checkpointing

Checkpoint/Restart (C/R) is widely used to provide fault tolerance on High-Performance Computing (HPC) systems. However, Parallel File System (PFS) overhead and failure uncertainty cause significant application overhead. This paper develops an adaptive multi-level C/R model that incorporates a failure prediction and analysis model, which orchestrates failure prediction, checkpointing, checkpoint frequency, and proactive live migration along with the additional benefit of Burst Buffers (BB). It effectively reduces the overheads due to failures, checkpointing, and recovery. Simulation results for the Summit supercomputer yield a reduction of ~20%-86% in application overhead due to BBs, orchestrated failure prediction, and migration. We also observe a ~29% decrease in checkpoint writes to BBs, which can increase the longevity of the BB storage devices.

Behera, Subhendu↗

Open Source Fault-tolerant Grid Frequency Measurement for Solar Inverters

The Discrete Fourier transform (DFT) based measurement algorithms are one of the most common measurement algorithms for grid parameter estimation such as rms, phase angle, frequency. Over the past few years, many DFT based algorithms have been developed to enhance its measurement accuracy under steady-state and/or dynamic grid conditions. For example, an adaptive band-pass filter utilizing exponential modulation filter has been proposed to reduce measurement errors at the presence of large frequency deviations. Measurement accuracy of different algorithms including FIR filter, extended Kalman filtering (EKF), and enhanced DFT method have been compared in detail under different grid conditions. Two artificial signals that have 90-degree phase difference were constructed by the Clarke transformation to address the frequency spectrum leakage of DFT. A multi-module approach was developed to enhance both steady-state and dynamic measurement accuracies, in which each module was developed to eliminate some specific errors. Besides DFT-based measurement algorithms, some signal model-based algorithms have been developed to further improve the accuracy under dynamic conditions. However, a key drawback of the state-of-the-art algorithms is that they cannot perform measurements accurately during system transient faults. In the Blue Cut Fire event, there was a phase angle jump of about 26 degrees in the voltage waveform during the transient fault. The phase angle jump fault will cause waveform discontinuity, and these algorithms will fail to provide reliable measurements during this period because they typically assume the waveform to be measured is continuous, no matter what method (DFT, PLL, EKF, FIR, or Taylor WLS) is used for estimation. In fact, the measurement errors during the system transient faults like phase-jump is not required in the IEEE Standard. As a result, although a measurement instrument can pass the strict IEEE Standard, it could still be the source of the problem in the future if we have similar system transient faults, which could happen again. Therefore, developing the fault-tolerant measurement technology is the key to solve the problem.

14 SOLAR ENERGY↗

Quantum computation of stopping power for inertial fusion target design

Stopping power is the rate at which a material absorbs the kinetic energy of a charged particle passing through it—one of many properties needed over a wide range of thermodynamic conditions in modeling inertial fusion implosions. First-principles stopping calculations are classically challenging because they involve the dynamics of large electronic systems far from equilibrium, with accuracies that are particularly difficult to constrain and assess in the warm-dense conditions preceding ignition. Here, we describe a protocol for using a fault-tolerant quantum computer to calculate stopping power from a first-quantized representation of the electrons and projectile. Our approach builds upon the electronic structure block encodings of Su et al. [ PRX Quant. 2 , 040332 (2021)], adapting and optimizing those algorithms to estimate observables of interest from the non-Born–Oppenheimer dynamics of multiple particle species at finite temperature. We also work out the constant factors associated with an implementation of a high-order Trotter approach to simulating a grid representation of these systems. Ultimately, we report logical qubit requirements and leading-order Toffoli costs for computing the stopping power of various projectile/target combinations relevant to interpreting and designing inertial fusion experiments. We estimate that scientifically interesting and classically intractable stopping power calculations can be quantum simulated with roughly the same number of logical qubits and about one hundred times more Toffoli gates than is required for state-of-the-art quantum simulations of industrially relevant molecules such as FeMoco or P450.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗