Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Adaptive Fault Tolerance”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

2nd & 3rd Generation Vehicle Subsystems

This paper contains viewgraph presentation on the "2nd & 3rd Generation Vehicle Subsystems" project. The objective behind this project is to design, develop and test advanced avionics, power systems, power control and distribution components and subsystems for insertion into a highly reliable and low-cost system for a Reusable Launch Vehicles (RLV). The project is divided into two sections: 3rd Generation Vehicle Subsystems and 2nd Generation Vehicle Subsystems. The following topics are discussed under the first section, 3rd Generation Vehicle Subsystems: supporting the NASA RLV program; high-performance guidance & control adaptation for future RLVs; Evolvable Hardware (EHW) for 3rd generation avionics description; Scaleable, Fault-tolerant Intelligent Network or X(trans)ducers (SFINIX); advance electric actuation devices and subsystem technology; hybrid power sources and regeneration technology for electric actuators; and intelligent internal thermal control. Topics discussed in the 2nd Generation Vehicle Subsystems program include: design, development and test of a robust, low-maintenance avionics with no active cooling requirements and autonomous rendezvous and docking systems; design and development of a low maintenance, high reliability, intelligent power systems (fuel cells and battery); and design of a low cost, low maintenance high horsepower actuation systems (actuators).

Source record↗

Fault-Tolerant Operation of Bosonic Qubits with Discrete-Variable Ancillae

Fault-tolerant quantum computation with bosonic qubits often necessitates the use of noisy discrete-variable ancillae. In this work, we establish a comprehensive and practical fault-tolerance framework for such a hybrid system and synthesize it with fault-tolerant protocols by combining bosonic quantum error correction (QEC) and advanced quantum control techniques. We introduce essential building blocks of error-corrected gadgets by leveraging ancilla-assisted bosonic operations using a generalized variant of path-independent quantum control. Using these building blocks, we construct a universal set of error-corrected gadgets that tolerate a single-photon loss and an arbitrary ancilla fault for four-legged cat qubits. Notably, our construction requires only dispersive coupling between bosonic modes and ancillae, as well as beam-splitter coupling between bosonic modes, both of which have been experimentally demonstrated with strong strengths and high accuracy. Moreover, each error-corrected bosonic qubit is comprised of only a single bosonic mode and a three-level ancilla, featuring the hardware efficiency of bosonic QEC in the full fault-tolerant setting. We numerically demonstrate the feasibility of our schemes using current experimental parameters in the circuit-QED platform. Finally, we present a hardware-efficient architecture for fault-tolerant quantum computing by concatenating the four-legged cat qubits with an outer qubit code utilizing only beam-splitter couplings. Our estimates suggest that the overall noise threshold can be reached using existing hardware. These developed fault-tolerant schemes extend beyond their applicability to four-legged cat qubits and can be adapted for other rotation-symmetrical codes, offering a promising avenue toward scalable and robust quantum computation with bosonic qubits. Published by the American Physical Society 2024

Physics↗

Opportunity-adaptive QoS enhancement in satellite constellations: a case study

Systems that are formed by massively distributed mobile resources, such as satellite constellations, often provide mission-critical functions. However, many existing fault tolerance schemes and quality-of-service (QoS) management concepts cannot be applied to those systems in a traditional way, due to the dynamically and continuously changing readiness-to-serve of their mobile resources. In this paper, we describe a case study that investigates a method called opportunity-adaptive QoS enhancement (QAQ).

QAQ algorithm↗

Fault-tolerant grid frequency measurement algorithm during transients

Many critical electric grid operations rely on accurate grid frequency measurements. Unfortunately, the measurement accuracy can be easily undermined by power system transient faults. During a power system transient fault, the power grid voltages and currents are usually highly distorted by high-frequency components. What is worse, the power grid signals could have discontinuity during some system transient faults such as phase angle jump, and the discontinuity could result in large measurement errors to state-of-the-art grid measurement algorithms. In this study, a fault-tolerant grid frequency measurement algorithm during transients is proposed. The new algorithm consists of two stages. The first stage is a transient detector, and it can detect the occurrence of system transient faults instantaneously. The second stage is the intelligent frequency estimator, and it will adapt its measurements according to the transient detector. The performance of the algorithm is evaluated under different steady-state and transient conditions. Both dependability and security of the fault-tolerant algorithm are assessed by using PSCAD simulation data and IEEE Standard test data.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Automated ultrareliability models - A review

Analytic models are required to assess the reliability of systems designed to ultrareliability requirements. This paper reviews the capabilities and limitations of five currently available automated reliability models which are applicable to fault-tolerant flight control systems. 'System' includes sensors, computers, and actuators. A set of review criteria including validation, configuration adaptability, and resource requirements for model evaluation are described. Five models, ARIES, CARE II, CARE III, CARSRA, and CAST, are assessed against the criteria, thereby characterizing their capabilities and limitations. This review should be helpful to potential users of the models.

Bridgman, M. S.↗

Federated Learning for Efficient Condition Monitoring and Anomaly Detection in Industrial Cyber-Physical Systems

Detecting and localizing anomalies in cyber-physical systems (CPS) has become increasingly challenging as systems grow in complexity, particularly due to varying sensor reliability and node failures in distributed environments. While federated learning (FL) offers a foundation for distributed model training, existing approaches lack mechanisms to handle these CPS-specific challenges. This paper presents an enhanced FL framework that introduces three key innovations: adaptive model aggregation based on sensor reliability, dynamic node selection for resource optimization, and Weibull-based checkpointing for fault tolerance. Our framework enables reliable condition monitoring while addressing the computational and reliability challenges of industrial CPS deployments. Experiments on NASA Bearing and Hydraulic System Datasets demonstrate superior performance over state-of-the-art FL methods, achieving 99.5% AUC-ROC in anomaly detection and maintaining accuracy under node failures. Statistical validation using Mann-Whitney (U) test confirms significant improvements (p < 0.05) in both detection accuracy and computational efficiency across diverse operational scenarios.1

Marfo, William [University of Texas at El Paso,Dep↗

Fault diagnosis and fault tolerant control for T-S fuzzy stochastic distribution systems subject to sensor and actuator faults

The problem of fault diagnosis (FD) and fault tolerant control (FTC) for a class of Takagi-Sugeno (T-S) fuzzy stochastic distribution control (SDC) systems subject to sensor and actuator faults is discussed in this paper. First, fuzzy logic models are used to approximate the output probability density function (PDF). Next, an adaptive augmented state/fault diagnosis observer is proposed to estimate the system state, sensor and the actuator faults simultaneously. New expected weights based on the sensor fault estimation information and a PI-type fuzzy feedback fault tolerant (FT) controller are designed to compensate the effect of sensor fault and actuator fault simultaneously. When the sensor fault occurs, the expected objective is redesigned to compensate the sensor fault. Meanwhile, the PI controller can compensate the effect of actuator fault, and the output PDF of the system can still track the desired PDF after the fault occurs. Finally, an example of quality distribution control in chemical reaction process is given to confirm the effectiveness of the algorithm.

42 ENGINEERING↗

5G integrated edge computing platform for efficient component monitoring in coal-fired power plants

This project developed a cutting-edge 5G-integrated edge computing framework to enhance operational efficiency and reliability in coal-fired power plants through real-time component monitoring and anomaly detection. The initiative focused on leveraging distributed machine learning, federated learning, and 5G-based dynamic network slicing to support scalable, fault-tolerant monitoring environments to meet the operational requirements in industrial control systems. With a Distributed Edge Computing Service (DECS) orchestration, this project enabled federated learning at edge for condition monitoring and introduced adaptive client selection strategies to minimize communication overhead. Scalable distributed training was achieved using the Horovod framework, thus enhancing performance across edge nodes. In the realm of 5G networking, the project designed and deployed reconfigurable, QoS-aware network slicing tailored for operational technology (OT) environments, integrating software-defined networks to bolster cyber-resilience and enabling dynamic slicing for federated learning workloads. A significant milestone was the development of a virtualized ICS environment with 5G core integration—which allowed elastic and fault tolerant distributed training on real-world datasets such as NASA Bearings, Hydraulic Systems, and TEP. To broaden the impact of the project, a TRL-3 virtualized ICS testbed for research and education was designed. This project engaged several graduate and undergraduate students to conduct research on the cutting-edge technology, and it resulted in one PhD dissertation, one MS thesis, and over 14 peer-reviewed publications. With the support of this project students also participated in national cybersecurity competitions to improve their professional development skills.

20 FOSSIL-FUELED POWER PLANTS↗

Investigation and Feasibility Assessment of TOPAZ-2 Derivations for Space Power Applications

The ability to provide continuous power at significant levels is of utmost importance for many space missions, from simple satellite operations to manned Mars missions. One of the main problems faced in delivering solar or chemical space power in the tens of kW range, is the increasingly massive nature of the power source and the costs associated with its launch, operation and maintenance. A national program had been initiated to study the feasibility of using certain advanced technologies in developing an efficient lightweight space power source. The starting point for these studies has been the Russian TOPAZ-2 space reactor system, with the ultimate goal to aid in the development of a TOPAZ-2 derivative which will be ready for flight by the year 2000. The main objective of this project has been to perform feasibility assessment and trade studies which would allow the development of an advanced space nuclear power system based on the in-core thermionic fuel element technology currently used in the Russian TOPAZ-2 reactor. Two of the important considerations in developing the concept are: (1) compliance of the current TOPAZ-2 and of any advanced designs with U.S. nuclear safety expectations, and (2) compliance of the design with the seven years lifetime requirement. The project was composed of two major phases. The initial phase of the project has concentrated on understanding the TOPAZ-2 thermionic reactor in sufficient detail to allow several follow-on tasks. The primary interest during this first phase has been given on identifying the potential of the TOPAZ-2 design for further improvements. The second phase of the project has focused on the feasibility of a TOPAZ-2 system capable of delivering 30-50 kWe. Towards the elimination of single-point failures in the load voltage regulation system an active voltage regulator has been designed to be used in conjunction with the available shunt load voltage regulator. The possible use of a dual-loop, model-based adaptive control system for load-following in the TOPAZ-2 has also been investigated. The objective of this fault-tolerant, autonomous control system is to deliver the demanded electric power at the desired voltage level, by appropriately manipulating the neutron power through the control drums. As a result, sufficient thermal power is produced to meet the required demand in the presence of dynamically changing system operating conditions and potential sensor failures. The designed controller is proposed for use in combination with the currently available shunt regulators, or as a back-up controller when other means of power system control, including some of the sensors, fail.

Parlos, Alexander G.↗

Inherent Problems in Designing Two-Failure Tolerant Electromechanical Actuators

An electromechanical ac-powered rotary actuated four-bar linkage system for rotating the Shuttle/Centaur deployment adapter is described. The essential features of the deployment adapter rotation system (DARS) are increased reliability for mission success and maximum practical hazard control for safety. The requirements, concept development, hardware configuration, quality assurance provisions, and techniques used to meet two-fault tolerance requirements are highlighted. The rationale used to achieve a degree of safety equivalent of that of two-failure tolerance is presented. Conditions that make this approach acceptable, including single failure point components with regard to redundancy versus credibility of failure modes, are also discussed.

Hornyak, S.↗

RAID Unbound: Storage Fault Tolerance in a Distributed Environment

Mirroring, data replication, backup, and more recently, redundant arrays of independent disks (RAID) are all technologies used to protect and ensure access to critical company data. A new set of problems has arisen as data becomes more and more geographically distributed. Each of the technologies listed above provides important benefits; but each has failed to adapt fully to the realities of distributed computing. The key to data high availability and protection is to take the technologies' strengths and 'virtualize' them across a distributed network. RAID and mirroring offer high data availability, which data replication and backup provide strong data protection. If we take these concepts at a very granular level (defining user, record, block, file, or directory types) and them liberate them from the physical subsystems with which they have traditionally been associated, we have the opportunity to create a highly scalable network wide storage fault tolerance. The network becomes the virtual storage space in which the traditional concepts of data high availability and protection are implemented without their corresponding physical constraints.

Ritchie, Brian↗

Toggle release

The invention relates to a pyrotechnic actuated release mechanism which is mechanically two fault tolerant for effecting release. It is particularly well suited for releasably connecting structures to be used in the space environment or in other aerospace applications. The device comprises a fastener plate and fastener body, each attachable to either one of a pair of structures to be joined. The fastener plate and the body are fastenable by a toggle supported at one end on the fastener plate and mounted for universal pivotal movement thereon. At its other end, which is received in a central opening in the fastener body and adapted for limited pivotal movement therein, the toggle is restrained by three retractable latching pins. Each pin is individually retractable by combustion of a pyrotechnic charge. While retraction of all three pins releases the toggle, the fastener is mechanically two fault tolerant since the failure of any single or pair of the latch pins to retract results in an asymmetrical loading on the toggle and its pivotal movement to effect a release. An annular bolt is mounted on the fastener plate as a support for the socket mounting of the toggle whereby its selective axial movement provides a means for pre-loading the toggle.

Graves, Thomas Joseph↗

An Architectural Concept for Intrusion Tolerance in Air Traffic Networks

The goal of an intrusion tolerant network is to continue to provide predictable and reliable communication in the presence of a limited num ber of compromised network components. The behavior of a compromised network component ranges from a node that no longer responds to a nod e that is under the control of a malicious entity that is actively tr ying to cause other nodes to fail. Most current data communication ne tworks do not include support for tolerating unconstrained misbehavio r of components in the network. However, the fault tolerance communit y has developed protocols that provide both predictable and reliable communication in the presence of the worst possible behavior of a limited number of nodes in the system. One may view a malicious entity in a communication network as a node that has failed and is behaving in an arbitrary manner. NASA/Langley Research Center has developed one such fault-tolerant computing platform called SPIDER (Scalable Proces sor-Independent Design for Electromagnetic Resilience). The protocols and interconnection mechanisms of SPIDER may be adapted to large-sca le, distributed communication networks such as would be required for future Air Traffic Management systems. The predictability and reliabi lity guarantees provided by the SPIDER protocols have been formally v erified. This analysis can be readily adapted to similar network stru ctures.

Maddalon, Jeffrey M.↗

Towards Precision-Aware Fault Tolerance Approaches for Mixed-Precision Applications

Graphics Processing Units (GPUs), the dominantly adopted accelerators in HPC systems, are susceptible to transient hardware fault. New generation of GPUs feature mixed-precision architectures such as NVIDIA Tensor Cores to accelerate matrix multiplications. While being widely adapted, how would they behave under transient hardware faults remain unclear. In this study, we conduct a large-scale fault injection experiments on GEMM kernels implemented with different floating-point data types on the V100 and A100 Tensor Cores, and show distinct error resilience characteristics for the GEMMS with different formats. In the future, we plan to explore this space by building precision-aware floating-point fault tolerance techniques for applications such as DNNs that exercise low-precision computations.

Fang, Bo↗

NIAC Phase II Orbiting Rainbows: Future Space Imaging with Granular Systems

Inspired by the light scattering and focusing properties of distributed optical assemblies in Nature, such as rainbows and aerosols, and by recent laboratory successes in optical trapping and manipulation, we propose a unique combination of space optics and autonomous robotic system technology, to enable a new vision of space system architecture with applications to ultra-lightweight space optics and, ultimately, in-situ space system fabrication. Typically, the cost of an optical system is driven by the size and mass of the primary aperture. The ideal system is a cloud of spatially disordered dust-like objects that can be optically manipulated: it is highly reconfigurable, fault-tolerant, and allows very large aperture sizes at low cost. This new concept is based on recent understandings in the physics of optical manipulation of small particles in the laboratory and the engineering of distributed ensembles of spacecraft swarms to shape an orbiting cloud of micron-sized objects. In the same way that optical tweezers have revolutionized micro- and nano-manipulation of objects, our breakthrough concept will enable new large scale NASA mission applications and develop new technology in the areas of Astrophysical Imaging Systems and Remote Sensing because the cloud can operate as an adaptive optical imaging sensor. While achieving the feasibility of constructing one single aperture out of the cloud is the main topic of this work, it is clear that multiple orbiting aerosol lenses could also combine their power to synthesize a much larger aperture in space to enable challenging goals such as exo-planet detection. Furthermore, this effort could establish feasibility of key issues related to material properties, remote manipulation, and autonomy characteristics of cloud in orbit. There are several types of endeavors (science missions) that could be enabled by this type of approach, i.e. it can enable new astrophysical imaging systems, exo-planet search, large apertures allow for unprecedented high resolution to discern continents and important features of other planets, hyperspectral imaging, adaptive systems, spectroscopy imaging through limb, and stable optical systems from Lagrange-points. Furthermore, future micro-miniaturization might hold promise of a further extension of our dust aperture concept to other more exciting smart dust concepts with other associated capabilities. Our objective in Phase II was to experimentally and numerically investigate how to optically manipulate and maintain the shape of an orbiting cloud of dust-like matter so that it can function as an adaptable ultra-lightweight surface. Our solution is based on the aperture being an engineered granular medium, instead of a conventional monolithic aperture. This allows building of apertures at a reduced cost, enables extremely fault-tolerant apertures that cannot otherwise be made, and directly enables classes of missions for exoplanet detection based on Fourier spectroscopy with tight angular resolution and innovative radar systems for remote sensing. In this task, we have examined the advanced feasibility of a crosscutting concept that contributes new technological approaches for space imaging systems, autonomous systems, and space applications of optical manipulation. The proposed investigation has matured the concept that we started in Phase I to TRL 3, identifying technology gaps and candidate system architectures for the space-borne cloud as an aperture.

Quadrelli, Marco B.↗

ATTNChecker: Highly-Optimized Fault Tolerant Attention for Large Language Model Training

Large Language Models (LLMs) have demonstrated remarkable performance in various natural language processing tasks. However, the training of these models is computationally intensive and susceptible to faults, particularly in the attention mechanism, which is a critical component of transformer-based LLMs. In this paper, we investigate the impact of faults on LLM training, focusing on INF, NaN, and near-INF values in the computation results with systematic fault injection experiments. We observe the propagation patterns of these errors, which can trigger non-trainable states in the model and disrupt training, forcing the procedure to load from checkpoints. To mitigate the impact of these faults, we propose ATTNChecker, the first Algorithm-Based Fault Tolerance (ABFT) technique tailored for the attention mechanism in LLMs. ATTNChecker is designed based on fault propagation patterns of LLM and incorporates performance optimization to adapt to both system reliability and model vulnerability while providing lightweight protection for fast LLM training. Evaluations on four LLMs show that ATTNChecker on average incurs on average 7% overhead on training while detecting and correcting all extreme errors. Compared with the state-of-the-art checkpoint/restore approach, ATTNChecker reduces recovery overhead by up to 49×.

Liang, Yuhang [University of Alabama - Birmingham]↗

Orchestrating Fault Prediction with Live Migration and Checkpointing

Checkpoint/Restart (C/R) is widely used to provide fault tolerance on High-Performance Computing (HPC) systems. However, Parallel File System (PFS) overhead and failure uncertainty cause significant application overhead. This paper develops an adaptive multi-level C/R model that incorporates a failure prediction and analysis model, which orchestrates failure prediction, checkpointing, checkpoint frequency, and proactive live migration along with the additional benefit of Burst Buffers (BB). It effectively reduces the overheads due to failures, checkpointing, and recovery. Simulation results for the Summit supercomputer yield a reduction of ~20%-86% in application overhead due to BBs, orchestrated failure prediction, and migration. We also observe a ~29% decrease in checkpoint writes to BBs, which can increase the longevity of the BB storage devices.

Behera, Subhendu↗

Reliability of voting in fault-tolerant software systems for small output spaces

Under a voting strategy in a fault-tolerant software system there is a difference between correctness and agreement. An independent N-version programming reliability model is proposed for treating small output spaces which distinguishes between correctness and agreement. System reliability is investigated using analytical relationships and simulation. A consensus majority voting strategy is proposed and its performance is analyzed and compared with other voting strategies. Consensus majority strategy automatically adapts the voting to different component reliability and output space cardinality characteristics. It is shown that absolute majority voting strategy provides a lower bound on the reliability provided by the consensus majority, and 2-of-n voting strategy an upper bound. If r is the cardinality of the output space it is proved the 1/r is a lower bound on the average reliability of fault-tolerant system components below which the system reliability begins to deteriorate as more versions are added.

Mcallister, David F.↗