Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “fault tolerance challenges”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Diagnosing a Failed Proof in Fault-Tolerance: A Disproving Challenge Problem

This paper proposes a challenge problem in disproving. We describe a fault-tolerant distributed protocol designed at NASA for use in a fly-by-wire system for next-generation commercial aircraft. An early design of the protocol contains a subtle bug that is highly unlikely to be caught in fault injection testing. We describe a failed proof of the protocol's correctness in a mechanical theorem prover (PVS) with a complex unfinished proof conjecture. We use a model checking suite (SAL) to generate a concrete counterexample to the unproven conjecture to demonstrate the existence of a bug. However, we argue that the effort required in our approach is too high and propose what conditions a better solution would satisfy. We carefully describe the protocol and bug to provide a challenging but feasible case study for disproving research.

Pike, Lee↗

A cryogenic muon tagging system based on kinetic inductance detectors for superconducting quantum processors

Ionizing radiation has emerged as a potential limiting factor for superconducting quantum processors, inducing quasiparticle bursts and correlated errors that challenge fault-tolerant operation. Atmospheric muons are particularly problematic due to their high energy and penetration power, making passive shielding ineffective. Therefore, monitoring the real-time muon flux is crucial to guide the development of alternative error-correction or mitigation strategies. We present the design, simulation, and first operation of a cryogenic muon-tagging system based on kinetic inductance detectors (KIDs), developed as a stand-alone cryogenic particle-tagging module for superconducting quantum processors. The system consists of two KIDs arranged in a vertical stack and operated at ∼20 mK. Monte Carlo simulations based on Geant4 guided the prototype design and provided reference expectations for muon-tagging efficiency and accidental coincidences due to ambient γ-rays. We observed a muon-induced coincidence rate among the top and bottom detectors of (192 ± 9) $\times\,10^{-3}$ events s$^{−1}$, in excellent agreement with the Monte Carlo prediction. The prototype achieves a muon-tagging efficiency of about 90% with negligible dead time. These results demonstrate the feasibility of operating a muon-tagging system at millikelvin temperatures and represent a key step toward the integration of cryogenic veto systems with multi-qubit chips to mitigate muon-induced errors.

Mariani, Ambra [INFN, Rome] (ORCID:000000028184857↗

Real-Time Detection of Charge Jumps in Superconducting Qubits with a Convolutional Neural Network

Ionizing radiation from cosmic rays and gammas can induce discontinuous jumps in the environmental charge of superconducting qubits (charge jumps), causing correlated errors that challenge fault-tolerant quantum computing while simultaneously providing a detection signature for quantum sensing applications. Current detection methods operate offline, introducing latency incompatible with in-the-loop qubit control. In this paper, an online detector of charge jumps for superconducting qubits, based on a dilated causal convolutional neural network (DCCNN) designed for in-the-loop deployment on the Quantum Instrumentation Control Kit (QICK) platform, is presented. The network is trained on synthetic Ramsey tomography scans generated from qubit templates measured at the Northwestern Experimental Underground Site (NEXUS) at Fermilab, and translated to FPGA firmware via hls4ml with ap_fixed$\langle 16,6 \rangle$ quantization, reaching a per-inference latency of $6.19 μ$s on the Zynq UltraScale+ RFSoC ZCU216. At this operating point the DCCNN matches the detection efficiency of the established offline $χ^2$ algorithm ($0.843 \pm 0.022$ vs. $0.866 \pm 0.020$ on $|Δq| \in [0.1, 0.5] e$ at matched false-positive rate), while requiring no per-qubit hyperparameter tuning. This shifts charge-jump detection from a post-hoc diagnostic to a control-loop primitive, enabling adaptive protocols that respond to radiation-induced events in situ, with applications to quantum-computing error mitigation and to the use of superconducting qubits as particle detectors.

Gaytan-Villarreal, Daniel [Carnegie Mellon U.]↗

Computer aided reliability, availability, and safety modeling for fault-tolerant computer systems with commentary on the HARP program

Many of the most challenging reliability problems of our present decade involve complex distributed systems such as interconnected telephone switching computers, air traffic control centers, aircraft and space vehicles, and local area and wide area computer networks. In addition to the challenge of complexity, modern fault-tolerant computer systems require very high levels of reliability, e.g., avionic computers with MTTF goals of one billion hours. Most analysts find that it is too difficult to model such complex systems without computer aided design programs. In response to this need, NASA has developed a suite of computer aided reliability modeling programs beginning with CARE 3 and including a group of new programs such as: HARP, HARP-PC, Reliability Analysts Workbench (Combination of model solvers SURE, STEM, PAWS, and common front-end model ASSIST), and the Fault Tree Compiler. The HARP program is studied and how well the user can model systems using this program is investigated. One of the important objectives will be to study how user friendly this program is, e.g., how easy it is to model the system, provide the input information, and interpret the results. The experiences of the author and his graduate students who used HARP in two graduate courses are described. Some brief comparisons were made with the ARIES program which the students also used. Theoretical studies of the modeling techniques used in HARP are also included. Of course no answer can be any more accurate than the fidelity of the model, thus an Appendix is included which discusses modeling accuracy. A broad viewpoint is taken and all problems which occurred in the use of HARP are discussed. Such problems include: computer system problems, installation manual problems, user manual problems, program inconsistencies, program limitations, confusing notation, long run times, accuracy problems, etc.

Shooman, Martin L.↗

Fault-Tolerant Software-Defined Radio on Manycore

Software-defined radio (SDR) platforms generally rely on field-programmable gate arrays (FPGAs) and digital signal processors (DSPs), but such architectures require significant software development. In addition, application demands for radiation mitigation and fault tolerance exacerbate programming challenges. MaXentric Technologies, LLC, has developed a manycore-based SDR technology that provides 100 times the throughput of conventional radiationhardened general purpose processors. Manycore systems (30-100 cores and beyond) have the potential to provide high processing performance at error rates that are equivalent to current space-deployed uniprocessor systems. MaXentric's innovation is a highly flexible radio, providing over-the-air reconfiguration; adaptability; and uninterrupted, real-time, multimode operation. The technology is also compliant with NASA's Space Telecommunications Radio System (STRS) architecture. In addition to its many uses within NASA communications, the SDR can also serve as a highly programmable research-stage prototyping device for new waveforms and other communications technologies. It can also support noncommunication codes on its multicore processor, collocated with the communications workload-reducing the size, weight, and power of the overall system by aggregating processing jobs to a single board computer.

Ricketts, Scott↗

Evolution of shuttle avionics redundancy management/fault tolerance

The challenge of providing redundancy management (RM) and fault tolerance to meet the Shuttle Program requirements of fail operational/fail safe for the avionics systems was complicated by the critical program constraints of weight, cost, and schedule. The basic and sometimes false effectivity of less than pure RM designs is addressed. Evolution of the multiple input selection filter (the heart of the RM function) is discussed with emphasis on the subtle interactions of the flight control system that were found to be potentially catastrophic. Several other general RM development problems are discussed, with particular emphasis on the inertial measurement unit RM, indicative of the complexity of managing that three string system and its critical interfaces with the guidance and control systems.

Boykin, J. C.↗

Offset Charge Dependence of Measurement-Induced Transitions in Transmons

A key challenge in achieving scalable fault tolerance in superconducting quantum processors is readout fidelity, which lags behind one- and two-qubit gate fidelity. A major limitation in improving qubit readout is measurement-induced transitions, also referred to as qubit ionization, caused by multiphoton qubit-resonator excitation occurring at specific photon numbers. Since ionization can involve highly excited states, it has been predicted that in transmons—the most widely used superconducting qubit—the photon number at which measurement-induced transitions occur is gate-charge dependent. This dependence is expected to persist deep in the transmon regime where the qubit frequency is gate-charge insensitive. We experimentally confirm this prediction by characterizing measurement-induced transitions with increasing resonator photon population while actively calibrating the transmon’s gate charge. Furthermore, because highly excited states are involved, achieving quantitative agreement between theory and experiment requires accounting for higher-order harmonics in the transmon Hamiltonian.

Quantum circuits↗

Reconfigurable robots for all terrain exploration

While significant recent progress has been made in development of mobile robots for planetary suface exploration,there remain major challenges. These include increased autonomy of operation, traverse of challenging terrain, and fault-tolerance under long, unattended periods of use.

mobile robots multi-robot cooperation robotic arch↗

Fault Tolerance in ZigBee Wireless Sensor Networks

Wireless sensor networks (WSN) based on the IEEE 802.15.4 Personal Area Network standard are finding increasing use in the home automation and emerging smart energy markets. The network and application layers, based on the ZigBee 2007 PRO Standard, provide a convenient framework for component-based software that supports customer solutions from multiple vendors. This technology is supported by System-on-a-Chip solutions, resulting in extremely small and low-power nodes. The Wireless Connections in Space Project addresses the aerospace flight domain for both flight-critical and non-critical avionics. WSNs provide the inherent fault tolerance required for aerospace applications utilizing such technology. The team from Ames Research Center has developed techniques for assessing the fault tolerance of ZigBee WSNs challenged by radio frequency (RF) interference or WSN node failure.

Alena, Richard↗

Fault Mitigation Schemes for Future Spaceflight Multicore Processors

Future planetary exploration missions demand significant advances in on-board computing capabilities over current avionics architectures based on a single-core processing element. The state-of-the-art multi-core processor provides much promise in meeting such challenges while introducing new fault tolerance problems when applied to space missions. Software-based schemes are being presented in this paper that can achieve system-level fault mitigation beyond that provided by radiation-hard-by-design (RHBD). For mission and time critical applications such as the Terrain Relative Navigation (TRN) for planetary or small body navigation, and landing, a range of fault tolerance methods can be adapted by the application. The software methods being investigated include Error Correction Code (ECC) for data packet routing between cores, virtual network routing, Triple Modular Redundancy (TMR), and Algorithm-Based Fault Tolerance (ABFT). A robust fault tolerance framework that provides fail-operational behavior under hard real-time constraints and graceful degradation will be demonstrated using TRN executing on a commercial Tilera(R) processor with simulated fault injections.

software based↗

Power and Limitations of Linear Programming Decoder for Quantum LDPC Codes

Decoding quantum error-correcting codes is a key challenge in enabling fault-tolerant quantum computation. In the classical setting, linear programming (LP) decoders offer provable performance guarantees and can leverage fast practical optimization algorithms. Although LP decoders have been proposed for quantum codes, their performance and limitations remain relatively underexplored. In this work, we uncover a key limitation of LP decoding for quantum low-density parity-check (LDPC) codes: certain constant-weight error patterns lead to ambiguous fractional solutions that cannot be resolved through independent rounding. To address this issue, we incorporate a post-processing technique known as ordered statistics decoding (OSD), which significantly enhances LP decoding performance in practice. Our results show that LP decoding, when augmented with OSD, can outperform belief propagation with the same post-processing for intermediate code sizes of up to hundreds of qubits. These findings suggest that LP-based decoders, equipped with effective post-processing, offer a promising approach for decoding near-term quantum LDPC codes.

Gu, Shouzhen [Yale U.]↗

Power and Limitations of Linear Programming Decoder for Quantum LDPC Codes

Decoding quantum error-correcting codes is a key challenge in enabling fault-tolerant quantum computation. In the classical setting, linear programming (LP) decoders offer provable performance guarantees and can leverage fast practical optimization algorithms. Although LP decoders have been proposed for quantum codes, their performance and limitations remain relatively underexplored. In this work, we uncover a key limitation of LP decoding for quantum low-density parity-check (LDPC) codes: certain constant-weight error patterns lead to ambiguous fractional solutions that cannot be resolved through independent rounding. To address this issue, we incorporate a post-processing technique known as ordered statistics decoding (OSD), which significantly enhances LP decoding performance in practice. Our results show that LP decoding, when augmented with OSD, can outperform belief propagation with the same post-processing for intermediate code sizes of up to hundreds of qubits. These findings suggest that LP-based decoders, equipped with effective post-processing, offer a promising approach for decoding near-term quantum LDPC codes.

Gu, Shouzhen [Yale U.]↗

Power and Limitations of Linear Programming Decoder for Quantum LDPC Codes

Decoding quantum error-correcting codes is a key challenge in enabling fault-tolerant quantum computation. In the classical setting, linear programming (LP) decoders offer provable performance guarantees and can leverage fast practical optimization algorithms. Although LP decoders have been proposed for quantum codes, their performance and limitations remain relatively underexplored. In this work, we uncover a key limitation of LP decoding for quantum low-density parity-check (LDPC) codes: certain constant-weight error patterns lead to ambiguous fractional solutions that cannot be resolved through independent rounding. To address this issue, we incorporate a post-processing technique known as ordered statistics decoding (OSD), which significantly enhances LP decoding performance in practice. Our results show that LP decoding, when augmented with OSD, can outperform belief propagation with the same post-processing for intermediate code sizes of up to hundreds of qubits. These findings suggest that LP-based decoders, equipped with effective post-processing, offer a promising approach for decoding near-term quantum LDPC codes.

Gu, Shouzhen [Yale U.]↗

System on a Chip (SoC) Overview

System-on-a-chip or system on chip (SoC or SOC) refers to integrating all components of a computer or other electronic system into a single integrated circuit (chip). It may contain digital, analog, mixed-signal, and often radio-frequency functions all on a single chip substrate. Complexity drives it all: Radiation tolerance and testability are challenges for fault isolation, propagation, and validation. Bigger single silicon die than flown before and technology is scaling below 90nm (new qual methods). Packages have changed and are bigger and more difficult to inspect, test, and understand. Add in embedded passives. Material interfaces are more complex (underfills, processing). New rules for board layouts. Mechanical and thermal designs, etc.

LaBel, Kenneth A.↗

A Blueprint for Demonstrating Quantum Supremacy with Superconducting Qubits

Long coherence times and high fidelity control recently achieved in scalable superconducting circuits paved the way for the growing number of experimental studies of many-qubit quantum coherent phenomena in these devices. Albeit full implementation of quantum error correction and fault tolerant quantum computation remains a challenge the near term pre-error correction devices could allow new fundamental experiments despite inevitable accumulation of errors. One such open question foundational for quantum computing is achieving the so called quantum supremacy, an experimental demonstration of a computational task that takes polynomial time on the quantum computer whereas the best classical algorithm would require exponential time and/or resources. It is possible to formulate such a task for a quantum computer consisting of less than a 100 qubits. The computational task we consider is to provide approximate samples from a non-trivial quantum distribution. This is a generalization for the case of superconducting circuits of ideas behind boson sampling protocol for quantum optics introduced by Arkhipov and Aaronson. In this presentation we discuss a proof-of-principle demonstration of such a sampling task on a 9-qubit chain of superconducting gmon qubits developed by Google. We discuss theoretical analysis of the driven evolution of the device resulting in output approximating samples from a uniform distribution in the Hilbert space, a quantum chaotic state. We analyze quantum chaotic characteristics of the output of the circuit and the time required to generate a sufficiently complex quantum distribution. We demonstrate that the classical simulation of the sampling output requires exponential resources by connecting the task of calculating the output amplitudes to the sign problem of the Quantum Monte Carlo method. We also discuss the detailed theoretical modeling required to achieve high fidelity control and calibration of the multi-qubit unitary evolution in the device. We use a novel cross-entropy statistical metric as a figure of merit to verify the output and calibrate the device controls. Finally, we demonstrate the statistics of the wave function amplitudes generated on the 9-gmon chain and verify the quantum chaotic nature of the generated quantum distribution. This verifies the implementation of the quantum supremacy protocol.

Kechedzhi, Kostyantyn↗

Disentangling the impact of quasiparticles and two-level systems on the statistics of superconducting-qubit lifetime

Temporal fluctuations in the superconducting qubit lifetime, T 1 , present additional challenges in the pursuit of fault-tolerant quantum computing. Although the exact mechanisms remain unclear, T 1 fluctuations are generally attributed to strong coupling between the qubit and a few near-resonant two-level systems (TLSs), which can exchange energy with an ensemble of thermally fluctuating two-level fluctuators (TLFs) at low frequencies. Here, we report T 1 measurements of qubits with varying geometrical footprints and surface dielectrics as a function of temperature. By analyzing the noise spectrum of the qubit depolarization rate, Γ 1 = 1 / T 1 , we disentangle the contributions of TLSs, nonequilibrium quasiparticles (QPs), and equilibrium (thermally excited) QPs to the variance in Γ 1 . We find that the Γ 1 variance in qubits with smaller footprints is more susceptible to QP and TLS fluctuations than that in larger-footprint qubits. Furthermore, the QP-induced variances in all qubits align with the theoretical framework of QP diffusion and fluctuation. These findings offer valuable insights for future qubit design and engineering optimization.

Zhu, Shaojiang [Fermilab] (ORCID:0000000293180092)↗

LLRF System Analysis for the Fermilab PIP-II LINAC

Developing long-lived quantum processing units (QPUs) capable of supporting high-fidelity quantum operations is a crucial challenge on the path toward fault-tolerant quantum computing. TESLA-shaped superconducting RF (SRF) cavities, known for photon relaxation times on the order of seconds, provide an excellent foundation for 3D QPUs and quantum memory. This talk presents a novel design that leverages TESLA cavity modes coupled to ancillary transmon qubits, optimized to preserve coherence and control. By carefully engineering the package geometry, optimizing Hamiltonian parameters, and minimizing lossy participation ratios, we achieve photon relaxation times of over 16 ms and 20 ms for the two cavity modes, representing a significant improvement over previous multimode quantum memories. Despite the reduced coupling between the qubit and cavity modes, which is necessary to preserve long lifetimes, the platform supports robust and universal control schemes that are not limited by low coupling strength. We will also discuss how this architecture can lead to scalable, modular quantum computing systems.

Varghese, P. [Fermilab]↗