Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Fault tolerance”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Low-overhead transversal fault tolerance for universal quantum computation

Fast, reliable logical operations are essential for realizing useful quantum computers. By redundantly encoding logical qubits into many physical qubits and using syndrome measurements to detect and correct errors, we can achieve low logical error rates. However, for many practical quantum error correction codes such as the surface code, owing to syndrome measurement errors, standard constructions require multiple extraction rounds—of the order of the code distance d—for fault-tolerant computation, particularly considering fault-tolerant state preparation. Here we show that logical operations can be performed fault-tolerantly with only a constant number of extraction rounds for a broad class of quantum error correction codes, including the surface code with magic state inputs and feedforward, to achieve ‘transversal algorithmic fault tolerance’. Through the combination of transversal operations7 and new strategies for correlated decoding, despite only having access to partial syndrome information, we prove that the deviation from the ideal logical measurement distribution can be made exponentially small in the distance, even if the instantaneous quantum state cannot be made close to a logical codeword because of measurement errors. We supplement this proof with circuit-level simulations in a range of relevant settings, demonstrating the fault tolerance and competitive performance of our approach. Furthermore, our work sheds new light on the theory of quantum fault tolerance and has the potential to reduce the space–time cost of practical fault-tolerant quantum computation by over an order of magnitude.

Zhou, Hengyun [QuEra Computing, Boston, MA (United

Fault-tolerant operation and materials science with neutral atom logical qubits

We report on the fault-tolerant operation of logical qubits on a neutral atom quantum computer, with logical performance surpassing physical performance for multiple circuits including Bell state preparation (12x error reduction), random circuits (15x), and a prototype Anderson Impurity Model ground state solver for materials science applications (up to 6x, non-fault-tolerantly). The logical qubits are implemented via the [[4, 2, 2]] code (C 4 ). Our work constitutes the first complete realization of the benchmarking protocol proposed by Gottesman 2016 demonstrating results consistent with fault tolerance. In light of recent advances on applying concatenated C 4 /C 6 detection codes to achieve error correction with high code rates and thresholds, our work can be regarded as a building block towards a practical scheme for fault tolerant quantum computation. Our demonstration of a materials science application with logical qubits particularly demonstrates the immediate value of these techniques on current experiments.

36 MATERIALS SCIENCE

Blockchain-based fault tolerant grid operations system

To maintain/improve the distribution system fault-tolerance for a grid with high penetration of distributed energy resources requires improvement, Blockchain can add value to improve fault-tolerant grid operations. That can be achieved using blockchain's core features of distributed consensus-based decision-making process and immutability.

Bereta dos Reis, Fernando

Fault-tolerant optical interconnects for neutral-atom arrays

We analyze the use of photonic links to enable large-scale fault-tolerant connectivity of locally error-corrected modules based on neutral atom arrays. Our approach makes use of recent theoretical results showing the robustness of surface codes to boundary noise and combines recent experimental advances in atom-array quantum computing with logical qubits with optical quantum networking techniques. We find the conditions for fault tolerance can be achieved with local two-qubit Rydberg gate and nonlocal Bell-pair errors below 1% and 10%, respectively, without requiring distillation or space-time overheads. Realizing the interconnects with a lens, a single optical cavity, or an array of cavities enables—with sufficient multiplexing—a Bell-pair generation rate in the 1–50 MHz range. When directly interfacing logical qubits, this rate translates to error-correction cycles in the 25–2000 kHz range, satisfying all requirements for fault tolerance and in the upper range fast enough for 100 kHz logical clock cycles. Published by the American Physical Society 2025

Sinclair, Josiah (ORCID:0000000215238295)

Constant-Overhead Fault-Tolerant Bell-Pair Distillation Using High-Rate Codes

We present a fault-tolerant Bell-pair distillation scheme achieving constant overhead through high-rate quantum low-density parity-check (qLDPC) codes. Our approach maintains a constant distillation rate equal to the code rate while requiring no additional overhead beyond the physical qubits of the code. Full circuit-level analysis demonstrates fault-tolerance for input Bell-pair infidelities below a threshold ∼10%, readily achievable with near-term capabilities. Unlike previous proposals, our scheme keeps the output Bell pairs encoded in qLDPC codes at each node, eliminating unencoding overhead and enabling direct use in distributed quantum applications through recent advances in qLDPC computation. These results establish qLDPC-based distillation as a practical route toward resource-efficient quantum networks and distributed quantum computing.

quantum communication, protocols & technology

A fault-tolerant neutral-atom architecture for universal quantum computation

Quantum error correction (QEC) is essential for the realization of large-scale quantum computers. However, owing to the complexity of operating on the encoded ‘logical’ qubits, understanding the physical principles for building fault-tolerant quantum devices and combining them into efficient architectures is an outstanding scientific challenge. Here we use reconfigurable arrays of up to 448 neutral atoms to implement the key elements of a universal, fault-tolerant quantum processing architecture and experimentally explore their underlying working mechanisms. We first use surface codes to study how repeated QEC suppresses errors, demonstrating 2.14(13)x below-threshold performance in a four-round characterization circuit by leveraging atom loss detection and machine learning decoding. We then investigate logical entanglement using transversal gates and lattice surgery and extend it to universal logic through transversal teleportation with three-dimensional [[15,1,3]] codes, enabling arbitrary-angle synthesis with polylogarithmic overhead. Finally, we develop mid-circuit qubit reuse16, increasing experimental cycle rates by two orders of magnitude and enabling deep-circuit protocols with dozens of logical qubits and hundreds of logical teleportations with [[7,1,3]] and high-rate [[16,6,4]] codes while maintaining constant internal entropy. Our experiments show key principles for efficient architecture design, involving the interplay between quantum logic and entropy removal, judiciously using physical entanglement in logic gates and magic state generation, and leveraging teleportations for universality and physical qubit reset. These results establish foundations for scalable, universal error-corrected processing and its practical implementation in neutral atom systems.

atomic and molecular physics

Fault-Tolerant Deep Learning Cache with Hash Ring for Load Balancing in HPC Systems

Large-scale DL on HPC systems like Frontier and Summit uses distributed node-local caching to address scalability and performance challenges. However, as these systems grow more complex, the risk of node failures increases, and current caching approaches lack fault tolerance, jeopardizing large-scale training jobs. We analyzed six months of SLURM job logs from Frontier and found that over 30% of jobs failed after an average of 75 minutes. To address this, we propose fault-tolerance strategies that recache data lost from failed nodes using a hash ring technique for balanced data recaching in the distributed node-local caching, reducing reliance on the PFS. Our extensive evaluations on Frontier showed that the hash ring-based recaching approach reduced training time by approximately 25% compared to the approach that redirects I/O to the PFS after node failures and demonstrated effective load balancing of training data across nodes.

Lee, Seoyeong

Fault-Tolerant Decentralized Control for Large-Scale Inverter-Based Resources for Active Power Tracking

Integration of inverter-based resources (IBRs) which lack the intrinsic characteristics such as the inertial response of the traditional synchronous-generator (SG)-based sources presents a new challenge in the form of analyzing the grid stability under their presence. While the dynamic composition of IBRs differs from that of the SGs, the control objective remains similar in terms of tracking the desired active power. This letter presents a decentralized primal-dual-based fault-tolerant control framework for the power allocation in IBRs. Overall, a hierarchical control algorithm is developed with a lower level addressing the current control and the parameter estimation for the IBRs and the higher level acting as the reference power generator to the low level based on the desired active power profile. The decentralized network-based algorithm adaptively splits the desired power between the IBRs taking into consideration the health of the IBRs transmission lines. The proposed framework is tested through a simulation on the network of IBRs and the high-level controller performance is compared against the existing framework in the literature. The proposed algorithm shows significant performance improvement in the magnitude of power deviation and settling time to the nominal value under faulty conditions as compared to the algorithm in the literature.

24 POWER TRANSMISSION AND DISTRIBUTION

ATTNChecker: Highly-Optimized Fault Tolerant Attention for Large Language Model Training

Large Language Models (LLMs) have demonstrated remarkable performance in various natural language processing tasks. However, the training of these models is computationally intensive and susceptible to faults, particularly in the attention mechanism, which is a critical component of transformer-based LLMs. In this paper, we investigate the impact of faults on LLM training, focusing on INF, NaN, and near-INF values in the computation results with systematic fault injection experiments. We observe the propagation patterns of these errors, which can trigger non-trainable states in the model and disrupt training, forcing the procedure to load from checkpoints. To mitigate the impact of these faults, we propose ATTNChecker, the first Algorithm-Based Fault Tolerance (ABFT) technique tailored for the attention mechanism in LLMs. ATTNChecker is designed based on fault propagation patterns of LLM and incorporates performance optimization to adapt to both system reliability and model vulnerability while providing lightweight protection for fast LLM training. Evaluations on four LLMs show that ATTNChecker on average incurs on average 7% overhead on training while detecting and correcting all extreme errors. Compared with the state-of-the-art checkpoint/restore approach, ATTNChecker reduces recovery overhead by up to 49×.

Liang, Yuhang [University of Alabama - Birmingham]

Fault Tolerant Decoding of QLDPC-GKP Codes with Circuit Level Soft Information

Concatenated bosonic-stabilizer codes have recently gained prominence as promising candidates for achieving low-overhead fault-tolerant quantum computing in the long term. In such systems, analog information obtained from the syndrome measurements of an inner bosonic code is used to inform decoding for an outer code layer consisting of a discrete-variable stabilizer code such as a surface code. The use of Quantum Low-Density Parity Check (QLDPC) codes as an outer code is of particular interest due to the significantly higher encoding rates offered by these code families, leading to a further reduction in overhead for large-scale quantum computing. Recent works have investigated the performance of QLDPC-GKP codes in detail, and the use of analog information from the inner code significantly boosts decoder performance. However, the noise models assumed in these works are typically limited to depolarizing or phenomenological noise. In this paper, we investigate the performance of QLDPC-GKP concatenated codes under circuit-level noise, based on a model introduced by Noh et al. in the context of the surface-GKP code. To demonstrate the performance boost from analog information, we investigate three scenarios: (a) decoding without soft information, (b) decoding with precomputed error probabilities but without real-time soft information, and (c) decoding with real-time soft information obtained from round-to-round decoding of the inner GKP code. Results show minimal improvement between (a) and (b), but a significant boost in (c), indicating that real-time soft information is critical for concatenated decoding under circuit-level noise. We also study the effect of measurement schedules with varying depths and show that using a schedule with minimum depth is essential for obtaining reliable soft information from the inner code.

Borah, Shantom K. [Arizona U. (main)]

Fault-tolerant resource comparison of qudit and qubit encodings for diagonal quadratic operators

Finite local Hilbert-space truncations arise naturally in quantum simulations of lattice field theories and motivate qudit encodings, but their fault-tolerant advantage over qubit encodings remains unclear. We compare the non-Clifford cost of implementing quadratic diagonal evolutions, exemplified by 𝑈 = 𝑒$^{−𝑖⁢𝑡⁢𝜙^2_𝑥}$ in a uniform field-amplitude discretization of a real scalar field, using either one logical 𝑑-level qudit or 𝑛 𝑏 = ⌈log 2⁡ 𝑑⌉ logical qubits. We analyze two standard settings: product-formula simulation and linear combination of unitaries (LCU) per block encoding, taking the resource metric to be the number of non-Clifford gates after synthesis into a discrete logical gate set. Because tight synthesis bounds for general single-qudit rotations are not known, we express the qudit constructions in terms of embedded two-level SU⁡(2) rotations and derive explicit finite-𝑑 break-even conditions for their synthesis cost; these serve as compiler targets for when qudit encodings can outperform the qubit baseline. Within the constructive models studied here, product-formula implementations would require an exponentially stronger per-primitive synthesis advantage for qudits to win asymptotically, while in the LCU setting the qubit encoding is asymptotically cheaper in 𝑑. Nevertheless, the finite-𝑑 threshold analysis identifies low-dimensional regions in which qudits can yield meaningful constant-factor savings, particularly for LCU-based implementations. As a secondary analysis of the LCU construction, we use an idealized negligible-overhead qubit-qudit code-switching model to give an absolute 𝑇-count comparison and reinterpret the savings as an allowable per-switch overhead budget.

Godwood, Samuel [Univ. of Liverpool (United Kingdo

Leveraging Qubit Loss Detection in Fault-Tolerant Quantum Algorithms

Qubit loss errors constitute a dominant source of noise in many quantum hardware systems, particularly in neutral-atom quantum computers. We develop a theoretical framework to effectively detect and correct loss errors in logical algorithms and leverage such loss information in decoding. Considering general quantum error correction codes and logical circuits, we introduce a delayed-erasure decoder for experimentally motivated error models which leverages information from delayed loss detection to accurately correct loss errors, even when the precise moment of the error is unknown. Using this decoder, we identify strategies for detecting and correcting loss errors based on the logical circuit structure. For deep circuits prior to logical measurement, we explore methods to integrate loss detection into syndrome extraction with minimal overhead, identifying optimal strategies depending on the qubit loss fraction in the noise and hardware capabilities. In contrast, we find that many key algorithmic subroutines involve frequent gate teleportation, shortening the circuit depth before logical measurement and naturally replacing qubits with no additional experimental overhead. We simulate this setting using a toy model algorithm for small-angle synthesis and find a significant performance improvement as the loss fraction increases. These results provide a path forward for advancing large-scale fault-tolerant quantum computation in systems with loss error detection.

atoms

Opportunities in full-stack design of low-overhead fault-tolerant quantum computation

Quantum error correction provides a route to realizing large-scale quantum computation but incurs substantial resource overheads. Here, in this work, we highlight recent advances that reduce these overheads by co-designing different levels of the computational stack, including algorithms, quantum-error-correction strategies and hardware architecture. We then discuss opportunities for further optimization such as leveraging flexible qubit connectivity and quantum low-density parity check codes. These strategies can bring useful quantum computation closer to reality as experiments advance in the coming years.

quantum information

Limitations of Fault-Tolerant Quantum Linear System Solvers for Quantum Power Flow

Quantum computers hold promise for solving problems intractable for classical computers, especially those with high time or space complexity. Practical quantum advantage can be said to exist for such problems when the end-to-end time for solving such a problem using a classical algorithm exceeds that required by a quantum algorithm. Reducing the power flow (PF) problem into a linear system of equations allows for the formulation of quantum PF (QPF) algorithms, which are based on solving methods for quantum linear systems such as the Harrow-Hassidim-Lloyd (HHL) algorithm. Speedup from using QPF algorithms is often claimed to be exponential when compared to classical PF solved by state-of-the-art algorithms. Here, we investigate the potential for practical quantum advantage in solving QPF compared to classical methods on gate-based quantum computers. Notably, this paper does not present a new QPF solving algorithm but scrutinizes the end-to-end complexity of the QPF approach, providing a nuanced evaluation of the purported quantum speedup in this problem. Our analysis establishes a best-case bound for the HHL-based quantum power flow complexity, conclusively demonstrating that the HHL-based method has higher runtime complexity compared to the classical algorithm for solving the direct current power flow (DCPF) and fast decoupled load flow (FDLF) problem. Notably, our analysis and conclusions can be extended to any quantum linear system solver with rigorous performance guarantees, based on the known complexity lower bounds for this problem. Additionally, we establish that for potential practical quantum advantage (PQA) to exist it is necessary to consider DCPF-type problems with a very narrow range of condition number values and readout requirements.

29 ENERGY PLANNING, POLICY, AND ECONOMY

Quantum materials for nanosensing and fault-tolerant quantum computing

New concepts of symmetry related to topological order emerged from the discovery of the fractional quantum Hall effect and high-temperature superconductivity in strongly correlated electron systems. This led to the study of quantum materials-- materials exhibiting emergent quantum phenomena with no classical analogues. While these materials have engendered exciting basic materials science and physics, realizing novel devices is a key challenge in the field. The goal of this proposal is to harness the unique properties of topological materials for quantum computing and quantum sensing applications. In this project, we investigated a variety of topological superconducting platforms and identified three technologies that can benefit from their quantum properties: quantum memory, single-photon detection, and non reciprocal electronics. The platforms developed in this work will be broadly useful to National Security and Basic Science.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND