Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “fault tolerance challenges”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

A cryogenic muon tagging system based on kinetic inductance detectors for superconducting quantum processors

Ionizing radiation has emerged as a potential limiting factor for superconducting quantum processors, inducing quasiparticle bursts and correlated errors that challenge fault-tolerant operation. Atmospheric muons are particularly problematic due to their high energy and penetration power, making passive shielding ineffective. Therefore, monitoring the real-time muon flux is crucial to guide the development of alternative error-correction or mitigation strategies. We present the design, simulation, and first operation of a cryogenic muon-tagging system based on kinetic inductance detectors (KIDs), developed as a stand-alone cryogenic particle-tagging module for superconducting quantum processors. The system consists of two KIDs arranged in a vertical stack and operated at ∼20 mK. Monte Carlo simulations based on Geant4 guided the prototype design and provided reference expectations for muon-tagging efficiency and accidental coincidences due to ambient γ-rays. We observed a muon-induced coincidence rate among the top and bottom detectors of (192 ± 9) $\times\,10^{-3}$ events s$^{−1}$, in excellent agreement with the Monte Carlo prediction. The prototype achieves a muon-tagging efficiency of about 90% with negligible dead time. These results demonstrate the feasibility of operating a muon-tagging system at millikelvin temperatures and represent a key step toward the integration of cryogenic veto systems with multi-qubit chips to mitigate muon-induced errors.

Mariani, Ambra [INFN, Rome] (ORCID:000000028184857↗

Real-Time Detection of Charge Jumps in Superconducting Qubits with a Convolutional Neural Network

Ionizing radiation from cosmic rays and gammas can induce discontinuous jumps in the environmental charge of superconducting qubits (charge jumps), causing correlated errors that challenge fault-tolerant quantum computing while simultaneously providing a detection signature for quantum sensing applications. Current detection methods operate offline, introducing latency incompatible with in-the-loop qubit control. In this paper, an online detector of charge jumps for superconducting qubits, based on a dilated causal convolutional neural network (DCCNN) designed for in-the-loop deployment on the Quantum Instrumentation Control Kit (QICK) platform, is presented. The network is trained on synthetic Ramsey tomography scans generated from qubit templates measured at the Northwestern Experimental Underground Site (NEXUS) at Fermilab, and translated to FPGA firmware via hls4ml with ap_fixed$\langle 16,6 \rangle$ quantization, reaching a per-inference latency of $6.19 μ$s on the Zynq UltraScale+ RFSoC ZCU216. At this operating point the DCCNN matches the detection efficiency of the established offline $χ^2$ algorithm ($0.843 \pm 0.022$ vs. $0.866 \pm 0.020$ on $|Δq| \in [0.1, 0.5] e$ at matched false-positive rate), while requiring no per-qubit hyperparameter tuning. This shifts charge-jump detection from a post-hoc diagnostic to a control-loop primitive, enabling adaptive protocols that respond to radiation-induced events in situ, with applications to quantum-computing error mitigation and to the use of superconducting qubits as particle detectors.

Gaytan-Villarreal, Daniel [Carnegie Mellon U.]↗

Offset Charge Dependence of Measurement-Induced Transitions in Transmons

A key challenge in achieving scalable fault tolerance in superconducting quantum processors is readout fidelity, which lags behind one- and two-qubit gate fidelity. A major limitation in improving qubit readout is measurement-induced transitions, also referred to as qubit ionization, caused by multiphoton qubit-resonator excitation occurring at specific photon numbers. Since ionization can involve highly excited states, it has been predicted that in transmons—the most widely used superconducting qubit—the photon number at which measurement-induced transitions occur is gate-charge dependent. This dependence is expected to persist deep in the transmon regime where the qubit frequency is gate-charge insensitive. We experimentally confirm this prediction by characterizing measurement-induced transitions with increasing resonator photon population while actively calibrating the transmon’s gate charge. Furthermore, because highly excited states are involved, achieving quantitative agreement between theory and experiment requires accounting for higher-order harmonics in the transmon Hamiltonian.

Quantum circuits↗

Power and Limitations of Linear Programming Decoder for Quantum LDPC Codes

Decoding quantum error-correcting codes is a key challenge in enabling fault-tolerant quantum computation. In the classical setting, linear programming (LP) decoders offer provable performance guarantees and can leverage fast practical optimization algorithms. Although LP decoders have been proposed for quantum codes, their performance and limitations remain relatively underexplored. In this work, we uncover a key limitation of LP decoding for quantum low-density parity-check (LDPC) codes: certain constant-weight error patterns lead to ambiguous fractional solutions that cannot be resolved through independent rounding. To address this issue, we incorporate a post-processing technique known as ordered statistics decoding (OSD), which significantly enhances LP decoding performance in practice. Our results show that LP decoding, when augmented with OSD, can outperform belief propagation with the same post-processing for intermediate code sizes of up to hundreds of qubits. These findings suggest that LP-based decoders, equipped with effective post-processing, offer a promising approach for decoding near-term quantum LDPC codes.

Gu, Shouzhen [Yale U.]↗

Power and Limitations of Linear Programming Decoder for Quantum LDPC Codes

Decoding quantum error-correcting codes is a key challenge in enabling fault-tolerant quantum computation. In the classical setting, linear programming (LP) decoders offer provable performance guarantees and can leverage fast practical optimization algorithms. Although LP decoders have been proposed for quantum codes, their performance and limitations remain relatively underexplored. In this work, we uncover a key limitation of LP decoding for quantum low-density parity-check (LDPC) codes: certain constant-weight error patterns lead to ambiguous fractional solutions that cannot be resolved through independent rounding. To address this issue, we incorporate a post-processing technique known as ordered statistics decoding (OSD), which significantly enhances LP decoding performance in practice. Our results show that LP decoding, when augmented with OSD, can outperform belief propagation with the same post-processing for intermediate code sizes of up to hundreds of qubits. These findings suggest that LP-based decoders, equipped with effective post-processing, offer a promising approach for decoding near-term quantum LDPC codes.

Gu, Shouzhen [Yale U.]↗

Power and Limitations of Linear Programming Decoder for Quantum LDPC Codes

Decoding quantum error-correcting codes is a key challenge in enabling fault-tolerant quantum computation. In the classical setting, linear programming (LP) decoders offer provable performance guarantees and can leverage fast practical optimization algorithms. Although LP decoders have been proposed for quantum codes, their performance and limitations remain relatively underexplored. In this work, we uncover a key limitation of LP decoding for quantum low-density parity-check (LDPC) codes: certain constant-weight error patterns lead to ambiguous fractional solutions that cannot be resolved through independent rounding. To address this issue, we incorporate a post-processing technique known as ordered statistics decoding (OSD), which significantly enhances LP decoding performance in practice. Our results show that LP decoding, when augmented with OSD, can outperform belief propagation with the same post-processing for intermediate code sizes of up to hundreds of qubits. These findings suggest that LP-based decoders, equipped with effective post-processing, offer a promising approach for decoding near-term quantum LDPC codes.

Gu, Shouzhen [Yale U.]↗

Disentangling the impact of quasiparticles and two-level systems on the statistics of superconducting-qubit lifetime

Temporal fluctuations in the superconducting qubit lifetime, T 1 , present additional challenges in the pursuit of fault-tolerant quantum computing. Although the exact mechanisms remain unclear, T 1 fluctuations are generally attributed to strong coupling between the qubit and a few near-resonant two-level systems (TLSs), which can exchange energy with an ensemble of thermally fluctuating two-level fluctuators (TLFs) at low frequencies. Here, we report T 1 measurements of qubits with varying geometrical footprints and surface dielectrics as a function of temperature. By analyzing the noise spectrum of the qubit depolarization rate, Γ 1 = 1 / T 1 , we disentangle the contributions of TLSs, nonequilibrium quasiparticles (QPs), and equilibrium (thermally excited) QPs to the variance in Γ 1 . We find that the Γ 1 variance in qubits with smaller footprints is more susceptible to QP and TLS fluctuations than that in larger-footprint qubits. Furthermore, the QP-induced variances in all qubits align with the theoretical framework of QP diffusion and fluctuation. These findings offer valuable insights for future qubit design and engineering optimization.

Zhu, Shaojiang [Fermilab] (ORCID:0000000293180092)↗

LLRF System Analysis for the Fermilab PIP-II LINAC

Developing long-lived quantum processing units (QPUs) capable of supporting high-fidelity quantum operations is a crucial challenge on the path toward fault-tolerant quantum computing. TESLA-shaped superconducting RF (SRF) cavities, known for photon relaxation times on the order of seconds, provide an excellent foundation for 3D QPUs and quantum memory. This talk presents a novel design that leverages TESLA cavity modes coupled to ancillary transmon qubits, optimized to preserve coherence and control. By carefully engineering the package geometry, optimizing Hamiltonian parameters, and minimizing lossy participation ratios, we achieve photon relaxation times of over 16 ms and 20 ms for the two cavity modes, representing a significant improvement over previous multimode quantum memories. Despite the reduced coupling between the qubit and cavity modes, which is necessary to preserve long lifetimes, the platform supports robust and universal control schemes that are not limited by low coupling strength. We will also discuss how this architecture can lead to scalable, modular quantum computing systems.

Varghese, P. [Fermilab]↗

Measuring quasiparticle dynamics for particle impact reconstruction in a superconducting qubit chip

Quasiparticle poisoning following particle impacts poses a significant challenge to the development of fault-tolerant superconducting quantum computers, as a sudden excess of quasiparticles can simultaneously degrade the coherence of multiple qubits across large device arrays. In this work, we present a statistical analysis that models the time evolution of radiation-induced qubit energy relaxation through quasiparticle density dynamics. This study provides insight into quasiparticle loss processes by distinguishing between recombination and trapping decay channels and assessing their respective impact on qubit performance. We precisely measure quasiparticle recombination in multiple transmon qubits and uncover an unexpected dependence of qubit relaxation dynamics on deposited energy. By linking correlated relaxation events across qubits to ballistic phonon propagation, we introduce a statistical localization approach to extract the energy deposited in the substrate, which is in good agreement with Monte Carlo simulation. This work establishes the quantitative framework for using an arbitrary subset of superconducting transmon qubits in a QPU as energy-resolving witness particle detectors.

Celi, E. [Northwestern U.]↗

A fault-tolerant neutral-atom architecture for universal quantum computation

Quantum error correction (QEC) is essential for the realization of large-scale quantum computers. However, owing to the complexity of operating on the encoded ‘logical’ qubits, understanding the physical principles for building fault-tolerant quantum devices and combining them into efficient architectures is an outstanding scientific challenge. Here we use reconfigurable arrays of up to 448 neutral atoms to implement the key elements of a universal, fault-tolerant quantum processing architecture and experimentally explore their underlying working mechanisms. We first use surface codes to study how repeated QEC suppresses errors, demonstrating 2.14(13)x below-threshold performance in a four-round characterization circuit by leveraging atom loss detection and machine learning decoding. We then investigate logical entanglement using transversal gates and lattice surgery and extend it to universal logic through transversal teleportation with three-dimensional [[15,1,3]] codes, enabling arbitrary-angle synthesis with polylogarithmic overhead. Finally, we develop mid-circuit qubit reuse16, increasing experimental cycle rates by two orders of magnitude and enabling deep-circuit protocols with dozens of logical qubits and hundreds of logical teleportations with [[7,1,3]] and high-rate [[16,6,4]] codes while maintaining constant internal entropy. Our experiments show key principles for efficient architecture design, involving the interplay between quantum logic and entropy removal, judiciously using physical entanglement in logic gates and magic state generation, and leveraging teleportations for universality and physical qubit reset. These results establish foundations for scalable, universal error-corrected processing and its practical implementation in neutral atom systems.

atomic and molecular physics↗

Fault-Tolerant Deep Learning Cache with Hash Ring for Load Balancing in HPC Systems

Large-scale DL on HPC systems like Frontier and Summit uses distributed node-local caching to address scalability and performance challenges. However, as these systems grow more complex, the risk of node failures increases, and current caching approaches lack fault tolerance, jeopardizing large-scale training jobs. We analyzed six months of SLURM job logs from Frontier and found that over 30% of jobs failed after an average of 75 minutes. To address this, we propose fault-tolerance strategies that recache data lost from failed nodes using a hash ring technique for balanced data recaching in the distributed node-local caching, reducing reliance on the PFS. Our extensive evaluations on Frontier showed that the hash ring-based recaching approach reduced training time by approximately 25% compared to the approach that redirects I/O to the PFS after node failures and demonstrated effective load balancing of training data across nodes.

Lee, Seoyeong↗

QECC-Synth: A Layout Synthesizer for Quantum Error Correction Codes on Sparse Architectures

Quantum Error Correction (QEC) codes are essential for achieving fault-tolerant quantum computing (FTQC). However, their implementation faces significant challenges due to disparity between required dense qubit connectivity and sparse hardware architectures. Current approaches often either underutilize QEC circuit features or focus on manual designs tailored to specific codes and architectures, limiting their capability and generality. In response, we introduce QECC-Synth, an automated compiler for QEC code implementation that addresses these challenges. We leverage the ancilla bridge technique tailored to the requirements of QEC circuits and introduces a systematic classification of its design space flexibilities. We then formalize this problem using the MaxSAT framework to optimize these flexibilities. Evaluation shows that our method significantly outperforms existing methods while demonstrating broader applicability across diverse QEC codes and hardware architectures.

Yin, Keyi [University of California, San Diego]↗

Federated Learning for Efficient Condition Monitoring and Anomaly Detection in Industrial Cyber-Physical Systems

Detecting and localizing anomalies in cyber-physical systems (CPS) has become increasingly challenging as systems grow in complexity, particularly due to varying sensor reliability and node failures in distributed environments. While federated learning (FL) offers a foundation for distributed model training, existing approaches lack mechanisms to handle these CPS-specific challenges. This paper presents an enhanced FL framework that introduces three key innovations: adaptive model aggregation based on sensor reliability, dynamic node selection for resource optimization, and Weibull-based checkpointing for fault tolerance. Our framework enables reliable condition monitoring while addressing the computational and reliability challenges of industrial CPS deployments. Experiments on NASA Bearing and Hydraulic System Datasets demonstrate superior performance over state-of-the-art FL methods, achieving 99.5% AUC-ROC in anomaly detection and maintaining accuracy under node failures. Statistical validation using Mann-Whitney (U) test confirms significant improvements (p < 0.05) in both detection accuracy and computational efficiency across diverse operational scenarios.1

Marfo, William [University of Texas at El Paso,Dep↗

First-Principles Assessment of ZnTe and CdSe as Prospective Tunnel Barriers at the InAs/Al Interface

Majorana zero modes are predicted to emerge in semiconductor/ superconductor interfaces, such as InAs/Al. Majorana modes could be utilized for fault tolerant topological qubits. However, their realization is hindered by materials challenges. The coupling between the superconductor and the semiconductor may be too strong for Majorana modes to emerge, due to effective doping of the semiconductor by the metallic contact. This could be mediated by adding a tunnel barrier of controlled thickness. We use density functional theory (DFT) with Hubbard U corrections, whose values are machine-learned via Bayesian optimization (BO), to assess ZnTe and CdSe as prospective tunnel barriers for the InAs/Al interface. The results of DFT +U(BO) for ZnTe are validated by comparison to angle resolved photoemission spectroscopy (ARPES). We then study bilayer interfaces of the three semiconductors with each other and with Al, as well as trilayer interfaces with a varying number of ZnTe or CdSe layers inserted between InAs and Al. We find that 16 atomic layers of either material completely insulate the InAs from metal induced gap states (MIGS). However, ZnTe and CdSe differ significantly in their band alignment, such that ZnTe forms an effective barrier for electrons, whereas CdSe forms a barrier for holes. Because of Fermi level pinning in the conduction band at the interface, only electron transport is relevant for InAs-based Majorana devices. Therefore, ZnTe is the better choice. Based on the results of our simulations, we suggest conducting experiments with ZnTe barriers in the thickness range of 6–18 atomic layers.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Scalable Circuit Cutting and Scheduling in a Resource-constrained and Distributed Quantum System

Despite quantum computing's rapid development, current systems remain limited in practical applications due to their limited qubit count and quality. Various technologies, such as superconducting, trapped ions, and neutral atom quantum computing technologies are progressing towards a fault tolerant era, however they all face a diverse set of challenges in scalability and control. Recent efforts have focused on multi-node quantum systems that connect multiple smaller quantum devices to execute larger circuits. Future demonstrations hope to use quantum channels to couple systems, however current demonstrations can leverage classical communication with circuit cutting techniques. This involves cutting large circuits into smaller subcircuits and reconstructing them post-execution. However, existing cutting methods are hindered by lengthy search times as the number of qubits and gates increases. Additionally, they often fail to effectively utilize the resources of various worker configurations in a multi-node system. To address these challenges, we introduce FitCut, a novel approach that transforms quantum circuits into weighted graphs and utilizes a community-based, bottom-up approach to cut circuits according to resource constraints, e.g., qubit counts, on each worker. FitCut also includes a scheduling algorithm that optimizes resource utilization across workers. Implemented with Qiskit and evaluated extensively, FitCut significantly outperforms the Qiskit Circuit Knitting Toolbox, reducing time costs by factors ranging from 3 to 2000 and improving resource utilization rates by up to 3.88 times on the worker side, achieving a system-wide improvement of 2.86 times.

Kan, Shuwen [Fordham University]↗

Quantum Zeno Monte Carlo for computing observables

The recent development of logical quantum processors marks a pivotal transition from the noisy intermediate-scale quantum (NISQ) era to the fault-tolerant quantum computing (FTQC) era. These devices have the potential to address classically challenging problems with polynomial computational time using quantum properties. However, they remain susceptible to noise, necessitating noise resilient algorithms. We introduce Quantum Zeno Monte Carlo (QZMC), a classical-quantum hybrid algorithm that demonstrates resilience to device noise and Trotter errors while showing polynomial computational cost for a gapped system. QZMC computes static and dynamic properties without requiring initial state overlap or variational parameters, offering reduced quantum circuit depth.

Han, Mancheon [Korea Institute for Advanced Study ↗

Fault-Tolerant Decentralized Control for Large-Scale Inverter-Based Resources for Active Power Tracking

Integration of inverter-based resources (IBRs) which lack the intrinsic characteristics such as the inertial response of the traditional synchronous-generator (SG)-based sources presents a new challenge in the form of analyzing the grid stability under their presence. While the dynamic composition of IBRs differs from that of the SGs, the control objective remains similar in terms of tracking the desired active power. This letter presents a decentralized primal-dual-based fault-tolerant control framework for the power allocation in IBRs. Overall, a hierarchical control algorithm is developed with a lower level addressing the current control and the parameter estimation for the IBRs and the higher level acting as the reference power generator to the low level based on the desired active power profile. The decentralized network-based algorithm adaptively splits the desired power between the IBRs taking into consideration the health of the IBRs transmission lines. The proposed framework is tested through a simulation on the network of IBRs and the high-level controller performance is compared against the existing framework in the literature. The proposed algorithm shows significant performance improvement in the magnitude of power deviation and settling time to the nominal value under faulty conditions as compared to the algorithm in the literature.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Exploiting Kubernetes to Simplify the Deployment and Management of the Multi-purpose CMS Pilot Job Factory

GlideinWMS, a widely utilized workload management system in high-energy physics (HEP) research, serves as the backbone for efficient job provisioning across distributed computing resources. It is utilized by various experiments and organizations, including CMS, OSG, Dune, and FIFE, to create HTCondor pools as large as 600k cores. In particular, a shared factory service historically deployed at UCSD has been configured to interface with more than 500 routes to compute clusters. As part of our team’s initiative to modernize infrastructure and enhance scalability, we undertook the migration of the GlideinWMS factory service into the Kubernetes environment. Leveraging the flexibility and orchestration capabilities of Kubernetes, we successfully deployed the factory service within the OSG Tiger Kubernetes cluster. The major benefits Kubernetes gives us is it streamlines the management and monitoring of the factory infrastructure, and improves fault tolerance through its resilient deployment strategies. Through this case study, we aim to share insights, challenges, and best practices encountered during the migration process. Our experience underscores the benefits of embracing containerization and Kubernetes orchestration for HEP computing infrastructure, paving the way for scalability and resilience in distributed computing environments.

Dost, Jeffrey Michael [UC, San Diego (main)]↗