Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Error detection and correction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Sensor-Assisted Fault Mitigation in Quantum Computation

We propose a method for assisting fault mitigation in quantum computation through the use of sensors, co-located near physical qubits. Specifically, we consider using transition edge sensors co-located on silicon substrates hosting superconducting qubits to monitor for energy injection from ionizing radiation that has been demonstrated to increase decoherence in transmon qubits. We generalize from these two physical device concepts and explore the potential advantages of co-located sensors to assist fault mitigation in quantum computation. In the simplest scheme, co-located sensors beneficially assist rejection of calculations potentially affected by environmental disturbances. Investigating the potential computational advantage further required development of an extension to the standard formulation of quantum error correction. In a specific case the standard three-qubit, bit flip quantum correction code, we show that given a 20% overall error probability per qubit, approximately 90% of repeated calculation attempts are correctable. However, when sensor-detectable errors account for 45% of overall error probability, the use of colocated sensors uniquely associated with independent qubits, boosts the fraction of correct final-state calculations to 96% at the cost of rejecting 7% of repeated calculation attempts.

Orrell, John L.↗

Self-Checking Memory Interface

Memory-interface integrated circuit not only detects errors in data from other circuits but also detects errors within itself. Memory-interface chip encodes 16-bit words with Hamming code for single-error correction or double-error detection. Chip used in fault-tolerant computers under development by NASA.

Sievers, M. W.↗

Preparing and Analyzing Iced Airfoils

SmaggIce version 1.2 is a computer program for preparing and analyzing iced airfoils. It includes interactive tools for (1) measuring ice-shape characteristics, (2) controlled smoothing of ice shapes, (3) curve discretization, (4) generation of artificial ice shapes, and (5) detection and correction of input errors. Measurements of ice shapes are essential for establishing relationships between characteristics of ice and effects of ice on airfoil performance. The shape-smoothing tool helps prepare ice shapes for use with already available grid-generation and computational-fluid-dynamics software for studying the aerodynamic effects of smoothed ice on airfoils. The artificial ice-shape generation tool supports parametric studies since ice-shape parameters can easily be controlled with the artificial ice. In such studies, artificial shapes generated by this program can supplement simulated ice obtained from icing research tunnels and real ice obtained from flight test under icing weather condition. SmaggIce also automatically detects geometry errors such as tangles or duplicate points in the boundary which may be introduced by digitization and provides tools to correct these. By use of interactive tools included in SmaggIce version 1.2, one can easily characterize ice shapes and prepare iced airfoils for grid generation and flow simulations.

Vickerman, Mary B.↗

Performance of concatenated Reed-Solomon trellis-coded modulation over Rician fading channels

A concatenated coding scheme for providing very reliable data over mobile-satellite channels at power levels similar to those used for vocoded speech is described. The outer code is a shorter Reed-Solomon code which provides error detection as well as error correction capabilities. The inner code is a 1-D 8-state trellis code applied independently to both the inphase and quadrature channels. To achieve the full error correction potential of this inner code, the code symbols are multiplexed with a pilot sequence which is used to provide dynamic channel estimation and coherent detection. The implementation structure of this scheme is discussed and its performance is estimated.

Moher, Michael L.↗

Erasure conversion in a high-fidelity Rydberg quantum simulator

Minimizing and understanding errors is critical for quantum science, both in noisy intermediate scale quantum (NISQ) devices and for the quest towards fault-tolerant quantum computation. Rydberg arrays have emerged as a prominent platform in this context with impressive system sizes and proposals suggesting how error-correction thresholds could be significantly improved by detecting leakage errors with single-atom resolution, a form of erasure error conversion. However, two-qubit entanglement fidelities in Rydberg atom arrays have lagged behind competitors and this type of erasure conversion is yet to be realized for matter-based qubits in general. Here we demonstrate both erasure conversion and high-fidelity Bell state generation using a Rydberg quantum simulator. When excising data with erasure errors observed via fast imaging of alkaline-earth atoms, we achieve a Bell state fidelity of $\ge 0.997{1}_{-13}^{+10}$ , which improves to $\ge 0.998{5}_{-12}^{+7}$ when correcting for remaining state-preparation errors. We further apply erasure conversion in a quantum simulation experiment for quasi-adiabatic preparation of long-range order across a quantum phase transition, and reveal the otherwise hidden impact of these errors on the simulation outcome. Our work demonstrates the capability for Rydberg-based entanglement to reach fidelities in the 0.999 regime, with higher fidelities a question of technical improvements, and shows how erasure conversion can be utilized in NISQ devices. These techniques could be translated directly to quantum-error-correction codes with the addition of long-lived qubits.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Experimental demonstration of continuous quantum error correction

The storage and processing of quantum information are susceptible to external noise, resulting in computational errors. A powerful method to suppress these effects is quantum error correction. Typically, quantum error correction is executed in discrete rounds, using entangling gates and projective measurement on ancillary qubits to complete each round of error correction. Here we use direct parity measurements to implement a continuous quantum bit-flip correction code in a resource-efficient manner, eliminating entangling gates, ancillary qubits, and their associated errors. An FPGA controller actively corrects errors as they are detected, achieving an average bit-flip detection efficiency of up to 91%. Furthermore, the protocol increases the relaxation time of the protected logical qubit by a factor of 2.7 over the relaxation times of the bare comprising qubits. Our results showcase resource-efficient stabilizer measurements in a multi-qubit architecture and demonstrate how continuous error correction codes can address challenges in realizing a fault-tolerant system.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

FPDetect: Efficient Reasoning About Stencil Programs Using Selective Direct Evaluation

We present FPDetect, a low-overhead approach for detecting logical errors and soft errors affecting stencil computations without generating false positives. We develop an offline analysis that tightly estimates the number of floating-point bits preserved across stencil applications. This estimate rigorously bounds the values expected in the data space of the computation. Violations of this bound can be attributed with certainty to errors. FPDetect helps synthesize error detectors customized for user-specified levels of accuracy and coverage. FPDetect also enables overhead reduction techniques based on deploying these detectors coarsely in space and time. Experimental evaluations demonstrate the practicality of our approach.

97 MATHEMATICS AND COMPUTING↗

Human factors process failure modes and effects analysis (HF PFMEA) software tool

Methods, computer-readable media, and systems for automatically performing Human Factors Process Failure Modes and Effects Analysis for a process are provided. At least one task involved in a process is identified, where the task includes at least one human activity. The human activity is described using at least one verb. A human error potentially resulting from the human activity is automatically identified, the human error is related to the verb used in describing the task. A likelihood of occurrence, detection, and correction of the human error is identified. The severity of the effect of the human error is identified. The likelihood of occurrence, and the severity of the risk of potential harm is identified. The risk of potential harm is compared with a risk threshold to identify the appropriateness of corrective measures.

Chandler, Faith T.↗

ATTNChecker: Highly-Optimized Fault Tolerant Attention for Large Language Model Training

Large Language Models (LLMs) have demonstrated remarkable performance in various natural language processing tasks. However, the training of these models is computationally intensive and susceptible to faults, particularly in the attention mechanism, which is a critical component of transformer-based LLMs. In this paper, we investigate the impact of faults on LLM training, focusing on INF, NaN, and near-INF values in the computation results with systematic fault injection experiments. We observe the propagation patterns of these errors, which can trigger non-trainable states in the model and disrupt training, forcing the procedure to load from checkpoints. To mitigate the impact of these faults, we propose ATTNChecker, the first Algorithm-Based Fault Tolerance (ABFT) technique tailored for the attention mechanism in LLMs. ATTNChecker is designed based on fault propagation patterns of LLM and incorporates performance optimization to adapt to both system reliability and model vulnerability while providing lightweight protection for fast LLM training. Evaluations on four LLMs show that ATTNChecker on average incurs on average 7% overhead on training while detecting and correcting all extreme errors. Compared with the state-of-the-art checkpoint/restore approach, ATTNChecker reduces recovery overhead by up to 49×.

Liang, Yuhang [University of Alabama - Birmingham]↗

3D Coded SUMMA: Communication-Efficient and Robust Parallel Matrix Multiplication

In this paper, we propose a novel fault-tolerant parallel matrix multiplication algorithm called 3D Coded SUMMA that achieves higher failure-tolerance than replication-based schemes for the same amount of redundancy. This work bridges the gap between recent developments in coded computing and fault-tolerance in high-performance computing (HPC). The core idea of coded computing is the same as algorithm-based fault-tolerance (ABFT), which is weaving redundancy in the computation using error-correcting codes. In particular, we show that MatDot codes, an innovative code construction for parallel matrix multiplications, can be integrated into three-dimensional SUMMA (Scalable Universal Matrix Multiplication Algorithm [30]) in a communication-avoiding manner. To tolerate any two node failures, the proposed 3D Coded SUMMA requires ~50% less redundancy than replication, while the overhead in execution time is only about 5–10%.

97 MATHEMATICS AND COMPUTING↗

Fault-tolerant operation and materials science with neutral atom logical qubits

We report on the fault-tolerant operation of logical qubits on a neutral atom quantum computer, with logical performance surpassing physical performance for multiple circuits including Bell state preparation (12x error reduction), random circuits (15x), and a prototype Anderson Impurity Model ground state solver for materials science applications (up to 6x, non-fault-tolerantly). The logical qubits are implemented via the [[4, 2, 2]] code (C 4 ). Our work constitutes the first complete realization of the benchmarking protocol proposed by Gottesman 2016 demonstrating results consistent with fault tolerance. In light of recent advances on applying concatenated C 4 /C 6 detection codes to achieve error correction with high code rates and thresholds, our work can be regarded as a building block towards a practical scheme for fault tolerant quantum computation. Our demonstration of a materials science application with logical qubits particularly demonstrates the immediate value of these techniques on current experiments.

36 MATERIALS SCIENCE↗

Design Primer for Reed-Solomon Encoders

Design and operation of Reed-Solomon (RS) encoders discussed in document prepared as instruction manual for computer designers and others in dataprocessing field. Conventional and Berlekamp architectures compared. Engineers who equip computer memory chips with burst-error and dropout detection and correction find report especially useful.

Perlman, M.↗

Fast decoding of a d(min) = 6 RS code

A method for high speed decoding a d sub min = 6 Reed-Solomon (RS) code is presented. Properties of the two byte error correcting and three byte error detecting RS code are discussed. Decoding using a quadratic equation is shown. Theorems and concomitant proofs are included to substantiate this decoding method.

Deng, H.↗

Decoding of DBEC-TBED Reed-Solomon codes

A problem in designing semiconductor memories is to provide some measure of error control without requiring excessive coding overhead or decoding time. In LSI and VLSI technology, memories are often organized on a multiple bit (or byte) per chip basis. For example, some 256 K bit DRAM's are organized in 32 K x 8 bit-bytes. Byte-oriented codes such as Reed-Solomon (RS) codes can provide efficient low overhead error control for such memories. However, the standard iterative algorithm for decoding RS codes is too slow for these applications. The paper presents a special decoding technique for double-byte-error-correcting, triple-byte-error-detecting RS codes which is capable of high-speed operation. This technique is designed to find the error locations and the error values directly from the syndrome without having to use the iterative algorithm to find the error locator polynomial.

Deng, Robert H.↗

A simplified procedure for decoding the (23,12) and (24,12) Golay codes

A simplified procedure is developed to decode the three possible erors in a (23,12) Golay codeword. A computer simulation shows that this algorithm is modular, regular and naturally suitable for both Very Large Scale Integration (VLSI) and software implementation. An extension of this new decoding procedure is used also to decode the 1/2-rate (24,12) Golay code, thereby correcting three and detecting four errors.

Truong, T. K.↗

Channel coding for digital HDTV terrestrial broadcasting

The Federal Communications Commission of the United States has ruled that high-definition television (HDTV) will occupy no more than 6 MHz of the VHF and UHF bands now used for conventional TV. In order to transmit the HDTV signal in 6 MHz, the four United States digital HDTV proponents, the DigiCipher, DSC-HDTV, ADTV, and ATVA-P systems, are reducing the video data rate of HDTV to 15-17 Mb/s, a compression ratio of approximately 60-70 times. The high compression dictates that channel coding be used to avoid block errors and multiframe error propagation. High efficiency in channel utilization required by the 6-MHz limitation means that the channel must be properly equalized and that the multipath and interfering signals must be severely limited. The channel coding techniques used for error reduction include data interleaving, error detection and replacement, and error correction at different levels of protection for bits and blocks of unequal importance.

Beakley, Guy W.↗

FPGA-Based, Self-Checking, Fault-Tolerant Computers

A proposed computer architecture would exploit the capabilities of commercially available field-programmable gate arrays (FPGAs) to enable computers to detect and recover from bit errors. The main purpose of the proposed architecture is to enable fault-tolerant computing in the presence of single-event upsets (SEUs). [An SEU is a spurious bit flip (also called a soft error) caused by a single impact of ionizing radiation.] The architecture would also enable recovery from some soft errors caused by electrical transients and, to some extent, from intermittent and permanent (hard) errors caused by aging of electronic components. A typical FPGA of the current generation contains one or more complete processor cores, memories, and highspeed serial input/output (I/O) channels, making it possible to shrink a board-level processor node to a single integrated-circuit chip. Custom, highly efficient microcontrollers, general-purpose computers, custom I/O processors, and signal processors can be rapidly and efficiently implemented by use of FPGAs. Unfortunately, FPGAs are susceptible to SEUs. Prior efforts to mitigate the effects of SEUs have yielded solutions that degrade performance of the system and require support from external hardware and software. In comparison with other fault-tolerant- computing architectures (e.g., triple modular redundancy), the proposed architecture could be implemented with less circuitry and lower power demand. Moreover, the fault-tolerant computing functions would require only minimal support from circuitry outside the central processing units (CPUs) of computers, would not require any software support, and would be largely transparent to software and to other computer hardware. There would be two types of modules: a self-checking processor module and a memory system (see figure). The self-checking processor module would be implemented on a single FPGA and would be capable of detecting its own internal errors. It would contain two CPUs executing identical programs in lock step, with comparison of their outputs to detect errors. It would also contain various cache local memory circuits, communication circuits, and configurable special-purpose processors that would use self-checking checkers. (The basic principle of the self-checking checker method is to utilize logic circuitry that generates error signals whenever there is an error in either the checker or the circuit being checked.) The memory system would comprise a main memory and a hardware-controlled check-pointing system (CPS) based on a buffer memory denoted the recovery cache. The main memory would contain random-access memory (RAM) chips and FPGAs that would, in addition to everything else, implement double-error-detecting and single-error-correcting memory functions to enable recovery from single-bit errors.

Some, Raphael↗