Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “error detection and error correction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

GCF HSD error control

A selective repeat automatic repeat request (ARQ) system was implemented under software control in the Ground Communications Facility error detection and correction (EDC) assembly at JPL and the comm monitor and formatter (CMF) assembly at the DSSs. The CMF and EDC significantly improved real time data quality and significantly reduced the post-pass time required for replay of blocks originally received in error. Since the remote mission operation centers (RMOCs) do not provide compatible error correction equipment, error correction will not be used on the RMOC-JPL high speed data (HSD) circuits. The real time error correction capability will correct error burst or outage of two loop-times or less for each DSS-JPL HSD circuit.

Hung, C. K.↗

Understanding Mixed Precision GEMM with MPGemmFI: Insights into Fault Resilience

Emerging deep learning workloads urgently need fast general matrix multiplication (GEMM). Thus, one of the critical features of machine-learning-specific accelerators such as NVIDIA Tensor Cores, AMD Matrix Cores, and Google TPUs is the support of mixed-precision enabled GEMM. For DNN models, lower-precision FP data formats and computation offer acceptable correctness but significant performance, area, and memory footprint improvement. While promising, the mixed-precision computation on error resilience remains unexplored. To this end, we develop a fault injection framework that systematically injects fault into the mixed-precision computation results. We investigate how the faults affect the accuracy of machine learning applications. Based on the characteristics of error resilience, we offer lightweight error detection and correction solutions that significantly improve the overall model accuracy by 75% if the models experience hardware faults. The solutions can be efficiently integrated into the accelerator's pipelines.

Fang, Bo↗

Self-dual (48,24;12) codes

Two self-dual (48,24;12) codes are constructed as 6 x 8 matrices whose columns add up to form an extended BCH-Hamming (8,4;4) code and whose rows sum to odd or even parity. The codes constructed have the identical weight structure of the extended quadratic residue code of length 48. Algebraic isomorphisms may exist between pairs of these three codes. However, because of their matrix form, the newly constructed codes are easily correctable for all five-error and many six-error patterns. The first code comes from restricting a binary cyclic (63,18;36) code to a 6 x 7 matrix and then adjoining six dimensions to the extended 6 x 8 matrix. These six dimensions are generated by linear combinations of row permutations of a 6 x 8 matrix of weight 12, whose sums of rows and columns add to one. The second code comes from a slight modification in the parity (eighth) dimension of the Reed-Solomon (8,4;5) code over GF(64). Error correction in both codes uses the row sum parity information to detect errors in the correction algorithm.

Solomon, G.↗

Low-overhead transversal fault tolerance for universal quantum computation

Fast, reliable logical operations are essential for realizing useful quantum computers. By redundantly encoding logical qubits into many physical qubits and using syndrome measurements to detect and correct errors, we can achieve low logical error rates. However, for many practical quantum error correction codes such as the surface code, owing to syndrome measurement errors, standard constructions require multiple extraction rounds—of the order of the code distance d—for fault-tolerant computation, particularly considering fault-tolerant state preparation. Here we show that logical operations can be performed fault-tolerantly with only a constant number of extraction rounds for a broad class of quantum error correction codes, including the surface code with magic state inputs and feedforward, to achieve ‘transversal algorithmic fault tolerance’. Through the combination of transversal operations7 and new strategies for correlated decoding, despite only having access to partial syndrome information, we prove that the deviation from the ideal logical measurement distribution can be made exponentially small in the distance, even if the instantaneous quantum state cannot be made close to a logical codeword because of measurement errors. We supplement this proof with circuit-level simulations in a range of relevant settings, demonstrating the fault tolerance and competitive performance of our approach. Furthermore, our work sheds new light on the theory of quantum fault tolerance and has the potential to reduce the space–time cost of practical fault-tolerant quantum computation by over an order of magnitude.

Zhou, Hengyun [QuEra Computing, Boston, MA (United↗

Processing and Archiving Camera Data Effectively for Operando Neutron Measurement of Metal Additive Manufacturing

As the name suggests, The Operando Neutron Measurement of Metal Additive Manufacturing project conducted by ORNL’s Manufacturing Demonstration Facility (MDF) is experimenting with advanced additive metal manufacturing techniques while analyzing the process using the Spallation Neutron Source’s (SNS) beamline. As part of this experiment, the MDF is seeking to employ 2 XIMEA visible light cameras and a single infrared camera to analyze and correct manufacturing in real-time. The MDF requires a solution for capturing the high-resolution data feed from the cameras with compression while preserving enough detail for their software to detect and correct errors in real time. Our solution was to develop a Robot Operating System (ROS) driver to feed the camera data into ROS. From ROS, the feed is compressed and temporarily stored locally to a stripped 4 NVMe SSD RAID array. Post-experiment, the data is transitioned to long-term storage for archival purposes.

36 MATERIALS SCIENCE↗

Self testing and repairing computer - A concept

STAR computer has five redundant modular function units, fixed store, arithmetic, memory, input, and output. Each unit is connected to a diagnostic control unit, each is coded for error detection and error correction. Separation into function units permits assembly of many different systems from the set of units.

Avizienis, A. A.↗

Development and testing of the infrared interferometer spectrometer for the Mariner Mars 1971 spacecraft

The design, development and testing of the infrared interferometer spectrometer is reported with emphasis on the unique features of the Mariner instrument as compared to previous IRIS instruments flown on the Nimbus meteorological research satellites. The interferometer functions in the spectral range from 50 microns to 6.3 microns. A noise equivalent radiance of 0.5 X 10 to the -7th power W/sq cm/ster/cm has been achieved. Major improvements that were implemented included the cesium iodide beamsplitter and electronic features to suppress the effect of vibration on the Michelson mirror motion and digital filtering through the summation of increased sampling of the infrared signal. A bit error detection and correction scheme was also implemented in order to recover the science data with a higher level of confidence over the telecommunication link.

Hanel, R. H.↗

A fault-tolerant information processing concept for space vehicles.

A distributed fault-tolerant information processing system is proposed, comprising a central multiprocessor, dedicated local processors, and multiplexed input-output buses connecting them together. The processors in the multiprocessor are duplicated for error detection, which is felt to be less expensive than using coded redundancy of comparable effectiveness. Error recovery is made possible by a triplicated scratchpad memory in each processor. The main multiprocessor memory uses replicated memory for error detection and correction. Local processors use any of three conventional redundancy techniques: voting, duplex pairs with backup, and duplex pairs in independent subsystems.

Hopkins, A. L., Jr.↗

Shift register generators and applications to coding

The most important properties of shift register generated sequences are exposed. The application of shift registers as multiplication and division circuits leads to the generation of some error correcting and detecting codes.

Morakis, J. C.↗

Interleaved cyclic codes

Analytical approach for development burst error correction and detection cyclic codes does not require shortening techniques.

Hockenberger, R. W.↗

Fault tolerant computing: A preamble for assuring viability of large computer systems

The need for fault-tolerant computing is addressed from the viewpoints of (1) why it is needed, (2) how to apply it in the current state of technology, and (3) what it means in the context of the Phoenix computer system and other related systems. To this end, the value of concurrent error detection and correction is described. User protection, program retry, and repair are among the factors considered. The technology of algebraic codes to protect memory systems and arithmetic codes to protect memory systems and arithmetic codes to protect arithmetic operations is discussed.

Lim, R. S.↗

The proposed coding standard at GSFC

As part of the continuing effort to introduce standardization of spacecraft and ground equipment in satellite systems, NASA's Goddard Space Flight Center and other NASA facilities have supported the development of a set of standards for the use of error control coding in telemetry subsystems. These standards are intended to ensure compatibility between spacecraft and ground encoding equipment, while allowing sufficient flexibility to meet all anticipated mission requirements. The standards which have been developed to date cover the application of block codes in error detection and error correction modes, as well as short and long constraint length convolutional codes decoded via the Viterbi and sequential decoding algorithms, respectively. Included are detailed specifications of the codes, and their implementation. Current effort is directed toward the development of standards covering channels with burst noise characteristics, channels with feedback, and code concatenation.

Morakis, J. C.↗

Evaluation of the SPAR thermal analyzer on the CYBER-203 computer

The use of the CYBER 203 vector computer for thermal analysis is investigated. Strengths of the CYBER 203 include the ability to perform, in vector mode using a 64 bit word, 50 million floating point operations per second (MFLOPS) for addition and subtraction, 25 MFLOPS for multiplication and 12.5 MFLOPS for division. The speed of scalar operation is comparable to that of a CDC 7600 and is some 2 to 3 times faster than Langley's CYBER 175s. The CYBER 203 has 1,048,576 64-bit words of real memory with an 80 nanosecond (nsec) access time. Memory is bit addressable and provides single error correction, double error detection (SECDED) capability. The virtual memory capability handles data in either 512 or 65,536 word pages. The machine has 256 registers with a 40 nsec access time. The weaknesses of the CYBER 203 include the amount of vector operation overhead and some data storage limitations. In vector operations there is a considerable amount of time before a single result is produced so that vector calculation speed is slower than scalar operation for short vectors.

Robinson, J. C.↗

Undetected error probability and throughput analysis of a concatenated coding scheme

The performance of a proposed concatenated coding scheme for error control on a NASA telecommand system is analyzed. In this scheme, the inner code is a distance-4 Hamming code used for both error correction and error detection. The outer code is a shortened distance-4 Hamming code used only for error detection. Interleaving is assumed between the inner and outer codes. A retransmission is requested if either the inner or outer code detects the presence of errors. Both the undetected error probability and the throughput of the system are analyzed. Results indicate that high throughputs and extremely low undetected error probabilities are achievable using this scheme.

Costello, D. J.↗

High-density digital recording

The problems associated with high-density digital recording (HDDR) are discussed. Five independent users of HDDR systems and their problems, solutions, and insights are provided as guidance for other users of HDDR systems. Various pulse code modulation coding techniques are reviewed. An introduction to error detection and correction head optimization theory and perpendicular recording are provided. Competitive tape recorder manufacturers apply all of the above theories and techniques and present their offerings. The methodology used by the HDDR Users Subcommittee of THIC to evaluate parallel HDDR systems is presented.

Kalil, F.↗