Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “error detection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Implementation of an experimental fault-tolerant memory system

The experimental fault-tolerant memory system described in this paper has been designed to enable the modular addition of spares, to validate the theoretical fault-secure and self-testing properties of the translator/corrector, to provide a basis for experiments using the new testing and correction processes for recovery, and to determine the practicality of such systems. The hardware design and implementation are described, together with methods of fault insertion. The hardware/software interface, including a restricted single error correction/double error detection (SEC/DED) code, is specified. Procedures are carefully described which, (1) test for specified physical faults, (2) ensure that single error corrections are not miscorrections due to triple faults, and (3) enable recovery from double errors.

Carter, W. C.↗

The Effectiveness of TAG or Guard-Gates in SET Suppression Using Delay and Dual-Rail Configurations at 0.35 microns

Design options for decreasing the susceptibility of integrated circuits to Single Event Upset (SEU) fall into two categories: (1) increasing the critical charge to cause an upset at a particular node, and (2) employing redundancy to mask or correct errors. With decreasing device sizes on an Integrated Circuit (IC), the amount of charge required to represent a logic state has steadily reduced. Critical charge methods such as increasing drive strength or increasing the time required to change state as in capacitive or resistive hardening or delay based approaches extract a steadily increasing penalty as a percentage of device resources and performance. Dual redundancy is commonly assumed only to provide error detection with Triple Modular Redundancy (TMR) required for correction, but less well known methods employ dual redundancy to achieve full error correction by voting two inputs with a prior state to resolve ambiguity. This requires special circuits such as the Whitaker latch [1], or the guard-gate [2] which some of us have called a Transition AND Gate (TAG) [3]. A 2-input guard gate is shown in Figure 1. It is similar to a Muller Completion Element [4] and relies on capacitance at node "out" to retain the prior state when inputs disagree, while eliminating any output buffer which would be susceptible to radiation strikes. This paper experimentally compares delay based and dual rail flip-flop designs wherein both types of circuits employ guard-gates to optimize layout and performance, and draws conclusions about design criteria and suitability of each option. In both cases a design goal is protection against Single Event Transients (SET) in combinational logic as well as SEU in the storage elements. For the delay based design, it is also a goal to allow asynchronous clear or preset inputs on the storage elements, which are often not available in radiation tolerant designs.

Shuler, Robert L.↗

Securing against errors in an error correcting code (ECC) implemented in an automotive system

In general, data is susceptible to errors caused by faults in hardware (i.e. permanent faults), such as faults in the functioning of memory and/or communication channels. To detect errors in data caused by hardware faults, the error correcting code (ECC) was introduced, which essentially provides a sort of redundancy to the data that can be used to validate that the data is free from errors caused by hardware faults. In some cases, the ECC can also be used to correct errors in the data caused by hardware faults. However, the ECC itself is also susceptible to errors, including specifically errors caused by faults in the ECC logic. A method, computer readable medium, and system are thus provided for securing against errors in an ECC.

Saxena, Nirmal Raj↗

Segmented testing

The fraction of faults detected for a digital network is frequently high for the first few input combinations applied out of a set of test vectors. When the particular ordering of test patterns does not appreciably change the shape of the coverage curve, there appears to be an advantage to splitting the test into segments which are applied at different times. It is shown that the expected time to error detection and the probability of an undetected double error can be reduced. The amount of reduction is dependent on the shape of the fault coverage curve. It is conjectured that such a reduction can be obtained for VLSI networks.

Robinson, J. P.↗

Segmented testing

The fraction of faults detected for a digital network is frequently high for the first few input combinations applied out of a set of test vectors. When the particular ordering of test patterns does not appreciably change the shape of the coverage curve, there appears to be an advantage to splitting the test into segments which are applied at different times. It is shown that the expected time to error detection and the probability of an undetected double error can be reduced. The amount of reduction is dependent on the shape of the fault coverage curve. It is conjectured that such a reduction can be obtained for VLSI networks.

Robinson, J. P.↗

Spacelab high density digital recorders

The design and performance of the high-density digital recorder (HDDR) developed for use at the NASA centers (KSC, JSC, and GSFC) and at the JPL to store and retrieve 50-Mb/s PCM data streams from the Spacelab experiments are reported. The recording reproduction, and transport requirements are reviewed; and the design solutions adopted in the final version of the HDDR are described, incuding three-position-modulation and Y-phase encoding, microprocessor-controlled automatic bit synchronization and equalization, cyclic-redundancy-check error detection and correction, clock regeneration, data and clock variations, tape-speed control, and EEE-488 remote control. Reliable performance, with bit error rates 1 in 10 to the 10th forward and 1 in 10 to the 9th reverse or better and packing density up to 50 percent greater than that obtainable using conventional codes, is reported after 1.5 years of service.

Blais, R. A.↗

Simulation results for an innovative anti-multipath digital receiver

Simulation results are presented for the error rate performance of the recursive digital MAP detector for known M-ary signals in multiplicative and additive Gaussian noise. Plots of detection error rate versus additive signal to noise ratio are given, with multipath interference strength as a parameter. For comparison, the error rates of conventional coherent and noncoherent digital MAP detectors are simultaneously simulated and graphed. It is shown that with nonzero multiplicative noise, the error rates of the conventional detectors saturate at an irreducible level as additive signal to noise ratio increases. The error rate for the innovative detector continues to decrease rapidly with increasing additive signal to noise ratio. In the absence of multiplicative interference, the conventional coherent detector and the innovative detector are shown to exhibit identical performance.-

Painter, J. H.↗

Recursive ideal observer detection of known M-ary signals in multiplicative and additive Gaussian noise.

This paper presents the derivation of the recursive algorithms necessary for real-time digital detection of M-ary known signals that are subject to independent multiplicative and additive Gaussian noises. The motivating application is minimum probability of error detection of digital data-link messages aboard civil aircraft in the earth reflection multipath environment. For each known signal, the detector contains one Kalman filter and one probability computer. The filters estimate the multipath disturbance. The estimates and the received signal drive the probability computers. Outputs of all the computers are compared in amplitude to give the signal decision. The practicality and usefulness of the detector are extensively discussed.

Painter, J. H.↗

Dynamic assertion testing of flight control software

Digital Flight Control System (DFCS) software was used as a test case for assertion testing. The assertions were written and embedded in the code, then errors were inserted (seeded) one at a time and the code executed. Results indicate that assertion testing is an effective and efficient method of detecting errors in flight software. Most errors are eliminate at an earlier stage in the development than before.

Andrews, D. M.↗

The use of automatic programming techniques for fault tolerant computing systems

It is conjectured that the production of software for ultra-reliable computing systems such as required by Space Station, aircraft, nuclear power plants and the like will require a high degree of automation as well as fault tolerance. In this paper, the relationship between automatic programming techniques and fault tolerant computing systems is explored. Initial efforts in the automatic synthesis of code from assertions to be used for error detection as well as the automatic generation of assertions and test cases from abstract data type specifications is outlined. Speculation on the ability to generate truly diverse designs capable of recovery from errors by exploring alternate paths in the program synthesis tree is discussed. Some initial thoughts on the use of knowledge based systems for the global detection of abnormal behavior using expectations and the goal-directed reconfiguration of resources to meet critical mission objectives are given. One of the sources of information for these systems would be the knowledge captured during the automatic programming process.

Wild, C.↗

Investigating Resilience of Loops in HPC Programs: A Semantic Approach with LLMs

Soft errors have become one of the major concerns for the error resilience of the HPC applications as those errors may cause HPC applications to generate serious outcomes such as silent data corruptions (SDCs). Protecting the applications from soft errors is an essential while challenging task. Among different approaches, obtaining a profound understanding of the resilience proneness of an application is very important to devise efficient error detection and recovery strategies. Given the scale of the HPC applications both in the code size and execution time, there are often cases that the error propagation analysis on such applications would produce a massive volume of unstructured data, which requires a significant amount of efforts, to process and to obtain indicating actions towards error protection. In this paper, we present a control-flow based visual analysis framework to help the users conduct error propagation analysis and identify the critical sections of a program that may have a higher likelihood of leading to erroneous outcomes when affected by the control flow related errors. We also design and implement the scalable visualization framework - ResilienceVis that efficiently and effectively visualizes the affected program states under errors and the propagation traces for an application in a user-friendly manner, and eventually, we combine the analysis and visualization to exhibit the error-proneness of the different sections of applications.

Jiang, Hailong↗

Augmented burst-error correction for UNICON laser memory

A single-burst-error correction system is described for data stored in the UNICON laser memory. In the proposed system, a long fire code with code length n greater than 16,768 bits was used as an outer code to augment an existing inner shorter fire code for burst error corrections. The inner fire code is a (80,64) code shortened from the (630,614) code, and it is used to correct a single-burst-error on a per-word basis with burst length b less than or equal to 6. The outer code, with b less than or equal to 12, would be used to correct a single-burst-error on a per-page basis, where a page consists of 512 32-bit words. In the proposed system, the encoding and error detection processes are implemented by hardware. A minicomputer, currently used as a UNICON memory management processor, is used on a time-demanding basis for error correction. Based upon existing error statistics, this combination of an inner code and an outer code would enable the UNICON system to obtain a very low error rate in spite of flaws affecting the recorded data.

Lim, R. S.↗

A packet telemetry system employing ARQ error control

A proposed packet telemetry system employing automatic retransmission request (ARQ) mode of error control is characterized. Limitations of the present multiplexing/demultiplexing approach are considered, and the use of the proposed system in near-earth satellites in the 1980s is suggested. Onboard processing and an adaptive multiplexing technique are described, as is an elastic buffer, required because the instantaneous data rate will be different from the telemetry transmission rate. The telemetry packets would be encoded into a powerful error-detection block code. A mechanism involving temporary buffering in a long shift register will permit retransmission request from the ground station for packets received in error. The ARQ mode of operation should ensure essentially error-free transmission at lower signal-to-noise ratios and at considerably higher transmission rates than are usually used.

Greene, E. P.↗

An experiment in software reliability

The results of a software reliability experiment conducted in a controlled laboratory setting are reported. The experiment was undertaken to gather data on software failures and is one in a series of experiments being pursued by the Fault Tolerant Systems Branch of NASA Langley Research Center to find a means of credibly performing reliability evaluations of flight control software. The experiment tests a small sample of implementations of radar tracking software having ultra-reliability requirements and uses n-version programming for error detection, and repetitive run modeling for failure and fault rate estimation. The experiment results agree with those of Nagel and Skrivan in that the program error rates suggest an approximate log-linear pattern and the individual faults occurred with significantly different error rates. Additional analysis of the experimental data raises new questions concerning the phenomenon of interacting faults. This phenomenon may provide one explanation for software reliability decay.

Dunham, J. R.↗

Logical quantum processor based on reconfigurable atom arrays

Suppressing errors is the central challenge for useful quantum computing, requiring quantum error correction (QEC) for large-scale processing. However, the overhead in the realization of error-corrected ‘logical’ qubits, in which information is encoded across many physical qubits for redundancy, poses substantial challenges to large-scale logical quantum computing. Here we report the realization of a programmable quantum processor based on encoded logical qubits operating with up to 280 physical qubits. Using logical-level control and a zoned architecture in reconfigurable neutral-atom arrays, our system combines high two-qubit gate fidelities, arbitrary connectivity, as well as fully programmable single-qubit rotations and mid-circuit readout. Operating this logical processor with various types of encoding, we demonstrate improvement of a two-qubit logic gate by scaling surface-code distance from d = 3 to d = 7, preparation of colour-code qubits with break-even fidelities, fault-tolerant creation of logical Greenberger–Horne–Zeilinger (GHZ) states and feedforward entanglement teleportation, as well as operation of 40 colour-code qubits. Finally, using 3D [[8,3,2]] code blocks, we realize computationally complex sampling circuits with up to 48 logical qubits entangled with hypercube connectivity with 228 logical two-qubit gates and 48 logical CCZ gates. We find that this logical encoding substantially improves algorithmic performance with error detection, outperforming physical-qubit fidelities at both cross-entropy benchmarking and quantum simulations of fast scrambling. These results herald the advent of early error-corrected quantum computation and chart a path towards large-scale logical processors.

74 ATOMIC AND MOLECULAR PHYSICS↗

Multiple single event upsets in CMOS static rams

The occurrence of multiple upset errors during ground tests can contaminate the data and lead to error cross sections which are too high. In space, multiple errors may produce higher upsets than predicted and if they occur in single words they can defeat error detection and correction hardware. This investigation involves data which were taken during an experimental study of dose imprint effects in static memories. The results show that multiple errors occurred mainly for heavy ions with high linear energy transfers and with the majority of these in the soft upset sections of the dose-imprinted memory samples. The percentages of the total number of errors which were singles, doubles, and triplets, were determined as a function of LET, dose, and soft or hard section of the devices. The experimental observations are compared to the predictions of simple binomial statistics.

Stassinopoulos, E. G.↗

Software error data collection and categorization

Software errors detected during development of an interactive special purpose editor system were studied. This product was followed during nine months of coding, unit testing, function testing, and system testing. A new error categorization scheme was developed.

Ostrand, T. J.↗

Automation of assertion testing - Grid and adaptive techniques

Assertions can be used to automate the process of testing software. Two methods for automating the generation of input test data are described in this paper. One method selects the input values of variables at regular intervals in a 'grid'. The other, adaptive testing, uses assertion violations as a measure of errors detected and generates new test cases based on test results. The important features of assertion testing are that: it can be used throughout the entire testing cycle; it provides automatic notification of error conditions; and it can be used with automatic input generation techniques which eliminate the subjectivity in choosing test data.

Andrews, D. M.↗