Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Soft Error Detection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Whole program Adaptive Error Detection and Mitigation

Our project investigated methods to detect soft errors in computing systems. Such detection schemes must have low overheads and also generate no false positives. Our solutions focused on control-space error detection based on changing addressing schemes so that faults tend to cascade – rather than being masked. Our solutions also addressed data-space error detection by synthesizing error detectors based on rigorous floating-point error analysis.

42 ENGINEERING↗

Determining diagnostic coverage for memory using redundant execution

Memory, used by a computer to store data, is generally prone to faults, including permanent faults (i.e. relating to a lifetime of the memory hardware), and also transient faults (i.e. relating to some external cause) which are otherwise known as soft errors. Since soft errors can change the state of the data in the memory and thus cause errors in applications reading and processing the data, there is a desire to characterize the degree of vulnerability of the memory to soft errors. In particular, once the vulnerability for a particular memory to soft errors has been characterized, cost/reliability trade-offs can be determined, or soft error detection mechanisms (e.g. parity) may be selectively employed for the memory. In some cases, memory faults can be diagnosed by redundant execution and a diagnostic coverage may be determined.

Bramley, Richard Gavin↗

Liveness as a factor to evaluate memory vulnerability to soft errors

Memory, used by a computer to store data, is generally prone to faults, including permanent faults (i.e. relating to a lifetime of the memory hardware), and also transient faults (i.e. relating to some external cause) which are otherwise known as soft errors. Since soft errors can change the state of the data in the memory and thus cause errors in applications reading and processing the data, there is a desire to characterize the degree of vulnerability of the memory to soft errors. In particular, once the vulnerability for a particular memory to soft errors has been characterized, cost/reliability trade-offs can be determined, or soft error detection mechanisms (e.g. parity) may be selectively employed for the memory. In some cases, memory faults can be diagnosed by redundant execution and a diagnostic coverage may be determined.

Bramley, Richard Gavin↗

FPDetect: Efficient Reasoning About Stencil Programs Using Selective Direct Evaluation

We present FPDetect, a low-overhead approach for detecting logical errors and soft errors affecting stencil computations without generating false positives. We develop an offline analysis that tightly estimates the number of floating-point bits preserved across stencil applications. This estimate rigorously bounds the values expected in the data space of the computation. Violations of this bound can be attributed with certainty to errors. FPDetect helps synthesize error detectors customized for user-specified levels of accuracy and coverage. FPDetect also enables overhead reduction techniques based on deploying these detectors coarsely in space and time. Experimental evaluations demonstrate the practicality of our approach.

97 MATHEMATICS AND COMPUTING↗

Optical Modem Enabling Broadband Datacom Links for Crewed Cis-Lunar Missions

We report on the design, development, and testing of our high-power broadband optical modem supporting NASA’s crewed Artemis-2 mission. The O2O modem will be mounted in the crewed Orion module and provide a broadband 505,000 km bi-directional optical link back to earth while en route to the moon. The full-duplex modem consists of a high-power optical transmitter and receiver optimized for serially-concatenated pulse-position modulation (SCPPM). The transmitter is a master-oscillator power-amplifier optical architecture using efficient cladding-pumped amplification in erbium-ytterbium co-doped fiber. The transmitter outputs up to 1 W at ≈1550 nm (limited for eye safety) and supports 6 different user-rates ranging from 20.39 Mb/s to 260.95 Mb/s using PPM16 and PPM32 modulation formats. The optical receiver supports two user-rates: 10.19 Mb/s and 20.39 Mb/s with both rates employing PPM32. The narrowband receiver filtering is optimized to simultaneously accept four separate wavelength channels to mitigate atmospherics through spatial diversity. A configurable interleaver provides additional protection against atmospherics-based signal fading and a powerful soft-decision error correction scheme enables highly sensitive detection. The measured sensitivities at the two bit-rates are -73.8 and -71.8dBm, respectively. The architecture was designed for reliable operation in space, featuring automatic hardware interlocks, pump sparing for the amplifiers, and autonomous operation of all internal hardware and software control loops. The protoflight unit (PFU) was put through rigorous environmental testing which included pyroshock, vibration, electromagnetic interference/compatibility, and thermal-vacuum testing. The modem successfully passed all the environmental screening and has been declared at Technology Readiness Level (TRL) 6.

space communications↗

Quasi-optimal decoding of linear block codes using soft decision detection

A simple but effective decoding procedure, applicable to any (n,k) linear block code with symbols from GF(q), is described. The technique involves a transformation of the parity check equations which focuses the code's correction power on the soft symbol set while still retaining the capability to correct one symbol error from outside this set. The soft symbol set is defined to be the n-k least reliably detected code symbol positions whose parity check row-spaces are linearly independent. The process generates a number of error vector screening candidates, each a solution to the parity check equations, and the maximum-likelihood candidate is accepted.

Greene, E. P.↗

The impact of mismatch on the performance of coded narrow-band FM with limiter/discriminator detection

An examination of the impact of mismatch on the performance of convolutionally encoded/Viterbi decoded narrow-band FM with limiter/discriminator detection is presented. Attention was given to the potential gain available by the combination of this type of system in terms of hard and soft decision decoding. Soft decision decoding was demonstrated to offer only approximately 0.3 dB better performance than hard decision coding. It was also shown, through a technique involving the number of clicks occurring in each detection interval, that both soft and hard decision decoding bit error probability performance could be improved. It is concluded that the mismatch between the coding channel and the decoding metric of the Viterbi algorithm is responsible for reducing the difference between hard and soft decoding metrics.

Simon, M. K.↗

Soft Syndrome Decoding of Quantum LDPC Codes for Joint Correction of Data and Syndrome Errors

Quantum errors are primarily detected and corrected using the measurement of syndrome information which itself is an unreliable step in practical error correction implementations. Typically, such faulty or noisy syndrome measurements are modeled as a binary measurement outcome flipped with some probability. However, the measured syndrome is in fact a discretized value of the continuous voltage or current values obtained in the physical implementation of the syndrome extraction. In this paper, we use this "soft" or analog information without the conventional discretization step to benefit the iterative decoders for decoding quantum low-density parity-check (QLDPC) codes. Syndrome-based iterative belief propagation (BP) decoders are modified to utilize the syndrome-soft information to successfully correct both data and syndrome errors simultaneously, without repeated measurements. We demonstrate the advantages of extracting the soft information from the syndrome in our improved decoders, not only in terms of comparison of thresholds and logical error rates for quasi-cyclic lifted-product QLDPC code families, but also for faster convergence of iterative decoders. In particular, the new BP decoder with noisy syndrome performs as good as the standard BP decoder under ideal syndrome.

97 MATHEMATICS AND COMPUTING↗

Minimax decoding of cyclic block codes

A minimax decoding algorithm utilizing soft bit detection of an (n,k) cyclic block code is described which will permit the correction of up to n-k bit errors interspersed at random locations throughout the block. The decoding solution consists of: (1) identifying the ordered soft bit set and, (2) finding the minimum order solution to the resulting syndrome equations where the nonzero error vector components are constrained to be a subset of the soft bit set. An efficient implementation of the decoding operation is described. In essence, this algorithm focuses the correction capability of the code on those bit positions which have the lowest a posteriori probabilities of correct detection.

Greene, E. P.↗

A Fast-Response High-Accuracy Overvoltage Protection Circuit for Soft-Switching Current-Source Converters

Although voltage-source converters (VSCs) have been a focus of research for decades and are widely applied in numerous applications, they face great challenges in short-circuit failures, high dv/dt and electromagnetic interference (EMI), especially using wide bandgap devices. Instead, current-source converters (CSCs) are attracting increasing attention in recent years owing to their friendliness to short-circuit faults, improved EMI, etc. For CSCs, overvoltage is the most catastrophic failure since the semiconductor devices can hardly withstand an overvoltage for a short pulse. In this paper, a fast-response high-accuracy overvoltage protection (OVP) circuit is proposed to protect CSCs from overvoltage damage. It also features a small form factor, good noise-immunity, friendly retrofit capability, and no need for active switches. In this paper, the operating principle and design guideline of the proposed OVP circuit is introduced. Its effectiveness is validated in soft-switching solid-state transformer (S4T) at 500 V. In experiments, the voltage detection error of less than 5% and a propagation delay of fewer than 400 ns have been achieved.

30 DIRECT ENERGY CONVERSION↗

Fault-tolerant system considerations for a redundant strapdown inertial measurement unit

The development and evaluation of a fault-tolerant system for the Redundant Strapdown Inertial Measurement Unit (RSDIMU) being developed and evaluated by the NASA Langley Research Center was continued. The RSDIMU consists of four two-degree-of-freedom gyros and accelerometers mounted on the faces of a semi-octahedron which can be separated into two halves for damage protection. Compensated and uncompensated fault-tolerant system failure decision algorithms were compared. An algorithm to compensate for sensor noise effects in the fault-tolerant system thresholds was evaluated via simulation. The effects of sensor location and magnitude of the vehicle structural modes on system performance were assessed. A threshold generation algorithm, which incorporates noise compensation and filtered parity equation residuals for structural mode compensation, was evaluated. The effects of the fault-tolerant system on navigational accuracy were also considered. A sensor error parametric study was performed in an attempt to improve the soft failure detection capability without obtaining false alarms. Also examined was an FDI system strategy based on the pairwise comparison of sensor measurements. This strategy has the specific advantage of, in many instances, successfully detecting and isolating up to two simultaneously occurring failures.

Motyka, P.↗

Asymmetric soft-error resistant memory

A memory system is provided, of the type that includes an error-correcting circuit that detects and corrects, that more efficiently utilizes the capacity of a memory formed of groups of binary cells whose states can be inadvertently switched by ionizing radiation. Each memory cell has an asymmetric geometry, so that ionizing radiation causes a significantly greater probability of errors in one state than in the opposite state (e.g., an erroneous switch from '1' to '0' is far more likely than a switch from '0' to'1'. An asymmetric error correcting coding circuit can be used with the asymmetric memory cells, which requires fewer bits than an efficient symmetric error correcting code.

Buehler, Martin G.↗

Resiliency in numerical algorithm design for extreme scale simulations

Here this work is based on the seminar titled ‘Resiliency in Numerical Algorithm Design for Extreme Scale Simulations’ held March 1–6, 2020, at Schloss Dagstuhl, that was attended by all the authors. Advanced supercomputing is characterized by very high computation speeds at the cost of involving an enormous amount of resources and costs. A typical large-scale computation running for 48 h on a system consuming 20 MW, as predicted for exascale systems, would consume a million kWh, corresponding to about 100k Euro in energy cost for executing 10 23 floating-point operations. It is clearly unacceptable to lose the whole computation if any of the several million parallel processes fails during the execution. Moreover, if a single operation suffers from a bit-flip error, should the whole computation be declared invalid? What about the notion of reproducibility itself: should this core paradigm of science be revised and refined for results that are obtained by large-scale simulation? Naive versions of conventional resilience techniques will not scale to the exascale regime: with a main memory footprint of tens of Petabytes, synchronously writing checkpoint data all the way to background storage at frequent intervals will create intolerable overheads in runtime and energy consumption. Forecasts show that the mean time between failures could be lower than the time to recover from such a checkpoint, so that large calculations at scale might not make any progress if robust alternatives are not investigated. More advanced resilience techniques must be devised. The key may lie in exploiting both advanced system features as well as specific application knowledge. Research will face two essential questions: (1) what are the reliability requirements for a particular computation and (2) how do we best design the algorithms and software to meet these requirements? While the analysis of use cases can help understand the particular reliability requirements, the construction of remedies is currently wide open. One avenue would be to refine and improve on system- or application-level checkpointing and rollback strategies in the case an error is detected. Developers might use fault notification interfaces and flexible runtime systems to respond to node failures in an application-dependent fashion. Novel numerical algorithms or more stochastic computational approaches may be required to meet accuracy requirements in the face of undetectable soft errors. These ideas constituted an essential topic of the seminar. The goal of this Dagstuhl Seminar was to bring together a diverse group of scientists with expertise in exascale computing to discuss novel ways to make applications resilient against detected and undetected faults. In particular, participants explored the role that algorithms and applications play in the holistic approach needed to tackle this challenge. This article gathers a broad range of perspectives on the role of algorithms, applications and systems in achieving resilience for extreme scale simulations. The ultimate goal is to spark novel ideas and encourage the development of concrete solutions for achieving such resilience holistically.

79 ASTRONOMY AND ASTROPHYSICS↗

Failure detection and isolation analysis of a redundant strapdown inertial measurement unit

The objective of this study was to define and develop techniques for failure detection and isolation (FDI) algorithms for a dual fail/operational redundant strapdown inertial navigation system are defined and developed. The FDI techniques chosen include provisions for hard and soft failure detection in the context of flight control and navigation. Analyses were done to determine error detection and switching levels for the inertial navigation system, which is intended for a conventional takeoff or landing (CTOL) operating environment. In addition, investigations of false alarms and missed alarms were included for the FDI techniques developed, along with the analyses of filters to be used in conjunction with FDI processing. Two specific FDI algorithms were compared: the generalized likelihood test and the edge vector test. A deterministic digital computer simulation was used to compare and evaluate the algorithms and FDI systems.

Motyka, P.↗

Failure detection and isolation methods for redundant gimballed inertial measurement units.

Skewed alignment of two redundant conventional inertial measuring units permits nonambiguous detection and isolation of hard and soft failures in real time by an airborne computer. Accelerometer outputs and gimbal readouts are monitored periodically, and attitude rate and velocity error vectors are computed from these data. Magnitudes of these vectors provide failure detection, and projection of these error vectors onto the coordinate axes of the two clusters permits isolation. A detailed Monte Carlo simulation of one version of the mechanization as applied to Space Shuttle boost trajectories demonstrates effectiveness down to very low levels of inertial instrument performance failures. The results indicate that worst case overall navigation performance occurs when accelerometer failures are of the order of 20 sigma and gyro failures are about 100 sigma for conventional state-of-the-art IMU instruments.

Solov, E. G.↗

An intense soft high-latitude X-ray source H2156-304 - A new BL Lacertae object

The discovery of an intense high-latitude soft X-ray source (designated H2156-304) detected with the low-energy detectors of the HEAO A-2 experiment is reported. The error box of the source includes a 14th mag starlike object that may be coincident with the radio source PKS 2155-304 and has been suggested as a BL Lac candidate. Intensity variations of H2156-304 on time scales of about 1 sec to 1 day are discussed, along with a probable flare from the source. The derived energy spectrum of the source is shown to be well fitted by either a two-temperature thermal model with temperatures of 16 million and 1.6 million K or a single power-law model with a spectral index of 2.4 + or - 0.3. Hydrogen column densities of approximately 2.5 x 10 to the 20th and 2.0 x 10 to the 20th H atoms/sq cm are obtained for the two-temperature and power-law models, respectively. The emission mechanism of H2156-304 is considered, and the characteristics of this source are compared with those of three known soft X-ray-emitting BL Lac objects.

Agrawal, P. C.↗

Analysis of a Coded, M-ary Orthogonal Input Optical Channel with Random-gain Photomultiplier Detection

Performance of two coding systems is analyzed for a noisy optical channel with M(=2(L)-ary orthogonal signaling and random gain photomultiplier detection. The considered coding systems are the Reed Solomon (RS) coding with error only correction decoding and the interleaved binary convolutional system with soft decision Viterbi decoding. The required average number of received signal photons per information bit, N sub b, for a desired bit error of 0.000001 is found for a set of commonly used parameters and with a high background noise level. We find that the interleaved binary convolutional coding system is preferable to the RS coding system in performance complexity tradeoffs.

Lee, P. J.↗