Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Soft Error Detection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Whole program Adaptive Error Detection and Mitigation

Our project investigated methods to detect soft errors in computing systems. Such detection schemes must have low overheads and also generate no false positives. Our solutions focused on control-space error detection based on changing addressing schemes so that faults tend to cascade – rather than being masked. Our solutions also addressed data-space error detection by synthesizing error detectors based on rigorous floating-point error analysis.

42 ENGINEERING↗

Determining diagnostic coverage for memory using redundant execution

Memory, used by a computer to store data, is generally prone to faults, including permanent faults (i.e. relating to a lifetime of the memory hardware), and also transient faults (i.e. relating to some external cause) which are otherwise known as soft errors. Since soft errors can change the state of the data in the memory and thus cause errors in applications reading and processing the data, there is a desire to characterize the degree of vulnerability of the memory to soft errors. In particular, once the vulnerability for a particular memory to soft errors has been characterized, cost/reliability trade-offs can be determined, or soft error detection mechanisms (e.g. parity) may be selectively employed for the memory. In some cases, memory faults can be diagnosed by redundant execution and a diagnostic coverage may be determined.

Bramley, Richard Gavin↗

Liveness as a factor to evaluate memory vulnerability to soft errors

Memory, used by a computer to store data, is generally prone to faults, including permanent faults (i.e. relating to a lifetime of the memory hardware), and also transient faults (i.e. relating to some external cause) which are otherwise known as soft errors. Since soft errors can change the state of the data in the memory and thus cause errors in applications reading and processing the data, there is a desire to characterize the degree of vulnerability of the memory to soft errors. In particular, once the vulnerability for a particular memory to soft errors has been characterized, cost/reliability trade-offs can be determined, or soft error detection mechanisms (e.g. parity) may be selectively employed for the memory. In some cases, memory faults can be diagnosed by redundant execution and a diagnostic coverage may be determined.

Bramley, Richard Gavin↗

FPDetect: Efficient Reasoning About Stencil Programs Using Selective Direct Evaluation

We present FPDetect, a low-overhead approach for detecting logical errors and soft errors affecting stencil computations without generating false positives. We develop an offline analysis that tightly estimates the number of floating-point bits preserved across stencil applications. This estimate rigorously bounds the values expected in the data space of the computation. Violations of this bound can be attributed with certainty to errors. FPDetect helps synthesize error detectors customized for user-specified levels of accuracy and coverage. FPDetect also enables overhead reduction techniques based on deploying these detectors coarsely in space and time. Experimental evaluations demonstrate the practicality of our approach.

97 MATHEMATICS AND COMPUTING↗

Soft Syndrome Decoding of Quantum LDPC Codes for Joint Correction of Data and Syndrome Errors

Quantum errors are primarily detected and corrected using the measurement of syndrome information which itself is an unreliable step in practical error correction implementations. Typically, such faulty or noisy syndrome measurements are modeled as a binary measurement outcome flipped with some probability. However, the measured syndrome is in fact a discretized value of the continuous voltage or current values obtained in the physical implementation of the syndrome extraction. In this paper, we use this "soft" or analog information without the conventional discretization step to benefit the iterative decoders for decoding quantum low-density parity-check (QLDPC) codes. Syndrome-based iterative belief propagation (BP) decoders are modified to utilize the syndrome-soft information to successfully correct both data and syndrome errors simultaneously, without repeated measurements. We demonstrate the advantages of extracting the soft information from the syndrome in our improved decoders, not only in terms of comparison of thresholds and logical error rates for quasi-cyclic lifted-product QLDPC code families, but also for faster convergence of iterative decoders. In particular, the new BP decoder with noisy syndrome performs as good as the standard BP decoder under ideal syndrome.

97 MATHEMATICS AND COMPUTING↗

A Fast-Response High-Accuracy Overvoltage Protection Circuit for Soft-Switching Current-Source Converters

Although voltage-source converters (VSCs) have been a focus of research for decades and are widely applied in numerous applications, they face great challenges in short-circuit failures, high dv/dt and electromagnetic interference (EMI), especially using wide bandgap devices. Instead, current-source converters (CSCs) are attracting increasing attention in recent years owing to their friendliness to short-circuit faults, improved EMI, etc. For CSCs, overvoltage is the most catastrophic failure since the semiconductor devices can hardly withstand an overvoltage for a short pulse. In this paper, a fast-response high-accuracy overvoltage protection (OVP) circuit is proposed to protect CSCs from overvoltage damage. It also features a small form factor, good noise-immunity, friendly retrofit capability, and no need for active switches. In this paper, the operating principle and design guideline of the proposed OVP circuit is introduced. Its effectiveness is validated in soft-switching solid-state transformer (S4T) at 500 V. In experiments, the voltage detection error of less than 5% and a propagation delay of fewer than 400 ns have been achieved.

30 DIRECT ENERGY CONVERSION↗

Resiliency in numerical algorithm design for extreme scale simulations

Here this work is based on the seminar titled ‘Resiliency in Numerical Algorithm Design for Extreme Scale Simulations’ held March 1–6, 2020, at Schloss Dagstuhl, that was attended by all the authors. Advanced supercomputing is characterized by very high computation speeds at the cost of involving an enormous amount of resources and costs. A typical large-scale computation running for 48 h on a system consuming 20 MW, as predicted for exascale systems, would consume a million kWh, corresponding to about 100k Euro in energy cost for executing 10 23 floating-point operations. It is clearly unacceptable to lose the whole computation if any of the several million parallel processes fails during the execution. Moreover, if a single operation suffers from a bit-flip error, should the whole computation be declared invalid? What about the notion of reproducibility itself: should this core paradigm of science be revised and refined for results that are obtained by large-scale simulation? Naive versions of conventional resilience techniques will not scale to the exascale regime: with a main memory footprint of tens of Petabytes, synchronously writing checkpoint data all the way to background storage at frequent intervals will create intolerable overheads in runtime and energy consumption. Forecasts show that the mean time between failures could be lower than the time to recover from such a checkpoint, so that large calculations at scale might not make any progress if robust alternatives are not investigated. More advanced resilience techniques must be devised. The key may lie in exploiting both advanced system features as well as specific application knowledge. Research will face two essential questions: (1) what are the reliability requirements for a particular computation and (2) how do we best design the algorithms and software to meet these requirements? While the analysis of use cases can help understand the particular reliability requirements, the construction of remedies is currently wide open. One avenue would be to refine and improve on system- or application-level checkpointing and rollback strategies in the case an error is detected. Developers might use fault notification interfaces and flexible runtime systems to respond to node failures in an application-dependent fashion. Novel numerical algorithms or more stochastic computational approaches may be required to meet accuracy requirements in the face of undetectable soft errors. These ideas constituted an essential topic of the seminar. The goal of this Dagstuhl Seminar was to bring together a diverse group of scientists with expertise in exascale computing to discuss novel ways to make applications resilient against detected and undetected faults. In particular, participants explored the role that algorithms and applications play in the holistic approach needed to tackle this challenge. This article gathers a broad range of perspectives on the role of algorithms, applications and systems in achieving resilience for extreme scale simulations. The ultimate goal is to spark novel ideas and encourage the development of concrete solutions for achieving such resilience holistically.

79 ASTRONOMY AND ASTROPHYSICS↗

Investigating Resilience of Loops in HPC Programs: A Semantic Approach with LLMs

Soft errors have become one of the major concerns for the error resilience of the HPC applications as those errors may cause HPC applications to generate serious outcomes such as silent data corruptions (SDCs). Protecting the applications from soft errors is an essential while challenging task. Among different approaches, obtaining a profound understanding of the resilience proneness of an application is very important to devise efficient error detection and recovery strategies. Given the scale of the HPC applications both in the code size and execution time, there are often cases that the error propagation analysis on such applications would produce a massive volume of unstructured data, which requires a significant amount of efforts, to process and to obtain indicating actions towards error protection. In this paper, we present a control-flow based visual analysis framework to help the users conduct error propagation analysis and identify the critical sections of a program that may have a higher likelihood of leading to erroneous outcomes when affected by the control flow related errors. We also design and implement the scalable visualization framework - ResilienceVis that efficiently and effectively visualizes the affected program states under errors and the propagation traces for an application in a user-friendly manner, and eventually, we combine the analysis and visualization to exhibit the error-proneness of the different sections of applications.

Jiang, Hailong↗

Kelvin probe force microscopy under ambient conditions

Kelvin probe force microscopy (KPFM) is a technique derived from atomic force microscopy that provides maps of surface potential or work function differences across material systems, with nanometre-scale resolution. KPFM is a useful tool for investigating electrical phenomena such as dipole orientation, interfacial charge transfer, charge accumulation, band bending and doping levels. This Primer aims to provide an overview of typical ambient-condition KPFM measurements, covering their underlying principles, experimental implementations and wide-ranging applications. Key KPFM variants, including amplitude and frequency modulation, heterodyne detection schemes and innovative open loop and pulsed force techniques, are discussed, with practical guidance on optimizing signal acquisition and reducing errors. Specialized approaches, such as time-resolved KPFM and multimodal KPFM, are discussed for their ability to capture dynamic charge processes and chemical information, respectively. Here, we highlight recent advances in KPFM applications, spanning metal alloys, soft matter, ferroelectrics, photovoltaics and 2D materials, showcasing its versatility across research domains. By addressing current limitations and identifying future opportunities, this Primer underscores the transformative potential of KPFM in advancing the understanding of nanoscale electrical phenomena.

Zahmatkeshsaredorahi, Amirhossein [Lehigh Univ., B↗