Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “soft errors”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Latest trends in parts SEP susceptibility from heavy ions

JPL and Aerospace have collected a third set of heavy-ion single-event phenomena (SEP) test data since their last joint IEEE publications in December 1985 and December 1987. Trends in SEP susceptibility (e.g., soft errors and latchup) for state-of-the-art parts are presented. Results of the study indicate that hard technologies and unacceptably soft technologies can be flagged. In some instances, specific tested parts can be taken as candidates for key microprocessors or memories. As always with radiation test data, specific test data for qualified flight parts is recommended for critical applications.

Nichols, Donald K.↗

Resiliency in numerical algorithm design for extreme scale simulations

Here this work is based on the seminar titled ‘Resiliency in Numerical Algorithm Design for Extreme Scale Simulations’ held March 1–6, 2020, at Schloss Dagstuhl, that was attended by all the authors. Advanced supercomputing is characterized by very high computation speeds at the cost of involving an enormous amount of resources and costs. A typical large-scale computation running for 48 h on a system consuming 20 MW, as predicted for exascale systems, would consume a million kWh, corresponding to about 100k Euro in energy cost for executing 10 23 floating-point operations. It is clearly unacceptable to lose the whole computation if any of the several million parallel processes fails during the execution. Moreover, if a single operation suffers from a bit-flip error, should the whole computation be declared invalid? What about the notion of reproducibility itself: should this core paradigm of science be revised and refined for results that are obtained by large-scale simulation? Naive versions of conventional resilience techniques will not scale to the exascale regime: with a main memory footprint of tens of Petabytes, synchronously writing checkpoint data all the way to background storage at frequent intervals will create intolerable overheads in runtime and energy consumption. Forecasts show that the mean time between failures could be lower than the time to recover from such a checkpoint, so that large calculations at scale might not make any progress if robust alternatives are not investigated. More advanced resilience techniques must be devised. The key may lie in exploiting both advanced system features as well as specific application knowledge. Research will face two essential questions: (1) what are the reliability requirements for a particular computation and (2) how do we best design the algorithms and software to meet these requirements? While the analysis of use cases can help understand the particular reliability requirements, the construction of remedies is currently wide open. One avenue would be to refine and improve on system- or application-level checkpointing and rollback strategies in the case an error is detected. Developers might use fault notification interfaces and flexible runtime systems to respond to node failures in an application-dependent fashion. Novel numerical algorithms or more stochastic computational approaches may be required to meet accuracy requirements in the face of undetectable soft errors. These ideas constituted an essential topic of the seminar. The goal of this Dagstuhl Seminar was to bring together a diverse group of scientists with expertise in exascale computing to discuss novel ways to make applications resilient against detected and undetected faults. In particular, participants explored the role that algorithms and applications play in the holistic approach needed to tackle this challenge. This article gathers a broad range of perspectives on the role of algorithms, applications and systems in achieving resilience for extreme scale simulations. The ultimate goal is to spark novel ideas and encourage the development of concrete solutions for achieving such resilience holistically.

79 ASTRONOMY AND ASTROPHYSICS↗

VISILIENCE: An Interactive Visualization Framework for Resilience Analysis using Control-Flow Graph

Soft errors have become one of the major concerns for the error resilience of HPC applications, as those errors can cause HPC applications to generate serious outcomes such as Silent Data Corruptions (SDCs). A large body of approaches has been proposed to analyze the resilience of HPC applications. However, existing studies rarely address the challenges of the analysis result perception. Specifically, resilience analysis techniques often produce a massive volume of unstructured data, making it difficult for programmers to conduct the resilience analysis due to non-intuitive raw data. Furthermore, different analysis models produce diverse results with multiple levels of details, which may create hurdles to compare and explore the resilience of HPC program execution. To this end, we present VISILIENCE, an interactive VISual resILIENCE analysis framework to allow programmers to facilitate the resilience analysis of HPC applications. In particular, VISILIENCE leverages an effective visualization approach Control Flow Graph (CFG) to present a function execution. In addition, three widely-used models for resilience analysis (i.e., Y-Branch, IPAS, and TRIDENT) are seamlessly embedded into the framework for resilience analysis and result comparison. Multiple case studies have been conducted to demonstrate the effectiveness of our proposed framework VISILIENCE.

Jiang, Hailong↗

HAPPA: A Modular Platform for HPC Application Resilience Analysis with LLMs Embedded

High-performance computing (HPC) systems are increasingly vulnerable to soft errors, which pose significant challenges in maintaining computational accuracy and reliability. Predicting the resilience of HPC applications to these errors is crucial for robust code protection and detailed resilience analysis. In this study, we present HAppA, a modular platform designed for HPC Application Resilience Analysis. Embedding Large Language Models (LLMs), HAppA addresses understanding the context information of long code sequences typical in HPC applications. HAppA implements a novel code representation module that chunks the code into fixed-size segments and aggregates the embeddings of these segments. Three aggregation methods have been explored: MeanPooling, MaxPooling, and LSTM-based techniques. We built a DAtaset for REsilience analysis using Fault Injection (FI), named DARE. Using our DARE dataset, HAppA is trained for regression prediction tasks. Our evaluation results demonstrate the predictive accuracy of HAppA compared to other models, particularly noting that the LSTM-based aggregation method -- HAppA-LSTM -- achieves a mean squared error (MSE) of 0.078 for SDC prediction, surpassing the existing state-of-the-art PARIS model, which recorded an MSE of 0.1172. Additionally, HAppA with the KeyBERT model extracts a list of keywords representing the source code. A comprehensive importance analysis of these keywords further elucidates the code patterns contributing to the error rate. These findings highlight the effectiveness of HAppA in analyzing the resilience of HPC applications and establish a new benchmark for predictive accuracy in resilience.

Jiang, Hailong [Kent State University]↗

Prediction of Alpha-Particle-Immune Gate-All-Around Field-Effect Transistors (GAA-FET) Based SRAM Design

Alpha particles are known to be a major source of particles creating soft errors in semiconductor devices, such as content flipping in Static Random-Access Memory (SRAM). Recent advancements in transistor nodes have led to the introduction of Gate-All-Around Field Effect Transistors (GAA-FETs), which have better gate control, thus better electrostatics. Moreover, the introduction of bottom dielectric isolation (BDI) eliminates substrate leakage and thus is expected to enhance its radiation hardness. It is thus important to explore if one can design an SRAM that is completely radiation-hard to alpha particles. In this paper, using 3D Technology Computer-Aided-Design (TCAD) simulations, we show that it is possible to design an SRAM using GAA-FET technology so that it is immune to single alpha particle radiation error. In other words, with the design, there will be no single-event upset (SEU) due to alpha particles. We first use ab initio calculations in PHITS to show that there is a maximum linear energy transfer (LET), LET max , for the alpha particle in Si and Si x Ge 1-x . Based on that, by de signing a sub-7nm GAA-FET-based SRAM with BDI, we show that the SRAM does not flip even if the particle strike is in the worst-case scenario for LET > LET max .

42 ENGINEERING↗

Microcircuit radiation effects databank

This databank is the collation of radiation test data submitted by many testers and serves as a reference for engineers who are concerned with and have some knowledge of the effects of the natural radiation environment on microcircuits. It contains radiation sensitivity results from ground tests and is divided into two sections. Section A lists total dose damage information, and section B lists single event upset cross sections, I.E., the probability of a soft error (bit flip) or of a hard error (latchup).

Source record↗

Microcircuit radiation effects databank

Radiation test data submitted by many testers is collated to serve as a reference for engineers who are concerned with and have some knowledge of the effects of the natural radiation environment on microcircuits. Total dose damage information and single event upset cross sections, i.e., the probability of a soft error (bit flip) or of a hard error (latchup) are presented.

Source record↗

Single event upset vulnerability of selected 4K and 16K CMOS static RAM's

Upset thresholds for bulk CMOS and CMOS/SOS RAMS were deduced after bombardment of the devices with 140 MeV Kr, 160 MeV Ar, and 33 MeV O beams in a cyclotron. The trials were performed to test prototype devices intended for space applications, to relate feature size to the critical upset charge, and to check the validity of computer simulation models. The tests were run on 4 and 1 K memory cells with 6 transistors, in either hardened or unhardened configurations. The upset cross sections were calculated to determine the critical charge for upset from the soft errors observed in the irradiated cells. Computer simulations of the critical charge were found to deviate from the experimentally observed variation of the critical charge as the square of the feature size. Modeled values of series resistors decoupling the inverter pairs of memory cells showed that above some minimum resistance value a small increase in resistance produces a large increase in the critical charge, which the experimental data showed to be of questionable validity unless the value is made dependent on the maximum allowed read-write time.

Kolasinski, W. A.↗

The dependence of single event upset on proton energy /15-590 MeV/

Low earth orbit satellite and Jupiter orbiter probe semiconductor devices may incur soft errors or single event upsets, manifested as bit flips, during exposure to such nuclear particles or heavy ions as trapped protons with energies ranging up to 1000 MeV. Experimental data is given on the average proton fluence needed to cause a bit flip as a function of proton energy for isoplanar bipolar TTL RAMs. Error dependence data shape and threshold energy can be related to the existing body of theoretical data on energy deposition following proton nuclear reactions. Experimental data also show that the relative cross sectional amplitude for functionally identical devices can be related to the device's power consumption.

Nichols, D. K.↗

Experimental determination of single-event upset (SEU) as a function of collected charge in bipolar integrated circuits

Single-Event Upset (SEU) in bipolar integrated circuits (ICs) is caused by charge collection from ion tracks in various regions of a bipolar transistor. This paper presents experimental data which have been obtained wherein the range-energy characteristics of heavy ions (Br) have been utilized to determine the cross section for soft-error generation as a function of charge collected from single-particle tracks which penetrate a bipolar static RAM. The results of this work provide a basis for the experimental verification of circuit-simulation SEU modeling in bipolar ICs.

Zoutendyk, J. A.↗

Accelerators for critical experiments involving single-particle upset in solid-state microcircuits

Charged-particle interactions in microelectronic circuit chips (integrated circuits) present a particularly insidious problem for solid-state electronic systems due to the generation of soft errors or single-particle event upset (SEU) by either cosmic rays or other radiation sources. Particle accelerators are used to provide both light and heavy ions in order to assess the propensity of integrated circuit chips for SEU. Critical aspects of this assessment involve the ability to analytically model SEU for the prediction of error rates in known radiation environments. In order to accurately model SEU, the measurement and prediction of energy deposition in the form of an electron-hole plasma generated along an ion track is of paramount importance. This requires the use of accelerators which allow for ease in both energy control (change of energy) and change of ion species. This and other aspects of ion-beam control and diagnostics (e.g., uniformity and flux) are of critical concern for the experimental verification of theoretical SEU models.

Zoutendyk, J. A.↗

Update on parts SEE suspectibility from heavy ions

JPL and the Aerospace Corporation have collected a fourth set of heavy ion single event effects (SEE) test data. Trends in SEE susceptibility (including soft errors and latchup) for state-of-the-art parts are displayed. All data are conveniently divided into two tables: one for MOS devices, and one for a shorter list of recently tested bipolar devices. In addition, a new table of data for latchup tests only (invariably CMOS processes) is given.

Nichols, D. K.↗

Compiled Data On Single-Event Effects Caused By Heavy Ions

Report presents test data on susceptibility of new set of digital integrated circuits and other semiconductor products to single-event effects (soft errors and latchups) caused by heavy ions incident at high energies. Data used to develop generalizations for protecting electronic equipment from single-event effects. In some cases, tested parts selected as candidates for use in specific applications.

Nichols, Donald K.↗

Proton Effects and Test Issues for Satellite Designers: Ionization Effects - Section 4

This portion of the Short Course is divided into two segments to separately address the two major proton-related effects confronting satellite designers: ionization effects and displacement damage effects. While both of these topics are deeply rooted in "traditional" descriptions of space radiation effects, there are several factors at play to cause renewed concern for satellite systems being designed today. For example, emphasis on Commercial Off-The-Shelf (COTS) technologies in both commercial and government systems increases both Total Ionizing Dose (TID) and Single Event Effect (SEE) concerns. Scaling trends exacerbate the problems, especially with regard to SEEs where protons can dominate soft error rates and even cause destructive failure. In addition, proton-induced displacement damage at fluences encountered in natural space environments can cause degradation in modern bipolar circuitry as well as in many emerging electronic and opto-electronic technologies. A crude, but nevertheless telling, indication of the level of concern for proton effects follows from surveying the themes treated in papers presented at this conference. The table lists themes found in the IEEE Transaction on Nuclear Science (TNS) December issue from the past year and compares them with the December issue's content a decade earlier. Ten years ago there were nine papers, or about 10% of the total, dealing with the four indicated topics. At that time, single event effects from protons were the primary concern, and these were thought to be possible only when a nuclear reaction initiated energetic recoil atoms. This is shown in the table as the 'traditional" SEE subject. A decade later, submissions addressing this topic had doubled, while papers devoted to displacement damage studies had increased from one to nine! More importantly, displacement damage effects in the natural space environments have become a concern for degradation in modern devices (other than solar cells), and this was not so ten years earlier.

Marshall, Paul W.↗

1999 NSREC Short Course: Proton Effects and Test Issues for Satellite Designers: Displacement Effects

This portion of the Short Course is divided into two segments to separately address the two major proton-related effects confronting satellite designers: ionization effects and displacement damage effects. While both of these topics are deeply rooted in "traditional" descriptions of space radiation effects, there are several factors at play to cause renewed concern for satellite systems being designed today. For example, emphasis on Commercial Off-The-Shelf (COTS) technologies in both commercial and government systems increases both Total Ionizing Dose (TID) and Single Event Effect (SEE) concerns. Scaling trends exacerbate the problems, especially with regard to SEEs where protons can dominate soft error rates and even cause destructive failure. In addition, proton-induced displacement damage at fluences encountered in natural space environments can cause degradation in modern bipolar circuitry as well as in many emerging electronic and opto-electronic technologies.

Marshall, Cheryl J.↗

Overview of Device SEE Susceptibility from Heavy Ions

A fifth set of heavy ion single event effects (SEE) test data have been collected since the last IEEE publications (1,2,3,4) in December issues for 1985, 1987, 1989, and 1991. Trends in SEE susceptibility (including soft errors and latchup) for state-of-the-art parts are evaluated.

Nichols, D. K.↗

Reconfigurable Processing Module

To accommodate a wide spectrum of applications and technologies, NASA s Exploration System's Missions Directorate has called for reconfigurable and modular technologies to support future missions to the moon and Mars. In response, Langley Research Center is leading a program entitled Reconfigurable Scaleable Computing (RSC) that is centered on the development of FPGA-based computing resources in a stackable form factor. This paper details the architecture and implementation of the Reconfigurable Processing Module (RPM), which is the key element of the RSC system. The RPM is an FPGA-based, space-qualified printed circuit assembly leveraging terrestrial/commercial design standards into the space applications domain. The form factor is similar to, and backwards compatible with, the PCI-104 standard utilizing only the PCI interface. The size is expanded to accommodate the required functionality while still better than 30% smaller than a 3U CompactPCI(TradeMark)card and without the overhead of the backplane. The architecture is built around two FPGA devices, one hosting PCI and memory interfaces, and another hosting mission application resources; both of which are connected with a high-speed data bus. The PCI interface FPGA provides access via the PCI bus to onboard SDRAM, flash PROM, and the application resources; both configuration management as well as runtime interaction. The reconfigurable FPGA, referred to as the Application FPGA - or simply "the application" - is a radiation-tolerant Xilinx Virtex-4 FX60 hosting custom application specific logic or soft microprocessor IP. The RPM implements various SEE mitigation techniques including TMR, EDAC, and configuration scrubbing of the reconfigurable FPGA. Prototype hardware and formal modeling techniques are used to explore the performability trade space. These models provide a novel way to calculate quality-of-service performance measures while simultaneously considering fault-related behavior due to SEE soft errors.

Somervill, Kevin↗

Impact of Scaled Technology on Radiation Testing and Hardening

This presentation gives a brief overview of some of the radiation challenges facing emerging scaled digital technologies with implications on using consumer grade electronics and next generation hardening schemes. Commercial semiconductor manufacturers are recognizing some of these issues as issues for terrestrial performance. Looking at means of dealing with soft errors. The thinned oxide has indicated improved TID tolerance of commercial products hardened by "serendipity" which does not guarantee hardness or say if the trend will continue. This presentation also focuses one reliability implications of thinned oxides.

LaBel, Kenneth A.↗