Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Tolerance Bounds”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Reliability of voting in fault-tolerant software systems for small output spaces

Under a voting strategy in a fault-tolerant software system there is a difference between correctness and agreement. An independent N-version programming reliability model is proposed for treating small output spaces which distinguishes between correctness and agreement. System reliability is investigated using analytical relationships and simulation. A consensus majority voting strategy is proposed and its performance is analyzed and compared with other voting strategies. Consensus majority strategy automatically adapts the voting to different component reliability and output space cardinality characteristics. It is shown that absolute majority voting strategy provides a lower bound on the reliability provided by the consensus majority, and 2-of-n voting strategy an upper bound. If r is the cardinality of the output space it is proved the 1/r is a lower bound on the average reliability of fault-tolerant system components below which the system reliability begins to deteriorate as more versions are added.

Mcallister, David F.↗

Reliability of voting in fault-tolerant software systems for small output spaces

Under a voting strategy in a fault-tolerant software system there is a difference between correctness and agreement. An independent N-version programming reliability model is proposed for treating small output spaces which distinguishes between correctness and agreement. System reliability is investigated using analytical relationships and simulation. A consensus majority voting stratey is proposed and its performance is analyzed and compared with other voting strategies. A consensus voting strategy automatically adapts the voting to diffeerent component reliability and output space cardinality characteristics. It is shown that absolute majority voting strategy provides a lower bound on the reliability provided by the consensus majority, and the 2-of-n voting strategy an upper bound. If r is the cardinality of output space it is proved that 1/r is a lower bound on the average reliability of fault-tolerant system components below which the system reliability begins to deteriorate as more versions are added.

Mcallister, David F.↗

Control of linear uncertain systems utilizing mismatched state observers

The control of linear continuous dynamical systems is investigated as a problem of limited state feedback control. The equations which describe the structure of an observer are developed constrained to time-invarient systems. The optimal control problem is formulated, accounting for the uncertainty in the design parameters. Expressions for bounds on closed loop stability are also developed. The results indicate that very little uncertainty may be tolerated before divergence occurs in the recursive computation algorithms, and the derived stability bound yields extremely conservative estimates of regions of allowable parameter variations.

Goldstein, B.↗

Structural and functional characterization of NEMO cleavage by SARS-CoV-2 3CLpro

Abstract In addition to its essential role in viral polyprotein processing, the SARS-CoV-2 3C-like protease (3CLpro) can cleave human immune signaling proteins, like NF-κB Essential Modulator (NEMO) and deregulate the host immune response. Here, in vitro assays show that SARS-CoV-2 3CLpro cleaves NEMO with fine-tuned efficiency. Analysis of the 2.50 Å resolution crystal structure of 3CLpro C145S bound to NEMO 226–234 reveals subsites that tolerate a range of viral and host substrates through main chain hydrogen bonds while also enforcing specificity using side chain hydrogen bonds and hydrophobic contacts. Machine learning- and physics-based computational methods predict that variation in key binding residues of 3CLpro-NEMO helps explain the high fitness of SARS-CoV-2 in humans. We posit that cleavage of NEMO is an important piece of information to be accounted for, in the pathology of COVID-19.

60 APPLIED LIFE SCIENCES↗

Stochastic Error Cancellation in Analog Quantum Simulation

Analog quantum simulation is a promising path towards solving classically intractable problems in many-body physics on near-term quantum devices. However, the presence of noise limits the size of the system and the length of time that can be simulated. In our work, we consider an error model in which the actual Hamiltonian of the simulator differs from the target Hamiltonian we want to simulate by small local perturbations, which are assumed to be random and unbiased. We analyze the error accumulated in observables in this setting and show that, due to stochastic error cancellation, with high probability the error scales as the square root of the number of qubits instead of linearly. We explore the concentration phenomenon of this error as well as its implications for local observables in the thermodynamic limit. Moreover, we show that stochastic error cancellation also manifests in the fidelity between the target state at the end of time-evolution and the actual state we obtain in the presence of noise. This indicates that, to reach a certain fidelity, more noise can be tolerated than implied by the worst-case bound if the noise comes from many statistically independent sources.

Analog quantum simulation↗

Fault-tolerant wait-free shared objects

A concurrent system consists of processes and shared objects. Previous research focused on the problem of tolerating process failure. We study the complementary problem of tolerating failures. We divide object failures into two broad classes: responsive and non-responsive. With responsive failures, a faulty object responds to every invocation, but responses may be incorrect. With non-responsive failures, a faulty object may also 'hang' without responding. For each class, we consider crash, and arbitrary types of failures. For each type of failure, we are seeking a universal implementation for fault-tolerant wait-free shared objects. We present (deterministic) implementations for all types of responsive failures, including arbitrary failures. In contrast, we show that even the most benign type of non-responsive failures requires the use of randomization. Of special interest is the problem of implementing fault-tolerant objects using only objects of the same type. We present such fault-tolerant self-implementations for many common object types. Graceful degradation is a desirable property of fault-tolerant implementations: the implemented object never fails more severely than the base objects it is derived from, even if all the base objects fail. For several failure models, we show whether this property can be achieved, and, if so, how. In addition to the above possibility/impossibility results, we also consider the resources complexity of fault-tolerant implementations. In many cases, we present lower bounds and give matching algorithms.

Jayanti, Prasad↗

Requirements Flowdown for Prognostics and Health Management

Prognostics and Health Management (PHM) principles have considerable promise to change the game of lifecycle cost of engineering systems at high safety levels by providing a reliable estimate of future system states. This estimate is a key for planning and decision making in an operational setting. While technology solutions have made considerable advances, the tie-in into the systems engineering process is lagging behind, which delays fielding of PHM-enabled systems. The derivation of specifications from high level requirements for algorithm performance to ensure quality predictions is not well developed. From an engineering perspective some key parameters driving the requirements for prognostics performance include: (1) maximum allowable Probability of Failure (PoF) of the prognostic system to bound the risk of losing an asset, (2) tolerable limits on proactive maintenance to minimize missed opportunity of asset usage, (3) lead time to specify the amount of advanced warning needed for actionable decisions, and (4) required confidence to specify when prognosis is sufficiently good to be used. This paper takes a systems engineering view towards the requirements specification process and presents a method for the flowdown process. A case study based on an electric Unmanned Aerial Vehicle (e-UAV) scenario demonstrates how top level requirements for performance, cost, and safety flow down to the health management level and specify quantitative requirements for prognostic algorithm performance.

video acuity↗

Modifying the Asynchronous Jacobi Method for Data Corruption Resilience

Moving scientific computation from high-performance computing (HPC) and cloud computing (CC) environments to devices on the edge, i.e., physically near instruments of interest, has received tremendous interest in recent years. Such edge computing environments can operate on data in situ, offering enticing benefits over data aggregation to HPC and CC facilities that include avoiding costs of transmission, increased data privacy, and real-time data analysis. Because of the inherent unreliability of edge computing environments, new fault-tolerant approaches must be developed before the benefits of edge computing can be realized. Motivated by algorithm-based fault tolerance, a variant of the asynchronous Jacobi (ASJ) method is developed that achieves resilience to data corruption by rejecting solution approximations from neighbor devices according to a bound derived from convergence theory. Numerical results on a two-dimensional Poisson problem show that the new rejection criterion, along with a novel approximation to the shortest path length on which the criterion depends, restores convergence for the ASJ variant in the presence of certain types data corruption. Numerical results are obtained for when the singular values in the analytic bound are approximated. Additional linear systems are also explored, one with a more dense sparsity pattern and one that includes advection. All results indicate that successful resilience to data corruption depends on whether the bound tightens fast enough to reject corrupted data before the iteration evolution deviates significantly from that predicted by the convergence theory defining the bound. This observation generalizes to future work on algorithm-based fault tolerance for other asynchronous algorithms, including upcoming approaches that leverage Krylov subspaces.

97 MATHEMATICS AND COMPUTING↗

An Analysis of the Johnson-Lindenstrauss Lemma with the Bivariate Gamma Distribution

Probabilistic proofs of the Johnson-Lindenstrauss lemma imply that random projection can reduce the dimension of a data set and approximately preserve pairwise distances. If a distance being approximately preserved is called a success, and the complement of this event is called a failure, then such a random projection likely results in no failures. Assuming a Gaussian random projection, the lemma is proved by showing that the no-failure probability is positive using a combination of Bonferroni's inequality and Markov's inequality. This paper modifies this proof in two ways to obtain a greater lower bound on the no-failure probability. First, Bonferroni's inequality is applied to pairs of failures instead of individual failures. Second, since a pair of projection errors has a bivariate gamma distribution, this probability of a pair of successes is bounded using an inequality from [Jensen, 1969]. If n is the number of points to be embedded and μ is the probability of success, then this leads to an increase in the lower bound on the no-failure probability of $\frac{1}{2}$ ($\genfrac{}{}{0pt}{}{n}{2}$) (1- μ ) 2 is ($\genfrac{}{}{0pt}{}{n}{2}$) is even and $\frac{1}{2}$ (($\genfrac{}{}{0pt}{}{n}{2}$)-1) (1- μ ) 2 if ($\genfrac{}{}{0pt}{}{n}{2}$) is odd. For example, if n =10 5 points are to be embedded in k =10 4 dimensions with a tolerance of ϵ=0.1, then the improvement in the lower bound is on the order of 10 -14 . We also show that further improvement is possible if the inequality in [Jensen, 1969] extends to three successes, though we do not have a proof of this result.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

ROBUS-2: A Fault-Tolerant Broadcast Communication System

The Reliable Optical Bus (ROBUS) is the core communication system of the Scalable Processor-Independent Design for Enhanced Reliability (SPIDER), a general-purpose fault-tolerant integrated modular architecture currently under development at NASA Langley Research Center. The ROBUS is a time-division multiple access (TDMA) broadcast communication system with medium access control by means of time-indexed communication schedule. ROBUS-2 is a developmental version of the ROBUS providing guaranteed fault-tolerant services to the attached processing elements (PEs), in the presence of a bounded number of faults. These services include message broadcast (Byzantine Agreement), dynamic communication schedule update, clock synchronization, and distributed diagnosis (group membership). The ROBUS also features fault-tolerant startup and restart capabilities. ROBUS-2 is tolerant to internal as well as PE faults, and incorporates a dynamic self-reconfiguration capability driven by the internal diagnostic system. This version of the ROBUS is intended for laboratory experimentation and demonstrations of the capability to reintegrate failed nodes, dynamically update the communication schedule, and tolerate and recover from correlated transient faults.

Torres-Pomales, Wilfredo↗

Design of the Protocol Processor for the ROBUS-2 Communication System

The ROBUS-2 Protocol Processor (RPP) is a custom-designed hardware component implementing the functionality of the ROBUS-2 fault-tolerant communication system. The Reliable Optical Bus (ROBUS) is the core communication system of the Scalable Processor-Independent Design for Enhanced Reliability (SPIDER), a general-purpose fault tolerant integrated modular architecture currently under development at NASA Langley Research Center. ROBUS is a time-division multiple access (TDMA) broadcast communication system with medium access control by means of time-indexed communication schedule. ROBUS-2 is a developmental version of the ROBUS providing guaranteed fault-tolerant services to the attached processing elements (PEs), in the presence of a bounded number of faults. These services include message broadcast (Byzantine Agreement), dynamic communication schedule update, time reference (clock synchronization), and distributed diagnosis (group membership). ROBUS also features fault-tolerant startup and restart capabilities. ROBUS-2 tolerates internal as well as PE faults, and incorporates a dynamic self-reconfiguration capability driven by the internal diagnostic system. ROBUS consists of RPPs connected to each other by a lower-level physical communication network. The RPP has a pipelined architecture and the design is parameterized in the behavioral and structural domains. The design of the RPP enables the bus to achieve a PE-message throughput that approaches the available bandwidth at the physical layer.

Torres-Pomales, Wilfredo↗

Structural insights into protection against a SARS-CoV-2 spike variant by T cell receptor diversity

T cells play a crucial role in combatting SARS-CoV-2 and forming long-term memory responses to this coronavirus. The emergence of SARS-CoV-2 variants that can evade T cell immunity has raised concerns about vaccine efficacy and the risk of reinfection. Some SARS-CoV-2 T cell epitopes elicit clonally restricted CD8 + T cell responses characterized by T cell receptors (TCRs) that lack structural diversity. Mutations in such epitopes can lead to loss of recognition by most T cells specific for that epitope, facilitating viral escape. Here, we studied an HLA-A2–restricted spike protein epitope (RLQ) that elicits CD8 + T cell responses in COVID-19 convalescent patients characterized by highly diverse TCRs. We previously reported the structure of an RLQ-specific TCR (RLQ3) with greatly reduced recognition of the most common natural variant of the RLQ epitope (T1006I). Opposite to RLQ3, TCR RLQ7 recognizes T1006I with even higher functional avidity than the WT epitope. To explain the ability of RLQ7, but not RLQ3, to tolerate the T1006I mutation, we determined structures of RLQ7 bound to RLQ–HLA-A2 and T1006I–HLA-A2. These complexes show that there are multiple structural solutions to recognizing RLQ and thereby generating a clonally diverse T cell response to this epitope that assures protection against viral escape and T cell clonal loss.

60 APPLIED LIFE SCIENCES↗

Fault tolerant sequential machines

Memory cell fault tolerant sequential machine synthesis, considering masking feasibility and lower bounds on minimum redundancy

Meyer, J. F.↗

The relationship between water binding and desiccation tolerance in tissues

In an effort to define the nature of desiccation tolerance, a comparison of the water sorption characteristics was made between tissues that were resistant and tissues that were sensitive to desiccation. Water sorption isotherms were constructed for germinated and ungerminated soybean axes and also for fronds of several species of Polypodium with varying tolerance to dehydration. The strength of water binding was determined by van't Hoff as well as D'Arcy/Watt analyses of the isotherms at 5, 15, and/or 25 degrees C. Tissues which were sensitive to desiccation had a poor capacity to bind water tightly. Tightly bound water can be removed from soybean and pea seeds by equilibration at 35 degrees C over very low relative humidities; this results in a reduction in the viability of the seed. We suggest that region 1 water (i.e. water bound with very negative enthalpy values) is an important component of desiccation tolerance.

NASA Discipline Number 40-10↗

Validation of the SURE Program, phase 1

Presented are the results of the first phase in the validation of the SURE (Semi-Markov Unreliability Range Evaluator) program. The SURE program gives lower and upper bounds on the death-state probabilities of a semi-Markov model. With these bounds, the reliability of a semi-Markov model of a fault-tolerant computer system can be analyzed. For the first phase in the validation, fifteen semi-Markov models were solved analytically for the exact death-state probabilities and these solutions compared to the corresponding bounds given by SURE. In every case, the SURE bounds covered the exact solution. The bounds, however, had a tendency to separate in cases where the recovery rate was slow or the fault arrival rate was fast.

Dotson, Kelly J.↗

Lieb-Robinson Bounds with Exponential-in-Volume Tails

Lieb-Robinson bounds demonstrate the emergence of locality in many-body quantum systems. Intuitively, Lieb-Robinson bounds state that, with local or exponentially decaying interactions, the correlation that can be built up between two sites separated by distance 𝑟 after a time 𝑡 decays as exp (𝑣⁢𝑡 −𝑟), where 𝑣 is the emergent Lieb-Robinson velocity. In many problems, it is important to also capture how much of an operator grows to act on 𝑟 𝑑 sites in 𝑑 spatial dimensions. Perturbation theory and cluster expansion methods suggest that, at short times, these volume-filling operators are suppressed as exp (−𝑟 𝑑 ). We confirm this intuition, showing that, for 𝑟 >𝑣⁢𝑡, the volume-filling operator is suppressed by exp (−(𝑟−𝑣⁢𝑡) 𝑑 /(𝑣⁢𝑡) 𝑑−1 ). This closes a conceptual and practical gap between the cluster expansion and the Lieb-Robinson bound. We then present two very different applications of this new bound. Firstly, we obtain improved bounds on the classical computational resources necessary to simulate many-body dynamics with error tolerance 𝜀 for any finite time 𝑡: as 𝜀 becomes sufficiently small, only 𝜀 −O⁡(𝑡 𝑑−1 ) resources are needed. A protocol that likely saturates this bound is given. Secondly, we prove that disorder operators have volume-law suppression near the “solvable (Ising) point” in quantum phases with spontaneous symmetry breaking, which implies a new diagnostic for distinguishing many-body phases of quantum matter.

computational complexity↗

Robustness of Vacancy-Bound Non-Abelian Anyons in the Kitaev Model in a Magnetic Field

Non-Abelian anyons in quantum spin liquids (QSLs) provide a promising route to fault-tolerant topological quantum computation. In the exactly solvable Kitaev honeycomb model, such anyons of the QSL state can be bound to nonmagnetic spin vacancies and endowed with non-Abelian statistics by an infinitesimal magnetic field. Here, we investigate how this approach for stabilizing non-Abelian anyons extends to a finite magnetic field represented by a proper Zeeman term. Through large-scale density-matrix renormalization group simulations, we compute the vacancy-anyon binding energy as a function of magnetic field for both the ferromagnetic and antiferromagnetic Kitaev models. Here, we find that anyon binding remains robust within the entire QSL phase for the ferromagnetic Kitaev model but breaks down already inside this phase for the antiferromagnetic Kitaev model. To compute a binding energy several orders of magnitude below the magnetic energy scale, we introduce both a refined definition and an extrapolation scheme based on carefully tailored perturbations.

Xiao, Bo [Oak Ridge National Laboratory (ORNL), Oa↗

Truncated Gaussians as tolerance sets

This work focuses on the use of truncated Gaussian distributions as models for bounded data measurements that are constrained to appear between fixed limits. The authors prove that the truncated Gaussian can be viewed as a maximum entropy distribution for truncated bounded data, when mean and covariance are given. The characteristic function for the truncated Gaussian is presented; from this, algorithms are derived for calculation of mean, variance, summation, application of Bayes rule and filtering with truncated Gaussians. As an example of the power of their methods, a derivation of the disparity constraint (used in computer vision) from their models is described. The authors' approach complements results in Statistics, but their proposal is not only to use the truncated Gaussian as a model for selected data; they propose to model measurements as fundamentally in terms of truncated Gaussians.

Cozman, Fabio↗