Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “rollback”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Self-checking self-repairing computer nodes using the mirror processor

Circuitry added to fault-tolerant systems for concurrent error deduction usually reduces performance. Using a technique called micro rollback, it is possible to eliminate most of the performance penalty of concurrent error detection. Error detection is performed in parallel with intermodule communication, and erroneous state changes are later undone. The author reports on the design and implementation of a VLSI RISC microprocessor, called the Mirror Processor (MP), which is capable of micro rollback. In order to achieve concurrent error detection, two MP chips operate in lockstep, comparing external signals and a signature of internal signals every clock cycle. If a mismatch is detected, both processors roll back to the beginning of the cycle when the error occurred. In some cases the erroneous state is corrected by copying a value from the fault-free processor to the faulty processor. The architecture, microarchitecture, and VLSI implementation of the MP, emphasizing its error-detection, error-recovery, and self-diagnosis capabilities, are described.

Tamir, Yuval↗

Read buffer optimizations to support compiler-assisted multiple instruction retry

Multiple instruction retry is a recovery mechanism for transient processor faults. We previously developed a compiler-assisted approach to multiple instruction ferry in which a read buffer of size 2N (where N represents the maximum instruction rollback distance) was used to resolve some data hazards while the compiler resolved the remaining hazards. The compiler-assisted scheme was shown to reduce the performance overhead and/or hardware complexity normally associated with hardware-only retry schemes. This paper examines the size and design of the read buffer. We establish a practical lower bound and average size requirement for the read buffer by modifying the scheme to save only the data required for rollback. The study measures the effect on the performance of a DECstation 3100 running ten application programs using six read buffer configurations with varying read buffer sizes. Two alternative configurations are shown to be the most efficient and differed depending on whether split-cycle-saves are assumed. Up to a 55 percent read buffer size reduction is achievable with an average reduction of 39 percent given the most efficient read buffer configuration and a variety of applications.

Alewine, N. J.↗

Experimental evaluation of multiprocessor cache-based error recovery

Several variations of cache-based checkpointing for rollback error recovery in shared-memory multiprocessors have been recently developed. By modifying the cache replacement policy, these techniques use the inherent redundancy in the memory hierarchy to periodically checkpoint the computation state. Three schemes, different in the manner in which they avoid rollback propagation, are evaluated. By simulation with address traces from parallel applications running on an Encore Multimax shared-memory multiprocessor, the performance effect of integrating the recovery schemes in the cache coherence protocol are evaluated. The results indicate that the cache-based schemes can provide checkpointing capability with low performance overhead but uncontrollable high variability in the checkpoint interval.

Janssens, Bob↗

Implementing forward recovery using checkpointing in distributed systems

The paper describes the implementation of a forward recovery scheme using checkpoints and replicated tasks. The implementation is based on the concept of lookahead execution and rollback validation. In the experiment, two tasks are selected for the normal execution and one for rollback validation. It is shown that the recovery strategy has nearly error-free execution time and an average redundancy lower than TMR.

Long, Junsheng↗

STS-39 Compiled Orbiter Footage

Live footage shows the rollback of STS-39 to the VAB (Vehicle Assembly Building), the rollback of Discovery to the OPF (Orbiter Processing Facility) High Bay 2, Discovery ET Disconnect Door Hinges (Cracks), Discovery ET Disconnect Door Hinges (Edited) and Discovery in the VAB.

Source record↗

Orbiter Processing Overview

Footage of orbiter processing is shown. The Space Shuttle Discovery Landing Operations, Rollback to Orbiter Processing Facility After Landing, Orbiter Processing Facility, Vehicle Assembly Building (VAB), Orbiter Rollover from the Orbiter Processing Facility to the VAB, Atlantis Orbiter Lift and Mate, Atlantis Rollout to Launch Pad 39, Payload Canister Door Opened and Payload Move into Payload Ground Handling Mechanism, Payload into Orbiter, Orbiter Payload Bay Door Closure, Launch Pad Processing at Launch Complex 39, Rotating Service Structure Rollback at Launch Pad, Discovery Lift-off and SRB separation is presented.

Source record↗

Preliminary Results From a Heavily Instrumented Engine Ice Crystal Icing Test in a Ground Based Altitude Test Facility

Preliminary results from the heavily instrumented ALF502R-5 engine test conducted in the NASA Glenn Research Center Propulsion Systems Laboratory are discussed. The effects of ice crystal icing on a full scale engine is examined and documented. This same model engine, serial number LF01, was used during the inaugural icing test in the Propulsion Systems Laboratory facility. The uncommanded reduction of thrust (rollback) events experienced by this engine in flight were simulated in the facility. Limited instrumentation was used to detect icing on the LF01 engine. Metal temperatures on the exit guide vanes and outer shroud and the load measurement were the only indicators of ice formation. The current study features a similar engine, serial number LF11, which is instrumented to characterize the cloud entering the engine, detect/characterize ice accretion, and visualize the ice accretion in the region of interest. Data were acquired at key LF01 test points and additional points that explored: icing threshold regions, low altitude, high altitude, spinner heat effects, and the influence of varying the facility and engine parameters. For each condition of interest, data were obtained from some selected variations of ice particle median volumetric diameter, total water content, fan speed, and ambient temperature. For several cases the NASA in-house engine icing risk assessment code was used to find conditions that would lead to a rollback event. This study further helped NASA develop necessary icing diagnostic instrumentation, expand the capabilities of the Propulsion Systems Laboratory, and generate a dataset that will be used to develop and validate in-house icing prediction and risk mitigation computational tools. The ice accretion on the outer shroud region was acquired by internal video cameras. The heavily instrumented engine showed good repeatability of icing responses when compared to the key LF01 test points and during day-to-day operation. Other noticeable observations are presented.

Enigine Icing↗

Preliminary Results From a Heavily Instrumented Engine Ice Crystal Icing Test in a Ground Based Altitude Test Facility

Preliminary results from the heavily instrumented ALF502R-5 engine test conducted in the NASA Glenn Research Center Propulsion Systems Laboratory are discussed. The effects of ice crystal icing on a full scale engine is examined and documented. This same model engine, serial number LF01, was used during the inaugural icing test in the Propulsion Systems Laboratory facility. The uncommanded reduction of thrust (rollback) events experienced by this engine in flight were simulated in the facility. Limited instrumentation was used to detect icing on the LF01 engine. Metal temperatures on the exit guide vanes and outer shroud and the load measurement were the only indicators of ice formation. The current study features a similar engine, serial number LF11, which is instrumented to characterize the cloud entering the engine, detect/ characterize ice accretion, and visualize the ice accretion in the region of interest. Data were acquired at key LF01 test points and additional points that explored: icing threshold regions, low altitude, high altitude, spinner heat effects, and the influence of varying the facility and engine parameters. For each condition of interest, data were obtained from some selected variations of ice particle median volumetric diameter, total water content, fan speed, and ambient temperature. For several cases the NASA in-house engine icing risk assessment code was used to find conditions that would lead to a rollback event. This study further helped NASA develop necessary icing diagnostic instrumentation, expand the capabilities of the Propulsion Systems Laboratory, and generate a dataset that will be used to develop and validate in-house icing prediction and risk mitigation computational tools. The ice accretion on the outer shroud region was acquired by internal video cameras. The heavily instrumented engine showed good repeatability of icing responses when compared to the key LF01 test points and during day-to-day operation. Other noticeable observations are presented.

turbomachinery↗

Modeling of a Turbofan Engine with Ice Crystal Ingestion in the NASA Propulsion System Laboratory

The main focus of this study is to apply a computational tool for the flow analysis of the turbine engine that has been tested with ice crystal ingestion in the Propulsion Systems Laboratory (PSL) at NASA Glenn Research Center. The PSL has been used to test a highly instrumented Honeywell ALF502R-5A (LF11) turbofan engine at simulated altitude operating conditions. Test data analysis with an engine cycle code and a compressor flow code was conducted to determine the values of key icing parameters, that can indicate the risk of ice accretion, which can lead to engine rollback (un-commanded loss of engine thrust). The full engine aerothermodynamic performance was modeled with the Honeywell Customer Deck specifically created for the ALF502R-5A engine. The mean-line compressor flow analysis code, which includes a code that models the state of the ice crystal, was used to model the air flow through the fan-core and low pressure compressor. The results of the compressor flow analyses included calculations of the ice-water flow rate to air flow rate ratio (IWAR), the local static wet bulb temperature, and the particle melt ratio throughout the flow field. It was found that the assumed particle size had a large effect on the particle melt ratio, and on the local wet bulb temperature. In this study the particle size was varied parametrically to produce a non-zero calculated melt ratio in the exit guide vane (EGV) region of the low pressure compressor (LPC) for the data points that experienced a growth of blockage there, and a subsequent engine called rollback (CRB). At data points where the engine experienced a CRB having the lowest wet bulb temperature of 492 degrees Rankine at the EGV trailing edge, the smallest particle size that produced a non-zero melt ratio (between 3 percent - 4 percent) was on the order of 1 micron. This value of melt ratio was utilized as the target for all other subsequent data points analyzed, while the particle size was varied from 1 micron - 9.5 microns to achieve the target melt ratio. For data points that did not experience a CRB which had static wet bulb temperatures in the EGV region below 492 degrees Rankine, a non-zero melt ratio could not be achieved even with a 1 micron ice particle size. The highest value of static wet bulb temperature for data points that experienced engine CRB was 498 degrees Rankine with a particle size of 9.5 microns. Based on this study of the LF11 engine test data, the range of static wet bulb temperature at the EGV exit for engine CRB was in the narrow range of 492 degrees Rankine - 498 degrees Rankine , while the minimum value of IWAR was 0.002. The rate of blockage growth due to ice accretion and boundary layer growth was estimated by scaling from a known blockage growth rate that was determined in a previous study. These results obtained from the LF11 engine analysis formed the basis of a unique “icing wedge.”

Turbo engine↗

Ice Crystal Icing Research at NASA

Ice crystals found at high altitude near convective clouds are known to cause jet engine power-loss events. These events occur due to ice crystals entering a propulsion systems core flowpath and accreting ice resulting in events such as uncommanded loss of thrust (rollback), engine stall, surge, and damage due to ice shedding. As part of a community with a growing need to understand the underlying physics of ice crystal icing, NASA has been performing experimental efforts aimed at providing datasets that can be used to generate models to predict the ice accretion inside current and future engine designs. Fundamental icing physics studies on particle impacts, accretion on a single airfoil, and ice accretions observed during a rollback event inside a full-scale engine in the Propulsion Systems Laboratory are summarized. Low fidelity code development using the results from the engine tests which identify key parameters for ice accretion risk and the development of high fidelity codes are described. These activities have been conducted internal to NASA and through collaboration efforts with industry, academia, and other government agencies. The details of the research activities and progress made to date in addressing ice crystal icing research challenges are discussed.

icing↗

Ice Crystal Icing Research at NASA

Ice crystals found at high altitude near convective clouds are known to cause jet engine power-loss events. These events occur due to ice crystals entering a propulsion systems core flowpath and accreting ice resulting in events such as uncommanded loss of thrust (rollback), engine stall, surge, and damage due to ice shedding. As part of a community with a growing need to understand the underlying physics of ice crystal icing, NASA has been performing experimental efforts aimed at providing datasets that can be used to generate models to predict the ice accretion inside current and future engine designs. Fundamental icing physics studies on particle impacts, accretion on a single airfoil, and ice accretions observed during a rollback event inside a full-scale engine in the Propulsion Systems Laboratory are summarized. Low fidelity code development using the results from the engine tests which identify key parameters for ice accretion risk and the development of high fidelity codes are described. These activities have been conducted internal to NASA and through collaboration efforts with industry, academia, and other government agencies. The details of the research activities and progress made to date in addressing ice crystal icing research challenges are discussed.

upwind schemes↗

Ice Crystal Icing Research at NASA

Ice crystals found at high altitude near convective clouds are known to cause jet engine power-loss events. These events occur due to ice crystals entering a propulsion system's core flowpath and accreting ice resulting in events such as uncommanded loss of thrust (rollback), engine stall, surge, and damage due to ice shedding. As part of a community with a growing need to understand the underlying physics of ice crystal icing, NASA has been performing experimental efforts aimed at providing datasets that can be used to generate models to predict the ice accretion inside current and future engine designs. Fundamental icing physics studies on particle impacts, accretion on a single airfoil, and ice accretions observed during a rollback event inside a full-scale engine in the Propulsion Systems Laboratory are summarized. Low fidelity code development using the results from the engine tests which identify key parameters for ice accretion risk and the development of high fidelity codes are described. These activities have been conducted internal to NASA and through collaboration efforts with industry, academia, and other government agencies. The details of the research activities and progress made to date in addressing ice crystal icing research challenges are discussed.

crysta↗

Virtual Time III, Part 3: Throttling and Message Cancellation

This is Part 3 of a trio of papers that unify in a natural way the two historically distinct parallel discrete event synchronization paradigms, optimistic and conservative, combining the best properties of both into a single framework called Unified Virtual Time (UVT). In this part, we survey the synchronization effects that can be achieved by restricting to corner cases the relationships permitted among the control variables, GVT, CVT, TVT, and LVT, which were defined in Part 1. Here we also survey various throttling policies from the literature and describe how they can be implemented in UVT by controlling the value of TVT, including policies that can take advantage of rollback in addition to LP blocking. A significant result is a new category of efficient and higher precision throttling algorithms for optimistic execution that are based on optimistic lookahead, defined in a way that is symmetric to what we now call the conservative lookahead information that is traditionally used for conservative synchronization. Finally, we present a novel algorithm allowing the choice between lazy and aggressive cancellation to be made on a message-by-message basis using either external logic expressed in the model code, or policy code internal to the simulator, or a mixture of both.

throttling↗

Towards Low-Overhead Resilience for Data Parallel Deep Learning

Data parallel techniques have been widely adopted both in academia and industry as a tool to enable scalable training of deep learning models. At scale, DL training jobs can fail due to software or hardware bugs, may need to be preempted or terminated due to unexpected events, or may perform suboptimally because they were misconfigured. Under such circumstances, there is a need to recover and/or reconfigure data-parallel DL training jobs on-the-fly, while minimizing the impact on the accuracy of the DNN model and the runtime overhead. In this regard, state-of-art techniques adopted by the HPC community mostly rely on checkpoint-restart, which inevitably leads to loss of progress, thus increasing the runtime overhead. In this paper we explore alternative techniques that exploit the properties of modern deep learning frameworks (overlapping of gradient averaging and weight updates with local gradient computations through pipeline parallelism) to reduce the overhead of resilience/elasticity. To this end we introduce a failure simulation framework and two resilience strategies (immediate mini-batch rollback and lossy forward recovery), which we study compared with checkpoint-restart approaches in a variety of settings in order to understand the trade-offs between the accuracy loss of the DNN model and the runtime overhead.

data-parallel training↗

SPADES (Scalable Parallel Discrete Events Simulation) [SWR-24-99]

SPADES (Solver for PArallel Discrete Event Simulation) is an open-source parallel discrete event simulation (PDES) package built on the AMReX library. Targeted at solving discrete event systems in parallel, this software package aims to be performance portable and scalable on heterogeneous computing architectures, e.g., graphic processing units (GPU). SPADES implements optimistic synchronization with rollback through an implementation of the Time Warp algorithm. An alternative conservative synchronization approach is also implemented using the Lower Bound on Incoming Time Stamp. In our implementation, logical processes are represented as cells in a grid and event messages are represented as particles. SPADES supports various parallel decomposition strategies, including the use of the Message Passing Interface (MPI) and OpenMP threading. All major GPU architectures (e.g., Intel, AMD, NVIDIA) are supported through the use of performance portability functionalities implemented in AMReX. The SPADES software is released in NREL Software Record SWR-24-99 “SPADES (Scalable Parallel Discrete Events Simulation)”.

Henry de Frahan, Marc [National Renewable Energy L↗

Resiliency in numerical algorithm design for extreme scale simulations

Here this work is based on the seminar titled ‘Resiliency in Numerical Algorithm Design for Extreme Scale Simulations’ held March 1–6, 2020, at Schloss Dagstuhl, that was attended by all the authors. Advanced supercomputing is characterized by very high computation speeds at the cost of involving an enormous amount of resources and costs. A typical large-scale computation running for 48 h on a system consuming 20 MW, as predicted for exascale systems, would consume a million kWh, corresponding to about 100k Euro in energy cost for executing 10 23 floating-point operations. It is clearly unacceptable to lose the whole computation if any of the several million parallel processes fails during the execution. Moreover, if a single operation suffers from a bit-flip error, should the whole computation be declared invalid? What about the notion of reproducibility itself: should this core paradigm of science be revised and refined for results that are obtained by large-scale simulation? Naive versions of conventional resilience techniques will not scale to the exascale regime: with a main memory footprint of tens of Petabytes, synchronously writing checkpoint data all the way to background storage at frequent intervals will create intolerable overheads in runtime and energy consumption. Forecasts show that the mean time between failures could be lower than the time to recover from such a checkpoint, so that large calculations at scale might not make any progress if robust alternatives are not investigated. More advanced resilience techniques must be devised. The key may lie in exploiting both advanced system features as well as specific application knowledge. Research will face two essential questions: (1) what are the reliability requirements for a particular computation and (2) how do we best design the algorithms and software to meet these requirements? While the analysis of use cases can help understand the particular reliability requirements, the construction of remedies is currently wide open. One avenue would be to refine and improve on system- or application-level checkpointing and rollback strategies in the case an error is detected. Developers might use fault notification interfaces and flexible runtime systems to respond to node failures in an application-dependent fashion. Novel numerical algorithms or more stochastic computational approaches may be required to meet accuracy requirements in the face of undetectable soft errors. These ideas constituted an essential topic of the seminar. The goal of this Dagstuhl Seminar was to bring together a diverse group of scientists with expertise in exascale computing to discuss novel ways to make applications resilient against detected and undetected faults. In particular, participants explored the role that algorithms and applications play in the holistic approach needed to tackle this challenge. This article gathers a broad range of perspectives on the role of algorithms, applications and systems in achieving resilience for extreme scale simulations. The ultimate goal is to spark novel ideas and encourage the development of concrete solutions for achieving such resilience holistically.

79 ASTRONOMY AND ASTROPHYSICS↗

The STAR /self-testing and repairing/ computer - An investigation of the theory and practice of fault-tolerant computer design.

This paper presents the results obtained in a continuing investigation of fault-tolerant computing which is being conducted at the Jet Propulsion Laboratory. Initial studies led to the decision to design and construct an experimental computer with dynamic (standby) redundancy, including replaceable subsystems and a program rollback provision to eliminate transient errors. This system, called the STAR computer, began operation in 1969. The following aspects of the STAR system are described: architecture, reliability analysis, software, automatic maintenance of peripheral systems, and adaptation to serve as the central computer of an outer-planet exploration spacecraft.

A Avizienis↗