Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “rollback”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Relaxing consistency in recoverable distributed shared memory

Relaxed memory consistency models have recently been proposed to tolerate memory access latency in both hardware and software distributed shared memory systems. In recoverable shared memory multiprocessors, relaxing consistency has the added benefit of reducing the number of checkpoints needed to avoid rollback propagation. In this paper, we introduce new checkpointing algorithms that take advantage of relaxed consistency to reduce the performance overhead of checkpointing. We also introduce a scheme based on lazy relaxed consistency, that reduces both checkpointing overhead and the overhead of avoiding error propagation in systems with error latency. Multiprocessor address traces are used to evaluate the relaxed consistency approach to checkpointing with distributed shared memory.

Janssens, Bob↗

Flexure and the role of inplane force around coronae on Venus

Large coronae on Venus, such as Artemis and Latona, are rimmed by conspicuous trenches and associated outer rises. Sandwell and Schubert have observed that these systems resemble terrestrial subduction zones in planform and have succeeded in fitting an elastic plate bending equation to the inferred flexural topography. However, the first zero crossing bending moments required are -2.5 x 10(exp 17) N for Artemis and -5.0 x 10(exp 16) N for Latona. Since these moments are similar in magnitude to those of subducting slabs on Earth, a rollback subduction mechanism was proposed to explain the flexure around the largest coronae, although a differential thermal subsidence model is sufficient to account for the topography around some coronae. The purpose is to investigate the effect of inplane force as a possible alternative to large applied moments in producing flexure at Artemis and Latona. The close correlation of gravity to topography on Venus implies the absence of a low viscosity zone and the strong coupling of the lithosphere to mantle convection. If coronae are the surface manifestations of mantle plumes, they may be the sites of active convective stress coupling. As the upwelling reaches the lithosphere, it spreads radially outward, inducing shear tractions on the base of the plate. In addition, the hot, expanding corona may load the surrounding plate horizonally. Both the basal shear stresses and radial loading can be treated as an equivalent compressive inplane force in the mechanical lithosphere, which contributes to the bending of the outlying plate. Using a model that relates inplane force to the measured gravity anomalies, a rough value of the inplane force at Artemis was calculated. Recent Pioneer Venus spherical harmonic gravity models indicate a geoid anomaly of about 75 m over Artemis, which corresponds to an estimated inplane force on the order of -1x10(exp 13) N/m. The gravity model is unable to resolve Latona, but an inplane force of similar dimensions is assumed. The maximum possible inplane force based on the expected rheology can be constrained by using the approximate 5 K/km thermal gradient inferred from the best fit 30 km elastic plate at Artemis and Latona. For a dry olivine flow law in the upper mantle, the compressional load limit of the 60 km thick mechanical lithosphere is -4 x 10(exp 13) N/m. This value is equivalent to a load of -8 x 10(exp 13) N/m on a 30 km thick elastic plate.

Brown, C. David↗

Efficient massively parallel simulation of dynamic channel assignment schemes for wireless cellular communications

Fast, efficient parallel algorithms are presented for discrete event simulations of dynamic channel assignment schemes for wireless cellular communication networks. The driving events are call arrivals and departures, in continuous time, to cells geographically distributed across the service area. A dynamic channel assignment scheme decides which call arrivals to accept, and which channels to allocate to the accepted calls, attempting to minimize call blocking while ensuring co-channel interference is tolerably low. Specifically, the scheme ensures that the same channel is used concurrently at different cells only if the pairwise distances between those cells are sufficiently large. Much of the complexity of the system comes from ensuring this separation. The network is modeled as a system of interacting continuous time automata, each corresponding to a cell. To simulate the model, conservative methods are used; i.e., methods in which no errors occur in the course of the simulation and so no rollback or relaxation is needed. Implemented on a 16K processor MasPar MP-1, an elegant and simple technique provides speedups of about 15 times over an optimized serial simulation running on a high speed workstation. A drawback of this technique, typical of conservative methods, is that processor utilization is rather low. To overcome this, new methods were developed that exploit slackness in event dependencies over short intervals of time, thereby raising the utilization to above 50 percent and the speedup over the optimized serial code to about 120 times.

Greenberg, Albert G.↗

Relaxing consistency in recoverable distributed shared memory

Relaxed memory consistency models tolerate increased memory access latency in both hardware and software distributed shared memory systems. In recoverable systems, relaxing consistency has the added benefit of reducing the number of checkpoints needed to avoid rollback propagation. In this paper, we introduce new checkpointing algorithms that take advantage of relaxed consistency to reduce the performance overhead of checkpointing. We also introduce a scheme based on lazy relaxed consistency, that reduces both checkpointing overhead and the overhead of avoiding error propagation in systems with error latency. We use multiprocessor address traces to evaluate the relaxed consistency approach to checkpointing with distributed shared memory.

Janssens, Bob↗

A Scheduling Algorithm for Replicated Real-Time Tasks

We present an algorithm for scheduling real-time periodic tasks on a multiprocessor system under fault-tolerant requirement. Our approach incorporates both the redundancy and masking technique and the imprecise computation model. Since the tasks in hard real-time systems have stringent timing constraints, the redundancy and masking technique are more appropriate than the rollback techniques which usually require extra time for error recovery. The imprecise computation model provides flexible functionality by trading off the quality of the result produced by a task with the amount of processing time required to produce it. It therefore permits the performance of a real-time system to degrade gracefully. We evaluate the algorithm by stochastic analysis and Monte Carlo simulations. The results show that the algorithm is resilient under hardware failures.

Yu, Albert C.↗

Reducing Interprocessor Dependence in Recoverable Distributed Shared Memory

Checkpointing techniques in parallel systems use dependency tracking and/or message logging to ensure that a system rolls back to a consistent state. Traditional dependency tracking in distributed shared memory (DSM) systems is expensive because of high communication frequency. In this paper we show that, if designed correctly, a DSM system only needs to consider dependencies due to the transfer of blocks of data, resulting in reduced dependency tracking overhead and reduced potential for rollback propagation. We develop an ownership timestamp scheme to tolerate the loss of block state information and develop a passive server model of execution where interactions between processors are considered atomic. With our scheme, dependencies are significantly reduced compared to the traditional message-passing model.

Janssens, Bob↗

Efficient Tracing for On-the-Fly Space-Time Displays in a Debugger for Message Passing Programs

In this work we describe the implementation of a practical mechanism for collecting and displaying trace information in a debugger for message passing programs. We introduce a trace format that is highly compressible while still providing information adequate for debugging purposes. We make the mechanism convenient for users to access by incorporating the trace collection in a set of wrappers for the MPI (message passing interface) communication library. We implement several debugger operations that use the trace display: consistent stoplines, undo, and rollback. They all are implemented using controlled replay, which executes at full speed in target processes until the appropriate position in the computation is reached. They provide convenient mechanisms for getting to places in the execution where the full power of a state-based debugger can be brought to bear on isolating communication errors.

Hood, Robert↗

A Domain Decomposition Parallelization of the Fast Marching Method

In this paper, the first domain decomposition parallelization of the Fast Marching Method for level sets has been presented. Parallel speedup has been demonstrated in both the optimal and non-optimal domain decomposition case. The parallel performance of the proposed method is strongly dependent on load balancing separately the number of nodes on each side of the interface. A load imbalance of nodes on either side of the domain leads to an increase in communication and rollback operations. Furthermore, the amount of inter-domain communication can be reduced by aligning the inter-domain boundaries with the interface normal vectors. In the case of optimal load balancing and aligned inter-domain boundaries, the proposed parallel FMM algorithm is highly efficient, reaching efficiency factors of up to 0.98. Future work will focus on the extension of the proposed parallel algorithm to higher order accuracy. Also, to further enhance parallel performance, the coupling of the domain decomposition parallelization to the G(sub 0)-based parallelization will be investigated.

Herrmann, M.↗

STS-114: Discovery Post MMT Press Conference

This press conference focuses on the outcome of the Mission Management Team (MMT) meeting. The launch and status of the Space Shuttle Discovery is discussed. George Diller from NASA Public Affairs introduces the panel which consists of: Wayne Hale, Space Shuttle Program Deputy Manager and Mike Wetmore, Director of Space Shuttle Processing at Nasa Kennedy Space Center. The news media asks questions about the history of the low level sensors in the hydrogen tank, the cryogenic atmosphere around the sensors, troubleshooting, astronaut activities, possible rollback procedures.

Source record↗

Launch Pad 39 Hail Monitor Array System

Weather conditions at Kennedy Space Center are extremely dynamic, and they greatly affect the safety of the Space Shuttles sitting on the launch pads. For example, on May 13, 1999, the foam on the External Tank (ET) of STS-96 was significantly damaged by hail at the launch pad, requiring rollback to the Vehicle Assembly Building. The loss of ET foam on STS-114 in 2005 intensified interest in monitoring and measuring damage to ET foam, especially from hail. But hail can be difficult to detect and monitor because it is often localized and obscured by heavy rain. Furthermore, the hot Florida climate usually melts the hail even before the rainfall subsides. In response, the hail monitor array (HMA) system, a joint effort of the Applied Physics Laboratory operated by NASA and ASRC Aerospace at KSC, was deployed for operational testing in the fall of 2006. Volunteers from the Community Collaborative Rain, Hail, and Snow (CoCoRaHS) network, in conjunction with Colorado State University, continue to test duplicate hail monitor systems deployed in the high plains of Colorado.

Source record↗

Software for Fault-Tolerant Matrix Multiplication

Formal Linear Algebra Recovery Environment is a computer program for high-performance, fault-tolerant matrix multiplication. The program is based on an extension of the prior theory and practice of fault-tolerant matrix matrix multiplication of the form C = AB. This extension provides low-overhead methods for detecting errors, not only in C, but also in A and/or B. These methods enable the detection of all errors as long as, in a given case, only one entry in A, B, or C is corrupted. The program also provides for following a low-overhead rollback approach to correct errors once detected. Results of computational experiments have demonstrated that the methods implemented in this program work well in practice while imposing an acceptably low level of overhead, relative to high-performance matrix-multiplication methods that do not afford fault tolerance.

Katz, Daniel↗

Fault Tolerance Middleware for a Multi-Core System

Fault Tolerance Middleware (FTM) provides a framework to run on a dedicated core of a multi-core system and handles detection of single-event upsets (SEUs), and the responses to those SEUs, occurring in an application running on multiple cores of the processor. This software was written expressly for a multi-core system and can support different kinds of fault strategies, such as introspection, algorithm-based fault tolerance (ABFT), and triple modular redundancy (TMR). It focuses on providing fault tolerance for the application code, and represents the first step in a plan to eventually include fault tolerance in message passing and the FTM itself. In the multi-core system, the FTM resides on a single, dedicated core, separate from the cores used by the application. This is done in order to isolate the FTM from application faults and to allow it to swap out any application core for a substitute. The structure of the FTM consists of an interface to a fault tolerant strategy module, a responder module, a fault manager module, an error factory, and an error mapper that determines the severity of the error. In the present reference implementation, the only fault tolerant strategy implemented is introspection. The introspection code waits for an application node to send an error notification to it. It then uses the error factory to create an error object, and at this time, a severity level is assigned to the error. The introspection code uses its built-in knowledge base to generate a recommended response to the error. Responses might include ignoring the error, logging it, rolling back the application to a previously saved checkpoint, swapping in a new node to replace a bad one, or restarting the application. The original error and recommended response are passed to the top-level fault manager module, which invokes the response. The responder module also notifies the introspection module of the generated response. This provides additional information to the introspection module that it can use in generating its next response. For example, if the responder triggers an application rollback and errors are still occurring, the introspection module may decide to recommend an application restart.

Some, Raphael R.↗

A Systems-Level Perspective on Engine Ice Accretion

Talk covers: (1) Problem of Engine Power Loss;(2) Modeling Engine Icing Effects; (3) Simulation of Engine Rollback; (4) Icing/Engine Control System Interaction; (5) Detection of Ice Accretion; (6) Potential Mitigation Strategies.

May, Ryan David↗

Hail Disrometer Array for Launch Systems Support

Prior to launch, the space shuttle might be described as a very large thermos bottle containing substantial quantities of cryogenic fuels. Because thermal insulation is a critical design requirement, the external wall of the launch vehicle fuel tank is covered with an insulating foam layer. This foam is fragile and can be damaged by very minor impacts, such as that from small- to medium-size hail, which may go unnoticed. In May 1999, hail damage to the top of the External Tank (ET) of STS-96 required a rollback from the launch pad to the Vehicle Assembly Building (VAB) for repair of the insulating foam. Because of the potential for hail damage to the ET while exposed to the weather, a vigilant hail sentry system using impact transducers was developed as a hail damage warning system and to record and quantify hail events. The Kennedy Space Center (KSC) Hail Monitor System, a joint effort of the NASA and University Affiliated Spaceport Technology Development Contract (USTDC) Physics Labs, was first deployed for operational testing in the fall of 2006. Volunteers from the Community Collaborative Rain. Hail, and Snow Network (CoCoRaHS) in conjunction with Colorado State University were and continue to be active in testing duplicate hail monitor systems at sites in the hail prone high plains of Colorado. The KSC Hail Monitor System (HMS), consisting of three stations positioned approximately 500 ft from the launch pad and forming an approximate equilateral triangle (see Figure 1), was deployed to Pad 39B for support of STS-115. Two months later, the HMS was deployed to Pad 39A for support of STS-116. During support of STS-117 in late February 2007, an unusual hail event occurred in the immediate vicinity of the exposed space shuttle and launch pad. Hail data of this event was collected by the HMS and analyzed. Support of STS-118 revealed another important application of the hail monitor system. Ground Instrumentation personnel check the hail monitors daily when a vehicle is on the launch pad, with special attention after any storm suspected of containing hail. If no hail is recorded by the HMS, the vehicle and pad inspection team has no need to conduct a thorough inspection of the vehicle immediately following a storm. On the afternoon of July 13, 2007, hail on the ground was reported by observers at the VAB, about three miles west of Pad 39A, as well as at several other locations around Kennedy Space Center. The HMS showed no impact detections, indicating that the shuttle had not been damaged by any of the numerous hail events which occurred that day.

Lane, John E.↗

Preliminary Results From a Heavily Instrumented Engine Ice Crystal Icing Test in a Ground Based Altitude Test Facility

Preliminary results from the Heavily Instrumented ALF503R-5 Engine test conducted in the NASA Glenn Research Center Propulsion Systems Laboratory will be discussed. The effects of ice crystal icing on a full scale engine is examined and documented. This model engine, serial number LF01, was used during the inaugural icing test in the PSL facility. The reduction of thrust (rollback) events experienced by this engine in flight were replicated in the facility. Limited instrumentation was used to detect icing. Metal temperature on the exit guide vanes and outer shroud and the load measurement were the only indicators of ice formation. The current study features a similar engine, serial number LF11, which is instrumented to characterize the cloud entering the engine, detect characterize ice accretion, and visualize the ice accretion in the region of interest.

enigine icing↗

A Dynamic Model for the Evaluation of Aircraft Engine Icing Detection and Control-Based Mitigation Strategies

Aircraft flying in regions of high ice crystal concentrations are susceptible to the buildup of ice within the compression system of their gas turbine engines. This ice buildup can restrict engine airflow and cause an uncommanded loss of thrust, also known as engine rollback, which poses a potential safety hazard. The aviation community is conducting research to understand this phenomena, and to identify avoidance and mitigation strategies to address the concern. To support this research, a dynamic turbofan engine model has been created to enable the development and evaluation of engine icing detection and control-based mitigation strategies. This model captures the dynamic engine response due to high ice water ingestion and the buildup of ice blockage in the engines low pressure compressor. It includes a fuel control system allowing engine closed-loop control effects during engine icing events to be emulated. The model also includes bleed air valve and horsepower extraction actuators that, when modulated, change overall engine operating performance. This system-level model has been developed and compared against test data acquired from an aircraft turbofan engine undergoing engine icing studies in an altitude test facility and also against outputs from the manufacturers customer deck. This paper will describe the model and show results of its dynamic response under open-loop and closed-loop control operating scenarios in the presence of ice blockage buildup compared against engine test cell data. Planned follow-on use of the model for the development and evaluation of icing detection and control-based mitigation strategies will also be discussed. The intent is to combine the model and control mitigation logic with an engine icing risk calculation tool capable of predicting the risk of engine icing based on current operating conditions. Upon detection of an operating region of risk for engine icing events, the control mitigation logic will seek to change the engines operating point to a region of lower risk through the modulation of available control actuators while maintaining the desired engine thrust output. Follow-on work will assess the feasibility and effectiveness of such control-based mitigation strategies.

engine control↗

A Dynamic Model for the Evaluation of Aircraft Engine Icing Detection and Control-Based Mitigation Strategies

Aircraft flying in regions of high ice crystal concentrations are susceptible to the buildup of ice within the compression system of their gas turbine engines. This ice buildup can restrict engine airflow and cause an uncommanded loss of thrust, also known as engine rollback, which poses a potential safety hazard. The aviation community is conducting research to understand this phenomena, and to identify avoidance and mitigation strategies to address the concern. To support this research, a dynamic turbofan engine model has been created to enable the development and evaluation of engine icing detection and control-based mitigation strategies. This model captures the dynamic engine response due to high ice water ingestion and the buildup of ice blockage in the engines low pressure compressor. It includes a fuel control system allowing engine closed-loop control effects during engine icing events to be emulated. The model also includes bleed air valve and horsepower extraction actuators that, when modulated, change overall engine operating performance. This system-level model has been developed and compared against test data acquired from an aircraft turbofan engine undergoing engine icing studies in an altitude test facility and also against outputs from the manufacturers customer deck. This paper will describe the model and show results of its dynamic response under open-loop and closed-loop control operating scenarios in the presence of ice blockage buildup compared against engine test cell data. Planned follow-on use of the model for the development and evaluation of icing detection and control-based mitigation strategies will also be discussed. The intent is to combine the model and control mitigation logic with an engine icing risk calculation tool capable of predicting the risk of engine icing based on current operating conditions. Upon detection of an operating region of risk for engine icing events, the control mitigation logic will seek to change the engines operating point to a region of lower risk through the modulation of available control actuators while maintaining the desired engine thrust output. Follow-on work will assess the feasibility and effectiveness of such control-based mitigation strategies.

engine control↗

Total Temperature Measurements Using a Rearward Facing Probe in Supercooled Liquid Droplet and Ice Crystal Clouds

Engine Icing Performance loss: rollback, surge, flameout, and even internal engine damagePartial melting and refreeze of ice inside engine core (Mason et al., 2006). Ingestion of ice crystals and aggregates, mixed-phase droplets, or supercooled liquid dropletsNeed to better understand the conditions and properties that lead to engine icing.Simulation and analysis (physical and computational, and modeling)Test facilities (PSL, NRC, ...). Thermal and computational models and analysisProbesMultiple probes (aerothermal probes and ice cloud characterization probes and techniques). Total temperatureTraditional total temperature probes (vented forward facing)Heated total temperature probes (Goodyear). Rearward facing (developmental). Total temperature relevance. Thermal interaction between the icing cloud and air flow impinging particles contribute to kinetic heating effect (Gent et al., 2000). Measurement considerations Temperature sensor accuracy. Incomplete recovery of total temperature. Thermal surfaces (sources and sinks). Flow effects (viscous losses). Debris contamination, including icing and ice ingestion.

Agui, Juan H.↗