Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Checkpoint”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Logic design for dynamic and interactive recovery.

Recovery in a fault-tolerant computer means the continuation of system operation with data integrity after an error occurs. This paper delineates two parallel concepts embodied in the hardware and software functions required for recovery; detection, diagnosis, and reconfiguration for hardware, data integrity, checkpointing, and restart for the software. The hardware relies on the recovery variable set, checking circuits, and diagnostics, and the software relies on the recovery information set, audit, and reconstruct routines, to characterize the system state and assist in recovery when required. Of particular utility is a handware unit, the recovery control unit, which serves as an interface between error detection and software recovery programs in the supervisor and provides dynamic interactive recovery.

Carter, W. C.↗

A study of the utilization of ERTS-1 data from the Wabash River Basin

The author has identified the following significant results. The crop identification effort produced results indicating difficulty in discriminating cotton and soybeans in the Missouri area in the late August, September-October period. The use of prior probabilities in classification was shown to increase performance significantly. The use of temporal data was observed to improve classification accuracy for September 14 and October 2 data but not for August 26 data. Geometric correction and temporal overlay processing of ERTS-1 data proved to be very valuable for field location and relationship of fields from one time to another. Precision geometric correction of ERTS-1 data was achieved using manually derived ground control checkpoints.

Landgrebe, D. A.↗

System safety as applied to Skylab

Procedural and organizational guidelines used in accordance with NASA safety policy for the Skylab missions are outlined. The basic areas examined in the safety program for Skylab were the crew interface, extra-vehicular activity (EVA), energy sources, spacecraft interface, and hardware complexity. Fire prevention was a primary goal, with firefighting as backup. Studies of the vectorcardiogram and sleep monitoring experiments exemplify special efforts to prevent fire and shock. The final fire control study included material review, fire detection capability, and fire extinguishing capability. Contractors had major responsibility for system safety. Failure mode and effects analysis (FMEA) and equipment criticality categories are outlined. Redundancy was provided on systems that were critical to crew survival (category I). The five key checkpoints in Skylab hardware development are explained. Skylab rescue capability was demonstrated by preparations to rescue the Skylab 3 crew after their spacecraft developed attitude control problems.

Kleinknecht, K. S.↗

Transport of lunar material to the sites of the colonies

An 'existence proof' is attempted for the feasibility of transport of lunar material to colonies in space. Masses of lunar material are accelerated to lunar escape by a tracked magnetically levitated mass driver; aim precision is to 1 km miss distance at L5 per mm/sec velocity error at the lunar surface. Mass driver design and linear synchronous motor drive design are discussed; laser-sensed checkpoints aid in velocity and directional precision. Moon-L5 trajectories are calculated. The design of the L5 construction station, or 'catcher vehicle,' is described; loads are received by chambers operating in a 'Venus flytrap' mode. Further research studies needed to round out the concept are listed explicitly.

Heppenheimer, T. A.↗

X10: A FORTRAN direct access data management system

The XIO system is a set of subroutines that provide generalized data management capability for FORTRAN programs using a direct access file. Arrays of integer, real, double precision, and character data may be stored, each logical group of data identified by a unique matrix number. A matrix may be organized and stored as batches to reduce core requirements. Batches may be accessed randomly or sequentially. The file may be checkpointed and retained, allowing for restarts with stored values. The XIO subroutines operate on either IBM 360-370/OS/VS or DEC PDP-11/RSX computing systems.

Roland, D. P.↗

Steady, oscillatory, and unsteady subsonic Aerodynamics, production version 1.1 (SOUSSA-P1.1). Volume 2: User/programmer manual

A user/programmer manual for the computer program SOUSSA P 1.1 is presented. The program was designed to provide accurate and efficient evaluation of steady and unsteady loads on aircraft having arbitrary shapes and motions, including structural deformations. These design goals were in part achieved through the incorporation of the data handling capabilities of the SPAR finite element Structural Analysis computer program. As a further result, SOUSSA P possesses an extensive checkpoint/ restart facility. The programmer's portion of this manual includes overlay/subroutine hierarchy, logical flow of control, definition of SOUSSA P 1.1 FORTRAN variables, and definition of SOUSSA P 1.1 subroutines. Purpose of the SOUSSA P 1.1 modules, input data to the program, output of the program, hardware/software requirements, error detection and reporting capabilities, job control statements, a summary of the procedure for running the program and two test cases including input and output and listings are described in the user oriented portion of the manual.

Smolka, S. A.↗

Enhancements to Sperry/NASTRAN

Reviewed is the enhancement to NASTRAN program performed by NUK (Nippon Univac Kaisha, Ltd.) added to Level 15.5. Features discussed include intermediate checkpoint-restart in triangular decomposition, I/O improvement, multibanked memory and new plate element. The first three improvements provide the capability to solve significantly large size problems, while the new elements release the analyst from the cumbersome work to constrain the singularities caused by the lack of stiffness of inplane rotation of old plate elements.

Koga, T.↗

X-29 flight - Acid test for design predictions

The X-29 flight test data are being disseminated to interested industrial and military users as fast as it becomes available. The aircraft is extensively instrumented with accelerometers and pressure sensors and optical sensors for measuring wing deflection. The thoroughness of preflight preparations permitted a rapid advance through initial test checkpoints, which have both confirmed many predictions and revealed several discrepancies. The flight envelope had been expanded to Mach 1.1 and an altitude of 40,000 ft by December 1985. Notably, the X-29 has provided in-flight data which could not be faithfully depicted in a simulator, e.g., flare procedures during landing, and has shown that the stability adjustments, although adequate for controlling the aircraft, are not rapid enough to offer a satisfactory margin of harmony. The tests are now being performed in the transonic regime, where supercritical airfoil and forward swept wing drag reduction become significant factors.

Putnam, T. W.↗

Ground-based time-guidance algorithm for control of airplanes in a time-metered air traffic control environment: A piloted simulation study

The rapidly increasing costs of flight operations and the requirement for increased fuel conservation have made it necessary to develop more efficient ways to operate airplanes and to control air traffic for arrivals and departures to the terminal area. One concept of controlling arrival traffic through time metering has been jointly studied and evaluated by NASA and ONERA/CERT in piloted simulation tests. From time errors attained at checkpoints, airspeed and heading commands issued by air traffic control were computed by a time-guidance algorithm for the pilot to follow that would cause the airplane to cross a metering fix at a preassigned time. These tests resulted in the simulated airplane crossing a metering fix with a mean time error of 1.0 sec and a standard deviation of 16.7 sec when the time-metering algorithm was used. With mismodeled winds representing the unknown in wind-aloft forecasts and modeling form, the mean time error attained when crossing the metering fix was increased and the standard deviation remained approximately the same. The subject pilots reported that the airspeed and heading commands computed in the guidance concept were easy to follow and did not increase their work load above normal levels.

Knox, C. E.↗

The cometary nucleus: Current concepts

Ideas concerning the nature of cometary nuclei as modified by observations of recent years, particularly by those of Halley's comet during the 1986 apparition are reviewed. The major trend is to increase the postulated dimensions of the nuclei and reduce their albedos. The Vega and Giotto missions establish an invaluable checkpoint with regard to comet nuclei. These and other observations are affecting opinions concerning many facets of cometary nature, origin, evolution, and decay.

Whipple, Fred L.↗

The Raid distributed database system

Raid, a robust and adaptable distributed database system for transaction processing (TP), is described. Raid is a message-passing system, with server processes on each site to manage concurrent processing, consistent replicated copies during site failures, and atomic distributed commitment. A high-level layered communications package provides a clean location-independent interface between servers. The latest design of the package delivers messages via shared memory in a configuration with several servers linked into a single process. Raid provides the infrastructure to investigate various methods for supporting reliable distributed TP. Measurements on TP and server CPU time are presented, along with data from experiments on communications software, consistent replicated copy control during site failures, and concurrent distributed checkpointing. A software tool for evaluating the implementation of TP algorithms in an operating-system kernel is proposed.

Bhargava, Bharat↗

Investigation of the applicability of a functional programming model to fault-tolerant parallel processing for knowledge-based systems

In a fault-tolerant parallel computer, a functional programming model can facilitate distributed checkpointing, error recovery, load balancing, and graceful degradation. Such a model has been implemented on the Draper Fault-Tolerant Parallel Processor (FTPP). When used in conjunction with the FTPP's fault detection and masking capabilities, this implementation results in a graceful degradation of system performance after faults. Three graceful degradation algorithms have been implemented and are presented. A user interface has been implemented which requires minimal cognitive overhead by the application programmer, masking such complexities as the system's redundancy, distributed nature, variable complement of processing resources, load balancing, fault occurrence and recovery. This user interface is described and its use demonstrated. The applicability of the functional programming style to the Activation Framework, a paradigm for intelligent systems, is then briefly described.

Harper, Richard↗

Analyzing the I/O behavior of supercomputer applications

Supercomputer applications of I/Os are classified here into three categories - required, checkpoint, and data staging - and it is shown how memory size and CPU speed are likely to affect each category. A data analysis shows that data staging I/O dominates when it is present. If an application requires data staging I/O, its read/write ratio is greater than one; otherwise, its read/write ratio is less than one. This information can be used to anticipate a program's I/O demands and thus reduce the load on the I/O system.

Miller, Ethan L.↗

Application of statistical process control and process capability analysis procedures in orbiter processing activities at the Kennedy Space Center

Successful ground processing at KSC requires that flight hardware and ground support equipment conform to specifications at tens of thousands of checkpoints. Knowledge of conformance is an essential requirement for launch. That knowledge of conformance at every requisite point does not, however, enable identification of past problems with equipment, or potential problem areas. This paper describes how the introduction of Statistical Process Control and Process Capability Analysis identification procedures into existing shuttle processing procedures can enable identification of potential problem areas and candidates for improvements to increase processing performance measures. Results of a case study describing application of the analysis procedures to Thermal Protection System processing are used to illustrate the benefits of the approaches described in the paper.

Safford, Robert R.↗

Progressive retry for software error recovery in distributed systems

In this paper, we describe a method of execution retry for bypassing software errors based on checkpointing, rollback, message reordering and replaying. We demonstrate how rollback techniques, previously developed for transient hardware failure recovery, can also be used to recover from software faults by exploiting message reordering to bypass software errors. Our approach intentionally increases the degree of nondeterminism and the scope of rollback when a previous retry fails. Examples from our experience with telecommunications software systems illustrate the benefits of the scheme.

Wang, Yi-Min↗

Fault-tolerant, embedded CLIPS applications

The enhancements to CLIPS4.3 presented in this paper provide an embedded CLIPS application with the ability to continue operation with minimal to no loss of information in the event of a hardware or a software failure. Given an arbitrary failure, the CLIPS application's environment (fact-list, agenda, and pattern matching network) will be reconstructed to the point at which the failure was experienced. The environment reconstruction is based on state files to which the application periodically checks environment information (fact-list and agenda). The routine for checkpointing the state of the application is as efficient as possible so that the overhead introduced to normal execution of the application is minimal. The only assumptions made by the CLIPS application are that it is running under an operating system that guarantees it access to uncorrupt state files and that the application will be automatically restarted should it terminate abnormally.

Hicks, Jaye↗

Software dependability in the Tandem GUARDIAN system

Based on extensive field failure data for Tandem's GUARDIAN operating system this paper discusses evaluation of the dependability of operational software. Software faults considered are major defects that result in processor failures and invoke backup processes to take over. The paper categorizes the underlying causes of software failures and evaluates the effectiveness of the process pair technique in tolerating software faults. A model to describe the impact of software faults on the reliability of an overall system is proposed. The model is used to evaluate the significance of key factors that determine software dependability and to identify areas for improvement. An analysis of the data shows that about 77% of processor failures that are initially considered due to software are confirmed as software problems. The analysis shows that the use of process pairs to provide checkpointing and restart (originally intended for tolerating hardware faults) allows the system to tolerate about 75% of reported software faults that result in processor failures. The loose coupling between processors, which results in the backup execution (the processor state and the sequence of events) being different from the original execution, is a major reason for the measured software fault tolerance. Over two-thirds (72%) of measured software failures are recurrences of previously reported faults. Modeling, based on the data, shows that, in addition to reducing the number of software faults, software dependability can be enhanced by reducing the recurrence rate.

Lee, Inhwan↗

Lidar measurements at Lauder, NZ

In March of 1994, the GSFC Stratospheric Ozone Lidar was deployed to the Network for the Detection of Stratospheric Change (NDSC) site at Lauder, NZ. This was in conjunction with a series of NASA ER-2 flights from Christchurch, NZ south to the Antarctic Circle. These flights were organized to study the chemistry of the stratosphere before, during and after the formation of the well-known 'ozone hole'. Lidar measurements were made at four different time periods corresponding to the times of the ER-2 flights. Lauder is situated nearly along the flight path as the aircraft flew south and so the lidar measurements provide a checkpoint for the ozone, aerosol and temperature instruments onboard the aircraft. Whenever the weather permitted, lidar measurements were made as near to dawn, prior to the flight, and as near to sunset, after the flight. This provided data as close to the aircraft transit time as possible. More than 70 individual lidar measurements were made, each consisting of a vertical profile of ozone, temperature, and aerosol. These were made over three different seasons and show seasonal variation. Of particular interest in the lidar data base is the wintertime stratospheric - mesospheric temperature profiles, which show large variations at the stratopause and also some significant wave activity.

McGee, Thomas J.↗