Engineering PapersSearch

SEARCH · Engineering Papers

Results for “communication failure”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Abnormal/Emergency Situations. Impact of Unmanned Aircraft Systems Emergency and Abnormal Events on the National Airspace System

Access 5 analyzed the differences between UAS and manned aircraft operations under five categories of abnormal or emergency situations: Link Failure, Lost Communications, Onboard System Failures, Control Station Failures and Abnormal Weather. These analyses were made from the vantage point of the impact that these operations have on the US air traffic control system, with recommendations for new policies and procedures included where appropriate.

Source record

Work Practice Simulation of Complex Human-Automation Systems in Safety Critical Situations: The Brahms Generalized berlingen Model

The transition from the current air traffic system to the next generation air traffic system will require the introduction of new automated systems, including transferring some functions from air traffic controllers to on­-board automation. This report describes a new design verification and validation (V&V) methodology for assessing aviation safety. The approach involves a detailed computer simulation of work practices that includes people interacting with flight-critical systems. The research is part of an effort to develop new modeling and verification methodologies that can assess the safety of flight-critical systems, system configurations, and operational concepts. The 2002 Ueberlingen mid-air collision was chosen for analysis and modeling because one of the main causes of the accident was one crew's response to a conflict between the instructions of the air traffic controller and the instructions of TCAS, an automated Traffic Alert and Collision Avoidance System on-board warning system. It thus furnishes an example of the problem of authority versus autonomy. It provides a starting point for exploring authority/autonomy conflict in the larger system of organization, tools, and practices in which the participants' moment-by-moment actions take place. We have developed a general air traffic system model (not a specific simulation of Überlingen events), called the Brahms Generalized Ueberlingen Model (Brahms-GUeM). Brahms is a multi-agent simulation system that models people, tools, facilities/vehicles, and geography to simulate the current air transportation system as a collection of distributed, interactive subsystems (e.g., airports, air-traffic control towers and personnel, aircraft, automated flight systems and air-traffic tools, instruments, crew). Brahms-GUeM can be configured in different ways, called scenarios, such that anomalous events that contributed to the Überlingen accident can be modeled as functioning according to requirements or in an anomalous condition, as occurred during the accident. Brahms-GUeM thus implicitly defines a class of scenarios, which include as an instance what occurred at Überlingen. Brahms-GUeM is a modeling framework enabling "what if" analysis of alternative work system configurations and thus facilitating design of alternative operations concepts. It enables subsequent adaption (reusing simulation components) for modeling and simulating NextGen scenarios. This project demonstrates that BRAHMS provides the capacity to model the complexity of air transportation systems, going beyond idealized and simple flights to include for example the interaction of pilots and ATCOs. The research shows clearly that verification and validation must include the entire work system, on the one hand to check that mechanisms exist to handle failures of communication and alerting subsystems and/or failures of people to notice, comprehend, or communicate problematic (unsafe) situations; but also to understand how people must use their own judgment in relating fallible systems like TCAS to other sources of information and thus to evaluate how the unreliability of automation affects system safety. The simulation shows in particular that distributed agents (people and automated systems) acting without knowledge of each others' actions can create a complex, dynamic system whose interactive behavior is unexpected and is changing too quickly to comprehend and control.

complex systems

Reliable communication in the presence of failures

The design and correctness of a communication facility for a distributed computer system are reported on. The facility provides support for fault-tolerant process groups in the form of a family of reliable multicast protocols that can be used in both local- and wide-area networks. These protocols attain high levels of concurrency, while respecting application-specific delivery ordering constraints, and have varying cost and performance that depend on the degree of ordering desired. In particular, a protocol that enforces causal delivery orderings is introduced and shown to be a valuable alternative to conventional asynchronous communication protocols. The facility also ensures that the processes belonging to a fault-tolerant process group will observe consistant orderings of events affecting the group as a whole, including process failures, recoveries, migration, and dynamic changes to group properties like member rankings. A review of several uses for the protocols is the ISIS system, which supports fault-tolerant resilient objects and bulletin boards, illustrates the significant simplification of higher level algorithms made possible by our approach.

Birman, Kenneth P.

Fallible humans and vulnerable systems - Lessons learned from aviation

It is suggested that the problems being experienced in complex automatic systems are essentially due to the failure of information management and communication. The failure covers the entire spectrum: display devices and techniques, coding information so as to reduce human error, and information economy, i.e., resisting the temptation to bombard the operator with unlimited information simply because the system possesses the capability to do so. Since there has been great progress in hardware engineering, it is suggested that further attention is needed in the 'soft' side of systems. The approach should focus on (1) preventing human cognitive slips and (2) making the systems less vulnerable to such slips when they do occur. Most of the examples are taken from studies of cockpit automation.

Wiener, Earl L.

Virtually-synchronous communication based on a weak failure suspector

Failure detectors (or, more accurately Failure Suspectors (FS)) appear to be a fundamental service upon which to build fault-tolerant, distributed applications. This paper shows that a FS with very weak semantics (i.e., that delivers failure and recovery information in no specific order) suffices to implement virtually-synchronous communication (VSC) in an asynchronous system subject to process crash failures and network partitions. The VSC paradigm is particularly useful in asynchronous systems and greatly simplifies building fault-tolerant applications that mask failures by replicating processes. We suggest a three-component architecture to implement virtually-synchronous communication: (1) at the lowest level, the FS component; (2) on top of it, a component (2a) that defines new views; and (3) a component (2b) that reliably multicasts messages within a view. The issues covered in this paper also lead to a better understanding of the various membership service semantics proposed in recent literature.

Schiper, Andre

Communication on the Flight Deck

The importance of good verbal communication by airplane crews is discussed. The common understanding of what it takes to fly a plane serves the communication and coordination. It is shown that accidents and incidents often are a result of failure to communicate. People's conceptions of what they are doing effect performance of a task and communication about performance of that task. It is known that if what they are doing is complex, they must find a simple framework to represent it to avoid being overwhelmed by complexity. A model that reduces the complexity of flying a jet to a representation composed of a few relatively independent dimensions which capture the major features of this task was developed. The model allows assessment of what crew members know, to look at actual performance, and to develop training recommendations.

Siesfeld, A.

An Autonomous Distributed Fault-Tolerant Local Positioning System

We describe a fault-tolerant, GPS-independent (Global Positioning System) distributed autonomous positioning system for static/mobile objects and present solutions for providing highly-accurate geo-location data for the static/mobile objects in dynamic environments. The reliability and accuracy of a positioning system fundamentally depends on two factors; its timeliness in broadcasting signals and the knowledge of its geometry, i.e., locations and distances of the beacons. Existing distributed positioning systems either synchronize to a common external source like GPS or establish their own time synchrony using a scheme similar to a master-slave by designating a particular beacon as the master and other beacons synchronize to it, resulting in a single point of failure. Another drawback of existing positioning systems is their lack of addressing various fault manifestations, in particular, communication link failures, which, as in wireless networks, are increasingly dominating the process failures and are typically transient and mobile, in the sense that they typically affect different messages to/from different processes over time.

Mahyar R Malekpour

Auxiliary engine digital interface unit (DIU)

This auxiliary propulsion engine digital unit controls both the valving of the fuel and oxidizer to the engine combustion chamber and the ignition spark required for timely and efficient engine burns. In addition to this basic function, the unit is designed to manage it's own redundancy such that it is still operational after two hard circuit failures. It communicates to the data bus system several selected information points relating to the operational status of the electronics as well as the engine fuel and burning processes.

Source record

Report of the Presidential Commission on the Space Shuttle Challenger Accident, Volume 1

The findings of the Commission regarding the circumstances surrounding the Challenger accident are reported and recommendations for corrective action are outlined. All available mission data, subsequent tests, and wreckage analyses were reviewed and specific failure scenarios were developed. The Commission concluded that the cause of the Mission 51-L accident was the failure of the pressure seal in the aft field joint of the right solid rocket motor. The failure was due to a faulty design unacceptably sensitive to a number of factors. These factors were the effects of temperature, physical dimensions, the character of materials, the effects of reuse, processing, and the reaction of the joint to dynamic loading. In addition to analyzing the material causes of the accident, the Commission examined the chain of decisions that culminated in approval of the launch. It concluded that the decision making process was flawed in several ways including (1) failure in communication resulting in a launch decision based on incomplete and misleading information, (2) a conflict between engineering data and management judgements, and (3) a NASA management structure that permitted flight safety problems to bypass key Shuttle managers.

Rogers, W. P.

A Multiple Model Based Approach for Deep Space Power System Fault Diagnosis

Improving protection and health management capabilities onboard the electrical power system (EPS) for spacecraft is essential for ensuring safe and reliable conditions for deep space human exploration. Electrical protection and control technologies on the National Aeronautics and Space Administration's (NASA's) current human space platform relies heavily on ground support to monitor and diagnose power systems and failures. As communication bandwidth diminishes for deep space applications, a transformation in system monitoring and control becomes necessary to maintain high reliability of electric power service. This paper presents a novel approach for on-line power system security monitoring for autonomous deep space spacecraft.

Autonomous Power Controller

Design and Verification of a Distributed Communication Protocol

The safety of remotely operated vehicles depends on the correctness of the distributed protocol that facilitates the communication between the vehicle and the operator. A failure in this communication can result in catastrophic loss of the vehicle. To complicate matters, the communication system may be required to satisfy several, possibly conflicting, requirements. The design of protocols is typically an informal process based on successive iterations of a prototype implementation. Yet distributed protocols are notoriously difficult to get correct using such informal techniques. We present a formal specification of the design of a distributed protocol intended for use in a remotely operated vehicle, which is built from the composition of several simpler protocols. We demonstrate proof strategies that allow us to prove properties of each component protocol individually while ensuring that the property is preserved in the composition forming the entire system. Given that designs are likely to evolve as additional requirements emerge, we show how we have automated most of the repetitive proof steps to enable verification of rapidly changing designs.

Munoz, Cesar A.

Full-Scale Wind-Tunnel Investigation of Wing-Cooling Ducts Effects of Propeller Slipstream, Special Report

The safety of remotely operated vehicles depends on the correctness of the distributed protocol that facilitates the communication between the vehicle and the operator. A failure in this communication can result in catastrophic loss of the vehicle. To complicate matters, the communication system may be required to satisfy several, possibly conflicting, requirements. The design of protocols is typically an informal process based on successive iterations of a prototype implementation. Yet distributed protocols are notoriously difficult to get correct using such informal techniques. We present a formal specification of the design of a distributed protocol intended for use in a remotely operated vehicle, which is built from the composition of several simpler protocols. We demonstrate proof strategies that allow us to prove properties of each component protocol individually while ensuring that the property is preserved in the composition forming the entire system. Given that designs are likely to evolve as additional requirements emerge, we show how we have automated most of the repetitive proof steps to enable verification of rapidly changing designs.

Nickle, F. R.

Autonomous Assessment and Predictive Capabilities for Low-Altitude Urban Flight Operations

The integration of unmanned aerial vehicles in the national airspace will introduce new vehicle types, technologies, and operational paradigms for which safety must be maintained and hazards mitigated. One approach is to attempt to design for possible hazards and unsafe incidents that can occur at different phases of flight (pre-flight, in-flight, and post-flight) and during ground operations. Another is to mitigate safety incidents by implementing changes to policies, procedures, regulations, and design to cover personnel, equipment, and aircraft during operations. These and other techniques, not described herein, are typically conservative or adhoc in that they reduce the likelihood of risk after safety incidents have occurred. In this work, the goal is to develop a more predictive capability to monitor and mitigate risk and hazards to safety “in-time” enough for decisions to be made. In line with NASA’s Aeronautics Mission Directorate Strategic Thrust 5 [1] (In-Time System-Wide Safety Assurance), the System-Wide Safety (SWS) project under which this work falls, is developing and demonstrating innovative and safety-oriented solutions that enable modernization and aviation transformation. To that effect, this work will detail data-driven efforts on the SWS project to develop a number of safety-critical services for in-time monitoring and mitigation of hazards to low-altitude flight operations. First, hazards to these operations are identified based on previous work by NASA [2,3] and others in the aerospace industry. These hazards include (i) unsafe proximity to other vehicles, property, and people on the ground, (ii) critical system failures such as communication signal/GPS loss, unexpected propulsion system degradation, engine/power failure, and (iii) operational/environmental issues such as severe weather and gusty winds. For these hazards, safety metrics, which can be quantified and assessed are defined, models to monitor and predict them are developed, and flight test data is generated to develop, validate, and test these models, considering the complex interplay of the different hazards that define them [4-6]. In addition, the uncertainty in the non-deterministic effects that cannot be modeled nor predicted and unknown unknowns that arise after design/testing and during operations must be handled in rigorous manner. As a result, for each of the developed safety metrics, their dependencies on one another are characterized and a framework for handling the uncertainties inherent in the modeling, algorithms, and measurements required for prediction is also developed [7]. To that effect, this presentation will describe the safety metrics and services already developed and underway under the System-Wide Safety project that utilize data-driven techniques for the identification of anomalies, precursors, and trends (APTs) to monitor and mitigate hazards to safety, in-time, for urban flight operations in low-altitude airspace.

Okolo, Wendy A.

Ayame/PAM-D apogee kick motor nozzle failure analysis

The failure of two communication satellites during firing sequence were examined. The correlation/comparison of the circumstances of the Ayame incidents and the failure of the STAR 48 (DM-2) motor are reviewed. The massive nozzle failure of the AKM to determine the impact on spacecraft performance is examined. It is recommended that a closer watch is kept on systems techniques,

Source record

Briefcase Communicator

In the photo at bottom right, a U.S. Park Police officer is demonstrating a battery-powered communications system, sufficiently compact to be packed in a briefcase-size container, which can send and receive signals over great distances by means of satellite relay. Key to the system's efficacy is the high-powered transmitting and receiving equipment aboard such NASA satellites as the Applications Technology Satellite6 (ATS-6) and the joint U.S.-Canadian Communications Technology Satellite (CTS); this enables the briefcase communicator to pick up satellite-relayed signals by means of the small hook-on antenna shown instead of the more elaborate-ground equipment customarily needed. Developed by NASA's Goddard Space Flight Center, the communicator is intended for use in emergency situations. It has utility, for example, in disasters, such as floods and hurricanes, where power failure disrupts conventional communications; for on-the-spot transmissions from major accident sites; or in remote areas where no other means of communication exists

Source record

Cyber-Threat Assessment for the Air Traffic Management System: A Network Controls Approach

Air transportation networks are being disrupted with increasing frequency by failures in their cyber- (computing, communication, control) systems. Whether these cyber- failures arise due to deliberate attacks or incidental errors, they can have far-reaching impact on the performance of the air traffic control and management systems. For instance, a computer failure in the Washington DC Air Route Traffic Control Center (ZDC) on August 15, 2015, caused nearly complete closure of the Centers airspace for several hours. This closure had a propagative impact across the United States National Airspace System, causing changed congestion patterns and requiring placement of a suite of traffic management initiatives to address the capacity reduction and congestion. A snapshot of traffic on that day clearly shows the closure of the ZDC airspace and the resulting congestion at its boundary, which required augmented traffic management at multiple locations. Cyber- events also have important ramifications for private stakeholders, particularly the airlines. During the last few months, computer-system issues have caused several airlines fleets to be grounded for significant periods of time: these include United Airlines (twice), LOT Polish Airlines, and American Airlines. Delays and regional stoppages due to cyber- events are even more common, and may have myriad causes (e.g., failure of the Department of Homeland Security systems needed for security check of passengers, see [3]). The growing frequency of cyber- disruptions in the air transportation system reflects a much broader trend in the modern society: cyber- failures and threats are becoming increasingly pervasive, varied, and impactful. In consequence, an intense effort is underway to develop secure and resilient cyber- systems that can protect against, detect, and remove threats, see e.g. and its many citations. The outcomes of this wide effort on cyber- security are applicable to the air transportation infrastructure, and indeed security solutions are being implemented in the current system. While these security solutions are important, they only provide a piecemeal solution. Particular computers or communication channels are protected from particular attacks, without a holistic view of the air transportation infrastructure. On the other hand, the above-listed incidents highlight that a holistic approach is needed, for several reasons. First, the air transportation infrastructure is a large scale cyber-physical system with multiple stakeholders and diverse legacy assets. It is impractical to protect every cyber- asset from known and unknown disruptions, and instead a strategic view of security is needed. Second, disruptions to the cyber- system can incur complex propagative impacts across the air transportation network, including its physical and human assets. Also, these implications of cyber- events are exacerbated or modulated by other disruptions and operational specifics, e.g. severe weather, operator fatigue or error, etc. These characteristics motivate a holistic and strategic perspective on protecting the air transportation infrastructure from cyber- events. The analysis of cyber- threats to the air traffic system is also inextricably tied to the integration of new autonomy into the airspace. The replacement of human operators with cyber functions leaves the network open to new cyber threats, which must be modeled and managed. Paradoxically, the mitigation of cyber events in the airspace will also likely require additional autonomy, given the fast time scale and myriad pathways of cyber-attacks which must be managed. The assessment of new vulnerabilities upon integration of new autonomy is also a key motivation for a holistic perspective on cyber threats.

Complex Networks

Automated Operations for Galileo Communications

Following the deployment failure of Galileo's high gain antenna, the downlink had to be redesigned so as to effectively use the low gain antenna. The downlink was redesigned to maximize the data return and increase the reliability which required the reconfiguration of the onboard software and the deep space network. The revised downlink features: data compression; antenna arraying; the recoding and reprocessing of telemetry; suppressed carrier tracking, and error-correction coding. The deep space network Galileo telemetry (DGT) subsystem was developed and deployed at three sites in Australia, Spain and the U.S. The DGT was designed as an automated system that continuously monitors and adjusts its parameters and environment in response to either pre-loaded sequences or changes in the internal status.

Statman, Joseph I.