Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “software resilience”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Risk-Significant Adverse Condition Awareness Strengthens Assurance of Fault Management Systems

As spaceflight systems increase in complexity, Fault Management (FM) systems are ranked high in risk-based assessment of software criticality, emphasizing the importance of establishing highly competent domain expertise to provide assurance. Adverse conditions (ACs) and specific vulnerabilities encountered by safety- and mission-critical software systems have been identified through efforts to reduce the risk posture of software-intensive NASA missions. Acknowledgement of potential off-nominal conditions and analysis to determine software system resiliency are important aspects of hazard analysis and FM. A key component of assuring FM is an assessment of how well software addresses susceptibility to failure through consideration of ACs. Focus on significant risk predicted through experienced analysis conducted at the NASA Independent Verification & Validation (IV&V) Program enables the scoping of effective assurance strategies with regard to overall asset protection of complex spaceflight as well as ground systems. Research efforts sponsored by NASAs Office of Safety and Mission Assurance (OSMA) defined terminology, categorized data fields, and designed a baseline repository that centralizes and compiles a comprehensive listing of ACs and correlated data relevant across many NASA missions. This prototype tool helps projects improve analysis by tracking ACs and allowing queries based on project, mission type, domain/component, causal fault, and other key characteristics. Vulnerability in off-nominal situations, architectural design weaknesses, and unexpected or undesirable system behaviors in reaction to faults are curtailed with the awareness of ACs and risk-significant scenarios modeled for analysts through this database. Integration within the Enterprise Architecture at NASA IV&V enables interfacing with other tools and datasets, technical support, and accessibility across the Agency. This paper discusses the development of an improved workflow process utilizing this database for adaptive, risk-informed FM assurance that critical software systems will safely and securely protect against faults and respond to ACs in order to achieve successful missions.

IV&V↗

Risk-Significant Adverse Condition Awareness Strengthens Assurance of Fault Management Systems

As spaceflight systems increase in complexity, Fault Management (FM) systems are ranked high in risk-based assessment of software criticality, emphasizing the importance of establishing highly competent domain expertise to provide assurance. Adverse conditions (ACs) and specific vulnerabilities encountered by safety- and mission-critical software systems have been identified through efforts to reduce the risk posture of software-intensive NASA missions. Acknowledgement of potential off-nominal conditions and analysis to determine software system resiliency are important aspects of hazard analysis and FM. A key component of assuring FM is an assessment of how well software addresses susceptibility to failure through consideration of ACs. Focus on significant risk predicted through experienced analysis conducted at the NASA Independent Verification Validation (IVV) Program enables the scoping of effective assurance strategies with regard to overall asset protection of complex spaceflight as well as ground systems. Research efforts sponsored by NASA's Office of Safety and Mission Assurance defined terminology, categorized data fields, and designed a baseline repository that centralizes and compiles a comprehensive listing of ACs and correlated data relevant across many NASA missions. This prototype tool helps projects improve analysis by tracking ACs and allowing queries based on project, mission type, domaincomponent, causal fault, and other key characteristics. Vulnerability in off-nominal situations, architectural design weaknesses, and unexpected or undesirable system behaviors in reaction to faults are curtailed with the awareness of ACs and risk-significant scenarios modeled for analysts through this database. Integration within the Enterprise Architecture at NASA IVV enables interfacing with other tools and datasets, technical support, and accessibility across the Agency. This paper discusses the development of an improved workflow process utilizing this database for adaptive, risk-informed FM assurance that critical software systems will safely and securely protect against faults and respond to ACs in order to achieve successful missions.

Fault management↗

Resilient Autonomy in the Face of Adversity

The NASA Resilient Autonomy Project developed a software framework that implemented a Run Time Assurance (RTA) architecture that leveraged ASTM International’s F3269 Industry Standard for safely bounding complex behavior in aircraft. This framework was called the Expandable Variable Autonomy Architecture, or EVAA. EVAA was developed during the height of the Covid-19 lockdown that caused the Resilient Autonomy team to pivot from flight test to distributed simulator testing. EVAA was developed to be platform and mission agnostic where platform specifics were behind a hardware abstraction layer that EVAA called a Coupler. EVAA was able to host multiple safety monitors that could resolve individual safety hazards. EVAA was able to resolve priority conflicts when multiple safety hazards needed to be resolved simultaneously and was able to resolve highly complex situations in a safe manner that could exceed human capabilities.

Ethan Williams↗

Historical Aerospace Software Errors Categorized to Influence Fault Tolerance

Since the first use of computers in space and aircraft, software errors have occurred. These errors can manifest as loss-of-life or less catastrophically. As the demand for automation increases, software in mission or safety-critical systems should be designed to be tolerant to the most likely software faults. This paper categorizes a set of 55 historic aerospace software error incidents from 1962 to 2023 to determine trends of how and where automation is most likely to fail, behaving unexpectedly. A distinction between software producing unexpected (erroneous) output versus no output (failsilent) is introduced. Of the historical incidents analyzed, 85% were from software producing wrong output rather than simply stopping. Rebooting was found to be ineffective to clear erroneous behavior, and not reliable to recover from silent failures. Error origin was within the code/logic itself in 58% of cases, 16% from configurable data, 15% from unexpected sensor input, and 11% from command/operator input. A substantial forty percent (40%) of unexpected software behavior was indicated by the absence of code, arising from unanticipated situations and missing requirements, and 16% of incidents were subjectively deemed “unknown-unknowns”. No incidents were found to be the result of programming language, compiler, tool, or operating system; and only sixteen percent (16%) of all incidents were considered errors traditional computer science/programming in nature. These findings indicate that for fault tolerance, erroneous automation behavior must be a primary consideration especially at critical moments, and reboot recoverability may not be viable. Special care should be taken to validate configurable data and commands prior to use. “Test-like-you-fly”, including hardware-in-the-loop combined with robust off-nominal testing should be used to uncover missing logic arising from unanticipated situations not covered by requirements alone. This study uniquely focuses on manifestations of unexpected flight software behavior, independent of ultimate root cause. We characterize software error behavior and origin to improve software design, test, and operations for resilience to the most common manifestations, and provide a rich dataset for further study.

Aerospace↗

Model-Driven Development For PDS4 Software And Services

Software and services that access Planetary Data System (PDS) PDS4 data products need to parse product labels to retrieve, interpret, and process the referenced digital objects. Under PDS4 a driving principle is that the product label provide all of the information necessary for these functions to be performed accurately. However, significantly more information is available in the PDS4 Information Model (IM)[1], the controlling document used to define, create, and syntactically and semantically verify the product labels. This additional information in the IM is made available for use, by both software and services, to configure, promote resiliency, and improve interoperability.

Padams, Jordan↗

Alternative Metrics for Evaluating the Resilence of Advanced Life Support Systems

Ensuring the safety of the crew is a key performance requirement of a life support system. However, a number of conceptual and practical difficulties arise when devising metrics to concretely measure the ability of a life support system to maintain critical functions in the presence of anticipated and unanticipated faults. Resilience is a dynamic property of a life support system that depends on the complex interactions between faults, controls and system hardware. We review some of the approaches to understanding the robustness or resilience of complex systems being developed in diverse fields such as ecology, software engineering and cell biology and discuss their applicability to regenerative life support systems. We also consider how approaches to measuring resilience vary depending on system design choices such as the definition and choice of the nominal operating regime. Finally, we explore data collection and implementation issues such as the key differences between the instantaneous or conditional and average or overall measures of resilience. Extensive simulation of a hybrid computational model of a water revitalization subsystem (WRS) with probabilistic, component-level faults provides data about off-nominal behavior of the system. The data are used to consider alternative measures of resilience as predictors of the system's ability to recover from component-level faults.

Bell, Ann Maria↗

Spacecraft Avionics Software Development Then and Now: Different but the Same

NASA has always been in the business of balancing new technologies and techniques to achieve human space travel objectives. NASA s historic Software Production Facility (SPF) was developed to serve complex avionics software solutions during an era dominated by mainframes, tape drives, and lower level programming languages. These systems have proven themselves resilient enough to serve the Shuttle Orbiter Avionics life cycle for decades. The SPF and its predecessor the Software Development Lab (SDL) at NASA s Johnson Space Center (JSC) hosted flight software (FSW) engineering, development, simulation, and test. It was active from the beginning of Shuttle Orbiter development in 1972 through the end of the shuttle program in the summer of 2011 almost 40 years. NASA s Kedalion engineering analysis lab is on the forefront of validating and using many contemporary avionics HW/SW development and integration techniques, which represent new paradigms to NASA s heritage culture in avionics software engineering. Kedalion has validated many of the Orion project s HW/SW engineering techniques borrowed from the adjacent commercial aircraft avionics environment, inserting new techniques and skills into the Multi-Purpose Crew Vehicle (MPCV) Orion program. Using contemporary agile techniques, COTS products, early rapid prototyping, in-house expertise and tools, and customer collaboration, NASA has adopted a cost effective paradigm that is currently serving Orion effectively. This paper will explore and contrast differences in technology employed over the years of NASA s space program, due largely to technological advances in hardware and software systems, while acknowledging that the basic software engineering and integration paradigms share many similarities.

Mangieri, Mark L.↗

Radiation Tolerance and Mitigation for Neuromorphic Processors

Neuromorphic processors are designed to execute Deep Neural Networks (DNNs) at very high speed using only a fraction of the electrical power needed to run a DNN on a traditional CPU or GPU. This unique capability makes Neuromorphic processors a prime candidate for space systems, where advanced computational tasks like image analysis, depth map reconstruction, or rover control need to be executed in a power-starved environment. In contrast to the growing number of applications of Neuromorphic processors in smart phones, the automotive and robotics domain, the space environment is unforgiving because of extreme temperatures and high levels of radiation. Any space system, operating beyond LEO requires computing hardware that is resilient against radiation effects. However, Neuromorphic processors have not yet been designed or tested for their radiation tolerance. In this report, we consider traditional methods of detection of radiation events and mitigation via redundancy and gauge their effectiveness on DNNs. In contrast to traditional flight software, however, neural networks represent a statistical algorithm, which might affect its resilience against radiation events. We will focus on the analysis of the tolerance of DNNs with respect to radiation events and discuss techniques to detect radiation hits using on-chip triple modular redundancy (TMR) on an Intel Loihi neuromorphic processor and to mitigate radiation damage. We describe an architecture for on-chip TMR for the Intel Loihi and present results of initial experiments.

Neural Networks↗

Resilient Space Habitat Design Using Safety Controls

Space habitats will involve a complex and tightly coupled combination of hardware, software, and humans, while operating in challenging environments that pose many risks, both known and unknown. It will not be possible to design habitats that are immune to failure, nor will it be possible to foresee all possible failures. Rather than aiming for designs where ―failure is not an option,‖ habitats must be resilient to disruptions. We propose an approach to resilient design for space habitats based on the concept of safety controls from system safety engineering. We model disruptions using a state-and-trigger approach, where the space habitat is in one of three distinct states at each time instance: nominal, hazardous, or accident. We use safety controls as ways of preventing a system from entering or remaining in a hazardous or accident state. We develop a safety control option space for the habitat, from which designers can select the set of safety controls that best meet resilience, performance, and other system goals. The safety control option space is likely to be large, accordingly, we design a database that links safety controls to the applicable states and triggers. We demonstrate our approach on the early design stage of a Martian space habitat.

Safety↗

Overview of NASA's Air Traffic Management - eXploration (ATM-X) Project

Projected increases in new vehicle types, new missions, and the continual growth in traditional (e.g., airlines, general aviation) aviation will require changes to the current air traffic system, particularly to accommodate the desire of operators to be more involved in air traffic decisions. To address these challenges, the National Airspace System needs to undergo a transformation to a more scalable, flexible, user-focused system that addresses safety and security requirements and resiliency for current and new users. A system designed to integrate modular software services, provided by users, third parties and government for air traffic management functions, will be scalable and more easily allow modernization and for collaboration between users and service providers. ATM-X is responding to NASA's pivot towards integrating projected new, diverse entrants into the NAS, while also leveraging NASA's prior ATM achievements that continue to improve traditional airspace operations. This project is a two-phased approach to conduct research and focused evaluations to assess the feasibility of a service-based approach and to identify critical design considerations to enable airspace access for new entrants, integrated with current traditional operations. Phase 1 research will be conducted to determine what is needed to reach the ATM-X goals based on specific use-cases to enable large-scale, passenger-carrying Urban Air Mobility operations in a metroplex environment, and also to improve traditional operations in the Northeast Region leveraging mature NASA technologies. Some of these evaluations will be conducted in simulations and field activities. Phase 2 will build upon Phase 1 towards more defined, focused research and field demonstrations in real-world environments to integrate multiple elements of a scalable, service-based ATM-X concept.

air traffic management↗

Contingency Software in Autonomous Systems: Technical Level Briefing

Contingency management is essential to the robust operation of complex systems such as spacecraft and Unpiloted Aerial Vehicles (UAVs). Automatic contingency handling allows a faster response to unsafe scenarios with reduced human intervention on low-cost and extended missions. Results, applied to the Autonomous Rotorcraft Project and Mars Science Lab, pave the way to more resilient autonomous systems.

autonomous systems↗

Research in computer science

Several short summaries of the work performed during this reporting period are presented. Topics discussed in this document include: (1) resilient seeded errors via simple techniques; (2) knowledge representation for engineering design; (3) analysis of faults in a multiversion software experiment; (4) implementation of parallel programming environment; (5) symbolic execution of concurrent programs; (6) two computer graphics systems for visualization of pressure distribution and convective density particles; (7) design of a source code management system; (8) vectorizing incomplete conjugate gradient on the Cyber 203/205; (9) extensions of domain testing theory and; (10) performance analyzer for the pisces system.

Ortega, J. M.↗

Towards Real-Time, On-Board, Hardware-Supported Sensor and Software Health Management for Unmanned Aerial Systems

For unmanned aerial systems (UAS) to be successfully deployed and integrated within the national airspace, it is imperative that they possess the capability to effectively complete their missions without compromising the safety of other aircraft, as well as persons and property on the ground. This necessity creates a natural requirement for UAS that can respond to uncertain environmental conditions and emergent failures in real-time, with robustness and resilience close enough to those of manned systems. We introduce a system that meets this requirement with the design of a real-time onboard system health management (SHM) capability to continuously monitor sensors, software, and hardware components. This system can detect and diagnose failures and violations of safety or performance rules during the flight of a UAS. Our approach to SHM is three-pronged, providing: (1) real-time monitoring of sensor and software signals; (2) signal analysis, preprocessing, and advanced on-the-fly temporal and Bayesian probabilistic fault diagnosis; and (3) an unobtrusive, lightweight, read-only, low-power realization using Field Programmable Gate Arrays (FPGAs) that avoids overburdening limited computing resources or costly re-certification of flight software. We call this approach rt-R2U2, a name derived from its requirements. Our implementation provides a novel approach of combining modular building blocks, integrating responsive runtime monitoring of temporal logic system safety requirements with model-based diagnosis and Bayesian network-based probabilistic analysis. We demonstrate this approach using actual flight data from the NASA Swift UAS.

Unmanned Aerial System↗

Robot Hand

Robots are limited only by the dexterity of the hand. Dr. Salisbury, in conjunction with Stanford, Caltech and Jet Propulsion Laboratory, developed the Salisbury Hand which has three, three-jointed human-like fingers. The tips are covered with a resilient, high friction material for gripping. The robot hand can manipulate objects by finger motion, and adapts to different aims. Advanced software allows the hand to interpret information from fingertip sensors. Further development is expected. A company has been formed to reproduce the device; copies have been delivered to several laboratories.

Source record↗

Systems Architecture for Fully Autonomous Space Missions

The NASA Goddard Space Flight Center is working to develop a revolutionary new system architecture concept in support of fully autonomous missions. As part of GSFC's contribution to the New Millenium Program (NMP) Space Technology 7 Autonomy and on-Board Processing (ST7-A) Concept Definition Study, the system incorporates the latest commercial Internet and software development ideas and extends them into NASA ground and space segment architectures. The unique challenges facing the exploration of remote and inaccessible locales and the need to incorporate corresponding autonomy technologies within reasonable cost necessitate the re-thinking of traditional mission architectures. A measure of the resiliency of this architecture in its application to a broad range of future autonomy missions will depend on its effectiveness in leveraging from commercial tools developed for the personal computer and Internet markets. Specialized test stations and supporting software come to past as spacecraft take advantage of the extensive tools and research investments of billion-dollar commercial ventures. The projected improvements of the Internet and supporting infrastructure go hand-in-hand with market pressures that provide continuity in research. By taking advantage of consumer-oriented methods and processes, space-flight missions will continue to leverage on investments tailored to provide better services at reduced cost. The application of ground and space segment architectures each based on Local Area Networks (LAN), the use of personal computer-based operating systems, and the execution of activities and operations through a Wide Area Network (Internet) enable a revolution in spacecraft mission formulation, implementation, and flight operations. Hardware and software design, development, integration, test, and flight operations are all tied-in closely to a common thread that enables the smooth transitioning between program phases. The application of commercial software development techniques lays the foundation for delivery of product-oriented flight software modules and models. Software can then be readily applied to support the on-board autonomy required for mission self-management. An on-board intelligent system, based on advanced scripting languages, facilitates the mission autonomy required to offload ground system resources, and enables the spacecraft to manage itself safely through an efficient and effective process of reactive planning, science data acquisition, synthesis, and transmission to the ground. Autonomous ground systems in turn coordinate and support schedule contact times with the spacecraft. Specific autonomy software modules on-board include mission and science planners, instrument and subsystem control, and fault tolerance response software, all residing within a distributed computing environment supported through the flight LAN. Autonomy also requires the minimization of human intervention between users on the ground and the spacecraft, and hence calls for the elimination of the traditional operations control center as a funnel for data manipulation. Basic goal-oriented commands are sent directly from the user to the spacecraft through a distributed internet-based payload operations "center". The ensuing architecture calls for the use of spacecraft as point extensions on the Internet. This paper will detail the system architecture implementation chosen to enable cost-effective autonomous missions with applicability to a broad range of conditions. It will define the structure needed for implementation of such missions, including software and hardware infrastructures. The overall architecture is then laid out as a common thread in the mission life cycle from formulation through implementation and flight operations.

Esper, Jamie↗

Cyber Resiliency and the Implementation of a Host-Based Intrusion Detection System in an Urban Air Mobility Environment

With the growth in Urban Air Mobility systems and the increasing reliance on interconnected technologies, ensuring the security of these complex components has become critical. As cities evolve into smart urban centers, the vulnerability to cyber threats escalates, possibly endangering citizens safety and the efficiency of transportation networks.In response to these challenges, this paper presents a study on the need for cyber resilient techniques within future air traffic environments. It will pay specific attention to the implementation of a Host-Based Intrusion Detection System (HIDS) utilizing Atomic OSSEC software, tailored specifically to a NASA simulation of an UrbanAirMobility environments’ unique demands. Further, this study seeks to outline the rational for NASA’s recommendation for a HIDS in such environments. It explores the design, development, and deployment of the proposed HIDS, focusing on its adaptability to monitor the hybrid nature of the Urban Air Mobility environment. Leveraging machine learning algorithms and anomaly detection techniques, the HIDS is equipped to continuously monitor and analyze the behavior of individual host systems, vehicles, and devices, thereby providing a proactive approach to threat detection. Implementing a HIDS is a pivotal strategy for enhancing cyber resiliency, as it gives an organization granular visibility into internal system activities, enables rapid detection and response to anomalous behavior and cyber threats, and fortifies the organizations overall cybersecurity posture. Finally, this study aims to provide recommendations and include learned takeaways that the Urban Air Mobility industry should consider. In brief, this paper highlights the significance of host-based intrusion detection in UrbanAirMobility environments and underscores the necessity of tailored security solutions to safeguard against emerging cyber threats.

UAM↗

Unsupervised Anomaly Detection in High-Dimensional Flight Data Using Convolutional Variational Auto-Encoder

The modern National Airspace System (NAS) is an extremely safe system and the aviation industry has experienced a steady decrease in fatalities over the years. This can be attributed to both improved flight critical systems with redundant hardware and software protections, as well as an increased focus on active monitoring and response to real time and historically identified vulnerabilities by implementing more resilient procedures and protocols. The main approach for identifying vulnerabilities in operations leverages domain expertise using knowledge about how the system should behave within the expected tolerances to known safety margins. This approach works well when the system has a well-defined operating condition. However, the operations in the NAS can be highly complex with various nuances that render it difficult to clearly pre-define all known safety vulnerabilities. With the advancement of data science and machine learning techniques, the potential to automatically identify emerging vulnerabilities in the observed operations has become more practical in recent years. The state-of-the-art anomaly detection approaches in aerospace data usually rely on supervised or semi-supervised learning. However, in many real-world problems such as flight safety, creating labels for the data requires huge amount of effort and is largely impractical. To address this challenge, we developed a Convolutional Variational Auto-Encoder (CVAE), which is an unsupervised learning approach for anomaly detection in high-dimensional heterogeneous time-series data. We validate performance of CVAE compared to the state-of-the-art supervised learning approach as well as unsupervised clustering-based approach using KMeans++ and kernel-based approach using One-Class Support Vector Machine (OC-SVM) on Yahoo!'s benchmark time series anomaly detection data. Finally, we showcase performance of CVAE on a case study of identifying anomalies in the first 60 seconds of commercial flights' take-offs using Flight Operational Quality Assurance (FOQA) data.

Memarzadeh, Milad↗