Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Engineering Risk Management”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

SOFIA Program SE and I Lessons Learned

Once a "Troubled Project" threatened with cancellation, the Stratospheric Observatory for Infrared Astronomy (SOFIA) Program has overcome many difficult challenges and recently achieved its first light images. To achieve success, SOFIA had to overcome significant deficiencies in fundamental Systems Engineering identified during a major Program restructuring. This presentation will summarize the lessons learn in Systems Engineering on the SOFIA Program. After the Program was reformulated, an initial assessment of Systems Engineering established the scope of the problem and helped to set a list of priorities that needed to be work. A revised Systems Engineering Management Plan (SEMP) was written to address the new Program structure and requirements established in the approved NPR7123.1A. An important result of the "Technical Planning" effort was the decision by the Program and Technical Leadership team to re-phasing the lifecycle into increments. The reformed SOFIA Program Office had to quickly develop and establish several new System Engineering core processes including; Requirements Management, Risk Management, Configuration Management and Data Management. Implementing these processes had to consider the physical and cultural diversity of the SOFIA Program team which includes two Projects spanning two NASA Centers, a major German partnership, and sub-contractors located across the United States and Europe. The SOFIA Program experience represents a creative approach to doing "System Engineering in the middle" while a Program is well established. Many challenges were identified and overcome. The SOFIA example demonstrates it is never too late to benefit from fixing deficiencies in the System Engineering processes.

Ray, Ronald J.

Root Source Analysis/ValuStream[Trade Mark] - A Methodology for Identifying and Managing Risks

Root Source Analysis (RoSA) is a systems engineering methodology that has been developed at NASA over the past five years. It is designed to reduce costs, schedule, and technical risks by systematically examining critical assumptions and the state of the knowledge needed to bring to fruition the products that satisfy mission-driven requirements, as defined for each element of the Work (or Product) Breakdown Structure (WBS or PBS). This methodology is sometimes referred to as the ValuStream method, as inherent in the process is the linking and prioritizing of uncertainties arising from knowledge shortfalls directly to the customer's mission driven requirements. RoSA and ValuStream are synonymous terms. RoSA is not simply an alternate or improved method for identifying risks. It represents a paradigm shift. The emphasis is placed on identifying very specific knowledge shortfalls and assumptions that are the root sources of the risk (the why), rather than on assessing the WBS product(s) themselves (the what). In so doing RoSA looks forward to anticipate, identify, and prioritize knowledge shortfalls and assumptions that are likely to create significant uncertainties/ risks (as compared to Root Cause Analysis, which is most often used to look back to discover what was not known, or was assumed, that caused the failure). Experience indicates that RoSA, with its primary focus on assumptions and the state of the underlying knowledge needed to define, design, build, verify, and operate the products, can identify critical risks that historically have been missed by the usual approaches (i.e., design review process and classical risk identification methods). Further, the methodology answers four critical questions for decision makers and risk managers: 1. What s been included? 2. What's been left out? 3. How has it been validated? 4. Has the real source of the uncertainty/ risk been identified, i.e., is the perceived problem the real problem? Users of the RoSA methodology have characterized it as a true bottoms up risk assessment.

Brown, Richard Lee

Accounting for Point Estimate Uncertainty in Space Systems Reliability and Risk Analysis

Understanding and accounting for uncertainty in risk analysis is a critical step in the management and communication of risk in engineered systems. The component and system-level analysis to determine the probability of a negative outcome and its consequence is often quantified by a point estimate. Many Program and Enterprise decisions involving technical concerns and issues rely on reliability engineering activities to produce quantified risk analysis to inform the decision making process. At NASA, it is common to use a Probabilistic Risk Analysis (PRA) to inform the overall risk to Loss of Mission or Loss of Crew that involves integration across all spacecraft subsystem fault trees to produce an overall probability of mission failure. The point estimate is an estimate of this overall probability and is an immediate result of a fault tree model. It is the result of a model where the probability of each event is taken to be equal to its mean. The value provides an approximation of the overall mean without running any uncertainty calculations (e.g., no sampling). Using only the point estimate can lead to a false sense of precision and the point estimate may not match the resulting mean when uncertainty is taken into consideration. This paper will explore five conditions that can cause the PRA model mean to diverge from the point estimate and will provide engineers and managers insight into the importance of understanding uncertainty in the elements of PRA models.

Paul J Collier

Structured prototyping as risk management

A methodology is presented for integrating the systems-engineering management recommendation of prototyping into the traditional project-management process for developing large-scale systems. The suggested methodology begins with the identification of life-cycle risk areas, outlines the structure and conduct of the prototyping process, and defines the composition of the prototyping team. The methodology includes a step-by-step procedure for creating, executing, and documenting a prototyping test plan to evaluate design alternatives. It is argued that managers who adopt this methodology and apply it rigorously will increase the likelihood that the systems they build will be operationally effective and will be accepted by the intended users.

Hornstein, Rhoda SH.

Risk Management: A Practical Design Tool For Space Systems and Technology Development

Over the past two decades, risk management and risk analysis have emerged throughout the business community in the United States (US) as prominent planning and development strategies used to mitigate risk of failure and ensure a high return on investment (ROI) for business endeavors (financial and otherwise). They are generic tools that can be applied to any business regardless of the sector (i.e., government, university, private) and have been used by the Federal government in the form of institutional practices aimed at maximizing the probability of success in business activities. One US Federal agency that incorporates risk management and analysis techniques into business and/or engineering activities is the National Aeronautics and Space Administration (NASA). The present work is a discussion on mission, spacecraft and instrument design (as well as technology development) and the role of risk management, analysis and mitigation as a fundamental tool in the design process.

Silk, Eric A.

The Value of Identifying and Recovering Lost GN&C Lessons Learned: Aeronautical, Spacecraft, and Launch Vehicle Examples

Within the broad aerospace community the importance of identifying, documenting and widely sharing lessons learned during system development, flight test, operational or research programs/projects is broadly acknowledged. Documenting and sharing lessons learned helps managers and engineers to minimize project risk and improve performance of their systems. Often significant lessons learned on a project fail to get captured even though they are well known 'tribal knowledge' amongst the project team members. The physical act of actually writing down and documenting these lessons learned for the next generation of NASA GN&C engineers fails to happen on some projects for various reasons. In this paper we will first review the importance of capturing lessons learned and then will discuss reasons why some lessons are not documented. A simple proven approach called 'Pause and Learn' will be highlighted as a proven low-impact method of organizational learning that could foster the timely capture of critical lessons learned. Lastly some examples of 'lost' GN&C lessons learned from the aeronautics, spacecraft and launch vehicle domains are briefly highlighted. In the context of this paper 'lost' refers to lessons that have not achieved broad visibility within the NASA-wide GN&C CoP because they are either undocumented, masked or poorly documented in the NASA Lessons Learned Information System (LLIS).

Dennehy, Cornelius J.

Human System Risk Communication: Directed Acyclic Graphs

- The Human System Risk Board (HSRB) is responsible for the management of a portfolio of 30 human system risks that NASA tracks and configuration manages to mitigate for future crewed exploration missions. - The HSRB has been exploring the concept of causal diagrams (in the form of Directed Acyclic Graphs or DAGs) as an approach to creating knowledge graphs for each risk to enable shared mental models of causal flow from spaceflight hazards to mission outcomes among HSRB Stakeholders. - These diagrams are intended to improve insight and communication of risk across the myriad subject matter experts and management interested in human system risk reduction. This includes program managers, systems engineers, and operators in addition to the Human Health and Performance Directorate. - The DAG project was intended to create the foundation for composition of the 30 baselined DAGs into a single risk network and software is being developed in parallel to enable this forward work.

directed acrylic graph

Systems Engineering with a Focus on Failure Prevention

A discussion of Systems Engineering principles as related to failure prevention methodologies and technologies. The presentation outlines requirements development, design reviews, verification and validation, and finally risk management prectices.

Systems Engineering, Failure Prevention

NASA Applications and Lessons Learned in Reliability Engineering

Since the Shuttle Challenger accident in 1986, communities across NASA have been developing and extensively using quantitative reliability and risk assessment methods in their decision making process. This paper discusses several reliability engineering applications that NASA has used over the year to support the design, development, and operation of critical space flight hardware. Specifically, the paper discusses several reliability engineering applications used by NASA in areas such as risk management, inspection policies, components upgrades, reliability growth, integrated failure analysis, and physics based probabilistic engineering analysis. In each of these areas, the paper provides a brief discussion of a case study to demonstrate the value added and the criticality of reliability engineering in supporting NASA project and program decisions to fly safely. Examples of these case studies discussed are reliability based life limit extension of Shuttle Space Main Engine (SSME) hardware, Reliability based inspection policies for Auxiliary Power Unit (APU) turbine disc, probabilistic structural engineering analysis for reliability prediction of the SSME alternate turbo-pump development, impact of ET foam reliability on the Space Shuttle System risk, and reliability based Space Shuttle upgrade for safety. Special attention is given in this paper to the physics based probabilistic engineering analysis applications and their critical role in evaluating the reliability of NASA development hardware including their potential use in a research and technology development environment.

Safie, Fayssal M.

Continuous Risk Management: An Overview

Software risk management is important because it helps avoid disasters, rework, and overkill, but more importantly because it stimulates win-win situations. The objectives of software risk management are to identify, address, and eliminate software risk items before they become threats to success or major sources of rework. In general, good project managers are also good managers of risk. It makes good business sense for all software development projects to incorporate risk management as part of project management. The Software Assurance Technology Center (SATC) at NASA GSFC has been tasked with the responsibility for developing and teaching a systems level course for risk management that provides information on how to implement risk management. The course was developed in conjunction with the Software Engineering Institute at Carnegie Mellon University, then tailored to the NASA systems community. This is an introductory tutorial to continuous risk management based on this course. The rational for continuous risk management and how it is incorporated into project management are discussed. The risk management structure of six functions is discussed in sufficient depth for managers to understand what is involved in risk management and how it is implemented. These functions include: (1) Identify the risks in a specific format; (2) Analyze the risk probability, impact/severity, and timeframe; (3) Plan the approach; (4) Track the risk through data compilation and analysis; (5) Control and monitor the risk; (6) Communicate and document the process and decisions.

Rosenberg, Linda

NASA's J-2X Engine Builds on the Apollo Program for Lunar Return Missions

In January 2006, NASA streamlined its U.S. Vision for Space Exploration hardware development approach for replacing the Space Shuttle after it is retired in 2010. The revised CLV upper stage will use the J-2X engine, a derivative of NASA s Apollo Program Saturn V s S-II and S-IVB main propulsion, which will also serve as the Earth Departure Stage (EDS) engine. This paper gives details of how the J- 2X engine effort mitigates risk by building on the Apollo Program and other lessons learned to deliver a human-rated engine that is on an aggressive development schedule, with first demonstration flight in 2010 and human test flights in 2012. It is well documented that propulsion is historically a high-risk area. NASA s risk reduction strategy for the J-2X engine design, development, test, and evaluation is to build upon heritage hardware and apply valuable experience gained from past development efforts. In addition, NASA and its industry partner, Rocketdyne, which originally built the J-2, have tapped into their extensive databases and are applying lessons conveyed firsthand by Apollo-era veterans of America s first round of Moon missions in the 1960s and 1970s. NASA s development approach for the J-2X engine includes early requirements definition and management; designing-in lessons learned from the 5-2 heritage programs; initiating long-lead procurement items before Preliminary Desi& Review; incorporating design features for anticipated EDS requirements; identifying facilities for sea-level and altitude testing; and starting ground support equipment and logistics planning at an early stage. Other risk reduction strategies include utilizing a proven gas generator cycle with recent development experience; utilizing existing turbomachinery ; applying current and recent main combustion chamber (Integrated Powerhead Demonstrator) and channel wall nozzle (COBRA) advances; and performing rigorous development, qualification, and certification testing of the engine system, with a philosophy of "test what you fly, and fly what you test". These and other active risk management strategies are in place to deliver the J-2X engine for LEO and lunar return missions as outlined in the U.S. Vision for Space Exploration.

Snoddy, Jimmy R.

A Vehicle Management End-to-End Testing and Analysis Platform for Validation of Mission and Fault Management Algorithms to Reduce Risk for NASAs Space Launch System

The engineering development of the National Aeronautics and Space Administration's (NASA) new Space Launch System (SLS) requires cross discipline teams with extensive knowledge of launch vehicle subsystems, information theory, and autonomous algorithms dealing with all operations from pre-launch through on orbit operations. The nominal and off-nominal characteristics of SLS's elements and subsystems must be understood and matched with the autonomous algorithm monitoring and mitigation capabilities for accurate control and response to abnormal conditions throughout all vehicle mission flight phases, including precipitating safing actions and crew aborts. This presents a large and complex systems engineering challenge, which is being addressed in part by focusing on the specific subsystems involved in the handling of off-nominal mission and fault tolerance with response management. Using traditional model-based system and software engineering design principles from the Unified Modeling Language (UML) and Systems Modeling Language (SysML), the Mission and Fault Management (M&FM) algorithms for the vehicle are crafted and vetted in Integrated Development Teams (IDTs) composed of multiple development disciplines such as Systems Engineering (SE), Flight Software (FSW), Safety and Mission Assurance (S&MA) and the major subsystems and vehicle elements such as Main Propulsion Systems (MPS), boosters, avionics, Guidance, Navigation, and Control (GNC), Thrust Vector Control (TVC), and liquid engines. These model-based algorithms and their development lifecycle from inception through FSW certification are an important focus of SLS's development effort to further ensure reliable detection and response to off-nominal vehicle states during all phases of vehicle operation from pre-launch through end of flight. To test and validate these M&FM algorithms a dedicated test-bed was developed for full Vehicle Management End-to-End Testing (VMET). For addressing fault management (FM) early in the development lifecycle for the SLS program, NASA formed the M&FM team as part of the Integrated Systems Health Management and Automation Branch under the Spacecraft Vehicle Systems Department at the Marshall Space Flight Center (MSFC). To support the development of the FM algorithms, the VMET developed by the M&FM team provides the ability to integrate the algorithms, perform test cases, and integrate vendor-supplied physics-based launch vehicle (LV) subsystem models. Additionally, the team has developed processes for implementing and validating the M&FM algorithms for concept validation and risk reduction. The flexibility of the VMET capabilities enables thorough testing of the M&FM algorithms by providing configurable suites of both nominal and off-nominal test cases to validate the developed algorithms utilizing actual subsystem models such as MPS, GNC, and others. One of the principal functions of VMET is to validate the M&FM algorithms and substantiate them with performance baselines for each of the target vehicle subsystems in an independent platform exterior to the flight software test and validation processes. In any software development process there is inherent risk in the interpretation and implementation of concepts from requirements and test cases into flight software compounded with potential human errors throughout the development and regression testing lifecycle. Risk reduction is addressed by the M&FM group but in particular by the Analysis Team working with other organizations such as S&MA, Structures and Environments, GNC, Orion, Crew Office, Flight Operations, and Ground Operations by assessing performance of the M&FM algorithms in terms of their ability to reduce Loss of Mission (LOM) and Loss of Crew (LOC) probabilities. In addition, through state machine and diagnostic modeling, analysis efforts investigate a broader suite of failure effects and associated detection and responses to be tested in VMET to ensure reliable failure detection, and confirm responses do not create additional risks or cause undesired states through interactive dynamic effects with other algorithms and systems. VMET further contributes to risk reduction by prototyping and exercising the M&FM algorithms early in their implementation and without any inherent hindrances such as meeting FSW processor scheduling constraints due to their target platform - the ARINC 6535-partitioned Operating System, resource limitations, and other factors related to integration with other subsystems not directly involved with M&FM such as telemetry packing and processing. The baseline plan for use of VMET encompasses testing the original M&FM algorithms coded in the same C++ language and state machine architectural concepts as that used by FSW. This enables the development of performance standards and test cases to characterize the M&FM algorithms and sets a benchmark from which to measure their effectiveness and performance in the exterior FSW development and test processes. This paper is outlined in a systematic fashion analogous to a lifecycle process flow for engineering development of algorithms into software and testing. Section I describes the NASA SLS M&FM context, presenting the current infrastructure, leading principles, methods, and participants. Section II defines the testing philosophy of the M&FM algorithms as related to VMET followed by section III, which presents the modeling methods of the algorithms to be tested and validated in VMET. Its details are then further presented in section IV followed by Section V presenting integration, test status, and state analysis. Finally, section VI addresses the summary and forward directions followed by the appendices presenting relevant information on terminology and documentation.

Trevino, Luis

Practical Application of PRA as an Integrated Design Tool for Space Systems

This paper presents the application of the first comprehensive Probabilistic Risk Assessment (PRA) during the design phase of a joint NASA/NOAA weather satellite program, Geostationary Operational Environmental Satellite Series R (GOES-R). GOES-R is the next generation weather satellite primarily to help understand the weather and help save human lives. PRA has been used at NASA for Human Space Flight for many years. PRA was initially adopted and implemented in the operational phase of manned space flight programs and more recently for the next generation human space systems. Since its first use at NASA, PRA has become recognized throughout the Agency as a method of assessing complex mission risks as part of an overall approach to assuring safety and mission success throughout project lifecycles. PRA is now included as a requirement during the design phase of both NASA next generation manned space vehicles as well as for high priority robotic missions. The influence of PRA on GOES-R design and operation concepts are discussed in detail. The GOES-R PRA is unique at NASA for its early implementation. It also represents a pioneering effort to integrate risks from both Spacecraft (SC) and Ground Segment (GS) to fully assess the probability of achieving mission objectives. PRA analysts were actively involved in system engineering and design engineering to ensure that a comprehensive set of technical risks were correctly identified and properly understood from a design and operations perspective. The analysis included an assessment of SC hardware and software, SC fault management system, GS hardware and software, common cause failures, human error, natural hazards, solar weather and infrastructure (such as network and telecommunications failures, fire). PRA findings directly resulted in design changes to reduce SC risk from micro-meteoroids. PRA results also led to design changes in several SC subsystems, e.g. propulsion, guidance, navigation and control (GNC), communications, mechanisms, and command and data handling (C&DH). The fault tree approach assisted in the development of the fault management system design. Human error analysis, which examined human response to failure, indicated areas where automation could reduce the overall probability of gaps in operation by half. In addition, the PRA brought to light many potential root causes of system disruptions, including earthquakes, inclement weather, solar storms, blackouts and other extreme conditions not considered in the typical reliability and availability analyses. Ultimately the PRA served to identify potential failures that, when mitigated, resulted in a more robust design, as well as to influence the program's concept of operations. The early and active integration of PRA with system and design engineering provided a well-managed approach for risk assessment that increased reliability and availability, optimized lifecyc1e costs, and unified the SC and GS developments.

Kalia, Prince

Advancing the practice of systems engineering at JPL

In FY 2004, JPL launched an initiative to improve the way it practices systems engineering. The Lab's senior management formed the Systems Engineering Advancement (SEA) Project in order to "significantly advance the practice and organizational capabilities of systems engineering at JPL on flight projects and ground support tasks." The scope of the SEA Project includes the systems engineering work performed in all three dimensions of a program, project, or task: 1. the full life-cycle, i.e., concept through end of operations 2. the full depth, i.e., Program, Project, System, Subsystem, Element (SE Levels 1 to 5) 3. the full technical scope, e.g., the flight, ground and launch systems, avionics, power, propulsion, telecommunications, thermal, etc. The initial focus of their efforts defined the following basic systems engineering functions at JPL: systems architecture, requirements management, interface definition, technical resource management, system design and analysis, system verification and validation, risk management, technical peer reviews, design process management and systems engineering task management, They also developed a list of highly valued personal behaviors of systems engineers, and are working to inculcate those behaviors into members of their systems engineering community. The SEA Project is developing products, services, and training to support managers and practitioners throughout the entire system lifecycle. As these are developed, each one needs to be systematically deployed. Hence, the SEA Project developed a deployment process that includes four aspects: infrastructure and operations, communication and outreach, education and training, and consulting support. In addition, the SEA Project has taken a proactive approach to organizational change management and customer relationship management - both concepts and approaches not usually invoked in an engineering environment. This paper'3 describes JPL's approach to advancing the practice of systems engineering at the Lab. It describes the general approach used and how they addressed the three key aspects of change: people, process and technology. It highlights a list of highly valued personal behaviors of systems engineers, discusses the various products, services and training that were developed, describes the deployment approach used, and concludes with several lessons learned.

process improvement

Risk management for the Space Exploration Initiative

Probabilistic Risk Assessment (PRA) is a quantitative engineering process that provides the analytic structure and decision-making framework for total programmatic risk management. Ideally, it is initiated in the conceptual design phase and used throughout the program life cycle. Although PRA was developed for assessment of safety, reliability, and availability risk, it has far greater application. Throughout the design phase, PRA can guide trade-off studies among system performance, safety, reliability, cost, and schedule. These studies are based on the assessment of the risk of meeting each parameter goal, with full consideration of the uncertainties. Quantitative trade-off studies are essential, but without full identification, propagation, and display of uncertainties, poor decisions may result. PRA also can focus attention on risk drivers in situations where risk is too high. For example, if safety risk is unacceptable, the PRA prioritizes the risk contributors to guide the use of resources for risk mitigation. PRA is used in the Space Exploration Initiative (SEI) Program. To meet the stringent requirements of the SEI mission, within strict budgetary constraints, the PRA structure supports informed and traceable decision-making. This paper briefly describes the SEI PRA process.

Buchbinder, Ben

Managing Fault Management Development

As the complexity of space missions grows, development of Fault Management (FM) capabilities is an increasingly common driver for significant cost overruns late in the development cycle. FM issues and the resulting cost overruns are rarely caused by a lack of technology, but rather by a lack of planning and emphasis by project management. A recent NASA FM Workshop brought together FM practitioners from a broad spectrum of institutions, mission types, and functional roles to identify the drivers underlying FM overruns and recommend solutions. They identified a number of areas in which increased program and project management focus can be used to control FM development cost growth. These include up-front planning for FM as a distinct engineering discipline; managing different, conflicting, and changing institutional goals and risk postures; ensuring the necessary resources for a disciplined, coordinated approach to end-to-end fault management engineering; and monitoring FM coordination across all mission systems.

McDougal, John M.

A Vehicle Management End-to-End Testing and Analysis Platform for Validation of Mission and Fault Management Algorithms to Reduce Risk for NASA's Space Launch System

The engineering development of the new Space Launch System (SLS) launch vehicle requires cross discipline teams with extensive knowledge of launch vehicle subsystems, information theory, and autonomous algorithms dealing with all operations from pre-launch through on orbit operations. The characteristics of these spacecraft systems must be matched with the autonomous algorithm monitoring and mitigation capabilities for accurate control and response to abnormal conditions throughout all vehicle mission flight phases, including precipitating safing actions and crew aborts. This presents a large and complex system engineering challenge, which is being addressed in part by focusing on the specific subsystems involved in the handling of off-nominal mission and fault tolerance with response management. Using traditional model based system and software engineering design principles from the Unified Modeling Language (UML) and Systems Modeling Language (SysML), the Mission and Fault Management (M&FM) algorithms for the vehicle are crafted and vetted in specialized Integrated Development Teams (IDTs) composed of multiple development disciplines such as Systems Engineering (SE), Flight Software (FSW), Safety and Mission Assurance (S&MA) and the major subsystems and vehicle elements such as Main Propulsion Systems (MPS), boosters, avionics, Guidance, Navigation, and Control (GNC), Thrust Vector Control (TVC), and liquid engines. These model based algorithms and their development lifecycle from inception through Flight Software certification are an important focus of this development effort to further insure reliable detection and response to off-nominal vehicle states during all phases of vehicle operation from pre-launch through end of flight. NASA formed a dedicated M&FM team for addressing fault management early in the development lifecycle for the SLS initiative. As part of the development of the M&FM capabilities, this team has developed a dedicated testbed that integrates specific M&FM algorithms, specialized nominal and off-nominal test cases, and vendor-supplied physics-based launch vehicle subsystem models. Additionally, the team has developed processes for implementing and validating these algorithms for concept validation and risk reduction for the SLS program. The flexibility of the Vehicle Management End-to-end Testbed (VMET) enables thorough testing of the M&FM algorithms by providing configurable suites of both nominal and off-nominal test cases to validate the developed algorithms utilizing actual subsystem models such as MPS. The intent of VMET is to validate the M&FM algorithms and substantiate them with performance baselines for each of the target vehicle subsystems in an independent platform exterior to the flight software development infrastructure and its related testing entities. In any software development process there is inherent risk in the interpretation and implementation of concepts into software through requirements and test cases into flight software compounded with potential human errors throughout the development lifecycle. Risk reduction is addressed by the M&FM analysis group working with other organizations such as S&MA, Structures and Environments, GNC, Orion, the Crew Office, Flight Operations, and Ground Operations by assessing performance of the M&FM algorithms in terms of their ability to reduce Loss of Mission and Loss of Crew probabilities. In addition, through state machine and diagnostic modeling, analysis efforts investigate a broader suite of failure effects and associated detection and responses that can be tested in VMET to ensure that failures can be detected, and confirm that responses do not create additional risks or cause undesired states through interactive dynamic effects with other algorithms and systems. VMET further contributes to risk reduction by prototyping and exercising the M&FM algorithms early in their implementation and without any inherent hindrances such as meeting FSW processor scheduling constraints due to their target platform - ARINC 653 partitioned OS, resource limitations, and other factors related to integration with other subsystems not directly involved with M&FM such as telemetry packing and processing. The baseline plan for use of VMET encompasses testing the original M&FM algorithms coded in the same C++ language and state machine architectural concepts as that used by Flight Software. This enables the development of performance standards and test cases to characterize the M&FM algorithms and sets a benchmark from which to measure the effectiveness of M&FM algorithms performance in the FSW development and test processes.

Trevino, Luis

NASA Spacecraft Fault Management Workshop Results

Fault Management is a critical aspect of deep-space missions. For the purposes of this paper, fault management is defined as the ability of a system to detect, isolate, and mitigate events that impact, or have the potential to impact, nominal mission operations. The fault management capabilities are commonly distributed across flight and ground subsystems, impacting hardware, software, and mission operations designs. The National Aeronautics and Space Administration (NASA) Discovery & New Frontiers (D&NF) Program Office at Marshall Space Flight Center (MSFC) recently studied cost overruns and schedule delays for 5 missions. The goal was to identify the underlying causes for the overruns and delays, and to develop practical mitigations to assist the D&NF projects in identifying potential risks and controlling the associated impacts to proposed mission costs and schedules. The study found that 4 out of the 5 missions studied had significant overruns due to underestimating the complexity and support requirements for fault management. As a result of this and other recent experiences, the NASA Science Mission Directorate (SMD) Planetary Science Division (PSD) commissioned a workshop to bring together invited participants across government, industry, academia to assess the state of the art in fault management practice and research, identify current and potential issues, and make recommendations for addressing these issues. The workshop was held in New Orleans in April of 2008. The workshop concluded that fault management is not being limited by technology, but rather by a lack of emphasis and discipline in both the engineering and programmatic dimensions. Some of the areas cited in the findings include different, conflicting, and changing institutional goals and risk postures; unclear ownership of end-to-end fault management engineering; inadequate understanding of the impact of mission-level requirements on fault management complexity; and practices, processes, and tools that have not kept pace with the increasing complexity of mission requirements and spacecraft systems. This paper summarizes the findings and recommendations from that workshop, as well as opportunities identified for future investment in tools, processes, and products to facilitate the development of space flight fault management capabilities.

Newhouse, Marilyn