Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Reliability and Mitigation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Software Architecture of Sensor Data Distribution In Planetary Exploration

Data from mobile and stationary sensors will be vital in planetary surface exploration. The distribution and collection of sensor data in an ad-hoc wireless network presents a challenge. Irregular terrain, mobile nodes, new associations with access points and repeaters with stronger signals as the network reconfigures to adapt to new conditions, signal fade and hardware failures can cause: a) Data errors; b) Out of sequence packets; c) Duplicate packets; and d) Drop out periods (when node is not connected). To mitigate the effects of these impairments, a robust and reliable software architecture must be implemented. This architecture must also be tolerant of communications outages. This paper describes such a robust and reliable software infrastructure that meets the challenges of a distributed ad hoc network in a difficult environment and presents the results of actual field experiments testing the principles and actual code developed.

Lee, Charles↗

REE radiation fault model: a tool for organizing and communication radiation test data and construction COTS based spacebourne computing systems

The growth in data rates of instruments on future NASA spacecraft continues to outstrip the improvement in communications bandwidth and processing capabilities of radiation-hardened computers. Sophisticated autonomous operations strategies will further increase the processing workload. Given the reductions in spacecraft size and available power, standard radiation hardened computing systems alone will not be able to address the requirements of future missions. The REE project was intended to overcome this obstacle by developing a COTS- based supercomputer suitable for use as a science and autonomy data processor in most space environments. This development required a detailed knowledge of system behavior in the presence of Single Event Effect (SEE) induced faults so that mitigation strategies could be designed to recover system level reliability while maintaining the COTS throughput advantage. The REE project has developed a suite of tools and a methodology for predicting SEU induced transient fault rates in a range of natural space environments from ground-based radiation testing of component parts. In this paper we provide an overview of this methodology and tool set with a concentration on the radiation fault model and its use in the REE system development methodology. Using test data reported elsewhere in this and other conferences, we predict upset rates for a particular COTS single board computer configuration in several space environments.

Radiation Effects Modeling COTS computers REE SEU ↗

Human Factors and Habitability Challenges for Mars Missions

As NASA is planning to send humans deeper into space than ever before, adequate crew health and performance will be critical for mission success. Within the NASA Human Research Program (HRP), the Space Human Factors and Habitability (SHFH) team is responsible for characterizing the risks associated with human capabilities and limitations with respect to long-duration spaceflight, and for providing mitigations (e.g., guidelines, technologies, and tools) to promote safe, reliable and productive missions. SHFH research includes three domains: Advanced Environmental Health (AEH), Advanced Food Technology (AFT), and Space Human Factors Engineering (SHFE). The AEH portfolio focuses on understanding the risk of microbial contamination of the spacecraft and on the development of standards for exposure to potential toxins such as chemicals, bacteria, fungus, and lunar/Martian dust. The two risks that the environmental health project focuses on are adverse health effects due to changes in host-microbe interactions, and risks associated with exposure to dust in planetary surface habitats. This portfolio also proposes countermeasures to these risks by making recommendations that relate to requirements for environmental quality, foods, and crew health on spacecraft and space missions. The AFT portfolio focuses on reducing the mass, volume, and waste of the entire integrated food system to be used in exploration missions, and investigating processing methods to extend the shelf life of food items up to five years, while assuring that exploration crews will have nutritious and palatable foods. The portfolio also delivers improvements in both the food itself and the technologies for storing and preparing it. SHFE sponsors research to establish human factors and habitability standards and guidelines in five risk areas, and provides improved design concepts for advanced crew interfaces and habitability systems. These risk areas include: Incompatible vehicle/habitat design, inadequate human-computer interaction, inadequate critical task design, inadequate human-automation/robotic interaction, and performance errors due to training deficiencies. To address the identified research gaps within each risk, SHFH's research plan includes studies in the laboratory, in analogs, and on International Space Station (ISS). In addition to establishing and maintaining the risk-based research portfolio, SHFH is also implementing a qualitative approach to determine how we at NASA evaluate human performance. Via interviews with experts, such as trainers, flight controllers, and flight surgeons, we are collecting the metrics by which they assess human performance, evidence of performance issues, and potential or actual consequences. The Human Performance Data Project will determine what human performance data have been collected in the past at NASA, and what data should be collected in the future in order to complete our knowledgebase and reduce risks related to human factors and habitability.

Whitmore, Mihriban↗

The role of AI in detecting and mitigating human errors in safety-critical industries: A review

For safety-critical industries, human error (HE) presents continual risks to system productivity, reliability and safety. Artificial intelligence (AI) and machine learning (ML) methods have emerged as promising approaches to understand, categorize and mitigate the risk of HE in safety-critical industries. Furthermore, this review offers an examination of the current landscape regarding the utilization of AI/ML with regards to HE in safety-critical industries, categorizing literature into descriptive modeling, predictive modeling, prescriptive modeling, and generative modeling techniques. Additionally, the review aims to provide insights regarding themes in literature, challenges, and future research directions. Findings of the review suggest that AI/ML methods can prove useful in addressing the HE problem across safety-critical industries.

42 ENGINEERING↗

Limitations of Reliability for Long-Endurance Human Spaceflight

Long-endurance human spaceflight - such as missions to Mars or its moons - will present a never-before-seen maintenance logistics challenge. Crews will be in space for longer and be farther way from Earth than ever before. Resupply and abort options will be heavily constrained, and will have timescales much longer than current and past experience. Spare parts and/or redundant systems will have to be included to reduce risk. However, the high cost of transportation means that this risk reduction must be achieved while also minimizing mass. The concept of increasing system and component reliability is commonly discussed as a means to reduce risk and mass by reducing the probability that components will fail during a mission. While increased reliability can reduce maintenance logistics mass requirements, the rate of mass reduction decreases over time. In addition, reliability growth requires increased test time and cost. This paper assesses trends in test time requirements, cost, and maintenance logistics mass savings as a function of increase in Mean Time Between Failures (MTBF) for some or all of the components in a system. In general, reliability growth results in superlinear growth in test time requirements, exponential growth in cost, and sublinear benefits (in terms of logistics mass saved). These trends indicate that it is unlikely that reliability growth alone will be a cost-effective approach to maintenance logistics mass reduction and risk mitigation for long-endurance missions. This paper discusses these trends as well as other options to reduce logistics mass such as direct reduction of part mass, commonality, or In-Space Manufacturing (ISM). Overall, it is likely that some combination of all available options - including reliability growth - will be required to reduce mass and mitigate risk for future deep space missions.

Owens, Andrew C.↗

An Approach to Automate tools for the Risk Assessment of Digital Instrumentation and Control Systems

Reliable digital instrumentation and control systems (DI&C) are integral for sustaining the continued operation of nuclear power plants. These systems ensure that nuclear reactors operate safely, efficiently, and within regulatory requirements. Yet, the cost of designing and licensing new nuclear DI&C can be prohibitively expensive. Under the U.S. Department of Energy Light Water Reactor Sustainability Program, Idaho National Laboratory has developed a framework for supporting the risk-informed design of DI&C systems by offering methods to support the identification, quantification, and evaluation of risks for various DI&C design architectures. The framework indicates potential software failure modes and provides pathways for quantifying the potential for these software failures, including common cause failures. Using the framework’s systematic approach, challenges for assessing risks within new and existing nuclear DI&C systems can be reduced. Nevertheless, the current framework can be further improved using the convenience of automation. This paper introduces the development of Software for the Hazard Identification and Evaluation of Digital Systems (SHIELDS). SHIELDS is an engineering software package that enables the identification, elimination, and mitigation of potential risks and reduces the burden of deploying reliable DI&C systems. This work introduces plans and techniques to digitize and improve the manual risk assessment modules of the framework. These improvements will save time and increase the repeatability and usability of the framework, making it more accessible to a wider range of users. Ultimately, this introduces SHIELDS and how its modules support efficient development of safe and reliable DI&C systems.

46 - INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AN↗

Long-Term Impacts of Constrained Transmission Deployment on the Cost-Reliability Tradeoff

Traditional Resource Adequacy (RA) frameworks in the U.S. undervalue the contributions of inter-regional transmission to resource adequacy during stress periods, focusing on the availability of nameplate capacity instead. However, availability of nameplate capacity does not always translate into electricity delivery, especially during tail events. Moreover, the rapid deployment of energy-limited resources and increasing electricity demand challenge existing resource adequacy frameworks and couple regional electricity demand and availability of supply via transmission. We propose a two-stage framework that goes beyond the existing capacity-centered approaches to reveal the RA contributions of transmission. In the first stage we introduce a multi-objective optimization framework to quantify the merits of transmission expansion via Pareto Frontiers under alternative futures of no transmission investment, primary energy resources availability and demand growth. The second stage focuses on tail events and leverages the results of the first stage to characterize the risk profile of regional consumers across the U.S. under the alternative energy futures. We find that no new transmission can lead to a more expensive and less reliable national grid across scenarios, however, the impact on regional RA can vary. The probabilistic analysis reveals that transmission investments can alleviate the tail risk of consumers, however, the availability of fuel resources does not always alleviate regional tail risks. Our findings inform policymakers and utilities on the prioritization of transmission investments to mitigate the risk of widespread outages, also for tail events, and ensure reliable and affordable electricity delivery to all.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Space Launch System (SLS) Safety, Mission Assurance, and Risk Mitigation

SLS Driving Objectives: I. Safe: a) Human-rated to provide safe and reliable systems for human missions. b) Protecting the public, NASA workforce, high-value equipment and property, and the environment from potential harm. II. Affordable: a) Maximum use of common elements and existing assets, infrastructure, and workforce. b) Constrained budget environment. c) Competitive opportunities for affordability on-ramps. III. Sustainable: a) Initial capability: 70 metric tons (t), 2017-2021. 1) Serves as primary transportation for Orion and exploration missions. 2) Provides back-up capability for crew/cargo to ISS. b) Evolved capability: 105 t and 130 t, post-2021. 1) Offers large volume for science missions and payloads. 2) Modular and flexible, right-sized for mission requirements.

May, Todd↗

Risk-controlled Expansion Planning with Distributed Resources (REPAIR) v1.0

The Risk-controlled Expansion Planning with Distributed Resources (REPAIR) is an innovative tool to support decisions around utility grid planning to prevent and mitigate the impact of outages caused by routine equipment failures (reliability) or by extreme events (resilience), such as storms, earthquakes or wildfires that long term interruption of service. REPAIR is a risk-based optimization and decision-making model allowing informed and transparent "cost vs risk" decisions regarding infrastructural planning of electric utilities. The model considers long-term resilience and reliability planning strategies that rely on traditional infrastructure upgrade (e.g. circuit hardening, reinforcement, new substations, etc.) or new investment alternatives, such as DERs.

Heleno, Miguel [Lawrence Berkeley National Laborat↗

System concepts for a large UV/optical/IR telescope on the moon

To assess the systems and technological requirements for constructing lunar telescopes in conjunction with the buildup of a lunar base for scientific exploration and as a waypoint for travel to Mars, the NASA Marshall Space Flight Center conducted concept studies of a 16-m-aperture large lunar telescope (LLT) and a 4-m-aperture precursor telescope, both operating in the UV/visible/IR spectral region. The feasibility of constructing a large telescope on the lunar surface is assessed, and its systems and subsystems are analyzed. Telescope site selection, environmental effects, and launch and assembly scenarios are also evaluated. It is argued that key technical drivers for the LLT must be tested in situ by precursor telescopes to evaluate such areas as the operations and long-term reliability of active optics, radiation protection of instruments, lunar dust mitigation, and thermal shielding of the telescope systems. For a manned lunar outpost or an LLT to become a reality, a low-cost dependable transportation system must be developed.

Nein, Max E.↗

An Impact Sensor System for the Characterization of the Micrometeoroid and Lunar Secondary Ejecta Environment

The Impact Sensor for Micrometeoroid and Lunar Secondary Ejecta (IMMUSE) project aims to apply and integrate previously demonstrated impact sensing subsystems to characterize the micrometeoroid and lunar secondary (MMSE) environment on the surface of the Moon. Once deployed, data returned from IMMUSE will benefit: (1) Fundamental Lunar Science: providing data to improve the understanding of lunar cratering processes and dynamics of the lunar regolith. (2) Lunar Exploration Applied Science: providing an accurate MMSE environment definition for reliable impact risk assessments, cost-effective shielding designs, and mitigation measures for long-term lunar exploration activities. (3) Planetary Science: providing micrometeoroid data to aid the understanding of asteroidal collisions and the evolution of comets. A well-established link between micrometeoroid impacts and lunar regolith is also key to understanding other regolith-covered bodies from remote-sensing data. The IMMUSE system includes two components: (1) a large area (greater than or equal to 1 m2) micrometeoroid detector based on acoustic impact and fiber optic displacement sensors and (2) a 100 cm2 lunar secondary ejecta detector consisting of dual-layer laser curtain and acoustic impact sensors. The combinations of different detection mechanisms will allow for a better characterization of the MMSE environment, including flux, particle size/mass, and impact velocity. IMMUSE is funded by the NASA LASER Program through 2012. The project fs goal is to reach a Technical Readiness Level of 4 in preparation for a more advanced development beyond 2012. Several prototype subsystems have been constructed and subjected to low impact and hypervelocity impact tests. The presentation will include a status review and preliminary test results.

Liou, J.-C.↗

Supporting Technology at GRC to Mitigate Risk as Stirling Power Conversion Transitions to Flight

Stirling power conversion technology has been reaching more advanced levels of maturity during its development for space power applications. The current effort is in support of the Advanced Stirling Radioisotope Generator (ASRG), which is being developed by the U.S. Department of Energy (DOE), Lockheed Martin Space Systems Company (LMSSC), Sunpower Inc., and the NASA Glenn Research Center (GRC). This generator would use two high-efficiency Advanced Stirling Convertors (ASCs) to convert thermal energy from a radioisotope heat source into electricity. Of paramount importance is the reliability of the power system and as a part of this, the Stirling power convertors. GRC has established a supporting technology effort with tasks in the areas of reliability, convertor testing, high-temperature materials, structures, advanced analysis, organics, and permanent magnets. The project utilizes the matrix system at GRC to make use of resident experts in each of the aforementioned fields. Each task is intended to reduce risk and enhance reliability of the convertor as this technology transitions toward flight status. This paper will provide an overview of each task, outline the recent efforts and accomplishments, and show how they mitigate risk and impact the reliability of the ASC s and ultimately, the ASRG.

Schreiber, Jeffrey G.↗

Supporting Technology at GRC to Mitigate Risk as Stirling Power Conversion Transitions to Flight

Stirling power conversion technology has been reaching more advanced levels of maturity during its development for space power applications. The current effort is in support of the Advanced Stirling Radioisotope Generator (ASRG), which is being developed by the U.S. Department of Energy (DOE), Lockheed Martin Space Systems Company (LMSSC), Sunpower Inc., and the NASA Glenn Research Center (GRC). This generator would use two high-efficiency Advanced Stirling Convertors (ASCs) to convert thermal energy from a radioisotope heat source into electricity. Of paramount importance is the reliability of the power system and as a part of this, the Stirling power convertors. GRC has established a supporting technology effort with tasks in the areas of reliability, convertor testing, high-temperature materials, structures, advanced analysis, organics, and permanent magnets. The project utilizes the matrix system at GRC to make use of resident experts in each of the aforementioned fields. Each task is intended to reduce risk and enhance reliability of the convertor as this technology transitions toward flight status. This paper will provide an overview of each task, outline the recent efforts and accomplishments, and show how they mitigate risk and impact the reliability of the ASC s and ultimately, the ASRG.

Schreiber, Jeffrey G.↗

FPGA Mitigation Strategies for Critical Applications

Technology is changing at a fast pace. Transistor geometries are getting smaller, voltage thresholds are getting lower, design complexity is exponentially increasing, and user options are expanding. Consequently, reliable insertion of error detection and correction (EDAC) circuitry has become relatively challenging. As a response, a variety of mitigation techniques are being evaluated. They range from weak EDAC circuits that save area and power to strong mitigation strategies that are a great expense to systems. This presentation will focus on radiation induced susceptibilities for a variety of FPGA types and ASIC devices. In addition, the user will be provided information on applicable mitigation strategies per device.

Berg, Melanie↗

Practical Application of PRA as an Integrated Design Tool for Space Systems

This paper presents the application of the first comprehensive Probabilistic Risk Assessment (PRA) during the design phase of a joint NASA/NOAA weather satellite program, Geostationary Operational Environmental Satellite Series R (GOES-R). GOES-R is the next generation weather satellite primarily to help understand the weather and help save human lives. PRA has been used at NASA for Human Space Flight for many years. PRA was initially adopted and implemented in the operational phase of manned space flight programs and more recently for the next generation human space systems. Since its first use at NASA, PRA has become recognized throughout the Agency as a method of assessing complex mission risks as part of an overall approach to assuring safety and mission success throughout project lifecycles. PRA is now included as a requirement during the design phase of both NASA next generation manned space vehicles as well as for high priority robotic missions. The influence of PRA on GOES-R design and operation concepts are discussed in detail. The GOES-R PRA is unique at NASA for its early implementation. It also represents a pioneering effort to integrate risks from both Spacecraft (SC) and Ground Segment (GS) to fully assess the probability of achieving mission objectives. PRA analysts were actively involved in system engineering and design engineering to ensure that a comprehensive set of technical risks were correctly identified and properly understood from a design and operations perspective. The analysis included an assessment of SC hardware and software, SC fault management system, GS hardware and software, common cause failures, human error, natural hazards, solar weather and infrastructure (such as network and telecommunications failures, fire). PRA findings directly resulted in design changes to reduce SC risk from micro-meteoroids. PRA results also led to design changes in several SC subsystems, e.g. propulsion, guidance, navigation and control (GNC), communications, mechanisms, and command and data handling (C&DH). The fault tree approach assisted in the development of the fault management system design. Human error analysis, which examined human response to failure, indicated areas where automation could reduce the overall probability of gaps in operation by half. In addition, the PRA brought to light many potential root causes of system disruptions, including earthquakes, inclement weather, solar storms, blackouts and other extreme conditions not considered in the typical reliability and availability analyses. Ultimately the PRA served to identify potential failures that, when mitigated, resulted in a more robust design, as well as to influence the program's concept of operations. The early and active integration of PRA with system and design engineering provided a well-managed approach for risk assessment that increased reliability and availability, optimized lifecyc1e costs, and unified the SC and GS developments.

Kalia, Prince↗

Fault Tree Analysis Application for Safety and Reliability

Many commercial software tools exist for fault tree analysis (FTA), an accepted method for mitigating risk in systems. The method embedded in the tools identifies a root as use in system components, but when software is identified as a root cause, it does not build trees into the software component. No commercial software tools have been built specifically for development and analysis of software fault trees. Research indicates that the methods of FTA could be applied to software, but the method is not practical without automated tool support. With appropriate automated tool support, software fault tree analysis (SFTA) may be a practical technique for identifying the underlying cause of software faults that may lead to critical system failures. We strive to demonstrate that existing commercial tools for FTA can be adapted for use with SFTA, and that applied to a safety-critical system, SFTA can be used to identify serious potential problems long before integrator and system testing.

Wallace, Dolores R.↗

Quantification of Uncertainty and Risk Sensitivity for Safety of Emerging Operations

The growing need to develop and deploy small unmanned aerial vehicles (sUAVs) for various applications in the airspace necessitates reliable tools to accurately predict the flight trajectories of the sUAVs. The knowledge of the predicted trajectories help decision makers anticipate potential conflict, assess the risk, and take appropriate risk mitigation actions. In addition, uncertainties in vehicle models, weather, and controller action further highlights the need for reliable prediction tools. In this project, the application of mixed sparse grid-based quadrature and generalized polynomial chaos(gPC) expansion method for uncertainty quantification and collision assessment in air traffic consisting of fixed-wing small unmanned aerial vehicles (sUAV) was studied. From the results obtained, it can be concluded that this provides a reliable framework to carry out quantitative conflict assessment in an unmanned air traffic, which when employed, can improve the functionalities of the unmanned traffic management system. It was observed that the results from the gPC expansion framework developed in the project can be utilized to conduct rapid probabilistic collision assessment for near real-time unmanned traffic management in the airspace. From the vehicle models, position updates, and wind-field data, a priori gPC based 3-σcon-fidence ellipses can provide estimates of potential conflict at some future instants. The computational costs scaled linearly when the uncertain inputs were fewer. Further, the largest allowable distribution of para-metric uncertainties that leads to the smallest risk of collision in traffic of small unmanned aerial vehicles could be calculated. The time of closest approach between two sUAVs can be established paving way for development of proactive mitigation strategies. The separation between the sUAVs was found to be most significantly affected by uncertainties in the maximum available thrusts, zero-lift drag coefficients, and wing planform areas of the sUAVs. The study of uncertain wind-fields indicated that a heterogeneous traffic mix resulted in an increased probability of conflict. Increased measurement update rate reduced the uncertain-ties in the trajectories of the vehicles, further reducing the probability of conflict but rapid updates of all vehicles in the airspace poses a stringent communication limitation. The gPC framework also provided the means to analyze vehicle impact (crash region) due to loss of control resulting from actuator failure in sUAS traffic, essentially to predict impact and crash zones for representative vehicles. The predicted regions when compared with non-participant density, provides a means to develop an early mitigation strategy, should the sUAV detect an imminent actuator failure.

Rajnish Bhusal↗