Engineering PapersSearch

Engineering topics

Alonso Vera

Publications and source records attributed to Alonso Vera.

NASA Exploration Systems Maintainability Standards for Artemis and Beyond

This assessment was requested by the NASA Engineering and Safety Center (NESC), which, based on findings from the NESC study “Safe Human Expeditions Beyond Low Earth Orbit (LEO)” (Valinia et al., 2022), determined its topic to be an underrecognized critical and urgent Agency need, due to impending Artemis vehicle procurements. The principal objective of this assessment was to review and update current Agency-level maintainability requirements for space systems to support crew on expeditions beyond LEO in both preventive and corrective maintenance. This report contains the results of the NESC assessment.

Orbital Replacement Units

Risk Trade-Space Analysis for Safe Human Expeditions to Mars

We assessed the integrated safety, health, and performance risk to crews on long-duration missions, specifically to Mars. Using a systems approach rather than one focused on individual countermeasures, we examined the trade space around several such risks to identify high-potential risk mitigation strategies and characterize aspects of Mars mission architectures that could lower aggregated risk. Current Mars Design Reference missions would require durations well over two years and would increase crew exposure to radiation and microgravity well beyond ISS levels, likely resulting in significantly reduced performance beyond our current capability to mitigate that could jeopardize mission success. A “fast Mars transit” round-trip mission concept was studied using an innovative flight dynamics approach to quantify the minimum total mission energy required for a Mars transit with total mission duration less than 400 days. This approach holds promise for sending humans to Mars and returning them safely with acceptable, potentially mitigatable, exposure to microgravity and radiation using current or near-term technologies. The fast transit concept would also result in fewer time-driven vehicle failures and enable sustainable deployment of humans and infrastructure to Mars on a regular cadence, allowing steady exploration and colonization of Mars. Finally, we conclude that reliance on the Low Earth Orbit (LEO) mission operations paradigm – i.e., one of near-complete real-time dependence on experts at Mission Control to manage the combined state of the mission, vehicle, and crew – is high risk given the communication delays and limited resupply of any Mars mission, and this risk is not eliminated by the shorter missions durations of fast transit scenarios. Based on historical trends, it is highly likely that the crew will face a high-consequence problem of uncertain origin during Mars transit when ground support will be greatly reduced. While it may be possible to reduce anomaly rates through improved reliability analysis and testing, and to reduce anomaly impacts through added robustness, such mitigations address only known failure modes and known uncertainties. Therefore, a radical shift in the Human-Systems Integration Architecture (HSIA) that defines the operational paradigm, systems design, and human-systems interactions is required to improve the risk posture to an acceptable level regardless of mission duration.

Mars

Automating CPM-GOMS

CPM-GOMS is a modeling method that combines the task decomposition of a GOMS analysis with a model of human resource usage at the level of cognitive, perceptual, and motor operations. CPM-GOMS models have made accurate predictions about skilled user behavior in routine tasks, but developing such models is tedious and error-prone. We describe a process for automatically generating CPM-GOMS models from a hierarchical task decomposition expressed in a cognitive modeling tool called Apex. Resource scheduling in Apex automates the difficult task of interleaving the cognitive, perceptual, and motor resources underlying common task operators (e.g. mouse move-and-click). Apex's UI automatically generates PERT charts, which allow modelers to visualize a model's complex parallel behavior. Because interleaving and visualization is now automated, it is feasible to construct arbitrarily long sequences of behavior. To demonstrate the process, we present a model of automated teller interactions in Apex and discuss implications for user modeling. available to model human users, the Goals, Operators, Methods, and Selection (GOMS) method [6, 21] has been the most widely used, providing accurate, often zero-parameter, predictions of the routine performance of skilled users in a wide range of procedural tasks [6, 13, 15, 27, 28]. GOMS is meant to model routine behavior. The user is assumed to have methods that apply sequences of operators and to achieve a goal. Selection rules are applied when there is more than one method to achieve a goal. Many routine tasks lend themselves well to such decomposition. Decomposition produces a representation of the task as a set of nested goal states that include an initial state and a final state. The iterative decomposition into goals and nested subgoals can terminate in primitives of any desired granularity, the choice of level of detail dependent on the predictions required. Although GOMS has proven useful in HCI, tools to support the construction of GOMS models have not yet come into general use.

GOMS

NASA’s Identified Risks of Adverse Outcomes Due to Inadequate Human Systems Integration Architecture in Human Spaceflight

The NASA Human System Risk Board (HSRB) has the overall responsibility for tracking the evolution of the top ~30 human system risks that it has identified to be associated with human spaceflight. As part of this process, the Board is charged with maintaining a consistent, integrated process to mitigate those risks, and developing evidence-based risk posture recommendations. One of the identified risks is due to inadequate human systems integration architecture (HSIA) and a driving factor of this risk is that given decreasing real-time ground support for execution of complex operations during future exploration missions, there is a possibility of adverse performance outcomes including that crew are unable to adequately respond to unanticipated critical malfunctions or detect safety critical procedural errors. The HSRB uses Directed Acyclic Graphs (DAGs) as a communication tool for describing how astronaut exposure to spaceflight hazards leads to meaningful mission-level health and performance outcomes and as the basis for understanding intermediate causal relationships between risk contributing factors and countermeasures that link hazards to outcomes. The HSIA risk DAG will be presented and described. Historically, critical malfunctions requiring Crew/MCC management occurred at a rate of 1.7 times per year for ISS averaged over the lifetime and 3-4 times per year in the burn in phase for the vehicle. These averages do not include EVA data, which greatly increases the incident rate. Prior experience from the Apollo program showed 10/11 crewed missions experienced significant anomalies where crew relied heavily on MCC expertise in real-time. These failure patterns are in line with those observed in other complex engineered systems (e.g., oil rigs, launch systems, commercial aviation, etc.) It is likely that general malfunction and error rates are > 10% for short duration missions (<30 days), based on past and current spaceflight operations data. Likelihood of adverse outcomes has the potential to increase as crew conduct work with new, complex systems and with less ground support. For Low Earth Orbit missions and Lunar missions less than 30 days, assuming minimal comm delays, disruptions and bandwidth limitations, malfunctions and errors can affect mission objectives and crew health but may be mitigated by ground support. For Lunar missions greater than 30 days and any potential Mars mission malfunctions and errors can have Loss of Crew and Loss of Mission consequences due to reduced ground support (communication delays, constraints and blackouts) for more complex operations, as well as reduced resupply and evacuation options.

Daniel M Buckland

NASA's Identified Risk of Adverse Outcomes due to Inadequate Human Systems Integration Architecture

The NASA Human System Risk Board (HSRB) is responsible for tracking the evolution of the top ~30 human system risks identified to be associated with human spaceflight. As part of this process, the Board is charged with maintaining a consistent, integrated process to evaluate those risks and developing evidence-based risk posture recommendations. Risks are ranked by likelihood and consequence. Intermediate causal relationships between risk contributing factors and countermeasures that link hazards to outcomes are described using Directed Acyclic Graphs (DAGs). The DAGs are also useful for identifying common factors and countermeasures across the top 30 risks as well as communicating how astronaut exposure to spaceflight hazards leads to meaningful mission-level health and performance outcomes. One of the top risks tracked by the HSRB is The Risk of Adverse Outcomes Due to Inadequate Human-Systems Integration Architecture (HSIA). This risk captures the possibility that due to decreasing real-time ground support during missions beyond LEO, crew will be unable to adequately respond to unanticipated critical malfunctions or detect safety-critical procedural errors. The HSIA risk is ranked red (high) for Lunar surface and Mars missions due to the probability of Loss of Crew and Loss of Mission consequences. This paper describes the evidence that supports the HSIA risk ranking and presents the central narrative of the HSIA risk DAG-- i.e., anomaly detection, diagnosis, intervention, and task performance. Characterizations of the current state of practice for each of the DAG’s central nodes and the future tools needed for successful anomaly response are provided.

human-systems integration architecture

Risks from Decreasing Ground Support

This chart set describes data and analysis underlying the risk of reduced ground support for long duration, deep space missions, starting with Lunar surface stays and progressing to a crewed mission to Mars. The functional impacts for different comm delay regimes are described, along with a specific and generalized anomaly scenario to demonstrate the impact of reduced ground support to onboard operations.

comm delay

Effects of Communication Delay on Human Spaceflight Missions

Missions onboard the International Space Station rely on the real-time availability of a large ground team of system experts to command the vehicle, solve safety-critical problems, and guide the crew during complex operations. Also, in Low Earth Orbit (LEO), supplies can be sent and crews evacuated quite quickly if needed. Future missions Beyond Low Earth Orbit (BLEO) will not have this 24/7, real-time safety net as communication latency increases, resupply difficulty increases, and evacuation opportunities diminish. There are few, if any, terrestrial analogs for human spaceflight missions BLEO that reflect the conditions—including extreme environments, long mission durations, and small crew sizes – that make these missions so high risk. Studies on specific conditions, such as communication delays and asynchronous interactions, have been performed in NASA Earth-based analog missions and have found that communication delays can disrupt ground-crew interactions and adversely impact team performance. However, there are gaps and limitations in studies conducted to date, notably on human spacecraft system failure response and recovery, the impacts of shorter lunar-relevant communication delays on complex operations, and the effectiveness of countermeasures. The work presented here breaks down real anomalies that occurred on ISS and Apollo missions and creates example scenarios fort Lunar Surface and Mars missions to explore the impact of communication delays of varying length on onboard operations and mission outcomes. Our analyses indicate that short communication delays (e.g., seconds to a minute) adversely impact the ability for ground to provide real-time oversight and guidance and to catch quickly emerging problems in time. Longer communication delays (e.g., up to 40 minutes on Mars missions) call for a shift of responsibility for tactical operations from ground to crew; crew must make time-critical decisions independently and respond to time-critical vehicle anomalies to prevent consequences.

human-systems integration

The Impact of Delayed Communication on NASA’s Human-Systems Operations: Preliminary Results of a Systematic Review

Throughout the history of human spaceflight, NASA has relied on a team of ground-based experts on Earth to manage its missions, vehicles, and crews to ensure crew safety and mission success. However, as missions progress beyond low-Earth orbit (LEO), this paradigm of dependence on ground must evolve. Beyond LEO, in missions to the moon and Mars, crews will confront new challenges: limited evacuation options, reduced resupply capabilities, and significant communication delays that impede real-time support from experts on the ground. This reduction in ground support amplifies the likelihood that crews will be unable to adequately respond to unanticipated, safety-critical events. Understanding the scope of these risks and identifying effective countermeasures hinges on understanding the impact of communication delays on complex operations, especially in urgent, unforeseen events. Real-time communication currently provides the crew with continuous access to a large, extensively resourced ground team skilled in anomaly resolution. However, as communication delays grow, the need to transfer some responsibilities from ground experts to onboard crew becomes evident. NASA has been exploring this shift in operational responsibilities and its effectiveness in managing complex operations for decades. Nevertheless, a comprehensive understanding of the specific challenges posed by communication delays and the necessary countermeasures to mitigate them remains a gap. In this paper, we present an update on our systematic review of the literature on communication delays, the first in-depth review since 2013 (Rader et al.). We introduce a coding taxonomy to capture key constructs from papers of interest and discuss preliminary findings. These preliminary results suggest two significant research gaps: limited studies have been conducted 1) with lunar-like latencies and 2) on problem-solving strategies for the maximum latencies expected in Mars missions. We outline plans and propose recommendations to address these gaps through ongoing and future research.

human-systems integration

Information Systems for Crew-Led Operations Beyond Low-Earth Orbit

On past and present human space missions, the management of vehicle health and status has primarily been executed from Earth. Missions such as Apollo, Space Shuttle, and ISS have relied on a safety net of ground-based experts with access to real-time telemetry data, broad and deep systems expertise, and powerful analytical and computing capabilities. The ground team monitors and manages the vehicle’s health in real-time and responds quickly to critical situations and malfunctions. Ground operators also provide real-time oversight and verbal guidance to flight crew members, especially during complex procedure execution and high-risk activities like extra-vehicular activities. However, this operational paradigm, in place for 60 years, will not transfer to long duration exploration missions beyond low Earth orbit (LEO). Lunar and deep-space crewed missions will encounter delayed communications that prohibit real-time operational and medical support. Additionally, there will be infrequent resupply and a diminished capacity to evacuate or rescue crew members. A small crew must operate independently, managing the vehicle’s state, responding to time-critical events, and executing complex procedures, all without the safety net of real-time support.

data representation

Earth Independent Operations Development for NASA’s Mars Campaign Office

NASA’s Moon to Mars Objectives lay out a strategic vision for the future of human spaceflight that culminates in crewed missions to the surface of Mars. The Mars Campaign Office (MCO) of the Moon to Mars Program is responsible for developing the technologies and capabilities necessary to support Mars missions. One such set of technologies supports a need for increased independence from Earth driven primarily by a communications delay between Earth and Mars that can reach twenty minutes or more. This paper lays out the initial strategy and approach for the MCO Earth-Independent Operations (EIO) portfolio and identifies areas targeted for investment between now and Mars missions tentatively planned for the early 2040s.

Ian D Maddox

Content and Representation of Information Needed to Support Time-Constrained Problem Solving

NASA’s current mission-operations paradigm originated with Project Mercury and endured with minimum evolution through the Apollo Program, Space Shuttle Program, and ISS missions. At its foundation is a near-complete real-time dependence on a ground team to manage the combined state of the mission, vehicle, and crew. Utilizing many engineers and operators with broad and deep expertise; large, distributed datasets including extensive telemetry; and expansive analytical and computing power, this ground team has served as the safety net for crewed spaceflight missions over the past 60 years. This approach must change to address challenges associated with missions beyond low Earth orbit (BLEO), including infrequent resupply, reduced ability to evacuate, and delayed communications that prohibit real-time operational support. We anticipate that a necessary part of this change will be increased independence for the crew, as roles and responsibilities traditionally performed by ground teams move on board the vehicle. While many risks are associated with Earth-independent operations, one particular concern is ensuring that the crew will have adequate onboard support to perform urgent problem solving when communication with the ground is delayed or intermittent. A key resource that enables the ground team to respond to anomalies quickly and effectively is the extraordinary expertise and experience it possesses. It is comprised of 80+ experts on at any given time, with a combined 600+ years of system-specific experience across 22 unique console disciplines. A small crew will face the unprecedented challenge of independently responding to anomalies that have historically been handled by a team 20 times their size. Another important resource upon which the ground heavily relies to support procedure execution and anomaly response is data. The amount of telemetry data that each flight controller monitors is extensive. In addition, as the ground team works to further assess impacts, trouble shoot, identify workarounds, and oversee procedure execution, it accesses and synthesizes engineering and procedure information, as well as system build, test, and configuration documentation. It is not feasible nor useful to put all these data onboard as crews become more Earth independent. Each member of a small Mars mission small crew will have multiple roles beyond monitoring telemetry and data gathering, and multiple roles within anomaly resolution processes, thereby limiting their capacity for copious amounts of information. Moreover, while access is necessary, it alone is insufficient. Information will need to be compiled, refined, and represented appropriately to support the crew’s reduced attention and expertise. This work seeks to understand the content and representation of information needed to support time-constrained problem solving and decision making by the crew without real-time ground support. To build this understanding, we first surveyed the literature, focusing on how expert problem solvers construct and manipulate their mental models. Next, we interviewed expert problem solvers in spaceflight and analogous domains and surveyed industry solutions for data presentation. Finally, we analyzed current spaceflight operations by investigating flight controller anomaly resolution processes during ISS training simulations and real operational events. These methods led to creating a problem-solving framework that details common themes and features of attending to, assessing, analyzing, and acting on problems in complex, time-constrained domains. Using this framework and the results of our analysis, we identified conceptual data representations needed for crew-led problem-solving. Preliminary onboard user interface concepts to meet identified needs will be presented.

anomaly response