Engineering PapersSearch

SEARCH · Engineering Papers

Results for “outage management system”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Controller-Hardware-in-the-Loop Evaluation of a Microgrid Controller for a Microgrid System With Multiple Grid-Forming Inverters

This paper presents the laboratory evaluation of a commercial Microgrid Management System (MGMS) implemented in the real-world Bronzeville Microgrid which features a futuristic scenario with high renewable energy integration and the use of multiple Grid-Forming (GFM) inverters. The primary objective of the performance evaluation for the MGMS is to assess the MGMS's capability to dispatch GFM units, including a GFM PV unit and two GFM battery units, to maintain the system stability and ensure economic operation, thus guaranteeing the microgrid's resilience during prolonged outages and dynamic events. The laboratory controller hardware-in-the-loop provides realistic testing environment through detailed electromagnetic transient modeling of the microgrid system, hardware MGMS, and standard communication protocols (DNP3). This CHIL evaluation shows how the MGMS effectively manages the GFM inverters, highlighting its performance in maintaining stability, reliability, and survivability in a microgrid environment with a high penetration of renewable energy sources.

controller hardware-in-the-loop

Discrepancy Reporting Management System

Discrepancy Reporting Management System (DRMS) is a computer program designed for use in the stations of NASA's Deep Space Network (DSN) to help establish the operational history of equipment items; acquire data on the quality of service provided to DSN customers; enable measurement of service performance; provide early insight into the need to improve processes, procedures, and interfaces; and enable the tracing of a data outage to a change in software or hardware. DRMS is a Web-based software system designed to include a distributed database and replication feature to achieve location-specific autonomy while maintaining a consistent high quality of data. DRMS incorporates commercial Web and database software. DRMS collects, processes, replicates, communicates, and manages information on spacecraft data discrepancies, equipment resets, and physical equipment status, and maintains an internal station log. All discrepancy reports (DRs), Master discrepancy reports (MDRs), and Reset data are replicated to a master server at NASA's Jet Propulsion Laboratory; Master DR data are replicated to all the DSN sites; and Station Logs are internal to each of the DSN sites and are not replicated. Data are validated according to several logical mathematical criteria. Queries can be performed on any combination of data.

Cooper, Tonja M.

Enhanced Communication Network Solution for Positive Train Control Implementation

The commuter and freight railroad industry is required to implement Positive Train Control (PTC) by 2015 (2012 for Metrolink), a challenging network communications problem. This paper will discuss present technologies developed by the National Aeronautics and Space Administration (NASA) to overcome comparable communication challenges encountered in deep space mission operations. PTC will be based on a new cellular wireless packet Internet Protocol (IP) network. However, ensuring reliability in such a network is difficult due to the "dead zones" and transient disruptions we commonly experience when we lose calls in commercial cellular networks. These disruptions make it difficult to meet PTC s stringent reliability (99.999%) and safety requirements, deployment deadlines, and budget. This paper proposes innovative solutions based on space-proven technologies that would help meet these challenges: (1) Delay Tolerant Networking (DTN) technology, designed for use in resource-constrained, embedded systems and currently in use on the International Space Station, enables reliable communication over networks in which timely data acknowledgments might not be possible due to transient link outages. (2) Policy-Based Management (PBM) provides dynamic management capabilities, allowing vital data to be exchanged selectively (with priority) by utilizing alternative communication resources. The resulting network may help railroads implement PTC faster, cheaper, and more reliably.

Policy-Based Management (PBM)

The power reliability event simulator tool (PRESTO): A novel approach to distribution system reliability analysis and applications

The growing interest in onsite solar photovoltaic and energy storage systems is partially motivated by customer concerns regarding grid reliability. However, accurately assessing the effectiveness of PVESS in mitigating these interruptions requires a comprehensive understanding of location-specific outage patterns and the ability to simulate realistic scenarios. To address the gap, we introduce the Power Reliability Event Simulation TOol (PRESTO), the first publicly available tool that simulates location-specific power interruptions at the county level. PRESTO allows for a more realistic assessment of system reliability by considering the unpredictability and location-specific patterns of power interruptions. We applied PRESTO in a case study of a single-family home across three U.S. counties, examining the performance of a solar photovoltaic system with 10kWh of battery storage during short-duration power interruptions. Our findings show that this system reliably met 93% of energy demand for essential non-heating and cooling loads, fully serving these loads in 84% of events, despite the constraints of daily time-of-use bill management which limits the battery's state-of-charge reserve. However, when heating and cooling loads were included, system performance decreased significantly, with only 70% of demand met and full service in 43% of events. These results highlight the challenges of using solar photovoltaic and energy storage systems for short-duration outages, emphasizing the need to consider factors like battery size and grid charging strategies to improve reliability. Our study demonstrates the practical applications of PRESTO, providing valuable insights into potential mitigation strategies including grid charging and optimizing battery size.

14 SOLAR ENERGY

Dynamic Boundary Microgrids Under Privatization Considerations

Microgrids have physical, electrical, and logical (data, network, and ownership) boundaries. To power unserved customer loads during an outage, microgrids can extend the traditional operational boundaries. This can become complex when considering microgrid-to-microgrid (M2M) interactions where sensitive information such as competitive microgrid operational data is not shared. This work proposes an optimization method coordinated between microgrid controllers and distribution management systems that limits data sharing. The method involves a competitive bidding strategy that maximizes unserved load coverage while minimizing resource utilization and sensitive operational data sharing among entities. The work is validated on a two-microgrid system with photovoltaic and energy storage systems and curves of load derived from real world residential buildings datasets. Results show that the proposed method, when applied for three distinct use cases of energy storage sufficiency to cover the predefined boundary and/or the expanded boundary, can successfully select and bid the available load coverage.

Starke, Michael [ORNL] (ORCID:0000000221211195)

Modeling Combined Heat and Power Systems in REopt

This webinar will explore how the National Laboratory of the Rockies' REopt(R) web tool can evaluate the techno-economics of combined heat and power (CHP) technologies and systems to assess performance and financial viability. REopt can analyze both standalone CHP systems, and systems paired with other on-site energy resources. Participants will gain practical insights into using REopt to optimize site energy strategies, reduce energy costs, and improve site energy security and outage recovery.

29 ENERGY PLANNING, POLICY, AND ECONOMY

Operating and Managing a Backup Control Center

Due to the criticality of continuous mission operations, some control centers must plan for alternate locations in the event an emergency shuts down the primary control center. Johnson Space Center (JSC) in Houston, Texas is the Mission Control Center (MCC) for the International Space Station (ISS). Due to Houston s proximity to the Gulf of Mexico, JSC is prone to threats from hurricanes which could cause flooding, wind damage, and electrical outages to the buildings supporting the MCC. Marshall Space Flight Center (MSFC) has the capability to be the Backup Control Center for the ISS if the situation is needed. While the MSFC Huntsville Operations Support Center (HOSC) does house the BCC, the prime customer and operator of the ISS is still the JSC flight operations team. To satisfy the customer and maintain continuous mission operations, the BCC has critical infrastructure that hosts ISS ground systems and flight operations equipment that mirrors the prime mission control facility. However, a complete duplicate of Mission Control Center in another remote location is very expensive to recreate. The HOSC has infrastructure and services that MCC utilized for its backup control center to reduce the costs of a somewhat redundant service. While labor talents are equivalent, experiences are not. Certain operations are maintained in a redundant mode, while others are simply maintained as single string with adequate sparing levels of equipment. Personnel at the BCC facility must be trained and certified to an adequate level on primary MCC systems. Negotiations with the customer were done to match requirements with existing capabilities, and to prioritize resources for appropriate level of service. Because some of these systems are shared, an activation of the backup control center will cause a suspension of scheduled HOSC activities that may share resources needed by the BCC. For example, the MCC is monitoring a hurricane in the Gulf of Mexico. As the threat to MCC increases, HOSC must begin a phased activation of the BCC, while working resource conflicts with normal HOSC activities. In a long duration outage to the MCC, this could cause serious impacts to the BCC host facility s primary mission support activities. This management of a BCC is worked based on customer expectations and negotiations done before emergencies occur. I.

Marsh, Angela L.

AI-Enabled Operations at Fermi Complex: Multivariate Time Series Prediction for Outage Prediction and Diagnosis

The Main Control Room of the Fermilab accelerator complex continuously gathers extensive time-series data from thousands of sensors monitoring the beam. However, unplanned events such as trips or voltage fluctuations often result in beam outages, causing operational downtime. This downtime not only consumes operator effort in diagnosing and addressing the issue but also leads to unnecessary energy consumption by idle machines awaiting beam restoration. The current threshold-based alarm system is reactive and faces challenges including frequent false alarms and inconsistent outage-cause labeling. To address these limitations, we propose an AI-enabled framework that leverages predictive analytics and automated labeling. Using data from $2,703$ Linac devices and $80$ operator-labeled outages, we evaluate state-of-the-art deep learning architectures, including recurrent, attention-based, and linear models, for beam outage prediction. Additionally, we assess a Random Forest-based labeling system for providing consistent, confidence-scored outage annotations. Our findings highlight the strengths and weaknesses of these architectures for beam outage prediction and identify critical gaps that must be addressed to fully harness AI for transitioning downtime handling from reactive to predictive, ultimately reducing downtime and improving decision-making in accelerator management.

Jain, Milan [PNL, Richland] (ORCID:000000021676111

Demonstrating the data center as a flexible grid asset using a C-HIL setup

Increasing data center demand is outpacing grid infrastructure development. Artificial intelligence workloads and hyperscale cloud growth are creating unprecedented demand for power, while traditional grid expansion faces multiyear development timelines. Verrus is developing an innovative datacenter solution for this challenge, data centers that act as active grid-supportive assets rather than passive loads. Our approach integrates a novel grid-aware power flow management system with battery energy storage systems(BESS) into a microgrid-controlled, medium-voltage power distribution architecture that delivers critical capabilities, such as: * Fast response to grid disturbances such over/ under voltage or over/ under frequency * Demand flexibility that can service requests from the utility within 10 s * Uninterrupted transition to islanded operation during grid outages * Continuous uptime assurance for compute loads while maintaining all customer service level agreements. Through Verrus' strategic partnership with the National Renewable Energy Laboratory (NREL), these capabilities were validated using NREL's Advanced Research on Integrated Energy Systems (ARIES) virtual emulation environment to model a 70-MW grid-interactive data center. This paper outlines the design, methodology, and results of this emulated deployment, demonstrating that data centers can provide both critical load resilience and ancillary grid support without compromising uptime requirements. Specifically, we present a digital real time simulation of a 70 MW data center integrated with a physical microgrid controller, and demonstrate the data center response in the event of a grid voltage and frequency event, utility demand response request and utility outage.

24 POWER TRANSMISSION AND DISTRIBUTION

Trajectory Specification Applied to Terminal Airspace

Despite major efforts to automate air traffic control (ATC), it is still performed by humans today. The complexity and safety-criticality of ATC makes it very difficult to safely automate, but it must be automated to increase airspace capacity (the density of traffic that can be safely managed) and airport throughput (the number of arrivals and departures that an airport can safely handle in a given period of time) beyond what is possible with human controllers. This paper presents the Trajectory Specification (TS) concept, which can help to safely automate ATC. TS is a method of specifying aircraft trajectories such that the position at any given time in flight is restricted to a precisely defined bounding space, removing all ambiguity as to where the flight is allowed to be. The bounding space or volume is determined by tolerances relative to a reference trajectory (position as a function of time). The tolerances are dynamic and are based on the aircraft navigation capabilities and the traffic situation. The tolerances can be a piecewise linear function of time or distance along the route, allowing the tolerances to vary as needed, typically increasing with time for departures and decreasing for arrivals. A Trajectory Specification Language (TSL) is proposed for communicating trajectories from aircraft to ATC as requests and from ATC to aircraft as assignments. The TS concept requires a new generation of airborne Flight Management Systems (FMS) that understand the TSL and can fly the assigned trajectories, but this paper focuses on the ATC functions and the prototype ATC algorithms and software that were developed to test the TS concept. Assuming conformance, TS can guarantee safe separation for an arbitrary length of time even in the event of an ATC system or communication outage. It can help to achieve the high level of safety and reliability needed for ATC automation, and it can also reduce the reliance on ATC backup systems for tactical conflict detection and resolution during normal operation. TS can be applied to any controlled airspace, including enroute, terminal, and urban airspace, but this paper presents algorithms and software for arrival spacing and conflict detection and resolution in the terminal airspace serving a major airport. In a fast-time simulation of a full day of traffic in a major terminal airspace, all conflicts were resolved in near real time, demonstrating the computational feasibility and the preliminary operational feasibility of the TS concept. This paper is a compilation of previous papers, and it adds significant information that was omitted from those papers due to length limitations. It also updates some of the results of those earlier papers due to algorithm refinements and corrections of minor software errors.

air traffic control, trajectory

Mississippi's Strategic Resilience: A multi-systems approach to secure, reliable, and adaptable electric grid infrastructure

Mississippi’s electric grid resilience challenges are linked to an intersection of complex socioeconomic, ecological, technological, historical, and political challenges, exacerbated by increasing severe weather like flooding and tornado events. The state’s legacy of underinvestment in critical energy infrastructure, particularly in rural areas and vulnerable floodplains, have stressed an aging grid, creating long-lasting disruptions in electric service during weather-related outages. Effective emergency management and preparedness is further hampered by a lack of coordination across local, county, and regional scales. Using the TASTI-GRID platform and partnership with Oak Ridge National Laboratory (ORNL), Mississippi is developing a comprehensive regional resilience strategy to overcome energy security and reliability challenges, mitigating the impacts of natural hazards, and positioning Mississippi as a resilient and premier destination for residents, businesses, and economic development.

24 POWER TRANSMISSION AND DISTRIBUTION

Monitoring Airspace Complexity and Determining Contributing Factors

The national airspace has evolved over many years to accommodate increased traffic demand while simultaneously maintaining air travel as one of the safest forms of transportation. One of the reasons for this success is the ability of the air traffic control system and the operators to adapt and accommodate to situations that routinely disrupt normal operations. These situations may include: adverse weather, delays, early arrivals, equipment outages, and other factors that are outside the operators’ ability to control. These factors can lead to states where automation is unable to properly handle these issues, and therefore air traffic controllers and pilots have to intervene — ultimately increasing communication between operators resulting in higher workload. As controller workload increases to handle sub-optimal operating conditions, complexity increases. This is because, under these conditions humans are required to make tactical decisions in response to external factors. This results in a departure from the original strategic plan where operations would be more efficiently managed. Human operators manage airspace complexity under rigid regulations but in a constantly changing environment. The airspace is divided into sectors and the number of aircraft assigned to each controller is limited for safe handling. Some prior studies devised airspace complexity metrics in commercial aviation and related these metrics to controller workload. The upper bounds on the system load are pre-determined. Such bounds on complexity make for a safe system, but the system cannot scale and adapt to autonomous, dense, and heterogeneous traffic — including the many types of Unmanned Aerial Vehicles (UAVs) envisioned to be added to the operations. We hypothesize that, as traffic density and heterogeneity grow, and other key metrics change, there will be phase transitions at which the way traffic should be managed changes significantly. We offer a method for in-time detection of contributing factors that lead to phase transitions, characterized by increased complexity. To the best of our knowledge, there is no tool similar to ours that identifies such contributing factors or precursor patterns.

Precursor

Monitoring Airspace Complexity and Determining Contributing Factors

The national airspace has evolved over many years to accommodate increased traffic demand [1] while simultaneously maintaining one of the safest forms of transportation [2], [3]. One of the reasons for this success is the ability of the system and the operators to adapt and accommodate to situations that routinely disrupt optimal operations. These situations may include: adverse weather, delays, early arrivals, equipment outages, and other factors that are outside the operators ability to control. These factors can lead to states where automation is unable to properly handle these issues and therefore air traffic controllers and pilots have to intervene, ultimately increasing communication between operators resulting in higher workload. As controller workload increases to handle sub-optimal operating conditions this can be viewed as an increase in complexity. The reasoning for this is because humans are now required to make tactical decisions in response to external factors, resulting in a departure from the strategic plan where operations would be more efficiently managed. Human operators control airspace complexity under rigid regulations that are constantly changing. The airspace is divided into sectors and the number of aircraft assigned to each controller is limited for safe handling. There has been past work that devised airspace complexity metrics in commercial aviation and related these metrics to controller workload (e.g., [4],[5]). The upper bounds on the system load are pre-determined. Such bounds on complexity make for a safe system, but the system cannot scale and adapt to autonomous, dense, and heterogeneous traffic, including the many types of Unmanned Aerial Vehicles (UAVs) envisioned to be added to the operations. We hypothesize that, as traffic density and heterogeneity grow, and other key metrics change, there will be phase transitions at which the way traffic should be managed changes significantly [6]. We offer a method for in-time detection of contributing factors that lead to phase transitions, characterized by increased complexity. To the best of our knowledge, there is no tool similar to our proposed effort that identifies such contributing factors or precursor patterns. To define the scope we are proposing to measure complexity from the viewpoint of the Terminal Radar Approach Control Facilities (TRACON) controller’s perspective. In particular we are analyzing arrivals into KSFO. With safety as the top concern for airspace operators, it is important to recognize that as density and heterogeneity grow, the focus of the system will change. Times of the day when the airspace has low density and heterogeneity, the flights will follow more efficient paths where the aircraft move on established routes that are more or less directly to the destination. However, when density and heterogeneity increases, the system will begin changing focus to avoiding conflicts and collisions and route the flights in a more flexible way. Higher flexibility requires more communication and coordination between controllers and pilots which the current automation is unable to handle. This paper proposes a novel approach that monitors airspace complexity at multiple scales, uses a Machine Learning-based tool that predicts when operations will transition to a regime of greater complexity, and identifies actions that can reduce the complexity while still maintaining efficient and safe operations. We demonstrate our proposed approach using data from multiple complementary sources. This includes, but is not limited to: historical aircraft surveillance data from NASA’s Sherlock Data Warehouse [7], METAR weather data, and airport configuration data from Aviation System Performance Metrics (ASPM). The surveillance data flight paths are sampled at a variable sample rate — increasing as the aircraft approaches the airport. This is due to how Sherlock manages flight track stitching between different radar facilities which have different sampling rates. The weather and performance data are logged at defined intervals throughout the day at a courser refresh rate. In addition to the logged data and metrics, we leverage pre-defined Standard Terminal Arrival Routes (STARs) procedures to characterize the path of each flight. Each flight files for one of these routes in the flight plan well before entering the terminal airspace, and approximately follows the route until it leaves the STAR, typically on the final fix of a runway transition. However, most flights do not always fly the full STAR procedure to completion [8], but the majority do adhere to the fixes within the common route of the procedure. Our approach leverages fixes in the common route of each of the STARs to build a reference path to the airport. This allows us to characterize the flight paths in what we are defining as the “maneuvering area” (the airspace between the STAR and before the flight is lined up on the runway’s final approach) to determine how off nominal the flights are to calculate its complexity score. Determining airspace complexity is a concept that does not have a concrete answer. In designing this metric, we consider what increases the workload for the air traffic controllers. Consequently more specialized vectoring maneuvers results in higher workload. Accordingly, we start with a theory: each flight has a direct path it takes from the STAR’s common route to the final approach’s outer marker fix for the flight’s landing runway. It is important to note that the direct path is only used as a reference. If the majority of the flights have a large consistent offset as compared to other routes it does not necessarily mean that those flights have higher complexity. We are merely building a distribution based on this direct path for that particular STAR and runway pair to determine the normal mode of operations for that route. Flights that are in the upper tail of these distributions will result in higher complexity scores and flights that fly in the median will represent the normal mode of operations and therefore will have lower complexity scores. Since flights following each STAR route take different paths to the airport, we have a different distribution for each STAR route and therefore can model these distributions to compute a complexity score from their respective normalized distributions. To evaluate the effectiveness of our proposed airspace complexity metric we will compare against an established approach based on trajectory clustering [9]. This unsupervised learning technique consists of the following steps: (1) identify the general maneuvering areas (waypoints) by performing $\kappa$-means or DBSCAN clustering on locations where aircraft frequently turn based on the surveillance radar track data, (2) map flight trajectories onto sequences of waypoints, and (3) cluster the sequences based on their common subsequences. From a high-level perspective, this baseline model learns nominal operations in the airspace through the sequence of waypoints that are representative of where aircraft change direction and defines deviations from the nominal operations as “complex.” Therefore, more deviations from the nominal operations correspond to higher complexity values. For our validation, we re-implemented this technique and tune model hyper-parameters to correctly detect waypoints for the arrival traffic into the San Francisco bay area. We will compute the complexity measure over a one-year period using our proposed technique as well as the baseline. Our validation will be based on each technique’s ability to detect a set of undesirable outcomes (e.g., go-arounds, holding patterns, average time in the airspace, etc.). Since our current complexity metric is derived from the offset from the direct reference path, it’s important to understand what causes these offsets. In many of the flights with high offset distance, flights performing holding patterns and S turns can be observed. These maneuvering tactics are utilized to add distance between the aircraft and the destination runway to prevent multiple flights from having conflicting arrival times. In order to predict a rise in complexity (or the precursor to complexity), it’s necessary to be able to identify these potential conflicts (which in turn, result in higher offsets). To do this, we define a “representative flight” for each STAR route and runway pair. This flight is approximately the path the flight would take if there was a clear path with no other flights in the airspace — including the time remaining to the airport. We first identify the flights for a given STAR runway pair using the offset to the reference path distributions that fall between the 44-55 percentiles. This yields the flights that conform to the most normal mode of operation. Each of these flights is partitioned based on the percent complete from the entry point into the maneuvering areas from 0\% – 100\% complete. Then for each percent “bin”, we take the median value of the flight’s latitude/longitude coordinates, airspeed, and (non causal) time remaining to the airport to construct a lookup table for each percent complete bin on a given route. As a flight enters the maneuvering area, we can find the estimated arrival time of a flight to the airport by finding the closest point to the representative path’s percent complete bin (relative to the flight’s current position at any snapshot in the airspace) and therefore retrieve the corresponding remaining time left on the “representative path”. We assume that the flight will follow the representative path to completion when deriving these estimates. We can then compare these estimated arrival times against other flights for the same snapshot in time to identify potential conflicts. If more flights are estimated to arrive within a tolerance window than there are runways available, then we have a potential conflict. We can use this derived measure along with other factors expected to add disruption to the operation such as weather and runway configuration changes as an input to machine learning tools to detect precursors that increases in our complexity measure. This novel method will assist in uncovering insights into the contributing factors that lead to increased complexity that may allow for in-time responses to avoid reaching a high complexity state in the airspace.

complexity

Land and Atmosphere Near-Real-Time Capability for Earth Observing System

The past decade has seen a rapid increase in availability and usage of near-real-time data from satellite sensors. The EOSDIS (Earth Observing System Data and Information System) was not originally designed to provide data with sufficiently low latency to satisfy the requirements for near-real-time users. The EOS (Earth Observing System) instruments aboard the Terra, Aqua and Aura satellites make global measurements daily, which are processed into higher-level 'standard' products within 8-40 hours of observation and then made available to users, primarily earth science researchers. However, applications users, operational agencies, and even researchers desire EOS products in near-real-time to support research and applications, including numerical weather and climate prediction and forecasting, monitoring of natural hazards, ecological/invasive species, agriculture, air quality, disaster relief and homeland security. These users often need data much sooner than routine science processing allows, usually within 3 hours, and are willing to trade science product quality for timely access. While Direct Broadcast provides more timely access to data, it does not provide global coverage. In 2002, a joint initiative between NASA (National Aeronautics and Space Administration), NOAA (National Oceanic and Atmospheric Administration), and the DOD (Department of Defense) was undertaken to provide data from EOS instruments in near-real-time. The NRTPE (Near Real Time Processing Effort) provided products within 3 hours of observation on a best-effort basis. As the popularity of these near-real-time products and applications grew, multiple near-real-time systems began to spring up such as the Rapid Response System. In recognizing the dependence of customers on this data and the need for highly reliable and timely data access, NASA's Earth Science Division sponsored the Earth Science Data and Information System Project (ESDIS)-led development of a new near-real-time system called LANCE (Land, Atmosphere Near-Real-Time Capability for EOS) in 2009. LANCE consists of special processing elements, co-located with selected EOSDIS data centers and processing facilities. A primary goal of LANCE is to bring multiple near-real-time systems under one umbrella, offering commonality in data access, quality control, and latency. LANCE now processes and distributes data from the Moderate Resolution Imaging Spectroradiometer (MODIS), Atmospheric Infrared Sounder (AIRS), Advanced Microwave Scanning Radiometer Earth Observing System (AMSR-E), Microwave Limb Sounder (MLS) and Ozone Monitoring Instrument (OMI) instruments within 3 hours of satellite observation. The Rapid Response System and the Fire Information for Resource Management System (FIRMS) capabilities will be incorporated into LANCE in 2011. LANCE maintains a central website to facilitate easy access to data and user services. LANCE products are extensively tested and compared with science products before being made available to users. Each element also plans to implement redundant network, power and server infrastructure to ensure high availability of data and services. Through the user registration system, users are informed of any data outages and when new products or services will be available for access. Building on a significant investment by NASA in developing science algorithms and products, LANCE creates products that have a demonstrated utility for applications requiring near-real-time data. From lower level data products such as calibrated geolocated radiances to higher-level products such as sea ice extent, snow cover, and cloud cover, users have integrated LANCE data into forecast models and decision support systems. The table above shows the current near-real-time product categories by instrument. The ESDIS Project continues to improve the LANCE system and use the experience gained through practice to seek adjustments to improve the quality and performance of the system. For example, anGC-compliant Web Map Service (WMS) will be added shortly that will allow users to download geo-referenced MODIS images for arbitrary bounding boxes. Further, an OGC-compliant Web Coverage Service (WCS) will be added later this year that will expedite user access to arbitrary data subsets or re-formatted products. AIRS images are now served through WMS and available in multiple formats (PNG, GeoTIFF, KMZ). NASA has established a LANCE User Working Group to steer the development of the system and create a forum for sharing ideas and experiences that are expected to further improve the LANCE capabilities. The LANCE system has proved a success by satisfying the growing needs of the applications and operational communities for land and atmosphere data in near-real-time. NASA's Earth Sciences Division was able to leverage existing science research capabilities to provide the near-real-time community with products and imagery that support monitoring of disasters in a timely manner.

Murphy, Kevin J.

Designing an Alternate Mission Operations Control Room

The Huntsville Operations Support Center (HOSC) is a multi-project facility that is responsible for 24x7 real-time International Space Station (ISS) payload operations management, integration, and control and has the capability to support small satellite projects and will provide real-time support for SLS launches. The HOSC is a serviceoriented/ highly available operations center for ISS payloads-directly supporting science teams across the world responsible for the payloads. The HOSC is required to endure an annual 2-day power outage event for facility preventive maintenance and safety inspection of the core electro-mechanical systems. While complete system shut-downs are against the grain of a highly available sub-system, the entire facility must be powered down for a weekend for environmental and safety purposes. The consequence of this ground system outage is far reaching: any science performed on ISS during this outage weekend is lost. Engineering efforts were focused to maximize the ISS investment by engineering a suitable solution capable of continuing HOSC services while supporting safety requirements. The HOSC Power Outage Contingency (HPOC) System is a physically diversified compliment of systems capable of providing identified real-time services for the duration of a planned power outage condition from an alternate control room. HPOC was designed to maintain ISS payload operations for approximately three continuous days during planned HOSC power outages and support a local Payload Operations Team, International Partners, as well as remote users from the alternate control room located in another building. This paper presents the HPOC architecture and lessons learned during testing and the planned maiden operational commissioning. Additionally, this paper documents the necessity of an HPOC capability given the unplanned HOSC Facility power outage on April 27th, 2011, as a result of the tornado outbreak that damaged the electrical grid to such a degree that significantly inhibited the Tennessee Valley Authority's ability to transmit electricity throughout the North Alabama region.

Montgomery, Patty

Gateway Autonomy for Enabling Deep Space Exploration

The Gateway spacecraft is an important stepping-stone to exploration of the solar system, integrating commercial and international partners into a tightly coupled system, enabling cislunar activities, and implementing key technologies for missions to Mars. Autonomy is a capability area necessary to handle long communication outages where intervention from Earth is impossible, to prepare to operate with long communication delays that will be common in interplanetary travel, and to make spaceflight more affordable and accessible by reducing sustaining operations costs. The Gateway Concept of Operations states that one of Gateway’s goals is to “focus on infrastructure and systems that will allow autonomous operations aboard the Gateway with robotics, automated systems, advanced communications, and distributed computing.” Gateway’s Vehicle Systems Manager (VSM) and associated Autonomous Spacecraft Management Architecture (ASMA) are key products towards delivering autonomous capability. The primary functions of the control architecture are Mission Management and Timeline Execution, Resource Management, Fault Management, and Vehicle Control and Operation (VCO). In each of these areas, there is an initial level of capability to be delivered at launch, with plans to continue development and grow to greater capability. The initial deployment of VSM will focus on maintaining vehicle safety by focusing on full fault management capabilities and deploying only enough resource and timeline planning functionality to support that. The final deployment of VSM will add significant planning and control optimization functionality to support nominal operations for up to 21 days without ground support, even accommodating fault and failure conditions. While the VSM is the vehicle-level representation of autonomous reasoning, distributed automation is essential to provide the right scope and abstraction of information to process. Module and system support of automation and simplicity of interfaces are two important design paradigms that Gateway is focusing on to garner a systems approach to autonomy. Distribution of reasoning can increase complexity, so Gateway is also taking a strict hierarchical approach to information flow and decision making. VSM is not the only capability necessary to achieve an autonomous spacecraft. Robotics support for maintenance of the spacecraft will be essential to provide continued vehicle functionality even when crew is not present. Technical and programmatic challenges exist when implementing autonomous robotics operations. These challenges include sufficient network flexibility to support data transfer to the rest of the vehicle to coordinate module-to-module robotic walk-offs and finding the proper interfaces to allow sufficient dexterity. Communication system upgrades planned for Gateway include Delay Tolerant Networking to best utilize the complex network of relays that will be part of mature cislunar operations. Distributed computing and management will provide failure tolerance, robustness, and growth of capabilities while still allowing significant reuse of heritage software on heritage systems as well as reuse of common applications across a spacecraft to minimize new development, but this requires adherence to key standards and interfaces. The Gateway program has demonstrated significant progress towards these capabilities and has identified challenges other spacecraft developers should be aware of from the start.

Molly Anderson

Gateway Autonomy for Enabling Deep Space Exploration

The Gateway spacecraft is an important stepping-stone to exploration of the solar system, integrating commercial and international partners into a tightly coupled system, enabling cislunar activities, and implementing key technologies for missions to Mars. Autonomy is a capability area necessary to handle long communication outages where intervention from Earth is impossible, to prepare to operate with long communication delays that will be common in interplanetary travel, and to make spaceflight more affordable and accessible by reducing sustaining operations costs. The Gateway Concept of Operations states that one of Gateway’s goals is to “focus on infrastructure and systems that will allow autonomous operations aboard the Gateway with robotics, automated systems, advanced communications, and distributed computing.” Gateway’s Vehicle Systems Manager (VSM) and associated Autonomous Spacecraft Management Architecture (ASMA) are key products towards delivering autonomous capability. The primary functions of the control architecture are Mission Management and Timeline Execution, Resource Management, Fault Management, and Vehicle Control and Operation (VCO). In each of these areas, there is an initial level of capability to be delivered at launch, with plans to continue development and grow to greater capability. The initial deployment of VSM will focus on maintaining vehicle safety by focusing on full fault management capabilities and deploying only enough resource and timeline planning functionality to support that. The final deployment of VSM will add significant planning and control optimization functionality to support nominal operations for up to 21 days without ground support, even accommodating fault and failure conditions. While the VSM is the vehicle-level representation of autonomous reasoning, distributed automation is essential to provide the right scope and abstraction of information to process. Module and system support of automation and simplicity of interfaces are two important design paradigms that Gateway is focusing on to garner a systems approach to autonomy. Distribution of reasoning can increase complexity, so Gateway is also taking a strict hierarchical approach to information flow and decision making. VSM is not the only capability necessary to achieve an autonomous spacecraft. Robotics support for maintenance of the spacecraft will be essential to provide continued vehicle functionality even when crew is not present. Technical and programmatic challenges exist when implementing autonomous robotics operations. These challenges include sufficient network flexibility to support data transfer to the rest of the vehicle to coordinate module-to-module robotic walk-offs and finding the proper interfaces to allow sufficient dexterity. Communication system upgrades planned for Gateway include Delay Tolerant Networking to best utilize the complex network of relays that will be part of mature cislunar operations. Distributed computing and management will provide failure tolerance, robustness, and growth of capabilities while still allowing significant reuse of heritage software on heritage systems as well as reuse of common applications across a spacecraft to minimize new development, but this requires adherence to key standards and interfaces. The Gateway program has demonstrated significant progress towards these capabilities and has identified challenges other spacecraft developers should be aware of from the start.

Molly Anderson

Designing an Alternate Mission Operations Control Room

The Huntsville Operations Support Center (HOSC) is a multi-project facility that is responsible for 24x7 real-time International Space Station (ISS) payload operations management, integration, and control and has the capability to support small satellite projects and will provide real-time support for SLS launches. The HOSC is a service-oriented/ highly available operations center for ISS payloads-directly supporting science teams across the world responsible for the payloads. The HOSC is required to endure an annual 2-day power outage event for facility preventive maintenance and safety inspection of the core electro-mechanical systems. While complete system shut-downs are against the grain of a highly available sub-system, the entire facility must be powered down for a weekend for environmental and safety purposes. The consequence of this ground system outage is far reaching: any science performed on ISS during this outage weekend is lost. Engineering efforts were focused to maximize the ISS investment by engineering a suitable solution capable of continuing HOSC services while supporting safety requirements. The HOSC Power Outage Contingency (HPOC) System is a physically diversified compliment of systems capable of providing identified real-time services for the duration of a planned power outage condition from an alternate control room. HPOC was designed to maintain ISS payload operations for approximately three continuous days during planned HOSC power outages and support a local Payload Operations Team, International Partners, as well as remote users from the alternate control room located in another building.

Montgomery, Patty