Resilient Execution Spaces: Challenges of Large-Scale Software Research During Quarantine.
Abstract not provided.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Abstract not provided.
Resilience to faults, errors, and failures in extreme-scale high-performance computing (HPC) systems is a critical challenge. Resilience design patterns offer a new, structured hardware and software design approach for improving resilience. While prior work focused on developing performance, reliability, and availability models for resilience design patterns, this paper extends it by providing a Resilience Design Patterns Modeling (RDPM) tool which allows (1) exploring performance, reliability, and availability of each resilience design pattern, (2) offering customization of parameters to optimize performance, reliability, and availability, and (3) allowing investigation of trade-off models for combining multiple patterns for practical resilience solutions.
The vision of the ‘Smart Grid’ anticipates a distributed real-time embedded system that implements various monitoring and control functions. As the reliability of the power grid is critical to modern society, the software supporting the grid must support fault tolerance and resilience of the resulting cyber-physical system. This paper describes the fault-tolerance features of a software framework called Resilient Information Architecture Platform for Smart Grid (RIAPS). The framework supports various mechanisms for fault detection and mitigation and works in concert with the applications that implement the grid-specific functions. The paper discusses the design philosophy for and the implementation of the fault tolerance features and presents an application example to show how it can be used to build highly resilient systems.
Grid-edge devices are becoming increasingly important in the energy transition. Preserving privacy was not previously considered an important aspect for power grid operations, but with the increased proliferation of customer-owned assets, it is now an essential consideration. Several mechanisms have been proposed to provide privacy for non-utility owned assets in the power grid. Federated learning (FL) is one method gaining prominence in this area. Although FL has been used for other applications, such as auto-complete in phones, there has not been much investigation into whether these approaches are feasible for grid applications. In this work, we use a research platform with real-time simulators and hardware-in-the-loop capabilities to investigate how FL can be applied to grid-edge devices, and we present the potential grid services that can be derived for these devices. We discuss the computational challenges with deploying complex FL approaches, and we explore several grid services, including participation in retail electricity markets, voltage control, and resilience-driven reconfiguration.
Reliability is a serious concern for future extreme-scale high-performance computing (HPC) systems. Projections based on the current generation of HPC systems and technology roadmaps suggest the prevalence of very high fault rates in future systems. The errors resulting from these faults will propagate and generate various kinds of failures, which may result in outcomes ranging from result corruptions to catastrophic application crashes. Therefore, the resilience challenge for extreme-scale HPC systems requires coordination between various hardware and software technologies that are capable of handling a broad set of fault models at accelerated fault rates. Also, due to practical limits on power consumption in future HPC systems, they are likely to embrace innovative architectures, increasing the levels of hardware and software complexities. Therefore, the techniques that seek to improve resilience must navigate the complex trade-off space between resilience and the overheads to power consumption and performance. While the HPC community has developed various resilience solutions, application-level techniques as well as system-based solutions, the solution space of HPC resilience techniques remains fragmented. There are no formal methods to integrate the various HPC resilience techniques into composite solutions, nor are there methods to holistically evaluate the adequacy and efficacy of such solutions in terms of their protection coverage, and their performance & power efficiency characteristics. Additionally, few implementations of current resilience solutions are portable to newer architectures and software environments that will be deployed on future systems. We developed a new structured approach to the management of HPC resilience using the concept of resilience-based design patterns. In general, a design pattern is a repeatable solution to a commonly occurring problem. We identified the well-known solutions that are commonly used to deal with faults, errors and failures in HPC systems. In the initial design patterns specification (version 1.0), we described the various solutions, which address specific problems in the design of resilient HPC environments, in the form of patterns. Each pattern describes a problem caused by a fault, error or failure event in an HPC environment, and then describes the core of the solution of the problem in such a way that this solution may be adapted to different systems and implemented at different layers of the system stack. The catalog of these resilience design patterns provides designers with a collection of design elements. To construct complete resilience solutions using combinations of various patterns, we defined a framework that enhances HPC designers' understanding of the important constraints and the opportunities for the design patterns to be implemented and deployed at various layers of the system stack. The design framework is also useful for establishing interfaces and mechanisms to coordinate flexible fault management across hardware and software components, as well as to consider the trade-off between performance, resilience, and power consumption when constructing a solution. The resilience design patterns specification version 1.1 included more detailed explanations of the pattern solutions, the context in which the patterns are applicable, and the implications for hardware or software design. It also provided several additional examples and detailed case studies to demonstrate the use of patterns to build realistic solutions. In version 1.2 of the specification document, we have improved the pattern descriptions, including graphical representations of the pattern components. These improvements are largely based on critical comments, feedback and suggestions received from pattern experts and readers of the previous versions of the specification. The pattern classification has been modified to further clarify the relationships between pattern categories. This version of the specification also introduces a pattern language for resilience design patterns. The pattern language presents the patterns in the catalog as a network, revealing the relations among the resilience patterns. The language provides designers with the means to explore alternative techniques for handling a specific fault model that may have different efficiency and complexity characteristics. Using the pattern language also enables the design and implementation of comprehensive resilience solutions as a set of interconnected resilience patterns that can be instantiated across layers of the system stack. The overall goal of this work is to provide hardware and software designers, as well as the users and operators of HPC systems, a systematic methodology for the design and evaluation of resilience technologies in HPC systems that keep scientific applications running to a correct solution in a timely and cost-efficient manner despite frequent faults, errors, and failures of various types. Version 2.0 expands the resilience design pattern classification and catalog to include self-stabilization patterns and reliability, availability and performance models for each structural pattern.
For high-performance computing (HPC) system designers and users, meeting the myriad challenges of next-generation exascale supercomputing systems requires rethinking their approach to application and system software design. Among these challenges, providing resiliency and stability to the scientific applications in the presence of high fault rates requires new approaches to software architecture and design. As HPC systems become increasingly complex, they require intricate solutions for detection and mitigation for various modes of faults and errors that occur in these large-scale systems, as well as solutions for failure recovery. These resiliency solutions often interact with and affect other system properties, including application scalability, power and energy efficiency. Therefore, resilience solutions for HPC systems must be thoughtfully engineered and deployed.In previous work, we developed the concept of resilience design patterns, which consist of templated solutions based on well-established techniques for detection, mitigation and recovery. In this paper, we use these patterns as the foundation to propose new approaches to designing runtime systems for HPC systems. The instantiation of these patterns within a runtime system enables flexible and adaptable end-to-end resiliency solutions for HPC environments. The paper describes the architecture of the runtime system, named Plexus, and the strategies for dynamically composing and adapting pattern instances under runtime control. This runtime-based approach enables actively balancing the cost-benefit trade-off between performance overhead and protection coverage of the resilience solutions. Based on a prototype implementation of PLEXUS, we demonstrate the resiliency and performance gains achieved by the pattern-based runtime system for a parallel linear solver application.
The NASA Resilient Autonomy Project developed a software framework that implemented a Run Time Assurance (RTA) architecture that leveraged ASTM International’s F3269 Industry Standard for safely bounding complex behavior in aircraft. This framework was called the Expandable Variable Autonomy Architecture, or EVAA. EVAA was developed during the height of the Covid-19 lockdown that caused the Resilient Autonomy team to pivot from flight test to distributed simulator testing. EVAA was developed to be platform and mission agnostic where platform specifics were behind a hardware abstraction layer that EVAA called a Coupler. EVAA was able to host multiple safety monitors that could resolve individual safety hazards. EVAA was able to resolve priority conflicts when multiple safety hazards needed to be resolved simultaneously and was able to resolve highly complex situations in a safe manner that could exceed human capabilities.
SAND2024-01040O PyRoCS software synthesizes mathematical equations from several domains—including information theory, ecology, and engineering sciences—to support resilience analysis for complex systems. Resilience is the ability of the complex system being analyzed to withstand, operate through, and recover from a disruption. The complex system can be a physical system such as an electric grid, an organization such as a company, or even a subfunction of an organization. Existing mathematical equations for resilience analysis are found within multiple domains including information theory, biological sciences, and complex systems. This package synthesizes and refactors equations from these various domains to make them more generalizable for application across different types of complex systems relevant for resilience analysis. Users will be able to apply these equations to characterize different components of complex systems based on available data. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.
Since the first use of computers in space and aircraft, software errors have occurred. These errors can manifest as loss-of-life or less catastrophically. As the demand for automation increases, software in mission or safety-critical systems should be designed to be tolerant to the most likely software faults. This paper categorizes a set of 55 historic aerospace software error incidents from 1962 to 2023 to determine trends of how and where automation is most likely to fail, behaving unexpectedly. A distinction between software producing unexpected (erroneous) output versus no output (failsilent) is introduced. Of the historical incidents analyzed, 85% were from software producing wrong output rather than simply stopping. Rebooting was found to be ineffective to clear erroneous behavior, and not reliable to recover from silent failures. Error origin was within the code/logic itself in 58% of cases, 16% from configurable data, 15% from unexpected sensor input, and 11% from command/operator input. A substantial forty percent (40%) of unexpected software behavior was indicated by the absence of code, arising from unanticipated situations and missing requirements, and 16% of incidents were subjectively deemed “unknown-unknowns”. No incidents were found to be the result of programming language, compiler, tool, or operating system; and only sixteen percent (16%) of all incidents were considered errors traditional computer science/programming in nature. These findings indicate that for fault tolerance, erroneous automation behavior must be a primary consideration especially at critical moments, and reboot recoverability may not be viable. Special care should be taken to validate configurable data and commands prior to use. “Test-like-you-fly”, including hardware-in-the-loop combined with robust off-nominal testing should be used to uncover missing logic arising from unanticipated situations not covered by requirements alone. This study uniquely focuses on manifestations of unexpected flight software behavior, independent of ultimate root cause. We characterize software error behavior and origin to improve software design, test, and operations for resilience to the most common manifestations, and provide a rich dataset for further study.
The supply chain attack pathway is being increasingly used by adversaries to bypass security controls and gain unauthorized access to sensitive networks and equipment (e.g., Critical Digital Assets). Cyber-attacks targeting supply chain generally aim to compromise the environments, products, or services of vendors and suppliers to inject, add, or substitute authentic software and hardware with malicious elements. These malicious elements are deemed to be authentic as they arise from the vendor or supplier (i.e., the supply chain). This research aims at providing a survey of technologies that have the potential to reduce exposure of sensitive networks and equipment to these attacks, thereby improving tamper resistance. The recent advances in the performance and capabilities of these technologies in recent years has increased their potential applications to reduce or mitigate exposure of the supply chain attack pathway. The focus being on providing an analysis of the benefits and disadvantages of smart cards, secure tokens, and elements to provide root of trust. This analysis provides evidence that these roots of trust can increase the technical capability of equipment and networks to authenticate changes to software and configuration thereby increasing resilience to some supply chain attacks, such as those related to logistics and ICT channels, but not development environment attacks.
With the falling costs of solar arrays and battery storage and reduced reliability of the grid due to natural disasters, small-scale local generation and storage resources are beginning to proliferate. However, very few software options exist for integrated control of building loads, batteries and other distributed energy resources. The available software solutions on the market can force customers to adopt one particular ecosystem of products, thus limiting consumer choice, and are often incapable of operating independently of the grid during blackouts. In this software package, we present the "Solar+ Optimizer" (SPO), a control platform that provides demand flexibility, resiliency and reduced utility bills, built using open-source software. SPO employs Model Predictive Control (MPC) to produce real time optimal control strategies for the building loads and the distributed energy resources on site. SPO is designed to be vendor-agnostic, protocol-independent and resilient to loss of wide-area network connectivity. The software was evaluated in a real convenience store in northern California with on-site solar generation, battery storage and control of HVAC and commercial refrigeration loads. Preliminary tests showed price responsiveness of the building and cost savings of more than 10% in energy costs alone.
Hydra Autoconfig is an algorithm and software program to automatically configure a resilient network of independent ASICs that are capable of being connected as a Hydra network. The Hydra network itself is described in an invention disclosure titled: "Ad-Hoc Networks of Readout ASICs for Reliable Detector Instrumentation" (2020-091). The Hydra Autoconfig software works by first configuring a Hydra node. Then it recursively builds a Hydra network by attempting to link to each available upstream node in succession, keeping links that are verified. It then repeats on each newly generated node until no possible links remain. Once it is so configured, a reliable Hydra network of ASICs remains that can be reconfigured later to respond to damage or malfunction of one or more of the constituent ASICs.
ReNCAT is a software application that suggests microgrid portfolios that reduce the impact of large-scale disruptions to power, as measured by the Social Burden Metric. ReNCAT examines a power distribution network to identify regions that can be isolated into microgrids that enable critical services to be provided even if the remainder of the study area is left without power. ReNCAT operates on a simplified representation of the power grid, one that aggregates and approximates loads and conductors. Microgrids are formed within the power network by setting switch states to split or join portions of the grid. ReNCAT identifies candidate microgrid portfolios with varying tradeoffs between cost and service availability.
Modern scientific workflows process massive amounts of data from diverse instruments and sensors, leveraging geographically distributed, heterogeneous compute and storage resources—from leadership-class systems to edge devices—connected by high-performance networks. The diversity of resources introduces challenges in harnessing their full potential, with resilience issues arising across applications, system software, networks, storage, and hardware. Today, workflow management systems (WMS) coordinate the execution of computation and data management tasks across target resources. However, WMS’s centralized nature makes them vulnerable to faults and scalability issues that may result in failures of entire computational campaigns. In conclusion, this paper introduces a novel agentic framework for workflow management, fully distributing and decentralizing the WMS functions and modeling them as swarm intelligence agents infused with advanced artificial intelligence solutions and traditional distributed computing algorithms that can make coordinated decisions in the presence of failures of the underlying cyberinfrastructure.
Software and services that access Planetary Data System (PDS) PDS4 data products need to parse product labels to retrieve, interpret, and process the referenced digital objects. Under PDS4 a driving principle is that the product label provide all of the information necessary for these functions to be performed accurately. However, significantly more information is available in the PDS4 Information Model (IM)[1], the controlling document used to define, create, and syntactically and semantically verify the product labels. This additional information in the IM is made available for use, by both software and services, to configure, promote resiliency, and improve interoperability.
Smart cities depend on flexible and secure energy systems to ensure resilient power for critical infrastructure; however, recent weather-related events and cyberattacks have highlighted weaknesses in our energy systems, with the potential for widespread economic and security impacts. As stated by the Executive Office of the President, "the resilience of the US electric grid is a key part of the nation's defense against severe weather." To address the energy delivery security challenge, microgrids are rising as a viable solution that enhances the flexibility and resilience of the distribution grid and boosts the reliability of the local supply for the end-user. Traditionally, high capital investment has been a barrier to large-scale adoption of microgrid technology. Understanding the flexibility and resilience benefits of microgrids and accounting for the associated value streams can make the microgrid's proposition economically viable. In this chapter, microgrids' utility and their potential to serve as a flexible and resilient resource for the utility grid by providing capabilities such as peak shaving, demand response, and frequency regulation is presented. Moreover, other value streams, such as (1) their ability to island during a disaster and sustain critical loads which makes them a robust resilience solution for end-users, in the event of the utility grid outage and (2) microgrids also provide a flexible platform for integrating distributed energy resources in conjunction with storage and conventional generation technologies, strengthen microgrid's role in reducing the over-arching goal of emission reduction. Given the myriad of benefits associated with microgrids, we present strategies which can be employed for making microgrid itself resilient against physical and cyberthreats by employing hardware, software, and personnel training solutions to operate the microgrid before, during, and after a potential disaster. This chapter, thus, provides a holistic study of the microgrid as a resilience resource, for the utility grid, and a self-contained end-user for the end-user.
With the falling costs of solar arrays and battery storage and reduced reliability of the grid due to natural disasters, small-scale local generation and storage resources are beginning to proliferate. However, very few software options exist for integrated control of building loads, batteries and other distributed energy resources. The available software solutions on the market can force customers to adopt one particular ecosystem of products, thus limiting consumer choice, and are often incapable of operating independently of the grid during blackouts. In this paper, we present the “Solar+ Optimizer” (SPO), a control platform that provides demand flexibility, resiliency and reduced utility bills, built using open-source software. SPO employs Model Predictive Control (MPC) to produce real time optimal control strategies for the building loads and the distributed energy resources on site. SPO is designed to be vendor-agnostic, protocol-independent and resilient to loss of wide-area network connectivity. The software was evaluated in a real convenience store in northern California with on-site solar generation, battery storage and control of HVAC and commercial refrigeration loads. Preliminary tests showed price responsiveness of the building and cost savings of more than 10% in energy costs alone.
GronOR is a program package for nonorthogonal configuration interaction calculations. Electronic wave functions are constructed in terms of antisymmetrized products of multiconfiguration molecular fragment wave functions. The computational complexity of the nonorthogonal methodologies implemented in GronOR applied to large molecular assemblies requires a design that takes full advantage of massively parallel supercomputer architectures and accelerator technologies. This work describes the implementation strategy and resulting performance characteristics. In addition to parallelization and acceleration, the software development strategy includes aspects of fault resiliency and heterogeneous computing. The program was designed for large-scale supercomputers but also runs effectively on small clusters and workstations for small molecular systems. GronOR is available as open source to the scientific community.