Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “software resilience”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Implementing Software Resiliency in HPX for Extreme Scale Computing

The DOE Office of Science Exascale Computing Project (ECP) outlines the next milestones in the supercomputing domain. The target computing systems under the project will deliver 10x performance while keeping the power budget under 30 megawatts. With such large machines, the need to make applications resilient has become paramount. The benefits of adding resiliency to mission critical and scientific applications, includes the reduced cost of restarting the failed simulation both in terms of time and power. Most of the current implementation of resiliency at the software level makes use of a Coordinated Checkpoint and Restart (C/R). This technique of resiliency generates a consistent global snapshot, also called a checkpoint. Generating snapshots involves global communication and coordination and is achieved by synchronizing all running processes. The generated checkpoint is then stored in some form of persistent storage. On failure detection, the runtime initiates a global rollback to the most recent previously saved checkpoint. This involves aborting all running processes, rolling them back to the previous state and restarting them.

97 MATHEMATICS AND COMPUTING↗

VAC: A Software Approach to Resilient SCADA Automation

To better secure critical infrastructure, especially power systems, this paper introduces a virtual SCADA automation controller. The automation controller is a gateway into a power subsystem, making it a valuable target for cyber-attacks that could cut it off from the control center and cause a loss of view and control. To prevent this, the Virtual Automation Controller (VAC) is a backup device that mirrors the capabilities of the physical controller. It can communicate via Modbus and DNP3 and is containerized so it can be deployed on a variety of platforms. Furthermore, it utilizes software-defined networking to quickly disconnect a failed automation controller and preserve its state for forensics. The VAC gives system operators time to replace the failed controller and prevents dangerous and costly damage to power systems. The VAC is compared against the SEL 3505-3 RTAC and shown to have the necessary features to act as a failover controller.

Johnson, Jordan↗

AEDAM: Whole Program Adaptive Error Detection and Mitigation (Final Report)

The overall goals of the AEDAM project were to fundamentally transform software transient-error detection through the design of configurable specialized detectors, quantitative characterization of hardware resilience and software vulnerabilities, and composition of the specialized detectors to most efficiently protect the whole program. Within this scope, UT Austin's research contribution related to enabling and studying the composition of detectors and possible different hardware errors, specifically: (1) developed the Hamartia open-source error injection framework that is designed to make composition studies simple, and (2) develop the methodology and demonstrate the potential benefits of error detector composition.

97 MATHEMATICS AND COMPUTING↗

RE+Microgrid Conference Presentation

Conference presentation about Idaho National Laboratory tools for damage prediction and resilience, along with an overview of microgrid capabilities.

Grid↗

PROTEUS: Machine Learning Driven Resilience for Extreme-scale Systems

The objective of this project is to design, develop, and evaluate scalable software to enhance resilience, data checkpointing, program restart, and analysis. The proposed tasks are to 1) develop scalable machine learning techniques to learn temporal change patterns in a scalable and in-situ manner, and to minimize data movement and maximize learning locally closest to data; 2) design a concise data representation and indexing mechanism to capture the distribution of changes in data that can guarantee point-wise user-defined tolerable errors while reducing the data storage requirements by an order of magnitude or more; 3) develop data reduction techniques as library modules; 4) exploit local SSD for minimizing data movement in storage hierarchy; 5) develop anomaly detection algorithms that can predict corruptions based on learning of emerging patterns; 6) develop software libraries to be incorporated within widely used data formats and APIs; and 7) evaluate the proposed software using DOE scientific applications. The outcomes of the proposed work are to satisfy many synergistic data reduction and resilience requirements for large-scale data intensive applications executed on extreme-scale computing systems. The developed mechanism for error-bound data approximation is directly applicable to existing scientific applications. Through machine learning from historical events and change distribution, this work will enable anomaly detection for DOE computer facility.

97 MATHEMATICS AND COMPUTING↗

Power Distribution Designing For Resilience Application (powdder)

Power Distribution Designing for Resilience Application (PowDDeR) is a software application to succinctly capture the capabilities of a power system to respond to disturbances, including natural or human (malicious or errors) caused disturbances. The software provides a measure of resilience for power systems.

McJunkin, TimothyR↗

Distributed Software-Defined Network Architecture for Smart Grid Resilience to Denial-of-Service Attacks

An important challenge for smart grid security is designing a secure and robust smart grid communications architecture to protect against cyber-threats, such as Denial-of-Service (DoS) attacks, that can adversely impact the operation of the power grid. Researchers have proposed using Software Defined Network frameworks to enhance cybersecurity of the smart grid, but there is a lack of benchmarking and comparative analyses among the many techniques. In this work, a distributed three-controller software-defined networking (D3-SDN) architecture, benchmarking, and comparative analysis with other techniques is presented. The selected distributed flat SDN architecture divides the network horizontally into multiple areas or clusters, where each cluster is handled by a single Open Network Operating System (ONOS) controller. A case study using the IEEE 118-bus system is provided to compare the performance of the presented ONOS-managed D3-SDN, against the POX controller. In addition, the proposed architecture outperforms a single SDN controller framework by a tenfold increase in throughput; a reduction in latency of > 20%; and an increase in throughput of approximately 11% during the DoS attack scenarios.

Agnew Jr., Dennis↗

Accurate Consensus-based Distributed Averaging with Variable Time Delay in Support of Distributed Secondary Control Algorithms

We report that distributed secondary control has been widely used in hierarchical control structures, where multiple distributed generators (DGs) need to coordinate to regulate system voltage and frequency. In these systems, consensus algorithms determine the average of a group of dynamic states (e.g. voltages measured by a group of DGs). To be useful, consensus algorithms must be computationally efficient, stable and accurate. In practice, numerous practical implementation challenges significantly affect the consensus equilibrium. In this paper, we quantify the accuracy deviations of the distributed average observer algorithms proposed in the literature to demonstrate the problems with the state-of-the-art distributed averaging techniques. A novel approach is proposed that achieves accurate average tracking in the presence of time-varying communication delays among agents. In our implementation, time synchronization of all distributed controllers is enabled by a novel software platform, called Resilient Information Architecture Platform for the Smart Grid (RIAPS). The proposed distributed average observer is implemented on hardware controllers and its effectiveness is validated in a controller hardware-in-the-loop testbed.

24 POWER TRANSMISSION AND DISTRIBUTION↗

RDPM: An Extensible Tool for Resilience Design Patterns Modelling

Resilience to faults, errors, and failures in extreme-scale high-performance computing (HPC) systems is a critical challenge. Resilience design patterns offer a new, structured hardware and software design approach for improving resilience. While prior work focused on developing performance, reliability, and availability models for resilience design patterns, this paper extends it by providing a Resilience Design Patterns Modeling (RDPM) tool which allows (1) exploring performance, reliability, and availability of each resilience design pattern, (2) offering customization of parameters to optimize performance, reliability, and availability, and (3) allowing investigation of trade-off models for combining multiple patterns for practical resilience solutions.

Kumar, Mohit↗

Designing a decentralized fault-tolerant software framework for smart grids and its applications

The vision of the ‘Smart Grid’ anticipates a distributed real-time embedded system that implements various monitoring and control functions. As the reliability of the power grid is critical to modern society, the software supporting the grid must support fault tolerance and resilience of the resulting cyber-physical system. This paper describes the fault-tolerance features of a software framework called Resilient Information Architecture Platform for Smart Grid (RIAPS). The framework supports various mechanisms for fault detection and mitigation and works in concert with the applications that implement the grid-specific functions. The paper discusses the design philosophy for and the implementation of the fault tolerance features and presents an application example to show how it can be used to build highly resilient systems.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗