Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “software resilience”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Implementing Software Resiliency in HPX for Extreme Scale Computing

The DOE Office of Science Exascale Computing Project (ECP) outlines the next milestones in the supercomputing domain. The target computing systems under the project will deliver 10x performance while keeping the power budget under 30 megawatts. With such large machines, the need to make applications resilient has become paramount. The benefits of adding resiliency to mission critical and scientific applications, includes the reduced cost of restarting the failed simulation both in terms of time and power. Most of the current implementation of resiliency at the software level makes use of a Coordinated Checkpoint and Restart (C/R). This technique of resiliency generates a consistent global snapshot, also called a checkpoint. Generating snapshots involves global communication and coordination and is achieved by synchronizing all running processes. The generated checkpoint is then stored in some form of persistent storage. On failure detection, the runtime initiates a global rollback to the most recent previously saved checkpoint. This involves aborting all running processes, rolling them back to the previous state and restarting them.

97 MATHEMATICS AND COMPUTING↗

VAC: A Software Approach to Resilient SCADA Automation

To better secure critical infrastructure, especially power systems, this paper introduces a virtual SCADA automation controller. The automation controller is a gateway into a power subsystem, making it a valuable target for cyber-attacks that could cut it off from the control center and cause a loss of view and control. To prevent this, the Virtual Automation Controller (VAC) is a backup device that mirrors the capabilities of the physical controller. It can communicate via Modbus and DNP3 and is containerized so it can be deployed on a variety of platforms. Furthermore, it utilizes software-defined networking to quickly disconnect a failed automation controller and preserve its state for forensics. The VAC gives system operators time to replace the failed controller and prevents dangerous and costly damage to power systems. The VAC is compared against the SEL 3505-3 RTAC and shown to have the necessary features to act as a failover controller.

Johnson, Jordan↗

AEDAM: Whole Program Adaptive Error Detection and Mitigation (Final Report)

The overall goals of the AEDAM project were to fundamentally transform software transient-error detection through the design of configurable specialized detectors, quantitative characterization of hardware resilience and software vulnerabilities, and composition of the specialized detectors to most efficiently protect the whole program. Within this scope, UT Austin's research contribution related to enabling and studying the composition of detectors and possible different hardware errors, specifically: (1) developed the Hamartia open-source error injection framework that is designed to make composition studies simple, and (2) develop the methodology and demonstrate the potential benefits of error detector composition.

97 MATHEMATICS AND COMPUTING↗

Risk-Significant Adverse Condition Awareness Strengthens Assurance of Fault Management Systems

As spaceflight systems increase in complexity, Fault Management (FM) systems are ranked high in risk-based assessment of software criticality, emphasizing the importance of establishing highly competent domain expertise to provide assurance. Adverse conditions (ACs) and specific vulnerabilities encountered by safety- and mission-critical software systems have been identified through efforts to reduce the risk posture of software-intensive NASA missions. Acknowledgement of potential off-nominal conditions and analysis to determine software system resiliency are important aspects of hazard analysis and FM. A key component of assuring FM is an assessment of how well software addresses susceptibility to failure through consideration of ACs. Focus on significant risk predicted through experienced analysis conducted at the NASA Independent Verification & Validation (IV&V) Program enables the scoping of effective assurance strategies with regard to overall asset protection of complex spaceflight as well as ground systems. Research efforts sponsored by NASAs Office of Safety and Mission Assurance (OSMA) defined terminology, categorized data fields, and designed a baseline repository that centralizes and compiles a comprehensive listing of ACs and correlated data relevant across many NASA missions. This prototype tool helps projects improve analysis by tracking ACs and allowing queries based on project, mission type, domain/component, causal fault, and other key characteristics. Vulnerability in off-nominal situations, architectural design weaknesses, and unexpected or undesirable system behaviors in reaction to faults are curtailed with the awareness of ACs and risk-significant scenarios modeled for analysts through this database. Integration within the Enterprise Architecture at NASA IV&V enables interfacing with other tools and datasets, technical support, and accessibility across the Agency. This paper discusses the development of an improved workflow process utilizing this database for adaptive, risk-informed FM assurance that critical software systems will safely and securely protect against faults and respond to ACs in order to achieve successful missions.

IV&V↗

Risk-Significant Adverse Condition Awareness Strengthens Assurance of Fault Management Systems

As spaceflight systems increase in complexity, Fault Management (FM) systems are ranked high in risk-based assessment of software criticality, emphasizing the importance of establishing highly competent domain expertise to provide assurance. Adverse conditions (ACs) and specific vulnerabilities encountered by safety- and mission-critical software systems have been identified through efforts to reduce the risk posture of software-intensive NASA missions. Acknowledgement of potential off-nominal conditions and analysis to determine software system resiliency are important aspects of hazard analysis and FM. A key component of assuring FM is an assessment of how well software addresses susceptibility to failure through consideration of ACs. Focus on significant risk predicted through experienced analysis conducted at the NASA Independent Verification Validation (IVV) Program enables the scoping of effective assurance strategies with regard to overall asset protection of complex spaceflight as well as ground systems. Research efforts sponsored by NASA's Office of Safety and Mission Assurance defined terminology, categorized data fields, and designed a baseline repository that centralizes and compiles a comprehensive listing of ACs and correlated data relevant across many NASA missions. This prototype tool helps projects improve analysis by tracking ACs and allowing queries based on project, mission type, domaincomponent, causal fault, and other key characteristics. Vulnerability in off-nominal situations, architectural design weaknesses, and unexpected or undesirable system behaviors in reaction to faults are curtailed with the awareness of ACs and risk-significant scenarios modeled for analysts through this database. Integration within the Enterprise Architecture at NASA IVV enables interfacing with other tools and datasets, technical support, and accessibility across the Agency. This paper discusses the development of an improved workflow process utilizing this database for adaptive, risk-informed FM assurance that critical software systems will safely and securely protect against faults and respond to ACs in order to achieve successful missions.

Fault management↗

RE+Microgrid Conference Presentation

Conference presentation about Idaho National Laboratory tools for damage prediction and resilience, along with an overview of microgrid capabilities.

Grid↗

PROTEUS: Machine Learning Driven Resilience for Extreme-scale Systems

The objective of this project is to design, develop, and evaluate scalable software to enhance resilience, data checkpointing, program restart, and analysis. The proposed tasks are to 1) develop scalable machine learning techniques to learn temporal change patterns in a scalable and in-situ manner, and to minimize data movement and maximize learning locally closest to data; 2) design a concise data representation and indexing mechanism to capture the distribution of changes in data that can guarantee point-wise user-defined tolerable errors while reducing the data storage requirements by an order of magnitude or more; 3) develop data reduction techniques as library modules; 4) exploit local SSD for minimizing data movement in storage hierarchy; 5) develop anomaly detection algorithms that can predict corruptions based on learning of emerging patterns; 6) develop software libraries to be incorporated within widely used data formats and APIs; and 7) evaluate the proposed software using DOE scientific applications. The outcomes of the proposed work are to satisfy many synergistic data reduction and resilience requirements for large-scale data intensive applications executed on extreme-scale computing systems. The developed mechanism for error-bound data approximation is directly applicable to existing scientific applications. Through machine learning from historical events and change distribution, this work will enable anomaly detection for DOE computer facility.

97 MATHEMATICS AND COMPUTING↗

Power Distribution Designing For Resilience Application (powdder)

Power Distribution Designing for Resilience Application (PowDDeR) is a software application to succinctly capture the capabilities of a power system to respond to disturbances, including natural or human (malicious or errors) caused disturbances. The software provides a measure of resilience for power systems.

McJunkin, TimothyR↗

Distributed Software-Defined Network Architecture for Smart Grid Resilience to Denial-of-Service Attacks

An important challenge for smart grid security is designing a secure and robust smart grid communications architecture to protect against cyber-threats, such as Denial-of-Service (DoS) attacks, that can adversely impact the operation of the power grid. Researchers have proposed using Software Defined Network frameworks to enhance cybersecurity of the smart grid, but there is a lack of benchmarking and comparative analyses among the many techniques. In this work, a distributed three-controller software-defined networking (D3-SDN) architecture, benchmarking, and comparative analysis with other techniques is presented. The selected distributed flat SDN architecture divides the network horizontally into multiple areas or clusters, where each cluster is handled by a single Open Network Operating System (ONOS) controller. A case study using the IEEE 118-bus system is provided to compare the performance of the presented ONOS-managed D3-SDN, against the POX controller. In addition, the proposed architecture outperforms a single SDN controller framework by a tenfold increase in throughput; a reduction in latency of > 20%; and an increase in throughput of approximately 11% during the DoS attack scenarios.

Agnew Jr., Dennis↗

Accurate Consensus-based Distributed Averaging with Variable Time Delay in Support of Distributed Secondary Control Algorithms

We report that distributed secondary control has been widely used in hierarchical control structures, where multiple distributed generators (DGs) need to coordinate to regulate system voltage and frequency. In these systems, consensus algorithms determine the average of a group of dynamic states (e.g. voltages measured by a group of DGs). To be useful, consensus algorithms must be computationally efficient, stable and accurate. In practice, numerous practical implementation challenges significantly affect the consensus equilibrium. In this paper, we quantify the accuracy deviations of the distributed average observer algorithms proposed in the literature to demonstrate the problems with the state-of-the-art distributed averaging techniques. A novel approach is proposed that achieves accurate average tracking in the presence of time-varying communication delays among agents. In our implementation, time synchronization of all distributed controllers is enabled by a novel software platform, called Resilient Information Architecture Platform for the Smart Grid (RIAPS). The proposed distributed average observer is implemented on hardware controllers and its effectiveness is validated in a controller hardware-in-the-loop testbed.

24 POWER TRANSMISSION AND DISTRIBUTION↗