Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “software resilience”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Achieving American Leadership in Cybersecurity and Digital Components Factsheet

The Biden Administration’s efforts to meet 100% clean electricity by 2035 and net-zero greenhouse gas emissions by 2050 has created opportunities to rebuild American supply chains create new jobs, strengthen community engagement, and spur U.S. economic growth. DOE’s strategy for success in the transition to a clean energy economy hinges on building and maintaining technology supply chains that are advanced, secure, and resilient to cyber threats. As the energy sector grows increasingly globalized, complex, and digitized, the supply chain for digital components of energy systems – including software, virtual platforms and services, and data – is facing greater threats. Nearly all digital components of U.S. energy sector systems are vulnerable to cyber supply chain instability, stemming from a variety of causes and shared among a broad set of interdependent stakeholders. Overall, supply chain risks for digital components in energy sector systems will continue to evolve and likely increase as these systems are increasingly interconnected, digitized, and remotely operated.

CCE↗

Towards Real-Time, On-Board, Hardware-Supported Sensor and Software Health Management for Unmanned Aerial Systems

For unmanned aerial systems (UAS) to be successfully deployed and integrated within the national airspace, it is imperative that they possess the capability to effectively complete their missions without compromising the safety of other aircraft, as well as persons and property on the ground. This necessity creates a natural requirement for UAS that can respond to uncertain environmental conditions and emergent failures in real-time, with robustness and resilience close enough to those of manned systems. We introduce a system that meets this requirement with the design of a real-time onboard system health management (SHM) capability to continuously monitor sensors, software, and hardware components. This system can detect and diagnose failures and violations of safety or performance rules during the flight of a UAS. Our approach to SHM is three-pronged, providing: (1) real-time monitoring of sensor and software signals; (2) signal analysis, preprocessing, and advanced on-the-fly temporal and Bayesian probabilistic fault diagnosis; and (3) an unobtrusive, lightweight, read-only, low-power realization using Field Programmable Gate Arrays (FPGAs) that avoids overburdening limited computing resources or costly re-certification of flight software. We call this approach rt-R2U2, a name derived from its requirements. Our implementation provides a novel approach of combining modular building blocks, integrating responsive runtime monitoring of temporal logic system safety requirements with model-based diagnosis and Bayesian network-based probabilistic analysis. We demonstrate this approach using actual flight data from the NASA Swift UAS.

Unmanned Aerial System↗

Robot Hand

Robots are limited only by the dexterity of the hand. Dr. Salisbury, in conjunction with Stanford, Caltech and Jet Propulsion Laboratory, developed the Salisbury Hand which has three, three-jointed human-like fingers. The tips are covered with a resilient, high friction material for gripping. The robot hand can manipulate objects by finger motion, and adapts to different aims. Advanced software allows the hand to interpret information from fingertip sensors. Further development is expected. A company has been formed to reproduce the device; copies have been delivered to several laboratories.

Source record↗

Machine Learning Based Resilience Testing of an Address Randomization Cyber Defense

Moving target defenses (MTDs) are widely used as an active defense strategy for thwarting cyberattacks on cyber-physical systems by increasing diversity of software and network paths. Recently, machine Learning (ML) and deep Learning (DL) models have been demonstrated to defeat some of the cyber defenses by learning attack detection patterns and defense strategies. It raises concerns about the susceptibility of MTD to ML and DL methods. Here, in this article, we analyze the effectiveness of ML and DL models when it comes to deciphering MTD methods and ultimately evade MTD-based protections in real-time systems. Specifically, we consider a MTD algorithm that periodically randomizes address assignments within the MIL-STD-1553 protocol—a military standard serial data bus. Two ML and DL-based tasks are performed on MIL-STD-1553 protocol to measure the effectiveness of the learning models in deciphering the MTD algorithm: 1) determining whether there is an address assignments change i.e., whether the given system employs a MTD protocol and if it does 2) predicting the future address assignments. The supervised learning models (random forest and k-nearest neighbors) effectively detected the address assignment changes and classified whether the given system is equipped with a specified MTD protocol. On the other hand, the unsupervised learning model (K-means) was significantly less effective. The DL model (long short-term memory) was able to predict the future addresses with varied effectiveness based on MTD algorithm's settings.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Systems Architecture for Fully Autonomous Space Missions

The NASA Goddard Space Flight Center is working to develop a revolutionary new system architecture concept in support of fully autonomous missions. As part of GSFC's contribution to the New Millenium Program (NMP) Space Technology 7 Autonomy and on-Board Processing (ST7-A) Concept Definition Study, the system incorporates the latest commercial Internet and software development ideas and extends them into NASA ground and space segment architectures. The unique challenges facing the exploration of remote and inaccessible locales and the need to incorporate corresponding autonomy technologies within reasonable cost necessitate the re-thinking of traditional mission architectures. A measure of the resiliency of this architecture in its application to a broad range of future autonomy missions will depend on its effectiveness in leveraging from commercial tools developed for the personal computer and Internet markets. Specialized test stations and supporting software come to past as spacecraft take advantage of the extensive tools and research investments of billion-dollar commercial ventures. The projected improvements of the Internet and supporting infrastructure go hand-in-hand with market pressures that provide continuity in research. By taking advantage of consumer-oriented methods and processes, space-flight missions will continue to leverage on investments tailored to provide better services at reduced cost. The application of ground and space segment architectures each based on Local Area Networks (LAN), the use of personal computer-based operating systems, and the execution of activities and operations through a Wide Area Network (Internet) enable a revolution in spacecraft mission formulation, implementation, and flight operations. Hardware and software design, development, integration, test, and flight operations are all tied-in closely to a common thread that enables the smooth transitioning between program phases. The application of commercial software development techniques lays the foundation for delivery of product-oriented flight software modules and models. Software can then be readily applied to support the on-board autonomy required for mission self-management. An on-board intelligent system, based on advanced scripting languages, facilitates the mission autonomy required to offload ground system resources, and enables the spacecraft to manage itself safely through an efficient and effective process of reactive planning, science data acquisition, synthesis, and transmission to the ground. Autonomous ground systems in turn coordinate and support schedule contact times with the spacecraft. Specific autonomy software modules on-board include mission and science planners, instrument and subsystem control, and fault tolerance response software, all residing within a distributed computing environment supported through the flight LAN. Autonomy also requires the minimization of human intervention between users on the ground and the spacecraft, and hence calls for the elimination of the traditional operations control center as a funnel for data manipulation. Basic goal-oriented commands are sent directly from the user to the spacecraft through a distributed internet-based payload operations "center". The ensuing architecture calls for the use of spacecraft as point extensions on the Internet. This paper will detail the system architecture implementation chosen to enable cost-effective autonomous missions with applicability to a broad range of conditions. It will define the structure needed for implementation of such missions, including software and hardware infrastructures. The overall architecture is then laid out as a common thread in the mission life cycle from formulation through implementation and flight operations.

Esper, Jamie↗

Cyber Resiliency and the Implementation of a Host-Based Intrusion Detection System in an Urban Air Mobility Environment

With the growth in Urban Air Mobility systems and the increasing reliance on interconnected technologies, ensuring the security of these complex components has become critical. As cities evolve into smart urban centers, the vulnerability to cyber threats escalates, possibly endangering citizens safety and the efficiency of transportation networks.In response to these challenges, this paper presents a study on the need for cyber resilient techniques within future air traffic environments. It will pay specific attention to the implementation of a Host-Based Intrusion Detection System (HIDS) utilizing Atomic OSSEC software, tailored specifically to a NASA simulation of an UrbanAirMobility environments’ unique demands. Further, this study seeks to outline the rational for NASA’s recommendation for a HIDS in such environments. It explores the design, development, and deployment of the proposed HIDS, focusing on its adaptability to monitor the hybrid nature of the Urban Air Mobility environment. Leveraging machine learning algorithms and anomaly detection techniques, the HIDS is equipped to continuously monitor and analyze the behavior of individual host systems, vehicles, and devices, thereby providing a proactive approach to threat detection. Implementing a HIDS is a pivotal strategy for enhancing cyber resiliency, as it gives an organization granular visibility into internal system activities, enables rapid detection and response to anomalous behavior and cyber threats, and fortifies the organizations overall cybersecurity posture. Finally, this study aims to provide recommendations and include learned takeaways that the Urban Air Mobility industry should consider. In brief, this paper highlights the significance of host-based intrusion detection in UrbanAirMobility environments and underscores the necessity of tailored security solutions to safeguard against emerging cyber threats.

UAM↗

ECP Software Technology Capability Assessment Report

The Exascale Computing Project Software Technology (ECP ST) focus area represents the key bridge between Exascale systems and the scientists developing applications that will run on those platforms. ECP offers a unique opportunity to build a coherent set of software (often referred to as the "software stack") that will allow application developers to maximize their ability to write highly parallel applications, targeting multiple Exascale architectures with runtime environments that will provide high performance and resilience. But applications are only useful if they can provide scientific insight, and the unprecedented data produced by these applications require a complete analysis work ow that includes new technology to scalably collect, reduce, organize, curate, and analyze the data into actionable decisions. This requires approaching scientific computing in a holistic manner, encompassing the entire user workflow - from conception of a problem, setting up the problem with validated inputs, performing high-fidelity simulations, to the application of uncertainty quantification to the final analysis. The software stack plan defined here aims to address all of these needs by extending current technologies to Exascale where possible, by performing the research required to conceive of new approaches necessary to address unique problems where current approaches will not suffice, and by deploying high-quality and robust software products on the platforms developed in the Exascale systems project. The ECP ST portfolio has established a set of interdependent projects that will allow for the research, development, and delivery of a comprehensive software stack,

97 MATHEMATICS AND COMPUTING↗

Intelligent Hierarchical Resilient Operation of Distribution Systems: Implementation and Validation in a Power Hardware-in-the-Loop Simulation Testbed

This paper reports on the structure of a power hardware-in-the-loop (PHIL) simulation testbed that implements, tests, and validates a novel AI-based hierarchical resilient operation model for distribution systems. The testbed implements the central and distributed controllers of the hierarchical resilient operation model and integrates a Digital Real-Time Simulator (DRTS), protective relays, a Real-Time Automation Controller (RTAC), a Software Defined Network (SDN) switch, and a battery energy storage (BES) system. The testbed provides comprehensive real-time visualization and monitoring capability as an advanced situational awareness and operator interface solution. The IEEE 33-node system is used as a test case to test and validate the operation of the model in normal operation and recovery operation after major outages in a fully automated fashion.

Ganjkhani, Mehdi↗

Assessing Energy Infrastructure Devices for Vulnerabilities

Industrial control systems prove to be vital to the health and security of the nation in our critical infrastructure. Critical infrastructure includes the most foundational systems to support modern civilization which includes water and wastewater systems, communications, and the electricity we use to name a few sectors. However, these devices' overall composition remains largely unknown and are untested from a cyber security perspective. As part of the Cyber Testing for Resilient Industrial Control Systems (CyTRICS) program, I analyzed one such energy infrastructure device to better understand how it functions, what hardware and software components are present within it, and assess it for security vulnerabilities. To achieve this, I reverse engineered binary files using Ghidra to understand system functionality and learned more about how to collaborate with other researchers on a shared Ghidra project. I learned more about how web sockets function and how to interact with them through Python to test if they are secure or not. This work led me to assess possible vulnerabilities in this device and provide a better understanding of its composition and function, which are essential to INL's mission of securing our nation's energy infrastructure.

99 - GENERAL AND MISCELLANEOUS↗

Towards Low-Overhead Resilience for Data Parallel Deep Learning

Data parallel techniques have been widely adopted both in academia and industry as a tool to enable scalable training of deep learning models. At scale, DL training jobs can fail due to software or hardware bugs, may need to be preempted or terminated due to unexpected events, or may perform suboptimally because they were misconfigured. Under such circumstances, there is a need to recover and/or reconfigure data-parallel DL training jobs on-the-fly, while minimizing the impact on the accuracy of the DNN model and the runtime overhead. In this regard, state-of-art techniques adopted by the HPC community mostly rely on checkpoint-restart, which inevitably leads to loss of progress, thus increasing the runtime overhead. In this paper we explore alternative techniques that exploit the properties of modern deep learning frameworks (overlapping of gradient averaging and weight updates with local gradient computations through pipeline parallelism) to reduce the overhead of resilience/elasticity. To this end we introduce a failure simulation framework and two resilience strategies (immediate mini-batch rollback and lossy forward recovery), which we study compared with checkpoint-restart approaches in a variety of settings in order to understand the trade-offs between the accuracy loss of the DNN model and the runtime overhead.

data-parallel training↗

Autonomous Energy Systems: Building Reliable, Resilient, and Secure Electrified Communities

Technological changes across energy systems are forcing utilities and operators to reconsider their methods for managing power delivery, but few operators have adopted advanced controls and operational software. Their challenge is that every system has peculiar requirements, and the available solutions are relatively new, untested, and difficult to integrate into an operational environment. Through extensive collaboration with utilities and cooperatives, the National Renewable Energy Laboratory has realized the need for autonomous and optimized management of energy resources, leading to the development of Autonomous Energy Systems, a packaged set of controls that is ready to be integrated into existing control rooms.

automation↗

Online Analytics for Remedy Support at DOE Environmental Management Sites

Environmental data is important for managing environmental restoration/waste site remediation, planning of monitoring efforts, addressing climate resilience, and engaging with stakeholders and regulators. A major challenge is how to manage the many different types and the large volume of environmental data in a way that allows practitioners and site managers to understand data implications and support decisions. The Suite Of Comprehensive Rapid Analysis Tools for Environmental Sites (SOCRATES, https://www.pnnl.gov/projects/socrates) is a web application that provides data access, visualization, and rapid analytics to help make sense of environmental data, support remedy decisions, and communicate information. Development of SOCRATES has been funded through the DOE Richland Operations Office (RL) to support communication and decision making for the Hanford Site, thus is only tied into Hanford environmental data. However, the capabilities of SOCRATES are more broadly applicable to DOE-EM sites engaged in environmental remediation and management. This report describes the work to develop mechanisms for bringing non-Hanford data into SOCRATES so that other DOE-EM sites could make use of the visualization and analysis capabilities to support communication and decision making related to managing environmental restoration/waste site remediation, optimization/exit strategies for pump-and-treat systems, planning monitoring efforts, addressing climate resilience, and/or engaging with stakeholders and regulators. The background, approach, data transfer formats, examples, and next steps for this new SOCRATES-EM software are described in this report.

54 ENVIRONMENTAL SCIENCES↗

Unsupervised Anomaly Detection in High-Dimensional Flight Data Using Convolutional Variational Auto-Encoder

The modern National Airspace System (NAS) is an extremely safe system and the aviation industry has experienced a steady decrease in fatalities over the years. This can be attributed to both improved flight critical systems with redundant hardware and software protections, as well as an increased focus on active monitoring and response to real time and historically identified vulnerabilities by implementing more resilient procedures and protocols. The main approach for identifying vulnerabilities in operations leverages domain expertise using knowledge about how the system should behave within the expected tolerances to known safety margins. This approach works well when the system has a well-defined operating condition. However, the operations in the NAS can be highly complex with various nuances that render it difficult to clearly pre-define all known safety vulnerabilities. With the advancement of data science and machine learning techniques, the potential to automatically identify emerging vulnerabilities in the observed operations has become more practical in recent years. The state-of-the-art anomaly detection approaches in aerospace data usually rely on supervised or semi-supervised learning. However, in many real-world problems such as flight safety, creating labels for the data requires huge amount of effort and is largely impractical. To address this challenge, we developed a Convolutional Variational Auto-Encoder (CVAE), which is an unsupervised learning approach for anomaly detection in high-dimensional heterogeneous time-series data. We validate performance of CVAE compared to the state-of-the-art supervised learning approach as well as unsupervised clustering-based approach using KMeans++ and kernel-based approach using One-Class Support Vector Machine (OC-SVM) on Yahoo!'s benchmark time series anomaly detection data. Finally, we showcase performance of CVAE on a case study of identifying anomalies in the first 60 seconds of commercial flights' take-offs using Flight Operational Quality Assurance (FOQA) data.

Memarzadeh, Milad↗

Unsupervised Anomaly Detection in High-Dimensional Flight Data Using Convolutional Variational Auto-Encoder

The modern National Airspace System (NAS) is an extremely safe system. The industry has experienced a steady decrease in fatalities over the years. This can be contributed to both improved flight critical systems with redundant hardware and software protections as well as an increased focus on active monitoring and response to real time and historically identified vulnerabilities by implementing more resilient procedures and protocols. The main practice for identifying vulnerabilities in operations leverages domain expertise using knowledge about how the system should behave with the expected tolerances to known safety margins. This approach works well when the system has a well-defined operating condition. However, the operations in the NAS can be highly complex with various nuances that render it difficult to clearly pre-define all known safety vulnerabilities. With the advancement of data science and machine learning techniques, the potential to automatically identify emerging vulnerabilities in the observed operations has become more practical in recent years. The state-of-the-art anomaly detection approaches in aerospace data usually rely on supervised or semi-supervised learning. However, in many real-world problems such as flight safety creating labels for the data requires huge amount of efforts and is largely expensive. As a result, in this article, we develop a Convolutional Variational Auto-Encoder (CVAE), an unsupervised learning approach for anomaly detection in high-dimensional heterogeneous time-series data. We validate performance of CVAE compared to the state-of-the-art supervised learning approach (as an upper bound) as well as an supervised clustering based on K-Means (as a lower bound) on Yahoo!'s benchmark time series anomaly detection data. Finally, we showcase performance of CVAE on a case study of identifying anomalies in the first 60 seconds of commercial flights' take-offs using Flight Operational Quality Assurance (FOQA) data.

Milad Memarzadeh↗

Transmission of Images on High-Temperature Nuclear-Grade Metallic Pipe with Ultrasonic Elastic Waves

Transmission of information using elastic ultrasonic waves on existing metallic pipes provides an alternative communication option for a nuclear facility. The advantages of this approach consist of transmitting information through barriers, such as the containment building wall, with minimal modification of the existing hardware. Because bit rates on the order of kilobits per second are achievable, relatively large volumes of data, such as images, can be transmitted. A viable candidate for an ultrasonic communication channel is a stainless steel pipe of the chemical volume control system (CVCS) that penetrates through the reactor containment building wall through a sealed tunnel. To study ultrasonic communication under simulated nuclear facility conditions of high temperature, a test article was developed by installing heating tapes, temperature controllers, and thermal insulation on a laboratory CVCS-like stainless steel pipe. High temperature and radiation-resilient lithium niobate ultrasonic transducers were utilized for information transmission on the heated pipe. The amplitude shift keying (ASK) digital communication protocol was developed and implemented in a GNU Radio software-defined radio environment. A root-raised-cosine filter was introduced to suppress ultrasonic transducer ringing and thus reduce inter-symbol interference. This resulted in the enhancement of the data transmission bit rate compared to information encoding with square pulses. Demonstrations of communication at high temperature included transmission of a 90-KB image at the bit rate of 10 Kbps with a bit error rate of 10 -3 across a 6-ft-long straight pipe heated up to 230 degrees C. Additional preliminary studies were conducted to evaluate ultrasonic communication system resilience to environmental degradation and damage.

42 ENGINEERING↗

Data Center Facility Monitoring with Physics Aware Approach

U.S. Department of Energy's National Renewable Energy Laboratory (NREL) hosts one of the world's most energy-efficient HPC data centers; this system uses component-level warm-water liquid cooling to efficiently remove heat from the data center and capture it for reuse in the building or rejection to the atmosphere. Given the complexity of this system, building data-driven tools for holistically monitoring and operating the entire data center is a priority for ensuring maximal efficiency and resiliency. In this advanced smart facility, over one million metrics are recorded per minute using state-of-the-art streaming data architecture and software to capture and process the state of the system in real time. Here we detail two efforts to effectively analyze, visualize, and interpret this large volume streaming data. We have developed a novel, flexible system for identifying and visualizing individual metric anomalies and component performance across the data center through automatic metadata extraction and physically-motivated visualization for quick interpretation. Additionally, to directly connect system maintenance to data stream processing we explore a physics informed multi-metric drift and anomaly detection application to detect scale-build up in heat exchangers.

anomaly detection↗

Recovering from On-orbit Anomalies on the Astrobee Free Flyers and its Systems

Since 2019, NASA has been operating three Astrobee free flying robots on board the International Space Station (ISS) providing an autonomous and flexible research platform for national and international payload developers in microgravity and serving as a robotic assistant for astronauts on the ISS. During its use on the ISS, in particular with over 750 hours of free-flyer operation as of March 2022, Astrobee and its Docking Station have encountered multiple software and hardware anomalies. These anomalies were either resolved remotely via software and firmware updates, or, where not possible, with hardware replacements on orbit or by the return of the faulty unit to NASA’s ground facilities for its repair. Despite being inherently designed to be repaired or replaced on orbit, Astrobee and its systems can still suffer anomalies that would be complex enough to disassemble, cause risks of hardware damage, or use excessive crew time to perform the repair in orbit. That was the case for the anomaly the Astrobee unit ‘Honey’ encountered, reason why it needed to be down-massed for repair. One of the most common points of failure was found to be the SD card, which is used for the different Astrobee processors and for the Dock Station. Other comparable SD card anomalies were found also on the Astrobee ground units, which provided useful data in the effort of upgrading their systems. This presentation will focus on 1) The overview of the different faults and anomalies on Astrobee and its systems on orbit and on the ground 2) The processes and procedures implemented to resolve the anomalies 3) The implementation of software updates and hardware upgrades in order to reduce the risk on returning anomalies 4) The lessons learned in increasing Astrobee’s robustness and resilience to such anomalies.

International Space Station↗

DINGO: Digital assistant to grid operators for resilience management of power distribution system

With increasing adverse weather events and disasters, enabling resiliency of the power distribution system (PDS) is becoming increasingly important. Here in this work, resiliency is defined as the systems ability to keep supplying critical loads even with multiple contingencies. Resiliency may depend on: (a) advanced tools to assist operators in situational awareness and decision making with the increasing volume of data generated by the PDS, (b) visualization and ease of interaction with system resources and information, especially during extreme events and resulting human operator stress, and (c) flexible resources and autonomous control. Operators and support engineers need to interact with the system for key information and take action under stress, given the requirement for decisions in a short time. Integrated technological solutions are prevailing steps to support the most appropriate decision during critical times to serve essential loads. In order to meet the required goals, a Real-time Resiliency Monitoring and Operational Decision Support (RT-RMOD) tool have been developed. It supports various functionalities, including real-time monitoring, resilience assessment, and proactive decision support. However, this work makes advanced feature additions to the tool by developing data-enabled resilience management algorithms for (i) outage detection and localization, (ii) Resiliency-metric driven restoration and reconfiguration, and (iii) NLP based digital assistant for operators called DINGO (DIgital assistaNt to Grid Operators) to interact with Advanced Distribution Management System (ADMS) and RT-RMOD. The developed algorithm was validated for multiple cases of weather events using a real-world, off-grid microgrid system modeled in a real-time simulator, sensor data, and software tools.

24 POWER TRANSMISSION AND DISTRIBUTION↗