Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hardware failure”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Resilience Design Patterns: A Structured Approach to Resilience at Extreme Scale (V.2.0)

Reliability is a serious concern for future extreme-scale high-performance computing (HPC) systems. Projections based on the current generation of HPC systems and technology roadmaps suggest the prevalence of very high fault rates in future systems. The errors resulting from these faults will propagate and generate various kinds of failures, which may result in outcomes ranging from result corruptions to catastrophic application crashes. Therefore, the resilience challenge for extreme-scale HPC systems requires coordination between various hardware and software technologies that are capable of handling a broad set of fault models at accelerated fault rates. Also, due to practical limits on power consumption in future HPC systems, they are likely to embrace innovative architectures, increasing the levels of hardware and software complexities. Therefore, the techniques that seek to improve resilience must navigate the complex trade-off space between resilience and the overheads to power consumption and performance. While the HPC community has developed various resilience solutions, application-level techniques as well as system-based solutions, the solution space of HPC resilience techniques remains fragmented. There are no formal methods to integrate the various HPC resilience techniques into composite solutions, nor are there methods to holistically evaluate the adequacy and efficacy of such solutions in terms of their protection coverage, and their performance & power efficiency characteristics. Additionally, few implementations of current resilience solutions are portable to newer architectures and software environments that will be deployed on future systems. We developed a new structured approach to the management of HPC resilience using the concept of resilience-based design patterns. In general, a design pattern is a repeatable solution to a commonly occurring problem. We identified the well-known solutions that are commonly used to deal with faults, errors and failures in HPC systems. In the initial design patterns specification (version 1.0), we described the various solutions, which address specific problems in the design of resilient HPC environments, in the form of patterns. Each pattern describes a problem caused by a fault, error or failure event in an HPC environment, and then describes the core of the solution of the problem in such a way that this solution may be adapted to different systems and implemented at different layers of the system stack. The catalog of these resilience design patterns provides designers with a collection of design elements. To construct complete resilience solutions using combinations of various patterns, we defined a framework that enhances HPC designers' understanding of the important constraints and the opportunities for the design patterns to be implemented and deployed at various layers of the system stack. The design framework is also useful for establishing interfaces and mechanisms to coordinate flexible fault management across hardware and software components, as well as to consider the trade-off between performance, resilience, and power consumption when constructing a solution. The resilience design patterns specification version 1.1 included more detailed explanations of the pattern solutions, the context in which the patterns are applicable, and the implications for hardware or software design. It also provided several additional examples and detailed case studies to demonstrate the use of patterns to build realistic solutions. In version 1.2 of the specification document, we have improved the pattern descriptions, including graphical representations of the pattern components. These improvements are largely based on critical comments, feedback and suggestions received from pattern experts and readers of the previous versions of the specification. The pattern classification has been modified to further clarify the relationships between pattern categories. This version of the specification also introduces a pattern language for resilience design patterns. The pattern language presents the patterns in the catalog as a network, revealing the relations among the resilience patterns. The language provides designers with the means to explore alternative techniques for handling a specific fault model that may have different efficiency and complexity characteristics. Using the pattern language also enables the design and implementation of comprehensive resilience solutions as a set of interconnected resilience patterns that can be instantiated across layers of the system stack. The overall goal of this work is to provide hardware and software designers, as well as the users and operators of HPC systems, a systematic methodology for the design and evaluation of resilience technologies in HPC systems that keep scientific applications running to a correct solution in a timely and cost-efficient manner despite frequent faults, errors, and failures of various types. Version 2.0 expands the resilience design pattern classification and catalog to include self-stabilization patterns and reliability, availability and performance models for each structural pattern.

97 MATHEMATICS AND COMPUTING↗

SMC 2021 : Analyzing Resource Utilization and User Behavior on Titan Supercomputer

Resource utilization statistics of submitted jobs on a supercomputer can help us understand how users from various scientific domains use HPC platforms and better design a job scheduler. We explore to generate insight regarding workload distribution and usage pattern domains from job scheduler trace, GPU failure information, and project-specific information collected from Titan supercomputer. Furthermore, we want to know how the scheduler performance varies over time and how the users' scheduling behavior changes following a system failure. These observations have the potential to provide valuable insight, which is helpful to prepare for system failures. These practices will help us develop and apply novel machine learning algorithms in understanding system behavior, requirement, and better scheduling of HPC systems. There are two datasets, RUR and GPU. RUR: This dataset is the job scheduler traces collected from the Titan supercomputerfrom 01/01/2015 to 07/31/2019 (2015.csv - 2019.csv). These were collected usingResource Utilization Report (RUR), a Cray-developed resource-usage data collectionand reporting system. It contains the usage information of its critical resources (CPU,Memory, GPU, and I/O) of each running job on Titan during that period [2]. ProjectAreas: Every job is associated with a project ID. TheProjectAreas.csvdatasetprovides a mapping of the project ID to its domain science. GPU: There have been some hardware-related issues in the GPUs in Titan that caused some GPUs to fail, sometimes irrecoverably during some job runs. This dataset provides information regarding these failures during the execution of the submitted jobs. GPUs on Titan are uniquely identified by a serial number (SN), and they are installed in a location. A GPU can be installed in a location, then removed from that location following a failure, and then re-installed in a different location after fixing the problem. If the failure can't be recovered, the GPU might be removed entirely from Titan. There are two prominent types of failures that resulted in the removal of GPUs from Titan: Double Bit Error (DBE) and Out of the Bus (OTB). The dataset (gc_full.csv) has the following fields: 1. SN : Serial number of a GPU 2. location : The location where it is installed 3. insert : The time when it was inserted into that location 4. remove : The time when it was removed from that location 5. duration : Amount of time the GPU spent in this location 6. out : If the device was taken out entirely w/o a re-installment into a new location. 7. event : If the GPU was taken out entirely, the reason for its removal. To learn more about this dataset, please refer to the git repositoryhttps://github.com/olcf/TitanGPULifeand the related publication [1]. References [1] George Ostrouchov, Don Maxwell, Rizwan A Ashraf, Christian Engelmann, MallikarjunShankar, and James H Rogers. Gpu lifetimes on titan supercomputer: Survival analysisand reliability. InSC20: International Conference for High Performance Computing,Networking, Storage and Analysis, pages 1-14. IEEE, 2020. [2] Feiyi Wang, Sarp Oral, Satyabrata Sen, and Neena Imam. Learning from five-yearresource-utilization data of titan system. In2019 IEEE International Conference onCluster Computing (CLUSTER), pages 1-6. IEEE, 2019.

42 ENGINEERING↗

RDPM: An Extensible Tool for Resilience Design Patterns Modelling

Resilience to faults, errors, and failures in extreme-scale high-performance computing (HPC) systems is a critical challenge. Resilience design patterns offer a new, structured hardware and software design approach for improving resilience. While prior work focused on developing performance, reliability, and availability models for resilience design patterns, this paper extends it by providing a Resilience Design Patterns Modeling (RDPM) tool which allows (1) exploring performance, reliability, and availability of each resilience design pattern, (2) offering customization of parameters to optimize performance, reliability, and availability, and (3) allowing investigation of trade-off models for combining multiple patterns for practical resilience solutions.

Kumar, Mohit↗

A Secondary Control Framework for Microgrid Interoperability With Vendor-Agnostic Grid-Forming Units: Design, Implementation, and Demonstration via Large-Scale Hardware Setup

The reliable operation of islanded microgrids increasingly depends on secondary controls that restore voltage and frequency to nominal values and ensure accurate active and reactive power sharing. Centralized secondary control architectures achieve high accuracy through global coordination at the cost of single-point failures and limited scalability compared with decentralized/distributed approaches. But a critical gap remains in addressing the interoperability and vendor-agnostic operation of secondary controls in real-world microgrids where heterogeneous diesel generator(s) and grid-forming (GFM) inverter(s) from multiple manufacturers always coexist. Practical and vendor-agnostic interoperability guidelines for the secondary control architecture of microgrids with multiple GFM units have not yet been developed; therefore, this paper proposes an interoperable and vendor-agnostic secondary control framework that operates seamlessly across GFM units from different vendors without relying on proprietary controls and protocols, hardware, or lock-ins. The framework leverages existing communication infrastructures (e.g., Modbus TCP/IP) to enable cost-effective deployment while addressing practical challenges, such as packet loss and quantization errors. Mitigation strategies-including data averaging, situational event-triggered control, and finite-iteration execution-are introduced to enhance reliability under real-world conditions. A generalized modeling and design framework is also presented, supported by robustness analysis to demonstrate independence from vendor-specific implementations. The proposed framework is validated through a large-scale hardware demonstration using a 3-$\phi$, 480-V, 60-Hz, 713-kVA laboratory hardware microgrid involving a heterogeneous diesel generator and multiple GFM inverters, showcasing its effectiveness in achieving stable voltage and frequency restoration and accurate power sharing under practical constraints. The results highlight the framework's potential as a scalable and practical solution for next-generation microgrids requiring openness, standard framework, and interoperability.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Impact of Load Tap Changer Control Operation Under Microgrid Conditions

Every microgrid is unique in its diversity of assets, operation, and network structure. This uniqueness presents challenges for each microgrid implementation that need to be understood and addressed to prevent power quality concerns, stability issues, and blackouts. In this research work, we present one challenge faced during microgrid operation that could potentially cause a blackout in some microgrid configurations. Through the incorrect operation of the load tap changer controller during an islanding event for a microgrid with generation present on both primary and secondary busses, the primary voltage will increase or decrease beyond allowable bounds, causing a cascading failure. This work presents the issue first with simulation results to demonstrate the phenomenon and then presents results of load tap changer controller operation using a controller-hardware-in-the-loop experiment. Finally, details of the complete sequence of operation and the blackout scenarios are presented and summarized.

blackout↗

Automated Controller Hardware-In-The-Loop Testbed for EV Charger Resilience Analysis

This paper focuses on the development of a tool that includes an automated testbed with controls, protection, and communications integrated into a real-time system to provide a platform to generate data sets for failure modes and effects analysis. This tool establishes a value for automation of data generation for different scenarios and addresses the gap of nonexistent field data for different applications and use cases. The features of this tool can further be expanded to include multiple power electronics models, communication protocols, and scaled system architectures. This general framework was evaluated for a DC fast charger system use case to provide quantitative solution for resiliency.

Starke, Michael↗

Distributed Grid Control of Flexible Loads and DERs for Optimized Provision of Synthetic Regulating Reserves

Over the course of this project, we have successfully de-risked our distributed microgrid control architecture by tightly integrating its associated control algorithms into a unified software library, installing the software library on several industrial-grade target hardware platforms, and validating the performance of the resulting microgrid controller in a real-life microgrid. Upon completion of the project, we demonstrated that our distributed control architecture is resilient against (i) failures in control devices, (ii) unreliable communication links, (iii) delays in transmitted data, and (iv) imperfect knowledge of the number of (and state of) generation and load assets in the microgrid. In this final report, we present results from all the milestones that were accomplished over the course of this project.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Demonstrating Distribution System Resiliency through Grid-Edge Microgrids, on a Multi-Site Networked Hardware-in-Loop Platform

With the increasing penetration of Distributed Energy Resources (DERs) at the grid-edge, power systems include more energy storage, remote switches, relays, voltage regulators, and other intelligent electronic devices (IED). Effective control of these grid-edge devices by using Advanced Distribution Management Systems (ADMS) can yield substantial improvements to the resiliency and power quality of distribution systems. In this paper, improvements to the resiliency of a distribution system are demonstrated using a multi-site evaluation environment consisting of a real-time Hardware-in-Loop (HIL) setup in which DERs and other IEDs are modeled; and an ADMS which monitors and is able to control the distribution system assets. The HIL model and the ADMS are located 2400 km away, with communication between the sites enabled by a data manager using Distributed Network Protocol 3 (DNP3), demonstrating the system's capabilities even over long distances. After a simulated transmission system failure in the HIL demonstration setup, DERs and other devices are operated to restore critical loads and node voltage profile (to within the 'nominal +/-5%' band) in the distribution system.

ADMS↗

SWARM: Reimagining scientific workflow management systems in a distributed world

Modern scientific workflows process massive amounts of data from diverse instruments and sensors, leveraging geographically distributed, heterogeneous compute and storage resources—from leadership-class systems to edge devices—connected by high-performance networks. The diversity of resources introduces challenges in harnessing their full potential, with resilience issues arising across applications, system software, networks, storage, and hardware. Today, workflow management systems (WMS) coordinate the execution of computation and data management tasks across target resources. However, WMS’s centralized nature makes them vulnerable to faults and scalability issues that may result in failures of entire computational campaigns. In conclusion, this paper introduces a novel agentic framework for workflow management, fully distributing and decentralizing the WMS functions and modeling them as swarm intelligence agents infused with advanced artificial intelligence solutions and traditional distributed computing algorithms that can make coordinated decisions in the presence of failures of the underlying cyberinfrastructure.

Swarm intelligence↗

Next-Generation, High-temperature, High-frequency, High-efficiency, High-power-density Traction System

To meet performance and reliability requirements necessary for broader adoption of electric drive vehicles, the Electrical and Electronics Technical Team of the U.S. Drive partnership has established aggressive design goals for next-generation electric vehicle drivetrains. Specifically, the 2025 roadmap stipulates a 100 kW/L power density target and a $\$$2.7/kW cost target for power electronics, in addition to high-voltage operation (i.e., greater than 800 VDC). The additional targets for traction motor and the overall system performance impose further challenges on the power electronics design. For example, many high specific power machines have reduced iron content, and therefore reduced intrinsic filtering, thus requiring the inverter to supply a low-distortion drive current. These machines also typically have a high pole count, thus requiring drive current at a higher electrical frequency. Other motors, such as brush-less dc and switch reluctance machines, require a carefully-shaped, non-sinusoidal drive current (Yang, Shang, Brown, & Krishnamurthy, 2015), (Zhang, Bowman, O'Connel, & Haran, 2018), (Anderson, et al., 2018). Two- and three-level inverter topologies are the conventional framework for the power electronics design of the drivetrain, and some demonstrations have shown recent progress towards addressing cost, power density and efficiency goals (Gurpinar & Ozpineci, 2018), (Zhu, Kim, Chen, Erickson, & Maksimović, 2018), (Deshpande, Chen, Narayanasamy, Sathyanarayanan, & Luo, 2018), (Alizadeh, et al., 2019). However, an unconventional approach may be necessary to take the dramatic leap in power density necessitated by the roadmap—while simultaneously addressing the other system needs. Therefore, this project leverages the flying capacitor multilevel (FCML) topology, together with a scalable, modular approach, to address these needs. This type of hybrid converter has several advantages: lower voltage (i.e., less than 300 V) transistors can be used, energy-dense capacitors process most of the power, and the output current waveform is multilevel and exhibits a frequency multiplying effect—in other words, the output has reduced dv/dt and filtering requirements for the same high voltage dc bus. For example, in an electric vehicle with an 800 V bus, a 10-level FCML could leverage 100 V, commercially available GaN devices switching at 115 kHz to produce a ~1 MHz switching waveform (modulated according to the motor drive requirements) with one ninth of the dv/dt of a two-level converter. Prior work has already demonstrated promising performance and gravimetric power density figures for more electric aircraft applications (Pallo, Foulkes, Modeer, Coday, & Pilawa-Podgurski, 2018). This project leverages lessons learned to achieve the volumetric power density of 100 kW/L by employing advanced liquid cooling, address the 300,000 mile reliability challenge with redundant design, topology failure studies and online health monitoring, and reduce costs to $\$$2.7/kW through the use of low-cost GaN devices, modular converter assemblies, and modest modifications to traditional manufacturing methods. The project involved several hardware designs, each achieving increasing performance. At the conclusion of the project, a volumetric power density of 380 kW/L was achieved, in a 800V dc-ac converter, greatly surpassing even the aggressive target goal.

33 ADVANCED PROPULSION SYSTEMS↗

Design of a Robust Memristive Spiking Neuromorphic System with Unsupervised Learning in Hardware

Spiking neural networks (SNN) offer a power efficient, biologically plausible learning paradigm by encoding information into spikes. The discovery of the memristor has accelerated the progress of spiking neuromorphic systems, as the intrinsic plasticity of the device makes it an ideal candidate to mimic a biological synapse. Despite providing a nanoscale form factor, non-volatility, and low-power operation, memristors suffer from device-level non-idealities, which impact system-level performance. To address these issues, this article presents a memristive crossbar-based neuromorphic system using unsupervised learning with twin-memristor synapses, fully digital pulse width modulated spike-timing-dependent plasticity, and homeostasis neurons. Additionally, the implemented single-layer SNN was applied to a pattern-recognition task of classifying handwritten-digits. The performance of the system was analyzed by varying design parameters such as number of training epochs, neurons, and capacitors. Furthermore, the impact of memristor device non-idealities, such as device-switching mismatch, aging, failure, and process variations, were investigated and the resilience of the proposed system was demonstrated.

97 MATHEMATICS AND COMPUTING↗

Efficient Implementation of Artificial Neural Networks for Sensor Data Analysis Based on a Genetic Algorithm

The reliability of many industrial processes depends on the sensor system. However, these sensors can be affected by noise, perturbations and failures. Hence, sensor monitoring and diagnosis are fundamental to guarantee the quality of an industrial process. Nowadays, artificial neural networks (ANN) are widely used in sensor signal processing and diagnosis. However, those ANNs usually require many artificial neurons, being difficult to implement in software and hardware due to their high computational costs. This paper presents an optimized implementation of artificial neurons in ANNs for sensor data analysis using a Genetic Algorithm (GA). The objective of GA is to find an adequate segmentation to reduce the activation function approximation error. One of the advantages of the proposed approach is that the cost function used in GA considers the effect of factors such as the ANN architecture or the number of bits used in arithmetic operations. The proposed ANN implementation technique aims to get the best possible approximation for a specific ANN architecture, making easier its implementation in software and hardware. Simulation and experimental results using FPGA (Field Programmable Gate Array) prove the advantages of the proposed approach for implementing sensor data analysis systems based on ANNs.

D estefani, André↗

Relief Zones Enhance the Durability of Ultrathin Membranes in Electrochemical Conversion Devices

Premature failures in electrochemical conversion systems often result when membrane electrode assemblies (MEAs) use ultrathin (≤15 μm-thick) polymer electrolyte membranes, susceptible to mechanical degradation from stress concentrations arising from device-level integration. Herein, relief zones were developed to mitigate mechanical degradation by alleviating excess and nonuniform compression across active areas. Relief zones, created through ablation of carbonaceous diffusion media, enable seamless adaptation across MEA dimensions without need for hardware modifications. Demonstrated using fuel cells as a case study, accelerated stress tests revealed a 6-fold lifetime improvement (∼1500 h) compared to conventional edge-protected MEAs, decoupling device-level engineering effects from material limitations.

accelerated stress test↗

Programmable Digital Devices used in Advanced Reactors

This paper introduces the concepts of common cause failure, diversity, and defense-in-depth used by the nuclear industry to analyze resilience in reactors. A survey of publicly traded and private companies building advanced reactors and their licensing status is presented. Safety and non-safety systems found in the NuScale Power design are summarized and the likely hardware and software categories used by those systems are enumerated. The importance of industry partners is highlighted. This paper also identifies an alternate path forward without industry partners to advance the knowledge needed to use artificial intelligence to analyze HBOMs and SBOMs to better understand reactor resiliency.

cybersecurity↗

An ontology-based fault generation and fault propagation analysis approach for safety-critical computer systems at the design stage

Abstract Fault propagation analysis is a process used to determine the consequences of faults residing in a computer system. A typical computer system consists of diverse components (e.g., electronic and software components), thus, the faults contained in these components tend to possess diverse characteristics. How to describe and model such diverse faults, and further determine fault propagation through different components are challenging problems to be addressed in the fault propagation analysis. This paper proposes an ontology-based approach, which is an integrated method allowing for the generation, injection, and propagation through inference of diverse faults at an early stage of the design of a computer system. The results generated by the proposed framework can verify system robustness and identify safety and reliability risks with limited design level information. In this paper, we propose an ontological framework and its application to analyze an example safety-critical computer system. The analysis result shows that the proposed framework is capable of inferring fault propagation paths through software and hardware components and is effective in predicting the impact of faults.

97 MATHEMATICS AND COMPUTING↗

Transmission and Distribution Real-Time Analysis Software for Monitoring and Control: Design and Simulation Testing

The US electric grid is facing operational, stability, and security challenges. Transmission system operators need some measure of visibility into distribution system renewable generation. Distribution system generation needs to support transmission system voltage. The grid is experiencing an expansion in measurement systems. How to take full advantage of this expansion and defend against attacks, both cyber and physical, poses additional challenges. This paper introduces software designed to meet these challenges. At the center of the software is an Integrated System Model (ISM) that spans from transmission to secondary distribution. The ISM is employed in real-time abnormality detection, voltage stability forecasting, and multi-mode control. The software architecture along with selected analysis modules is presented. Testing results are presented for: 1—attacks on utility infrastructure; 2—energy savings from optimal control; 3—distribution system control response during a low voltage transmission system event; 4—cyber-attacks on PV inverters, where physical inverters are used in hardware-in-the-simulation-loop studies. Contributions of this work include real-time analysis that spans from three-phase transmission through secondary distribution; an approach for detecting abnormalities that employs measurements from three independent measurement systems; and a multi-mode distribution system control that responds to cyber-attacks, physical attacks, equipment failures, and transmission system needs.

14 SOLAR ENERGY↗

Predicting Wind Loading and Instability in Solar Tracking PV Arrays

Wind loading and the fluctuating pressure loads it creates on PV panel surfaces are associated with multiple degradation mechanisms and failures. Modest wind speeds create reversing loads that can initiate cell cracks and weather cracked cells. Stronger wind speeds and extreme weather events can lead to larger scale forces and the aerodynamic instability known as torsional galloping. All these effects are dependent on the complex coupling between wind speed, panel orientation, and a myriad of other hardware and site-specific factors. In this work, we present the latest developments from our work to build an open-source, high-performance computing (HPC) fluid dynamics solver to predict and mitigate these effects. This simulation package allows users to easily specify different array layouts, solar-tracking angles, panel geometries, and weather conditions before automatically generating a refined computational mesh and solving for the unsteady loading on each panel surface. Small domains (e.g., a single panel row in isolation) can be solved on a modern laptop, while larger domains or very high-fidelity studies can be solved on distributed or HPC resources with minimal modifications to the underlying problem specification. We present preliminary case studies obtained using this simulation package and highlight how increased wind speeds combined with sub-optimal tracking angles can exacerbate degradation drivers.

aerodynamics↗