Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hardware failure”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Investigation of the Use of Dynamic Probabilistic Risk Assessment Methodologies for Identifying Digital I&C System Common Cause Failures

Digital Instrumentation and Control (I&C) systems have a key role in nuclear power plants in the upgrade of aging analog systems. Digital systems improve plant safety and reliability through features such as increased hardware reliability and stability and improved failure detection capability. There is no consensus on which of the current probabilistic risk assessment methods are most suitable for use in the reliability analysis of digital I&C systems. While the traditional event-tree/fault-tree (ET/FT) approach is still used for their reliability modeling, there are concerns regarding this approach in properly accounting for dynamic interactions among system components since potentially significant dependencies among failure events may not be identified and/or their likelihood may not be properly quantified. Dynamic methodologies are expected to provide a much more accurate representation of probabilistic evolution of the I&C systems in time due to their capability to more properly account for complex interactions than the static approach. The applicability of dynamic PRA methodologies for digital I&C system is investigated using the criteria presented in the NUREG/CR-6901, and the comparisons made in NUREG/CR-6901 are updated in light of the latest studies. The Dynamic Event Tree (DET) approach has been identified as one of the top dynamic methods when evaluated against the requirements for the reliability modeling of digital I&C systems. The DET method is a strong candidate for integration into existing PRA studies, as it bears many similarities to the traditional ET approach. In this study, the DET approach has been applied to the Plant Protection System of the APR1400 design, and the results are compared to results from its available traditional ET/FT analysis. Possible approaches to evaluate and quantify the effects of common cause failures on system safety using dynamic methods are also examined.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Balance of Plant Modeling and Real-Time Hardware-in-the-Loop Integration with the Microreactor Automated Control System

The advent of novel microreactor technology has driven a focused effort to explore safety and efficiency improvements that can be achieved through the use of automated system control. Development of control strategies, especially for initial demonstration, requires an adequate surrogate environment to safely research failure modes and control integration with realistic hardware delay. However, efficiency gains from control strategies are improved when the scope of controller action is expanded to include system-level dynamics such as downstream heat extraction and mass flow. For this reason, a balance-of-plant (BOP) model of a representative microreactor system has been developed using the TRANsient Simulation Framework of Reconfigurable Models library in Modelica. This model captures a reactor and primary NaK coolant loop that represent corresponding system components of the Microreactor Applications Research Validation and EvaLuation (MARVEL) design as well as a secondary coolant loop and heat extraction representative of the Microreactor Agile Non-Nuclear Experimental Test Bed (MAGNET). This model configuration allows for hardware-in-the-loop (HIL) integration with microreactor automated control system (MACS) hardware in real time through a Python-based gRPC client. Real-time simulation of model performance with emulated hardware and communication delay suggests that under independent proportional-integral-derivative control of BOP model drum dynamics and downstream heat extraction, stable power load following is achievable. A slight delay in load following, filtering of high-frequency dynamics, and localized temperature fluctation suggest room for improvement through the development of higher-level control strategies. The simulated coupling of the MAGNET facility lays the groundwork for future digital twin analysis with a coupled MACS-MAGNET HIL demonstration.

McConnell, Jono [ORNL] (ORCID:0000000238984741)↗

SMC 2021 Data Challenge: Analyzing Resource Utilization and User Behavior on Titan Supercomputer

Resource utilization statistics of submitted jobs on a supercomputer can help us understand how users from various scientific domains use HPC platforms and better design a job scheduler. We explore to generate insight regarding workload distribution and usage pattern domains from job scheduler trace, GPU failure information, and project-specific information collected from Titan supercomputer. Furthermore, we want to know how the scheduler performance varies over time and how the users' scheduling behavior changes following a system failure. These observations have the potential to provide valuable insight, which is helpful to prepare for system failures. These practices will help us develop and apply novel machine learning algorithms in understanding system behavior, requirement, and better scheduling of HPC systems. There are two datasets, RUR and GPU: RUR dataset is the job scheduler traces collected from the Titan supercomputer from 01/01/2015 to 07/31/2019 (2015.csv - 2019.csv). These were collected using resource Utilization Report (RUR), a Cray-developed resource-usage data collection and reporting system. It contains the usage information of its critical resources (CPU, Memory, GPU, and I/O) of each running job on Titan during that period (https://ieeexplore.ieee.org/abstract/document/8891001). It includes ProjectAreas as additional information, every job is associated with a project ID. TheProjectAreas.csv dataset provides a mapping of the project ID to its domain science. GPU dataset has information regarding GPU failure on Titan. There have been some hardware-related issues in the GPUs in Titan that caused some GPUs to fail, sometimes irrecoverably during some job runs. This dataset provides information regarding these failures during the execution of the submitted jobs. GPUs on Titan are uniquely identified by a serial number (SN), and they are installed in a location. A GPU can be installed in a location, then removed from that location following a failure, and then re-installed in a different location after fixing the problem. If the failure can't be recovered, the GPU might be removed entirely from Titan. There are two prominent types of failures that resulted in the removal of GPUs from Titan: Double Bit Error (DBE) and Out of the Bus (OTB). The dataset (gc_full.csv) has seven attributes, we provided a short description of these attributes in the ReadMe file. To learn more about this dataset, please refer to the git repository https://github.com/olcf/TitanGPULife and the related publication (https://ieeexplore.ieee.org/abstract/document/9355319).

42 ENGINEERING↗

Hardware Aware Mitigation of Timing Side-Channel Vulnerabilities in Critical Infrastructure Software

Program runtime/timing attacks exploit variations in a program’s execution times to extract sensitive information from the program (e.g. encryption keys, sensitive variable data, intellectual property). State-of-the-art solutions to runtime sidechannel attacks attempt to balance the execution time of the sensitive code for different control flow paths to eliminate the timing leakage. However, during the mitigation process, most techniques do not consider the underlying hardware/device on which the target program is supposed to run on. This can lead to over-fixing (unnecessary extra operations), under-fixing (not solving the imbalance properly), and even failures. We propose DISARM, a joint hardware-software methodology (unlike any existing solution) for mitigating runtime side-channel vulnerabilities that utilizes timing values from real embedded devices to generate targeted software fixes. We implement DISARM to support C/C++/Java source codes and validate it across 22 standard benchmarks. DISARM outperforms state-of-the-art solutions such as PENDULUM and DifFuzzAR in terms of execution time overhead (up to −46%), code size overhead (up to −10%), and correctness (no failures) on five different embedded/edge devices.

Suha, Tasneem [University of Maine]↗

DISARM: Target Electronic Device Informed Mitigation of Software Runtime Side-Channel Vulnerabilities

Program runtime/timing attacks exploit variations in a program’s execution times to extract sensitive information from the program (e.g. encryption keys, sensitive variable data, intellectual property). State-of-the-art solutions to runtime side-channel attacks attempt to balance the execution time of the sensitive code for different control flow paths to eliminate the timing leakage. However, during the mitigation process, most techniques do not consider the underlying hardware/device on which the target program is supposed to run on. This can lead to over-fixing (unnecessary extra operations), under-fixing (not solving the imbalance properly), and even failures. Here, we propose DISARM, a joint hardware-software methodology (unlike any existing solution) for mitigating runtime side-channel vulnerabilities that utilizes timing values from real embedded devices to generate targeted software fixes. We implement DISARM to support C/C++/Java source codes and validate it across 22 standard benchmarks. DISARM outperforms state-of-the-art solutions such as PENDULUM and DifFuzzaR in terms of execution time overhead, code size overhead, and correctness on five different embedded/edge devices.

Timing/runtime side-channel↗

Evaluation of Hardware and Software Bill of Materials (HBOMs/SBOMs) Extraction Methods

Hardware and software bills of materials (HBOMs and SBOMs) provide important visibility into the components, dependencies, and supply chain relationships within programmable digital devices. This visibility is critical for advanced nuclear reactor applications, where use of common or shared hardware components, software libraries, suppliers, or manufacturing processes may create common cause failure (CCF) vulnerabilities despite apparent diversity. This paper evaluates current approaches for obtaining and analyzing HBOMs and SBOMs in support of CCF, diversity and defense-in-depth (D3) assessments, and begins to explore potential methods for artificial intelligence/machine learning-based analysis. The availability of BOM information from advanced reactor manufacturers and vendors, representative hardware and software categories found in advanced reactor systems continues to limit research [13]. This paper compares commonly used BOM formats, including CycloneDX, SPDX, and SWID. It also surveys publicly available tools for generating BOMs from source code, compiled binaries, and hardware-related information, noting limitations in language coverage, system age, and format interoperability. Finally, this paper evaluates methods for correlating BOM data with vulnerability and exploitability information, including VEX, CVE, and CWE resources. The findings indicate that publicly available nuclear-vendor BOMs are limited, making third-party extraction and research into novel analysis techniques necessary.

Cybersecurity↗

Seismic Resilience of Large Power Transformer Bushings & Non-SF6 Industrial Base Scan Review

Large (high voltage) power transformers (LPT), and more specifically, their bushings, are known to be susceptible to seismic failure. With bushing failure, a transformer will have to be replaced, which has a considerable lead time, adding to the power outage duration. Cost-efficient, proven solutions are not currently available to mitigate this risk, which can persist for the more than 30-year life of a particular transformer. This work will focus on developing and demonstrating a hardware solution to address seismic vulnerabilities and reduce outage risks from LPT failure. Sulfur Hexafluoride (SF6) is a specialty gas with excellent electrical insulation properties which has been used extensively in the power industry. This gas is unfortunately also one of the most potent greenhouse gases known to humanity. A 2014 report by the Intergovernmental Panel on Climate Change found that SF6 has a global warming potential (GWP) 23,000 times higher than Carbon Dioxide, and has the highest GWP of all gases assessed (Myhre 2013). SF6 is almost exclusively man-made and is produced for use as an insulator in high voltage electrical equipment. This makes the production and use of SF6 one of the leading sources of anthropogenic climate change. To fully eliminate the environmental impacts of SF6, alternative technology is needed. The ideal replacement would be a technology that can fulfil the same role as SF6, at the same cost or cheaper, but without adverse environmental effects. Currently, no technology fits this description, however several promising technologies have begun to enter the market. An industry scan was performed to assess the state of industry adoption and manufacturing capability for SF6-free alternative technologies for use at the high-voltage level, and the primary barriers to broader adoption.

10 SYNTHETIC FUELS↗

A Fast, Accurate Prediction for System-Wide Damage Due to Dynamic Wind Loading

The complex relationship between photovoltaic (PV) hardware configurations, overall system dynamics, and turbulent aerodynamic phenomena generates highly unsteady, non-uniform loads that can lead to damaging instabilities. These effects may result in glass breakage, cell cracking, and structural failures in frames and mounting systems, even under moderate wind conditions. Addressing industry concerns about premature system failures in field conditions deemed survivable, our research aims to develop a fast and accurate predictive model for system damage. This model integrates configurable hardware choices with advanced simulation tools to represent the overall system-specific dynamics effectively. Using this model, we predict responses under varying weather conditions and hardware setups, translating these predictions into pre-trained surrogate models capable of accurately identifying failure risks and rapidly testing new system hardening measures. In this presentation, we will showcase preliminary results in capturing system dynamics through our customizable library of PV hardware configurations. Additionally, we will highlight how these new tools build upon PVade's established wind load modeling capabilities and foster the development of advanced AI/ML surrogates for improving system robustness.

97 MATHEMATICS AND COMPUTING↗

Upgrade and Operation of the ATLAS Radiation Interlock System (ARIS)

ATLAS (the Argonne Tandem Linac Accelerator Sys-tem) is a superconducting heavy ion accelerator which can accelerate nearly all stable, and some unstable, iso-topes between hydrogen and uranium. Prompt radiation fields from gamma and or neutron are typically below 1 rem/hr at 30 cm, but are permitted up to 300 rem/hr at 30 cm. The original ATLAS Radiation Interlock System (ARIS), hereafter referred to as ARIS 1.0 was installed 30 years ago. While it has been a functional critical safe-ty system, its age has exposed the facility to high risk of temporary shutdown due to failure of obsolete compo-nents. Topics discussed will be architecture, hardware improvements, functional improvements, and operation permitting personnel access to areas with low levels of radiation.

Blomberg, B. R.↗

HAZARD ANALYSIS OF DIGITAL ENGINEERED SAFETY FEATURES ACTUATION SYSTEM IN ADVANCED NUCLEAR POWER PLANTS USING A REDUNDANCY-GUIDED APPROACH

Replacing the existing aging analog instrumentation and control (I&C) systems with modern safety control and protection digital technology offers one of the foremost means of performance improvements and cost reductions for the existing nuclear power plants (NPPs). However, the qualification of digital I&C systems remains a challenge, especially considering the issue of software common-cause failures (CCFs), which are difficult to address. With the application and upgrades of advanced digital I&C systems, software CCFs have become a potential threat to plant safety because most redundant designs use similar digital platforms or software in the operating and application systems. With complex designs of multilayer redundancy to meet the single-failure criterion, digital I&C safety systems (e.g., engineered safety-features actuation system [ESFAS]) are of a particular concern in the U.S. Nuclear Regulatory Commission (NRC) licensing procedures. This paper applies a modularized approach to conduct redundancy-guided systems-theoretic hazard analysis for an advanced digital ESFAS with multilevel redundancy designs. Systematic methods and risk-informed tools are incorporated to address both hardware and software CCFs, which provide guidance to eliminate the triggers of potential single points of failure in the design of digital safety systems in advanced plant designs.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Fault localization in a microfabricated surface ion trap using diamond nitrogen-vacancy center magnetometry

Here, as quantum computing hardware becomes more complex with ongoing design innovations and growing capabilities, the quantum computing community needs increasingly powerful techniques for fabrication failure root-cause analysis. This is especially true for trapped-ion quantum computing. As trapped-ion quantum computing aims to scale to thousands of ions, the electrode numbers are growing to several hundred, with likely integrated photonic components also adding to the electrical and fabrication complexity, making faults even harder to locate. In this work, we used a high-resolution quantum magnetic imaging technique, based on nitrogen-vacancy centers in diamond, to investigate short-circuit faults in an ion trap chip. We imaged currents from these short-circuit faults to ground and compared them to intentionally created faults, finding that the root cause of the faults was failures in the on-chip trench capacitors. This work, where we exploited the performance advantages of a quantum magnetic sensing technique to troubleshoot a piece of quantum computing hardware, is a unique example of the evolving synergy between emerging quantum technologies to achieve capabilities that were previously inaccessible.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Scattering phase shift in quantum mechanics on quantum computers

Here, we investigate the feasibility of extracting infinite volume scattering phase shift on quantum computers in a simple one-dimensional quantum mechanical model, using the formalism established in the work by Guo and Gasparian [Phys. Rev. D 108, 074504 (2023)] that relates the integrated correlation functions for a trapped system to the infinite volume scattering phase shifts through a weighted integral. The system is first discretized in a finite box with periodic boundary conditions, and the formalism in real time is verified by employing a contact interaction potential with exact solutions. Quantum circuits are then designed and constructed to implement the formalism on current quantum computing architectures. To overcome the fast oscillatory behavior of the integrated correlation functions in real-time simulation, different methods of postdata analysis are proposed and discussed. Test results on IBM hardware show that good agreement can be achieved with two qubits, but complete failure ensues with three qubits due to two-qubit gate operation errors and thermal relaxation errors.

Guo, Peng [Dakota State Univ., Madison, SD (United↗

Lifetime extension of legacy CEBAF LLRF hardware

A significant portion of the Low-Level Radio Frequency (LLRF) hardware in Jefferson Lab’s CEBAF is from the original construction of the facility using 1980’s CAMAC technology. Of the fifty-three zones in CEBAF, thirty-six of them are legacy hardware. The age of the legacy system has led to difficulties in maintaining the hardware due to parts going obsolete without suitable drop in replacements. Continued operation of the legacy system is required as the installation of LLRF 3.0 systems is costly and cannot be completed in a short period of time with the available resources. The most pressing failure in the legacy system was a failing buffer card, which is responsible for communication between the EPICs network and individual RF control modules. A new buffer card was designed as a transparent, drop in, replacement so that upgrades are simply a matter of swapping the existing legacy hardware. This buffer card upgrades a single point failure component and promises to extend the operable lifetime of CEBAF’s legacy systems.

Accelerator Physics↗

Microreactor Automated Control System Test Bed Digital Architecture for Real-Time, Hardware-in-the-Loop Simulation

This work describes progress made towards the development of a real-time hardware-in-the-loop (HIL) test bed for non-nuclear testing of microreactor control schemes and failure modes. Non-nuclear testing is a crucial step in developing robust control algorithms for managing microreactor dynamics. The creation of an HIL simulation harnesses the realistic dynamics of physical analogue systems while additionally considering the challenges of variable communication delay. This collaborative effort between Oak Ridge National Laboratory and Idaho National Laboratory has resulted in a LabVIEW-based gRPC communication protocol which couples a TRANSFORM Modelica simulation of nuclear components to the ViBRANT physical hardware for realistic feedback and visual representation of control action in real time. A modular python client structure is developed to manage FMU-based Modelica simulation and real-time gRPC communication. HIL testing suggests that the modeled reactor with natural convection molten salt loop coolant configuration responds well to PID control of drum positioning for modulation of reactor core power, however, future efforts will be made to explore the added thermal inertial delay of system level control and downstream demand changes. Development of this platform with a generalized methodology provides a foundation for exploring a variety of reactor configurations and failure modes in rapid order to provide insight into the most effective avenues of study for further research and development.

McConnell, Jono [ORNL] (ORCID:0000000238984741)↗

Hazard analysis for identifying common cause failures of digital safety systems using a redundancy-guided systems-theoretic approach

Replacing the existing aging analog instrumentation and control (I&C) systems with modern safety control and protection, digital technology offers one of the foremost means of performance improvements and cost reductions for the existing nuclear power plants (NPPs). However, the qualification of digital I&C systems remains a challenge, especially considering the issue of software common-cause failures (CCFs), which are difficult to address. With the application and upgrades of advanced digital I&C systems, software CCFs have become a potential threat to plant safety because most redundant designs use similar digital platforms or software in the operating and application systems. With complex designs of multilayer redundancy to meet the single-failure criterion, digital I&C safety systems (e.g., engineered safety-features actuation system [ESFAS]) are of a particular concern in the U.S. Nuclear Regulatory Commission (NRC) licensing procedures. Here, this paper applies a modularized approach to conduct redundancy-guided systems-theoretic hazard analysis for an advanced digital ESFAS with multilevel redundancy designs. Systematic methods and risk-informed tools are incorporated to address both hardware and software CCFs, which provide guidance to eliminate the causal factors of potential single points of failure in the design of digital safety systems in advanced plant designs.

42 ENGINEERING↗

Reliability Improvement by Fault-Tolerant Operation of NPC Inverter for Motor Driving

The high-reliability operation of three-level inverters is crucial to prevent equipment damage, process downtime, and economic losses. This article investigates a three-level neutral-point-clamped inverter under all possible combinations of open-circuit and short-circuit faults and proposes a new postfault operation method without adding extra hardware. This method provides comprehensive solution for operations after single and multiple-device failures, increasing the inverter reliability by 24%. This article classified postfault modulations and uncovered previously unknown fault scenarios that can be addressed using the proposed control method. A new postfault modulation based on space vector modulation with virtual vectors is proposed. The feasibility of the proposed control method is verified by simulation and an experiment for one fault scenario in a three-level neutral-point-clamped inverter with a 3.73 kW motor load. Furthermore, this article contributes to improving the reliability of three-level inverters.

42 ENGINEERING↗

HDBind: encoding of molecular structure with hyperdimensional binary representations

Traditional methods for identifying “hit” molecules from a large collection of potential drug-like candidates rely on biophysical theory to compute approximations to the Gibbs free energy of the binding interaction between the drug and its protein target. These approaches have a significant limitation in that they require exceptional computing capabilities for even relatively small collections of molecules. Increasingly large and complex state-of-the-art deep learning approaches have gained popularity with the promise to improve the productivity of drug design, notorious for its numerous failures. However, as deep learning models increase in their size and complexity, their acceleration at the hardware level becomes more challenging. Hyperdimensional Computing (HDC) has recently gained attention in the computer hardware community due to its algorithmic simplicity relative to deep learning approaches. The HDC learning paradigm, which represents data with high-dimension binary vectors, allows the use of low-precision binary vector arithmetic to create models of the data that can be learned without the need for the gradient-based optimization required in many conventional machine learning and deep learning methods. This algorithmic simplicity allows for acceleration in hardware that has been previously demonstrated in a range of application areas (computer vision, bioinformatics, mass spectrometery, remote sensing, edge devices, etc.). To the best of our knowledge, our work is the first to consider HDC for the task of fast and efficient screening of modern drug-like compound libraries. We also propose the first HDC graph-based encoding methods for molecular data, demonstrating consistent and substantial improvement over previous work. We compare our approaches to alternative approaches on the well-studied MoleculeNet dataset and the recently proposed LIT-PCBA dataset derived from high quality PubChem assays. We demonstrate our methods on multiple target hardware platforms, including Graphics Processing Units (GPUs) and Field Programmable Gate Arrays (FPGAs), showing at least an order of magnitude improvement in energy efficiency versus even our smallest neural network baseline model with a single hidden layer. Our work thus motivates further investigation into molecular representation learning to develop ultra-efficient pre-screening tools. We make our code publicly available at https://github.com/LLNL/hdbind.

59 BASIC BIOLOGICAL SCIENCES↗

Maturing Rational Design Methodologies and Industry Consensus Engineering Standards: Critical Fastened Joints - Solar PV Industry

Critical structural joints can be seen throughout a solar array and are called upon to secure modules and keep racking assembled and able to resist large demands from winds and snow loads. In the relatively new and fast-growing solar PV industry, the important role these hardware assemblies (e.g. clips, clamps, bolts, nuts, washers) play is not well understood by product designers. Failures with critical structural joints are surprisingly common and point to the need for maturing the engineering and assembly of these joints. The wide variety of design concepts (Figure 2&2) demonstrate interesting and innovative ideas but are lacking the basics of fastener engineering seen in matured industries (e.g. transportation, buildings). Complicating the maturing process for critical structural joints is that they are one component in rack supporting structures that exhibits a systems behavior; each component will affect the other and play a key role in maintaining structural integrity. When wind loads the surface of a module, the underlying racking members deflect and twist which in turn imparts forces back into the joints and into the mounted modules. Often, these supporting rack structures exhibit high deflections and low natural frequencies which amplify the demands placed into the joints even in moderate winds. Current engineering practices and associated structural conventions view solar racking support structures as they would a high mass building that exhibit more static behaviors in wind events. Solar structures are unique from high mass buildings and require the development of solar specific industry engineering consensus standards.

14 SOLAR ENERGY↗