Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hardware failure”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

CHESS 2025: Waveform LiDAR data from NEON AOP surveys

This dataset provides Level 1 (L1) full-waveform light detection and ranging (LiDAR) data collected for the 2025 Colorado Headwaters Ecological Spectroscopy Study (CHESS). These data were acquired to enable characterization of vegetation structure and other three-dimensional features of the land surface, and to evaluate structural changes that may have occurred between a prior LiDAR acquisition in 2018 and the 2025 overflight. Waveform LiDAR data can provide more detailed information about objects on the ground than discrete point clouds typically do, and they are often used for granular target segmentation and characterization of subcanopy vegetation. The data were acquired over three study domains in the Upper Gunnison river basin: the upper East River watershed (CRBU); Almont Triangle and Taylor Canyon (ALMO); and Upper Taylor River watershed (UPTA) between 2025-06-13 and 2025-07-15. LiDAR data were acquired using the Optech Galaxy Prime Airborne LiDAR Terrain Mapper onboard the National Ecological Observatory Network (NEON) Airborne Observation Platform (AOP). These are the primary waveform LiDAR data delivered by NEON and are provided per flightline in compressed Pulsewaves format, an open-source binary file standard. A Pulsewaves object comprises a two files: a pulse (.pls) file, which stores the geographic origin, outgoing vector, and metadata for every laser pulse emitted by the scanner, and a wave file (.wvs), which stores the sequential amplitude samples of the outgoing pulse and the returning signals. The files are published here in their compressed forms (.plz, .wvz). All waveform data were processed following the theoretical workflow described in the NEON L0-to-L1 Waveform LiDAR Algorithm Theoretical Basis Document (Krause and Goulden 2022a); however, the Pulsewaves output format differs from a legacy format described in that document. Waveform amplitude samples are recorded at 1 nanosecond intervals. All coordinates are provided in meters. Horizontal coordinates are referenced in Universal Transverse Mercator (UTM) zone 13N and the World Geodetic System (WGS) 1984 ensemble datum. Elevations are referenced to Geoid12A. Waveform data for the UPTA survey area were collected without incident and the published records are complete. However, both the ALMO and CRBU collections experienced issues that resulted in incomplete data for those areas. On collection day 2018-06-16 a hardware failure caused the waveform digitizer to lose data from the eastern edge of the ALMO site (Figure 22). The waveform data for flightlines 2–20 could not be extracted from the digitizer, and the data proved unrecoverable. As a result, a portion of the site does not have coverage with waveform data. Although no hardware failure was observed during collection over the CRBU area, final waveform files generated by vendor software contained only ~25% of the expected number of return pulses. After discovery, NEON initiated troubleshooting with the vendor. The root cause of the data ablation had not been identified at the time of publication. Additional data will be published in an update to this package if further recovery proves successful. CHESS Project Description: The Colorado Headwaters Ecological Spectroscopy Study (CHESS) comprised a multi-week airborne remote sensing and field observation campaign in the Upper Gunnison Basin, Colorado, conducted in June and July of 2025. Airborne remote sensing was conducted by the National Ecological Observatory Network Airborne Observation Platform (NEON AOP), concurrent with a field campaign run by the Rocky Mountain Biological Laboratory (RMBL), the Lawrence Berkeley National Laboratory (LBNL) and SLAC National Accelerator Laboratory Watershed Function Science Focus Area (SFA), and NASA-JPL (Jet Propulsion Laboratory) Earth Surface Mineral Dust Source Investigation (EMIT) program. Between June 10 and July 18, 2025, the NEON AOP flight team collected high-resolution aerial imaging spectroscopy and Light Detection and Ranging (LiDAR) data over three domains: the Upper East River (CRBU), Almont Triangle (ALMO), and the Upper Taylor Basin (UPTA). In coordination with the flights, a field campaign acquired ground-truth observations, including observations of vegetation composition, foliar traits, forest demography, and subsurface properties in 18 core sampling areas within the domains. Additional surface water observations were taken at over 380 point locations. All CHESS campaign datasets can be found within the CHESS ESS-DIVE data portal: https://data.ess-dive.lbl.gov/portals/chess. Funding Acknowledgement: Field and remote-sensing data acquisition was performed under a grant from the National Aeronautics and Space Administration (80NSSC24K1005). This work was also supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

2018 NEON and 2025 CHESS Campaigns↗

Resilience and Robustness of Spiking Neural Networks for Neuromorphic Systems

Though robustness and resilience are commonly quoted as features of neuromorphic computing systems, the expected performance of neuromorphic systems in the face of hardware failures is not clear. In this work, we study the effect of failures on the performance of four different training algo-rithms for spiking neural networks on neuromorphic systems: two back-propagation-based training approaches (Whetstone and SLAYER), a liquid state machine or reservoir computing approach, and an evolutionary optimization-based approach (EONS). We show that these four different approaches have very different resilience characteristics with respect to simulated hardware failures. We then analyze an approach for training more resilient spiking neural networks using the evolutionary optimization approach. We show how this approach produces more resilient networks and discuss how it can be extended to other spiking neural network training approaches as well.

Schuman, Catherine↗

Deep learning model to detect various synchrophasor data anomalies

High-density synchrophasors provide valuable information for power grid situational awareness, operation and control. Unfortunately, due to factors including communication instability and hardware failure, their data quality can be greatly deteriorated by anomalies. Since the anomalies can impact the performance of the synchrophasor applications, it is of paramount significance to propose a model to detect anomalies in synchrophasor. In this study, a convolutional neural network model is established to detect and classify the anomalies in the synchrophasor measurements. Additionally, four types of anomalies observed in actual synchrophasors including erroneous patterns, random spikes, missing points and high-frequency interferences are considered in this study. The proposed model is extensively evaluated via field-collected measurements from the synchrophasor network in Jiangsu grid, China. The superior performance of the proposed model indicates the great potential of using deep learning for the detection of abnormal synchrophasor measurements.

42 ENGINEERING↗

INTEGRATED RISK ASSESSMENT OF DIGITAL I&C SAFETY SYSTEMS FOR NUCLEAR POWER PLANTS

Upgrading the existing analog instrumentation and control (I&C) systems to state-of-the-art digital I&C (DI&C) systems provides the foremost means to improve performance and reduce costs for existing light-water reactors (LWRs). However, qualification of digital technologies remains a challenge?especially the issue of software common cause failure (CCF), which has been difficult to address. Existing analyses of CCFs in I&C systems mainly focus on hardware failures. With the application and upgrading of new DI&C systems, software CCFs due to design flaws might become a potential threat to plant safety, considering that most redundancy designs use similar digital platforms or software in their operating and application systems. With complex multi-layer redundancy designs to meet the single failure criterion, these I&C safety systems are of particular concern in U.S. Nuclear Regulatory Commission (NRC) licensing procedures. In 2019, the Risk-Informed Systems Analysis (RISA) Pathway of the U.S. Department of Energy?s (DOE?s) Light Water Reactor Sustainability (LWRS) Program initiated a project to develop a risk assessment strategy for delivering a strong technical basis to support effective, licensable, secure DI&C technologies for digital upgrades/designs. An integrated risk assessment for the DI&C (RADIC) process was proposed for this strategy to identify potential key digital-induced failures, implement reliability analyses of related digital safety I&C systems, and evaluate the unanalyzed sequences introduced by these failures (particularly software CCFs) at the plant level. This paper summarizes these RISA efforts in the risk analysis of safety-related DI& systems at Idaho National Laboratory.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

An Integrated Risk Assessment Process of Safety-Related Digital I&C Systems in Nuclear Power Plants

Upgrading the existing analog instrumentation and control (I&C) systems to state-of-the-art digital I&C (DI&C) systems will greatly benefit existing light water reactors. However, the issue of software common cause failure (CCF) remains an obstacle in terms of qualification for digital technologies. Existing analyses of CCFs in I&C systems mainly focus on hardware failures. With the application and upgrading of new DI&C systems, design flaws could cause software CCFs to become a potential threat to plant safety, considering that most redundancy designs use similar digital platforms or software in their operating and application systems. With complex multilayer redundancy designs to meet the single failure criterion, these I&C safety systems are of particular concern in U.S. Nuclear Regulatory Commission licensing procedures. In Fiscal Year 2019, the Risk-Informed Systems Analysis (RISA) Pathway of the U.S. Department of Energy’s Light Water Reactor Sustainability Program initiated a project to develop a risk assessment strategy for delivering a strong technical basis to support effective, licensable, and secure DI&C technologies for digital upgrades and designs. An integrated risk assessment for the DI&C process was proposed for this strategy to identify potential key digital-induced failures, implement reliability analyses of related digital safety I&C systems, and evaluate the unanalyzed sequences introduced by these failures (particularly software CCFs) at the plant level. Here this paper summarizes these RISA efforts in the risk analysis of safety-related DI&C systems at Idaho National Laboratory.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Asynchronous distributed-memory task-parallel algorithm for compressible flows on unstructured 3D Eulerian grids

Here, we discuss the implementation of a finite element method, used to numerically solve the Euler equations of compressible flows, using an asynchronous runtime system (RTS). The algorithm is implemented for distributed-memory machines, using stationary unstructured 3D meshes, combining data-, and task-parallelism on top of the Charm++ RTS. Charm++’s execution model is asynchronous by default, allowing arbitrary overlap of computation and communication. Task-parallelism allows scheduling parts of an algorithm independently of, or dependent on, each other. Built-in automatic load balancing enables continuous redistribution of computational load by migration of work units based on real-time CPU load measurement. The RTS also features automatic checkpointing, fault tolerance, resilience against hardware failure, and supports power-, and energy-aware computation. We demonstrate scalability up to 25 x 10 9 cells at $\mathscr{O}$10 4 compute cores and the benefits of automatic load balancing for irregular workloads. The full source code with documentation is available at https://quinoacomputing.org.

42 ENGINEERING↗

Anomaly Detection in Accelerator Facilities Using Machine Learning

Synchrotron light sources are user facilities and usually run about 5000 hours per year to support many beamlines operations in parallel. Reliability is a key parameter to evaluate machine performance. Even many facilities have achieved >95% beam reliability, there are still many hours of unscheduled downtime and every hour lost is a waste of operation costs along with a big impact on individual scheduled user experiments. Preventive maintenance on subsystems and quick recovery from machine trips are the basic strategies to achieve high reliability, which heavily depends on experts’ dedication. Recently, SLAC, APS, and NSLS-II collaborated to develop machine-learning-based approaches aiming to solve both situations, hardware failure prediction and machine failure diagnosis to find the root sources. In this paper, we report our facility operation status, development progress, and plans.

Accelerator Physics↗

ECE 4396 (Final Report)

In the summer of 2025, I was fortunate enough intern at Sandia National Laboratories in Albuquerque, New Mexico. I was hired into the Southwest Analysis Laboratories for Semiconductor Advancement (SALSA) intern program. In this internship, I applied my knowledge and skills in electrical engineering to conduct hardware failure analysis. I utilized various failure analyze techniques involving the use of Infrared Thermography (IRT) and Laser Scanning Microscopy (LSM) to test different Application-Specific Integrated Circuits (ASIC) chips that are available in the public market.

42 ENGINEERING↗

Analysis of Real-World Preignition Data Using Neural Networks

Increasing adoption of downsized, boosted, spark-ignition engines has improved vehicle fuel economy, and continued improvement is desirable to reduce carbon emissions in the near-term. However, this strategy is limited by damaging preignition events which can cause hardware failure. Research to date has shed light on various contributing factors related to fuel and lubricant properties as well as calibration strategies, but the causal factors behind an individual preignition cycle remain elusive. If actionable precursors could be identified, mitigation through active control strategies would be possible. This paper uses artificial neural networks to search for identifiable precursors in the cylinder pressure data from a large real-world data set containing many preignition cycles. It is found that while follow-up preignition cycles in clusters can be readily predicted, the initial preignition cycle is not predictable based on features of the cylinder pressure. Further, this indicates that the alternating pattern of preignition cycles within clusters is influenced by the thermodynamic state as reflected in the pressure, but that the trigger for the initial preignition cycle is not thermodynamic in nature, but more likely tied to a critical threshold in the chemistry of the fuel/lubricant mixture in the upper crevice or other factors related to the presence of an ignition source.

33 ADVANCED PROPULSION SYSTEMS↗

Assessing the Impact of Mirror Technology on Driver Perception and Safety: Traditional vs. Camera-Based Systems

Camera-based mirror systems (CBMS) are being adopted by commercial fleets based on the potential improvements to operational efficiency through improved aerodynamics, resulting in better fuel economy, improved maneuverability, and the potential improvement for overall safety. Until CBMS are widely adopted it will be expected that drivers will be required to adapt to both conventional glass mirrors and CBMS which could have potential impact on the safety and performance of the driver when moving between vehicles with and without CBMS. To understand the potential impact to driver perception and safety, along with other human factors related to CBMS, laboratory testing was performed to understand the impact of CBMS and conventional glass mirrors. Drivers were subjected to various, nominal driving scenarios using a truck equipped with conventional glass mirrors, CBMS, and both glass mirrors and CBMS, to observe the differences in metrics such as head and eye movement, reaction time, and perception of distance. The finds from this study will serve as the baseline measurements for future research regarding off-nominal driving scenarios and hardware failures of CBMS, as well as inform potential future policy regarding CBMS for the use in commercial vehicles in lieu of conventional glass mirrors.

Siekmann, Adam [ORNL] (ORCID:0000000284653935)↗

Replacement of Legacy Analytical Codes at the Advanced Test Reactor

For each operating cycle of the Advanced Test Reactor (ATR) at Idaho National Laboratory, a Core Safety Assurance Package (CSAP) is necessary to demonstrate compliance with the safety basis approved by the United State Department of Energy (DOE). Certain computer codes are used in CSAP development, most of them developed in-house. This work describes replacement of a large set of these codes and updates previous work. Replacement of legacy codes is necessary due to computer hardware failure but also has generally improved user-friendliness.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Replacement of Legacy Analytical Codes at the Advanced Test Reactor

For each operating cycle of the Advanced Test Reactor (ATR) at Idaho National Laboratory, a Core Safety Assurance Package (CSAP) is necessary to demonstrate compliance with the safety basis approved by the United State Department of Energy (DOE). Certain computer codes are used in CSAP development, most of them developed in-house. This work describes replacement of a large set of these codes and updates previous work. Replacement of legacy codes is necessary due to computer hardware failure but also has generally improved user-friendliness.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Analyzing Hardware and Software Common Cause Failures in Digital Instrumentation and Control Systems using Dual Error Propagation Method

This paper develops a methodology for quantifying software common cause failures (CCFs) in digital instrumentation and control (I&C) systems of nuclear power plants. To support the transition of analog I&C systems to digital in nuclear power plants, probabilistic risk assessment (PRA) techniques are used. The hardware components of the I&C systems have reliability databases that can be used in the PRA studies. However, the failure data for redundant software components of the systems is sparse. Failure of components constitutes a CCF, wherein two or more components or systems fail due to a single shared cause and coupling mechanism. This paper proposes a quantification approach that can simultaneously model hardware and software components, incorporate the CCFs of software systems in the models, and bridge the gap between the failure quantification of models and the development of CCF parametric databases. We demonstrate the dual error propagation method (DEPM) by developing I&C systems failure models for a representative digital reactor trip system. The DEPM models are built to simulate the control and data flows within the systems and can accommodate failure states. By expanding DEPM to software CCFs, we generated alpha factor parameter estimates for each of the modeled error propagation mechanisms.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Resilience and fault tolerance in high-performance computing for numerical weather and climate prediction

Progress in numerical weather and climate prediction accuracy greatly depends on the growth of the available computing power. As the number of cores in top computing facilities pushes into the millions, increased average frequency of hardware and software failures forces users to review their algorithms and systems in order to protect simulations from breakdown. This report surveys hardware, application-level and algorithm-level resilience approaches of particular relevance to time-critical numerical weather and climate prediction systems. A selection of applicable existing strategies is analysed, featuring interpolation-restart and compressed checkpointing for the numerical schemes, in-memory checkpointing, user-level failure mitigation and backup-based methods for the systems. Numerical examples showcase the performance of the techniques in addressing faults, with particular emphasis on iterative solvers for linear systems, a staple of atmospheric fluid flow solvers. The potential impact of these strategies is discussed in relation to current development of numerical weather prediction algorithms and systems towards the exascale. Trade-offs between performance, efficiency and effectiveness of resiliency strategies are analysed and some recommendations outlined for future developments.

54 ENVIRONMENTAL SCIENCES↗

Modeling interconnections of safety and financial performance of nuclear power plants, part 3: Spatiotemporal probabilistic physics-of-failure analysis and its connection to safety and financial performance

Here, this paper is a byproduct of a line of research by the authors to analyze interrelationships of safety and financial performance of nuclear power plants (NPPs). The result of this line of research is summarized in three parts: Part 1 covers a categorical review of relevant literature and the theoretical bases that support the methodological developments in Part 2. Part 2 introduces an Integrated Enterprise Risk Management (I-ERM) methodological framework to quantify the interconnections of safety and financial performance with a focus on operation and maintenance (O&M) of NPPs. Part 2 has also demonstrated the applicability and values of the I-ERM methodology through an NPP case study. This paper is Part 3, where detailed development and implementation of one of the I-ERM modules, i.e., probabilistic physics-of-failure (PPoF) analysis, and its connection with safety and financial performance is reported. In this article, the physical failure modeling for hardware components is advanced by incorporating finite element analysis (FEA) into PPoF analysis and coupling the FEA-based PPoF with the maintenance performance through a renewal process model. This article covers two scientific contributions: (i) first-of-its-kind incorporation of FEA into the PPoF model of thermal fatigue for NPP components; and (ii) advancing the interface between the PPoF analysis and the renewal process model in order to deal with spatiotemporal FEA outputs and to efficiently estimate the physical transition rates even when the PPoF outputs are dominated by success data. Through the incorporation of FEA, the resolution of the PPoF analysis is enhanced as spatiotemporal conditions such as stress and temperature can be considered explicitly instead of relying on simplified assumptions or analytical models with reduced spatiotemporal dimensions. To demonstrate an application of the FEA-based PPoF analysis and its coupling with maintenance through the renewal process model, a case study is conducted using excess letdown elbow piping in the chemical and volume control system of a Pressurized Water Reactor.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

CalWave Open Water Demo - FMEA Update Budget Period 2

The Failure Modes and Effects Analysis (FMEA) is a qualitative reliability technique for systematically analyzing each possible failure mode within a hardware system, and identifying the resulting effect on that system, the mission, and the personnel. This submission includes an updated FMEA summary for CalWave's open water demonstration including pre- and post-mitigation results, hazard identification (HAZID) analysis, and component/function rooted FMEA.

16 TIDAL AND WAVE POWER↗

Complete time-resolved X-ray shape capability for a 10 MJ implosion: Milestone Report for MRT 8818

This report documents completion of MRT 8818, which established an upgraded equatorial time-resolved X-ray imaging capability for high-yield inertial confinement fusion experiments at the National Ignition Facility (NIF) using the equatorial Dilation X-ray Imager (DIXI). The need for this work arose from the sustained increase in fusion yield at NIF, which progressively intensified the neutron and gamma-ray environment experienced by target diagnostics. Although the existing polar imaging capability, particularly the Polar Dilation X-ray Imager (PDIXI), remained viable at high yields, the equatorial capability became inadequate because of radiation-induced failure of electronic readout hardware and increasing background levels that degraded image quality and ultimately caused saturation. The work performed under MRT 8818 addressed this limitation through a combination of hardware replacement, background characterization, and targeted mitigation. The DIXI backend was converted from CCD-based electronic readout to photographic film, and the principal sources of internally generated background were investigated using prior analyses, dedicated tests, and comparison with PDIXI performance on high-yield shots. This effort led to implementation of two principal improvements: a multilayer optical coating to suppress broadband radiation-induced background generated in the fiber-optic extension cylinder, and replacement of Kodak TMAX 400 film with Agfa Copex Rapid to reduce background generated directly in the film. These upgrades were completed in 2025 and performance assessment based on DT shot data since then, comparing with Polar DIXI and scaling of these experiments to 10 MJ indicates an expected signal-to-background ratio of approximately 34 for a single pinhole image, substantially exceeding the MRT 8818 completion criterion of SNR > 5.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Investigation of the Use of Dynamic Probabilistic Risk Assessment Methodologies for Identifying Digital I&C System Common Cause Failures

Digital Instrumentation and Control (I&C) systems have a key role in nuclear power plants in the upgrade of aging analog systems. Digital systems improve plant safety and reliability through features such as increased hardware reliability and stability and improved failure detection capability. There is no consensus on which of the current probabilistic risk assessment methods are most suitable for use in the reliability analysis of digital I&C systems. While the traditional event-tree/fault-tree (ET/FT) approach is still used for their reliability modeling, there are concerns regarding this approach in properly accounting for dynamic interactions among system components since potentially significant dependencies among failure events may not be identified and/or their likelihood may not be properly quantified. Dynamic methodologies are expected to provide a much more accurate representation of probabilistic evolution of the I&C systems in time due to their capability to more properly account for complex interactions than the static approach. The applicability of dynamic PRA methodologies for digital I&C system is investigated using the criteria presented in the NUREG/CR-6901, and the comparisons made in NUREG/CR-6901 are updated in light of the latest studies. The Dynamic Event Tree (DET) approach has been identified as one of the top dynamic methods when evaluated against the requirements for the reliability modeling of digital I&C systems. The DET method is a strong candidate for integration into existing PRA studies, as it bears many similarities to the traditional ET approach. In this study, the DET approach has been applied to the Plant Protection System of the APR1400 design, and the results are compared to results from its available traditional ET/FT analysis. Possible approaches to evaluate and quantify the effects of common cause failures on system safety using dynamic methods are also examined.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗