Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hardware failure”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

A hardware-in-the-loop (HIL) testbed for cyber-physical energy systems in smart commercial buildings

In recent years, there has been a growing trend toward the development of smart buildings that rely on cyber-physical systems (CPS) to optimize occupant comfort, safety, and energy efficiency. To ensure the reliable and efficient operation of CPS with designed control strategies, it is important to evaluate their performance under various scenarios before deploying them in the real world. This is where a Hardware-in-the-loop (HIL) testbed designed for studying sensor and control-related studies in smart buildings can be highly valuable. With the growing threat of cyber-attacks and physical faults targeting smart buildings, it is essential to ensure the security of building operations. A HIL testbed can emulate cyber-attack and physical fault scenarios, allowing researchers to develop and test threat detection and mitigation algorithms. This enables researchers to identify potential issues and optimize the algorithms in a safe and controlled environment before they are deployed in real-world settings, reducing the risk of failures that can negatively impact occupant comfort, safety, and energy efficiency. Therefore, this paper developed a HIL testbed designed for cyber-physical energy systems (e.g. buildings automation system (BAS)) in smart commercial buildings. The HIL testbed is comprised of a real-time building and Heating, Ventilation, and Air-Conditioning (HVAC) emulator using Modelica-based dynamic models, a set of BAS controllers, and a BAS computer server. The data generation capability of the HIL testbed is demonstrated by tracking normal and faulty operating data in the BAS, as well as monitoring detailed network traffic in the local BAS network. Here, this study further demonstrates the HIL testbed’s capability by conducting case studies on real-time physical fault and cyber-attack experiments using a Department of Energy (DOE) prototype commercial building. It is anticipated that the fully functional HIL testbed will be utilized for a variety of sensor and control-related studies, including but not limited to testing, developing, validating of different HVAC control strategies, fault detection & diagnosis, energy monitoring and analysis, cyber security study, etc.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Automated, reliable, and efficient continental-scale replication of 7.3 petabytes of computational simulation data: A case study

We report on our experiences replicating 7.3 petabytes (PB) of Earth System Grid Federation (ESGF) computational simulation data from Lawrence Livermore National Laboratory (LLNL) in California to Argonne National Laboratory (ANL) in Illinois and Oak Ridge National Laboratory (ORNL) in Tennessee—a task motivated by a need for increased reliability, capacity, and performance. This task presented significant challenges: the need to move 29 million files twice under time pressure from aging storage hardware; a source file system bottleneck limiting throughput to 1.5 GB/s; frequent site maintenance windows; and the need for complete reliability at scale. We addressed these challenges using a simple replication tool that invoked Globus to transfer large bundles of files while tracking progress in a database, dynamically rerouting transfers to work around maintenance periods and file system limitations. Under the covers, Globus organized transfers to make efficient use of the high-speed Energy Sciences network (ESnet) and the data transfer nodes deployed at participating sites, and also addressed security, integrity checking, and recovery from a variety of transient failures. This success demonstrates the considerable benefits that can accrue from the adoption of performant data replication infrastructure. The replication tool is available at https://github.com/esgf2-us/data-replication-tools.

Globus↗

Dynamic Probabilistic Safety Assessment Studies for Advanced Reactor Using RAVEN

Probabilistic Safety Assessment (PSA) is used extensively to evaluate the risks associated with complex engineering systems like Nuclear Power Plants (NPPs). Current PSA models are based on the Event-Tree/Fault-Tree (ET/FT) methodology. ET and FT models are static and are based on Boolean logic approaches. In the past, concerns have been raised in the literature regarding the capability of the traditional static modelling approaches to adequately account for the impact of process, hardware, software, firmware and human interactions on the stochastic system behaviour. To overcome the limitations of the traditional approach to PSA, several dynamic PSA methodologies have been proposed. One of the dynamic PSA methodologies used for dynamic evaluations is Dynamic Event Tree (DET) framework which can be used to assess the impact of the parameter variability and scenario dynamics on the PSA model for the initiating event. The DET framework couples the stochastic model (number of component/trains that start on demand, operator action timing, etc.) with a Thermal-Hydraulic (TH) model of the plant. This paper explores the use of DET along with a case study on advanced reactor. The initiating event selected for the study was Class IV power supply failure event. The TH analysis considering uncertainty in various parameters was performed using RELAP5 and Reactor Analysis and Virtual control ENvironment (RAVEN) tool. Based on the uncertainty analysis, it is concluded that the peak clad temperatures (PCT) are within the limits in all the code runs implying a high-degree of safety margin. However, variation in time to reach the PCT was observed among the code runs and the mean time to reach the PCT was found to be around 8590sec (approximately 2.4 hours). Hence, sufficient time margin is available for human intervention and the operator might have a relatively stress-free state during such an accident scenario. Due to the static nature of the traditional PSA models, the safety margin available was lesser, whereas, with the help of dynamic PSA models, one can demonstrate that the actual available safety margin is more in the present case study and is valuable input from the design point of view.

99 GENERAL AND MISCELLANEOUS↗

Artificial Intelligence Thermostat to Detect Faults

Residential air conditioners and heat pumps often experience faults due to inadequate maintenance, which can severely reduce efficiency or even cause system failure. Common issues include dirty or clogged air filters and refrigerant leaks. These problems degrade performance and increase energy use and operating costs. This study presents a smart thermostat with embedded artificial intelligence to detect such faults and alert homeowners when maintenance is needed. The thermostat uses low-cost measurements—including return-air temperature, relative humidity, supply-air temperature, outdoor-air temperature, and condenser subcooling—to identify abnormal operations. Because different faults produce distinct response patterns, tailored algorithms are developed to recognize characteristic fault signatures. The investigation is built on a detailed co-simulation platform that couples EnergyPlus with the DOE/ORNL Heat Pump Design Model (HPDM). EnergyPlus represents the building’s dynamic environment, while HPDM is a high-fidelity, hardware-based model that can simulate fault-free performance as well as a wide range of faults, including gradual degradation such as minor refrigerant leakage. This platform provides a virtual training and testing environment that helps distinguish fault-induced behavior from normal operation and supports development of robust diagnostic algorithms. Using this framework, a Dynamic Bayesian Network was developed to identify two common faults—gradual refrigerant charge loss and indoor airflow blockage—and the AI-embedded thermostat was verified through annual building simulations.

Shen, Bo [ORNL] (ORCID:0000000336600393)↗

Towards Low-Overhead Resilience for Data Parallel Deep Learning

Data parallel techniques have been widely adopted both in academia and industry as a tool to enable scalable training of deep learning models. At scale, DL training jobs can fail due to software or hardware bugs, may need to be preempted or terminated due to unexpected events, or may perform suboptimally because they were misconfigured. Under such circumstances, there is a need to recover and/or reconfigure data-parallel DL training jobs on-the-fly, while minimizing the impact on the accuracy of the DNN model and the runtime overhead. In this regard, state-of-art techniques adopted by the HPC community mostly rely on checkpoint-restart, which inevitably leads to loss of progress, thus increasing the runtime overhead. In this paper we explore alternative techniques that exploit the properties of modern deep learning frameworks (overlapping of gradient averaging and weight updates with local gradient computations through pipeline parallelism) to reduce the overhead of resilience/elasticity. To this end we introduce a failure simulation framework and two resilience strategies (immediate mini-batch rollback and lossy forward recovery), which we study compared with checkpoint-restart approaches in a variety of settings in order to understand the trade-offs between the accuracy loss of the DNN model and the runtime overhead.

data-parallel training↗

Adaptive Cybersecurity for Distributed Energy Resources (AdCyDER): Online Reinforcement Learning with Stackelberg-Optimized Defenses — Pipeline Architecture, Evaluation Methodology, and Findings from a Synthetic-Data Evaluation

This report documents the design and evaluation of an integrated online-learning pipeline developed within the AdCyDER project for Distributed Energy Resource (DER) cybersecurity. The pipeline couples a Reinforcement Learning (RL) attack classifier — which produces an attack-type probability distribution — with a Stackelberg game-theoretic (GT) defense selector that consumes those distributions alongside SME-encoded priors over (defense, attack) effectiveness pairings and perdefense costs to choose grid-health-preserving defenses. The objective is not attack classification per se but production of distributions that drive effective defense selection through the Stackelberg layer, learned from delayed grid-health feedback rather than labeled attack data. AdCyDER as a whole is broader than the work presented here; this report covers the specific RL/GT loop integration and its evaluation. We present the integrated pipeline (SCADA telemetry with Fronius inverter physics, Suricata IDS, time-windowed aggregation, per-facility LSTM classifier, Stackelberg optimizer, OpenC2 actuators), an experimental campaign of 28 eight-hour iterations across three baseline modes, and a pipeline-ordered diagnostic protocol. The protocol identifies two distinct failure modes within the loop: paired supervised ceilings on the same features establish that the deployed online RL classifier (macro F1 ≈ 0.07) sits at least 4.7× below a same-architecture supervised LSTM (≈ 0.34) and 10–11× below a linear feature-signal ceiling (≈ 0.70–0.79 depending on per-facility isolation), localizing the dominant failure to the training procedure; and the reward signal driving online updates carries weak directional coupling with classifier correctness in the methodology-expected direction (multi-lens convergent: top-decile P(true) records produce more frequent state changes and slightly larger improvements, top-vs-bot Cohen’s 𝑑 ≈ −0.19), but at effect magnitudes too small to drive gradient-based learning at the campaign sample size. The original learning hypothesis is not supported by the data. The primary contributions are the diagnostic methodology — proposed as a transferable falsification protocol for online RL/GT defense pipelines learning from delayed environmental reward — and the open, reproducible experimental infrastructure. We outline reward reformulation as the highest-priority aspirational next step given the underpowered-but-aligned Q6 reading, with hardware-in-the-loop evaluation as the broadest scope-expansion option.

Blakely, Benjamin [Argonne National Laboratory (AN↗

Evaluation of Centralized Model Based FLISR in a Lab Setup

Utilities are installing advanced distribution management systems (ADMS) around the globe to improve the sensing and control of distribution systems. ADMS is becoming a critical component to improve the resiliency and reliability of distribution systems. These management systems host a multitude of applications that can be used to sense, control, and operate distribution systems. Fault location, isolation, and service restoration (FLISR) is one of the ADMS applications that is critical to improving resilience during fault conditions. FLISR applications use smart, controllable devices that are installed in the distribution system for FLISR operation. These controllable devices include distributed energy resources (DER) and reclosers. Evaluating such ADMS applications before installation in the field can help de-risk the field implementation and avoid costly failures in the field. This paper presents a background on experimental setup that are typically used to evaluate ADMS applications. This is followed by briefly presenting the setup used to evaluate the FLISR application in an off-the-shelf ADMS tool. Finally, results from the evaluation experiments are presented.

94 GMLC - Grid Modernization Laboratory Consortium↗

Evaluation of Centralized Model-Based FLISR in a Lab Setup: Preprint

Utilities are installing advanced distribution management systems (ADMS) around the globe to improve the sensing and control of distribution systems. ADMS is becoming a critical component to improve the resiliency and reliability of distribution systems. These management systems host a multitude of applications that can be used to sense, control, and operate distribution systems. Fault location, isolation, and service restoration (FLISR) is one of the ADMS applications that is critical to improving resilience during fault conditions. FLISR applications use smart, controllable devices that are installed in the distribution system for FLISR operation. These controllable devices include distributed energy resources (DER) and reclosers. Evaluating such ADMS applications before installation in the field can help de-risk the field implementation and avoid costly failures in the field. This paper presents a background on experimental setup that are typically used to evaluate ADMS applications. This is followed by briefly presenting the setup used to evaluate the FLISR application in an off-the-shelf ADMS tool. Finally, results from the evaluation experiments are presented.

94 GMLC - Grid Modernization Laboratory Consortium↗

AGR-5/6/7 Irradiation Experiment Fission Product Mass Balance

This report presents the fission product mass balance for the AGR-5/6/7 TRISO fuel irradiation experiment. The fission product inventories deposited on capsule components outside of the fuel (e.g., stainless-steel shells, graphite holders, Grafoil disks, and associated hardware) were quantified as part of the post-irradiation examination (PIE) to assess the performance of this fuel. Comparisons were made between these inventories and depletion calculations, non-destructive measurements of fuel fission product inventories in the intact fuel compacts, prior AGR experiments, and results from among each of the five distinct AGR-5/6/7 capsules. The data served as estimates of the condensable fission products released from the fuel during irradiation. Excluding Capsule 1 (which experienced accidental damage during irradiation) and Capsule 3 (which was tested at very high irradiation temperatures of >1300°C), the results indicate that the AGR-5/6/7 fuel performed comparably to fuel from earlier AGR experiments. The mass balance results were also used to estimate the number of particles with in-pile SiC failures. Subject to the assumptions made in these estimates, and excluding Capsules 1 and 3, the in-pile SiC failure rates for AGR-5/6/7 are comparable to those observed in AGR-2.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

The Durability of Piston Seals in Hydraulic Power Take-off Systems in Wave Energy Converters

Hydraulic cylinder seals are a critical component of hydraulic power take-off (PTO) systems in wave energy converters (WECs). Primary hydraulic piston seal wear is a major concern for the longevity of hydraulic PTOs, especially in the context of the effort and expense associated with seal replacement. Piston seals, made from polymeric elastomers, are used to contain and isolate high pressure fluids within PTO systems. A specific challenge for WEC designers is knowing, with confidence, the relative expected lifetimes of commercially available seals and seal materials for the unique long travel and continuous use case of WEC hydraulic systems. This information is critical to accurately determine operating expense (OPEX) and levelized cost of electricity (LCOE). If failures of seals occur earlier than their designed lifetime, the estimated operations, and maintenance (O&M) and LCOE costs may double based on estimation. Unfortunately, information from seal manufacturers on longevity in these applications is not generally available and quantitative performance comparison between different manufacturers is not available, creating significant uncertainty on use of hydraulic PTO system in wave power generation. In this project, PNNL, with advice from different WEC device and seal manufacturers, has created a framework to address the industry need for available, dependable and comparative data for seals and seals materials for WEC hydraulic applications including piston seals, glide rings, and shaft seals. Commonly used and candidate seal materials were identified and available information on the materials such as mechanical and fatigue performance, chemical (fluid) compatibility, and cost has been compiled. Hardware and strategy for bench scale measurement of key materials and seal performance and pathway for publicly available library of hydraulic seal materials, properties, suppliers, and options have also been identified for future implementation. The results of this project were presented at WPTO Seedling Symposium 2023 and OCEANS 2023. The results of the literature review including identified polymer seals and ideal operating conditions were summaried and compiled to a database WEC-SealsDB hosted locally at PNNL.

16 TIDAL AND WAVE POWER↗

Materials Challenges and Opportunities for Energy Generation, Conversion, Delivery, and Storage (Applied Energy Tri-Laboratory Consortium Workshop Report)

This report documents the outcomes of the Tri-Laboratory Materials Workshop that was held July 31 and August 1, 2019 to begin addressing the needs, opportunities, and challenges associated with the development, fabrication, and testing of the needed materials and components for integrated hybrid energy systems (i.e., incorporating nuclear, fossil, and renewables for electric and thermal applications). This was accomplished by assembling the research program leads and principal investigators at Idaho National Laboratory (INL), National Energy Technology Laboratory (NETL), and National Renewable Energy Laboratory (NREL), who support the research and development of new technology and system integration. The team then identified and prioritized key materials development needs. This effort was intended to enhance communications and synergy among the Tri-Lab partners. Advanced functional and structural materials are central to transformative energy technologies for energy generation, conversion, delivery, and storage. With that in mind, the workshop focused on identifying and assessing the foundational materials research needs at both the basic and applied levels. Materials challenges include the ability to withstand harsh environments, such as high temperatures and pressures, corrosion, oxidation, or irradiation while maintaining flexible mission profiles and long service lifespans. Advanced energy system material challenges and needs range from materials for the capture, upgrading/concentration, storage, and delivery of low-grade heat to materials for high temperature environments that involve liquid metals, molten salt, and very high temperature gas heat delivery and storage systems. Material improvements are needed for hybrid energy systems due to accelerated corrosion and stress-fatigue failure of materials and equipment, which results from increased frequency and amplitude of thermal, mechanical, and electrical cycling of systems components. Multifunctional materials are needed for high temperature solid-oxide fuel cells, advanced electrochemical reactors, and in-process separation. Relative to materials manufacturing, application of advanced additive and subtractive methods need to be understood to develop both thin-layer homogenous materials and materials of graded composition. Materials modeling and machine learning will be critical to accelerate the design and production of power electronics, and nuclear reactor materials and fuel, as well as to gain an understanding of beneficial materials phenomena or deleterious microstructure evolution. There is also a need for standardized models, computational structures, data reporting protocols and modeling tools across the three laboratories. This would allow consistent results, analysis, and data sharing. Combining computational capabilities between the three laboratories (e.g., hardware, software) would greatly increase computational capabilities and throughput. The workshop identified the need for laboratories to anticipate and address problems that will occur during scale-up. Laboratory work must connect with industry to ensure that research focuses on processes that are scalable and marketable. Industry input and perspective are essential to guide laboratory research to meet these requirements and deploy new technology in industrial demonstrations. Another aspect of scale-up is the integration of multiple systems since new challenges often arise at the subsystem interfaces. Establishing a scale-up manufacturing demonstration/pilot plant, potentially as an industrial user facility, would be beneficial to the laboratories and industry. That modular scale-up manufacturing demonstration/pilot plant would allow researchers to find and resolve interface problems that cannot be identified by focusing only on individual parts. Communication exchanges among the organizers, attendees, and workshop survey responses indicate that the workshop was successful in achieving its goal to identify key technology gaps and research needs. Strong positive feedback was received on the sharing of ideas, capabilities, talent, and passion to move forward on the materials-related action items.

36 MATERIALS SCIENCE↗

Solar Photovoltaics Resilient Fasteners Levelized Cost of Energy (LCOE) Tool

Solar photovoltaics (PV) module fasteners are one of the most common structural failure points on PV systems, particularly in high winds and coastal areas with ocean spray. Some fastener types have been shown to survive these conditions at higher rates than others. The fastener type, material, quantity, and placement all impact performance. Fasteners that fail less often typically have a higher upfront cost, but this investment can pay off in savings from less frequent torque audits (which reduces O&M costs), reduced system damage, and decreased system downtime. We developed an Excel-based tool to evaluate different module fasteners for a PV system - either a new or retrofit project - and compare differences in upfront and outyear costs to determine the expected life cycle costs and simple payback periods of different fastener options. The tool is site-specific, with inputs including system attributes (such as system size, location, price of power) and fastener attributes (such as design, washer type, nut type, use of locking hardware, materials, installation time, and torque audit requirements). A baseline fastener scenario can be compared to up to four proposed fastener scenarios. In addition to presenting expected life cycle cost implications of the different fastener options, the tool produces results showing the reductions in outyear costs needed to offset any initial cost premiums for more reliable fasteners across four categories: preventative O&M, avoided damage, reduced downtime, and reduced insurance premiums. These numbers can serve as decision aids for users when considering fastener options on new or existing projects. This poster will present the tool, methodology, and scenarios using example sites to highlight the tool capabilities and how it can inform different fastener decisions on different projects. Future work includes incorporating lifetime expected damage costs by embedding damage function curves that the authors are developing from field data.

14 SOLAR ENERGY↗

Characterization of Performance Degradation Mechanisms in Low-Cost High Throughput DI-O3 Layer for Passivated Contact Silicon Solar Cells

Characterization and mitigating performance limiting defects in Silicon (Si) PV is one of key areas to be addressed to improve PV hardware costs and energy yield in order to lower the levelized cost of energy (LCOE) of installed PV cost to $0.02/kWh. As Si PV cells efficiencies have surpassed 22% and approaching 23%, the recombination at the metal contacts have become the focus point to be addressed. Passivated contact technologies—having a heterojunction with a band-gap larger than silicon between the metal and silicon—have emerged as a great potential for future highand ultrahigh-efficiency solar cells, as it concurrently reduces recombination and increases carrier selectivity, by incorporating thin films within the contact structure. Passivated contact Si solar cell technologies use a wide variety of tunnel layers—playing a crucial role to passivate metal contacts and tunnel charge carriers—including stoichiometric silicon oxide (SiO 2 ) grown by thermal oxidation and Low-Pressure Chemical Vapor Deposition (LPCVD) technique and silicon oxide (SiO x ) by hot nitric acid. However, thorough investigations on understanding the failure and performance degradation mechanisms associated with tunnel layers are still limited to date. Unlocking those degradation characteristics in crucial tunnel layers could improve the reliability and energy yield of passivated contact Si solar cells. Besides, the technique of growing aforementioned tunneling layers are low throughput, and requires high temperature processes and/or a vacuum environment. In this project, we investigated the performance degradation mechanisms of a low-cost high-throughput ozonated oxide (DI-O 3 ) tunnel layer for the passivated contact Si solar cells.

14 SOLAR ENERGY↗

Distributed Quantum-Enhanced Optimization: A Topographical Preconditioning Approach for High-Dimensional Search

Optimization problems become fundamentally challenging as the number of variables increases. Because the volume of the search space grows exponentially, classical algorithms frequently fail to locate the global minimum of non-convex functions. While quantum optimization offers a potential alternative, mapping continuous problems onto near-term quantum hardware introduces severe scaling limits and barren plateaus. To bridge this gap, we propose the Distributed Quantum-Enhanced Optimization (D-QEO) framework. Instead of forcing the quantum processor to find the exact minimum, we use it simply as a topographical preconditioner. The QPU maps the landscape to locate the most promising basin of attraction, generating high-quality seed points for a classical GPU-accelerated solver to refine. To make this approach viable for utility-scale problems, we exploit the mathematical structure of separable functions. This allows us to cut a 50-qubit (i.e., $2^{50}$) global search space into independent and manageable sub-spaces using 5-qubit subcircuits. By executing these fragments concurrently with CUDA-Q, we completely bypass the overhead of cross-register entanglement and classical tensor knitting for separable functions. Benchmarks on the 10-dimensional Rastrigin and Ackley functions show that D-QEO prevents the exponential failure rates observed in purely classical algorithms. Furthermore, this quantum warm-start significantly reduces the number of classical BFGS iterations required to converge, providing a highly practical blueprint for utilizing near-term quantum resources in complex global search.

Soos, Dominik [Old Dominion U.]↗

Development of dispensing hardware for safe fueling of heavy duty vehicles

The development of safe dispensing equipment for the fueling of heavy duty (HD) vehicles is critical to the expansion of this newly and quickly expanding market. This paper discusses the development of a HD dispenser and nozzles assembly (nozzle, hose, breakaway) for these new, larger vehicles where flow rates are more than double compared to light duty (LD) vehicles. This equipment must operate at nominal pressures of 700 bar, -40o C gas temperature, and average flow rate of 5-10 kg/min at a high throughput commercial hydrogen fueling station without leaking hydrogen. The project surveyed HD vehicle manufacturers, station developers, and component suppliers to determine the basic specifications of the dispensing equipment and nozzle assembly. The team also examined existing codes and standards to determine necessary changes to accommodate HD components. From this information, the team developed a set of specifications which will be used to design the dispensing equipment. In order to meet these goals, the team performed computational fluid dynamic, pressure modelling, and temperature analysis in order to determine the necessary parameters to meet existing safety standards modified for HD fueling. The team also considered user, operational, and maintenance requirements, such as freeze lock which has been an issue which prevents the removal of the nozzle from LD vehicles. The team also performed a failure mode and effects analysis (FMEA) to identify the possible failures in the design. The dispenser and nozzle assembly will be tested separately, and then installed on an innovative, HD fueling station which will use a HD vehicle simulator to test the entire system.

08 HYDROGEN↗

Analytics-at-scale of Sensor Data for Digital Monitoring in Nuclear Plants (4th Annual Report)

Nuclear plant sites collect and store large volumes of data collected from various equipment and systems. These datasets typically include plant process parameters, maintenance records, technical logs, online monitoring data, and equipment failure data. The collection of such data affords an opportunity to leverage data-driven machine learning and artificial intelligence technologies to provide diagnostic and prognostic capabilities within the nuclear power industry to reduce operating and maintenance costs. In this way, nuclear energy can become more economically competitive with other energy sources, and premature closures can be avoided. From a maintenance standpoint, savings can be achieved by leveraging machine learning and artificial intelligence technologies to develop data-driven algorithms to better diagnose and predict potential faults within the system. Improved model accuracy can lead to reductions in unnecessary maintenance and more efficient planning of future maintenance, thus lowering the costs associated with parts, labor, and unnecessary planned, forced, or extended outages. From an operations perspective, cost savings can be generated by shifting from route-based monitoring to wireless technologies for online monitoring, and by transitioning from onsite- to cloud-based computing and storage services. Wireless monitoring would reduce the operator manhours required for taking routine measurements, while cloud computing services would generate cost savings by reducing the amount of hardware needing to be purchased and maintained—all while scaling to both computational and storage demands. This report summarizes this project’s effort to shift from costly, labor-intensive preventative maintenance to cheaper predictive maintenance.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Experiences with SYCL on AMD GPUs with Kokkos

With the recent diversification of the hardware landscape in the high-performance computing (HPC) community, performance-portability solutions are becoming more and more important. One of the most popular choices is Kokkos, which recently became a Linux Foundation project. Most of its development is supported by the US Department of Energy and the French Alternative Energies and Atomic Energy Commission. Kokkos is implemented as a C++ library with multiple backends to support CPUs as well as various GPU architectures. These backends include OpenMP, CUDA, HIP, and also SCYL. This approach enables users to leverage the preferred vendor toolchain for the respective platform (e.g. CUDA, ROCm, OneAPI). The SYCL backend is used to target Intel GPUs, in particular to support the Aurora exascale supercomputer. However, SYCL itself also offers a large degree of portability, and in fact Kokkos’ CI for SYCL has been running on NVIDIA hardware due to a lack of access to Intel GPUs. In this report, we describe our experience with using Kokkos SYCL backend on AMD GPUs targeting the Frontier supercomputer at Oak Ridge National Laboratory. The two major SYCL implementations are DPC++ and AdaptiveCpp. While the Kokkos SYCL backend has been implemented using the former, the latter was the first implementation to target AMD GPUs. We will discuss the experience with both of these SYCL implementations in terms of functionality and performance. Using Kokkos to evaluate SYCL toolchains has a number of benefits. Kokkos’ use of SYCL is fairly complex, exercising features such as graphs, relocatable device functions, atomics – including for non-arithmetic types, as well as pinned and page migratable memory allocations. Kokkos also needs to implement capabilities such as Kokkos’ hierarchical parallelism that are not a straight-forward mapping to SYCL capabilities. Furthermore, a large number of libraries and applications that represent diverse use cases are implemented in Kokkos, providing readily available test cases for a toolchain evaluation. Preliminary results show that support for AMD GPUs in DPC++ is much less mature than for NVIDIA GPUs or Intel GPUs. While the situation has improved significantly over the last year, we still encounter many runtime failures, dispatching problems, and code generation issues. With AdaptiveCpp the challenges arise even earlier in the evaluation process. Since Kokkos’ SYCL implementation is largely focused on supporting Intel GPUs, we opted to leverage SYCL extensions which are available in DPC++ but not in AdaptiveCpp. Furthermore, AdaptiveCpp appears to be less conformant with the SYCL2020 standard which Kokkos relies on. In some cases, we are able to work around the lack of feature support, in other cases we have to disable certain Kokkos capabilities to evaluate the toolchain. Our evaluation will leverage Kokkos’ unit tests to establish basic functionality and feature completeness. We then use simple benchmarks for components of a CG implementation as a measure of usability and performance of the SYCL toolchains.

97 MATHEMATICS AND COMPUTING↗

Upgrade of Gamma Spectrometry Systems for ORNL TRISO Fuel PIE

Gamma spectrometry is a key element in much of the post-irradiation examination (PIE) work performed under the Advanced Gas Reactor Fuel Development and Qualification (AGR) Program (Demkowicz et al. 2015; Stempien et al. 2021). Gamma spectrometers are integrated into three major capabilities used at the Oak Ridge National Laboratory (ORNL) Irradiated Fuels Examination Laboratory (IFEL) for PIE of tristructural-isotropic (TRISO) coated particles and fuel compacts: the Core Conduction Cooldown Test Facility (CCCTF), the Vertical Counting System (VCS), and the Irradiated Microsphere Gamma Analyzer (IMGA). The CCCTF includes liquid-nitrogen-cooled traps to extract 85 Kr out of the He sweep gas that passes through the furnace in which the fuel compacts are heated during safety testing. Analysis of the 85 Kr activity in the traps is the primary indicator for TRISO failure during safety testing. The VCS is a system used to accurately measure gamma emission from components placed in a lead-shielded chamber. It is used to count the CCCTF deposition cups after removal from furnace. Each cup resides in the CCCTF furnace for typically 12–24 h and is periodically replaced with a fresh cup throughout the safety test. Metallic fission products collect on the water-cooled cups and several gamma-emitting isotopes ( 110 mAg, 134 Cs, 137 Cs, 154 Eu, and 155 Eu) are often measured and provide indication of the retention performance of the TRISO coatings. The VCS is also used to measure the presence of these isotopes on the CCCTF tantalum liner and sweep gas inlet tube for the determination of cup collection efficiency, as well as support other gamma spectrometry needs related to calibration of the 85 Kr fission gas traps and various other special PIE tasks. The IMGA uses gamma spectrometry to measure the inventory of gamma-emitting isotopes in individual TRISO particles. An automated particle handling system within the IMGA hot cell removes each particle from a source vial and positions it in front of a gamma detector, and output from the gamma spectrometer is used by the IMGA software to determine a destination vial such that particles are sorted according to their inventory and retention characteristics. At the conclusion of the AGR-1 and AGR-2 PIE campaigns, the gamma spectrometer systems used at ORNL to support that PIE had reached the end of its life cycle due to gradual obsolescence of the hardware and software. Upgrade of the Canberra Genie 2000 software used by these systems to a Windows 10 version was not a viable option, because the newest Windows 10 version offered by Mirion (the new owner of the Canberra technology) did not include the dynamic-link libraries (DLLs) needed for integration with the custom PIE software used with the CCCTF and IMGA, and Mirion had no current plans for development and release of Windows 10 versions of these DLLs with the Model S560 Genie 2000 Programming Library. Ultimately a switch was made to ORTEC gamma spectrometry systems, which appeared to be a more sustainable solution due to more proactive vendor support. The ORTEC conversion involved replacing the aging detector preamplifier and multichannel analyzer (MCA) hardware, upgrading the obsolete Windows 7 computers to Windows 10 compatible models, adopting ORTEC GammaVision software, and extensive modification of the ORNL-developed Visual Basic .NET (VB.NET) programs that provide the CCCTF and IMGA user interfaces.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗