Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Reliability and Mitigation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Grid Reliability and U.S. Coal Fleet Attributes: Considerations for State Regulators

This briefing paper provides state utility regulators with a comprehensive and current understanding of the reliability attributes of coal as a generation resource, potential reliability impacts associated with near-term coal plant retirements, and possible mitigation strategies to ensure a stable and resilient energy system.

01 COAL, LIGNITE, AND PEAT↗

AI-Driven Crack Detection for Remanufacturing Cylinder Heads Using Deep Learning and Engineering-Informed Data Augmentation

Detecting cracks in cylinder heads traditionally relies on manual inspection, which is time-consuming and susceptible to human error. As an alternative, automated object detection utilizing computer vision and machine learning models has been explored. However, these methods often face challenges due to a lack of sufficiently annotated training data, limited image diversity, and the inherently small size of cracks. Addressing these constraints, this paper introduces a novel automated crack-detection method that enhances data availability through a synthetic data generation technique. Unlike general data augmentation practices, our method involves copying cracks from one location to another, guided by both random and informed engineering decisions about likely crack formations due to cyclic thermomechanical loads. The innovative aspect of our approach lies in the integration of domain-specific engineering knowledge into the synthetic generation process, which substantially improves detection accuracy. We evaluate our method’s effectiveness using two metrics: the F2 score, which emphasizes recall to prioritize detecting all potential cracks, and mean average precision (MAP), a standard measure in object detection. Experimental results demonstrate that, without engineering insights, our method increases the F2 score from 0.40 to 0.65, while maintaining a stable MAP. Incorporating detailed engineering knowledge further enhances the F2 score to 0.70 and improves MAP to 0.57, representing increases of 63% and 43%, respectively. These results confirm that our approach not only mitigates the limitations of traditional data augmentation but also significantly advances the reliability and precision of crack detection in industrial settings.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Performance of High Temperature Operational Amplifier, Type LM2904WH, under Extreme Temperatures

Operation of electronic parts and circuits under extreme temperatures is anticipated in NASA space exploration missions as well as terrestrial applications. Exposure of electronics to extreme temperatures and wide-range thermal swings greatly affects their performance via induced changes in the semiconductor material properties, packaging and interconnects, or due to incompatibility issues between interfaces that result from thermal expansion/contraction mismatch. Electronics that are designed to withstand operation and perform efficiently in extreme temperatures would mitigate risks for failure due to thermal stresses and, therefore, improve system reliability. In addition, they contribute to reducing system size and weight, simplifying its design, and reducing development cost through the elimination of otherwise required thermal control elements for proper ambient operation. A large DC voltage gain (100 dB) operational amplifier with a maximum junction temperature of 150 C was recently introduced by STMicroelectronics [1]. This LM2904WH chip comes in a plastic package and is designed specifically for automotive and industrial control systems. It operates from a single power supply over a wide range of voltages, and it consists of two independent, high gain, internally frequency compensated operational amplifiers. Table I shows some of the device manufacturer s specifications.

Patterson, Richard↗

Bringing Single-Event Effects Down to Earth

In the 47 years since single-event effects were first observed in spacecraft electronics, radiation experts have developed an effective methodology supported by a nationwide infrastructure. A highly skilled workforce of radiation engineers has developed test facilities and methods, modeling and simulation techniques, and mitigation and design strategies to ensure space missions meet their performance and reliability requirements even in the harsh radiation environments of space. Now, increasing performance demands of space missions, the continued disruptive evolution of microcircuit technologies and growth and changes of the space industry have combined with an aging infrastructure are placing increasing strain on the radiation effects community, and the community is responding.

Single-event effects↗

Technical Assistance for Digital Assurance: DistribuTECH Workshop

Idaho National Laboratory, the Department of Energy’s Grid Deployment Office, and other national laboratories are collaborating to ensure the U.S. energy infrastructure is reliable, resilient, and secure. This involves strategically leveraging digital technologies to modernize the grid, enhance its resilience against all hazards and disruptions, and fortify national energy security across a diverse energy portfolio. In this workshop you will learn about Idaho National Lab’s Technical Assistance programs, where organizations will be matched with a national laboratory subject matter expert to focus on their key topical area. The technical assistance offered through this track is designed to be responsive to a rapidly changing regulatory landscape and cutting-edge technologies that enhance grid reliability and efficiency. Users will be guided through a tailored analysis and mitigation program to determine their current security posture and given assistance in evaluating supply chain and protection choices against potential consequences.

Digital assurance↗

Risk Trade-Space Analysis for Safe Human Expeditions to Mars

We assessed the integrated safety, health, and performance risk to crews on long-duration missions, specifically to Mars. Using a systems approach rather than one focused on individual countermeasures, we examined the trade space around several such risks to identify high-potential risk mitigation strategies and characterize aspects of Mars mission architectures that could lower aggregated risk. Current Mars Design Reference missions would require durations well over two years and would increase crew exposure to radiation and microgravity well beyond ISS levels, likely resulting in significantly reduced performance beyond our current capability to mitigate that could jeopardize mission success. A “fast Mars transit” round-trip mission concept was studied using an innovative flight dynamics approach to quantify the minimum total mission energy required for a Mars transit with total mission duration less than 400 days. This approach holds promise for sending humans to Mars and returning them safely with acceptable, potentially mitigatable, exposure to microgravity and radiation using current or near-term technologies. The fast transit concept would also result in fewer time-driven vehicle failures and enable sustainable deployment of humans and infrastructure to Mars on a regular cadence, allowing steady exploration and colonization of Mars. Finally, we conclude that reliance on the Low Earth Orbit (LEO) mission operations paradigm – i.e., one of near-complete real-time dependence on experts at Mission Control to manage the combined state of the mission, vehicle, and crew – is high risk given the communication delays and limited resupply of any Mars mission, and this risk is not eliminated by the shorter missions durations of fast transit scenarios. Based on historical trends, it is highly likely that the crew will face a high-consequence problem of uncertain origin during Mars transit when ground support will be greatly reduced. While it may be possible to reduce anomaly rates through improved reliability analysis and testing, and to reduce anomaly impacts through added robustness, such mitigations address only known failure modes and known uncertainties. Therefore, a radical shift in the Human-Systems Integration Architecture (HSIA) that defines the operational paradigm, systems design, and human-systems interactions is required to improve the risk posture to an acceptable level regardless of mission duration.

Mars↗

Advanced reliability modeling of fault-tolerant computer-based systems

Two methodologies for the reliability assessment of fault tolerant digital computer based systems are discussed. The computer-aided reliability estimation 3 (CARE 3) and gate logic software simulation (GLOSS) are assessment technologies that were developed to mitigate a serious weakness in the design and evaluation process of ultrareliable digital systems. The weak link is based on the unavailability of a sufficiently powerful modeling technique for comparing the stochastic attributes of one system against others. Some of the more interesting attributes are reliability, system survival, safety, and mission success.

Bavuso, S. J.↗

Understanding and Estimating Error Propagation in Neural Networks for Scientific Data Analysis

Neural networks are increasingly integrated into scientific discovery, where input data reduction and model quantization play a key role in accelerating inference. However, understanding and mitigating the impact of these techniques on output error is critical for ensuring reliable results, particularly in tasks demanding high numerical precision. This paper introduces a comprehensive framework for optimizing neural network inference in scientific computing by combining data reduction and weight quantization while maintaining error-controlled outcomes. We develop theoretical analyses to bound error propagation under these reductions and propose a framework that balances computational performance with error constraints. Evaluation on real-world learning-based combustion simulations and satellite image classification demonstrates that our derived error bounds accurately predict observed errors while enabling significant computational speedup under our framework. This work highlights the potential for further leveraging advancements in modern lossy compression algorithms and hardware accelerators that support lower-precision formats.

He, Weiming [New Jersey Institute of Technology]↗

Optimal Droop Setting for Congestion Reduction in a 100% Grid-Forming Inverter-based Power System

he high penetration of inverter-based resources (IBRs) introduces new challenges to power systems due to the complex inverter control. However, IBRs can be configured to maximize their benefits to improve system resilience and reliability. This paper proposes a steady-state optimization model that aims to mitigate transmission congestion in a 100% grid- forming (GFM) IBR-based power system. This goal is achieved by determining the optimal droop settings for the GFM IBRs under different congestion conditions due to renewable energy and load variations. The numerical solution is rigorously verified by a high-fidelity model of the IEEE 39-bus test system with detailed GFM IBR control in the time-domain electromagnetic transient (EMT) simulation tool PSCAD. The numerical solution and simulation results show a significant congestion reduction while meeting all other operating requirements. It is also observed that the numerical solving time is substantially less compared to the EMT simulation time.

Nguyen, Quan H.↗

Fundamentals of Solar PV Bolted Joint Loosening and Prevention

The photovoltaic (PV) industry has long reported anecdotal accounts of systems exhibiting intermittent or chronic fastener loosening, including joints that fail to maintain preload despite multiple re-tightening attempts. These occurrences are frequently, and often incorrectly, attributed to installer errors, vibration, or loading beyond design expectations (e.g., extreme weather events). Loose fasteners have serious implications and can significantly impact solar PV systems’ performance, reliability, and safety. However, effective bolted joint design and proper assembly practices can mitigate or eliminate loosening. Until the US Department of Energy’s Solar Energy Technologies Office (SETO) funded research on fasteners and solar PV structures, there was a notable gap in understanding the causes of loosening in solar PV systems.

14 SOLAR ENERGY↗

Market Assessment of Forward-Looking Turbulence Sensing Systems

In recognition of the importance of turbulence mitigation as a tool to improve aviation safety, NASA's Aviation Safety Program developed a Turbulence Detection and Mitigation Sub-element. The objective of this effort is to develop highly reliable turbulence detection technologies for commercial transport aircraft to sense dangerous turbulence with sufficient time warning so that defensive measures can be implemented and prevent passenger and crew injuries. Current research involves three forward sensing products to improve the cockpit awareness of possible turbulence hazards. X-band radar enhancements will improve the capabilities of current weather radar to detect turbulence associated with convective activity. LIDAR (Light Detection and Ranging) is a laser-based technology that is capable of detecting turbulence in clear air. Finally, a possible Radar-LIDAR hybrid sensor is envisioned to detect the full range of convective and clear air turbulence. To support decisions relating to the development of these three forward-looking turbulence sensor technologies, the objective of this study was defined as examination of cost and implementation metrics. Tasks performed included the identification of cost factors and certification issues, the development and application of an implementation model, and the development of cost budget/targets for installing the turbulence sensor and associated software devices into the commercial transport fleet.

Kauffmann, Paul↗

Ballistic Puncture Self-Healing Polymeric Materials

Space exploration launch costs on the order of $10,000 per pound provide an incentive to seek ways to reduce structural mass while maintaining structural function to assure safety and reliability. Damage-tolerant structural systems provide a route to avoiding weight penalty while enhancing vehicle safety and reliability. Self-healing polymers capable of spontaneous puncture repair show promise to mitigate potentially catastrophic damage from events such as micrometeoroid penetration. Effective self-repair requires these materials to quickly heal following projectile penetration while retaining some structural function during the healing processes. Although there are materials known to possess this capability, they are typically not considered for structural applications. Current efforts use inexpensive experimental methods to inflict damage, after which analytical procedures are identified to verify that function is restored. Two candidate self-healing polymer materials for structural engineering systems are used to test these experimental methods.

Gordon, Keith L.↗

Truncated ARQ Statistical Link Analysis for Dynamic Links

The future deep space links are migrating towards higher frequency bands such as Ka band and optical. These links are susceptible to non Gaussian and non linear effects such as atmospheric turbulence, scintillation, antenna mis-pointing, jitter, etc. These dynamic links thus will experience various degrees of fading loss, and some of these link disruptions cannot be effectively mitigated by forward error correction coding and/or interleaving. One effective way to ensure reliable communication is by using Automatic Repeat Request (ARQ) protocol, where the receiver acknowledges to the transmitter whether or not a data unit is successfully received. If a data unit is not successfully received (such as after a pre-set time-out), the transmitter would then re-transmit the lost data unit to the receiver. In a previous paper, we derived a statistical link analysis method of finding the optimal operating Signal-to-Noise Ratio (SNR) and estimating the latency of an ARQ scheme. In a more recent paper, we demonstrated the above method using the SNR distribution constructed from the Ka-band (32 GHz) flight data. To simplify the discussion, we considered the academic approach that the ARQ scheme allows for an infinite number of retransmissions. In this paper, we consider the more practical case of a truncated ARQ scheme, where there is a limit on the number of retransmissions. We derive the error probability, the optimal SNR setting, and the latency statistics of the correctly received frames of the truncated ARQ schemes. We first discuss the truncated ARQ link analysis principles using the Gaussian assumption for SNR distribution with a large variance. Next, we demonstrate the statistical truncated ARQ link analysis using the SNR distribution constructed from the Ka-band flight data. The results in this paper can be applied in the design of reliable communication systems such as the Consultative Committee for Space Data System (CCSDS) File Transfer Protocol (CFTP) and the Delay Tolerant Network (DTN).

Morabito, David↗

Threat Landscape for BESS and IBR

The cyber risk landscape for BESS and IBR can be broken up by threats, vulnerabilities, and consequences for these systems. This presentation walks through the cyber risk landscape for BESS through the lens of consequence-informed awareness and mitigation for each risk factor. Threats with varying capabilities have been demonstrated in real-world events. Though threat actors can rarely be directly influenced by organizations, exposure of systems to adversaries can be limited (a known issue with IBR systems) to reduce likelihood of adversaries accessing systems with disruptive consequences. Common trends in disclosed IBR vulnerabilities include weak password generation or managements for various devices or services and web portal vulnerabilities that provide unauthorized access to data or capabilities or elevated user privileges. Understanding these common vulnerabilities and considering the consequences if these types of vulnerabilities were to occur can help mitigate risk. Consequences range from loss-of-view events that have no reliability impact to asset damage or grid stability impacts. Five case studies are briefly shared to highlight trends in real-world events affecting IBR.

14 - SOLAR ENERGY↗

Electric Drive Technologies Consortium (EDTC)/ Cost competitive, high-Performance, highly Reliable (CPR) Power Devices on 4H-SiC (Final Report)

4H-Silicon carbide (4H-SiC) is a wide bandgap semiconductor that offers superior material properties over silicon, including higher critical electric field, thermal conductivity, and electron saturation velocity. These advantages make 4H-SiC highly attractive for high-voltage, high-efficiency power electronics. However, realizing the full potential of SiC requires device technologies that are not only high-performing but also manufacturable and reliable under real-world operating conditions. This report summarizes the outcomes of a five-year R&D effort funded by the U.S. Department of Energy (DOE) under the Electric Drive Technologies Consortium (EDTC), focused on developing cost-competitive, high-performance, and highly reliable (CPR) power devices on 4H-SiC substrates. The program targeted scalable and manufacturable 1.2 kV-class SiC MOSFETs optimized for next-generation electric vehicles, renewable energy systems, and industrial power conversion. The project delivered transformative advancements in SiC power device performance and ruggedness. Particularly, Specific on-resistance (R on,sp ) was reduced by up to 37%, from ~4.0 m$\Omega \cdot$cm 2 in earlier designs to an industry-leading 2.40 m$\Omega \cdot$cm 2 , driven by optimized doping, refined JFET widths, and layout engineering. Breakdown voltages (BV) exceeded 1600 V, marking improvement over legacy baselines, and demonstrating the robustness of newly implemented junction profiles and edge terminations. Short-circuit withstand time (SCWT) saw a remarkable 4$\times$ increase, from ~2 $\mu$s to over 8 $\mu$s, achieved through the successful deployment of deep P-well structures (~1.8–2.0 $\mu$m) via channeling implantation. This innovative process breakthrough enabled precise junction formation without MeV-class implantation tools, reduced leakage under high field stress, and allowed even the shortest-channel devices (down to 0.3 $\mu$m) to achieve both high BV and excellent ruggedness—breaking the traditional trade-off between conduction efficiency and blocking capability. Several novel architectures pushed the performance envelope further. JBSFETs—featuring embedded Schottky portions—eliminated bipolar degradation and drastically reduced third-quadrant leakage, while Ladder MOSFETs introduced a clever orthogonal conduction path that achieved a 15.4% reduction in R on,sp over standard linear designs. Switching performance reached new benchmarks: short-channel devices showed a 31% reduction in total switching energy compared to 0.5 $\mu$m counterparts, while maintaining manageable gate drive requirements. Layout-optimized structures not only improved transconductance but also accelerated switching transitions, pointing to real-world benefits in converter-level efficiency. The devices also passed rigorous reliability validation. Stress-tested across TDDB, HTGB, HTRB, HVP, and burn-in, the devices screened under 30 V/10 hr and 43 V/1 s protocols consistently exhibited tighter lifetime distributions and long-term oxide robustness. These screening techniques proved effective in identifying latent defects and ensuring deployment-grade reliability. Meanwhile, advanced 3D TCAD simulations revealed and resolved electric field hotspots—particularly in HEXFET corners—where fields exceeding 4.8 MV/cm were mitigated through geometry-aware layout corrections. Overall, the results of this project demonstrate a manufacturable and scalable SiC power device platform that addresses key DOE performance targets for efficient, robust, and reliable 1.2kV 4H-SiC Power Devices. The developed technologies represent a meaningful step forward in the commercial readiness of high-voltage SiC solutions and provide a strong foundation for continued advancement in wide bandgap power electronics.

42 ENGINEERING↗

Physics-informed heterogeneous graph neural networks for DC blocker placement

The threat of geomagnetic disturbances (GMDs) to the reliable operation of the bulk energy system has spurred the development of effective strategies for mitigating their impacts. One such approach involves placing transformer neutral blocking devices, which interrupt the path of geomagnetically induced currents (GICs) to limit their impact. The high cost of these devices and the sparsity of transformers that experience high GICs during GMD events, however, calls for a sparse placement strategy that involves high computational cost. To address this challenge, we developed a physics-informed heterogeneous graph neural network (PIHGNN) for solving the graph-based dc-blocker placement problem. Our approach combines a heterogeneous graph neural network (HGNN) with a physics-informed neural network (PINN) to capture the diverse types of nodes and edges in ac/dc networks and incorporates the physical laws of the power grid. We train the PIHGNN model using a surrogate power flow model and validate it using case studies. Results demonstrate that PIHGNN can effectively and efficiently support the deployment of GIC dc-current blockers, ensuring the continued supply of electricity to meet societal demands. Furthermore, our approach has the potential to contribute to the development of more reliable and resilient power grids capable of withstanding the growing threat that GMDs pose.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Floating photovoltaic power plants: A review of energy yield, reliability, and operation and maintenance

Photovoltaic (PV) systems are essential for the transition to sustainable energy, reducing fossil fuel dependence and mitigating climate change. Although PV requires minimal land area — PV can meet the European Union's energy needs using only 0.26% of its land — space for deployment is often scarce in densely populated regions. Floating photovoltaics (FPV) offer an effective solution to land-use challenges by installing PV systems on floating structures in water bodies. FPV is a growing niche within PV with a cumulative installed capacity reaching 7.7 GW globally by 2023. Almost 90% of the installed FPV capacity is in Asia, with close to 50% of in China alone, while the Netherlands and France are the largest markets outside Asia. FPV shows strong potential to support climate targets, but still faces challenges like regulatory barriers, cost competitiveness compared to ground-based PV (GPV), and uncertainties about environmental impacts and system reliability. FPV systems are currently installed mainly on sheltered inland waters, such as quarry lakes, irrigation ponds and reservoirs. FPV technical standards are still being developed. Guidelines have been published by the World Bank, DNV, and Solar Power Europe, and emerging national standards from South Korea, China, and Singapore address design, components, and safety. The International Electrotechnical Commission (IEC) is working on formal standards for floats, mooring systems, and electrical connectors. However, the published best practices lack quantitative guidance for yield modelling and reliability, which this report aims to address. It provides data-driven insights, models, and parameters essential for accurate energy yield, reliability, and maintenance predictions over FPV systems' lifetimes.

14 SOLAR ENERGY↗