Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Reliability and Mitigation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Dispersion Loss Counteracts Embedding Condensation and Improves Generalization in Small Language Models

Large language models (LLMs) achieve remarkable performance through ever-increasing parameter counts, but scaling incurs steep computational costs. To better understand LLM scaling, we study representational differences between LLMs and their smaller counterparts, with the goal of replicating the representational qualities of larger models in smaller models. We observe a geometric phenomenon which we term embedding condensation, where token embeddings collapse into a narrow cone-like subspace in some language models. Through systematic analyses across multiple Transformer families, we show that small models such as GPT2 and Qwen3-0.6B exhibit severe condensation, whereas larger models such as GPT2-x1 and Qwen3-32B are more resistant to this phenomenon. Additional observations show that embedding condensation is not reliably mitigated by knowledge distillation from larger models. To fight against it, we formulate a dispersion loss that explicitly encourages embedding dispersion during training. Experiments demonstrate that it mitigates condensation, recovers dispersion patterns seen in larger models, and yields performance gains across 10 benchmarks. We believe this work offers a principled path toward improving smaller Transformers without additional parameters.

Xiao, Xi [ORNL] (ORCID:0009000009316982)

Cracking in polymer substrates for flexible electronic devices and its mitigation

Mechanical reliability plays a critical role in determining the durability of flexible electronic devices because of the significant mechanical stresses they experience during manufacturing and operation. Many such devices are built on sheets comprising stiff transparent-conducting oxide (TCO) electrode films on compliant polymer substrates, and it is generally assumed that the high-toughness polymer substrates do not crack. Contrary to this assumption, here we show extensive cracking in the polymer substrates during bending of a variety of TCO/polymer sheets, and a device example — flexible perovskite solar cells. Such substrate cracking, which compromises the overall mechanical integrity of the entire device, is driven by the amplified stress-intensity factor caused by the elastic mismatch at the film/substrate interface. To mitigate this substrate cracking, an interlayer-engineering approach is designed and experimentally demonstrated. This approach is potentially applicable to myriad flexible electronic devices, with stiff films on compliant substrates, for improving their durability and reliability.

42 ENGINEERING

Field-based AFDD for refrigerant undercharge in residential HVAC systems: enhancing reliability through false alarm mitigation

This study evaluated rule-based and machine learning (ML) based automated fault detection and diagnostics (AFDD) algorithms for detecting refrigerant undercharge faults in residential heating, ventilation, and air conditioning (HVAC) systems, using actual building data and a minimal set of features. The ML-based algorithms included Decision Tree (DT) and K-Nearest Neighbors (KNN). Both the rule-based and ML-based algorithms demonstrated the capability to detect refrigerant undercharge faults of -30% or more. Both types of algorithms exhibited false alarms before the implementation of a false alarm mitigation algorithm, which motivated the development of such a mitigation strategy. After applying the mitigation, false alarms were substantially reduced, with the rule-based algorithm decreasing to 0.6% and the ML-based algorithms reaching 0%, while maintaining strong detection performance. Although the rule-based algorithm initially showed lower performance compared to the ML-based algorithms, its detection accuracy improved after mitigation to a level comparable to the ML-based algorithms. These results confirm that combining false alarm mitigation with both rule-based and ML-based AFDD algorithms significantly enhances practical reliability while preserving robust fault detection capabilities. Furthermore, the findings demonstrate the potential for field deployment of these algorithms in residential HVAC systems and highlight the importance of minimizing false alarms.

False Alarm

Signal and Power Integrity Design Methodology for High-Performance Flight Computing Systems

Computing capabilities of space systems have in-creased onboard performance by orders of magnitude with the use of radiation-tolerant field-programmable gate arrays (FPGA)and processors. The incorporation of signal and power integrity analysis with printed circuit board (PCB) design in reliable computing architectures for space systems has become critical to enable future mission capabilities. Developers launch high-performance processors into a breadth of orbits and missions, running varying applications that create challenges for designing reliable computing hardware. Specifically, for these designs, academic and industry research has focused on component radiation performance, fault mitigation, and reliable architectures. How-ever, other design parameters including electromagnetic interference (EMI), PCB stackup, signal integrity (SI), voltage regulator module (VRM) design, and power distribution network (PDN)are often deprioritized or disregarded as the design matures. Since these characteristics are becoming more significant in high-performance processor designs, this research presents a hardware design and analysis methodology for high-performance, space-computing systems that focuses on a holistic design approach and PDN reliability. While these challenges exist across all space hardware, the reduced PCB dimensions imposed by SmallSats and CubeSats introduce additional hurdles, specifically to VRM and decoupling design. By examining the relationship between the PDN and radiation performance, an analytical relationship is developed that incorporates Total Ionizing Dose and Single-Event Transients to ensure reliability throughout the mission duration. The presented design methodology is applied to the SpaceCube v3.0 Mini, an FPGA-based on-board science data processing system developed at NASA Goddard Space Flight Center.

Advanced avionics

Operating Klystrons at the Spallation Neutron Source – Two Decades of Perspective

Klystron amplifiers have been operated at the Spallation Neutron Source (SNS) in support of the user program since 2006. SNS tubes have amassed over 100,000 hours of high-voltage and filament-on time providing a significant source of operational statistics for high power, high-duty klystrons. Additionally, the SNS Radiofrequency (RF) Systems Group has developed operational methods to mitigate specific reliability challenges, such as cathode arcing and output power instability. Despite much progress, many of these challenges persist, prompting the SNS to update the procedures used for klystron processing as well as develop a higher throughput test stand.

Moss, John [ORNL] (ORCID:0009000085988916)

High-Temperature Gas Sensor Materials with Properties Predicted via First-Principles Calculations with Machine Learning Modeling and Experimental Corroboration

Understanding the temperature dependence of functional properties of sensing materials is vital for their applications in combustion environments. The electron-phonon coupling that derives the electronic structure change with temperatures is a key property of interest as it affects other sensing responses. Herein, we first assess the temperature dependence of band gap renormalization in sensing materials by employing Allen-Heine-Cardona (AHC) theory with density functional theory (DFT) simulations corroborated with experimental observation. As the AHC calculations are impractical for high-throughput screening of materials, we employ data-driven Gaussian process regression to predict the parameters employed in the O’Donnell empirical model from a set of physical features. To mitigate the reliability issues arising from the small size of the dataset, we apply a Bayesian technique to improve the generalizability of the data-driven models as well as to quantify the uncertainty associated with theoretical predictions. These models capture well the overall trend of the O’Donnell parameters with respect to a reduced feature set obtained by transforming the available physical features. Quantifying the associated uncertainty helps us understand the reliability of the predictions and, therefore, the variation of bandgap as a function of temperature for other novel materials. The predicted candidates from machine learning models are further validated by experiments and DFT calculations.

bandgap renormalization

Mississippi's Strategic Resilience: A multi-systems approach to secure, reliable, and adaptable electric grid infrastructure

Mississippi’s electric grid resilience challenges are linked to an intersection of complex socioeconomic, ecological, technological, historical, and political challenges, exacerbated by increasing severe weather like flooding and tornado events. The state’s legacy of underinvestment in critical energy infrastructure, particularly in rural areas and vulnerable floodplains, have stressed an aging grid, creating long-lasting disruptions in electric service during weather-related outages. Effective emergency management and preparedness is further hampered by a lack of coordination across local, county, and regional scales. Using the TASTI-GRID platform and partnership with Oak Ridge National Laboratory (ORNL), Mississippi is developing a comprehensive regional resilience strategy to overcome energy security and reliability challenges, mitigating the impacts of natural hazards, and positioning Mississippi as a resilient and premier destination for residents, businesses, and economic development.

24 POWER TRANSMISSION AND DISTRIBUTION

Fly-by-light flight control system technology development plan

The results of a four-month, phased effort to develop a Fly-by-Light Technology Development Plan are documented. The technical shortfalls for each phase were identified and a development plan to bridge the technical gap was developed. The production configuration was defined for a 757-type airplane, but it is suggested that the demonstration flight be conducted on the NASA Transport Systems Research Vehicle. The modifications required and verification and validation issues are delineated in this report. A detailed schedule for the phased introduction of fly-by-light system components has been generated. It is concluded that a fiber-optics program would contribute significantly toward developing the required state of readiness that will make a fly-by-light control system not only cost effective but reliable without mitigating the weight and high-energy radio frequency related benefits.

Chakravarty, A.

Reliability of CGA/LGA/HDI Package Board/Assembly (Revision A)

This follow-up report presents reliability test results conducted by thermal cycling of five CGA assemblies evaluated under two extreme cycle profiles, representative of use for high-reliability applications. The thermal cycles ranged from a low temperature of 55 C to maximum temperatures of either 100 C or 125 C with slow ramp-up rate (3 C/min) and dwell times of about 15 minutes at the two extremes. Optical photomicrographs that illustrate key inspection findings of up to 200 thermal cycles are presented. Other information presented include an evaluation of the integrity of capacitors on CGA substrate after thermal cycling as well as process evaluation for direct assembly of an LGA onto PCB. The qualification guidelines, which are based on the test results for CGA/LGA/HDI packages and board assemblies, will facilitate NASA projects' use of very dense and newly available FPGA area array packages with known reliably and mitigation risks, allowing greater processing power in a smaller board footprint and lower system weight.

Ghaffarian, Reza

Reliability of CGA/LGA/HDI Package Board/Assembly

This follow-up report presents reliability test results conducted by thermal cycling of five CGA assemblies evaluated under two extreme cycle profiles, representative of use for high-reliability applications. The thermal cycles ranged from a low temperature of −55°C to maximum temperatures of either 100°C or 125°C with slow ramp-up rate (3°C/min) and dwell times of about 15 minutes at the two extremes. Optical photomicrographs that illustrate key inspection findings of up to 200 thermal cycles are presented. Other information presented include an evaluation of the integrity of capacitors on CGA substrate after thermal cycling as well as process evaluation for direct assembly of an LGA onto PCB. The qualification guidelines, which are based on the test results for CGA/LGA/HDI packages and board assemblies, will facilitate NASA projects’ use of very dense and newly available FPGA area array packages with known reliably and mitigation risks, allowing greater processing power in a smaller board footprint and lower system weight.

Ghaffarian, Reza

Risk-informed Graded Approach for Reliability and Performance Assessment of Sensor and Instrumentation Systems within Advanced Condition Monitoring Technologies

Advanced condition monitoring (ACM) technologies, such as digital twins, are innovative strategies designed to provide real-time health insights, including the remaining useful life of components. The primary goal of ACM is to predict and alert operators to potential functional failures before they occur. ACM systems achieve this by integrating predictive models with various sensor instrumentation, analog-to-digital converters, data warehouses, and data pre-processors. These sensor and instrumentation systems (SIS) are essential for forming a comprehensive understanding of component conditions and ensuring the predictive success of ACM programs. Introducing new technologies like ACM involves varying degrees of risk that can impact plant reliability. Therefore, risk mitigation should be commensurate with the performance and reliability of the developed technology, following a risk-informed graded approach (RIGA). Establishing a RIGA process requires a clear understanding of the hazards and reliability of all subsystems, including their interdependencies and potential impacts on the overall system. Given the critical role of SIS in ACM, this work reviews hazard identification and reliability quantification methods for SIS. It also considers these methods' implications when developing a RIGA process for ACM.

22 - GENERAL STUDIES OF NUCLEAR REACTORS

Backup power or bill savings? How electricity tariffs impact residential solar-plus-storage usage in the United States

Adoption of paired solar-plus-storage systems has accelerated in recent years, driven by both the demand for backup power and a desire to manage utility bills. Tradeoffs between those two uses can arise through the reserve setting on the battery storage system, which serves to maintain a minimum state of charge in case of a power interruption. Our paper applies an economic framework to evaluate this tradeoff in terms of changes in bill savings and customer reliability value across reserve levels, considering how those tradeoffs depend on the underlying electricity rate structure and levels. The analysis is based on a representative set of load profiles, solar profiles, tariff designs, and stochastic power interruption events across ten different regions in the United States. We find that the opportunity cost of holding storage capacity in reserve, in terms of foregone bill reductions, outweighs any gains in reliability value from mitigated power interruptions in the majority of customer situations. Higher storage reserve levels increase total customer value only in specific circumstances, such as for customers with inferior reliability (10x average interruptions), with a very high value of lost load ($50/kWh), and with tariff or interconnection rules that disallow grid charging. However, even this result is dampened when considering tariff designs with higher price differentials that increase the opportunity cost of holding storage in reserve (e.g. import/export or time-of-use rates). Allowing grid charging in tariffs essentially eliminates the necessity to hold any storage in reserve in all sensitivity cases explored.

Electric resilience

Modeling of Transient Flow Mixing of Streams Injected into a Mixing Chamber

Ignition is recognized as one the critical drivers in the reliability of multiple-start rocket engines. Residual combustion products from previous engine operation can condense on valves and related structures thereby creating difficulties for subsequent starting procedures. Alternative ignition methods that require fewer valves can mitigate the valve reliability problem, but require improved understanding of the spatial and temporal propellant distribution in the pre-ignition chamber. Current design tools based mainly on one-dimensional analysis and empirical models cannot predict local details of the injection and ignition processes. The goal of this work is to evaluate the capability of the modern computational fluid dynamics (CFD) tools in predicting the transient flow mixing in pre-ignition environment by comparing the results with the experimental data. This study is a part of a program to improve analytical methods and methodologies to analyze reliability and durability of combustion devices. In the present paper we describe a series of detailed computational simulations of the unsteady mixing events as the cold propellants are first introduced into the chamber as a first step in providing this necessary environmental description. The present computational modeling represents a complement to parallel experimental simulations' and includes comparisons with experimental results from that effort. A large number of rocket engine ignition studies has been previously reported. Here we limit our discussion to the work discussed in Refs. 2, 3 and 4 which is both similar to and different from the present approach. The similarities arise from the fact that both efforts involve detailed experimental/computational simulations of the ignition problem. The differences arise from the underlying philosophy of the two endeavors. The approach in Refs. 2 to 4 is a classical ignition study in which the focus is on the response of a propellant mixture to an ignition source, with emphasis on the level of energy needed for ignition and the ensuing flame propagation issues. Our focus in the present paper is on identifying the unsteady mixing processes that provide the propellant mixture in which the ignition source is to be placed. In particular, we wish to characterize the spatial and temporal mixture distribution with a view toward identifying preferred spatial and temporal locations for the ignition source. As such, the present work is limited to cold flow (pre-ignition) conditions

Voytovych, Dmytro M.

An Assessment of Technical Hydropower Potential at Non-Powered Dams in the United States

Historically, dams have been constructed for a variety of purposes, such as providing a more secure and reliable water supply, mitigating impacts from variations in river flow, allowing continuous navigability, and harnessing mechanical power. A relatively small portion of dams have been designed to store or regulate flows for the purpose of generating electricity (roughly 3% of nationally inventoried dams in the US and 17% of the dams included in the World Register of Dams). The remaining population of existing non-powered dams (NPDs) presents both an opportunity to generate renewable energy and a need to modernize aging infrastructure. This report describes an assessment of more than 2,600 NPDs in the US that have a collective potential of nearly 4 GW in new power capacity. Previous national-scale assessments were aimed at evaluating the theoretical maximum power potential at existing dams in the United States. These estimates were based on the best available information at the time for water availability, hydraulic head, and representative regional capacity factors. This study revisits a subset of 3,299 dams identified in the most recent theoretical resource assessment and uses more detailed and updated hydrologic data to produce estimates of technical potential. These improvements in data enable estimates that more realistically reflect what is physically possible given simple assumptions about the existing structure and constraints on flow and head (Figure 1).

13 HYDRO ENERGY

Identifying Decoherence Mechanisms in Superconducting Qubits through Advanced Materials Characterization

Although superconducting qubits have emerged as a leading technology platform for quantum computing through large improvements in device coherence times and gate fidelity in recent years, the presence of defects and impurities at the interfaces and surfaces in the constituent materials continue to limit performance and serve as a critical barrier in achieving scalable quantum systems. Understanding and eliminating these sources of quantum decoherence in superconducting qubit devices requires dedicated studies aimed at establishing robust structure-property relationships that will enable researchers to target and eliminate defects strategically. As part of the Superconducting Materials and Systems (SQMS) center, we have extensively employed state-of-the-art materials characterization techniques, including scanning/transmission electron microscopy, secondary ion mass spectrometry, atom probe tomography, x-ray diffraction, and x-ray photoelectron spectroscopy in conjunction with device measurements to elucidate such relationships. In this talk, I will discuss some of our recent findings, including linking atomic defects to microwave loss in surface oxides, linking impurities in the Josephson Junction to qubit parameters, and linking low temperature precipitates to device performance. By applying these insights, we have been able to strategically develop and implement mitigation strategies for reliable fabrication of high coherence superconducting qubits.

Murthy, A. [Fermilab] (ORCID:0000000176776866)

Identifying Decoherence Mechanisms in Superconducting Qubits through Advanced Materials Characterization

Although superconducting qubits have emerged as a leading technology platform for quantum computing through large improvements in device coherence times and gate fidelity in recent years, the presence of defects and impurities at the interfaces and surfaces in the constituent materials continue to limit performance and serve as a critical barrier in achieving scalable quantum systems. Understanding and eliminating these sources of quantum decoherence in superconducting qubit devices requires dedicated studies aimed at establishing robust structure-property relationships that will enable researchers to target and eliminate defects strategically. As part of the Superconducting Materials and Systems (SQMS) center, we have extensively employed state-of-the-art materials characterization techniques, including scanning/transmission electron microscopy, secondary ion mass spectrometry, atom probe tomography, x-ray diffraction, and x-ray photoelectron spectroscopy in conjunction with device measurements to elucidate such relationships. In this talk, I will discuss some of our recent findings, including linking atomic defects to microwave loss in surface oxides, linking impurities in the Josephson Junction to qubit parameters, and linking low temperature precipitates to device performance. By applying these insights, we have been able to strategically develop and implement mitigation strategies for reliable fabrication of high coherence superconducting qubits.

Murthy, A. [Fermilab] (ORCID:0000000176776866)