Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “error mitigation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Improving UAVSAR results with GPS, radiometry, and QUAKES topographic imager

UAVSAR is NASA’s airborne interferometricsynthetic aperture radar (InSAR) platform. The instrument has been used to detect deformation from earthquakes, volcanoes, oil pumping, landslides, water withdrawal, landfill compaction, and glaciers. It has been used to detect scars from wildfires and damage from debris flows. The instrument performs well for large changes or for local small changes. Determining subtle changes over large areas requires improved instrumentation and processing. We are working to improve the utility of UAVSAR by including GPS station position results in the processing chain, and adding a topographic imager to improve estimates of topography, 3D change, and damage. We are also exploring the benefit of microwave radiometry to mitigating error from water vapor path delay. A goal is to determine 3D tectonic deformation to millimeters per year at ~100 km plate boundary scales and to understand surface processes in areas of decorrelated radar imagery.

Muellerschoen, Ronald↗

Hardware-Efficient Quantum Optimization Layered Algorithms and Experiments

Quantum optimization algorithms, such as QAOA, that implement parametrized stochastic optimization solvers attempt to identify low-energy solutions of Ising systems by exploiting available quantum effects in noisy-intermediate scale machines. Engineering a well-performing parametrized quantum optimization circuit is indeed an exercise in balancing the trade-off between expressivity and implementation complexity. We show that, for MaxCut QAOA circuits defined on native hardware topology (Rigetti’s Aspen Quantum Processors), error-mitigation techniques recover simulated features of the noiseless theory. Moreover, we explore a design space for QAOA-like ansatze that perform well in theory as well as in hardware for fully-connected problems. We also discuss how efficient coherence and entanglement detection methods that could be coupled with quantum optimization experiments require only linear overhead in benchmarking time.

quantum computing↗

Assessing and Advancing the Potential of Quantum Computing: A NASA Case Study

Quantum computing is one of the most enticing computational paradigms with the potential to revolutionize diverse areas of future-generation computational systems. While quantum computing hardware has advanced rapidly, from tiny laboratory experiments to quantum chips that can outperform even the largest supercomputers on specialized computational tasks, these noisy- intermediate scale quantum (NISQ) processors are still too small and non-robust to be directly useful for any real-world applications. In this paper, we describe NASA’s work in assessing and advancing the potential of quantum computing. We discuss advances in algorithms, both near- and longer-term, and the results of our explorations on current hardware as well as with simulations, including illustrating the benefits of algorithm-hardware codesign in the NISQ era. This work also includes physics-inspired classical algorithms that can be used at application scale today. We discuss innovative tools supporting the assessment and advancement of quantum computing, and describe improved methods for simulating quantum systems of various types on high performance computing systems that incorporate realistic error models. We provide an overview of recent methods for benchmarking, evaluating, and characterizing quantum hardware for error mitigation, computational purposes.

quantum computing↗

Quantum Networks for High Energy Physics

Quantum networks of quantum objects promise to be exponentially more powerful than the objects considered independently. To live up to this promise will require the development of error mitigation and correction strategies to preserve quantum information as it is initialized, stored, transported, utilized, and measured. The quantum information could be encoded in discrete variables such as qubits, in continuous variables, or anything in-between. Quantum computational networks promise to enable simulation of physical phenomena of interest to the HEP community. Quantum sensor networks promise new measurement capability to test for new physics and improve upon existing measurements of fundamental constants. Such networks could exist at multiple scales from the nano-scale to a global-scale quantum network.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Is Knowledge About Running Applications Helping Improve Runtime Prediction of HPC Jobs?

High-performance computing systems rely upon scheduling algorithms to achieve high utilization. These schedulers rely upon user estimates of job resource requirements, such as runtime, to determine optimal scheduling of incoming jobs. These user estimates, however, are prone to error. To mitigate this error, significant research has been directed at providing better estimates of job runtime, usually employing machine learning techniques. These techniques are dependent upon the input features selected. Among the possible features is the primary application used by the job. In a survey of more than 20 papers directed at improving runtime prediction, only four included primary application as an input feature. We focus this investigation specifically on the value of adding primary application as an input feature, and find that it does improve model performance, especially for jobs with longer runtimes, though this improvement varies based on the application used. We recommend further research to determine the cause of this variability as well as an optimal strategy for employing a mixture of models both including and not including primary application as a feature.

feature selection↗

HydroDCM: Hydrological Domain-Conditioned Modulation for Cross-Reservoir Inflow Prediction

Deep learning models have shown promise in reservoir inflow prediction, yet their performance often deteriorates when applied to different reservoirs due to distributional differences, referred to as the domain shift problem. Domain generalization (DG) solutions aim to address this issue by extracting domain-invariant representations that mitigate errors in unseen domains. However, in hydrological settings, each reservoir exhibits unique inflow patterns, while some metadata beyond observations like spatial information exerts indirect but significant influence. This mismatch limits the applicability of conventional DG techniques to many-domain hydrological systems. To overcome these challenges, we propose HydroDCM, a scalable DG framework for cross-reservoir inflow forecasting. Spatial metadata of reservoirs is used to construct pseudo-domain labels that guide adversarial learning of invariant temporal features. During inference, HydroDCM adapts these features through light-weight conditioning layers informed by the target reservoir’s metadata, reconciling DG’s invariance with location-specific adaptation. Experiment results on 30 real-world reservoirs in the Upper Colorado River Basin demonstrate that our method substantially outperforms state-of-the-art DG baselines under many-domain conditions and remains computationally efficient.

Hu, Pengfei [ORNL] (ORCID:0009000367130950)↗

Real-Time Detection of Charge Jumps in Superconducting Qubits with a Convolutional Neural Network

Ionizing radiation from cosmic rays and gammas can induce discontinuous jumps in the environmental charge of superconducting qubits (charge jumps), causing correlated errors that challenge fault-tolerant quantum computing while simultaneously providing a detection signature for quantum sensing applications. Current detection methods operate offline, introducing latency incompatible with in-the-loop qubit control. In this paper, an online detector of charge jumps for superconducting qubits, based on a dilated causal convolutional neural network (DCCNN) designed for in-the-loop deployment on the Quantum Instrumentation Control Kit (QICK) platform, is presented. The network is trained on synthetic Ramsey tomography scans generated from qubit templates measured at the Northwestern Experimental Underground Site (NEXUS) at Fermilab, and translated to FPGA firmware via hls4ml with ap_fixed$\langle 16,6 \rangle$ quantization, reaching a per-inference latency of $6.19 μ$s on the Zynq UltraScale+ RFSoC ZCU216. At this operating point the DCCNN matches the detection efficiency of the established offline $χ^2$ algorithm ($0.843 \pm 0.022$ vs. $0.866 \pm 0.020$ on $|Δq| \in [0.1, 0.5] e$ at matched false-positive rate), while requiring no per-qubit hyperparameter tuning. This shifts charge-jump detection from a post-hoc diagnostic to a control-loop primitive, enabling adaptive protocols that respond to radiation-induced events in situ, with applications to quantum-computing error mitigation and to the use of superconducting qubits as particle detectors.

Gaytan-Villarreal, Daniel [Carnegie Mellon U.]↗

Error Sources and Mitigation Strategies for Thermocouples Integrated in Flexible Thermal Protection System Materials

Brief Presenter Biography:Ruth Miller is an aer-ospace systems engineer in the Entry Systems and Ve-hicle Development Branch at NASA Ames Research Center.Introduction:The flexible thermal protection sys-tem (FTPS) on NASA’s Low-Earth Orbit Flight Test of an Inflatable Decelerator (LOFTID) vehicle will be in-strumented with thermocouples (TCs) to measure the in-depth thermal response during entry into Earth’s atmos-phere[1, 2]. Accurate flight temperature measurements are critical for verifying vehicle performance during the flight test and reducing uncertainties in the thermal models.However,the deployable nature of inflatable decelerator technology presents challengesfor integrat-ing TCs, specifically the TCsneed to be compactableand cannot damage the FTPSnor the inflatable structure(IS).Unlike traditional rigid aeroshells, routing TCsthrough the thickness of the FTPS could cause signifi-cant damage during packing of the deployable aeroshellbecause the different layers may shift small amounts in relation to each other imparting strain on the TCsand FTPS materials. The LOFTID TCleads are routed from the measurement location back to the data acquisition system in the vehicle centerbody within the same FTPS layer that they are monitoring the temperature. This ap-proach eliminates the need to put holes in the FTPS lay-ers,butas a consequence,the insulated TCleads travel for an appreciable distance through a region that will expose them to high temperaturesand large thermal gra-dients.LOFTID’s TCswere baselined to be commercially available Type K TCswith a binder impregnated glass braid insulation. These TCswere chosen because they had been used successfully on IRVE-3 and in ground-based arc jet testing. Additionally, these TCsdid notdamage the FTPS nor IS during packing and deploy-ment testing. However, the glass braid insulation is only rated to a maximum continuous use temperature of 482ºC. For reference, the duration of the heat pulse on the LOFTID vehicle is on the order of minutes and themaximum predicted temperaturebeneaththe outermost FTPS layersis 1350ºC.Ground-based testing in a tube furnace at NASA Ames Research Centerwas conducted to determine if the baseline TCsrouted through FTPS samples would survive and provide accurate temperature measure-mentsat LOFTID flight-relevant temperatures[3].The test results showed large measurement errors occurred beginning at approximately 400°C due to conductive deposits on the TCinsulation electrically shorting the TCleads.The conductive deposits and thus electrical shorting weredetermined to be caused by twoerror sources:1.The organic binder on the TCinsulation carbon-izingin a high temperature, low oxygen envi-ronment2.Decomposition products from the FTPSperme-atingthe braided TC insulationFurther testing in the tube furnace demonstrated that heat cleaning the TCinsulation effectivelyremovedthe organic binder andeliminatedthe first error source.The second error sourcewas shown to be mitigated by the addition of amica wraparound each individual TCleadtoact as an impermeable barrier.To understand the applicability of theground-based tube furnace test results to flight,anarc jet testseriesat Boeing’s Large Core Arc Tunnel (LCAT) facilitywas conducted[4].The error sourcesand mitigation strate-giesidentified in the tube furnace testing were substan-tiatedin the arc jet testing.However, the arc jet testing also revealed three new error sources:1.Glass TCinsulation meltingwhich resultsin electrical shorting of the TCeither through di-rect contact between the two leads or through the electrically conductive FTPS materials2.TCwire meltingwhich results in a noisy and/or open-loop TCresponse3.Type K TCwire green-rotwhich results in large calibration errors Scope of the Presentation:This presentation will include a brief discussion onthe effect of electrical shorting on the output of a TC(i.e. how to identify elec-trical shorting in TCdataand what the associated erroris).The tube furnace and arc jet test resultswill be dis-cussedand the solutions LOFTID is implementing to mitigate the error sourcesidentified in the tube furnace and arc jet testingwill be presented.Additionally, futureresearch and development work to eliminateTCerror sourcesfor future missionswill be recommended.

R A Miller↗

diffReplication - An Energy-Aware Fault Tolerance Model for Silent Error Detection and Mitigation in Heterogeneous Extreme-scale Computing Environment

At extreme scale, the frequency of silent errors – a class of errors that remain undetected by low-level error detection mechanisms – increases significantly with the computational complexity of the application and the scale of the computing infrastructure. As hardware and software advances are made to usher in the next scientific era of computing, developing new approaches to mitigate the impact of silent errors remains a challenging problem. In this work, we propose an energy-aware fault-tolerance model, referred to diffReplication to overcome silent errors. In the proposed model, the main process is associated with one replica that executes at the same rate as the main process, and one diffReplica that is executed at a fraction of the main process' execution rate. If the main and its replica reach consensus at the end of a computation phase, the state of the diffReplica is updated and computation is resumed. If the synchronization attempt results in a disagreement, however, the diffReplica increases its execution speed to complete the computation and quickly reach the synchronization barrier. Assuming a single error over any given synchronization interval, a majority voting is used to reach consensus and tolerate silent errors. To further enhance its performance, diffReplication is augmented with speculative execution, whereby the main or its fast replica is selected to continue execution without waiting for the diffReplica. The selection process is based on the previous behaviour of the main and its replica. A performance analysis study is carried out to assess the performance of diffReplication, in terms of the energy saving and time-to-completion reduction achieved by the diffReplication scheme. The experiment shows that speculative execution reduces the time to completion with additional energy, and dynamic decision-making balances the energy consumption and time to completion.

97 MATHEMATICS AND COMPUTING↗

AEDAM: Whole Program Adaptive Error Detection and Mitigation (Final Report)

The overall goals of the AEDAM project were to fundamentally transform software transient-error detection through the design of configurable specialized detectors, quantitative characterization of hardware resilience and software vulnerabilities, and composition of the specialized detectors to most efficiently protect the whole program. Within this scope, UT Austin's research contribution related to enabling and studying the composition of detectors and possible different hardware errors, specifically: (1) developed the Hamartia open-source error injection framework that is designed to make composition studies simple, and (2) develop the methodology and demonstrate the potential benefits of error detector composition.

97 MATHEMATICS AND COMPUTING↗

Realtime mitigation of GPS SA errors using Loran-C

The hybrid use of Loran-C with the Global Positioning System (GPS) was shown capable of providing a sole-means of enroute air radionavigation. By allowing pilots to fly direct to their destinations, use of this system is resulting in significant time savings and therefore fuel savings as well. However, a major error source limiting the accuracy of GPS is the intentional degradation of the GPS signal known as Selective Availability (SA). SA-induced position errors are highly correlated and far exceed all other error sources (horizontal position error: 100 meters, 95 percent). Realtime mitigation of SA errors from the position solution is highly desirable. How that can be achieved is discussed. The stability of Loran-C signals is exploited to reduce SA errors. The theory behind this technique is discussed and results using bench and flight data are given.

Braasch, Soo Y.↗

Error Characterization and Mitigation for 16Nm MLC NAND Flash Memory Under Total Ionizing Dose Effect

A data device includes a memory having a plurality of memory cells configured to store data values in accordance with a predetermined rank modulation scheme that is optional and a memory controller that receives a current error count from an error decoder of the data device for one or more data operations of the flash memory device and selects an operating mode for data scrubbing in accordance with the received error count and a program cycles count.

Li, Yue↗

Analyzing Software Specifications for Mode Confusion Potential

Increased automation in complex systems has led to changes in the human controller's role and to new types of technology-induced human error. Attempts to mitigate these errors have primarily involved giving more authority to the automation, enhancing operator training, or changing the interface. While these responses may be reasonable under many circumstances, an alternative is to redesign the automation in ways that do not reduce necessary or desirable functionality or to change functionality where the tradeoffs are judged to be acceptable. This paper describes an approach to detecting error-prone automation features early in the development process while significant changes can still be made to the conceptual design of the system. The information about such error-prone features can also be useful in the design of the operator interface, operational procedures, or operator training.

Nancy G. Leveson↗

Error Sources and Mitigation Strategies for Thermocouples Integrated in Flexible Thermal Protection System Materials

• The flexible thermal protection system (FTPS) on NASA’s Low-Earth Orbit Flight Test of an Inflatable Decelerator (LOFTID) vehicle will be instrumented with thermocouples (TCs) to measure the in-depth thermal response during entry into Earth’s atmosphere. Accurate flight temperature measurements are critical for verifying vehicle performance and reducing thermal model uncertainties. • The TC leads are routed from the measurement location to the data acquisition system within the same FTPS layer that they are monitoring the temperature. This approach eliminates the need to put holes in the FTPS layers. However, the insulated TC leads are exposed to high temperatures and large thermal gradients. • The baseline TCs were commercially available 30 AWG Type K TCs with a binder impregnated glass braid insulation. The glass braid is rated to a maximum continuous use temperature of 482°C. LOFTID’s heat pulse will be on the order of minutes and the maximum predicted temperature beneath the outermost FTPS layers is 1350°C.

R. A. Miller↗

Masked fault detection for reliable low voltage cache operation

Systems, apparatuses, and methods for implementing masked fault detection for reliable low voltage cache operation are disclosed. A processor includes a cache that can operate at a relatively low voltage level to conserve power. However, at low voltage levels, the cache is more likely to suffer from bit errors. To mitigate the bit errors occurring in cache lines at low voltage levels, the cache employs a strategy to uncover masked faults during runtime accesses to data by actual software applications. For example, on the first read of a given cache line, the data of the given cache line is inverted and written back to the same data array entry. Also, the error correction bits are regenerated for the inverted data. On a second read of the given cache line, if the fault population of the given cache line changes, then the given cache line's error protection level is updated.

Ganapathy, Shrikanth↗

Application of Artificial Intelligence in Detection and Mitigation of Human Factor Errors in Nuclear Power Plants: A Review

Human factors and ergonomics have played an essential role in increasing the safety and performance of operators in the nuclear energy industry. In this critical review, we examine how artificial intelligence (AI) technologies can be leveraged to mitigate human errors, thereby improving the safety and performance of operators in nuclear power plants (NPPs). First, we discuss the various causes of human errors in NPPs. Next, we examine the ways in which AI has been introduced to and incorporated into different types of operator support systems to mitigate these human errors. We specifically examine (1) operator support systems, including decision support systems, (2) sensor fault detection systems, (3) operation validation systems, (4) operator monitoring systems, (5) autonomous control systems, (6) predictive maintenance systems, (7) automated text analysis systems, and (8) safety assessment systems. Finally, we provide some of the shortcomings of the existing AI technologies and discuss the challenges still ahead for their further adoption and implementation to provide future research directions.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Enhanced MPM framework with multipatch isogeometric analysis for geotechnical applications

Achieving stable stress solutions at large strains using the Material Point Method (MPM) is challenging due to the accumulation of errors associated with geometry discretization, cell-crossing noise, and volumetric locking. Several simplified attempts exist in the literature to mitigate these errors, including higher-order frameworks. However, the stability of the MPM solution in such frameworks has been limited to simple geometries and the single-phase formulation (i.e., neglecting pore fluid). Although never explored, multipatch isogeometric analysis offers desirable qualities to simulate complex geometries while mitigating errors in the MPM. The degree of required high-order spatial integration has also never been investigated to infer a minimum limit for the stability of the stress solution in MPM. This paper presents a general-purpose numerical framework for simulating stable stresses in porous media, capturing both near incompressibility and multiphase interactions. First, the numerical framework is presented considering Non-Uniform Rational B-splines (NURBS) to perform isogeometric analysis (IGA) in MPM. Additionally, a volumetric strain smoothing algorithm is used to alleviate errors associated with volumetric locking. Second, the manifestation of cell-crossing errors is assessed via a series of problems with orders ranging from linear to cubic interpolation functions. Third, the use of NURBS is investigated and verified for problems with circular geometries. Finally, multipatch analysis is deployed to simulate plane strain and 3D penetration in soils, considering nearly incompressible elastoplastic (total stress) analysis and fully-coupled hydro-mechanical (effective stress) analysis. The stability of the solution is also analyzed for different constitutive models. From the results, it can be concluded that the framework using cubic interpolation functions with strain smoothing is the most convenient, presenting stable stress solutions for a broad range of multiphase geotechnical applications.

58 GEOSCIENCES↗