Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “error detection and error correction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

ATTNChecker: Highly-Optimized Fault Tolerant Attention for Large Language Model Training

Large Language Models (LLMs) have demonstrated remarkable performance in various natural language processing tasks. However, the training of these models is computationally intensive and susceptible to faults, particularly in the attention mechanism, which is a critical component of transformer-based LLMs. In this paper, we investigate the impact of faults on LLM training, focusing on INF, NaN, and near-INF values in the computation results with systematic fault injection experiments. We observe the propagation patterns of these errors, which can trigger non-trainable states in the model and disrupt training, forcing the procedure to load from checkpoints. To mitigate the impact of these faults, we propose ATTNChecker, the first Algorithm-Based Fault Tolerance (ABFT) technique tailored for the attention mechanism in LLMs. ATTNChecker is designed based on fault propagation patterns of LLM and incorporates performance optimization to adapt to both system reliability and model vulnerability while providing lightweight protection for fast LLM training. Evaluations on four LLMs show that ATTNChecker on average incurs on average 7% overhead on training while detecting and correcting all extreme errors. Compared with the state-of-the-art checkpoint/restore approach, ATTNChecker reduces recovery overhead by up to 49×.

Liang, Yuhang [University of Alabama - Birmingham]↗

3D Coded SUMMA: Communication-Efficient and Robust Parallel Matrix Multiplication

In this paper, we propose a novel fault-tolerant parallel matrix multiplication algorithm called 3D Coded SUMMA that achieves higher failure-tolerance than replication-based schemes for the same amount of redundancy. This work bridges the gap between recent developments in coded computing and fault-tolerance in high-performance computing (HPC). The core idea of coded computing is the same as algorithm-based fault-tolerance (ABFT), which is weaving redundancy in the computation using error-correcting codes. In particular, we show that MatDot codes, an innovative code construction for parallel matrix multiplications, can be integrated into three-dimensional SUMMA (Scalable Universal Matrix Multiplication Algorithm [30]) in a communication-avoiding manner. To tolerate any two node failures, the proposed 3D Coded SUMMA requires ~50% less redundancy than replication, while the overhead in execution time is only about 5–10%.

97 MATHEMATICS AND COMPUTING↗

Evaluating the mutagenicity of N-nitrosodimethylamine in 2D and 3D HepaRG cell cultures using error-corrected next generation sequencing

Human liver-derived metabolically competent HepaRG cells have been successfully employed in both two-dimensional (2D) and 3D spheroid formats for performing the comet assay and micronucleus (MN) assay. In the present study, we have investigated expanding the genotoxicity endpoints evaluated in HepaRG cells by detecting mutagenesis using two error-corrected next generation sequencing (ecNGS) technologies, Duplex Sequencing (DS) and High-Fidelity (HiFi) Sequencing. Both HepaRG 2D cells and 3D spheroids were exposed for 72 h to N-nitrosodimethylamine (NDMA), followed by an additional incubation for the fixation of induced mutations. NDMA-induced DNA damage, chromosomal damage, and mutagenesis were determined using the comet assay, MN assay, and ecNGS, respectively. The 72-h treatment with NDMA resulted in concentration-dependent increases in cytotoxicity, DNA damage, MN formation, and mutation frequency in both 2D and 3D cultures, with greater responses observed in the 3D spheroids compared to 2D cells. The mutational spectrum analysis showed that NDMA induced predominantly A:T → G:C transitions, along with a lower frequency of G:C → A:T transitions, and exhibited a different trinucleotide signature relative to the negative control. These results demonstrate that the HepaRG 2D cells and 3D spheroid models can be used for mutagenesis assessment using both DS and HiFi Sequencing, with the caveat that severe cytotoxic concentrations should be avoided when conducting DS. With further validation, the HepaRG 2D/3D system may become a powerful human-based metabolically competent platform for genotoxicity testing.

3D spheroids↗

Fault-tolerant operation and materials science with neutral atom logical qubits

We report on the fault-tolerant operation of logical qubits on a neutral atom quantum computer, with logical performance surpassing physical performance for multiple circuits including Bell state preparation (12x error reduction), random circuits (15x), and a prototype Anderson Impurity Model ground state solver for materials science applications (up to 6x, non-fault-tolerantly). The logical qubits are implemented via the [[4, 2, 2]] code (C 4 ). Our work constitutes the first complete realization of the benchmarking protocol proposed by Gottesman 2016 demonstrating results consistent with fault tolerance. In light of recent advances on applying concatenated C 4 /C 6 detection codes to achieve error correction with high code rates and thresholds, our work can be regarded as a building block towards a practical scheme for fault tolerant quantum computation. Our demonstration of a materials science application with logical qubits particularly demonstrates the immediate value of these techniques on current experiments.

36 MATERIALS SCIENCE↗

Spectral Properties and Coding Transitions of Haar-Random Quantum Codes

A quantum error-correcting code with a nonzero error threshold undergoes a mixed-state phase transition when the error rate reaches that threshold. We explore this phase transition for Haar-random quantum codes, in which the logical information is encoded in a random subspace of the physical Hilbert space. We focus on the spectrum of the encoded system density matrix as a function of the rate of uncorrelated, single-qudit errors. For low error rates, this spectrum consists of well-separated bands, representing errors of different weights. As the error rate increases, the bands for high-weight errors merge. The evolution of these bands with increasing error rate is well described by a simple analytic ansatz. Using this ansatz, as well as an explicit calculation, we show that the threshold for Haar-random quantum codes saturates the hashing bound, and thus coincides with that for random stabilizer codes. For error rates that exceed the hashing bound, typical errors are uncorrectable, but postselected error correction remains possible until a much higher detection threshold. Postselection can in principle be implemented by projecting onto subspaces corresponding to low-weight errors, which remain correctable past the hashing bound.

decoherence↗

Author Correction: US oil and gas system emissions from nearly one million aerial site measurements

Correction to: Naturehttps://doi.org/10.1038/s41586-024-07117-5 Published online 13 March 2024 In the version of the article initially published, several errors were present and have been corrected in the HTML and PDF versions of the article and Supplementary Information. The main results, conclusions, and our interpretations of the data remain unchanged. See the new Supplementary Information Section S15 for a more detailed description of the errors corrected and the resulting effects on the analysis. Data processing and methods corrections Overflight count correction: We previously used pre-computed source coverage data for some Carbon Mapper campaigns that was computed differently than was required for our analysis. We have re-computed Carbon Mapper source coverage based on flightline polygons and source coordinates. Transition point computation, well sites: The updated version now correctly compares the cumulative emissions distribution of simulated well site emissions with that of aerially detected sources (rather than plumes) when computing the transition point. Transition point computation, midstream: Additionally, the transition point calculation has been corrected to exclude aerially detected midstream emissions below the transition point, which was previously leading to double counting of these emissions. This error was not present for upstream (well site) emissions. Calculation errors Unit error: We corrected a specific unit conversion error affecting well site emissions in the Kairos Fort Worth dataset. Across all datasets, we also correct the conversion factor for converting from standard volume to mass for midstream emissions. Sorting error: We correct code that was applying incorrect sorting when computing correction factors to account for partial detection at well sites. Small typographical corrections were made in Fig. 1b and SI Section S4.1. Data processing and methods corrections Overflight count correction: We previously used pre-computed source coverage data for some Carbon Mapper campaigns that was computed differently than was required for our analysis. We have re-computed Carbon Mapper source coverage based on flightline polygons and source coordinates. Transition point computation, well sites: The updated version now correctly compares the cumulative emissions distribution of simulated well site emissions with that of aerially detected sources (rather than plumes) when computing the transition point. Transition point computation, midstream: Additionally, the transition point calculation has been corrected to exclude aerially detected midstream emissions below the transition point, which was previously leading to double counting of these emissions. This error was not present for upstream (well site) emissions. Calculation errors Unit error: We corrected a specific unit conversion error affecting well site emissions in the Kairos Fort Worth dataset. Across all datasets, we also correct the conversion factor for converting from standard volume to mass for midstream emissions. Sorting error: We correct code that was applying incorrect sorting when computing correction factors to account for partial detection at well sites. Small typographical corrections were made in Fig. 1b and SI Section S4.1. The following practices may help researchers conducting similar analyses avoid making similar errors: 1, Clear, accessible documentation explaining the interpretation of all columns in data input tables and all internal variables within the model, 2, Simple cross-check calculations computed before and after unit conversions.

Sherwin, Evan D↗

Analysis of Long-Term Quality Control Data for a 137 Cs Dosimetry Calibration Source

Strict quality assurance programs are required for many radiological applications, but these seldom exist for verifying dosimetry calibration sources. After initial characterization of a dosimetry calibration facility, quality control procedures are recommended to ensure the early detection of any changes or malfunctions. These also result in refined knowledge about average dose rate and experimental variations in dose delivery. This paper describes the implementation of a phase I quality control protocol for a 137 Cs dosimetry calibration source and includes an analysis of the resulting data collected over a 24-mo period. During this time, substantial data was collected to establish trial control limits. Air kerma rate measurements were obtained using an ion chamber and were adjusted for decay, corrected for ambient temperature, pressure and humidity, and then analyzed using quality control charts. Three variations of rational subgrouping methods were used in order to find assignable causes of error, and Nelson's Rules were followed to detect any non-random statistical variations. Measurements were subgrouped according to same-day measurements in order to detect positional errors as well as atmospheric correction errors. Additionally, measurements were subgrouped according to analogous experimental setups in order to detect failure in equipment or incorrect settings. Both were analyzed using the X-bar and R chart method. Similarly, individuals and moving ranges charts were used to carefully examine each position in order to observe any situational errors that may occur which include timing, positional, or interference errors. Each method was successful in identifying unique out-of-control data points that occurred during the phase I application of forming control limits. Furthermore, over the 24-mo period, enough data points were deemed in-control to establish reliable trial limits. Future experiments will include the phase II application of gaining more reliable measurements in order to fine-tune the limits, as well as performing a designed experiment, where variables are purposefully changed in order to test the variation of the data.

61 RADIATION PROTECTION AND DOSIMETRY↗

A robust dynamic state estimation approach against model errors caused by load changes

Dynamic state estimation (DSE) plays an important role in power system security monitoring and online control. In practice, there are two approaches to implementing DSE. The first approach is distributed DSE, which is based on the assumption that the terminal bus of each generator can be measured by PMUs (phasor measurement units). The assumption cannot be satisfied currently, however, because PMUs usually are installed at important high-voltage buses such as 500-kV buses installed in portions of the grid overseen by the Western Electricity Coordinating Council. Another issue of this approach is that performance of DSE is vulnerable to bad measurement data. The reason for this vulnerability is that DSE is performed separately through measurements at each terminal bus, and measurements at terminal buses are the only measurement upon which DSE can rely. Therefore, important redundant measurements are not included in this approach. The second approach is centralized DSE. This approach does not have the requirement for PMU location, and redundant measurements can be considered fully. However, load changes and grid topology changes impact centralized DSE. In this paper, we propose a new approach for handling the impact of load changes on DSE. We have developed a new algorithm that includes two sequential steps. In the first step, errors caused by load changes are detected by analyzing the difference between prediction results and measured results. In the second step, once model error is detected, a model optimization procedure is run to correct the error so the state estimation error can be mitigated. Simulation results from the IEEE 68 bus system show that the proposed approach can effectively handle model errors caused by load changes.

robust dynamic state estimation, load change, powe↗

MSU IETC ML for Modbus (AN EDGE)

This study explores machine learning for decoding Modbus RTU data using K-Nearest Neighbors (KNN) models. An initial KNN model trained on 8,000 packets achieved 95.15% accuracy. Although ML improves generalization, accuracy still falls short of deterministic methods. These findings have implications for Modbus traffic analysis, intrusion detection in industrial networks, and adaptive error correction in real-time monitoring systems. By refining ML-based decoding, future work could enable more efficient anomaly detection and predictive maintenance in industrial automation and cybersecurity applications.

Communication Protocol↗

Anomaly detection for MPC forecast in Fleet of Water Heaters

Among residential devices, water heaters consume 20% of home energy use in the United States. Water heaters possess the capability to store energy within their reservoirs, enabling the ability to decouple energy use from hot water use. This capability can be used to reduce energy usage and costs while also supporting grid services. This requires accurate forecasting of the parameters of the water heater such as upper and lower temperatures. In this study, we analyzed the performance and behavior of a water heater model used in the real-world to predict a control mechanism that is implemented in a smart residential neighborhood. The model forecasts are accurate in most cases but not all. In such scenarios, error correction of the model is necessary to further improve model predictive control accuracy. Anomaly detection is the first step of error correction. This study complements existing research by grouping time series data into two clusters one with anomalies and another without anomalies. To achieve this task, we explored and compared multiple unsupervised machine learning algorithms to perform clustering. Among these algorithms, Ward clustering has the lowest running time and identified the highest number of anomalies for the upper temperature limit. The proposed approach is tested based on the data collected in a neighborhood with 46 townhomes located in Atlanta, GA.

Lebakula, Viswadeep↗

Error-Detectable Bosonic Entangling Gates with a Noisy Ancilla

Bosonic quantum error correction has proven to be a successful approach for extending the coherence of quantum memories, but to execute deep quantum circuits, high-fidelity gates between encoded qubits are needed. To that end, we present a family of error-detectable two-qubit gates for a variety of bosonic encodings. From a new geometric framework based on a “Bloch sphere” of bosonic operators, we construct ZZ L ⁡(θ) and exponential-swap(θ) gates for the binomial, four-legged cat, dual-rail, and several other bosonic codes. The gate Hamiltonian is simple to engineer, requiring only a programmable beam splitter between two bosonic qubits and an ancilla dispersively coupled to one qubit. This Hamiltonian can be realized in circuit QED hardware with ancilla transmons and microwave cavities. The proposed theoretical framework was developed for circuit QED but is generalizable to any platform that can effectively generate this Hamiltonian. Crucially, one can also detect first-order errors in the ancilla and the bosonic qubits during the gates. We show that this allows one to reach error-detected gate fidelities at the 0.01% level with today’s hardware, limited only by second-order hardware errors.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Securing against errors in an error correcting code (ECC) implemented in an automotive system

In general, data is susceptible to errors caused by faults in hardware (i.e. permanent faults), such as faults in the functioning of memory and/or communication channels. To detect errors in data caused by hardware faults, the error correcting code (ECC) was introduced, which essentially provides a sort of redundancy to the data that can be used to validate that the data is free from errors caused by hardware faults. In some cases, the ECC can also be used to correct errors in the data caused by hardware faults. However, the ECC itself is also susceptible to errors, including specifically errors caused by faults in the ECC logic. A method, computer readable medium, and system are thus provided for securing against errors in an ECC.

Saxena, Nirmal Raj↗

A Logarithmic Bayesian Approach to Quantum Error Detection

We consider the problem of continuous quantum error correction from a Bayesian perspective, proposing a pair of digital filters using logarithmic probabilities that are able to achieve near-optimal performance on a three-qubit bit-flip code, while still being reasonable to implement on low-latency hardware. These practical filters are approximations of an optimal filter that we derive explicitly for finite time steps, in contrast with previous work that has relied on stochastic differential equations such as the Wonham filter. By utilizing logarithmic probabilities, we are able to eliminate the need for explicit normalization and can reduce the Gaussian noise distribution to a simple quadratic expression. The state transitions induced by the bit-flip errors are modeled using a Markov chain, which for log-probabilities must be evaluated using a LogSumExp function. We develop the two versions of our filter by constraining this LogSumExp to have either one or two inputs, which favors either simplicity or accuracy, respectively. Using simulated data, we demonstrate that the single-term and two-term filters are able to significantly outperform both a double threshold scheme and a linearized version of the Wonham filter in tests of error detection under a wide variety of error rates and time steps.

97 MATHEMATICS AND COMPUTING↗

Center-of-Mass Corrections in Associated Particle Imaging

Associated particle imaging (API) utilizes the inelastic scattering of neutrons produced in deuterium-tritium (DT) fusion reactions to obtain 3-D isotopic distributions within an object. The locations of the inelastic scattering centers are calculated by measuring the arrival time and position of the associated alpha particle produced in the fusion reactions, and the arrival time of the prompt gamma created in the neutron scattering event. While the neutron and its associated particle move in opposite directions in the center-of-mass (COM) system, in the laboratory system the angle is slightly less than 180°, and the COM movement must be taken into account in the reconstruction of the scattering location. Furthermore, the fusion reactions are produced by ions of different momenta, and thus the COM velocity varies, resulting in an uncertainty in the reconstructed positions. In this article, we analyze the COM corrections to this reconstruction by simulating the energy loss of beam ions in the target material and identifying sources of uncertainty in these corrections. Here we show that an average COM velocity calculated using the ion beam direction and energy can be used in the reconstruction and discuss errors as a function of ion beam energy, composition, and alpha detection location. When accounting for the COM effect, the mean of the reconstructed locations can be considered a correctable systematic error leading to a shift/tilt in the reconstruction. However, the distribution of reconstructed locations also have a spread that will introduce an error in the reconstruction that cannot be corrected. In this article, we will use the known stopping powers of ions in materials and reaction cross sections to examine the reconstruction uncertainties. We also discuss the impact of this effect on our API system.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Joint Channel Equalization and Symbol Detection for IoT Devices in Severe Multipath Channels

Internet of Things (IoT) devices often transmit very short message blocks without strong error correction coding or pilots for channel equalization. When the transmitted signals encounter a severe multipath channel, the receiver is often unable to equalize the resulting inter-symbol interference (ISI) via traditional methods, leading to a high retransmission rate. This paper proposes a joint channel equalization and symbol detection scheme for this scenario by utilizing the preamble sequence as a training pilot for equalization and the relatively weak cyclic redundancy check (CRC) 8-Dallas code for error correction. By assuming a sparse channel impulse response, the proposed joint channel equalization and symbol detection scheme formulates an l1 norm constrained optimization problem and solves it by a saddle-point algorithm. The proposed algorithm is applied to a set of real-world underwater animal tracking IoT data, and the results show improvement of correct message detection rate from 21:5% to 41:6%.

Jing, Shusen↗

Dual-rail encoding with superconducting cavities

The design of quantum hardware that reduces and mitigates errors is essential for practical quantum error correction (QEC) and useful quantum computation. To this end, we introduce the circuit-Quantum Electrodynamics (QED) dual-rail qubit in which our physical qubit is encoded in the single-photon subspace, { | 01 〉 , | 10 〉 } , of two superconducting microwave cavities. The dominant photon loss errors can be detected and converted into erasure errors, which are in general much easier to correct. In contrast to linear optics, a circuit-QED implementation of the dual-rail code offers unique capabilities. Using just one additional transmon ancilla per dual-rail qubit, we describe how to perform a gate-based set of universal operations that includes state preparation, logical readout, and parametrizable single and two-qubit gates. Moreover, first-order hardware errors in the cavities and the transmon can be detected and converted to erasure errors in all operations, leaving background Pauli errors that are orders of magnitude smaller. Hence, the dual-rail cavity qubit exhibits a favorable hierarchy of error rates and is expected to perform well below the relevant QEC thresholds with today’s coherence times.

97 MATHEMATICS AND COMPUTING↗

Error detection on quantum computers improving the accuracy of chemical calculations

The ultimate goal of quantum error correction is to achieve the fault-tolerance threshold beyond which quantum computers can be made arbitrarily accurate. This requires extraordinary resources and engineering efforts. We show that even without achieving full fault-tolerance, quantum error detection is already useful on the current generation of quantum hardware. We demonstrate this by executing an end-to-end chemical calculation for the hydrogen molecule encoded in the [[4, 2, 2]] quantum error-detecting code. The encoded calculation with logical qubits significantly improves the accuracy of the molecular ground-state energy.

97 MATHEMATICS AND COMPUTING↗

GEDet: Detecting Erroneous Nodes with A Few Examples

Detecting nodes with erroneous values in real-world graphs re- mains challenging due to the lack of examples and various error scenarios. We demonstrate GEDet, an error detection engine that can detect erroneous nodes in graphs with a few examples. The GEDet framework tackles error detection as a few-shot node classification problem. We invite the attendees to experience the following unique features. (1) Few-shot detection. Users only need to provide a few examples of erroneous nodes to perform error detection with GEDet. GEDet achieves desirable accuracy with (a) a graph augmentation module, which automatically generates synthetic examples to learn the classifier, and (b) an adversarial detection module, which improves classifiers to better distinguish erroneous nodes from both cleaned nodes and synthetic examples. We show that GEDet significantly improves the state-of-the-art error detection methods. (2) Diverse error scenarios. GEDet profiles data errors with a built-in library of transformation functions from correct values to errors. Users can also easily “plug in” new error types or examples. (3) User-centric detection. GEDet supports (a) an active learning mode to engage users to verify detected results, and adapts the error detection process accordingly; and (b) visual interfaces to interpret and track detected errors.

Guan, Sheng↗