Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “reading error”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Reading Error Analysis (REA) Software Version 1

This is training presentation for REA (Reading Error Analysis) software. REA is designed for analyzing UGT scope reading errors and for developing reading points optimization algorithms.

97 MATHEMATICS AND COMPUTING↗

An Accurate, Error-Tolerant, and Energy-Efficient Neural Network Inference Engine Based on SONOS Analog Memory

In this work, we demonstrate SONOS (silicon-oxide-nitrideoxide- silicon) analog memory arrays that are optimized for neural network inference. The devices are fabricated in a 40nm process and operated in the subthreshold regime for in-memory matrix multiplication. Subthreshold operation enables low conductances to be implemented with low error, which matches the typical weight distribution of neural networks, which is heavily skewed toward near-zero values. This leads to high accuracy in the presence of programming errors and process variations. We simulate the end-to-end neural network inference accuracy, accounting for the measured programming error, read noise, and retention loss in a fabricated SONOS array. Evaluated on the ImageNet dataset using ResNet50, the accuracy using a SONOS system is within 2.16% of floating-point accuracy without any retraining. The unique error properties and high On/Off ratio of the SONOS device allow scaling to large arrays without bit slicing, and enable an inference architecture that achieves 20 TOPS/W on ResNet50, a >10× gain in energy efficiency over state-of-the-art digital and analog inference accelerators.

97 MATHEMATICS AND COMPUTING↗

CMS HGCAL ECON-D ASIC : Impact of CMOS fabrication process tuning on performance and radiation tolerance

The CMS experiment’s High Granularity Calorimeter (HGCAL) upgrade will replace CMS’s existing endcap calorimeters in preparation for the High Luminosity LHC. To effectively use over 6 million channels of this “imaging”calorimeter, CMS has developed two novel Endcap Concentrator (ECON) ASICs to perform data compression/selection on detector. The ECON-D ASIC operates on the 750kHz data path, and the ECON-T ASIC on the 40MHz trigger path. These 65 nm CMOS ASICs are radiation tolerant to 200 Mrad and low power, operating at less than 2.5 mW/channel. The first full-functionality prototype ECONs were produced and characterized in 2021-23, and an initial engineering run was performed in 2024. ECON-D radiation testing for the engineering run revealed that the chip’s internal SRAMs produce intermittent read errors for a non-negligible fraction of chips. Further investigation indicated that the SRAM performance is highly sensitive to the exact parameters of the CMOS fabrication process. To both study this process sensitivity and mitigate SRAM performance issues, twenty ECON wafers were produced in 2025 with a range of doping concentrations designed to tune the underlying transistor threshold voltage by 0%, 5%, 10%, and 15% from nominal. This talk will present first measurements of ECON-D performance, power consumption, and radiation tolerance for these four variations of CMOS process.

Syal, Chinar [Fermilab]↗

Radiation-Induced Error Mitigation by Read-Retry Technique for MLC 3-D NAND Flash Memory

Here, we have evaluated the Read-Retry (RR) functionality of the 3-D NAND chip of multilevel-cell (MLC) configuration after total ionization dose (TID) exposure. The RR function is typically offered in the high-density state-of-the-art NAND memory chips to recover data once the default memory read method fails to correct data with error correction codes (ECCs). In this work, we have applied the RR method on the irradiated 3-D NAND chip that was exposed with a Co-60 gamma-ray source for TID up to 50 krad (Si). Based on our experimental evaluation results, we have proposed an algorithm to efficiently implement the RR method to extend the radiation tolerance of the NAND memory chip. Our experimental evaluation shows that the RR method coupled with ECC can ensure data integrity of MLC 3-D NAND for TID up to 50 krad (Si).

3-D NAND↗

Techniques for storing data to enhance recovery and detection of data corruption errors

Often there are errors when reading data from computer memory. To detect and correct these errors, there are multiple types of error correction codes. Disclosed is an error correction architecture that creates a codeword having a data portion and an error correction code portion. Swizzling rearranges the order of bits and distributes the bits among different codewords. Because the data is redistributed, a potential memory error of up to N contiguous bits, where N for example equals 2 times the number of codewords swizzled together, only affects up to, at most, two bits per swizzled codeword. This keeps the error within the error detecting capabilities of the error correction architecture. Furthermore, this can allow improved error correction and detection without requiring a change to error correcting code generators and checkers.

Mills, Peter↗

Techniques for storing data to enhance recovery and detection of data corruption errors

Often there are errors when reading data from computer memory. To detect and correct these errors, there are multiple types of error correction codes. Disclosed is an error correction architecture that creates a codeword having a data portion and an error correction code portion. Swizzling rearranges the order of bits and distributes the bits among different codewords. Because the data is redistributed, a potential memory error of up to N contiguous bits, where N for example equals 2 times the number of codewords swizzled together, only affects up to, at most, two bits per swizzled codeword. This keeps the error within the error detecting capabilities of the error correction architecture. Furthermore, this can allow improved error correction and detection without requiring a change to error correcting code generators and checkers.

Mills, Peter↗

Improving precision and accuracy of genetic mapping with genotyping‐by‐sequencing data in outcrossing species

Abstract Genotyping‐by‐sequencing (GBS) is a widely used strategy for obtaining large numbers of genetic markers in model and non‐model organisms. In crop plants, GBS‐derived marker datasets are frequently used to perform quantitative trait locus (QTL) mapping. In some plant species, however, high heterozygosity and complex genome structure mean that researchers must use care in handling GBS data to conduct QTL mapping most effectively. Such outbred crops include most of the perennial grass and tree species used for bioenergy. To identify strategies for increasing accuracy and precision of QTL mapping using GBS data in outbred crops, we conducted an empirical study of SNP‐calling and genetic map‐building pipeline parameters in a Miscanthus sinensis population, and a complementary simulation study to estimate the relationship between genome‐wide error rate, read depth, and marker number. The bioenergy grass Miscanthus is an obligate outcrossing species with a recent (diploidized) whole‐genome duplication. For the study of empirical M. sinensis data, we compared two SNP‐calling methods (one non‐reference‐based and one reference‐based), a series of depth filters (12×, 20×, 30×, and 40×) and two map‐construction methods (i.e., marker ordering: linkage‐only and order‐corrected based on a reference genome). We found that correcting the order of markers on a linkage map by using a high‐quality reference genome improved QTL precision (shorter confidence intervals). For typical GBS datasets of between 1000 and 5000 markers to build a genetic map for biparental populations, a depth filter set at 30× to 40× applied to outbred populations provided a genome‐wide genotype‐calling error rate of less than 1%, improved accuracy of QTL point estimates and minimized type I errors for identifying QTL. Based on these results, we recommend using a reference genome to correct the marker order of genetic maps and a robust genotype depth filter to improve QTL mapping for outbred crops.

59 BASIC BIOLOGICAL SCIENCES↗

VirION2: a short- and long-read sequencing and informatics workflow to study the genomic diversity of viruses in nature

Microbes play fundamental roles in shaping natural ecosystem properties and functions, but do so under constraints imposed by their viral predators. However, studying viruses in nature can be challenging due to low biomass and the lack of universal gene markers. Though metagenomic short-read sequencing has greatly improved our virus ecology toolkit—and revealed many critical ecosystem roles for viruses—microdiverse populations and fine-scale genomic traits are missed. Some of these microdiverse populations are abundant and the missed regions may be of interest for identifying selection pressures that underpin evolutionary constraints associated with hosts and environments. Though long-read sequencing promises complete virus genomes on single reads, it currently suffers from high DNA requirements and sequencing errors that limit accurate gene prediction. Here we introduce VirION2, an integrated short- and long-read metagenomic wet-lab and informatics pipeline that updates our previous method (VirION) to further enhance the utility of long-read viral metagenomics. Using a viral mock community, we first optimized laboratory protocols (polymerase choice, DNA shearing size, PCR cycling) to enable 76% longer reads (now median length of 6,965 bp) from 100-fold less input DNA (now 1 nanogram). Using a virome from a natural seawater sample, we compared viromes generated with VirION2 against other library preparation options (unamplified, original VirION, and short-read), and optimized downstream informatics for improved long-read error correction and assembly. VirION2 assemblies combined with short-read based data (‘enhanced’ viromes), provided significant improvements over VirION libraries in the recovery of longer and more complete viral genomes, and our optimized error-correction strategy using long- and short-read data achieved 99.97% accuracy. In the seawater virome, VirION2 assemblies captured 5,161 viral populations (including all of the virus populations observed in the other assemblies), 30% of which were uniquely assembled through inclusion of long-reads, and 22% of the top 10% most abundant virus populations derived from assembly of long-reads. Viral populations unique to VirION2 assemblies had significantly higher microdiversity means, which may explain why short-read virome approaches failed to capture them. These findings suggest the VirION2 sample prep and workflow can help researchers better investigate the virosphere, even from challenging low-biomass samples. Our new protocols are available to the research community on protocols.io as a ‘living document’ to facilitate dissemination of updates to keep pace with the rapid evolution of long-read sequencing technology.

Long-reads↗

polishCLR: A Nextflow Workflow for Polishing PacBio CLR Genome Assemblies

Long-read sequencing has revolutionized genome assembly, yielding highly contiguous, chromosome-level contigs. However, assemblies from some third generation long read technologies, such as Pacific Biosciences (PacBio) continuous long reads (CLR), have a high error rate. Such errors can be corrected with short reads through a process called polishing. Although best practices for polishing non-model de novo genome assemblies were recently described by the Vertebrate Genome Project (VGP) Assembly community, there is a need for a publicly available, reproducible workflow that can be easily implemented and run on a conventional high performance computing environment. Here, we describe polishCLR (https://github.com/isugifNF/polishCLR), a reproducible Nextflow workflow that implements best practices for polishing assemblies made from CLR data. PolishCLR can be initiated from several input options that extend best practices to suboptimal cases. It also provides re-entry points throughout several key processes, including identifying duplicate haplotypes in purge_dups, allowing a break for scaffolding if data are available, and throughout multiple rounds of polishing and evaluation with Arrow and FreeBayes. PolishCLR is containerized and publicly available for the greater assembly community as a tool to complete assemblies from existing, error-prone long-read data.

59 BASIC BIOLOGICAL SCIENCES↗

Determining diagnostic coverage for memory using redundant execution

Memory, used by a computer to store data, is generally prone to faults, including permanent faults (i.e. relating to a lifetime of the memory hardware), and also transient faults (i.e. relating to some external cause) which are otherwise known as soft errors. Since soft errors can change the state of the data in the memory and thus cause errors in applications reading and processing the data, there is a desire to characterize the degree of vulnerability of the memory to soft errors. In particular, once the vulnerability for a particular memory to soft errors has been characterized, cost/reliability trade-offs can be determined, or soft error detection mechanisms (e.g. parity) may be selectively employed for the memory. In some cases, memory faults can be diagnosed by redundant execution and a diagnostic coverage may be determined.

Bramley, Richard Gavin↗

Liveness as a factor to evaluate memory vulnerability to soft errors

Memory, used by a computer to store data, is generally prone to faults, including permanent faults (i.e. relating to a lifetime of the memory hardware), and also transient faults (i.e. relating to some external cause) which are otherwise known as soft errors. Since soft errors can change the state of the data in the memory and thus cause errors in applications reading and processing the data, there is a desire to characterize the degree of vulnerability of the memory to soft errors. In particular, once the vulnerability for a particular memory to soft errors has been characterized, cost/reliability trade-offs can be determined, or soft error detection mechanisms (e.g. parity) may be selectively employed for the memory. In some cases, memory faults can be diagnosed by redundant execution and a diagnostic coverage may be determined.

Bramley, Richard Gavin↗

NovaDemux v39.07

This program is a sequence demultiplexer intended primarily for, but not limited to, Illumina sequencing machines. Typically, multiple experiments ("libraries") are pooled together and sequenced at once, with genetic molecules of these libraries tagged with a synthetic DNA "barcode". After sequencing, the data is demultiplexed into one file per library based on the barcode. However, errors in barcode reading cause misassignment and decrease yield. NovaDemux uses advanced statistical methods to maximize yield while minimizing misassignment compared to existing software.

Bushnell, Brian [Lawrence Berkeley National Labora↗

A Conceptual Design for a Mobile Application to Support Infield Inventory Activities

This paper introduces an inventory assistant being developed at Oak Ridge National Laboratory (ORNL) that we believe will empower users to perform inventory activities at nuclear facilities more accurately, reliably, and quickly. Inventory activities at nuclear facilities are often conducted using pen and paper, which can be time-consuming, tedious, and susceptible to reading or transcription errors. The proposed inventory assistant would replace the paper-based process used by International Atomic Energy Agency (IAEA) inspectors, nuclear facility operators, or verification monitors to complete an inventory of nuclear and non-nuclear items. In general, the inventory assistant would ingest an inventory list, distribute assigned items from the inventory list to one or more mobile devices, enable inventory teams to record their observations in the field, and then enable an inventory lead to integrate and reconcile the observations to produce a final report. The assistant consists of two software components—one for the inventory teams to record observations in the field (In-Field Observations App [IFOA]) and one for the inventory lead to reconcile the inventory list with observations (Distribution, Integration, and Reconciliation Application [DIRA]). This paper introduces the overall workflow of the inventory assistant and describes the IFOA user experience in more detail. To demonstrate the concept, the authors present a use case of IAEA inspectors conducting item counting and tag checking activities of UF6 cylinders at a gas centrifuge enrichment plant with a large number of UF6 cylinders (e.g., thousands). These activities can currently require 30–40 person-days of inspection to complete. Based on experiences during an exercised performed at the IAEA by the ORNL team in 2016, we believe an inventory assistant could allow the IAEA to complete item counting and tag checking using the global identifier or the operator’s barcode in 8–10 person-days of inspection. We would expect other users (e.g., facility operators or verification monitors) to also benefit from significant time savings.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

A Conceptual Design for a Desktop Application to Support Inventory Reconciliation Activities

This paper describes desktop software being developed at Oak Ridge National Laboratory (ORNL) to empower users to reconcile observations about items at nuclear facilities more accurately, reliably, and quickly. Inventory activities at nuclear facilities are often conducted using pen and paper, which can be time-consuming, tedious, and susceptible to reading or transcription errors. The proposed inventory assistant would be a replacement for the paper-based process that could be used by International Atomic Energy Agency (IAEA) inspectors, nuclear facility operators, or treaty verification monitors to conduct an inventory of nuclear and non-nuclear items in a more timely and accurate manner. Garner, McGirl, and Whitaker previously reported their work about a conceptual design for a mobile app to assist users in the field. In general, the inventory assistant would ingest an inventory list, distribute assigned items from the inventory list to one or more mobile devices, enable inventory teams to record their observations in the field, and then enable an inventory lead to integrate and reconcile the observations to produce a final report. The assistant consists of two software components - one for the inventory teams to record observations in the field (In-Field Observations App [IFOA]) and one for the inventory lead to (1) distribute the inventory list to each team, (2) integrate the observations from each team, and (3) reconcile the inventory list with observations (Distribution, Integration, and Reconciliation Application [DIRA]). This paper reviews the inventory assistant workflow and describes the DIRA user experience in more detail. As an example, the authors have chosen to follow the use case of IAEA inspectors conducting item counting and tag checking activities of UF 6 cylinders at a gas centrifuge enrichment plant with a large number of UF 6 cylinders (e.g., thousands). These activities can currently require 30 - 40 person-days of inspection to complete. The authors believe an inventory assistant could allow the IAEA to complete item counting and tag checking using the global identifier or the operator’s barcode in 8 - 10 person-days of inspection. We would expect other users (e.g., facility operators or treaty verification monitors) to also benefit from significant time savings.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Distributed Berkeley Efficient Long-Read to Long-Read Aligner and Overlapper (DiBELLA) v1.0.0

We present a parallel algorithm and scalable implementation for genome analysis, specifically the problem of finding overlaps and alignments for data from "third generation" long read sequencers. While long sequences of DNA offer enormous advantages for biological analysis and insight, current long read sequencing instruments have high error rates and therefore require different approaches to analysis than their short read counterparts. Our work focuses on an efficient distributed-memory parallelization of an accurate single-node algorithm for overlapping and aligning long reads. We achieve scalability of this irregular algorithm by addressing the competing issues of increasing parallelism, minimizing communication, constraining the memory footprint, and ensuring good load balance. The resulting application, DiBELLA, is the first distributed memory overlapper and aligner specifically designed for long reads and parallel scalability.

Ellis, Marquita↗

Recalculation of Soil Bulk Density Used in the Scorpius MCNP6 Model

The soil composition used in the Scorpius MCNP6.2 (Ref. 1) model was recently recalculated. Originally, it was “determined using the United States Geological Survey (USGS) Mercury Core Library and the Nevada National Security Site U.S. Geological Survey Databases (NNSS USGS) and Nevada National Security Site Petrographic, Geochemical, and Geophysical Database (NNSS PGG),” but the details had been lost. Reference 2 documents the updated soil composition determination. Reference 2 estimated the grain density of the soil instead of the bulk density. This report corrects that error. I am indebted to Garrett Euler (EES-17) for reading Ref. 2, pointing out its errors, and guiding me through the correct calculations. This report is organized as follows. For completeness, the composition calculations from Ref. 2 are repeated: Sec. II discusses the data that are reported in the PGG Access database, and Sec. III uses the data in the PGG Access database to compute the composition of the soil. Section IV estimates the bulk density of the soil. Section V evaluates the soil wall thickness in the Scorpius model. Section VI estimates the effect of the soil density change on previously calculated results. Section VII presents a summary and conclusions. Input files and output files are listed in appendices.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

DNABERT-S: pioneering species differentiation with species-aware DNA embeddings

SUMMARY: We introduce DNABERT-S, a tailored genome model that develops species-aware embeddings to naturally cluster and segregate DNA sequences of different species in the embedding space. Differentiating species from genomic sequences (i.e. DNA and RNA) is vital yet challenging, since many real-world species remain uncharacterized, lacking known genomes for reference. Embedding-based methods are therefore used to differentiate species in an unsupervised manner. DNABERT-S builds upon a pre-trained genome foundation model named DNABERT-2. To encourage effective embeddings to error-prone long-read DNA sequences, we introduce Manifold Instance Mixup (MI-Mix), a contrastive objective that mixes the hidden representations of DNA sequences at randomly selected layers and trains the model to recognize and differentiate these mixed proportions at the output layer. We further enhance it with the proposed Curriculum Contrastive Learning (C2LR) strategy. Empirical results on 28 diverse datasets show DNABERT-S's effectiveness, especially in realistic label-scarce scenarios. For example, it identifies twice more species from a mixture of unlabeled genomic sequences, doubles the Adjusted Rand Index (ARI) in species clustering, and outperforms the top baseline's performance in 10-shot species classification with just a 2-shot training. AVAILABILITY AND IMPLEMENTATION: Model, codes, and data are publically available at https://github.com/MAGICS-LAB/DNABERT_S.

Zhou, Zhihan↗

High-Throughput Custom Monitoring for the Mu2e TDAQ System

In this project we are studying the application of programmable network hardware to provide a custom monitoring capability for the Mu2e Trigger and Data Acquisition System (TDAQ) system. The goal of the Mu2e experiment is to search for a charged-lepton flavor violating processes where a negative muon converts into an electron in the field of an aluminum nucleus. This experiment is intended to improve by four orders of magnitude the search sensitivity reached so far. We have a working prototype of a system that provides high-throughput, custom monitoring for the Mu2e TDAQ system. The custom Mu2e network packet header format is parsed as it crosses the network switch. Parsing extracts bits that convey information about error states at read-out controllers (ROCs). This information is periodically relayed to the switch controller, which in turn alerts experiment operators.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗