Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Sequencing data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Sequence data - Magnitude and implications of some ambiguities.

A stochastic model is applied to the divergence of the horse-pig lineage from a common ansestor in terms of the alpha and beta chains of hemoglobin and fibrinopeptides. The results are compared with those based on the minimum mutation distance model of Fitch (1972). Buckwheat and cauliflower cytochrome c sequences are analyzed to demonstrate their ambiguities. A comparative analysis of evolutionary rates for various proteins of horses and pigs shows that errors of considerable magnitude are introduced by Glx and Asx ambiguities into evolutionary conclusions drawn from sequences of incompletely analyzed proteins.

Holmquist, R.

Methods and apparatus for extraction and tracking of objects from multi-dimensional sequence data

An object tracking technique is provided which, given: (i) a potentially large data set; (ii) a set of dimensions along which the data has been ordered; and (iii) a set of functions for measuring the similarity between data elements, a set of objects are produced. Each of these objects is defined by a list of data elements. Each of the data elements on this list contains the probability that the data element is part of the object. The method produces these lists via an adaptive, knowledge-based search function which directs the search for high-probability data elements. This serves to reduce the number of data element combinations evaluated while preserving the most flexibility in defining the associations of data elements which comprise an object.

Hill, Matthew L.

Methods and apparatus for extraction and tracking of objects from multi-dimensional sequence data

An object tracking technique is provided which, given: (i) a potentially large data set; (ii) a set of dimensions along which the data has been ordered; and (iii) a set of functions for measuring the similarity between data elements, a set of objects are produced. Each of these objects is defined by a list of data elements. Each of the data elements on this list contains the probability that the data element is part of the object. The method produces these lists via an adaptive, knowledge-based search function which directs the search for high-probability data elements. This serves to reduce the number of data element combinations evaluated while preserving the most flexibility in defining the associations of data elements which comprise an object.

Hill, Matthew L.

LevSeq: Rapid Generation of Sequence-Function Data for Directed Evolution and Machine Learning

Sequence-function data provides valuable information about the protein functional landscape but is rarely obtained during directed evolution campaigns. Here, we present Long-read every variant Sequencing (LevSeq), a pipeline that combines a dual barcoding strategy with nanopore sequencing to rapidly generate sequence-function data for entire protein-coding genes. LevSeq integrates into existing protein engineering workflows and comes with open-source software for data analysis and visualization. The pipeline facilitates data-driven protein engineering by consolidating sequence-function data to inform directed evolution and provide the requisite data for machine learning-guided protein engineering (MLPE). LevSeq enables quality control of mutagenesis libraries prior to screening, which reduces time and resource costs. Simulation studies demonstrate LevSeq’s ability to accurately detect variants under various experimental conditions. Lastly, we show LevSeq’s utility in engineering protoglobins for new-to-nature chemistry. Widespread adoption of LevSeq and sharing of the data will enhance our understanding of protein sequence-function landscapes and empower data-driven directed evolution.

59 BASIC BIOLOGICAL SCIENCES

Bayesian estimation of HIV acquisition dates for prevention trials

Accurate timing estimates of when participants acquire HIV in HIV prevention trials are necessary for determining antibody levels at acquisition. The Antibody-Mediated Prevention (AMP) Studies showed that a passively administered broadly neutralizing antibody can prevent the acquisition of HIV from a neutralization-sensitive virus. We developed a pipeline for estimating the date of detectable HIV acquisition (DDA) in AMP Study participants using diagnostic and viral sequence data. Using a Bayesian strategy that combines three streams of data (REN [rev/vpu/env/Δnef] sequence, GP [gag/Δpol] sequence, and diagnostic) where their 95% credible intervals overlap based on pre-specified criteria and decision rules. We evaluated the performance of our AMP pipeline using PacBio viral sequence data from 41 participants across two prospective acute HIV acquisition cohort studies, FRESH and RV217, with twice-weekly sampling. These cohort studies enrolled young women in South Africa and men and women in Kenya and Thailand, respectively, with a high likelihood of HIV acquisition. In evaluating performance, “true DDA” was the center of bounds between last-negative and first-positive RNA diagnostic tests (median time 4 days, range 2–7 days); bias was the mean difference between estimated and true DDA. Using diagnostic data alone yielded timing estimates with a bias of 2.4 days and root mean square error (RMSE) of 7.9 days. These results were improved using sequence + diagnostic data (bias 1.5 days, RMSE 6.9 days), as well as by restricting sequence-based estimation to samples from ≤5 weeks post-DDA (bias 0.2 days, RMSE 7.8 days).

59 BASIC BIOLOGICAL SCIENCES

Omega flight-test data reduction sequence

Computer programs for Omega data conversion, summary, and preparation for distribution are presented. Program logic and sample data formats are included, along with operational instructions for each program. Flight data (or data collected in flight format in the laboratory) is provided by the Ohio University Omega receiver base in the form of 6-bit binary words representing the phase of an Omega station with respect to the receiver's local clock. All eight Omega stations are measured in each 10-second Omega time frame. In addition, an event-marker bit and a time-slot D synchronizing bit are recorded. Program FDCON is used to remove data from the flight recorder tape and place it on data-processing cards for later use. Program FDSUM provides for computer plotting of selected LOP's, for single-station phase plots, and for printout of basic signal statistics for each Omega channel. Mean phase and standard deviation are printed, along with data from which a phase distribution can be plotted for each Omega station. Program DACOP simply copies the Omega data deck a controlled number of times, for distribution to users.

Lilley, R. W.

Archaeal phylogeny: reexamination of the phylogenetic position of Archaeoglobus fulgidus in light of certain composition-induced artifacts

A major and too little recognized source of artifact in phylogenetic analysis of molecular sequence data is compositional difference among sequences. The problem becomes particularly acute when alignments contain ribosomal RNAs from both mesophilic and thermophilic species. Among prokaryotes the latter are considerably higher in G + C content than the former, which often results in artificial clustering of thermophilic lineages and their being placed artificially deep in phylogenetic trees. In this communication we review archaeal phylogeny in the light of this consideration, focusing in particular on the phylogenetic position of the sulfate reducing species Archaeoglobus fulgidus, using both 16S rRNA and 23S rRNA sequences. The analysis shows clearly that the previously reported deep branching of the A. fulgidus lineage (very near the base of the euryarchaeal side of the archaeal tree) is incorrect, and that the lineage actually groups with a previously recognized unit that comprises the Methanomicrobiales and extreme halophiles.

NASA Discipline Exobiology

The rate and efficiency of high-mass star formation along the Hubble sequence

Data obtained with IRAS are used to compare and contrast the global star formation rates for a galactic sample which represents essentially all known noninteracting spiral and lenticular galaxies within 40 Mpc. The distribution of 60 micron luminosity is similar for spirals of types Sa-Scd inclusively, although the luminosities of the very early and very late types are, on average, one order of magnitude lower. High-mass star formation rates are similar for early, intermediate, and late type spirals, and the average high-mass star formation rate per unit molecular gas mass is independent of type for spiral galaxies. A remarkable homogeneity exists in the high-mass star-forming capabilities of spiral galaxies, particularly among the Sa-Scd types. The Hubble sequence is therefore not a sequence in the present-day rate or production efficiency of high-mass stars.

Devereux, Nicholas A.

Dataset for the Danczak et al., 2025 manuscript about bacterial-fungal interactions

We generated genome-resolved multiomics data from a series of metagenomic and metatranscriptomic sequencing. Specifically, we acquired, functionally annotated, and taxonomically classified both bacterial and eukaryotic metagenome assembled genomes (MAGs). For bacterial MAGs, we assembled eukaryotic float metagenomic sequencing data from JGI using MEGAHIT, binned and refined MAGs using MetaWRAP and dRep, functionally annotated MAGs using eggNOG mapper, and assigned taxonomy using GTDB-tk. For eukaryotic MAGs, we first identified potentially eukaryotic contigs from a coassembly of eukaryotic float metagenomic sequencing data from JGI using EukRep and Whokaryote, binned MAGs using MetaBAT2, functionally annotated MAGs using eggNOG mapper, and assigned taxonomy using Eukulele. Bulk metatranscriptomic reads were mapped to bacterial MAGs and polyA-metatranscriptomic read were mapped to eukaryotic MAGs using bbmap.

Danczak, Robert E. [Pacific Northwest National Lab

Cascade Error Projection with Low Bit Weight Quantization for High Order Correlation Data

In this paper, we reinvestigate the solution for chaotic time series prediction problem using neural network approach. The nature of this problem is such that the data sequences are never repeated, but they are rather in chaotic region. However, these data sequences are correlated between past, present, and future data in high order. We use Cascade Error Projection (CEP) learning algorithm to capture the high order correlation between past and present data to predict a future data using limited weight quantization constraints. This will help to predict a future information that will provide us better estimation in time for intelligent control system. In our earlier work, it has been shown that CEP can sufficiently learn 5-8 bit parity problem with 4- or more bits, and color segmentation problem with 7- or more bits of weight quantization. In this paper, we demonstrate that chaotic time series can be learned and generalized well with as low as 4-bit weight quantization using round-off and truncation techniques. The results show that generalization feature will suffer less as more bit weight quantization is available and error surfaces with the round-off technique are more symmetric around zero than error surfaces with the truncation technique. This study suggests that CEP is an implementable learning technique for hardware consideration.

Duong, Tuan A.

Delay Tolerant, Radio Frequency Identification (RFID )-enabled Sensing

Radio Frequency Identification (RFID) technology offers a completely passive method to transmit fixed data sequences from an RFID tag, which typically doesn’t have its own power supply, to an interrogator. Radio frequency (RF) energy harvested from the interrogator is rectified by the tag and used to charge an integrated circuit (IC). The IC then modulates the received signal with the data stored on the tag and reflects the energy back to the interrogator. RFID has seen great proliferation in terrestrial inventory management applications, and it has recently made the jump to spaceflight applications onboard the International Space Station, augmenting an existing optical bar‐code infrastructure for tracking supplies. A number of advanced automated logistics management (ALM) concepts employing RFID are currently being developed and evaluated, including so‐called “smart” shelves, cubbies, and trash receptacles using low‐power, embeddable RFID interrogators. An infrastructure where both crew members and robotic assistants, such as autonomous free flyers, are similarly equipped with small RFID interrogators seems likely. It therefore behooves us to consider extending this infrastructure beyond ALM to applications such as low power, embedded sensing. Typically, data on an RFID tag can only be written by an RFID interrogator, using interrogator energy. In recent years, however, a few efforts have focused on using that energy to drive data acquisition from the tag IC, allowing the tag to modify its stored data sequence with sensor data before replying to an interrogator. In this way, the tag can act as a completely passive sensing device. One problem exists with this approach, however: the tag cannot gather data when an interrogator is not present. Thus, strictly passive RFID sensing tags cannot gather data at regular intervals, in the manner of a typical wireless sensor network, without careful, and impractical, planning of mobile interrogator movements. To address this shortcoming, we look to a recent advance in RFID technology which allows an external microprocessor to power the tag IC and write directly into its RFID memory using a wired serial interface. In this paradigm, data gathering is driven by a small, on‐board power supply (using batteries or harvested energy), and data transfer is provided passively through the RFID interrogation service. Since communication typically consumes the lion’s share of power in WSNs, such a technique has the potential to enable extremely long‐lived, embedded wireless sensing when used with extremely low‐current microcontrollers. Since the communication channel is only open when an interrogator is present and actively interrogating the RFID sensing tag, transport of periodically‐sampled sensor data presents itself as a delay/disruption‐tolerant networking (DTN) problem. In this paper, we present the design of a DTN‐like overlay on the common EPC Global, Class 1, Generation 2 RFID standard. This overlay allows seamless, guaranteed data transfer using the EPC Global protocol, supporting extremely low‐power, embedded sensing using an infrastructure likely to be already in place for ALM applications. We evaluate a prototype implementation of a complete end‐to‐end system using a robotic RFID interrogation agent, and we present future directions for the development of this sensing technique.

Raymond S. Wagner

Implied alignment: a synapomorphy-based multiple-sequence alignment method and its use in cladogram search

A method to align sequence data based on parsimonious synapomorphy schemes generated by direct optimization (DO; earlier termed optimization alignment) is proposed. DO directly diagnoses sequence data on cladograms without an intervening multiple-alignment step, thereby creating topology-specific, dynamic homology statements. Hence, no multiple-alignment is required to generate cladograms. Unlike general and globally optimal multiple-alignment procedures, the method described here, implied alignment (IA), takes these dynamic homologies and traces them back through a single cladogram, linking the unaligned sequence positions in the terminal taxa via DO transformation series. These "lines of correspondence" link ancestor-descendent states and, when displayed as linearly arrayed columns without hypothetical ancestors, are largely indistinguishable from standard multiple alignment. Since this method is based on synapomorphy, the treatment of certain classes of insertion-deletion (indel) events may be different from that of other alignment procedures. As with all alignment methods, results are dependent on parameter assumptions such as indel cost and transversion:transition ratios. Such an IA could be used as a basis for phylogenetic search, but this would be questionable since the homologies derived from the implied alignment depend on its natal cladogram and any variance, between DO and IA + Search, due to heuristic approach. The utility of this procedure in heuristic cladogram searches using DO and the improvement of heuristic cladogram cost calculations are discussed. c2003 The Willi Hennig Society. Published by Elsevier Science (USA). All rights reserved.

Non-NASA Center

Produced Water DNA Database (PW-DNA): Utilizing KBase to generate an environmental specific curated molecular database

The deep subsurface is estimated to host the majority of Earth’s microbial biomass yet remains one of the most challenging environments to access and study. One common approach to investigate these microbial communities is through the analysis of produced water from subsurface reservoirs, where researchers can assess water and gas chemistry along with molecular (DNA/RNA) sequence data. Advances in high-throughput sequencing have greatly expanded our understanding of these environments and their biotechnological potential. However, further progress requires large-scale, integrative meta-analyses across diverse datasets. To address this need, we developed the Produced Water-DNA (PW-DNA) Database, a curated, publicly available resource that consolidates microbial DNA/RNA sequences, geochemical data, and relevant metadata from in situ hydrocarbon environments such as coal beds, oil reservoirs, and natural gas systems. The PW-DNA database delivers three core benefits to the research community: (1) it improves data sharing by linking environmental microbial datasets with corresponding geochemical parameters, enabling more robust filtering and analysis; (2) it connects with complementary research databases to promote broader dissemination and interoperability; and (3) it supports technological innovation by serving as a resource for identifying microbial trends and exploring genetic potential. While individual studies have highlighted basin-specific microbial communities and functional redundancy in biogeochemical cycling, a comprehensive, system-wide perspective is needed to better understand connectivity and novelty across subsurface ecosystems. By designing the PW-DNA in the KBase platform, we provide a reproducible, visual framework for integrating large-scale genomic and geochemical data, enabling researchers to perform more informed analyses and experimental design. Ultimately, this resource enhances the ability to identify, characterize, and interpret microbial functions across diverse subsurface environments, thereby accelerating discovery in subsurface microbiology and biotechnology.

59 BASIC BIOLOGICAL SCIENCES

Modularization of EDGE Workflows Using Nextflow: Improving the Efficiency and Maintainability of Bioinformatics Software

EDGE is a bioinformatics platform developed in 2016 by researchers at Los Alamos National Laboratory (LANL) to facilitate the analysis of next-generation sequencing data by researchers with varying levels of experience in bioinformatics (Li et al., 2017). Users with single-end, paired-end or long-read sequencing data can provide their reads as input to EDGE and select the combination of workflows to run that are most useful for their research (e.g., quality control of reads, genome assembly, or the taxonomic classification of input reads). Table 1 summarizes the modules available in EDGE. EDGE is available as a web platform at https://edgebioinformatics.org, as installable source code maintained on GitHub under a GPLv3 license, and as a publicly hosted Docker image.

59 BASIC BIOLOGICAL SCIENCES

Parallel VLSI Architecture

Fermat number transformation convolutes two digital data sequences. Very-large-scale integration (VLSI) applications, such as image and radar signal processing, X-ray reconstruction, and spectrum shaping, linear convolution of two digital data sequences of arbitrary lenghts accomplished using Fermat number transform (ENT).

Truong, T. K.