Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “DNA sequencing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Build Optimization Software Tools (BOOST) v2.0.0

Build Optimization Software Tools (BOOST) accelerate the design of DNA, RNA and protein sequences for their synthesis and assembly in an automated and scalable fashion. The BOOST Juggler automates the following two design tasks: - Reverse-Translation of protein sequences into DNA sequences - Codon Juggling of DNA sequences The BOOST Polisher provides the following design tasks: - The Verification of DNA sequences against DNA synthesis constraints - The Modification of DNA sequences in case they violate any DNA synthesis constraint In addition, the BOOST Polisher can be instructed to verify and modify DNA sequences against sequence patterns, such as restriction sites. The BOOST Partitioner supports both Gibson/chewback and Yeast assembly (also known as Transformation-Associated Recombination, or TAR) methods when performing the decomposition of large sequences that exceed the maximum length of DNA synthesis into synthesizable building blocks. The Partitioner can be instructed to find building blocks with appropriate overlap sequences, making their assembly into the complete construct more efficient. BOOST Workflows enable, in a customized fashion, the execution of the Juggler, Polisher, and Partitioner tools on a batch of sequences.

Sharkey, Michael↗

Prevalence and diversity of TAL effector-like proteins in fungal endosymbiotic Mycetohabitans spp.

EndofungalMycetohabitans(formerlyBurkholderia) spp. rely on a type III secretion system to deliver mostly unidentified effector proteins when colonizing their host fungus,Rhizopus microsporus. The one known secreted effector family fromMycetohabitansconsists of homologues of transcription activator-like (TAL) effectors, which are used by plant pathogenicXanthomonasandRalstoniaspp. to activate host genes that promote disease. These ‘BurkholderiaTAL-like (Btl)’ proteins bind corresponding specific DNA sequences in a predictable manner, but their genomic target(s) and impact on transcription in the fungus are unknown. Recent phenotyping of Btl mutants of twoMycetohabitansstrains revealed that the single Btl in oneMycetohabitans endofungorumstrain enhances fungal membrane stress tolerance, while others in aMycetohabitans rhizoxinicastrain promote bacterial colonization of the fungus. The phenotypic diversity underscores the need to assess the sequence diversity and, given that sequence diversity translates to DNA targeting specificity, the functional diversity of Btl proteins. Using a dual approach to maximize capture of Btl protein sequences for our analysis, we sequenced and assembled nineMycetohabitansspp. genomes using long-read PacBio technology and also mined available short-read Illumina fungal–bacterial metagenomes. We show thatbtlgenes are present across diverseMycetohabitansstrains from Mucoromycota fungal hosts yet vary in sequences and predicted DNA binding specificity. Phylogenetic analysis revealed distinct clades of Btl proteins and suggested thatMycetohabitansmight contain more species than previously recognized. Within our data set, Btl proteins were more conserved acrossM. rhizoxinicastrains than acrossM. endofungorum, but there was also evidence of greater overall strain diversity within the latter clade. Overall, the results suggest that Btl proteins contribute to bacterial–fungal symbioses in myriad ways.

Genetics & Heredity↗

EPBD-BERT

This framework aids in Transcription factor binding site prediction given DNA sequence information and DNA biophysical characteristics computed by pyDNA-EPBD. This framework utilizes a large language model for binding site prediction.

Kabir, Anowarul↗

Fine-scale evaluation of two standard 16S rRNA gene amplicon primer pairs for analysis of total prokaryotes and archaeal nitrifiers in differently managed soils

The advance of high-throughput molecular biology tools allows in-depth profiling of microbial communities in soils, which possess a high diversity of prokaryotic microorganisms. Amplicon-based sequencing of 16S rRNA genes is the most common approach to studying the richness and composition of soil prokaryotes. To reliably detect different taxonomic lineages of microorganisms in a single soil sample, an adequate pipeline including DNA isolation, primer selection, PCR amplification, library preparation, DNA sequencing, and bioinformatic post-processing is required. Besides DNA sequencing quality and depth, the selection of PCR primers and PCR amplification reactions arguably have the largest influence on the results. This study tested the performance and potential bias of two primer pairs, i.e., 515F (Parada)-806R (Apprill) and 515F (Parada)-926R (Quince) in the standard pipelines of 16S rRNA gene Illumina amplicon sequencing protocol developed by the Earth Microbiome Project (EMP), against shotgun metagenome-based 16S rRNA gene reads. The evaluation was conducted using five differently managed soils. We observed a higher richness of soil total prokaryotes by using reverse primer 806R compared to 926R, contradicting to in silico evaluation results. Both primer pairs revealed various degrees of taxon-specific bias compared to metagenome-derived 16S rRNA gene reads. Nonetheless, we found consistent patterns of microbial community variation associated with different land uses, irrespective of primers used. Total microbial communities, as well as ammonia oxidizing archaea (AOA), the predominant ammonia oxidizers in these soils, shifted along with increased soil pH due to agricultural management. In the unmanaged low pH plot abundance of AOA was dominated by the acid-tolerant NS-Gamma clade, whereas limed agricultural plots were dominated by neutral-alkaliphilic NS-Delta/NS-Alpha clades. This study stresses how primer selection influences community composition and highlights the importance of primer selection for comparative and integrative studies, and that conclusions must be drawn with caution if data from different sequencing pipelines are to be compared.

16S rRNA gene amplicon Illumina sequencing↗

Epigenetic Control of Drought Response in Sorghum (EPICON)

Genetic manipulation of crops to increase the presence of desirable traits has been critical to increasing agricultural productivity. These changes have primarily involved modification of the plant’s DNA sequence. However, there is increasing evidence that environmental responses are also mediated by epigenetics, which involves heritable changes without changes in DNA sequence. Epigenetic changes have been shown to play a major role in regulating plant responses to drought, an increasing problem worldwide due to climate change. In general, exposure of plants to water limitation triggers epigenetic changes, which include remodeling of chromatin, the network of DNA, RNA and various proteins making up chromosomes, and related changes in regulatory mechanisms. EPICON’s efforts focus on unraveling the role epigenetic signals play in acclimation to and recovery from drought through effects on individual transcription factors or transcriptional networks that direct entire metabolic pathways. To achieve this goal we will follow responses to water deprivation in sorghum, a widely cultivated cereal with recognized drought tolerance. In EPICON’s field trials, sorghum will be grown under controlled irrigation conditions. Leaf and root samples will be taken to perform molecular phenotyping to track changes in epigenetic, transcriptomic, metabolomic and proteomic footprints. Analysis of this data will provide a better understanding of the epigenetic processes related to drought tolerance, leading to our ultimate goal of identifying transcriptional regulators and pathways controlling drought resistance. The identified genetic targets and their regulatory pathways will be used in future efforts to improve growth of sorghum and other crops in the field and in marginal lands under water-limiting conditions.

59 BASIC BIOLOGICAL SCIENCES↗

Genomad v1.0

Genomad aims to identify mobile genetic elements (namely, viruses and plasmids) from DNA sequence data. It uses a combination of marker gene identification and machine learning models to find likely virus/plasmid candidates in environmental DNA sequencing data. It has an improved classification performance over similar tools.

Camargo, Antonio↗

Remote sensing of cytotype and its consequences for canopy damage in quaking aspen

Abstract Mapping geographic mosaics of genetic variation and their consequences via genotype x environment interactions at large extents and high resolution has been limited by the scalability of DNA sequencing. Here, we address this challenge for cytotype (chromosome copy number) variation in quaking aspen, a drought‐impacted foundation tree species. We integrate airborne imaging spectroscopy data with ground‐based DNA sequencing data and canopy damage data in 391 km 2 of southwestern Colorado. We show that (1) aspen cover and cytotype can be remotely sensed at 1 m spatial resolution, (2) the geographic mosaic of cytotypes is heterogeneous and interdigitated, (3) triploids have higher leaf nitrogen, canopy water content, and carbon isotope shifts (δ 13 C) than diploids, and (4) canopy damage varies among cytotypes and depends on interactions with topography, canopy height, and trait variables. Triploids are at higher risk in hotter and drier conditions.

Blonder, Benjamin↗

Performance of methods for SARS-CoV-2 variant detection and abundance estimation within mixed population samples

The accurate identification of SARS-CoV-2 (SC2) variants and estimation of their abundance in mixed population samples (e.g., air or wastewater) is imperative for successful surveillance of community level trends. Assessing the performance of SC2 variant composition estimators (VCEs) should improve our confidence in public health decision making. Here, we introduce a linear regression based VCE and compare its performance to four other VCEs: two re-purposed DNA sequence read classifiers (Kallisto and Kraken2), a maximum-likelihood based method (Lineage deComposition for Sars-Cov-2 pooled samples (LCS)), and a regression based method (Freyja). We simulated DNA sequence datasets of known variant composition from both Illumina and Oxford Nanopore Technologies (ONT) platforms and assessed the performance of each VCE. We also evaluated VCEs performance using publicly available empirical wastewater samples collected for SC2 surveillance efforts. Bioinformatic analyses were performed with a custom NextFlow workflow (C-WAP, CFSAN Wastewater Analysis Pipeline). Relative root mean squared error (RRMSE) was used as a measure of performance with respect to the known abundance and concordance correlation coefficient (CCC) was used to measure agreement between pairs of estimators. Based on our results from simulated data, Kallisto was the most accurate estimator as it had the lowest RRMSE, followed by Freyja. Kallisto and Freyja had the most similar predictions, reflected by the highest CCC metrics. We also found that accuracy was platform and amplicon panel dependent. For example, the accuracy of Freyja was significantly higher with Illumina data compared to ONT data; performance of Kallisto was best with ARTICv4. However, when analyzing empirical data there was poor agreement among methods and variations in the number of variants detected (e.g., Freyja ARTICv4 had a mean of 2.2 variants while Kallisto ARTICv4 had a mean of 10.1 variants). This work provides an understanding of the differences in performance of a number of VCEs and how accurate they are in capturing the relative abundance of SC2 variants within a mixed sample (e.g., wastewater). Such information should help officials gauge the confidence they can have in such data for informing public health decisions.

60 APPLIED LIFE SCIENCES↗

Nanopore Activity Assays for Detection of Biomarker Protease Activity: Design and Testing of Substrates for Both Nanopore Sequencing and PCR-Based Detection Methods

The work performed in this project has demonstrated the ability to construct proteolytic enzyme substrates that are PCR and sequencing-readable reporter molecules. Specifically, the goal was to detect those reporter molecules via PCR and Oxford Nanopore Technologies MinION sequencing methods following exposure to the biomarker protease thrombin. The assay development focused on binding the constructed peptide-oligonucleotide chimera to immobilized streptavidin. The action of thrombin on the peptide portion of the molecule released the oligonucleotide for detection. Detection of protease activity was demonstrated in a concentration-dependent manner using MALDI-MS, RT-PCR and DNA sequencing. Additional steps to remove background release of reporter molecules during the assay was used to improve the difference in detected oligonucleotide reporter following protease activity. Additional steps in assay development will be to (1) test the assay in an appropriate matrix, (2) investigate detection using additional DNA sequencing platforms and (3) demonstrate multiplexed detection of multiple protease markers in a single reaction.

59 BASIC BIOLOGICAL SCIENCES↗

Highly efficient and simple SSPER and rrPCR approaches for the accurate site-directed mutagenesis of large and small plasmids

Advances are needed in the site-directed mutagenesis of large plasmids for protein structure-function studies, as current methods are often inefficient, complicated and time-consuming. Here two new methods are reported that overcome these difficulties, namely the single primer extension reaction (SSPER) strategy that reaches 100% efficiency and the reduce recycle PCR (rrPCR) method that is advantageous in generating single and pairwise combinations of mutations. Both methods are distinguished from current technologies by the addition of a step that easily removes the oligonucleotide primer(s) after the first reaction, thus allowing for the addition of a second reaction in chronological sequence to generate and isolate the appropriate DNA product with the site-directed mutation(s). High efficiency of the methods is demonstrated by generating single and paired combinations of the 11 site-directed mutations targeted on 5 different plasmid DNA templates ranging from 10 to 12 kb and 57–60% GC-content at a rate of 50–100%. Overall, the methods are demonstrated to be (i) highly accurate, allowing for screening of plasmids by DNA sequencing, (ii) streamlined to generate the mutations within a single day, (iii) cost-effective in requiring only two primers and two enzymes (DpnI and a proofreading DNA polymerase), (iv) straightforward in primer design, (v) applicable for both large and small plasmids, and (vi) easily implemented by entry level researchers.

59 BASIC BIOLOGICAL SCIENCES↗

Sequence-specific dynamic DNA bending explains mitochondrial TFAM’s dual role in DNA packaging and transcription initiation

Abstract Mitochondrial transcription factor A (TFAM) employs DNA bending to package mitochondrial DNA (mtDNA) into nucleoids and recruit mitochondrial RNA polymerase (POLRMT) at specific promoter sites, light strand promoter (LSP) and heavy strand promoter (HSP). Herein, we characterize the conformational dynamics of TFAM on promoter and non-promoter sequences using single-molecule fluorescence resonance energy transfer (smFRET) and single-molecule protein-induced fluorescence enhancement (smPIFE) methods. The DNA-TFAM complexes dynamically transition between partially and fully bent DNA conformational states. The bending/unbending transition rates and bending stability are DNA sequence-dependent—LSP forms the most stable fully bent complex and the non-specific sequence the least, which correlates with the lifetimes and affinities of TFAM with these DNA sequences. By quantifying the dynamic nature of the DNA-TFAM complexes, our study provides insights into how TFAM acts as a multifunctional protein through the DNA bending states to achieve sequence specificity and fidelity in mitochondrial transcription while performing mtDNA packaging.

59 BASIC BIOLOGICAL SCIENCES↗

Genome dependent Cas9/gRNA search time underlies sequence dependent gRNA activity

Abstract CRISPR-Cas9 is a powerful DNA editing tool. A gRNA directs Cas9 to cleave any DNA sequence with a PAM. However, some gRNA sequences mediate cleavage at higher efficiencies than others. To understand this, numerous studies have screened large gRNA libraries and developed algorithms to predict gRNA sequence dependent activity. These algorithms do not predict other datasets as well as their training dataset and do not predict well between species. Here, to better understand these discrepancies, we retrospectively examine sequence features that impact gRNA activity in 44 published data sets. We find strong evidence that gRNA sequence dependent activity is largely influenced by the ability of the Cas9/gRNA complex to find the target site rather than activity at the target site and that this drives sequence dependent differences in gRNA activity between different species. This understanding will help guide future work to understand Cas9 activity as well as efforts to identify optimal gRNAs and improve Cas9 variants.

59 BASIC BIOLOGICAL SCIENCES↗

DL-TODA: A Deep Learning Tool for Omics Data Analysis

Metagenomics is a technique for genome-wide profiling of microbiomes; this technique generates billions of DNA sequences called reads. Given the multiplication of metagenomic projects, computational tools are necessary to enable the efficient and accurate classification of metagenomic reads without needing to construct a reference database. The program DL-TODA presented here aims to classify metagenomic reads using a deep learning model trained on over 3000 bacterial species. A convolutional neural network architecture originally designed for computer vision was applied for the modeling of species-specific features. Using synthetic testing data simulated with 2454 genomes from 639 species, DL-TODA was shown to classify nearly 75% of the reads with high confidence. The classification accuracy of DL-TODA was over 0.98 at taxonomic ranks above the genus level, making it comparable with Kraken2 and Centrifuge, two state-of-the-art taxonomic classification tools. DL-TODA also achieved an accuracy of 0.97 at the species level, which is higher than 0.93 by Kraken2 and 0.85 by Centrifuge on the same test set. Application of DL-TODA to the human oral and cropland soil metagenomes further demonstrated its use in analyzing microbiomes from diverse environments. Compared to Centrifuge and Kraken2, DL-TODA predicted distinct relative abundance rankings and is less biased toward a single taxon.

59 BASIC BIOLOGICAL SCIENCES↗

ULTRA-effective labeling of tandem repeats in genomic sequence

In the age of long read sequencing, genomics researchers now have access to accurate repetitive DNA sequence (including satellites) that, due to the limitations of short read-sequencing, could previously be observed only as unmappable fragments. Tools that annotate repetitive sequence are now more important than ever, so that we can better understand newly uncovered repetitive sequences, and also so that we can mitigate errors in bioinformatic software caused by those repetitive sequences. To that end, we introduce the 1.0 release of our tool for identifying and annotating locally repetitive sequence, ULTRA Locates Tandemly Repetitive Areas (ULTRA). ULTRA is fast enough to use as part of an efficient annotation pipeline, produces state-of-the-art reliable coverage of repetitive regions containing many mutations, and provides interpretable statistics and labels for repetitive regions.

59 BASIC BIOLOGICAL SCIENCES↗

Characterizing Integrase-Attachment Site Pairs: A Machine Learning Approach

Genomic islands (GIs) are mobile genetic elements that integrate into host genomes via self-encoded integrases at specific DNA sequences known as attachment (att) sites. The ability to predict the target att site from an integrase's protein sequence is a central challenge in genomics due to the sequence diversity of integrases and the subtlety of their DNA recognition motifs.

59 BASIC BIOLOGICAL SCIENCES↗

Mitochondrial Genomes of the United States Distribution of Gray Fox (Urocyon cinereoargenteus) Reveal a Major Phylogeographic Break at the Great Plains Suture Zone

We examined phylogeographic structure in gray fox (Urocyon cinereoargenteus) across the United States to identify the location of secondary contact zone(s) between eastern and western lineages and investigate the possibility of additional cryptic intraspecific divergences. We generated and analyzed complete mitochondrial genome sequence data from 75 samples and partial control region mitochondrial DNA sequences from 378 samples to investigate levels of genetic diversity and structure through population- and individual-based analyses including estimates of divergence (FST and SAMOVA), median joining networks, and phylogenies. We used complete mitochondrial genomes to infer phylogenetic relationships and date divergence times of major lineages of Urocyon in the United States. Despite broad-scale sampling, we did not recover additional major lineages of Urocyon within the United States, but identified a deep east-west split (~0.8 million years) with secondary contact at the Great Plains Suture Zone and confirmed the Channel Island fox (Urocyon littoralis) is nested within U. cinereoargenteus. Genetic diversity declined at northern latitudes in the eastern United States, a pattern concordant with post-glacial recolonization and range expansion. Beyond the east-west divergence, morphologically-based subspecies did not form monophyletic groups, though unique haplotypes were often geographically limited. Gray foxes in the United States displayed a deep, cryptic divergence suggesting taxonomic revision is needed. Secondary contact at a common phylogeographic break, the Great Plains Suture Zone, where environmental variables show a sharp cline, suggests ongoing evolutionary processes may reinforce this divergence. Follow-up study with nuclear markers should investigate whether hybridization is occurring along the suture zone and characterize contemporary population structure to help identify conservation units. Comparative work on other wide-ranging carnivores in the region should test whether similar evolutionary patterns and processes are occurring.

54 ENVIRONMENTAL SCIENCES↗