Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Illumina”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

A Two-Step PCR Protocol Enabling Flexible Primer Choice and High Sequencing Yield for Illumina MiSeq Meta-Barcoding

High-throughput amplicon sequencing that primarily targets the 16S ribosomal DNA (rDNA) (for bacteria and archaea) and the Internal Transcribed Spacer rDNA (for fungi) have facilitated microbial community discovery across diverse environments. A three-step PCR that utilizes flexible primer choices to construct the library for Illumina amplicon sequencing has been applied to several studies in forest and agricultural systems. The three-step PCR protocol, while producing high-quality reads, often yields a large number (up to 46%) of reads that are unable to be assigned to a specific sample according to its barcode. Here, we improve this technique through an optimized two-step PCR protocol. We tested and compared the improved two-step PCR meta-barcoding protocol against the three-step PCR protocol using four different primer pairs (fungal ITS: ITS1F-ITS2 and ITS1F-ITS4, and bacterial 16S: 515F-806R and 341F-806R). We demonstrate that the sequence quantity and recovery rate were significantly improved with the two-step PCR approach (fourfold more read counts per sample; determined reads ≈90% per run) while retaining high read quality (Q30 > 80%). Given that synthetic barcodes are incorporated independently from any specific primers, this two-step PCR protocol can be broadly adapted to different genomic regions and organisms of scientific interest.

16S rDNA↗

Altering translation allows E. coli to overcome chemically stabilized G-quadruplexes

Genomic DNA from each sample was prepared using the Wizard Genomic DNA Purification Kit (Promega) and after, DNA was quantified using the QuantiFluor ONE dsDNA System (Promega). Genomic DNA underwent shearing to ~200 bp fragments via sonication and the gDNA fragments were prepared for sequencing using the NEBNext Ultra II DNA Library Prep Kit for Illumina (NEB). Bead-based size selection was used to select ~200 bp fragments and the fragments then underwent a splinkerette PCR using a Tn5-enriching forward primer and custom reverse primers for multiplexing. A final bead-based size selection was used to select for the correct length DNA. DNA was sequenced at the University of Michigan Advanced Genomics Core using Illumina sequencing with a custom read primer reading the last 10 nt of the transposon. PhiX174 DNA spike was added to the run to ensure sufficient sequence diversity on the flow cell. Then, a custom index read primer and standard Illumina primer were used to sequence the index reads and PhiX174, respectively.

Keck, James L.↗

EDGE COVID-19: a web platform to generate submission-ready genomes from SARS-CoV-2 sequencing efforts

Abstract Summary Genomics has become an essential technology for surveilling emerging infectious disease outbreaks. A range of technologies and strategies for pathogen genome enrichment and sequencing are being used by laboratories worldwide, together with different and sometimes ad hoc, analytical procedures for generating genome sequences. A fully integrated analytical process for raw sequence to consensus genome determination, suited to outbreaks such as the ongoing COVID-19 pandemic, is critical to provide a solid genomic basis for epidemiological analyses and well-informed decision making. We have developed a web-based platform and integrated bioinformatic workflows that help to provide consistent high-quality analysis of SARS-CoV-2 sequencing data generated with either the Illumina or Oxford Nanopore Technologies (ONT). Using an intuitive web-based interface, this workflow automates data quality control, SARS-CoV-2 reference-based genome variant and consensus calling, lineage determination and provides the ability to submit the consensus sequence and necessary metadata to GenBank, GISAID and INSDC raw data repositories. We tested workflow usability using real world data and validated the accuracy of variant and lineage analysis using several test datasets, and further performed detailed comparisons with results from the COVID-19 Galaxy Project workflow. Our analyses indicate that EC-19 workflows generate high-quality SARS-CoV-2 genomes. Finally, we share a perspective on patterns and impact observed with Illumina versus ONT technologies on workflow congruence and differences. Availability and implementation https://edge-covid19.edgebioinformatics.org, and https://github.com/LANL-Bioinformatics/EDGE/tree/SARS-CoV2. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

Metagenomes from Eastern Brazilian Amazonian floodplains in the wet and dry seasons

Brief sample description Sediment samples from 0 to 10 cm depth were collected in triplicate from two floodplains of the Eastern Brazilian Amazon, one located on the Amazonas River (FP2, 2°28'11.2"S 54°38'49.9"W) and the other at the intersection between the Amazonas and the Tapajós rivers (FP3, 2°22'44.8"S 54°44'21.1"W), in the wet and dry seasons (May and October 2016, respectively). Total DNA was extracted in duplicate from 0.25 g of sediment using PowerLyzer PowerSoil DNA Isolation Kit. The metagenomic libraries were constructed using NEBNext Ultra II DNA Library Prep Kit for Illumina and paired-end sequenced (2 x 150 bp) on an Illumina HiSeq 2500 platform. Detailed information about the study sites, sampling, sediment physicochemical properties, DNA extraction and quantification have been previously described by Gontijo et al. (2021). Sample IDs: M1, M2 and M3: FP2, wet season M4, M5 and M6: FP3, wet season M7, M8 and M9: FP2, dry season M10, M11 and M12: FP3, dry season

59 BASIC BIOLOGICAL SCIENCES↗

Whole genome resequencing data from a collection of Clostridium Thermocellum strains

Clostridium thermocellum is an anaerobic thermophilic bacterium that natively ferments cellulose to ethanol and organic acids. This data set is a collection of whole genome resequencing data for several hundred strains of Clostridium thermocellum. It includes strains that have been engineered to increase ethanol production, strains that have been engineered to understand microbial physiology, and strains that have been adapted for desired phenotypes including increased ethanol tolerance. Resequencing data consists of paired Illumina reads, 100-150 bp on each end, with a ~500 bp insert size. One data file containing raw Illumina data (interleaved) is available for each strain. We also provide data describing the mutations identified in each strain, and distinguish between inherited and newly observed mutations. In addition to resequencing data, we also provide metadata describing the lineage of each strain, and any targeted genetic modifications.

resequencing bio energy fermentation↗

Performance of methods for SARS-CoV-2 variant detection and abundance estimation within mixed population samples

The accurate identification of SARS-CoV-2 (SC2) variants and estimation of their abundance in mixed population samples (e.g., air or wastewater) is imperative for successful surveillance of community level trends. Assessing the performance of SC2 variant composition estimators (VCEs) should improve our confidence in public health decision making. Here, we introduce a linear regression based VCE and compare its performance to four other VCEs: two re-purposed DNA sequence read classifiers (Kallisto and Kraken2), a maximum-likelihood based method (Lineage deComposition for Sars-Cov-2 pooled samples (LCS)), and a regression based method (Freyja). We simulated DNA sequence datasets of known variant composition from both Illumina and Oxford Nanopore Technologies (ONT) platforms and assessed the performance of each VCE. We also evaluated VCEs performance using publicly available empirical wastewater samples collected for SC2 surveillance efforts. Bioinformatic analyses were performed with a custom NextFlow workflow (C-WAP, CFSAN Wastewater Analysis Pipeline). Relative root mean squared error (RRMSE) was used as a measure of performance with respect to the known abundance and concordance correlation coefficient (CCC) was used to measure agreement between pairs of estimators. Based on our results from simulated data, Kallisto was the most accurate estimator as it had the lowest RRMSE, followed by Freyja. Kallisto and Freyja had the most similar predictions, reflected by the highest CCC metrics. We also found that accuracy was platform and amplicon panel dependent. For example, the accuracy of Freyja was significantly higher with Illumina data compared to ONT data; performance of Kallisto was best with ARTICv4. However, when analyzing empirical data there was poor agreement among methods and variations in the number of variants detected (e.g., Freyja ARTICv4 had a mean of 2.2 variants while Kallisto ARTICv4 had a mean of 10.1 variants). This work provides an understanding of the differences in performance of a number of VCEs and how accurate they are in capturing the relative abundance of SC2 variants within a mixed sample (e.g., wastewater). Such information should help officials gauge the confidence they can have in such data for informing public health decisions.

60 APPLIED LIFE SCIENCES↗

Intensity of sample processing methods impacts wastewater SARS-CoV-2 whole genome amplicon sequencing outcomes

Wastewater SARS-CoV-2 surveillance has been deployed since the beginning of the COVID-19 pandemic to monitor the dynamics in virus burden in local communities. Genomic surveillance of SARS-CoV-2 in wastewater, particularly efforts aimed at whole genome sequencing for variant tracking and identification, are still challenging due to low target concentration, complex microbial and chemical background, and lack of robust nucleic acid recovery experimental procedures. The intrinsic sample limitations are inherent to wastewater and are thus unavoidable. Here, we use a statistical approach that couples correlation analyses to a random forest-based machine learning algorithm to evaluate potentially important factors associated with wastewater SARS-CoV-2 whole genome amplicon sequencing outcomes, with a specific focus on the breadth of genome coverage. We collected 182 composite and grab wastewater samples from the Chicago area between November 2020 to October 2021. Samples were processed using a mixture of processing methods reflecting different homogenization intensities (HA + Zymo beads, HA + glass beads, and Nanotrap), and were sequenced using one of the two library preparation kits (the Illumina COVIDseq kit and the QIAseq DIRECT kit). Technical factors evaluated using statistical and machine learning approaches include sample types, certain sample intrinsic features, and processing and sequencing methods. The results suggested that sample processing methods could be a predominant factor affecting sequencing outcomes, and library preparation kits was considered a minor factor. Finally, a synthetic SARS-CoV-2 RNA spike-in experiment was performed to validate the impact from processing methods and suggested that the intensity of the processing methods could lead to different RNA fragmentation

60 APPLIED LIFE SCIENCES↗

Fine-scale evaluation of two standard 16S rRNA gene amplicon primer pairs for analysis of total prokaryotes and archaeal nitrifiers in differently managed soils

The advance of high-throughput molecular biology tools allows in-depth profiling of microbial communities in soils, which possess a high diversity of prokaryotic microorganisms. Amplicon-based sequencing of 16S rRNA genes is the most common approach to studying the richness and composition of soil prokaryotes. To reliably detect different taxonomic lineages of microorganisms in a single soil sample, an adequate pipeline including DNA isolation, primer selection, PCR amplification, library preparation, DNA sequencing, and bioinformatic post-processing is required. Besides DNA sequencing quality and depth, the selection of PCR primers and PCR amplification reactions arguably have the largest influence on the results. This study tested the performance and potential bias of two primer pairs, i.e., 515F (Parada)-806R (Apprill) and 515F (Parada)-926R (Quince) in the standard pipelines of 16S rRNA gene Illumina amplicon sequencing protocol developed by the Earth Microbiome Project (EMP), against shotgun metagenome-based 16S rRNA gene reads. The evaluation was conducted using five differently managed soils. We observed a higher richness of soil total prokaryotes by using reverse primer 806R compared to 926R, contradicting to in silico evaluation results. Both primer pairs revealed various degrees of taxon-specific bias compared to metagenome-derived 16S rRNA gene reads. Nonetheless, we found consistent patterns of microbial community variation associated with different land uses, irrespective of primers used. Total microbial communities, as well as ammonia oxidizing archaea (AOA), the predominant ammonia oxidizers in these soils, shifted along with increased soil pH due to agricultural management. In the unmanaged low pH plot abundance of AOA was dominated by the acid-tolerant NS-Gamma clade, whereas limed agricultural plots were dominated by neutral-alkaliphilic NS-Delta/NS-Alpha clades. This study stresses how primer selection influences community composition and highlights the importance of primer selection for comparative and integrative studies, and that conclusions must be drawn with caution if data from different sequencing pipelines are to be compared.

16S rRNA gene amplicon Illumina sequencing↗

High-Throughput Microbial Community Analyses to Establish a Natural Fungal and Bacterial Consortium from Sewage Sludge Enriched with Three Pharmaceutical Compounds

Emerging and unregulated contaminants end up in soils via stabilized/composted sewage sludges, paired with possible risks associated with the development of microbial resistance to antimicrobial agents or an imbalance in the microbial communities. An enrichment experiment was performed, fortifying the sewage sludge with carbamazepine, ketoprofen and diclofenac as model compounds, with the aim to obtain strains with the capability to transform these pollutants. Culturable microorganisms were obtained at the end of the experiment. Among fungi, Cladosporium cladosporioides, Alternaria alternata and Penicillium raistrickii showed remarkable degradation rates. Population shifts in bacterial and fungal communities were also studied during the selective pressure using Illumina MiSeq. These analyses showed a predominance of Ascomycota (Dothideomycetes and Aspergillaceae) and Actinobacteria and Proteobacteria, suggesting the possibility of selecting native microorganisms to carry out bioremediation processes using tailored techniques.

59 BASIC BIOLOGICAL SCIENCES↗

Benchmarking second and third-generation sequencing platforms for microbial metagenomics

Shotgun metagenomic sequencing is a common approach for studying the taxonomic diversity and metabolic potential of complex microbial communities. Current methods primarily use second generation short read sequencing, yet advances in third generation long read technologies provide opportunities to overcome some of the limitations of short read sequencing. Here, we compared seven platforms, encompassing second generation sequencers (Illumina HiSeq 300, MGI DNBSEQ-G400 and DNBSEQ-T7, ThermoFisher Ion GeneStudio S5 and Ion Proton P1) and third generation sequencers (Oxford Nanopore Technologies MinION R9 and Pacific Biosciences Sequel II). We constructed three uneven synthetic microbial communities composed of up to 87 genomic microbial strains DNAs per mock, spanning 29 bacterial and archaeal phyla, and representing the most complex and diverse synthetic communities used for sequencing technology comparisons. Our results demonstrate that third generation sequencing have advantages over second generation platforms in analyzing complex microbial communities, but require careful sequencing library preparation for optimal quantitative metagenomic analysis. Our sequencing data also provides a valuable resource for testing and benchmarking bioinformatics software for metagenomics.

59 BASIC BIOLOGICAL SCIENCES↗

Coassembly and binning of a twenty-year metagenomic time-series from Lake Mendota

Abstract The North Temperate Lakes Long-Term Ecological Research (NTL-LTER) program has been extensively used to improve understanding of how aquatic ecosystems respond to environmental stressors, climate fluctuations, and human activities. Here, we report on the metagenomes of samples collected between 2000 and 2019 from Lake Mendota, a freshwater eutrophic lake within the NTL-LTER site. We utilized the distributed metagenome assembler MetaHipMer to coassemble over 10 terabases (Tbp) of data from 471 individual Illumina-sequenced metagenomes. A total of 95,523,664 contigs were assembled and binned to generate 1,894 non-redundant metagenome-assembled genomes (MAGs) with ≥50% completeness and ≤10% contamination. Phylogenomic analysis revealed that the MAGs were nearly exclusively bacterial, dominated by Pseudomonadota (Proteobacteria, N = 623) and Bacteroidota (N = 321). Nine eukaryotic MAGs were identified by eukCC with six assigned to the phylum Chlorophyta. Additionally, 6,350 high-quality viral sequences were identified by geNomad with the majority classified in the phylum Uroviricota. This expansive coassembled metagenomic dataset provides an unprecedented foundation to advance understanding of microbial communities in freshwater ecosystems and explore temporal ecosystem dynamics.

59 BASIC BIOLOGICAL SCIENCES↗

A high-throughput skim-sequencing approach for genotyping, dosage estimation and identifying translocations

The development of next-generation sequencing (NGS) enabled a shift from array-based genotyping to directly sequencing genomic libraries for high-throughput genotyping. Even though whole-genome sequencing was initially too costly for routine analysis in large populations such as breeding or genetic studies, continued advancements in genome sequencing and bioinformatics have provided the opportunity to capitalize on whole-genome information. As new sequencing platforms can routinely provide high-quality sequencing data for sufficient genome coverage to genotype various breeding populations, a limitation comes in the time and cost of library construction when multiplexing a large number of samples. Here we describe a high-throughput whole-genome skim-sequencing (skim-seq) approach that can be utilized for a broad range of genotyping and genomic characterization. Using optimized low-volume Illumina Nextera chemistry, we developed a skim-seq method and combined up to 960 samples in one multiplex library using dual index barcoding. With the dual-index barcoding, the number of samples for multiplexing can be adjusted depending on the amount of data required, and could be extended to 3,072 samples or more. Panels of doubled haploid wheat lines ( Triticum aestivum , CDC Stanley x CDC Landmark), wheat-barley ( T . aestivum x Hordeum vulgare ) and wheat-wheatgrass ( Triticum durum x Thinopyrum intermedium ) introgression lines as well as known monosomic wheat stocks were genotyped using the skim-seq approach. Bioinformatics pipelines were developed for various applications where sequencing coverage ranged from 1 × down to 0.01 × per sample. Using reference genomes, we detected chromosome dosage, identified aneuploidy, and karyotyped introgression lines from the skim-seq data. Leveraging the recent advancements in genome sequencing, skim-seq provides an effective and low-cost tool for routine genotyping and genetic analysis, which can track and identify introgressions and genomic regions of interest in genetics research and applied breeding programs.

60 APPLIED LIFE SCIENCES↗

Environmental RNA as a Tool for Marine Community Biodiversity Assessments

Abstract Microscopic organisms are often overlooked in traditional diversity assessments due to the difficulty of identifying them based on morphology. Metabarcoding is a method for rapidly identifying organisms where Environmental DNA (eDNA) is used as a template. However, legacy DNA is problematically detected from organisms no longer in the environment during sampling. Environmental RNA (eRNA), which is only produced by living organisms, can also be collected from environmental samples and used for metabarcoding. The aim of this study was to determine differences in community composition and diversity between eRNA and eDNA templates for metabarcoding. Using mesocosms containing field-collected communities from an estuary, RNA and DNA were co-extracted from sediment, libraries were prepared for two loci (18S and COI), and sequenced using an Illumina MiSeq. Results show a higher number of unique sequences detected from eRNA in both markers and higher α-diversity compared to eDNA. Significant differences between eRNA and eDNA for all β-diversity metrics were also detected. This study is the first to demonstrate community differences detected with eRNA compared to eDNA from an estuarine system and illustrates the broad applications of eRNA as a tool for assessing benthic community diversity, particularly for environmental conservation and management applications.

Giroux, Marissa S.↗

Full genome sequence for the African swine fever virus outbreak in the Dominican Republic in 1980

African swine fever is a lethal disease of domestic pigs, geographically expanding as a pandemic, that is affecting countries across Eurasia and severely damaging their swine production industry. After more than 40 years of being absent in the Western hemisphere, in 2020 ASF reappeared in the Dominican Republic and Haiti. The recent outbreak strain in the Dominican Republic has been identified as a genotype II ASFV a derivative of the ASF strain circulating in Asia and Europe. However, to date no full-length genome sequence from either the 1978–1980 Here we report the complete genome sequence of an African swine fever virus (ASFV) (DR-1980) that was previously isolated from blood collected in 1980 from the Dominican Republic at the end of the last outbreak, before culling of all swine on the island of Hispaniola and stored in the Plum Island Animal Disease Center ASFV repository. A contig representing the full-length genome (183,687 base pairs) was de novo assembled into a single contig using both Nanopore and Illumina sequences. DR-1980 was determined to belong to genotype I and, as determined by full genome comparison, a close relative to the sequenced Sardinia viruses that were causing outbreaks at this time.

60 APPLIED LIFE SCIENCES↗

Community RNA-Seq: multi-kingdom responses to living versus decaying roots in soil

Abstract Roots are a primary source of organic carbon input in most soils. The consumption of living and detrital root inputs involves multi-trophic processes and multiple kingdoms of microbial life, but typical microbial ecology studies focus on only one or two major lineages. We used Illumina shotgun RNA sequencing to conduct PCR-independent SSU rRNA community analysis (“community RNA-Seq”) and simultaneously assess the bacteria, archaea, fungi, and microfauna surrounding both living and decomposing roots of the annual grass, Avena fatua. Plants were grown in 13CO2-labeled microcosms amended with 15N-root litter to identify the preferences of rhizosphere organisms for root exudates (13C) versus decaying root biomass (15N) using NanoSIMS microarray imaging (Chip-SIP). When litter was available, rhizosphere and bulk soil had significantly more Amoebozoa, which are potentially important yet often overlooked top-down drivers of detritusphere community dynamics and nutrient cycling. Bulk soil containing litter was depleted in Actinobacteria but had significantly more Bacteroidetes and Proteobacteria. While Actinobacteria were abundant in the rhizosphere, Chip-SIP showed Actinobacteria preferentially incorporated litter relative to root exudates, indicating this group’s more prominent role in detritus elemental cycling in the rhizosphere. Our results emphasize that decomposition is a multi-trophic process involving complex interactions, and our methodology can be used to track the trajectory of carbon through multi-kingdom soil food webs.

Nuccio, Erin E. (ORCID:000000030189183X)↗

A global phylogenomic analysis of the shiitake genus Lentinula

Lentinula is a broadly distributed group of fungi that contains the cultivated shiitake mushroom, L. edodes . We sequenced 24 genomes representing eight described species and several unnamed lineages of Lentinula from 15 countries on four continents. Lentinula comprises four major clades that arose in the Oligocene, three in the Americas and one in Asia–Australasia. To expand sampling of shiitake mushrooms, we assembled 60 genomes of L. edodes from China that were previously published as raw Illumina reads and added them to our dataset. Lentinula edodes sensu lato (s. lat.) contains three lineages that may warrant recognition as species, one including a single isolate from Nepal that is the sister group to the rest of L. edodes s. lat., a second with 20 cultivars and 12 wild isolates from China, Japan, Korea, and the Russian Far East, and a third with 28 wild isolates from China, Thailand, and Vietnam. Two additional lineages in China have arisen by hybridization among the second and third groups. Genes encoding cysteine sulfoxide lyase ( lecsl ) and γ-glutamyl transpeptidase ( leggt ), which are implicated in biosynthesis of the organosulfur flavor compound lenthionine, have diversified in Lentinula . Paralogs of both genes that are unique to Lentinula ( lecsl 3 and leggt 5b) are coordinately up-regulated in fruiting bodies of L. edodes . The pangenome of L. edodes s. lat. contains 20,308 groups of orthologous genes, but only 6,438 orthogroups (32%) are shared among all strains, whereas 3,444 orthogroups (17%) are found only in wild populations, which should be targeted for conservation.

59 BASIC BIOLOGICAL SCIENCES↗