Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Illumina”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Mapping crown rust resistance in the oat diploid accession PI 258731 ( Avena strigosa )

Oat crown rust, caused by Puccinia coronata Corda f. sp. avenae Eriks. (Pca), is a major biotic impediment to global oat production. Crown rust resistance has been described in oat diploid species A. strigosa accession PI 258731 and resistance from this accession has been successfully introgressed into hexaploid A. sativa germplasm. The current study focuses on 1) mapping the location of QTL containing resistance and evaluating the number of quantitative trait loci (QTL) conditioning resistance in PI 258731; 2) understanding the relationship between the original genomic location in A. strigosa and the location of the introgression in the A. sativa genome; 3) identifying molecular markers tightly linked with PI 258731 resistance loci that could be used for marker assisted selection and detection of this resistance in diverse A. strigosa accessions. To achieve this, A. strigosa accessions, PI 258731 and PI 573582 were crossed to produce 168 F5:6 recombinant inbred lines (RILs) through single seed descent. Parents and RILs were genotyped with the 6K Illumina SNP array which generated 168 segregating SNPs. Seedling reactions to two isolates of Pca (races TTTG, QTRG) were conditioned by two genes (0.6 cM apart) in this population. Linkage mapping placed these two resistant loci to 7.7 (QTRG) to 8 (TTTG) cM region on LG7. Field reaction data was used for QTL analysis and the results of interval mapping (MIM) revealed a major QTL (QPc.FD-AS-AA4) for field resistance. SNP marker assays were developed and tested in 125 diverse A. strigosa accessions that were rated for crown rust resistance in Baton Rouge, LA and Gainesville, FL and as seedlings against races TTTG and QTRG. Our data proposed SNP marker GMI_ES17_c6425_188 as a candidate for use in marker-assisted selection, in addition to the marker GMI_ES02_c37788_255 suggested by Rine’s group, which provides an additional tool in facilitating the utilization of this gene in oat breeding programs.

60 APPLIED LIFE SCIENCES↗

Single-nuclei transcriptome analysis of channel catfish spleen provides insight into the immunome of an aquaculture-relevant species

The catfish industry is the largest sector of U.S. aquaculture production. Given its role in food production, the catfish immune response to industry-relevant pathogens has been extensively studied and has provided crucial information on innate and adaptive immune function during disease progression. To further examine the channel catfish immune system, we performed single-cell RNA sequencing on nuclei isolated from whole spleens, a major lymphoid organ in teleost fish. Libraries were prepared using the 10X Genomics Chromium X with the Next GEM Single Cell 3’ reagents and sequenced on an Illumina sequencer. Each demultiplexed sample was aligned to the Coco_2.0 channel catfish reference assembly, filtered, and counted to generate feature-barcode matrices. From whole spleen samples, outputs were analyzed both individually and as an integrated dataset. The three splenic transcriptome libraries generated an average of 278,717,872 reads from a mean 8,157 cells. The integrated data included 19,613 cells, counts for 20,121 genes, with a median 665 genes/cell. Cluster analysis of all cells identified 17 clusters which were classified as erythroid, hematopoietic stem cells, B cells, T cells, myeloid cells, and endothelial cells. Subcluster analysis was carried out on the immune cell populations. Here, distinct subclusters such as immature B cells, mature B cells, plasma cells, γδ T cells, dendritic cells, and macrophages were further identified. Differential gene expression analyses allowed for the identification of the most highly expressed genes for each cluster and subcluster. This dataset is a rich cellular gene expression resource for investigation of the channel catfish and teleost splenic immunome.

Science & Technology - Other Topics↗

Depth-resolved sagebrush root metabolomics, rhizosphere microbial communities, and geochemistry at the East River Watershed

This data set consists of results from soil nutrient profile, untargeted metabolomics, mass spec imaging, and amplicon sequencing. Data for soil nutrient profile includes common cations (Ca, Mg, Na, and K etc.) extracted from 3 digesting steps – ammonia acetate (for exchangeable cations), nitric acid (for acid dissolved fraction), and hydrofluoric acid/perchloric acid (HF/HClO4) for whole soil digestion. It also includes concentration of organic carbon, inorganic nitrogen (ammonia and nitrate) and phosphorus (Bray-1 P and nitric acid extract), and total nitrogen and phosphorus. Data for untargeted metabolomics includes metabolomic profile for root exudate/tissues and soil extracts from depths at surface soil to saprolite, that were measured using gas chromatography – mass spectrometry (GC-MS), and liquid chromatography – tandem mass spectrometry (LC-MS/MS). Data for mass spec imaging includes spatial distribution of metabolites that were detected and annotated with Fourier transformation ion cyclotron resonance mass spectrometer (FTICR-MS). Data for amplicon sequencing includes the base paired 16S and ITS ribosomal RNA sequences from Miseq Illumina sequencing. All samples were collected from 2 sampling campaign October 2022 and June 2023. Collectively, these datasets enable a mechanistic evaluation of how nutrient acquisition, especially nitrogen and phosphorus, differs between shallow roots operating in soil and deep roots functioning within the fractured bedrock zone. All files are provided as comma-separated values (CSV) fies (.csv) and (GZIP) file (.gz). The compressed .gz FASTQ files can be read directly in R using the dada2 package as part of the amplicon sequence analysis workflow. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. This research was performed on a project award 60563 (https://dx.doi.org/10.46936/expl.proj.2022.60563/60008727) from the Environmental Molecular Sciences Laboratory, a DOE Office of Science User Facility sponsored by the Biological and Environmental Research program under Contract No. DE-AC05-76RL01830.

EARTH SCIENCE > AGRICULTURE > SOILS > CARBON↗

Sample Type Adaptable RNA depletion performance data [Slides]

Sequencing data was aligned to rRNA data sets QIIME_16S_MiDAS_4.8.1 and SILVA_138.1_LSURef_NR99 using BWA. RNA used was a composite (pool) of wastewater RNA samples collected at LANL between April 2022 and December 2022. RNA sample was split into four aliquots (STAR depletion and STAR depletion no-probe-control as well as RiboZero and RiboZero no-probe-control). All sequencing libraries were prepared from the depleted RNA samples using the same library prep method and sequenced on Illumina platforms.

59 BASIC BIOLOGICAL SCIENCES↗

Improved Biofuel Production through Discovery and Engineering of Terpene Metabolism in Switchgrass

Project Objectives - Of the myriad specialized metabolites that plants deploy to adapt to environmental challenges, terpenes form the largest group. In many major crops, unique terpene blends serve as key stress defenses that directly impact plant fitness and yield. In addition, terpenes, such as bisabolene and pinene, are used for producing renewable biofuels. Essential to advancing a broader use of terpenes for biofuel feedstock engineering is a system-wide knowledge of the diverse biosynthetic machinery and defensive potential of often species-specific terpene blends. The proposed project would merge genome-wide enzyme discovery with comparative –omics, protein structural and plant microbiome studies to define the biosynthesis and stress-defensive functions of the switchgrass (Panicum virgatum) terpene network. These insights would be combined with developing and applying non-transgenic genome editing tools to design plants with desirable terpene blends for higher productivity and biofuel production on marginal lands. As a dedicated lignocellulosic feedstock for U.S. biofuel production with high net energy yield, stress tolerance, and available genome resources, switchgrass is well-suited for devising new avenues for biofuel production. Project Description – The diversity of plant terpene defenses is governed by species-specific families of terpene synthase (TPS) and cytochrome P450 monooxygenase (P450) enzymes. Mining of the switchgrass genome (genotype Alamo) identified ~100 TPS and P450 candidate genes, and combinatorial biochemical analysis of synthesized TPSs and P450s revealed more than a dozen enzymes with common and novel activities. In addition, several identified terpene metabolites and the corresponding transcripts were up-regulated in response to abiotic stressors. These findings demonstrate a unique switchgrass terpene network with probable importance to abiotic stress tolerance, thus providing a large chemical portfolio for optimizing crop resistance, yield, and biofuel composition. Leveraging these preliminary data, we propose to generate a genome-wide map of the switchgrass terpene metabolic network through multi-gene co-expression analyses that allow the efficient cross-validation of TPS and P450 functions. Key enzymes would further be applied to structure-function studies via X-ray protein crystallography, homology modeling and site-directed mutagenesis to gain mechanistic insight into the catalytic specificity of switchgrass terpene metabolism and provide gene and amino acid targets for genome editing. In tandem with terpene pathway discovery, system-wide metabolomics, transcriptomics and proteomics studies in switchgrass accessions of contrasting drought tolerance would define the role of switchgrass terpene metabolism in conferring abiotic stress resilience. Metabolic changes would be assessed in a combined approach of targeted (terpenes) and untargeted metabolite profiling using a high-resolution LC-MS/MS approach, differential gene expression analyses through multiplexed Illumina RNA sequencing, and quantitative analysis of high-priority pathway enzymes using multiple reaction monitoring (MRM). Drawing on these insights, knock-down/out mutants of stress-associated pathway nodes would be generated by optimizing transient virus-induced gene silencing (VIGS) and CRISPR/Cas9 systems under control of the Tobacco Rattle Virus (TRV). The resulting mutant lines would then be analyzed for stress susceptibility and the impact on the root microbiome to define gene functions in planta. Knowledge of terpene pathways, enzyme mechanisms and bioactivities would be applied to enhance switchgrass stress resilience and to tailor-make terpene blends for biofuel production. Here, TRV-enabled CRISPR/Cas9 genome editing, including allele-specific knock-out of redundant genes, engineering of enzyme specificity via structure-guided point mutations, and overexpression of terpene genes relevant to stress-protection or biofuel production, would be used to increase metabolic flux toward desired pathways. Broader Impacts - Integrating the system-wide discovery, mechanistic analysis and non-transgenic genome engineering of the switchgrass terpene network aligns the required steps to unlock the chemical potential of this important metabolite class to generate crops that are more resistant to stress and provide advanced biofuel production in light of rising climate pressures as foreseeable challenges for bioenergy crop cultivation. The proposed project would further offer interdisciplinary student training through active involvement in the project and integration of research concepts and outcomes into newly-developed graduate and undergraduate courses on Plant Biotechnology.

09 BIOMASS FUELS↗

16s Amplicon Analysis of Soil Data for Interactive effects of depth and differential irrigation on soil microbiome composition and functioning

Genomic DNA was isolated from soil and rhizosphere samples using the Zymo Quick-DNA fecal/soil microbe miniprep kit (catalog no. D6010) according to the manufacturer’s instructions (Zymo Research; Irvine, CA) with the modification of eluting in 100 uL elution buffer. Sample concentrations were quantified using the Qubit dsDNA HS assay kit (Thermo Fisher). For rhizosphere samples only, DNA was subsequently purified using Zymo’s ZR-96 DNA Clean & Concentrator kit (catalog no. D4024) to account for low DNA concentrations of these samples. . In each replicate block, there were five drip irrigation treatments (T1 = 100% normal irrigation, T2 = 56.25%, T3 = 37.5%, T4 = 18.75% and T5=no irrigation. On July 20, the strength of the drought treatments was increased: T1 remained at 100%, whereas T2 changed from 75% to 56.25%, T2 changed from 50% to 37.5%, and T4 changed from 25% to 18.75%. Sequencing was performed as described previously (Naylor, Fansler, et al. 2020). Sequences were amplified on the MiSeq platform (Illumina, San Diego, CA) using 16S primers (515F and 806R) specific to the V4 region. Raw sequence data was processed with the pipeline Hundo for amplicon quality control and annotation. Downstream statistical analyses on 16S datasets were performed using the program R and the packages ‘phyloseq’ and ‘vegan’.

Soil microbiome, metatranscriptomics↗

Metatransciptomic Analysis Data for Interactive effects of depth and differential irrigation on soil microbiome composition and functioning

RNA was collected from soil at different depths and after three different levels of irrigation T1 100% of normal field irrigation, T4: 18.75% or normal irrigation and T5: unirrigated controls. Total RNA was isolated using the Zymo Quick-RNA fecal/soil microbe miniprep (catalog no. R2040), incorporating the DNase I treatment using Zymo’s DNase I kit (catalog no. E1010). To increase the yield of RNA, we modified the manufacturer’s instructions by first doubling the amount of soil per extraction (from 0.25 g to 0.5 g) and by performing extractions in triplicate before pooling separate extractions together. Certain soil samples (largely those from deeper soil layers) had low yield (< 100 ng per extraction) so additional rounds of extraction were performed to obtain sufficient RNA. RNA concentration was assessed using a Qubit RNA HS assay kit (Thermo Fisher) and RNA quality was determined using an Agilent 2100 BioAnalyzer (Agilent; Santa Clara, CA). The resultant RNA samples were then sequenced by GENEWIZ using Illumina technology (GENEWIZ; South Plainfield, NJ). Sequences were then aligned to a soil metagenome previously obtained from the same site using the Burrows-Wheeler aligner (BWA). SAM files were then converted to raw counts using HTSeq.

Soil microbiome, metatranscriptomics↗

PPI DataHub Project Data Package: S. elongatus PCC 7942 Circadian Control Bioproduction Transcriptomics (PB-DP3)

The purpose of this experiment was to evaluate how circadian clock regulation impacts carbon partitioning between storage, growth, and product synthesis in Synechococcus elongatus PCC 7942 in providing insights to strategies for enhanced bioproduction. Sample data was acquired using a Illumina HiSeq sequencer system and processed for RNA sequencing (RNA-Seq) expression analysis. Transcriptomic differential expression analysis revealed coordinated circadian clock-driven adjustment of the cell cycle and rewiring of energy and carbon metabolism. Processed RNA-Seq datasets are openly accessible from the PNNL DataHub project dataset download page and contain secondary processed RNA-seq results files and supporting metadata materials linked to relevant source code information supporting data transparency and reuse.

59 BASIC BIOLOGICAL SCIENCES↗

Human Liver Epithelial Cells (HuH7) Response to HCoV-229E Infection Epigenomics (ATAC-Seq) (ACS-DP4)

The purpose of this experiment was to evaluate how wild-type Human coronavirus strain 229E (HCoV-299E) infection alters chromatin accessibility in infected cells. Sample data was obtained from mock-infected cells, UV-inactivated virus treated cells, and replication competent HCoV-229E infected immortalized human liver cells (HuH7) at 24 hours post infection. Samples were processed using ATAC-seq methods for reported bar coded libraries. Sample data was acquired using an Illumina Hi-Seq 2500 sequencer system and further processed for ATAC-Seq expression analysis.

59 BASIC BIOLOGICAL SCIENCES↗

Human Host Cellular Response to HCoV-229E Infection Transcriptomics (ACS-DP1)

The purpose of this experiment was to evaluate the human host cellular response to wild-type Human coronavirus strain 229E (HCoV-229E) infection. Sample data was obtained for mock and infected immortalized human lung epithelial cells (A549) (MOI 5), immortalized human lung fibroblasts cells (MRC5) (MOI5), and primary human airway epithelial (HAE) (MOI 3) cells from lung tissue. Sample data was acquired using an Illumina HiSeq 2000 sequencer system and processed for RNA sequencing (RNA-Seq) expression analysis.

59 BASIC BIOLOGICAL SCIENCES↗

Human Primary Airway Epithelium +/- Macrophages Response to HCoV-229E Infection Transcriptomics (ACS-DP3)

The purpose of this experiment was to evaluate the human host cellular response to wild-type Human coronavirus strain 229E (HCoV-299E) infection. Sample data was obtained for mock and infected (MOI 3) primary human airway epithelial cells with and without macrophages and grown in air-liquid interface conditions. Sample data was acquired using an Illumina Hi-Seq 4000 sequencer system and further processed for RNA sequencing (RNA-Seq) expression analysis.

59 BASIC BIOLOGICAL SCIENCES↗

Genomes_SET2_178 - Shared

The dataset consists of 178 bacterial genomes in total. Of these, 145 genomes were drawn from published literature linked to Biolog Gen III growth data, while 33 genomes came from soil isolates collected at the Cedar Creek Ecosystem Science Research Reserve (CCESR) in Minnesota. The isolates were sequenced on an Illumina HiSeq platform (~50× coverage), assembled with Megahit, and annotated using RAST at PATRIC, with strict quality filtering (<90% completeness or >5% contamination).

59 BASIC BIOLOGICAL SCIENCES↗

Five draft genome assemblies from Bacillaceae isolated from a degraded wetland environment

Abstract We isolated 5 Bacillaceae from a degraded wetland environment and sequenced their genomes using Illumina NextSeq. Here, we report draft genome sequences of Bacillus velezensus-SC119, Priestia megaterium-SC120, Bacillus zhangzhouensis-SC123, Bacillus pumilis-SC124, and Bacillus idriensis-SC127. The genomes range between 3,657,353 and 5,772,725 base pairs with %GC between 37.62% and 46.38%. Introduction Wetland environments play critical roles in the terrestrial carbon and water cycles and microbial communities are key players in healthy ecosystem function. Endospore forming bacteria in the Bacillaceae family are metabolically and genomically diverse soil heterotrophs that influence plant health, carbon and nitrogen cycling, and often produce diverse natural products that influence other bacterial and non-bacterial species in their environment (1). We collected two soil samples on January 19, 2023 from 42°43'12.7"N 73°45'01.4"W. One was highly hydrated and within a patch of invasive common reeds (Phragmites sp.) and the other was near the base of an Eastern cottonwood tree (Populus deltoides). Bacillus pumilis strain SC124 was isolated from the soil from near the cottonwood tree, while Bacillus velezensis strain SC119, Priestia megaterium strain SC120, Bacillus zhangzhouensis strain SC123, Bacillus idriensis strain SC127 and were isolated from the marshy soil.

isolate, wetlands, genome announcement↗

Metabacillus indicus EGFCL74

Chromosomal DNA of bacterium isolated from fermented cider (originally named m74) was sequenced on an Illumina NovaSeq platform 2x150bp. Raw paired end reads were trimmed and processed with BBDUK v39.01. Trimmed FASTQ files were uploaded to the KBase narrative. Within KBase, the paired-end reads were assembled using SPAdes v3.15.3 (output data file m74_SPAdes.Assembly). Completeness was evaluated with CheckM v1.0.18. The assembled genome was annotated using RAStk v1.073 (output data file m74_RAStkgenomeassembly). Comparison of the isolate genome to other available genomes was performed two ways. The InsertGenomeIntoSpeciesTree function placed the isolate in the same clade as Bacillus indicus (later renamed Metabacillus indicus). FastANI was then used to compare the average nucleotide identity of isolate (m74) to 3 available strains of Metabacillus indicus (ASM70993v2, ASM70875v2, 4-1317).

59 BASIC BIOLOGICAL SCIENCES↗

Seven soil endospore forming bacteria from campus woodland fragments

We isolated 7 endospore forming bacteria from campus woodland and sequenced their genomes using Illumina NextSeq. We share the draft genome assemblies for strains Bacillus wiedmanii_SC129, Bacillus pseudomycoides_SC131, Bacillus pumilis_SC133, Peribacillus butanolivorans_SC135, Bacillus thuringiensis_SC136, Priestia megaterium_SC138, and Bacillus wiedmanii_SC141. Draft genomes are between 3645032-5969865 bp and 34.8-41.2 % GC.

59 BASIC BIOLOGICAL SCIENCES↗

Genome collection processing for “Conserved upper thermal limits and small safety margins in soil copiotrophic bacteria”

We extracted the genomic DNA of 400 randomly selected isolates using a Quick-DNA Microprep Kit (Zymo Research D3020) according to the manufacturer’s protocol. We then submitted the extracted gDNA samples for short-read Illumina sequencing (200 Mbp) at SeqCoast Genomics (Portsmouth, NH, USA). After preprocessing the sequences using Trimmommatic (Bolger et al. 2014), we assembled the genomes using SPADES (Bankevich et al. 2012) and checked the quality of each assembly using QUAST (Gurevich et al. 2013). We processed the genome assemblies using a KBase (v1.4.0) pipeline (Allen et al. 2017; Arkin et al. 2018). Briefly, we used DRAM (v0.1.2) with default settings to annotate the genome assemblies. We then evaluated genome quality and possible contamination levels using CheckM (v1.0.18) (Parks et al. 2015) and retained genomes with completeness above 98% and contamination below 5% (n = 354), following the authors' guidelines. We then obtained taxonomic assignments for all remaining isolates using the Genome Taxonomy Database tool GTDB-Tk (v2.3.2, database version r214) (Chaumeil et al. 2019). We constructed a phylogenetic tree using the tool SpeciesTree (v2.2.0). We then trimmed the tree (using Trim SpeciesTree to GenomeSet- v1.4.0), retaining only tips within our collection with measured thermal performance.

59 BASIC BIOLOGICAL SCIENCES↗

Comprehensive SNP Data for 1,323 GWAS Population in Populus trichocarpa and Combined Annotation Files for P. trichocarpa v3.0 and v3.1

The VCF dataset includes genetic variations found in 1,323 Populus trichocarpa genotypes, providing valuable information for scientists studying plant genetics. Researchers have generated this dataset using whole-genome DNA short-read sequencing on the Illumina Genome Analyzer, HiSeq 2000, and HiSeq 2500 platforms. This sequencing effort ensured a minimum expected sequencing depth of 15×. The dataset comprises more than 9.7 million single nucleotide polymorphisms (SNPs) and indel variants. The combined annotation files are derived from P. trichocarpa v3.0 and v3.1. We merged these files to create a comprehensive annotation file used for GWAS analysis. In total, 38,830 genes overlapped between the two versions. For overlapping genes, we defined the start as the smaller and the end as the larger among the two versions to increase the likelihood of locating candidate genetic loci. Additionally, we included 2,505 unique genes from v3.0 and 4,120 unique genes from v3.1, resulting in a total of 45,455 genes in the updated annotation file.

09 BIOMASS FUELS↗

Next-generation sequencing dataset of genome-scale CRISPRi in Synechococcus sp. PCC 7002 across seven conditions

A 33,298-member sgRNA library developed for Synechococus sp. PCC 7002 was screened with two replicates across seven growth conditions and sequenced with Illumina NextSeq (paired end, 2x150 bp) for a total of ~700M reads. The original plasmid library and the library after transformation into a dCas9-containing and dCas9-absent strain were also sequenced as a reference for initial sgRNA abundance.

genome- wide screens environmental acclimation spe↗