Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “DNA sequencing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

kb_DRAM: annotation and metabolic profiling of genomes with DRAM in KBase

Microbial genome annotation is the process of identifying structural and functional elements in DNA sequences and subsequently attaching biological information to those elements. DRAM is a tool developed to annotate bacterial, archaeal, and viral genomes derived from pure cultures or metagenomes. DRAM goes beyond traditional annotation tools by distilling multiple gene annotations to genome level summaries of functional potential. Despite these benefits, a downside of DRAM is the requirement of large computational resources, which limits its accessibility. Further, it did not integrate with downstream metabolic modeling tools that require genome annotation. To alleviate these constraints, DRAM and the viral counterpart, DRAM-v, are now available and integrated with the freely accessible KBase cyberinfrastructure. With kb_DRAM users can generate DRAM annotations and functional summaries from microbial or viral genomes in a point-and-click interface, as well as generate genome-scale metabolic models from DRAM annotations.

59 BASIC BIOLOGICAL SCIENCES↗

Taxogenomic analysis of Pichia senei sp. nov. and new insights into hybridization events in the Pichia cactophila species complex

Three strains of a novel yeast species were isolated from necrotic cactus tissues of Cereus saddianus and Micranthocereus dolichospermaticus and from phytotelmata of Bromelia karatas. DNA sequence analysis of the Internal Transcribed Spacer (ITS) region and D1/D2 domains of the large subunit ribosomal RNA, along with whole genome phylogenomic analysis, showed that this yeast is most closely related to Pichia insulana, Pichia cactophila, and Pichia inconspicua. The new species differs by 10–13 nucleotide substitutions from these species in D1/D2 sequences and exhibits <90% genome-wide average nucleotide identity to them. The name Pichia senei sp. nov. is proposed for the novel species, which is homothallic and produces asci with one to four hat-shaped ascospores. The holotype is CBS 16311 (MycoBank MB 858723). Taxogenomic analyses of the P. cactophila species complex, including P. senei, provide new insights about the hybridizations events that shaped this group. Pichia insulana and P. inconspicua are identified as the parental lineages that originated P. cactophila, and P. senei also appears closely related to one of the progenitors of P. inconspicua. We assess phylogeny, heterozygosity, and ploidy to explore the processes shaping diversity, showing how genomic data support yeast species delimitation and reveal complex hybridization.

59 BASIC BIOLOGICAL SCIENCES↗

Robust, versatile DNA FISH probes for chromosome-specific repeats in Caenorhabditis elegans and Pristionchus pacificus

Repetitive DNA sequences are useful targets for chromosomal fluorescence in situ hybridization. We analyzed recent genome assemblies of Caenorhabditis elegans and Pristionchus pacificus to identify tandem repeats with a unique genomic localization. Based on these findings, we designed and validated sets of oligonucleotide probes for each species targeting at least 1 locus per chromosome. These probes yielded reliable fluorescent signals in different tissues and can easily be combined with the immunolocalization of cellular proteins. Synthesis and labeling of these probes are highly cost-effective and require no hands-on labor. The methods presented here can be easily applied in other model and nonmodel organisms with a sequenced genome.

59 BASIC BIOLOGICAL SCIENCES↗

High-resolution mapping reveals hotspots and sex-biased recombination in Populus trichocarpa

Abstract Fine-scale meiotic recombination is fundamental to the outcome of natural and artificial selection. Here, dense genetic mapping and haplotype reconstruction were used to estimate recombination for a full factorial Populus trichocarpa cross of 7 males and 7 females. Genomes of the resulting 49 full-sib families (N = 829 offspring) were resequenced, and high-fidelity biallelic SNP/INDELs and pedigree information were used to ascertain allelic phase and impute progeny genotypes to recover gametic haplotypes. The 14 parental genetic maps contained 1,820 SNP/INDELs on average that covered 376.7 Mb of physical length across 19 chromosomes. Comparison of parental and progeny haplotypes allowed fine-scale demarcation of cross-over regions, where 38,846 cross-over events in 1,658 gametes were observed. Cross-over events were positively associated with gene density and negatively associated with GC content and long-terminal repeats. One of the most striking findings was higher rates of cross-overs in males in 8 out of 19 chromosomes. Regions with elevated male cross-over rates had lower gene density and GC content than windows showing no sex bias. High-resolution analysis identified 67 candidate cross-over hotspots spread throughout the genome. DNA sequence motifs enriched in these regions showed striking similarity to those of maize, Arabidopsis, and wheat. These findings, and recombination estimates, will be useful for ongoing efforts to accelerate domestication of this and other biomass feedstocks, as well as future studies investigating broader questions related to evolutionary history, perennial development, phenology, wood formation, vegetative propagation, and dioecy that cannot be studied using annual plant model systems.

59 BASIC BIOLOGICAL SCIENCES↗

Disentangling the effects of sulfate and other seawater ions on microbial communities and greenhouse gas emissions in a coastal forested wetland

Seawater intrusion into freshwater wetlands causes changes in microbial communities and biogeochemistry, but the exact mechanisms driving these changes remain unclear. Here we use a manipulative laboratory microcosm experiment, combined with DNA sequencing and biogeochemical measurements, to tease apart the effects of sulfate from other seawater ions. We examined changes in microbial taxonomy and function as well as emissions of carbon dioxide, methane, and nitrous oxide in response to changes in ion concentrations. Greenhouse gas emissions and microbial richness and composition were altered by artificial seawater regardless of whether sulfate was present, whereas sulfate alone did not alter emissions or communities. Surprisingly, addition of sulfate alone did not lead to increases in the abundance of sulfate reducing bacteria or sulfur cycling genes. Similarly, genes involved in carbon, nitrogen, and phosphorus cycling responded more strongly to artificial seawater than to sulfate. These results suggest that other ions present in seawater, not sulfate, drive ecological and biogeochemical responses to seawater intrusion and may be drivers of increased methane emissions in soils that received artificial seawater addition. A better understanding of how the different components of salt water alter microbial community composition and function is necessary to forecast the consequences of coastal wetland salinization.

54 ENVIRONMENTAL SCIENCES↗

Novel candidate taxa contribute to key metabolic processes in Fennoscandian Shield deep groundwaters

The continental deep biosphere contains a vast reservoir of microorganisms, although a large proportion of its diversity remains both uncultured and undescribed. In this study, the metabolic potential (metagenomes) and activity (metatranscriptomes) of the microbial communities in Fennoscandian Shield deep subsurface groundwaters were characterized with a focus on novel taxa. DNA sequencing generated 1270 de-replicated metagenome-assembled genomes and single-amplified genomes, containing 7 novel classes, 34 orders, and 72 families. The majority of novel taxa were affiliated with Patescibacteria, whereas among novel archaea taxa, Thermoproteota and Nanoarchaeota representatives dominated. Metatranscriptomes revealed that 30 of the 112 novel taxa at the class, order, and family levels were active in at least one investigated groundwater sample, implying that novel taxa represent a partially active but hitherto uncharacterized deep biosphere component. The novel taxa genomes coded for carbon fixation predominantly via the Wood–Ljungdahl pathway, nitrogen fixation, sulfur plus hydrogen oxidation, and fermentative pathways, including acetogenesis. These metabolic processes contributed significantly to the total community’s capacity, with up to 9.9% of fermentation, 6.4% of the Wood–Ljungdahl pathway, 6.8% of sulfur plus 8.6% of hydrogen oxidation, and energy conservation via nitrate (4.4%) and sulfate (6.0%) reduction. Key novel taxa included the UBA9089 phylum, with representatives having a prominent role in carbon fixation, nitrate and sulfate reduction, and organic and inorganic electron donor oxidation. These data provided insights into deep biosphere microbial diversity and their contribution to nutrient and energy cycling in this ecosystem.

Candidatus↗

The solution structures of higher-order human telomere G-quadruplex multimers

Human telomeres contain the repeat DNA sequence 5'-d(TTAGGG), with duplex regions that are several kilobases long terminating in a 3' single-stranded overhang. The structure of the single-stranded overhang is not known with certainty, with disparate models proposed in the literature. We report here the results of an integrated structural biology approach that combines small-angle X-ray scattering, circular dichroism (CD), analytical ultracentrifugation, size-exclusion column chromatography and molecular dynamics simulations that provide the most detailed characterization to date of the structure of the telomeric overhang. We find that the single-stranded sequences 5'-d(TTAGGG)n, with n = 8, 12 and 16, fold into multimeric structures containing the maximal number (2, 3 and 4, respectively) of contiguous G4 units with no long gaps between units. The G4 units are a mixture of hybrid-1 and hybrid-2 conformers. In the multimeric structures, G4 units interact, at least transiently, at the interfaces between units to produce distinctive CD signatures. Global fitting of our hydrodynamic and scattering data to a worm-like chain (WLC) model indicates that these multimeric G4 structures are semi-flexible, with a persistence length of ~34 Å. Investigations of its flexibility using MD simulations reveal stacking, unstacking, and coiling movements, which yield unique sites for drug targeting.

59 BASIC BIOLOGICAL SCIENCES↗

The secondary metabolism collaboratory: a database and web discussion portal for secondary metabolite biosynthetic gene clusters

Secondary metabolites are small molecules produced by all corners of life, often with specialized bioactive functions with clinical and environmental relevance. Secondary metabolite biosynthetic gene clusters (BGCs) can often be identified within DNA sequences by various sequence similarity tools, but determining the exact functions of genes in the pathway and predicting their chemical products can often only be done by careful, manual comparative analysis. To facilitate this, we report the first release of the secondary metabolism collaboratory (SMC), which aims to provide a comprehensive, tool-agnostic repository of BGC sequence data drawn from all publicly available and user-submitted bacterial and archaeal genome and contig sources. On the website, users are provided a searchable catalog of putative BGCs identified from each source, along with visualizations of gene and domain annotations derived from multiple sequence analysis tools. SMC’s data is also available through publicly-accessible application programming interface (API) endpoints to facilitate programmatic access. Users are encouraged to share their findings (and search for others’) through comment posts on BGC and source pages. At the time of writing, SMC is the largest repository of BGC information, holding 13.1M BGC regions from 1.3M source sequences and growing, and can be found at https://smc.jgi.doe.gov.

59 BASIC BIOLOGICAL SCIENCES↗

Interaction of N-methylmesoporphyrin IX with a hybrid left-/right-handed G-quadruplex motif from the promoter of the SLC2A1 gene

Abstract Left-handed G-quadruplexes (LHG4s) belong to a class of recently discovered noncanonical DNA structures under the larger umbrella of G-quadruplex DNAs (G4s). The biological relevance of these structures and their ability to be targeted with classical G4 ligands is underexplored. Here, we explore whether the putative LHG4 DNA sequence from the SLC2A1 oncogene promoter maintains its left-handed characteristics upon addition of nucleotides in the 5′- and 3′-direction from its genomic context. We also investigate whether this sequence interacts with a well-established G4 binder, N-methylmesoporphyrin IX (NMM). We employed biophysical and X-ray structural studies to address these questions. Our results indicate that the sequence d[G(TGG)3TGA(TGG)4] (termed here as SLC) adopts a two-subunit, four-tetrad hybrid left-/right-handed G4 (LH/RHG4) topology. Addition of 5′-G or 5′-GG abolishes the left-handed fold in one subunit, while the addition of 3′-C or 3′-CA maintains the original fold. X-ray crystal structure analyses show that SLC maintains the same hybrid LH/RHG4 fold in the solid state and that NMM stacks onto the right-handed subunit of SLC. NMM binds to SLC with a 1:1 stoichiometry and a moderate-to-tight binding constant of 15 μM−1. This work deepens our understanding of LHG4 structures and their binding with traditional G4 ligands.

Seth, Paul↗

Enhancers in Plant Development, Adaptation and Evolution

Understanding plant responses to developmental and environmental cues is crucial for studying morphological divergence and local adaptation. Gene expression changes, governed by cis-regulatory modules (CRMs) including enhancers, are a major source of plant phenotypic variation. However, while genome-wide approaches have revealed thousands of putative enhancers in mammals, far fewer have been identified and functionally characterized in plants. This review provides an overview of how enhancers function to control gene regulation, methods to predict DNA sequences that may have enhancer activity, methods utilized to functionally validate enhancers and the current knowledge of enhancers in plants, including how they impact plant development, response to environment and evolutionary adaptation.

59 BASIC BIOLOGICAL SCIENCES↗

Machine learning-enabled phenotyping for GWAS and TWAS of WUE traits in 869 field-grown sorghum accessions

Abstract Sorghum (Sorghum bicolor) is a model C4 crop made experimentally tractable by extensive genomic and genetic resources. Biomass sorghum is studied as a feedstock for biofuel and forage. Mechanistic modeling suggests that reducing stomatal conductance (gs) could improve sorghum intrinsic water use efficiency (iWUE) and biomass production. Phenotyping to discover genotype-to-phenotype associations remains a bottleneck in understanding the mechanistic basis for natural variation in gs and iWUE. This study addressed multiple methodological limitations. Optical tomography and a machine learning tool were combined to measure stomatal density (SD). This was combined with rapid measurements of leaf photosynthetic gas exchange and specific leaf area (SLA). These traits were the subject of genome-wide association study and transcriptome-wide association study across 869 field-grown biomass sorghum accessions. The ratio of intracellular to ambient CO2 was genetically correlated with SD, SLA, gs, and biomass production. Plasticity in SD and SLA was interrelated with each other and with productivity across wet and dry growing seasons. Moderate-to-high heritability of traits studied across the large mapping population validated associations between DNA sequence variation or RNA transcript abundance and trait variation. A total of 394 unique genes underpinning variation in WUE-related traits are described with higher confidence because they were identified in multiple independent tests. This list was enriched in genes whose Arabidopsis (Arabidopsis thaliana) putative orthologs have functions related to stomatal or leaf development and leaf gas exchange, as well as genes with nonsynonymous/missense variants. These advances in methodology and knowledge will facilitate improving C4 crop WUE.

54 ENVIRONMENTAL SCIENCES↗

FUSARIUM-ID v.3.0: An Updated, Downloadable Resource for Fusarium Species Identification

Species within Fusarium are of global agricultural, medical, and food/feed safety concern and have been extensively characterized. However, accurate identification of species is challenging and usually requires DNA sequence data. FUSARIUM-ID ( http://isolate.fusariumdb.org/blast.php ) is a publicly available database designed to support the identification of Fusarium species using sequences of multiple phylogenetically informative loci, especially the highly informative ~680-bp 5' portion of the translation elongation factor 1-alpha (TEF1) gene that has been adopted as the primary barcoding locus in the genus. However, FUSARIUM-ID v.1.0 and 2.0 had several limitations, including inconsistent metadata annotation for the archived sequences and poor representation of some species complexes and marker loci. Here, we present FUSARIUM-ID v.3.0, which provides the following improvements: (i) additional and updated annotation of metadata for isolates associated with each sequence, (ii) expanded taxon representation in the TEF1 sequence database, (iii) availability of the sequence database as a downloadable file to enable local BLAST queries, and (iv) a tutorial file for users to perform local BLAST searches using either freely available software, such as SequenceServer, BLAST+ executable in the command line, and Galaxy, or the proprietary Geneious software. FUSARIUM-ID will be updated on a regular basis by archiving sequences of TEF1 and other loci from newly identified species and greater in-depth sampling of currently recognized species.

Plant Sciences↗

Escaping the fate of Sisyphus: assessing resistome hybridization baits for antimicrobial resistance gene capture

Finding, characterizing and monitoring reservoirs for antimicrobial resistance (AMR) is vital to protecting public health. Hybridization capture baits are an accurate, sensitive and cost-effective technique used to enrich and characterize DNA sequences of interest, including antimicrobial resistance genes (ARGs), in complex environmental samples. We demonstrate the continued utility of a set of 19 933 hybridization capture baits designed from the Comprehensive Antibiotic Resistance Database (CARD)v1.1.2 and Pathogenicity Island Database (PAIDB)v2.0, targeting 3565 unique nucleotide sequences that confer resistance. We demonstrate the efficiency of our bait set on a custom-made resistance mock community and complex environmental samples to increase the proportion of on-target reads as much as >200-fold. However, keeping pace with newly discovered ARGs poses a challenge when studying AMR, because novel ARGs are continually being identified and would not be included in bait sets designed prior to discovery. Here we provide imperative information on how our bait set performs against CARDv3.3.1, as well as a generalizable approach for deciding when and how to update hybridization capture bait sets. This research encapsulates the full life cycle of baits for hybridization capture of the resistome from design and validation (both in silico and in vitro) to utilization and forecasting updates and retirement.

59 BASIC BIOLOGICAL SCIENCES↗

An integrative framework reveals widespread gene flow during the early radiation of oaks and relatives in Quercoideae (Fagaceae)

ABSTRACT Although the frequency of ancient hybridization across the Tree of Life is greater than previously thought, little work has been devoted to uncovering the extent, timeline, and geographic and ecological context of ancient hybridization. Using an expansive new dataset of nuclear and chloroplast DNA sequences, we conducted a multifaceted phylogenomic investigation to identify ancient reticulation in the early evolution of oaks (Quercus). We document extensive nuclear gene tree and cytonuclear discordance among major lineages ofQuercusand relatives in Quercoideae. Our analyses recovered clear signatures of gene flow against a backdrop of rampant incomplete lineage sorting, with gene flow most prevalent among major lineages ofQuercusand relatives in Quercoideae during their initial radiation, dated to the Early‐Middle Eocene. Ancestral reconstructions including fossils suggest ancestors ofCastanea + Castanopsis,Lithocarpus, and the Old World oak clade probably co‐occurred in North America and Eurasia, while the ancestors ofChrysolepis, Notholithocarpus, and the New World oak clade co‐occurred in North America, offering ample opportunity for hybridization in each region. Our study shows that hybridization—perhaps in the form of ancient syngameons like those seen today—has been a common and important process throughout the evolutionary history of oaks and their relatives. Concomitantly, this study provides a methodological framework for detecting ancient hybridization in other groups.

Biochemistry & Molecular Biology↗

Comparative physiological and genomic characterization of a novel Nitrobacter vulgaris strain from a nitrate-contaminated subsurface

Nitrite-oxidizing bacteria (NOB) represent a crucial node in the global nitrogen cycle. By catalyzing the second step of nitrification—the oxidation of nitrite to nitrate to generate energy for growth—NOB activity controls the fate of nitrite (NO 2 - ) in aerobic environments. Despite thriving in diverse environments, including soils, freshwater, marine ecosystems, subsurface habitats, and water treatment systems, organisms capable of nitrite oxidation are confined to Nitrobacter, Nitrospira, Nitrospina, Nitrotoga, and a few other specific lineages. The genus Nitrobacter, recognized for its facultative heterotrophic metabolism, is often associated with high-nitrogen environments. Here, we report the physiological characterization of a novel strain, Nitrobacter vulgaris strain MLSD-S22, isolated from a nitrate- and heavy-metal-contaminated subsurface. Growth inhibition experiments revealed that strain MLSD-S22 and the N. vulgaris type strain Z exhibited similar sensitivities to nitrite and nitrate, with nitrite being the most inhibitory. Microrespirometry demonstrated that the two N. vulgaris strains and Nitrobacter winogradskyi Nb-255 possessed higher affinities for nitrite and oxygen than previously reported for Nitrobacter, suggesting potential to compete in low-substrate environments. Long-read DNA sequencing provided a complete genome for strain MLSD-S22, revealing two plasmids and an intact nitrous oxide (N 2 O) reduction operon—an unexpected feature for Nitrobacter. While N 2 O reduction activity was not observed under the tested conditions, this discovery raises questions about the contribution of Nitrobacter NOB to the N 2 O sink. These findings broaden the physiological and genomic diversity of Nitrobacter, offering new insights into their adaptation strategies and providing a framework for future evaluation of their potential roles in nitrogen loss.

Nitrobacter↗

Synthetic overlapping genes stabilize genetic systems

Overlapping genes—wherein two different proteins are translated from alternative reading frames of the same DNA sequence—provide a means to stabilize an engineered gene by directly linking its evolutionary fate with that of an overlapping gene. However, creating overlapping gene pairs is challenging, as it requires redesigning both protein products to accommodate overlap constraints. Here, we present a new “overlapping, alternate-frame insertion” (OAFI) method for creating synthetic overlapping genes by inserting an “inner” gene, encoded in an alternate frame, into a flexible region of an “outer” gene. Using OAFI, we create new overlapping gene pairs of genetic reporters and bacterial toxins within an antibiotic resistance gene. We show that both the inner and outer genes retain function despite redesign, with translation of the inner gene influenced by its overlap position in the outer gene. Importantly, we show that, despite these inner gene sequences not contributing to outer gene function, selection for the outer gene alters the permitted inactivating mutations in the inner gene, and that overlapping toxins can restrict horizontal gene transfer of the antibiotic resistance gene. Overall, OAFI offers a versatile tool for synthetic biology, expanding the applications of overlapping genes in gene stabilization and biocontainment.

Biological and medical sciences↗

Geochemistry and Multiomics Data Differentiate Streams in Pennsylvania Based on Unconventional Oil and Gas Activity

Unconventional oil and gas (UOG) extraction is increasing exponentially around the world, as new technological advances have provided cost-effective methods to extract hard-to-reach hydrocarbons. While UOG has increased the energy output of some countries, past research indicates potential impacts in nearby stream ecosystems as measured by geochemical and microbial markers. Here, we utilized a robust data set that combines 16S rRNA gene amplicon sequencing (DNA), metatranscriptomics (RNA), geochemistry, and trace element analyses to establish the impact of UOG activity in 21 sites in northern Pennsylvania. These data were also used to design predictive machine learning models to determine the UOG impact on streams. We identified multiple biomarkers of UOG activity and contributors of antimicrobial resistance within the order Burkholderiales. Furthermore, we identified expressed antimicrobial resistance genes, land coverage, geochemistry, and specific microbes as strong predictors of UOG status. Of the predictive models constructed (n = 30), 15 had accuracies higher than expected by chance and area under the curve values above 0.70. The supervised random forest models with the highest accuracy were constructed with 16S rRNA gene profiles, metatranscriptomics active microbial composition, metatranscriptomics active antimicrobial resistance genes, land coverage, and geochemistry (n = 23). The models identified the most important features within those data sets for classifying UOG status. These findings identified specific shifts in gene presence and expression, as well as geochemical measures, that can be used to build robust models to identify impacts of UOG development.

16S rRNA↗

Auto_PDI

Protein-DNA Interaction Workflow (PDI Workflow), a pipeline that focuses on generating high-quality docking and molecular dynamics simulations for Protein-DNA complexes. This allows us to take DNA sequences with unknown tertiary structures, accurately predict their structure, dock them with the desired target protein, and then simulate their interactions using molecular dynamics simulations

Kumar, Neeraj↗