Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “DNA sequencing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Inhibition of DNA methylation in P. soloecismus alters algae productivity

Eukaryotic organisms regulate the organization, structure, and accessibility of their genomes through chromatin remodeling that can be inherited as epigenetic modifications. These DNA and histone protein modifications are ultimately responsible for an organism’s molecular adaptation to the environment, resulting in distinctive phenotypes. Epigenetic manipulation of algae holds yet untapped potential for the optimization of biofuel production and bioproduct formation; however, epigenetic machinery and modes-of-action have not been well characterized in algae. We sought to determine the extent to which the biofuel platform species Picochlorum soloecismus utilizes DNA methylation to regulate its genome. We found candidate genes with domains for DNA methylation in the P. soloecismus genome. Whole-genome bisulfite sequencing revealed DNA methylation in all three cytosine contexts (CpG, CHH, and CHG). While global DNA methylation is low overall (~1.15%), it occurs in appreciable quantities (12.1%) in CpG dinucleotides in a bimodal distribution in all genomic contexts, though terminators contain the greatest number of CpG sites per kilobase. The P. soloecismus genome becomes hypomethylated during the growth cycle in response to nitrogen starvation. Algae cultures were treated daily across the growth cycle with 20 μM 5-aza-2'-deoxycytidine (5AZA) to inhibit propagation of DNA methylation in daughter cells. 5AZA treatment significantly increased optical density and forward and side scatter of cells across the growth cycle (16 days). This increase in cell size and complexity correlated with a significant increase (~66%) in lipid accumulation. Site specific CpG DNA methylation was significantly altered with 5AZA treatment over the time course, though nitrogen starvation itself induced significant hypomethylation in CpG contexts. Genes involved in several biological processes, including fatty acid synthesis, had altered methylation ratios in response to 5AZA; we hypothesize that these changes are potentially responsible for the phenotype of early induction of carbon storage as lipids. This is the first report to utilize epigenetic manipulation strategies to alter algal physiology and phenotype. Collectively, these data suggest these strategies can be utilized to fine-tune metabolic responses, alter growth, and enhance environmental adaption of microalgae for desired outcomes.

5-aza-2'-deoxycytidine↗

DIVA/DeviceEditor v6.1.2

DIVA is an end-to-end DNA design and construction management platform that streamlines how researchers design, build, and receive sequence-verified DNA constructs. Through a web-based BioCAD interface (DeviceEditor), researchers independently design DNA constructs and submit them to a centralized queue with a single action. Designs progress transparently through standardized states which allow researchers to track status and access finished constructs via a central DNA repository. Submitted designs are reviewed by dedicated staff for feasibility and optimization, reducing costly failures and improving downstream execution. Automated DNA assembly software optimizes construction strategies by reusing existing parts where possible and sourcing synthetic DNA only when needed. Standardized, sequence-agnostic assembly methods enable many independent constructs to be built in parallel using lab automation, dramatically increasing throughput. High-throughput next-generation sequencing is used to verify construct accuracy, with flexible platforms selected based on task requirements. Throughout the process, detailed success and failure data are captured and analyzed, enabling continuous improvement of assembly protocols. Compared to traditional, manual DNA construction workflows, DIVA offers higher scalability, transparency, reproducibility, and data-driven optimization.

Plahar, Hector [Lawrence Berkeley National Laborat↗

The ChiS-Family DNA-Binding Domain Contains a Cryptic Helix-Turn-Helix Variant

Sequence-specific DNA-binding domains (DBDs) are conserved in all domains of life. These proteins carry out a variety of cellular functions, and there are a number of distinct structural domains already described that allow for sequence-specific DNA binding, including the ubiquitous helix-turn-helix (HTH) domain. In the facultative pathogen Vibrio cholerae, the chitin sensor ChiS is a transcriptional regulator that is critical for the survival of this organism in its marine reservoir. We recently showed that ChiS contains a cryptic DBD in its C terminus. This domain is not homologous to any known DBD, but it is a conserved domain present in other bacterial proteins. Here, we present the crystal structure of the ChiS DBD at a resolution of 1.28 Å. We find that the ChiS DBD contains an HTH domain that is structurally similar to those found in other DNA-binding proteins, like the LacI repressor. However, one striking difference observed in the ChiS DBD is that the canonical tight turn of the HTH is replaced with an insertion containing a β-sheet, a variant which we term the helix-sheet-helix. Through systematic mutagenesis of all positively charged residues within the ChiS DBD, we show that residues within and proximal to the ChiS helix-sheet-helix are critical for DNA binding. Finally, through phylogenetic analyses we show that the ChiS DBD is found in diverse proteobacterial proteins that exhibit distinct domain architectures. Together, these results suggest that the structure described here represents the prototypical member of the ChiS-family of DBDs.

59 BASIC BIOLOGICAL SCIENCES↗

Distributed Berkeley Efficient Long-Read to Long-Read Aligner and Overlapper (DiBELLA) v1.0.0

We present a parallel algorithm and scalable implementation for genome analysis, specifically the problem of finding overlaps and alignments for data from "third generation" long read sequencers. While long sequences of DNA offer enormous advantages for biological analysis and insight, current long read sequencing instruments have high error rates and therefore require different approaches to analysis than their short read counterparts. Our work focuses on an efficient distributed-memory parallelization of an accurate single-node algorithm for overlapping and aligning long reads. We achieve scalability of this irregular algorithm by addressing the competing issues of increasing parallelism, minimizing communication, constraining the memory footprint, and ensuring good load balance. The resulting application, DiBELLA, is the first distributed memory overlapper and aligner specifically designed for long reads and parallel scalability.

Ellis, Marquita↗

Genome-wide Transcription Factor DNA Binding Sites and Gene Regulatory Networks in Clostridium thermocellum

Clostridium thermocellum is a thermophilic bacterium recognized for its natural ability to effectively deconstruct cellulosic biomass. While there is a large body of studies on the genetic engineering of this bacterium and its physiology to-date, there is limited knowledge in the transcriptional regulation in this organism and thermophilic bacteria in general. The study herein is the first report of a high-throughput application of DNA-affinity purification sequencing (DAP-seq) to transcription factors (TFs) from a thermophile. We applied DAP-seq to >90 TFs in C. thermocellum and detected genome-wide binding sites for 11 of them. We then compiled and aligned DNA binding sequences from these TFs to deduce the primary DNA-binding sequence motifs for each TF. These binding motifs are further validated with electrophoretic mobility shift assay (EMSA) and are used to identify individual TFs’ regulatory targets in C. thermocellum. Our results led to the discovery of novel, uncharacterized TFs as well as homologues of previously studied TFs including RexA-, LexA- and LacI-type TFs. We then used these data to reconstruct gene regulatory networks for the 11 TFs individually, which resulted in a global network encompassing the TFs with some interconnections. As gene regulation governs and constrains how bacteria behave, our findings shed light on the roles of TFs delineated by their regulons, and potentially provides a means to enable rational, advanced genetic engineering of C. thermocellum and other organisms alike towards a desired phenotype.

09 BIOMASS FUELS↗

Genome-Wide Transcription Factor DNA Binding Sites and Gene Regulatory Networks in Clostridium thermocellum

Clostridium thermocellum is a thermophilic bacterium recognized for its natural ability to effectively deconstruct cellulosic biomass. While there is a large body of studies on the genetic engineering of this bacterium and its physiology to-date, there is limited knowledge in the transcriptional regulation in this organism and thermophilic bacteria in general. The study herein is the first report of a large-scale application of DNA-affinity purification sequencing (DAP-seq) to transcription factors (TFs) from a bacterium. We applied DAP-seq to > 90 TFs in C. thermocellum and detected genome-wide binding sites for 11 of them. We then compiled and aligned DNA binding sequences from these TFs to deduce the primary DNA-binding sequence motifs for each TF. These binding motifs are further validated with electrophoretic mobility shift assay (EMSA) and are used to identify individual TFs’ regulatory targets in C. thermocellum . Our results led to the discovery of novel, uncharacterized TFs as well as homologues of previously studied TFs including RexA-, LexA-, and LacI-type TFs. We then used these data to reconstruct gene regulatory networks for the 11 TFs individually, which resulted in a global network encompassing the TFs with some interconnections. As gene regulation governs and constrains how bacteria behave, our findings shed light on the roles of TFs delineated by their regulons, and potentially provides a means to enable rational, advanced genetic engineering of C. thermocellum and other organisms alike toward a desired phenotype.

59 BASIC BIOLOGICAL SCIENCES↗

Metagenomic clustering links specific metabolic functions to globally relevant ecosystems

ABSTRACT Metagenomic sequencing has advanced our understanding of biogeochemical processes by providing an unprecedented view into the microbial composition of different ecosystems. While the amount of metagenomic data has grown rapidly, simple-to-use methods to analyze and compare across studies have lagged behind. Thus, tools expressing the metabolic traits of a community are needed to broaden the utility of existing data. Gene abundance profiles are a relatively low-dimensional embedding of a metagenome’s functional potential and are, thus, tractable for comparison across many samples. Here, we compare the abundance of KEGG Ortholog Groups (KOs) from 6,539 metagenomes from the Joint Genome Institute’s Integrated Microbial Genomes and Metagenomes (JGI IMG/M) database. We find that samples cluster into terrestrial, aquatic, and anaerobic ecosystems with marker KOs reflecting adaptations to these environments. For instance, functional clusters were differentiated by the metabolism of antibiotics, photosynthesis, methanogenesis, and surprisingly GC content. Using this functional gene approach, we reveal the broad-scale patterns shaping microbial communities and demonstrate the utility of ortholog abundance profiles for representing a rapidly expanding body of metagenomic data. IMPORTANCE Metagenomics, or the sequencing of DNA from complex microbiomes, provides a view into the microbial composition of different environments. Metagenome databases were created to compile sequencing data across studies, but it remains challenging to compare and gain insight from these large data sets. Consequently, there is a need to develop accessible approaches to extract knowledge across metagenomes. The abundance of different orthologs (i.e., genes that perform a similar function across species) provides a simplified representation of a metagenome’s metabolic potential that can easily be compared with others. In this study, we cluster the ortholog abundance profiles of thousands of metagenomes from diverse environments and uncover the traits that distinguish them. This work provides a simple to use framework for functional comparison and advances our understanding of how the environment shapes microbial communities.

54 ENVIRONMENTAL SCIENCES↗

Mechanistic modeling of in vitro transcription incorporating effects of magnesium pyrophosphate crystallization

The in vitro transcription (IVT) reaction used in the production of messenger RNA vaccines and therapies remains poorly quantitatively understood. Mechanistic modeling of IVT could inform reaction design, scale-up, and control. In this work, we develop a mechanistic model of IVT to include nucleation and growth of magnesium pyrophosphate crystals and subsequent agglomeration of crystals and DNA. To help generalize this model to different constructs, a novel quantitative description is included for the rate of transcription as a function of target sequence length, DNA concentration, and T7 RNA polymerase concentration. The model explains previously unexplained trends in IVT data and quantitatively predicts the effect of adding the pyrophosphatase enzyme to the reaction system. The model is validated on additional literature data showing an ability to predict transcription rates as a function of RNA sequence length.

59 BASIC BIOLOGICAL SCIENCES↗

Genomic Metrics Applied to Rhizobiales ( Hyphomicrobiales ): Species Reclassification, Identification of Unauthentic Genomes and False Type Strains

Taxonomic decisions within the order Rhizobiales have relied heavily on the interpretations of highly conserved 16S rRNA sequences and DNA–DNA hybridizations (DDH). Currently, bacterial species are defined as including strains that present 95–96% of average nucleotide identity (ANI) and 70% of digital DDH (dDDH). Thus, ANI values from 520 genome sequences of type strains from species of Rhizobiales order were computed. From the resulting 270,400 comparisons, a ≥95% cut-off was used to extract high identity genome clusters through enumerating maximal cliques. Coupling this graph-based approach with dDDH from clusters of interest, it was found that: (i) there are synonymy between Aminobacter lissarensis and Aminobacter carboxidus, Aurantimonas manganoxydans and Aurantimonas coralicida, “Bartonella mastomydis,” and Bartonella elizabethae, Chelativorans oligotrophicus, and Chelativorans multitrophicus, Rhizobium azibense, and Rhizobium gallicum, Rhizobium fabae, and Rhizobium pisi, and Rhodoplanes piscinae and Rhodoplanes serenus; (ii) Chelatobacter heintzii is not a synonym of Aminobacter aminovorans; (iii) “Bartonella vinsonii” subsp. arupensis and “B. vinsonii” subsp. berkhoffii represent members of different species; (iv) the genome accessions GCF_003024615.1 (“Mesorhizobium loti LMG 6125T”), GCF_003024595.1 (“Mesorhizobium plurifarium LMG 11892T”), GCF_003096615.1 (“Methylobacterium organophilum DSM 760T”), and GCF_000373025.1 (“R. gallicum R-602 spT”) are not from the genuine type strains used for the respective species descriptions; and v) “Xanthobacter autotrophicus” Py2 and “Aminobacter aminovorans” KCTC 2477T represent cases of misuse of the term “type strain”. Aminobacter heintzii comb. nov. and the reclassification of Aminobacter ciceronei as A. heintzii is also proposed. To facilitate the downstream analysis of large ANI matrices, we introduce here ProKlust (“Prokaryotic Clusters”), an R package that uses a graph-based approach to obtain, filter, and visualize clusters on identity/similarity matrices, with settable cut-off points and the possibility of multiple matrices entries.

59 BASIC BIOLOGICAL SCIENCES↗

Diversity of Sordariales Fungi: Identification of Seven New Species of Naviculisporaceae Through Morphological Analyses and Genome Sequencing

Thanks to next-generation sequencing (NGS) technologies, the diversity of fungi can now be investigated through the analysis of their genome sequences. Naviculisporaceae is a family within the Sordariales, whose diversity is not well-known, with only one genome sequence published for this family. Here, we report on the isolation and cultivation of 20 new strains of Naviculisporaceae. Their genome sequences, as well as those of the five commercially available strains, were determined, thus providing complete genome sequences for 25 new Naviculisporaceae strains. Species delimitation was conducted using a combination of (1) ITS + LSU phylogenetic analysis of the new isolates along with other known species of the family, (2) comparisons between DNA barcode sequences of the new strains with those of the known species, and (3) average genome-wide nucleotide identity calculation. We built a phylogenomic tree and studied the organization of the mating-type locus. In vitro fruiting was obtained for 16 strains, enabling the definition of seven new species, namely Pseudorhypophila gallica, Pseudorhypophila guyanensis Rhypophila alpibus, Rhypophila brasiliensis, Rhypophila camarguensis, Rhypophila reunionensis and Rhypophila thailandica, as well as two new combinations, namely Pseudorhypophila latipes and Pseudorhypophila oryzae. Eight strains for which in vitro fruiting was not obtained may belong to additional new species. These results expand the known diversity of the Naviculisporaceae and greatly enlarge the genomic data available for the family.

Naviculisporaceae↗

Molecular and morphological characterization of Tylenchus zeae n. sp. (Nematoda: Tylenchida) from Corn ( Zea mays ) in South Carolina

Specimens of a tylenchid nematode were recovered in 2019 from soil samples collected from a corn field, located in Pickens County, South Carolina, USA. A moderate number of Tylenchus sp. adults (females and males) were recovered. Extracted nematodes were examined morphologically and molecularly for species identification, which indicated that the specimens of the tylenchid adults were a new species, described herein as Tylenchus zeae n. sp. Morphological examination and the morphometric details of the specimens were very close to the original descriptions of Tylenchus sherianus and T. rex. However, females of the new species can be differentiated from these species by body shape and length, shape of excretory duct, distance between anterior end and esophageal intestinal valve, and a few other characteristics given in the diagnosis. Males of the new species can be differentiated from the two closely related species by tail, spicules, and gubernaculum length. Cryo-scanning electron microscopy confirmed head bearing five or six annules; four to six cephalic sensilla represented by small pits at the rounded corners of the labial plate; a small, round oral plate; and a large, pit-like amphidial opening confined to the labial plate and extending three to four annules beyond it. Phylogenetic analysis of 18S rRNA gene sequences placed Tylenchus zeae n. sp. in a clade with Tylenchus arcuatus and several Filenchus spp., and the mitochondrial cytochrome oxidase c subunit 1 (COI) gene region separated the new species from T. arcuatus and other tylenchid species. In the 28S tree, T. zeae n. sp. showed a high level of sequence divergence and was positioned outside of the main Tylenchus-Filenchus clade.

18S↗

Vaccine Advances against Venezuelan, Eastern, and Western Equine Encephalitis Viruses

Vaccinations are a crucial intervention in combating infectious diseases. The three neurotropic Alphaviruses, Eastern (EEEV), Venezuelan (VEEV), and Western (WEEV) equine encephalitis viruses, are pathogens of interest for animal health, public health, and biological defense. In both equines and humans, these viruses can cause febrile illness that may progress to encephalitis. Currently, there are no licensed treatments or vaccines available for these viruses in humans. Experimental vaccines have shown variable efficacy and may cause severe adverse effects. Here, we outline recent strategies used to generate vaccines against EEEV, VEEV, and WEEV with an emphasis on virus-vectored and plasmid DNA delivery. Despite candidate vaccines protecting against one of the three viruses, few studies have demonstrated an effective trivalent vaccine. We evaluated the potential of published vaccines to generate cross-reactive protective responses by comparing DNA vaccine sequences to a set of EEEV, VEEV, and WEEV genomes and determining the vaccine coverages of potential epitopes. Finally, we discuss future directions in the development of vaccines to combat EEEV, VEEV, and WEEV.

59 BASIC BIOLOGICAL SCIENCES↗

Directed evolution expands CRISPR–Cas12a genome-editing capacity

CRISPR-Cas12a enzymes are versatile RNA-guided genome-editing tools with applications encompassing viral diagnosis, agriculture, and human therapeutics. However, their dependence on a 5'-TTTV-3' protospacer adjacent motif (PAM) next to DNA target sequences restricts Cas12a's gene targeting capability to only ∼1% of a typical genome. To mitigate this constraint, we used a bacterial-based directed evolution assay combined with rational engineering to identify variants of Lachnospiraceae bacterium Cas12a with expanded PAM recognition. The resulting Cas12a variants use a range of noncanonical PAMs while retaining recognition of the canonical 5'-TTTV-3' PAM. In particular, biochemical and cell-based assays show that the variant Flex-Cas12a utilizes 5'-NYHV-3' PAMs that expand DNA recognition sites to ∼25% of the human genome. With enhanced targeting versatility, Flex-Cas12a unlocks access to previously inaccessible genomic loci, providing new opportunities for both therapeutic and agricultural genome engineering.

Ma, Enbo↗

Multiomics and deep learning dissect regulatory syntax in human development

Transcription factors establish cell identity during development by binding regulatory DNA in a sequence-specific manner, often promoting local chromatin accessibility and regulating gene expression1. Mapping accessible chromatin offers critical insights into transcriptional control, but available datasets for human development are restricted to bulk tissue, single organs or single modalities2. Here we present the Human Development Multiomic Atlas, a single-cell atlas of chromatin accessibility and gene expression from 817,740 fetal cells across 12 organs, spanning 203 cell types and more than 1 million candidate cis-regulatory elements, many of which exhibit organ-specific in vivo enhancer activity. Deep learning models trained to predict accessibility from local DNA sequence unravel a comprehensive lexicon of motifs that influence accessibility, including composite motifs exhibiting distinct syntactic constraints that are predicted to mediate transcription factor cooperativity. We identify ‘hard’ syntactic rules requiring precise motif spacing and orientation, ‘soft’ rules allowing flexible motif arrangements, and ubiquitous motifs inhibiting accessibility. Model-based interpretation of genetic variants reveals that disruption of motifs with positive and negative effects is associated with concordant effects on gene expression. Our work delineates how motif syntax governs cell-type-specific chromatin accessibility and provides a foundational resource for decoding cis-regulatory logic and interpreting genetic variation during human development.

59 BASIC BIOLOGICAL SCIENCES↗