Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “DNA sequencing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Deeplasmid: deep learning accurately separates plasmids from bacterial chromosomes

Plasmids are mobile genetic elements that play a key role in microbial ecology and evolution by mediating horizontal transfer of important genes, such as antimicrobial resistance genes. Many microbial genomes have been sequenced by short read sequencers and have resulted in a mix of contigs that derive from plasmids or chromosomes. New tools that accurately identify plasmids are needed to elucidate new plasmid-borne genes of high biological importance. We have developed Deeplasmid, a deep learning tool for distinguishing plasmids from bacterial chromosomes based on the DNA sequence and its encoded biological data. It requires as input only assembled sequences generated by any sequencing platform and assembly algorithm and its runtime scales linearly with the number of assembled sequences. Deeplasmid achieves an AUC–ROC of over 89%, and it was more accurate than five other plasmid classification methods. Finally, as a proof of concept, we used Deeplasmid to predict new plasmids in the fish pathogen Yersinia ruckeri ATCC 29473 that has no annotated plasmids. Deeplasmid predicted with high reliability that a long assembled contig is part of a plasmid. Using long read sequencing we indeed validated the existence of a 102 kb long plasmid, demonstrating Deeplasmid's ability to detect novel plasmids.

59 BASIC BIOLOGICAL SCIENCES↗

Five key aspects of metaproteomics as a tool to understand functional interactions in host-associated microbiomes

Host-associated microbial communities (microbiomes) play critical roles in human, animal, and plant health and development. However, interactions between the host, members of the microbiome, and invading pathogens are in most cases still poorly understood. Such interactions are multidimensional and can alter the taxonomic composition and/or the functional metabolic activities of the microbiome in response to disease or treatment conditions. For example, after 2 days of antibiotic treatment, the mouse gut microbiome is altered and more susceptible to invasion by the pathogen Clostridioides difficile. Studies of these multidimensional interactions have been fueled by the ability to use high-throughput sequencing of phylogenetic marker genes to profile microbial community composition and shotgun metagenomics to profile functional potential. However, many protein-coding genes predicted from metagenomes are not necessarily expressed under a given condition, and thus, it is difficult to assess the activities and functional interactions in microbial communities based on DNA sequencing data alone. The physiological and pathological processes expressed in these communities under specific conditions are better reflected by the abundances of transcripts or proteins. In this Pearl, we provide a brief introduction to metaproteomics, which is a tool for the large-scale analysis of proteins in microbiomes that allows researchers to address a diversity of questions related to functions and interactions in microbiomes. The term “metaproteomics” was first used in 2004 for “the large-scale characterization of the entire protein complement of environmental microbiota at a given point in time”, and since then, a large array of metaproteomics approaches have been developed. Our objective in this Pearl is to highlight what we feel are 5 essential elements to be considered for a metaproteomics research campaign and to introduce nonexpert readers to the topic without going into too much technical detail.

59 BASIC BIOLOGICAL SCIENCES↗

Structural basis of DNA recognition of the Campylobacter jejuni CosR regulator

Campylobacter jejuni is a foodborne pathogen commonly found in the intestinal tracts of animals. This pathogen is a leading cause of gastroenteritis in humans. Besides its highly infectious nature, C. jejuni is increasingly resistant to a number of clinically administrated antibiotics. As a consequence, the Centers for Disease Control and Prevention has designated antibiotic-resistant Campylobacter as a serious antibiotic resistance threat in the United States. The C. jejuni CosR regulator is essential to the viability of this bacterium and is responsible for regulating the expression of a number of oxidative stress defense enzymes. Importantly, it also modulates the expression of the CmeABC multidrug efflux system, the most predominant and clinically important system in C. jejuni that mediates resistance to multiple antimicrobials. Here, we report structures of apo-CosR and CosR bound with a 21 bp DNA sequence located at the cmeABC promotor region using both single-particle cryo-electron microscopy and X-ray crystallography. These structures allow us to propose a novel mechanism for CosR regulation that involves a long-distance conformational coupling and rearrangement of the secondary structural elements of the regulator to bind target DNA.

CosR-DNA complex↗

Warmer incubation temperatures and later lay–orders lead to shorter telomere lengths in wood duck ( Aix sponsa ) ducklings

The environment that animals experience during development shapes phenotypic expression. In birds, two important aspects of the early-developmental environment are lay-order sequence and incubation. Later-laid eggs tend to produce weaker offspring, sometimes with compensatory mechanisms to accelerate their growth rate to catch-up to their siblings. Further, small decreases in incubation temperature slow down embryonic growth rates and lead to wide-ranging negative effects on many post-hatch traits. Recently, telomeres, non-coding DNA sequences at the end of chromosomes, have been recognized as a potential proxy for fitness because longer telomeres are positively related to lifespan and individual quality in many animals, including birds. Although telomeres appear to be mechanistically linked to growth rate, little is known about how incubation temperature and lay-order may influence telomere length. We incubated wood duck (Aix sponsa) eggs at two ecologically-relevant temperatures (34.9 and 36.2ºC) and measured telomere length at hatch and one week after. We found that ducklings incubated at the lower temperature had longer telomeres than those incubated at the higher temperature both at hatch and one week later. Further, we found that later-laid eggs produced ducklings with shorter telomeres than those laid early in the lay-sequence, although lay-order was not related to embryonic developmental rate. Furthermore, this study contributes to our broader understanding of how parental effects can affect telomere length early in life. More work is needed to determine if these effects on telomere length persist until adulthood, and if they are associated with effects on fitness in this precocial species.

59 BASIC BIOLOGICAL SCIENCES↗

Strong parallel evidence of selection during switchgrass sward establishment in hybrid and lowland ecotypes

Switchgrass sward establishment results in up to 90% seedling mortality. The degree of selection during sward establishment has not been reported using modern genetic methods. Pooled leaf samples were sequenced from replicated swards of 46 half-sib families from two breeding groups (lowland and hybrid) before and through 3 years of stand establishment. Pooled allele frequencies were then assessed using fixation indices (Fst) and an independent data set was used to predict the polygenic impact of establishment selection on two traits (heading date and winter survivorship). Last, the DNA pools were assigned survival rankings to predict the sward survival genomically estimated breeding values within the training data set. Strong and parallel selection occured in both breeding groups. Five genomic regions exceeded the significant threshold of 99.9% in >10 families, indicating consistent selection across families and breeding groups. Polygenic trait predictions determined that establishment selection was partially associated with winter survivorship but resulted in variable heading date alterations. The genomewide variation is consistent with selection for a small number of related parental lines. This study observed strong selection for a small number of hybrid and coastal ecotype individuals which are promising germplasm sources for improved sward survival. This confirms prior reports of sward selection during grassland establishment and highlights the strength of pooled DNA sequencing for survival traits.

54 ENVIRONMENTAL SCIENCES↗

Precision genome editing in plants using gene targeting and prime editing: existing and emerging strategies

Precise modification of plant genomes, such as seamless insertion, deletion, or replacement of DNA sequences at a predefined site, is a challenging task. Gene targeting (GT) and prime editing are currently the best approaches for this purpose. However, these techniques are inefficient in plants, which limits their applications for crop breeding programs. Recently, substantial developments have been made to improve the efficiency of these techniques in plants. Several strategies, such as RNA donor templating, chemically modified donor DNA template, and tandem-repeat homology-directed repair, are aimed at improving GT. Additionally, improved prime editing gRNA design, use of engineered reverse transcriptase enzymes, and splitting prime editing components have improved the efficacy of prime editing in plants. These emerging strategies and existing technologies are reviewed along with various perspectives on their future improvement and the development of robust precision genome editing technologies for plants.

59 BASIC BIOLOGICAL SCIENCES↗

Predictive links between microbial communities and biological oxygen utilization in the Arctic Ocean

Microbial metabolism influences rates of net community production (NCP), exerting a direct biological control on marine oxygen and carbon fluxes. In the Arctic, it is increasingly important to understand and quantify this process, as ecological and oceanographic conditions shift due to changing climate. Here, we describe potential ecological links between pelagic microbial diversity and an NCP precursor, biological oxygen utilization, using machine learning and paired observations of community structure and metabolic activity from a seasonally and spatially variable transect of the Arctic Ocean (2019–2020 MOSAiC Expedition). Community structure was determined using 16S (prokaryotic) and 18S (eukaryotic) rRNA gene amplicon sequencing, and metabolic activity was derived from ΔO 2 /Ar. Using self-organizing maps, we identified clear successional patterns in observed microbial community structure that were seasonally driven in the upper ocean and vertically stratified with depth. Metabolic activity was also stratified, with a primarily net heterotrophic water column (median −1.5% biological oxygen saturation), excepting periodic oxygen supersaturation (maximum: 13.6%) within the mixed layer. Using DNA sequences as predictor variables, we then constructed a random forest regression model that reliably reconstructed biological oxygen concentrations (root mean squared error = 4.14 μmol kg −1 ). Top predictors from this model were from heterotrophic (bacteria) or potentially mixotrophic (dinoflagellate) taxa. These analyses highlight biologically driven diagnostic tools that can be used to expand biogeochemical datasets and improve the microbial perspectives and metabolisms represented in ecological models of net productivity and carbon flux in a changing Arctic Ocean.

Chamberlain, Emelia J. [Univ. of San Diego, San Di↗

Nucleic Acid-Based Detection Protease Activity

Proteases include clinically relevant markers for clotting disorders, certain cancers as well as toxins. Assays for protease activity often use designed peptides mimicking natural substrates and detection with colorometric and fluorescence-based detection that is difficult to multiplex without expensive and resource demanding instruments. This work demonstrates detection of proteolytic activity using PCR and sequencing-readable reporter molecules. The assay development focused on binding the constructed peptide-oligonucleotide chimera to immobilized streptavidin. Thrombin, an essential component of the clotting cascade, was used as a model system for testing peptide substrate recognition and release of a designed oligonucleotide for detection. Detection of protease activity was demonstrated in a concentration-dependent manner using MALDI-MS, RT-PCR and DNA sequencing.

Wunschel, David S [Pacific Northwest National Labo↗

Manipulating the 3D organization of the largest synthetic yeast chromosome

Whether synthetic genomes can power life has attracted broad interest in the synthetic biology field. Here, we report de novo synthesis of the largest eukaryotic chromosome thus far, synIV, a 1,454,621-bp yeast chromosome resulting from extensive genome streamlining and modification. We developed megachunk assembly combined with a hierarchical integration strategy, which significantly increased the accuracy and flexibility of synthetic chromosome construction. Besides the drastic sequence changes, we further manipulated the 3D structure of synIV to explore spatial gene regulation. Surprisingly, we found few gene expression changes, suggesting that positioning inside the yeast nucleoplasm plays a minor role in gene regulation. Lastly, we tethered synIV to the inner nuclear membrane via its hundreds of loxPsym sites and observed transcriptional repression of the entire chromosome, demonstrating chromosome-wide transcription manipulation without changing the DNA sequences. Our manipulation of the spatial structure of synIV sheds light on higher-order architectural design of the synthetic genomes.

59 BASIC BIOLOGICAL SCIENCES↗

Cas3-Mediated Genome Reduction: Demonstration in Cupriavidus Necator H16 Improves Growth on Heterotrophic and Autotrophic Carbon Sources

Genome reduction is widely used to improve microbial bioprocessing hosts by reducing the burden of inessential physiology. Rationally identifying genomic regions that are dispensable or even detrimental to bioprocessing is challenged by our inability to map genome sequence to function across complex regulation and physiology. Thus, there is a need for tools that rapidly generate reduced genome strains with improved performance in process-relevant conditions. Here, we report a Cascade-Cas3-enabled method called TRIM3 that generates large deletions by targeting a randomly integrated transposon, enabling facile generation of a genome-reduced mutant library. Mutants with improved performance were isolated following growth-coupled selection and analyzed by long-read DNA sequencing to identify deletions in their genomes. We deploy this system iteratively in the industrial host Cupriavidus necator H16 on fructose and on formate. After two rounds of TRIM3, we isolate a strain containing a total reduction of 1.4 Mb (18.4% of the genome) that grows 25% faster in a bioreactor on fructose and a strain with a total reduction of 0.5 Mb (7.3% of the genome) that grows 14% faster on formate. This work demonstrates a method for random, iterative, growth-selectable genome reduction that represents a new avenue for large-scale genome modifications and the development of improved bioprocessing hosts.

09 BIOMASS FUELS↗

The complex polyploid genome architecture of sugarcane

Sugarcane, the world’s most harvested crop by tonnage, has shaped global history, trade and geopolitics, and is currently responsible for 80% of sugar production worldwide. While traditional sugarcane breeding methods have effectively generated cultivars adapted to new environments and pathogens, sugar yield improvements have recently plateaued. The cessation of yield gains may be due to limited genetic diversity within breeding populations, long breeding cycles and the complexity of its genome, the latter preventing breeders from taking advantage of the recent explosion of whole-genome sequencing that has benefited many other crops. Thus, modern sugarcane hybrids are the last remaining major crop without a reference-quality genome. Here we take a major step towards advancing sugarcane biotechnology by generating a polyploid reference genome for R570, a typical modern cultivar derived from interspecific hybridization between the domesticated species (Saccharum officinarum) and the wild species (Saccharum spontaneum). In contrast to the existing single haplotype (‘monoploid’) representation of R570, our 8.7 billion base assembly contains a complete representation of unique DNA sequences across the approximately 12 chromosome copies in this polyploid genome. Using this highly contiguous genome assembly, we filled a previously unsized gap within an R570 physical genetic map to describe the likely causal genes underlying the single-copy Bru1 brown rust resistance locus. This polyploid genome assembly with fine-grain descriptions of genome architecture and molecular targets for biotechnology will help accelerate molecular and transgenic breeding and adaptation of sugarcane to future environmental conditions.

59 BASIC BIOLOGICAL SCIENCES↗

Targeted seed EMS mutagenesis reveals a basic helix–loop–helix transcription factor underlying male sterility in sorghum

Abstract Forward genetic screens of mutant populations are fundamental for functional genomics studies. However, isolating independent mutant alleles to molecularly identify causal genes is challenging in species recalcitrant to genetic manipulation. Here, we demonstrate that classic seed ethyl methanesulfonate (EMS) mutagenesis coupled with genome sequencing can overcome this limitation in sorghum. We used this method to generate new mutant alleles of sorghum MALE STERILE 8 (MS8) and identified the causal locus for the ms8 phenotype as Sobic.004G270900, which encodes the sorghum ortholog of maize bhlh122, a basic helix–loop–helix (bHLH) transcription factor required for male fertility in maize. Bulked segregant analysis mapped ms8-1 to a region on chromosome 4 containing Sobic.004G270900. Seeds from heterozygous MS8/ms8-1 plants were mutagenized and screened for chimeric inflorescences containing sectors with white, sterile anthers resembling the ms8-1 homozygous phenotype. DNA sequencing of sterile and fertile sectors from a single chimeric inflorescence revealed two mutations in Sobic.004G270900 within the sterile sector, but not the fertile sector. Isolation of this loss-of-function allele (ms8-2) established Sobic.004G270900 as the causative locus for male sterility in the ms8 mutant. We generated additional alleles of MS8 in a different genetic background using CRISPR/Cas9-based gene editing, where deletions in Sobic.004G270900 also resulted in male sterility. Our work identified a gene underlying male sterility in sorghum and provides a novel and straightforward genetic tool for researchers who lack access to advanced transformation facilities to validate gene candidates. Unlike gene editing, no prior knowledge of candidate genes is required for targeted seed EMS mutagenesis to aid identification of causal loci.

Genetics & Heredity↗

The role of chromatin state in intron retention: A case study in leveraging large scale deep learning models

Complex deep learning models trained on very large datasets have become key enabling tools for current research in natural language processing and computer vision. By providing pre-trained models that can be fine-tuned for specific applications, they enable researchers to create accurate models with minimal effort and computational resources. Large scale genomics deep learning models come in two flavors: the first are large language models of DNA sequences trained in a self-supervised fashion, similar to the corresponding natural language models; the second are supervised learning models that leverage large scale genomics datasets from ENCODE and other sources. We argue that these models are the equivalent of foundation models in natural language processing in their utility, as they encode within them chromatin state in its different aspects, providing useful representations that allow quick deployment of accurate models of gene regulation. We demonstrate this premise by leveraging the recently created Sei model to develop simple, interpretable models of intron retention, and demonstrate their advantage over models based on the DNA language model DNABERT-2. Our work also demonstrates the impact of chromatin state on the regulation of intron retention. Using representations learned by Sei, our model is able to discover the involvement of transcription factors and chromatin marks in regulating intron retention, providing better accuracy than a recently published custom model developed for this purpose.

Biochemistry & Molecular Biology↗

Structure of a 10-23 deoxyribozyme exhibiting a homodimer conformation

Deoxyribozymes (DNAzymes) are in vitro evolved DNA sequences capable of catalyzing chemical reactions. The RNA-cleaving 10-23 DNAzyme was the first DNAzyme to be evolved and possesses clinical and biotechnical applications as a biosensor and a knockdown agent. DNAzymes do not require the recruitment of other components to cleave RNA and can turnover, thus they have a distinct advantage over other knockdown methods (siRNA, CRISPR, morpholinos). Despite this, a lack of structural and mechanistic information has hindered the optimization and application of the 10-23 DNAzyme. Here, we report a 2.7 Å crystal structure of the RNA-cleaving 10-23 DNAzyme in a homodimer conformation. Although proper coordination of the DNAzyme to substrate is observed along with intriguing patterns of bound magnesium ions, the dimer conformation likely does not capture the true catalytic form of the 10-23 DNAzyme.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Salt-Induced Polymorphs Observed in Colloidal Single Crystals

Polymorphs are solid materials with the same chemical composition but different crystallographic structures. A unique aspect of polymorphs is that they exhibit different physical properties, such as solubility, melting point, density, color, hardness, and bioavailability. Here, we synthesized polymorphs of colloidal crystals engineered with DNA by slow-cooling gold nanoparticle-core programmable atom equivalents (PAEs, particles with DNA sequences that control their bonding characteristics) under salt concentrations ranging from 0.5 to 4 M NaCl. This approach yielded a diverse set of single-crystalline phases with cubic, tetragonal, and hexagonal lattice symmetries. The structural transitions observed here arise solely from the modulation of interparticle repulsion via ionic strength and thermal processing. Notably, in certain cases, we observed diffusionless phase transformations, wherein the superlattices evolve from cubic to lower-symmetry tetragonal lattices. By tuning the thermal stability and salt concentration, we captured intermediate, metastable body-centered tetragonal structures during the slow-cool process, indicating that subtle changes in free energy can direct crystallization to low-symmetry phases. In conclusion, this study demonstrates that thermal and ionic parameters can be tuned to access and stabilize colloidal crystal polymorphs with emergent structures and interesting functional properties.

Colloidal crystallization↗

Predicting ecosystem metaphenome from community metagenome: A grand challenge for environmental biology

Abstract Elucidating how an organism's characteristics emerge from its DNA sequence has been one of the great triumphs of biology. This triumph has cumulated in sophisticated computational models that successfully predict how an organism's detailed phenotype emerges from its specific genotype. Inspired by that effort's vision and empowered by its methodologies, a grand challenge is described here that aims to predict the biotic characteristics of an ecosystem, its metaphenome, from nucleic acid sequences of all the species in its community, its metagenome. Meeting this challenge would integrate rapidly advancing abilities of environmental nucleic acids (eDNA and eRNA) to identify organisms, their ecological interactions, and their evolutionary relationships with advances in mechanistic models of complex ecosystems. Addressing the challenge would help integrate ecology and evolutionary biology into a more unified and successfully predictive science that can better help describe and manage ecosystems and the services they provide to humanity.

59 BASIC BIOLOGICAL SCIENCES↗

Use of Fluorescent Protein Reporters for Assessing and Detecting Genome Editing Reagents and Transgene Expression in Plants

Fluorescent protein reporters have been widely used for monitoring the expression of target genes in various engineered organisms. Although a wide range of analytical approaches (e.g., genotyping PCR, digital PCR, DNA sequencing) have been utilized to detect and identify genome editing reagents and transgene expression in genetically modified plants, these methods are usually limited to use in the late stages of plant transformation and can only be used invasively. Here we describe GFP- and eYGFPuv-based strategies and methods for assessing and detecting genome editing reagents and transgene expression in plants, including protoplast transformation, leaf infiltration, and stable transformation. These methods and strategies enable easy, noninvasive screening of genome editing and transgenic events in plants.

Yuan, Guoliang↗

A Comparative Metagenomic Analysis of Specified Microorganisms in Groundwater for Non-Sterilized Pharmaceutical Products

In pharmaceutical manufacturing, ensuring product safety involves the detection and identification of microorganisms with human pathogenic potential, including Burkholderia cepacia complex (BCC), Escherichia coli, Pseudomonas aeruginosa, Salmonella enterica, Staphylococcus aureus, Clostridium sporogenes, Candida albicans, and Mycoplasma spp., some of which may be missed or not identified by traditional culture-dependent methods. In this study, we employed a metagenomic approach to detect these taxa, avoiding the limitations of conventional cultivation methods. We assessed the groundwater microbiome’s taxonomic and functional features from samples collected at two locations in the spring and summer. All datasets comprised 436–557 genera with Proteobacteria, Bacteroidota, Firmicutes, Actinobacteria, and Cyanobacteria accounting for > 95% of microbial DNA sequences. The aforementioned species constituted less than 18.3% of relative abundance. Escherichia and Salmonella were mainly detected in Hot Springs, relative to Jefferson, while Clostridium and Pseudomonas were mainly found in Jefferson relative to Hot Springs. Multidrug resistance efflux pumps and BlaR1 family regulatory sensor-transducer disambiguation dominated in Hot Springs and in Jefferson. These initial results provide insight into the detection of specified microorganisms and could constitute a framework for the establishment of comprehensive metagenomic analysis for the microbiological evaluation of pharmaceutical-grade water and other non-sterile pharmaceutical products, ensuring public safety.

59 BASIC BIOLOGICAL SCIENCES↗