Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “DNA sequencing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

FUSARIUM-ID v.3.0: An Updated, Downloadable Resource for Fusarium Species Identification

Species within Fusarium are of global agricultural, medical, and food/feed safety concern and have been extensively characterized. However, accurate identification of species is challenging and usually requires DNA sequence data. FUSARIUM-ID ( http://isolate.fusariumdb.org/blast.php ) is a publicly available database designed to support the identification of Fusarium species using sequences of multiple phylogenetically informative loci, especially the highly informative ~680-bp 5' portion of the translation elongation factor 1-alpha (TEF1) gene that has been adopted as the primary barcoding locus in the genus. However, FUSARIUM-ID v.1.0 and 2.0 had several limitations, including inconsistent metadata annotation for the archived sequences and poor representation of some species complexes and marker loci. Here, we present FUSARIUM-ID v.3.0, which provides the following improvements: (i) additional and updated annotation of metadata for isolates associated with each sequence, (ii) expanded taxon representation in the TEF1 sequence database, (iii) availability of the sequence database as a downloadable file to enable local BLAST queries, and (iv) a tutorial file for users to perform local BLAST searches using either freely available software, such as SequenceServer, BLAST+ executable in the command line, and Galaxy, or the proprietary Geneious software. FUSARIUM-ID will be updated on a regular basis by archiving sequences of TEF1 and other loci from newly identified species and greater in-depth sampling of currently recognized species.

Plant Sciences↗

Escaping the fate of Sisyphus: assessing resistome hybridization baits for antimicrobial resistance gene capture

Finding, characterizing and monitoring reservoirs for antimicrobial resistance (AMR) is vital to protecting public health. Hybridization capture baits are an accurate, sensitive and cost-effective technique used to enrich and characterize DNA sequences of interest, including antimicrobial resistance genes (ARGs), in complex environmental samples. We demonstrate the continued utility of a set of 19 933 hybridization capture baits designed from the Comprehensive Antibiotic Resistance Database (CARD)v1.1.2 and Pathogenicity Island Database (PAIDB)v2.0, targeting 3565 unique nucleotide sequences that confer resistance. We demonstrate the efficiency of our bait set on a custom-made resistance mock community and complex environmental samples to increase the proportion of on-target reads as much as >200-fold. However, keeping pace with newly discovered ARGs poses a challenge when studying AMR, because novel ARGs are continually being identified and would not be included in bait sets designed prior to discovery. Here we provide imperative information on how our bait set performs against CARDv3.3.1, as well as a generalizable approach for deciding when and how to update hybridization capture bait sets. This research encapsulates the full life cycle of baits for hybridization capture of the resistome from design and validation (both in silico and in vitro) to utilization and forecasting updates and retirement.

59 BASIC BIOLOGICAL SCIENCES↗

An integrative framework reveals widespread gene flow during the early radiation of oaks and relatives in Quercoideae (Fagaceae)

ABSTRACT Although the frequency of ancient hybridization across the Tree of Life is greater than previously thought, little work has been devoted to uncovering the extent, timeline, and geographic and ecological context of ancient hybridization. Using an expansive new dataset of nuclear and chloroplast DNA sequences, we conducted a multifaceted phylogenomic investigation to identify ancient reticulation in the early evolution of oaks (Quercus). We document extensive nuclear gene tree and cytonuclear discordance among major lineages ofQuercusand relatives in Quercoideae. Our analyses recovered clear signatures of gene flow against a backdrop of rampant incomplete lineage sorting, with gene flow most prevalent among major lineages ofQuercusand relatives in Quercoideae during their initial radiation, dated to the Early‐Middle Eocene. Ancestral reconstructions including fossils suggest ancestors ofCastanea + Castanopsis,Lithocarpus, and the Old World oak clade probably co‐occurred in North America and Eurasia, while the ancestors ofChrysolepis, Notholithocarpus, and the New World oak clade co‐occurred in North America, offering ample opportunity for hybridization in each region. Our study shows that hybridization—perhaps in the form of ancient syngameons like those seen today—has been a common and important process throughout the evolutionary history of oaks and their relatives. Concomitantly, this study provides a methodological framework for detecting ancient hybridization in other groups.

Biochemistry & Molecular Biology↗

Comparative physiological and genomic characterization of a novel Nitrobacter vulgaris strain from a nitrate-contaminated subsurface

Nitrite-oxidizing bacteria (NOB) represent a crucial node in the global nitrogen cycle. By catalyzing the second step of nitrification—the oxidation of nitrite to nitrate to generate energy for growth—NOB activity controls the fate of nitrite (NO 2 - ) in aerobic environments. Despite thriving in diverse environments, including soils, freshwater, marine ecosystems, subsurface habitats, and water treatment systems, organisms capable of nitrite oxidation are confined to Nitrobacter, Nitrospira, Nitrospina, Nitrotoga, and a few other specific lineages. The genus Nitrobacter, recognized for its facultative heterotrophic metabolism, is often associated with high-nitrogen environments. Here, we report the physiological characterization of a novel strain, Nitrobacter vulgaris strain MLSD-S22, isolated from a nitrate- and heavy-metal-contaminated subsurface. Growth inhibition experiments revealed that strain MLSD-S22 and the N. vulgaris type strain Z exhibited similar sensitivities to nitrite and nitrate, with nitrite being the most inhibitory. Microrespirometry demonstrated that the two N. vulgaris strains and Nitrobacter winogradskyi Nb-255 possessed higher affinities for nitrite and oxygen than previously reported for Nitrobacter, suggesting potential to compete in low-substrate environments. Long-read DNA sequencing provided a complete genome for strain MLSD-S22, revealing two plasmids and an intact nitrous oxide (N 2 O) reduction operon—an unexpected feature for Nitrobacter. While N 2 O reduction activity was not observed under the tested conditions, this discovery raises questions about the contribution of Nitrobacter NOB to the N 2 O sink. These findings broaden the physiological and genomic diversity of Nitrobacter, offering new insights into their adaptation strategies and providing a framework for future evaluation of their potential roles in nitrogen loss.

Nitrobacter↗

Community Structure of Arbuscular Mycorrhizal Fungi in Soils of Switchgrass Harvested for Bioenergy

We assessed the different species of beneficial fungi living in agricultural fields of switchgrass, a large grass grown for biofuels, using high-resolution DNA sequencing. Contrary to our expectations, the fungi were not greatly affected by fertilization. However, we found a positive relationship between plant productivity and the number of families of beneficial fungi at one site. Furthermore, we sequenced many species that could not be identified with existing reference databases. One group of fungi was highlighted in an earlier study for being widely distributed but of unknown taxonomy. We discovered that this group belonged to a family called Pervetustaceae , which may benefit switchgrass in stressful environments. To produce higher-yielding switchgrass in a more sustainable manner, it could help to study these undescribed fungi and the ways in which they may contribute to greater switchgrass yield in the absence of fertilization.

59 BASIC BIOLOGICAL SCIENCES↗

Synthetic overlapping genes stabilize genetic systems

Overlapping genes—wherein two different proteins are translated from alternative reading frames of the same DNA sequence—provide a means to stabilize an engineered gene by directly linking its evolutionary fate with that of an overlapping gene. However, creating overlapping gene pairs is challenging, as it requires redesigning both protein products to accommodate overlap constraints. Here, we present a new “overlapping, alternate-frame insertion” (OAFI) method for creating synthetic overlapping genes by inserting an “inner” gene, encoded in an alternate frame, into a flexible region of an “outer” gene. Using OAFI, we create new overlapping gene pairs of genetic reporters and bacterial toxins within an antibiotic resistance gene. We show that both the inner and outer genes retain function despite redesign, with translation of the inner gene influenced by its overlap position in the outer gene. Importantly, we show that, despite these inner gene sequences not contributing to outer gene function, selection for the outer gene alters the permitted inactivating mutations in the inner gene, and that overlapping toxins can restrict horizontal gene transfer of the antibiotic resistance gene. Overall, OAFI offers a versatile tool for synthetic biology, expanding the applications of overlapping genes in gene stabilization and biocontainment.

Biological and medical sciences↗

Geochemistry and Multiomics Data Differentiate Streams in Pennsylvania Based on Unconventional Oil and Gas Activity

Unconventional oil and gas (UOG) extraction is increasing exponentially around the world, as new technological advances have provided cost-effective methods to extract hard-to-reach hydrocarbons. While UOG has increased the energy output of some countries, past research indicates potential impacts in nearby stream ecosystems as measured by geochemical and microbial markers. Here, we utilized a robust data set that combines 16S rRNA gene amplicon sequencing (DNA), metatranscriptomics (RNA), geochemistry, and trace element analyses to establish the impact of UOG activity in 21 sites in northern Pennsylvania. These data were also used to design predictive machine learning models to determine the UOG impact on streams. We identified multiple biomarkers of UOG activity and contributors of antimicrobial resistance within the order Burkholderiales. Furthermore, we identified expressed antimicrobial resistance genes, land coverage, geochemistry, and specific microbes as strong predictors of UOG status. Of the predictive models constructed (n = 30), 15 had accuracies higher than expected by chance and area under the curve values above 0.70. The supervised random forest models with the highest accuracy were constructed with 16S rRNA gene profiles, metatranscriptomics active microbial composition, metatranscriptomics active antimicrobial resistance genes, land coverage, and geochemistry (n = 23). The models identified the most important features within those data sets for classifying UOG status. These findings identified specific shifts in gene presence and expression, as well as geochemical measures, that can be used to build robust models to identify impacts of UOG development.

16S rRNA↗

BigDNA

SAND2020-13696 M BigDNA takes a user input fast DNA sequence and outputs overlapping long PCR primers for Gibson Assembly. The software can design primers that remove gene products, insert new gene products, or completely circularize long DNAs. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Porter, Kelley↗

Auto_PDI

Protein-DNA Interaction Workflow (PDI Workflow), a pipeline that focuses on generating high-quality docking and molecular dynamics simulations for Protein-DNA complexes. This allows us to take DNA sequences with unknown tertiary structures, accurately predict their structure, dock them with the desired target protein, and then simulate their interactions using molecular dynamics simulations

Kumar, Neeraj↗

Inventory of Composable Elements (ICE) v6.0.0

The Inventory of Composable Elements (ICE) is an open source registry software platform for managing information about biological parts. It is capable of recording information about plasmids, microbial host strains and seeds, as well as DNA parts. Includes features such as DNA sequence visualization, editing and annotation, auto-aligning sequencing trace files against reference templates, SBOL XML/RDF support, and web-of-registries functionality. The web of registries functionality provides strong support for distributed interconnected use and enables sharing and transfer of biological parts across various independent ICE instances. ICE adopts modern software development principles, leveraging component-base frameworks, offering a REST API for convenient third-party integration and emphasizing scalability, security, and service integrations for dynamic content availability. The source code is hosted at https://github.com/JBEI/ice. A public instance is available at public-registry.jbei.org, where users can try out features, upload parts or simply use it for their projects.

Plahar, Hector↗

germs-lab/LAMPS-miscanthus-microbiome

Data and analysis to accompany our submitted paper characterizing the LAMPS Miscanthus Microbiome comparing 16S rRNA gene DNA sequencing of Miscanthus in a staggered-start experiment. Authors for this analysis include Fernando Igne Rocha, Lanying Ma, and Adina Howe.

Igne Rocha, Fernando↗

GenomeFace v1.0

GenomeFace is meta-genome binning software. Metagenomic binning, the process of grouping DNA sequences into taxonomic units, is critical for understanding the functions, interactions, and evolutionary dynamics of microbial communities. We propose a deep learning approach to binning using two neural networks, one based on composition and another on environmental abundance, dynamically weighting the contribution of each based on characteristics of the input data. Trained on over 43,000 prokaryotic genomes, our network for composition-based binning is inspired by metric learning techniques used for facial recognition. Using a task-specific, multi-GPU accelerated algorithm to cluster the embeddings produced by our network, our binner leverages marker genes observed to be universally present in nearly all taxa to grade and select optimal clusters of sequences from a hierarchy of candidates. We evaluate our approach on four simulated datasets with known ground truth. Our linear time integration of marker genes recovers more near complete genomes than state of the art but computationally infeasible solutions using them, while being over an order of magnitude faster. Finally, we demonstrate the scalability and acuity of our approach by testing it on three of the largest metagenome assemblies ever performed. Compared to other binners, we produced 47%-183% more near complete genomes. From these datasets, we find over the genomes of over 3000 new candidate species which have never been previously cataloged, representing a potential 4% expansion of the known bacterial tree of life.

Lettich, Richard [Lawrence Berkeley National Labor↗

Genomic Language model for Annotation of Repetitive Elements (GLARE) v1.0

GLARE (Genomic Language model for Annotation of Repetitive Elements) is a tool that classifies transposable elements (TEs)—the mobile, repetitive DNA sequences that make up large fractions of eukaryotic genomes. GLARE fine-tunes the NTv3-650M genomic language model on a harmonized collection of curated TE sequences from the PanTEon and Repbase reference databases, assigning each input sequence to one of 11 orders and 32 superfamilies in a Wicker-compatible taxonomy. Features. From nucleotide FASTA input, GLARE outputs per-sequence predictions, class summaries, composition figures, and an annotated FASTA. It provides calibrated confidence scores with optional abstention and runs on CPU or GPU. Uses. GLARE serves as a classification component in genome-annotation pipelines, downstream of TE discovery, supporting genome annotation and comparative and evolutionary genomics. Advantages. GLARE is the first repeat-element classifier to leverage a pretrained genomic language model. Combined with multi-database training, this approach outperformed all nine classifiers in the PanTEon benchmark, generalized better to unseen taxonomic clades, and remained robust to sequence orientation—a common failure mode of existing tools.

Bruna, Tomas [Lawrence Berkeley National Laborator↗

Nanobodies as potential tools for microbiological testing of live biotherapeutic products

Nanobodies are highly specific binding domains derived from naturally occurring single chain camelid antibodies. Live biotherapeutic products (LBPs) are biological products containing preparations of live organisms, such as Lactobacillus, that are intended for use as drugs, i.e. to address a specific disease or condition. Demonstrating potency of multi-strain LBPs can be challenging. The approach investigated here is to use strain-specific nanobody reagents in LBP potency assays. Llamas were immunized with radiation-killed Lactobacillus jensenii or L. crispatus whole cell preparations. A nanobody phage-display library was constructed and panned against bacterial preparations to identify nanobodies specific for each species. Nanobody-encoding DNA sequences were subcloned and the nanobodies were expressed, purified, and characterized. Colony immunoblots and flow cytometry showed that binding by Lj75 and Lj94 nanobodies were limited to a subset of L. jensenii strains while binding by Lc38 and Lc58 nanobodies were limited to L. crispatus strains. Mass spectrometry was used to demonstrate that Lj75 specifically bound a peptidase of L. jensenii, and that Lc58 bound an S-layer protein of L. crispatus. The utility of fluorescent nanobodies in evaluating multi-strain LBP potency assays was assessed by evaluating a L. crispatus and L. jensenii mixture by fluorescence microscopy, flow cytometry, and colony immunoblots. Our results showed that the fluorescent nanobody labelling enabled differentiation and quantitation of the strains in mixture by these methods. Development of these nanobody reagents represents a potential advance in LBP testing, informing the advancement of future LBP potency assays and, thereby, facilitation of clinical investigation of LBPs.

60 APPLIED LIFE SCIENCES↗

Microbial ecology and biogeochemistry of hypersaline sediments in Orca Basin

In deep ocean hypersaline basins, the combination of high salinity, unusual ionic composition and anoxic conditions represents significant challenges for microbial life. We used geochemical porewater characterization and DNA sequencing based taxonomic surveys to enable environmental and microbial characterization of anoxic hypersaline sediments and brines in the Orca Basin, the largest brine basin in the Gulf of Mexico. Full-length bacterial 16S rRNA gene clone libraries from hypersaline sediments and the overlying brine were dominated by the uncultured halophilic KB1 lineage, Deltaproteobacteria related to cultured sulfate-reducing halophilic genera, and specific lineages of heterotrophic Bacteroidetes. Archaeal clones were dominated by members of the halophilic methanogen genus Methanohalophilus, and the ammonia-oxidizing Marine Group I (MG-I) within the Thaumarchaeota. Illumina sequencing revealed higher phylum- and subphylum-level complexity, especially in lower-salinity sediments from the Orca Basin slope. Illumina and clone library surveys consistently detected MG-I Thaumarchaeota and halotolerant Deltaproteobacteria in the hypersaline anoxic sediments, but relative abundances of the KB1 lineage differed between the two sequencing methods. The stable isotopic composition of dissolved inorganic carbon and methane in porewater, and sulfate concentrations decreasing downcore indicated methanogenesis and sulfate reduction in the anoxic sediments. While anaerobic microbial processes likely occur at low rates near their maximal salinity thresholds in Orca Basin, long-term accumulation of reaction products leads to high methane concentrations and reducing conditions within the Orca Basin brine and sediments.

54 ENVIRONMENTAL SCIENCES↗

Bacterial diversity dynamics in microbial consortia selected for lignin utilization

Lignin is nature’s largest source of phenolic compounds. Its recalcitrance to enzymatic conversion is still a limiting step to increase the value of lignin. Although bacteria are able to degrade lignin in nature, most studies have focused on lignin degradation by fungi. To understand which bacteria are able to use lignin as the sole carbon source, natural selection over time was used to obtain enriched microbial consortia over a 12-week period. The source of microorganisms to establish these microbial consortia were commercial and backyard compost soils. Cultivation occurred at two different temperatures, 30°C and 37°C, in defined culture media containing either Kraft lignin or alkaline-extracted lignin as carbon source. iTag DNA sequencing of bacterial 16S rDNA gene was performed for each of the consortia at six timepoints (passages). The initial bacterial richness and diversity of backyard compost soil consortia was greater than that of commercial soil consortia, and both parameters decreased after the enrichment protocol, corroborating that selection was occurring. Bacterial consortia composition tended to stabilize from the fourth passage on. After the enrichment protocol, Firmicutes phylum bacteria were predominant when lignin extracted by alkaline method was used as a carbon source, whereas Proteobacteria were predominant when Kraft lignin was used. Bray-Curtis dissimilarity calculations at genus level, visualized using NMDS plots, showed that the type of lignin used as a carbon source contributed more to differentiate the bacterial consortia than the variable temperature. The main known bacterial genera selected to use lignin as a carbon source were Altererythrobacter , Aminobacter , Bacillus , Burkholderia , Lysinibacillus , Microvirga , Mycobacterium , Ochrobactrum , Paenibacillus , Pseudomonas , Pseudoxanthomonas , Rhizobiales and Sphingobium . These selected bacterial genera can be of particular interest for studying lignin degradation and utilization, as well as for lignin-related biotechnology applications.

59 BASIC BIOLOGICAL SCIENCES↗

Precision engineering of biological function with large-scale measurements and machine learning

As synthetic biology expands and accelerates into real-world applications, methods for quantitatively and precisely engineering biological function become increasingly relevant. This is particularly true for applications that require programmed sensing to dynamically regulate gene expression in response to stimuli. However, few methods have been described that can engineer biological sensing with any level of quantitative precision. Here, we present two complementary methods for precision engineering of genetic sensors: in silico selection and machine-learning-enabled forward engineering. Both methods use a large-scale genotype-phenotype dataset to identify DNA sequences that encode sensors with quantitatively specified dose response. First, we show that in silico selection can be used to engineer sensors with a wide range of dose-response curves. To demonstrate in silico selection for precise, multi-objective engineering, we simultaneously tune a genetic sensor’s sensitivity (EC 50 ) and saturating output to meet quantitative specifications. In addition, we engineer sensors with inverted dose-response and specified EC 50 . Second, we demonstrate a machine-learning-enabled approach to predictively engineer genetic sensors with mutation combinations that are not present in the large-scale dataset. We show that the interpretable machine learning results can be combined with a biophysical model to engineer sensors with improved inverted dose-response curves.

59 BASIC BIOLOGICAL SCIENCES↗

The genotype-phenotype landscape of an allosteric protein

Allostery is a fundamental biophysical mechanism that underlies cellular sensing, signaling, and metabolism. Yet a quantitative understanding of allosteric genotype-phenotype relationships remains elusive. Here, we report the large-scale measurement of the genotype-phenotype landscape for an allosteric protein: the lac repressor from Escherichia coli , LacI. Using a method that combines long-read and short-read DNA sequencing, we quantitatively measure the dose-response curves for nearly 10 5 variants of the LacI genetic sensor. The resulting data provide a quantitative map of the effect of amino acid substitutions on LacI allostery and reveal systematic sequence-structure-function relationships. We find that in many cases, allosteric phenotypes can be quantitatively predicted with additive or neural-network models, but unpredictable changes also occur. For example, we were surprised to discover a new band-stop phenotype that challenges conventional models of allostery and that emerges from combinations of nearly silent amino acid substitutions.

59 BASIC BIOLOGICAL SCIENCES↗