Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “DNA sequencing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

Sequence-directed dynamic covalent assembly of base-4-encoded oligomers

As an information-bearing biomacromolecule, DNA is encoded in base-4, where each residue site can be occupied by any one of four nucleobases. Mimicking the information dense, sequence-selective hybridization of DNA, we demonstrate two orthogonal dynamic covalent interactions to effect the selective assembly of molecular ladders and grids from base-4-encoded oligo(peptoid)s.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A standardized quantitative analysis strategy for stable isotope probing metagenomics

ABSTRACT Stable isotope probing (SIP) facilitates culture-independent identification of active microbial populations within complex ecosystems through isotopic enrichment of nucleic acids. Many DNA-SIP studies rely on 16S rRNA gene sequences to identify active taxa, but connecting these sequences to specific bacterial genomes is often challenging. Here, we describe a standardized laboratory and analysis framework to quantify isotopic enrichment on a per-genome basis using shotgun metagenomics instead of 16S rRNA gene sequencing. To develop this framework, we explored various sample processing and analysis approaches using a designed microbiome where the identity of labeled genomes and their level of isotopic enrichment were experimentally controlled. With this ground truth dataset, we empirically assessed the accuracy of different analytical models for identifying active taxa and examined how sequencing depth impacts the detection of isotopically labeled genomes. We also demonstrate that using synthetic DNA internal standards to measure absolute genome abundances in SIP density fractions improves estimates of isotopic enrichment. In addition, our study illustrates the utility of internal standards to reveal anomalies in sample handling that could negatively impact SIP metagenomic analyses if left undetected. Finally, we present SIPmg , an R package to facilitate the estimation of absolute abundances and perform statistical analyses for identifying labeled genomes within SIP metagenomic data. This experimentally validated analysis framework strengthens the foundation of DNA-SIP metagenomics as a tool for accurately measuring the in situ activity of environmental microbial populations and assessing their genomic potential. IMPORTANCE Answering the questions, “who is eating what?” and “who is active?” within complex microbial communities is paramount for our ability to model, predict, and modulate microbiomes for improved human and planetary health. These questions can be pursued using stable isotope probing to track the incorporation of labeled compounds into cellular DNA during microbial growth. However, with traditional stable isotope methods, it is challenging to establish links between an active microorganism’s taxonomic identity and genome composition while providing quantitative estimates of the microorganism’s isotope incorporation rate. Here, we report an experimental and analytical workflow that lays the foundation for improved detection of metabolically active microorganisms and better quantitative estimates of genome-resolved isotope incorporation, which can be used to further refine ecosystem-scale models for carbon and nutrient fluxes within microbiomes.

54 ENVIRONMENTAL SCIENCES↗

Methods for safely sharing dual-use genetic data

Background: Some genetic data has dual-use potential. Sharing pathogen data has shown tremendous value. For example therapeutic development and lineage tracking during the COVID pandemic. This data sharing is complicated by the fact that these data have the potential to be used for harm. The genome sequence of a pathogen can be used to enable malicious genetic engineering approaches or to recreate the pathogen from synthetic DNA. Standard data security methods can be applied to genetic data, but when data is shared between institutions, ensuring appropriate security can be difficult. Sensitive data that is shared internationally among a wide array of institutions can be especially difficult to control. Methods for securely storing and sharing genetic data with potential for dual-use are needed to mitigate this potential harm.Results: Here we propose new methods that allow genetic data to be shared in a data format that prevents a nefarious actor from accessing sensitive aspects of the data. Our methods obfuscate raw sequence data by pooling reads from different samples. This approach can ensure that data is secure while stored and during electronic transfer. We demonstrate that by pooling raw sequence data from multiple samples of the same organism, the ability to fully reconstruct any individual sample is prevented. In the pooled data, most genomic information remains, but reads or mutations cannot be directly attributed to any individual sample. To further restrict access to information, regions of a genome can be removed from the reads.Conclusion: Our methods obscure genomic information within raw sequence reads. This method can allow genetic data to be stored and shared while preventing a nefarious actor from being able to perfectly reconstruct an organism. Broad-scale sequence information remains, while fine scale details about specific samples are difficult or impossible to reconstruct. Our software is available at https://github.com/Geneinfosec-Inc/ReadMixer.

59 BASIC BIOLOGICAL SCIENCES↗

Platform for efficient large-scale storage and analysis of multi-omics data in plant and microbial systems (Final Technical Report)

Genomic variation at the sequence level fundamentally affects the phenotypic state of all organisms at all stages of development, while dynamic processes such as changes in the epigenome (e.g. DNA methylation state) and transcriptome regulate the specific phenotype expressed at any given state of development based upon that genomic variation. In plants, DNA methylation is a particularly important mechanism for both regulating transcriptomic expression and for management of genomic variations that could be deleterious to the organism due to the presence of active retrotransposons in plant genomes. While DNA methylation is heritable, it is also dynamic through a given plant’s development and life cycle, particularly during the development from seed to mature specimen suggesting variations in DNA methylation could be critical regulators of biologically and commercially important phenotypes such as time to flowering; in addition, plant DNA methylation is more complex than that of animals, with methylation of CHG and CHH trinucleotides evident in addition to the better-known CG methylation. The complexity of plant DNA methylation and its interplay with genomic sequence variation, transcriptomics and other epigenomic factors demand a storage and analysis framework that can cope with the complexity both within a single specimen and with analyses that span many individuals and even many species, such as attempts to extend models from model organisms to commercially relevant species. In addition to complexity, the rapid development and proliferation of sequencing technology has led to an explosion of data volume that conventional storage and analysis solutions will likely be unable to cope with in the long run. We proposed to study these with suitable distributed storage and computation and therefore for the application of cloud computing to biological analyses; integrate with existing data sources and compatible with virtually any interface use case, from fully automated shell scripts to notebooks and do all these at scale in this STTR grant.

60 APPLIED LIFE SCIENCES↗

High-throughput detection of T-DNA insertion sites for multiple transgenes in complex genomes

Abstract Background Genetic engineering of crop plants has been successful in transferring traits into elite lines beyond what can be achieved with breeding techniques. Introduction of transgenes originating from other species has conferred resistance to biotic and abiotic stresses, increased efficiency, and modified developmental programs. The next challenge is now to combine multiple transgenes into elite varieties via gene stacking to combine traits. Generating stable homozygous lines with multiple transgenes requires selection of segregating generations which is time consuming and labor intensive, especially if the crop is polyploid. Insertion site effects and transgene copy number are important metrics for commercialization and trait efficiency. Results We have developed a simple method to identify the sites of transgene insertions using T-DNA-specific primers and high-throughput sequencing that enables identification of multiple insertion sites in the T 1 generation of any crop transformed via Agrobacterium . We present an example using the allohexaploid oil-seed plant Camelina sativa to determine insertion site location of two transgenes. Conclusion This new methodology enables the early selection of desirable transgene location and copy number to generate homozygous lines within two generations.

59 BASIC BIOLOGICAL SCIENCES↗

High-throughput single-cell transcriptomics of bacteria using combinatorial barcoding

Microbial split-pool ligation transcriptomics (microSPLiT) is a high-throughput single-cell RNA sequencing method for bacteria. With four combinatorial barcoding rounds, microSPLiT can profile transcriptional states in hundreds of thousands of Gram-negative and Gram-positive bacteria in a single experiment without specialized equipment. As bacterial samples are fixed and permeabilized before barcoding, they can be collected and stored ahead of time. During the first barcoding round, the fixed and permeabilized bacteria are distributed into a 96-well plate, where their transcripts are reverse transcribed into cDNA and labeled with the first well-specific barcode inside the cells. The cells are mixed and redistributed two more times into new 96-well plates, where the second and third barcodes are appended to the cDNA via in-cell ligation reactions. Finally, the cells are mixed and divided into aliquot sub-libraries, which can be stored until future use or prepared for sequencing with the addition of a fourth barcode. It takes 4 days to generate sequencing-ready libraries, including 1 day for collection and overnight fixation of samples. Here, the standard plate setup enables single-cell transcriptional profiling of up to 1 million bacterial cells and up to 96 samples in a single barcoding experiment, with the possibility of expansion by adding barcoding rounds. The protocol requires experience in basic molecular biology techniques, handling of bacterial samples and preparation of DNA libraries for next-generation sequencing. It can be performed by experienced undergraduate or graduate students. Data analysis requires access to computing resources, familiarity with Unix command line and basic experience with Python or R.

59 BASIC BIOLOGICAL SCIENCES↗

Cryo-EM structures of DNA-free and DNA-bound BsaXI: architecture of a Type IIB restriction–modification enzyme

Abstract We have determined multiple cryogenic electron microscopy (cryo-EM) structures of the Type IIB restriction–modification enzyme BsaXI. Such enzymes cleave DNA on both sides of their recognition sequence and share features of Types I, II, and III restriction systems. BsaXI forms a heterotrimeric (RM)2S assemblage in the presence and absence of bound DNA. Two unique structural motifs—a multi-helical “knob” and a long antiparallel double-helical “paddle”—are involved in DNA binding and cleavage. Binding of the DNA target triggers a large conformational change from an ‘open’ to ‘closed’ configuration, resulting in a mixture of two different conformations with respect to the positioning of the S subunit and its target recognition domains on the enzyme’s bipartite DNA target site. Structure-guided mutagenesis studies implicated two clusters of residues in the RM subunit as being critical for DNA cleavage, both are located proximal to a DNA cleavage site. One corresponds to a canonical PD-(D/E)xK endonuclease site in the N-terminal endonuclease domain, while the other corresponds to residues clustered within the paddle motif (near to the C-terminal end of the RM subunit). This analysis facilitates a comparison of three potential mechanisms by which such enzymes cleave DNA on each side of the bound target.

Shen, Betty W.↗

Integrase-on-Demand

SAND2025-07449O Integrase-on-Demand is a software tool that allows users to identify regions in genomic sequences where genetic material can be integrated with high probability. It uses a database of integrases and their DNA attachment sites to search against any genomic sequence, producing a list of open sites, the integrase sequence, and the source of the genomic island. The program requires MASH software to be available on the system. It consists of a main script and a precomputed input file, with a taxonomy mode that searches closely related genomes and a search mode that looks for identical attachment site matches in the integrase/attachment input file. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Williams, Kelly [Sandia National Lab. (SNL-CA), Li↗

Structures of CTCF–DNA complexes including all 11 zinc fingers

The CCCTC-binding factor (CTCF) binds tens of thousands of enhancers and promoters on mammalian chromosomes by means of its 11 tandem zinc finger (ZF) DNA-binding domain. In addition to the 12–15-bp CORE sequence, some of the CTCF binding sites contain 5' upstream and/or 3' downstream motifs. Here, we describe two structures for overlapping portions of human CTCF, respectively, including ZF1–ZF7 and ZF3–ZF11 in complex with DNA that incorporates the CORE sequence together with either 3' downstream or 5' upstream motifs. Like conventional tandem ZF array proteins, ZF1–ZF7 follow the right-handed twist of the DNA, with each finger occupying and recognizing one triplet of three base pairs in the DNA major groove. ZF8 plays a unique role, acting as a spacer across the DNA minor groove and positioning ZF9–ZF11 to make cross-strand contacts with DNA. We ascribe the difference between the two subgroups of ZF1–ZF7 and ZF8–ZF11 to residues at the two positions -6 and -5 within each finger, with small residues for ZF1–ZF7 and bulkier and polar/charged residues for ZF8–ZF11. ZF8 is also uniquely rich in basic amino acids, which allows salt bridges to DNA phosphates in the minor groove. Highly specific arginine–guanine and glutamine–adenine interactions, used to recognize G:C or A:T base pairs at conventional base-interacting positions of ZFs, also apply to the cross-strand interactions adopted by ZF9–ZF11. The differences between ZF1–ZF7 and ZF8–ZF11 can be rationalized structurally and may contribute to recognition of high-affinity CTCF binding sites.

59 BASIC BIOLOGICAL SCIENCES↗

Structure of native four-repeat satellite III sequence with non-canonical base interactions

Abstract Tandem-repetitive DNA (where two or more DNA bases are repeated numerous times) can adopt non-canonical secondary structures. Many of these structures are implicated in important biological processes. Human Satellite III (HSat3) is enriched for tandem repeats of the sequence ATGGA and is located in pericentromeric heterochromatin in many human chromosomes. Here, we investigate the secondary structure of the four-repeat HSat3 sequence 5′-ATGGA ATGGA ATGGA ATGGA-3′ using X-ray crystallography, NMR, and biophysical methods. Circular dichroism spectroscopy, thermal stability, native PAGE, and analytical ultracentrifugation indicate that this sequence folds into a monomolecular hairpin with non-canonical base pairing and B-DNA characteristics at concentrations below 0.9 mM. NMR studies at 0.05–0.5 mM indicate that the hairpin is likely folded-over into a compact structure with high dynamics. Crystallographic studies at 2.5 mM reveal an antiparallel self-complementary duplex with the same base pairing as in the hairpin, extended into an infinite polymer. The non-canonical base pairing includes a G–G intercalation sandwiched by sheared A–G base pairs, leading to a cross-strand four guanine stack, so called guanine zipper. The guanine zippers are spaced throughout the structure by A–T/T–A base pairs. Our findings lend further insight into recurring structural motifs associated with the HSat3 and their potential biological functions.

59 BASIC BIOLOGICAL SCIENCES↗

Mycorrhiza better predict soil fungal community composition and function than aboveground traits in temperate forest ecosystems

We sampled soil in the organic and mineral horizons beneath two AM-associated (Fraxinus americana, Thuja occidentalis) and two ECM-associated tree species (Betula alleghaniensis, and Tsuga canadensis), with an evergreen and deciduous species in each mycorrhizal group. To characterize fungal communities and organic matter decomposition beneath each tree species, we sequenced the ITS1 region of fungal DNA and measured the potential activity of carbon and nitrogen-targeting extracellular enzymes. Each tree species harbored distinct fungal communities, supporting the need to consider both mycorrhizal type and leaf habit. However, between tree characteristics, mycorrhizal type better predicted fungal communities. Across fungal guilds, saprotrophic fungi were the most important group in shaping fungal community differences in soils beneath all tree species. The effect of leaf habit on carbon and nitrogen-targeting hydrolytic enzymes depended on tree mycorrhizal association in the organic horizon, while oxidative enzyme activities were higher beneath EcM-associated trees across both soil horizons and leaf habits. These data include extracellular hydrolytic and oxidative enzyme activities, ITS sequencing fungal community data, soil carbon to nitrogen, and soil pH. Site level data include climate (mean annual temperature, precipitation), elevation, and soil series and order information.

54 ENVIRONMENTAL SCIENCES↗

Structural dissection of sequence recognition and catalytic mechanism of human LINE-1 endonuclease

Abstract Long interspersed nuclear element-1 (L1) is an autonomous non-LTR retrotransposon comprising ∼20% of the human genome. L1 self-propagation causes genomic instability and is strongly associated with aging, cancer and other diseases. The endonuclease domain of L1’s ORFp2 protein (L1-EN) initiates de novo L1 integration by nicking the consensus sequence 5′-TTTTT/AA-3′. In contrast, related nucleases including structurally conserved apurinic/apyrimidinic endonuclease 1 (APE1) are non-sequence specific. To investigate mechanisms underlying sequence recognition and catalysis by L1-EN, we solved crystal structures of L1-EN complexed with DNA substrates. This showed that conformational properties of the preferred sequence drive L1-EN’s sequence-specificity and catalysis. Unlike APE1, L1-EN does not bend the DNA helix, but rather causes ‘compression’ near the cleavage site. This provides multiple advantages for L1-EN’s role in retrotransposition including facilitating use of the nicked poly-T DNA strand as a primer for reverse transcription. We also observed two alternative conformations of the scissile bond phosphate, which allowed us to model distinct conformations for a nucleophilic attack and a transition state that are likely applicable to the entire family of nucleases. This work adds to our mechanistic understanding of L1-EN and related nucleases and should facilitate development of L1-EN inhibitors as potential anticancer and antiaging therapeutics.

59 BASIC BIOLOGICAL SCIENCES↗

The Hyperthermophilic Restriction-Modification Systems of Thermococcus kodakarensis Protect Genome Integrity

Thermococcus kodakarensis (T. kodakarensis), a hyperthermophilic, genetically accessible model archaeon, encodes two putative restriction modification (R-M) defense systems, TkoI and TkoII. TkoI is encoded by TK1460 while TkoII is encoded by TK1158. Bioinformative analysis suggests both R-M enzymes are large, fused methyltransferase (MTase)-endonuclease polypeptides that contain both restriction endonuclease (REase) activity to degrade foreign invading DNA and MTase activity to methylate host genomic DNA at specific recognition sites. In this work, we demonsrate T. kodakarensis strains deleted for either or both R-M enzymes grow more slowly but display significantly increased competency compared to strains with intact R-M systems, suggesting that both TkoI and TkoII assist in maintenance of genomic integrity in vivo and likely protect against viral- or plasmid-based DNA transfers. Pacific Biosciences single molecule real-time (SMRT) sequencing of T. kodakarensis strains containing both, one or neither R-M systems permitted assignment of the recognition sites for TkoI and TkoII and demonstrated that both R-M enzymes are TypeIIL; TkoI and TkoII methylate the N 6 position of adenine on one strand of the recognition sequences GTGAAG and TTCAAG, respectively. Further in vitro biochemical characterization of the REase activities reveal TkoI and TkoII cleave the DNA backbone GTGAAG(N) 20 /(N) 18 and TTCAAG(N) 10 /(N) 8 , respectively, away from the recognition sequences, while in vitro characterization of the MTase activities reveal transfer of tritiated S-adenosyl methionine by TkoI and TkoII to their respective recognition sites. Together these results demonstrate TkoI and TkoII restriction systems are important for protecting T. kodakarensis genome integrity from invading foreign DNA.

59 BASIC BIOLOGICAL SCIENCES↗

A Two-Step PCR Protocol Enabling Flexible Primer Choice and High Sequencing Yield for Illumina MiSeq Meta-Barcoding

High-throughput amplicon sequencing that primarily targets the 16S ribosomal DNA (rDNA) (for bacteria and archaea) and the Internal Transcribed Spacer rDNA (for fungi) have facilitated microbial community discovery across diverse environments. A three-step PCR that utilizes flexible primer choices to construct the library for Illumina amplicon sequencing has been applied to several studies in forest and agricultural systems. The three-step PCR protocol, while producing high-quality reads, often yields a large number (up to 46%) of reads that are unable to be assigned to a specific sample according to its barcode. Here, we improve this technique through an optimized two-step PCR protocol. We tested and compared the improved two-step PCR meta-barcoding protocol against the three-step PCR protocol using four different primer pairs (fungal ITS: ITS1F-ITS2 and ITS1F-ITS4, and bacterial 16S: 515F-806R and 341F-806R). We demonstrate that the sequence quantity and recovery rate were significantly improved with the two-step PCR approach (fourfold more read counts per sample; determined reads ≈90% per run) while retaining high read quality (Q30 > 80%). Given that synthetic barcodes are incorporated independently from any specific primers, this two-step PCR protocol can be broadly adapted to different genomic regions and organisms of scientific interest.

16S rDNA↗

Structural snapshots of human DNA polymerase μ engaged on a DNA double-strand break

Genomic integrity is threatened by cytotoxic DNA double-strand breaks (DSBs), which must be resolved efficiently to prevent sequence loss, chromosomal rearrangements/translocations, or cell death. Polymerase μ (Polμ) participates in DSB repair via the nonhomologous end-joining (NHEJ) pathway, by filling small sequence gaps in broken ends to create substrates ultimately ligatable by DNA Ligase IV. Here we present structures of human Polμ engaging a DSB substrate. Synapsis is mediated solely by Polμ, facilitated by single-nucleotide homology at the break site, wherein both ends of the discontinuous template strand are stabilized by a hydrogen bonding network. The active site in the quaternary Pol μ complex is poised for catalysis and nucleotide incoporation proceeds in crystallo. These structures demonstrate that Polμ may address complementary DSB substrates during NHEJ in a manner indistinguishable from single-strand breaks.

59 BASIC BIOLOGICAL SCIENCES↗

The genome of the polyextremophilic yeast, Naganishia friedmannii, reveals adaptations involved in stress response pathways, carbohydrate metabolism expansion, and a limited DNA repair repertoire

Here we report the draft genome sequence of Naganishia friedmannii (formerly Cryptococcus friedmannii) isolate, a Basidiomycota yeast commonly found in some of the most extreme environments of the Earth's cryosphere. We isolated N. friedmannii strain Llullensis from soils at 6000 m above sea level on Volcán Llullaillaco, Argentina. The genome was 22.2 Mb with 6251 identified protein coding genes. Proteins known to be associated with thermal, osmotic, and radiation stress were identified in the genome. Comparative analysis with seven other Naganishia genomes revealed unique features underlying its polyextremophilic lifestyle. Naganishia friedmannii showed an expansion of genes involved in breaking down plant-derived carbohydrates, supporting the hypothesis that it survives at high elevations by metabolizing wind-deposited organic matter. Surprisingly, many genes involved in cell-cycle checkpoints and DNA repair were missing, as in several other Naganishia species. This extensive loss may be adaptive in extreme environments prone to abiotic stress, where a high mutation rate could generate advantageous traits, and reduced cell-cycle control may allow for faster reproduction that would be advantageous for rapid growth during brief periods of soil wetting following rare snow events.

Vimercati, Lara↗

Targeted mutagenesis with sequence–specific nucleases for accelerated improvement of polyploid crops: Progress, challenges, and prospects

Many of the world's most important crops are polyploid. The presence of more than two sets of chromosomes within their nuclei and frequently aberrant reproductive biology in polyploids present obstacles to conventional breeding. The presence of a larger number of homoeologous copies of each gene makes random mutation breeding a daunting task for polyploids. Genome editing has revolutionized improvement of polyploid crops as multiple gene copies and/or alleles can be edited simultaneously while preserving the key attributes of elite cultivars. Most genome–editing platforms employ sequence–specific nucleases (SSNs) to generate DNA double–stranded breaks at their target gene. Such DNA breaks are typically repaired via the error–prone nonhomologous end–joining process, which often leads to frame shift mutations, causing loss of gene function. Genome editing has enhanced the disease resistance, yield components, and end–use quality of polyploid crops. However, identification of candidate targets, genotyping, and requirement of high mutagenesis efficiency remain bottlenecks for targeted mutagenesis in polyploids. In this review, we will survey the tremendous progress of SSN–mediated targeted mutagenesis in polyploid crop improvement, discuss its challenges, and identify optimizations needed to sustain further progress.

60 APPLIED LIFE SCIENCES↗

Peatland Microbial Community Composition Is Driven by a Natural Climate Gradient

Peatlands are important players in climate change–biosphere feedbacks via long-term net carbon (C) accumulation in soil organic matter and as potential net C sources including the potent greenhouse gas methane (CH 4 ). Interactions of climate, site-hydrology, plant community, and groundwater chemical factors influence peatland development and functioning, including C dioxide (CO 2 ) and CH 4 fluxes, but the role of microbial community composition is not well understood. To assess microbial functional and taxonomic dissimilarities, we used high throughput sequencing of the small subunit ribosomal DNA (SSU rDNA) to determine bacterial and archaeal community composition in soils from twenty North American peatlands. Targeted DNA metabarcoding showed that although Proteobacteria, Acidobacteria, and Actinobacteria were the dominant phyla on average, intermediate and rich fens hosted greater diversity and taxonomic richness, as well as an array of candidate phyla when compared with acidic and nutrient-poor poor fens and bogs. Moreover, pH was revealed to be the strongest predictor of microbial community structure across sites. Predictive metagenome content (PICRUSt) showed increases in specific genes, such as purine/pyrimidine and amino-acid metabolism in mid-latitude peatlands from 38 to 45° N, suggesting a shift toward utilization of microbial biomass over utilization of initial plant biomass in these microbial communities. Overall, there appears to be noticeable differences in community structure between peatland classes, as well as differences in microbial metabolic activity between latitudes. Finally, these findings are in line with a predicted increase in the decomposition and accelerated C turnover, and suggest that peatlands north of 37° latitude may be particularly vulnerable to climate change.

59 BASIC BIOLOGICAL SCIENCES↗