Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “DNA sequencing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Homology and the optimization of DNA sequence data

Three methods of nucleotide character analysis are discussed. Their implications for molecular sequence homology and phylogenetic analysis are compared. The criterion of inter-data set congruence, both character based and topological, are applied to two data sets to elucidate and potentially discriminate among these parsimony-based ideas. c2001 The Willi Hennig Society.

Non-NASA Center↗

Bubble Relaxation Dynamics in Homopolymer DNA Sequences

Understanding the inherent timescales of large bubbles in DNA is critical to a thorough comprehension of its physicochemical characteristics, as well as their potential role on helix opening and biological function. In this work, we employ the coarse-grained Peyrard–Bishop–Dauxois model of DNA to study relaxation dynamics of large bubbles in homopolymer DNA, using simulations up to the microsecond time scale. By studying energy autocorrelation functions of relatively large bubbles inserted into thermalised DNA molecules, we extract characteristic relaxation times from the equilibration process for both adenine–thymine (AT) and guanine–cytosine (GC) homopolymers. Bubbles of different amplitudes and widths are investigated through extensive statistics and appropriate fittings of their relaxation. Characteristic relaxation times increase with bubble amplitude and width. We show that, within the model, relaxation times are two orders of magnitude longer in GC sequences than in AT sequences. Overall, our results confirm that large bubbles leave a lasting impact on the molecule’s dynamics, for times between 0.5–500 ns depending on the homopolymer type and bubble shape, thus clearly affecting long-time evolutions of the molecule.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Long-range correlation properties of coding and noncoding DNA sequences: GenBank analysis

An open question in computational molecular biology is whether long-range correlations are present in both coding and noncoding DNA or only in the latter. To answer this question, we consider all 33301 coding and all 29453 noncoding eukaryotic sequences--each of length larger than 512 base pairs (bp)--in the present release of the GenBank to dtermine whether there is any statistically significant distinction in their long-range correlation properties. Standard fast Fourier transform (FFT) analysis indicates that coding sequences have practically no correlations in the range from 10 bp to 100 bp (spectral exponent beta=0.00 +/- 0.04, where the uncertainty is two standard deviations). In contrast, for noncoding sequences, the average value of the spectral exponent beta is positive (0.16 +/- 0.05) which unambiguously shows the presence of long-range correlations. We also separately analyze the 874 coding and the 1157 noncoding sequences that have more than 4096 bp and find a larger region of power-law behavior. We calculate the probability that these two data sets (coding and noncoding) were drawn from the same distribution and we find that it is less than 10(-10). We obtain independent confirmation of these findings using the method of detrended fluctuation analysis (DFA), which is designed to treat sequences with statistical heterogeneity, such as DNA's known mosaic structure ("patchiness") arising from the nonstationarity of nucleotide concentration. The near-perfect agreement between the two independent analysis methods, FFT and DFA, increases the confidence in the reliability of our conclusion.

Non-NASA Center↗

Tracing the Path of Carbon Export in the Ocean though DNA Sequencing of Individual Sinking Particles

Surface phytoplankton communities were linked with the carbon they export into the deep ocean by comparing 18 S rRNA gene sequence communities from surface seawater and individually isolated sinking particles. Particles were collected in sediment traps deployed at locations in the North Pacific subtropical gyre and the California Current. DNA was isolated from individual particles, bulk-collected trap particles, and the surface seawater. The relative sequence abundance of exported phytoplankton taxa in the surface water varied across functional groups and ecosystems. Of the sequences detected in sinking particles, about half were present in large (>300 μm), individually isolated particles and primarily belonged to taxa with small cell sizes (<50 μm). Exported phytoplankton taxa detected only in bulk trap samples, and thus presumably packaged in the smaller sinking size fraction, contained taxa that typically have large cell sizes (>500 m). The effect of particle degradation on the detectable 18 S rRNA gene community differed across taxa, and differences in community composition among individual particles from the same location largely reflected differences in relative degradation state. Using these data and particle imaging, we present an approach that incorporates genetic diversity into mechanistic models of the ocean's biological carbon pump, which will lead to better quantification of the ocean’s carbon cycle.

Carbon Expert↗

DNA Sequence Analysis of an Inversion Hot Spot in Lobeliaceae Plastomes

The evolution of plastid genomes (plastomes) in land plants is typically conservative, with extensive structural rearrangements present in only a few groups. Early Southern blot analysis identified two Lobelia species that minimally required deletion of the plastid gene accD and five inversions to account for their plastome arrangement relative to the ancestral organization. Sixty alternative 5-step inversion scenarios could account for the observed arrangement, but only one scenario was consistent with the criterion of ‘common cause’ attributable to a putative rearrangement hot spot at the accD deletion-site. Plastome sequencing demonstrated that this previously hypothesized inversion order is historically accurate. Detailed reconstructions of the ancestral plastome organization before and after each inversion are presented herein. Stem-loop and disruption-rescue models were evaluated for each inversion. One inversion has an obvious stem-loop basis, but the other four inversions were primarily caused by serial insertion of foreign (extra-plastid) DNA bearing large open-reading frames that disrupted plastome organization at the accD deletion-site, and complete plastomes were rescued by seemingly arbitrary ligation or fortuitous recombination at the other inversion endpoint. Transposed copies of DNA segments from elsewhere in the plastome are frequently inserted at inversion junctions, and four junctions are consistent with the stem-loop ligation model.

59 BASIC BIOLOGICAL SCIENCES↗

Quantitation of Radiation Induced Deletion and Recombination Events Associated with Repeated DNA Sequences

Manned exploration of space exposes the explorers to a complex and novel radiation environment. The galactic cosmic ray and trapped belt radiation (predominantly proton) components of this environment are relatively constant, and the variations with the solar cycle are well understood and predictable. The level of radiation encountered in low earth orbits is determined by several factors, including altitude, inclination of orbit with respect to the equator, and spacecraft shielding. At higher altitudes, and on a Mars mission, the level of radiation exposure will increase significantly. A significant fraction of the dose may be delivered by solar particle events which vary dramatically in dose rate and incident particle spectrum. High-LET radiation is of particular concern. High-LET radiation, a component of galactic cosmic rays (GCR), is comprised of a variety of charged particles of various energies (10 MeV/n to 10 GeV/n), including about 87% photons, 12% helium ions, and heavy ions (including iron). These high energy particles can cause significant damage to target cells. The different particle types and energies result in different patterns of energy deposition at the molecular and cellular level in a primary target cell. They can also cause significant damage to other, nearby cells as a result of secondary particles. Protons, for instance produce secondaries that include photons, neutrons, pions, heavy particles, as well as gamma rays. Heavy ions deposit energy in a "track" in which the magnitude of the damage varies as the particle loses energy. Heavy ions produce secondary delta rays, or electrons. The distribution of damage through tissue is described by a Bragg curve which will be characteristic for different energies. Needless to say there are differences in the RBE of protons and a particles. High-LET heavy ions are particularly damaging to cells as they do continual damage throughout their track. Differences in these energy deposition patterns can significantly influence the nature of DNA damage and the ability of cellular systems to repair such damage. It has been suspected that these differences also affect the spatial distribution of damage within the DNA of the interphase cell nucleus and produce corresponding differences in endpoints related to health effects. The interaction of a single high-LET particle with chromatin has been suggested to cause multiple double strand breaks within a relatively short distance. In part this is due to the organization of DNA into chromatin fibers in which distant regions of the DNA helix can be physically juxtaposed by the various levels of coiling of the DNA. This prediction was confirmed by the detection of the generation of double strand DNA fragments of 100-2000 bp following exposure to high-LET ions (including iron).

Sinden, Richard R.↗

Supplementary table, figures and DNA sequences of sorghum gene models SbiRTx430.01G455400 and SbiRTx.02G006600 that feature primers, gRNAs and indels created

In-context promoter bashing via genome editing is a route to identify and characterize critical regulatory regions that govern expression of genes of interest. The outcomes of in-context promoter bashing can be used to inform editing strategies to modulate the expression of selected gene models in a desired fashion. Here we employed in-context promoter bashing to characterize the proximal upstream regulatory regions of sorghum genes encoding phosphoenolpyruvate carboxykinase (Sb.PEPCK.BS, SbiTx430.01G455400) and alanine aminotransferase (SbiTx430.02G006600, SbAlaAT.BS), two proteins involved in the PCK C4 pathway. Characterized germinal edits within the targeted regions upstream of these two genes ranged in size from 138 bp up to 1790 bp. A 138 bp within the Sb.PEPCK.BS upstream region and a 1643 bp element within the Sb.AlaAT.BS upstream region were determined to be important for maintenance of transcription levels. No change in development or various physiological parameters was observed in characterized lineages carrying promoter edits. However, significant changes in seed reserves and a reduction in 100 seed weight were consistently observed, under both greenhouse and field environments, in plants carrying an edit in the promoter of Sb.PEPCK.BS gene were significantly reduced in transcript accumulation for this gene.

Quach, Truyen [Center for Plant Science Innovation↗

The chemical structure of DNA sequence signals for RNA transcription

The proposed recognition sites for RNA transcription for E. coli NRA polymerase, bacteriophage T7 RNA polymerase, and eukaryotic RNA polymerase Pol II are evaluated in the light of the requirements for efficient recognition. It is shown that although there is good experimental evidence that specific nucleic acid sequence patterns are involved in transcriptional regulation in bacteria and bacterial viruses, among the sequences now available, only in the case of the promoters recognized by bacteriophage T7 polymerase does it seem likely that the pattern is sufficient. It is concluded that the eukaryotic pattern that is investigated is not restrictive enough to serve as a recognition site.

George, D. G.↗

Monitoring Astronaut Health with DNA Sequencing

In recent years microbe a plethora of microbe populations have been identified onboard the ISS (International Space Station). Approaches for real-time tracking of microbes for routine housekeeping and food/water safety monitoring will be critical for mission safety and crew health on future longer duration missions to the Moon or Mars. This work is a proof-of-concept study demonstrating an end-to-end phylogenetic identification and full genome sequencing effort of multiple microbial populations. Our methodology utilized the ISS flight-certified WetLab-2 molecular toolbox and the Biomolecule Sequencer projects for real-time end-to-end on-orbit microbial biological samples processing and molecular analysis with real time results generated utilizing only field "offline" analytic software. For this experiment we colony-cultured several ISS isolated microorganisms before generation of the pre-sequencing library via the automated VolTRAX device which enabled high library turnover with little wet-bench activity or potential future costly astronaut time. The pre-sequencing library is diluted in loading buffer and injected into the MinION sample port, drawn into the nanopore window by capillary action, and sequenced using the MinKnown. 16S and full genome alignment, nucleotide matching, gene identification, and phylogenetic sorting was accomplished utilizing the Epi2me software and the offline NCBI Blast viral, microbiome, and human somatic databases. In short, the methodologies developed herein replace the myriad of specific, often highly targeted microbiological tests used in the clinical laboratory, which would be difficult if not impossible to currently implement aboard the ISS or in deep space, with a single metagenomics test.

genomics↗

Does Collection Time Bias the Ecology of Cleanroom Air Samples?

Microbial monitoring of astromaterials collections has taken on increased importance with the return of biologically sensitive samples from the asteroids Ryugu and Bennu and the initiation of the Mars Sample Return Program. Terrestrial bacteria and fungi can alter the mineralogy and organic composition of our collections causing irreversible contamination of pristine samples and increasing the risk of false positives for life detection measurements. NASA has conducted routine microbial monitoring of its existing collections since 20181. Initial monitoring focused on surface samples collected with foam swabs. Although, airborne microbiology is often decoupled from surface microbiology in the built environment2 culture-based air sampling techniques like impactors were not compliant with existing contamination control requirements. Bringing organic rich media, gelatin or liquids into curation cleanrooms presents an unacceptable risk to pristine samples. In 2022 NASA purchased a materials complaint air sampler and began collecting air samples from the cleanrooms in addition to surface samples3. The new instrument uses an electret filter to collect samples that are suitable for cultivating organisms or for direct DNA sequencing. Preliminary DNA sequencing results appeared to indicate that longer sampling times biased the microbial community in favor of hearty, spore-forming bacteria3. We present the results of a study comparing overnight sampling (17 hours) to short (1 hour) sampling of unoccupied curation cleanrooms. The results will help us optimize our monitoring protocols and develop a more detailed inventory of the ecology of astromaterials curation cleanrooms. Methods: We analyzed 72 paired air samples from six different cleanrooms including the meteorite processing lab (ISO 7 equivalent, 16 samples), the lunar lab (ISO 6 equivalent, 10 samples), the stardust lab (ISO 5 equivalent 14 samples), the OSIRIS-REx lab (ISO 5 equivalent, 12 samples), the Hayabusa2 lab (ISO 5 equivalent, 14 samples), and the Genesis lab (ISO 4 equivalent, 6 samples). All the samples were collected with an InnovaPrep Bobcat air sampler operating at a sampling rate of 200 L/min. The sampler operates for 5 minutes out of every 20 minute period. Half of the samples were collected by filtering 3,000L (15 min. of active sampling) of air across an electret filter for one hour. The rest of the samples were collected by filtering approximately 51,000 L air across the filter overnight (~17 hours, 255 min. of active sampling). Cells were eluted from the filter using 6-7 ml of pressurized 0.15% tween 20 in PBS (phosphate buffered saline). This liquid was used to cultivate bacteria according to previously published methods1,4,5 and for DNA extraction and next generation sequencing. DNA was extracted with a Qiagen MagAttract PowerMicrobiome kit6. To identify bacteria and archaea, the 16S rRNA gene was amplified using Earth Microbiome primers for the V4 region 7. The amplified DNA was sequenced on an Illumina MiSeq using a V3 reagent kit. The resulting sequences were processed using DADA2 and QIIME2 as implemented on the EDGE bioinformatics platform8–10. Results: Only two of the 72 samples had no amplifiable DNA. Amplified DNA concentrations ranged from 2.67 – 0.272 ng/µl. The median concentration of amplified DNA for the 1 hour samples was 0.770 ± 0.368 ng/µl. The median concentration of amplified DNA for the overnight samples was 0.877 ± 0.434 ng/µl. On average the overnight samples had slightly more sequences (58,960 vs. 59,456) and ASV’s (amplicon sequence variants) (60 vs 64.5) than the one hour samples, but these differences are not statistically significant. The most abundant ASV in every sample mapped to the genus Cupravidus. ASV’s mapping to the genuses Bacillus, Schlegelella, Thermus, and Staphylococcus were also common. Discussion and Future Work: Alpha diversity statistics like Shannon Entropy and Faith Phylogenetic Diversity are used to describe the diversity of organisms in a single sample. If a longer sampling time was biasing the data, we would expect to see a change in these diversity statistics vs. sample time. However, we did not observe this in our data. The median Shannon entropy was slightly higher for the overnight samples (3.773 vs 3.611) as was the Faith Phylogenetic Diversity (4.042 vs 3.596), but both values were within a standard deviation of each other for the two sampling times (Fig. 1). It is unlikely, that the longer sampling time is introducing bias into our data. We do observe a significant decrease in diversity when comparing the air samples by lab. The Genesis lab (ISO 4 equivalent) has a lower median number of ASV’s (45.5) than the other labs (62). Median values for Shannon Entropy (3.717 vs. 3.430) and Faith Phylogenetic Diversity (3.796 vs. 3.548) are also lower for Genesis, but those values are with one standard deviation of each other for the different sampling times. This is consistent with previous culture-based results suggesting that the environment in cleanrooms tends to select for a core group of organisms capable of surviving under dry, low nutrient, conditions. The presence of the ASV’s mapping to Cupravidus and Thermus in our sequencing blanks and controls suggests that several of the most common organisms in our samples represent contaminants from the reagents used to perform the DNA extractions and sequencing. Further work is needed to identify these contaminants, remove them from our data and recalculate the diversity statistics. This is a systematic error. Therefore, we do not expect removing the sequencing contaminants to change our conclusions. Longer air sample collection times appear to result in slightly higher diversity and do not bias the results towards “hardy” bacteria like spore-formers. Based on these preliminary results we conclude that sampling at least 3,000 liters of air is sufficient to capture the microbial diversity of cleanrooms, and that air samples can also be collected overnight without negatively impacting diversity. These results allow us to be flexible when designing microbial monitoring plans so that they do not interfere with routine lab activity. References: 1. Regberg, A. B. et al. 49th Lunar and Planetary Science Conference (2018). 2. The United States Pharmacopeial Convention. USP General Chapter <1116> (2013). 3. Regberg, A. B., et al. 54th Lunar and Planetary Science Conference (2023). 4. Regberg, A. B. et al. 53rd Lunar and Planetary Science Conference ( 2022). 5. Davis, R. E.,et al. 50th Lunar and Planetary Science Conference (2019). 6. Qiagen. MagAttract® PowerMicrobiome® DNA/RNA EP Kit Handbook. (2018). 7. Walters, W. et al. mSystems 1, (2015). 8. Callahan, B. J. et al. Nat. Methods 13, 581–583 (2016). 9. Hall, M. & Beiko, R. G. Microbiome Analysis: Methods and Protocols113–129 (Springer, 2018). 10. Philipson, C. et al. Bio-Protoc. 7, e2622 (2017).

A. B. Regberg↗

CRITICA: coding region identification tool invoking comparative analysis

Gene recognition is essential to understanding existing and future DNA sequence data. CRITICA (Coding Region Identification Tool Invoking Comparative Analysis) is a suite of programs for identifying likely protein-coding sequences in DNA by combining comparative analysis of DNA sequences with more common noncomparative methods. In the comparative component of the analysis, regions of DNA are aligned with related sequences from the DNA databases; if the translation of the aligned sequences has greater amino acid identity than expected for the observed percentage nucleotide identity, this is interpreted as evidence for coding. CRITICA also incorporates noncomparative information derived from the relative frequencies of hexanucleotides in coding frames versus other contexts (i.e., dicodon bias). The dicodon usage information is derived by iterative analysis of the data, such that CRITICA is not dependent on the existence or accuracy of coding sequence annotations in the databases. This independence makes the method particularly well suited for the analysis of novel genomes. CRITICA was tested by analyzing the available Salmonella typhimurium DNA sequences. Its predictions were compared with the DNA sequence annotations and with the predictions of GenMark. CRITICA proved to be more accurate than GenMark, and moreover, many of its predictions that would seem to be errors instead reflect problems in the sequence databases. The source code of CRITICA is freely available by anonymous FTP (rdp.life.uiuc.edu in/pub/critica) and on the World Wide Web (http:/(/)rdpwww.life.uiuc.edu).

Non-NASA Center↗

DNABERT-S: pioneering species differentiation with species-aware DNA embeddings

SUMMARY: We introduce DNABERT-S, a tailored genome model that develops species-aware embeddings to naturally cluster and segregate DNA sequences of different species in the embedding space. Differentiating species from genomic sequences (i.e. DNA and RNA) is vital yet challenging, since many real-world species remain uncharacterized, lacking known genomes for reference. Embedding-based methods are therefore used to differentiate species in an unsupervised manner. DNABERT-S builds upon a pre-trained genome foundation model named DNABERT-2. To encourage effective embeddings to error-prone long-read DNA sequences, we introduce Manifold Instance Mixup (MI-Mix), a contrastive objective that mixes the hidden representations of DNA sequences at randomly selected layers and trains the model to recognize and differentiate these mixed proportions at the output layer. We further enhance it with the proposed Curriculum Contrastive Learning (C2LR) strategy. Empirical results on 28 diverse datasets show DNABERT-S's effectiveness, especially in realistic label-scarce scenarios. For example, it identifies twice more species from a mixture of unlabeled genomic sequences, doubles the Adjusted Rand Index (ARI) in species clustering, and outperforms the top baseline's performance in 10-shot species classification with just a 2-shot training. AVAILABILITY AND IMPLEMENTATION: Model, codes, and data are publically available at https://github.com/MAGICS-LAB/DNABERT_S.

Zhou, Zhihan↗

Methods for isothermal molecular amplification with nanoparticle-based reactions

The present method of detection involves increasing an amount of analyte molecules by an isothermal molecular amplification approach. In the present approach a starting molecule of interest may be amplified through a reaction it induces with specifically engineered and functionalized particles, namely protected particles A and storage particles B. This reaction may result in a set of output DNA molecules that is larger in number than the input DNA molecules. Thus the reaction between nanoparticles for amplification of a certain DNA sequence (input DNA molecules) may occur when there is a match with a targeted molecule (stored molecules on storage particles B) and if the DNA sequence of the input DNA molecules does not match (partially or completely) the targeted molecule the reaction may not occur. Without a certain molecular input of the input DNA molecule the reaction may not occur.

Gang, Oleg↗

Generalized Levy-walk model for DNA nucleotide sequences

We propose a generalized Levy walk to model fractal landscapes observed in noncoding DNA sequences. We find that this model provides a very close approximation to the empirical data and explains a number of statistical properties of genomic DNA sequences such as the distribution of strand-biased regions (those with an excess of one type of nucleotide) as well as local changes in the slope of the correlation exponent alpha. The generalized Levy-walk model simultaneously accounts for the long-range correlations in noncoding DNA sequences and for the apparently paradoxical finding of long subregions of biased random walks (length lj) within these correlated sequences. In the generalized Levy-walk model, the lj are chosen from a power-law distribution P(lj) varies as lj(-mu). The correlation exponent alpha is related to mu through alpha = 2-mu/2 if 2 < mu < 3. The model is consistent with the finding of "repetitive elements" of variable length interspersed within noncoding DNA.

NASA Discipline Number 14-10↗

Scar-less multi-part DNA assembly design automation

The present invention provides a method of a method of designing an implementation of a DNA assembly. In an exemplary embodiment, the method includes (1) receiving a list of DNA sequence fragments to be assembled together and an order in which to assemble the DNA sequence fragments, (2) designing DNA oligonucleotides (oligos) for each of the DNA sequence fragments, and (3) creating a plan for adding flanking homology sequences to each of the DNA oligos. In an exemplary embodiment, the method includes (1) receiving a list of DNA sequence fragments to be assembled together and an order in which to assemble the DNA sequence fragments, (2) designing DNA oligonucleotides (oligos) for each of the DNA sequence fragments, and (3) creating a plan for adding optimized overhang sequences to each of the DNA oligos.

Hillson, Nathan J.↗

Structural Characterization of Alzheimer DNA Promoter Sequences from the Amyloid Precursor Gene in the Presence of Thioflavin T and Analogs

Understanding DNA-ligand binding interactions requires ligand screening, crystallization, and structure determination. In order to obtain insights into the amyloid peptide precursor (APP) gene-Thioflavin T (ThT) interaction, single crystals of two DNA sequences 5'-GCCCACCACGGC-3' (PDB 8ASK) and d(CCGGGGTACCCCGG) 2 (PDB 8ASH) were grown in the presence of ThT or its analogue 2-((4-(dimethylamino)benzylidene)amino)-3,6-dimethylbenzo[d]thiazol-3-ium iodide (XRB). Both structures were solved by molecular replacement. In the case of 8ASK, the space group was H3 with unit cell dimensions of a = b = 64.49 Å, c = 46.19 Å. Phases were obtained using a model generated by X3DNA. The novel 12-base-pair B-DNA structure did not have extra density for the ThT ligand. The 14-base-pair A-DNA structure with bound ThT analog XRB was isomorphous with previously the obtained apo-DNA structure 5WV7 (space group was P 4 1 2 1 2 with unit cell dimensions a = b = 41.76 Å, c = 88.96 Å). Binding of XRB to DNA slightly changes the DNA's buckle parameters at the CpG regions. Comparison of the two conformations of the XRB molecule: alone and bound to DNA indicates that the binding results from the freedom of rotation of the two aromatic rings.

59 BASIC BIOLOGICAL SCIENCES↗