Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “genomic methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Phylodynamics of SARS-CoV-2 Lineages B.1.1.7, B.1.1.529 and B.1.617.2 in Nigeria Suggests Divergent Evolutionary Trajectories

Background: The early months of the COVID-19 pandemic were characterized by high transmission rates and mortality, compounded by the emergence of multiple SARS-CoV-2 lineages, including Variants of Concern (VOCs). This study investigates the phylodynamic and spatio-temporal trends of VOCs during the peak of the pandemic in Nigeria. Methods: Whole-genome sequencing (WGS) data from three major VOCs circulating in Nigeria, B.1.1.7 (Alpha), B.1.617.2 (Delta), and B.1.1.529 (Omicron), were analyzed using tools such as Nextclade, R Studio v 4.2.3, and BEAST X v 10.5.0. The spatial distribution, evolutionary history, viral ancestral introductions, and geographic dispersal patterns were characterized. Results: Three major lineages following WHO nomenclature were identified: Alpha, Delta, and Omicron. The Delta variant exhibited the widest geographic spread, detected in 14 states, while the Alpha variant was the least distributed, identified in only eight states but present across most epidemiological weeks studied. Evolutionary rates varied slightly, with Alpha exhibiting the slowest rate (2.66 × 10 −4 substitutions/site/year). Viral population analyses showed distinct patterns: Omicron sustained elevated population growth over time, while Delta declined after initial expansion. The earliest Times to Most Recent Common Ancestor (TMRCA) were consistent with the earliest outbreaks of SARS-CoV-2 globally. Geographic transmission analysis indicated a predominant coastal-to-inland spread for all variants, with Omicron showing the most diffuse dispersal, highlighting commercial routes as significant drivers of viral diffusion. Conclusion: The SARS-CoV-2 epidemic in Nigeria was characterized by multiple variant introductions and a dominant coastal-to-inland spread, emphasizing that despite lockdown measures, commercial trade routes played a critical role in viral dissemination. These findings provide insights into pandemic control strategies and future outbreak preparedness.

Nigeria↗

Quantitative analysis of bristle number in Drosophila mutants identifies genes involved in neural development

BACKGROUND: The identification of the function of all genes that contribute to specific biological processes and complex traits is one of the major challenges in the postgenomic era. One approach is to employ forward genetic screens in genetically tractable model organisms. In Drosophila melanogaster, P element-mediated insertional mutagenesis is a versatile tool for the dissection of molecular pathways, and there is an ongoing effort to tag every gene with a P element insertion. However, the vast majority of P element insertion lines are viable and fertile as homozygotes and do not exhibit obvious phenotypic defects, perhaps because of the tendency for P elements to insert 5' of transcription units. Quantitative genetic analysis of subtle effects of P element mutations that have been induced in an isogenic background may be a highly efficient method for functional genome annotation. RESULTS: Here, we have tested the efficacy of this strategy by assessing the extent to which screening for quantitative effects of P elements on sensory bristle number can identify genes affecting neural development. We find that such quantitative screens uncover an unusually large number of genes that are known to function in neural development, as well as genes with yet uncharacterized effects on neural development, and novel loci. CONCLUSIONS: Our findings establish the use of quantitative trait analysis for functional genome annotation through forward genetics. Similar analyses of quantitative effects of P element insertions will facilitate our understanding of the genes affecting many other complex traits in Drosophila.

Non-NASA Center↗

A Monte-Carlo Model for the Formation of Radiation-induced Chromosomal Aberrations

Purpose: To simulate radiation-induced chromosome aberrations in mammalian cells (e.g., rings, translocations, and dicentrics) and to calculate their frequency distributions following exposure to DNA double strand breaks (DSBs) produced by high-LET ions. Methods: The interphase genome was assumed to be comprised of a collection of 2 kbp rigid-block monomers following the random-walk geometry. Additional details for the modeling of chromosomal structure, such as chromosomal domains and chromosomal loops, were included. A radial energy profile for heavy ion tracks was used to simulate the high-LET pattern of induced DSBs. The induced DSB pattern depended on the ion charge and kinetic energy, but always corresponded to the DSB yield of 25 DSBs/cell/Gy. The sum of all energy contributions from Poisson-distributed particle tracks was taken to account for all possible one-track and multi-track effects. The relevant output of the model was DNA fragments produced by DSBs. The DSBs, or breakpoints, were defined by (x, y, z, l) positions, where x, y, z were the Euclidian coordinates of a DSB, and where l was the relative position along the genome. Results: The code was used to carry out Monte Carlo simulations for DSB rejoinings at low doses. The resulting fragments were analyzed to estimate the frequencies of specific types of chromosomal aberrations. Histograms for relative frequencies of chromosomal aberrations and P.D.F.s (probability density functions) of a given aberration type were produced. The relative frequency of dicentrics to rings was compared to empirical data to calibrate rejoining probabilities. Of particular interest was the predicted distribution of ring sizes, irrespective of their frequencies relative to other aberrations. Simulated ring sizes were . 4 kbp, which are far too small to be observed experimentally (i.e., by microscopy) but which, nevertheless, are conjectured to exist. Other aberrations, for example, inversions, translocations, as well as multi-centrics were also recorded. Conclusion: High-LET DNA damage affects the frequencies of chromosomal aberrations. The ratio of rings to dicentrics is correct for the genomic size cut-offs corresponding to available experimental data. The present work predicts a relative abundance of small rings following irradiation by heavy ions.

Ponomarev, Artem L.↗

A review of methods for the reconstruction and analysis of integrated genome-scale models of metabolism and regulation

The current survey aims to describe the main methodologies for extending the reconstruction and analysis of genome-scale metabolic models and phenotype simulation with Flux Balance Analysis mathematical frameworks, via the integration of Transcriptional Regulatory Networks and/or gene expression data. Although the surveyed methods are aimed at improving phenotype simulations obtained from these models, the perspective of reconstructing integrated genome-scale models of metabolism and gene expression for diverse prokaryotes is still an open challenge.

59 BASIC BIOLOGICAL SCIENCES↗

Congruity of genomic and epidemiological data in modelling of local cholera outbreaks

Cholera continues to be a global health threat. Understanding how cholera spreads between locations is fundamental to the rational, evidence-based design of intervention and control efforts. Traditionally, cholera transmission models have used cholera case-count data. More recently, whole-genome sequence data have qualitatively described cholera transmission. Integrating these data streams may provide much more accurate models of cholera spread; however, no systematic analyses have been performed so far to compare traditional case-count models to the phylodynamic models from genomic data for cholera transmission. Here, we use high-fidelity case-count and whole-genome sequencing data from the 1991 to 1998 cholera epidemic in Argentina to directly compare the epidemiological model parameters estimated from these two data sources. We find that phylodynamic methods applied to cholera genomics data provide comparable estimates that are in line with established methods. Our methodology represents a critical step in building a framework for integrating case-count and genomic data sources for cholera epidemiology and other bacterial pathogens.

59 BASIC BIOLOGICAL SCIENCES↗

Next-Generation Sequencing Data from a CUT&RUN Study of R. toruloides IFO0880 Cse4 and Orc1 Binding Sites

Rhodotorula toruloides has been increasingly explored as a host for bioproduction of lipids, fatty acid derivatives and terpenoids. Various genetic tools have been developed, but neither a centromere nor an autonomously replicating sequence (ARS), both necessary elements for stable episomal plasmid maintenance, has yet been reported. In this study, cleavage under targets and release using nuclease (CUT&RUN), a method used for genome-wide mapping of DNA–protein interactions, was used to identify R. toruloides IFO0880 genomic regions associated with the centromeric histone H3 protein Cse4, a marker of centromeric DNA. Fifteen putative centromeres ranging from 8 to 19 kb in length were identified and analyzed, and four were tested for, but did not show, ARS activity. These centromeric sequences contained below average GC content, corresponded to transcriptional cold spots, were primarily nonrepetitive and shared some vestigial transposon-related sequences but otherwise did not show significant sequence conservation. Future efforts to identify an ARS in this yeast can utilize these centromeric DNA sequences to improve the stability of episomal plasmids derived from putative ARS elements.

Genome Engineering↗

Methods and compositions for protection of cells and tissues from computed tomography radiation

Described are methods for preventing or inhibiting genomic instability and in cells affected by diagnostic radiology procedures employing ionizing radiation. Embodiments include methods of preventing or inhibiting genomic instability and in cells affected by computed tomography (CT) radiation. Subjects receiving ionizing radiation may be those persons suspected of having cancer, or cancer patients having received or currently receiving cancer therapy, and or those patients having received previous ionizing radiation, including those who are approaching or have exceeded the recommended total radiation dose for a person.

Grdina, David J.↗

Probing photosynthesis by altering chloroplast proteins (Final Technical Report)

The goal of this project was to explore the effect of altering the carbon fixing enzyme Rubisco in the model C3 plant tobacco on photosynthesis and plant growth and development. We introduced the gene sequences encoding Rubisco from a Limonium species and from a red alga into tobacco and found that the enzyme did not assemble properly, undoubtedly requiring additional assembly factors from the source organism. We then developed a method to make single amino-acid changes into the Rubisco subunit that is encoded by the chloroplast genome. We used the method to investigate the role of particular amino acid residues in kinetic properties of Rubisco. Finally, we improved a bacterial expression system that has allowed assembly of active tobacco Rubisco in E. coli. We used the improved system to investigate the kinetic properties of Rubisco enzymes that contain only one type of Rubisco small subunit along with the large subunit.

59 BASIC BIOLOGICAL SCIENCES↗

Benchmarking blockchain-based gene-drug interaction data sharing methods: A case study from the iDASH 2019 secure genome analysis competition blockchain track

Blockchain distributed ledger technology is just starting to be adopted in genomics and healthcare applications. Despite its increased prevalence in biomedical research applications, skepticism regarding the practicality of blockchain technology for real-world problems is still strong and there are few implementations beyond proof-of-concept. We focus on benchmarking blockchain strategies applied to distributed methods for sharing records of gene-drug interactions. We expect this type of sharing will expedite personalized medicine. We generated gene-drug interaction test datasets using the Clinical Pharmacogenetics Implementation Consortium (CPIC) resource. We developed three blockchain-based methods to share patient records on gene-drug interactions: Query Index, Index Everything, and Dual-Scenario Indexing. We achieved a runtime of about 60 s for importing 4,000 gene-drug interaction records from four sites, and about 0.5 s for a data retrieval query. Our results demonstrated that it is feasible to leverage blockchain as a new platform to share data among institutions.

60 APPLIED LIFE SCIENCES↗

Improved Microbial Community Characterization of 16S rRNA via Metagenome Hybridization Capture Enrichment

Environmental microbial diversity is often investigated from a molecular perspective using 16S ribosomal RNA (rRNA) gene amplicons and shotgun metagenomics. While amplicon methods are fast, low-cost, and have curated reference databases, they can suffer from amplification bias and are limited in genomic scope. In contrast, shotgun metagenomic methods sample more genomic regions with fewer sequence acquisition biases, but are much more expensive (even with moderate sequencing depth) and computationally challenging. Here, we develop a set of 16S rRNA sequence capture baits that offer a potential middle ground with the advantages from both approaches for investigating microbial communities. These baits cover the diversity of all 16S rRNA sequences available in the Greengenes (v. 13.5) database, with no sequence having <78% sequence identity to at least one bait for all segments of 16S. The use of our baits provide comparable results to 16S amplicon libraries and shotgun metagenomic libraries when assigning taxonomic units from 16S sequences within the metagenomic reads. We demonstrate that 16S rRNA capture baits can be used on a range of microbial samples (i.e., mock communities and rodent fecal samples) to increase the proportion of 16S rRNA sequences (average > 400-fold) and decrease analysis time to obtain consistent community assessments. Furthermore, our study reveals that bioinformatic methods used to analyze sequencing data may have a greater influence on estimates of community composition than library preparation method used, likely due in part to the extent and curation of the reference databases considered. Thus, enriching existing aliquots of shotgun metagenomic libraries and obtaining modest numbers of reads from them offers an efficient orthogonal method for assessment of bacterial community composition.

59 BASIC BIOLOGICAL SCIENCES↗

An archaeal genomic signature

Comparisons of complete genome sequences allow the most objective and comprehensive descriptions possible of a lineage's evolution. This communication uses the completed genomes from four major euryarchaeal taxa to define a genomic signature for the Euryarchaeota and, by extension, the Archaea as a whole. The signature is defined in terms of the set of protein-encoding genes found in at least two diverse members of the euryarchaeal taxa that function uniquely within the Archaea; most signature proteins have no recognizable bacterial or eukaryal homologs. By this definition, 351 clusters of signature proteins have been identified. Functions of most proteins in this signature set are currently unknown. At least 70% of the clusters that contain proteins from all the euryarchaeal genomes also have crenarchaeal homologs. This conservative set, which appears refractory to horizontal gene transfer to the Bacteria or the Eukarya, would seem to reflect the significant innovations that were unique and fundamental to the archaeal "design fabric." Genomic protein signature analysis methods may be extended to characterize the evolution of any phylogenetically defined lineage. The complete set of protein clusters for the archaeal genomic signature is presented as supplementary material (see the PNAS web site, www.pnas.org).

Non-NASA Center↗

Phage-based delivery of CRISPR-associated transposases for targeted bacterial editing

Phage λ, a well-characterized temperate phage, has been recently leveraged for bacterial genome editing by selectively delivering base editors into targeted bacterial species. We extend this concept by engineering phage λ to deliver CRISPR-guided transposases, accomplishing large insertions and targeted gene disruptions. To achieve this, we engineered phage λ using homologous recombination paired with Cas13a-based counterselection for precise phage modifications. Initially, we established the utility of Cas13a in phage λ by conducting minimal recoding edits, deletions, and insertions. Subsequently, we scaled up the engineering to embed the comprehensive DNA-editing CRISPR-Cas transposase (DART) system within the phage genome, creating λ-DART phages. These modified λ-DART phages were then employed to infectEscherichia coli, generating CRISPR RNA-guided transposition events in the host genome. Applying our engineered λ-DART phages to monocultures and a mixed bacterial community comprising three genera led to efficient, precise, and specific gene knockouts and insertions in the targetedE. colicells, achieving editing efficiencies surpassing 50% of the population. This research enhances phage-mediated genome editing by enabling efficient in situ gene integrations in bacteria, offering an avenue for further application in microbial community contexts. This scalable method enables flexible microbial genome editing in situ to manipulate the function and composition of diverse ecosystems.

Science & Technology - Other Topics↗

Genomic selection in algae with biphasic lifecycles: A Saccharina latissima (sugar kelp) case study

Introduction Sugar kelp ( Saccharina latissima ) has a biphasic life cycle, allowing selection on both thediploid sporophytes (SPs) and haploid gametophytes (GPs). Methods We trained a genomic selection (GS) model from farm-tested SP phenotypic data and used a mixed-ploidy additive relationship matrix to predict GP breeding values. Topranked GPs were used to make crosses for further farm evaluation. The relationship matrix included 866 individuals: a) founder SPs sampled from the wild; b) progeny GPs from founders; c) Farm-tested SPs crossed from b); and d) progeny GPs from farm-tested SPs. The complete pedigree-based relationship matrix was estimated for all individuals. A subset of founder SPs ( n = 58) and GPs ( n = 276) were genotyped with Diversity Array Technology and whole genome sequencing, respectively. We evaluated GS prediction accuracy via cross validation for SPs tested on farm in 2019 and 2020 using a basic GBLUP model. We also estimated the general combining ability (GCA) and specific combining ability (SCA) variances of parental GPs. A total of 11 yield-related and morphology traits were evaluated. Results The cross validation accuracies for dry weight per meter ( r ranged from 0.16 to 0.35) and wet weight per meter ( r ranged 0.19 to 0.35) were comparable to GS accuracy for yield traits in terrestrial crops. For morphology traits, cross validation accuracy exceeded 0.18 in all scenarios except for blade thickness in the second year. Accuracy in a third validation year (2021) was 0.31 for dry weight per meter over a confirmation set of 87 individuals. Discussion Our findings indicate that progress can be made in sugar kelp breeding by using genomic selection.

59 BASIC BIOLOGICAL SCIENCES↗

GWAS supported by computer vision identifies large numbers of candidate regulators of in planta regeneration in Populus trichocarpa

Plant regeneration is an important dimension of plant propagation and a key step in the production of transgenic plants. However, regeneration capacity varies widely among genotypes and species, the molecular basis of which is largely unknown. Association mapping methods such as genome-wide association studies (GWAS) have long demonstrated abilities to help uncover the genetic basis of trait variation in plants; however, the performance of these methods depends on the accuracy and scale of phenotyping. To enable a large-scale GWAS of in planta callus and shoot regeneration in the model tree Populus, we developed a phenomics workflow involving semantic segmentation to quantify regenerating plant tissues over time. We found that the resulting statistics were of highly non-normal distributions, and thus employed transformations or permutations to avoid violating assumptions of linear models used in GWAS. We report over 200 statistically supported quantitative trait loci (QTLs), with genes encompassing or near to top QTLs including regulators of cell adhesion, stress signaling, and hormone signaling pathways, as well as other diverse functions. Our results encourage models of hormonal signaling during plant regeneration to consider keystone roles of stress-related signaling (e.g. involving jasmonates and salicylic acid), in addition to the auxin and cytokinin pathways commonly considered. The putative regulatory genes and biological processes we identified provide new insights into the biological complexity of plant regeneration, and may serve as new reagents for improving regeneration and transformation of recalcitrant genotypes and species.

59 BASIC BIOLOGICAL SCIENCES↗

Investigating genomic prediction strategies for grain carotenoid traits in a tropical/subtropical maize panel

Abstract Vitamin A deficiency remains prevalent on a global scale, including in regions where maize constitutes a high percentage of human diets. One solution for alleviating this deficiency has been to increase grain concentrations of provitamin A carotenoids in maize (Zea mays ssp. mays L.)—an example of biofortification. The International Maize and Wheat Improvement Center (CIMMYT) developed a Carotenoid Association Mapping panel of 380 inbred lines adapted to tropical and subtropical environments that have varying grain concentrations of provitamin A and other health-beneficial carotenoids. Several major genes have been identified for these traits, 2 of which have particularly been leveraged in marker-assisted selection. This project assesses the predictive ability of several genomic prediction strategies for maize grain carotenoid traits within and between 4 environments in Mexico. Ridge Regression-Best Linear Unbiased Prediction, Elastic Net, and Reproducing Kernel Hilbert Spaces had high predictive abilities for all tested traits (β-carotene, β-cryptoxanthin, provitamin A, lutein, and zeaxanthin) and outperformed Least Absolute Shrinkage and Selection Operator. Furthermore, predictive abilities were higher when using genome-wide markers rather than only the markers proximal to 2 or 13 genes. These findings suggest that genomic prediction models using genome-wide markers (and assuming equal variance of marker effects) are worthwhile for these traits even though key genes have already been identified, especially if breeding for additional grain carotenoid traits alongside β-carotene. Predictive ability was maintained for all traits except lutein in between-environment prediction. The TASSEL (Trait Analysis by aSSociation, Evolution, and Linkage) Genomic Selection plugin performed as well as other more computationally intensive methods for within-environment prediction. The findings observed herein indicate the utility of genomic prediction methods for these traits and could inform their resource-efficient implementation in biofortification breeding programs.

59 BASIC BIOLOGICAL SCIENCES↗

Identification of transcribed sequences in Arabidopsis thaliana by using high-resolution genome tiling arrays

Using a maskless photolithography method, we produced DNA oligonucleotide microarrays with probe sequences tiled throughout the genome of the plant Arabidopsis thaliana. RNA expression was determined for the complete nuclear, mitochondrial, and chloroplast genomes by tiling 5 million 36-mer probes. These probes were hybridized to labeled mRNA isolated from liquid grown T87 cells, an undifferentiated Arabidopsis cell culture line. Transcripts were detected from at least 60% of the nearly 26,330 annotated genes, which included 151 predicted genes that were not identified previously by a similar genome-wide hybridization study on four different cell lines. In comparison with previously published results with 25-mer tiling arrays produced by chromium masking-based photolithography technique, 36-mer oligonucleotide probes were found to be more useful in identifying intron-exon boundaries. Using two-dimensional HPLC tandem mass spectrometry, a small-scale proteomic analysis was performed with the same cells. A large amount of strongly hybridizing RNA was found in regions "antisense" to known genes. Similarity of antisense activities between the 25-mer and 36-mer data sets suggests that it is a reproducible and inherent property of the experiments. Transcription activities were also detected for many of the intergenic regions and the small RNAs, including tRNA, small nuclear RNA, small nucleolar RNA, and microRNA. Expression of tRNAs correlates with genome-wide amino acid usage.

Arabidopsis/genetics↗

BayFlux: A Bayesian method to quantify metabolic Fluxes and their uncertainty at the genome scale

Metabolic fluxes, the number of metabolites traversing each biochemical reaction in a cell per unit time, are crucial for assessing and understanding cell function. 13 C Metabolic Flux Analysis ( 13 C MFA) is considered to be the gold standard for measuring metabolic fluxes. 13 C MFA typically works by leveraging extracellular exchange fluxes as well as data from 13 C labeling experiments to calculate the flux profile which best fit the data for a small, central carbon, metabolic model. However, the nonlinear nature of the 13 C MFA fitting procedure means that several flux profiles fit the experimental data within the experimental error, and traditional optimization methods offer only a partial or skewed picture, especially in “non-gaussian” situations where multiple very distinct flux regions fit the data equally well. Here, we present a method for flux space sampling through Bayesian inference (BayFlux), that identifies the full distribution of fluxes compatible with experimental data for a comprehensive genome-scale model. This Bayesian approach allows us to accurately quantify uncertainty in calculated fluxes. We also find that, surprisingly, the genome-scale model of metabolism produces narrower flux distributions (reduced uncertainty) than the small core metabolic models traditionally used in 13 C MFA. The different results for some reactions when using genome-scale models vs core metabolic models advise caution in assuming strong inferences from 13 C MFA since the results may depend significantly on the completeness of the model used. Based on BayFlux, we developed and evaluated novel methods (P- 13 C MOMA and P- 13 C ROOM) to predict the biological results of a gene knockout, that improve on the traditional MOMA and ROOM methods by quantifying prediction uncertainty.

59 BASIC BIOLOGICAL SCIENCES↗

Inference of Chromosome-Length Haplotypes Using Genomic Data of Three or a Few More Single Gametes

Compared with genomic data of individual markers, haplotype data provide higher resolution for DNA variants, advancing our knowledge in genetics and evolution. Although many computational and experimental phasing methods have been developed for analyzing diploid genomes, it remains challenging to reconstruct chromosome-scale haplotypes at low cost, which constrains the utility of this valuable genetic resource. Gamete cells, the natural packaging of haploid complements, are ideal materials for phasing entire chromosomes because the majority of the haplotypic allele combinations has been preserved. Therefore, compared with the current diploid-based phasing methods, using haploid genomic data of single gametes may substantially reduce the complexity in inferring the donor’s chromosomal haplotypes. In this study, we developed the first easy-to-use R package, Hapi, for inferring chromosome-length haplotypes of individual diploid genomes with only a few gametes. Hapi outperformed other phasing methods when analyzing both simulated and real single gamete cell sequencing data sets. The results also suggested that chromosome-scale haplotypes may be inferred by using as few as three gametes, which has pushed the boundary to its possible limit. The single gamete cell sequencing technology allied with the cost-effective Hapi method will make large-scale haplotype-based genetic studies feasible and affordable, promoting the use of haplotype data in a wide range of research.

59 BASIC BIOLOGICAL SCIENCES↗