Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data lineage”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Viral Envelope Evolution in Simian–HIV-Infected Neonate and Adult-Dam Pairs of Rhesus Macaques

We recently demonstrated that Simian–HIV (SHIV)-infected neonate rhesus macaques (RMs) generated heterologous HIV-1 neutralizing antibodies (NAbs) with broadly-NAb (bNAb) characteristics at a higher frequency compared with their corresponding dam. Here, we characterized genetic diversity in Env sequences from four neonate or adult/dam RM pairs: in two pairs, neonate and dam RMs made heterologous HIV-1 NAbs; in one pair, neither the neonate nor the dam made heterologous HIV-1 NAbs; and in another pair, only the neonate made heterologous HIV-1 NAbs. Phylogenetic and sequence diversity analyses of longitudinal Envs revealed that a higher genetic diversity, within the host and away from the infecting SHIV strain, was correlated with heterologous HIV-1 NAb development. We identified 22 Env variable sites, of which 9 were associated with heterologous HIV-1 NAb development; 3/9 sites had mutations previously linked to HIV-1 Env bNAb development. These data suggested that viral diversity drives heterologous HIV-1 NAb development, and the faster accumulation of viral diversity in neonate RMs may be a potential mechanism underlying bNAb induction in pediatric populations. Moreover, these data may inform candidate Env immunogens to guide precursor B cells to bNAb status via vaccination by the Env-based selection of bNAb lineage members with the appropriate mutations associated with neutralization breadth.

60 APPLIED LIFE SCIENCES↗

Deep Green: Structural and Functional Genomic Characterization of Conserved Unannotated Green Lineage Proteins

Our overall objective is to improve and increase functional and structural predictions for a growing number of plant proteins of unknown structure and function (the Deep Green proteins), and to make our predictive data useful and accessible to the larger research community. The project is divided into five major objectives which include 1) Assembly and curation of Deep Green candidate protein sets; 2) in silico structural and functional predictions and network analyses; 3) assembly and validation of reverse genetic resources in Chlamydomonas reinhardtii (Chlamydomonas); 4) high throughput functional genomics characterization and prioritization in Chlamydomonas; and 5) structural validation of selected candidates and functional validation in two important reference plant species, Arabidopsis thaliana (Arabidopsis) and Setaria viridis (Setaria).

59 BASIC BIOLOGICAL SCIENCES↗

Deep Green Unannotated Protein Structures

The Deep Green list is based on the identification and curation of conserved unannotated proteins in three green lineage (Viridiplantae) model organisms; Arabidopsis thaliana, Chlamydomonas reinhardtii, and Setaria viridis. Preliminary characterization of Deep Green proteins and genes was done using various informatics tools and published data sets and is presented in Knoshaug, Sun, et al., 2023, submitted. The structures of these unannotated proteins were also predicted using AlphaFold (Jumper et al., 2021). The data deposited here are the AlphaFold structural predictions having the highest pLDDT score and thus identified as the best folded structure (ranked_0). These data enable others to do in-depth structural characterizations to aid in functional characterization leading to deeper understanding of plant biology. References: Jumper, J., Evans, R., Pritzel, A., Green, T., Figurnov, M., Ronneberger, O., Tunyasuvunakool, K., Bates, R., Žídek, A., Potapenko, A., Bridgland, A., Meyer, C., Kohl, S. A. A., Ballard, A. J., Cowie, A., Romera-Paredes, B., Nikolov, S., Jain, R., Adler, J., Back, T., Petersen, S., Reiman, D., Clancy, E., Zielinski, M., Steinegger, M., Pacholska, M., Berghammer, T., Bodenstein, S., Silver, D., Vinyals, O., Senior, A. W., Kavukcuoglu, K., Kohli, P. and Hassabis, D. (2021) Highly accurate protein structure prediction with AlphaFold. Nature, 596:583-589. Knoshaug, E. P., Sun, P., Nag, A., Nguyen, H., Mattoon, E. M., Zhang, N., Liu, J., Chen, C., Cheng, J., Zhang, R., St. John, P., and Umen, J. (submitted) Identification and preliminary characterization of conserved uncharacterized proteins from Chlamydomonas reinhardtii, Arabidopsis thaliana, and Setaria viridis.

09 BIOMASS FUELS↗

Genome-resolved biogeography of Phaeocystales, cosmopolitan bloom-forming algae

Phaeocystales, comprising the genus Phaeocystis and an uncharacterized sister lineage, are nanoplanktonic haptophytes widespread in the global ocean. Several species form mucilaginous colonies and influence key biogeochemical cycles, yet their underlying diversity and ecological strategies remain underexplored. Here, we present new genomic data from 13 strains, including three high-quality reference genomes (N50 > 30 kbp), and integrate previous metagenome-assembled genomes to resolve a robust phylogeny. Divergence timing of P. antarctica aligns with Miocene cooling and Southern Ocean isolation. Genomic traits reveal metabolic flexibility, including mixotrophic nitrogen acquisition in temperate waters and gene expansions linked to polar nutrient adaptation. Concordantly, transcriptomic comparisons between temperate and polar Phaeocystis suggest Southern Ocean populations experience iron and B12 limitation. We also identify signatures of horizontal gene transfer and endogenous giant virus/virophage insertions. Together, these findings highlight Phaeocystales as an ecologically versatile and geographically widespread lineage shaped by evolutionary innovation and adaptation to contrasting environmental stressors.

Füssy, Zoltán↗

Mechanisms of SARS-CoV-2 neutralization by shark variable new antigen receptors elucidated through X-ray crystallography

Single-domain Variable New Antigen Receptors (VNARs) from the immune system of sharks are the smallest naturally occurring binding domains found in nature. Possessing flexible paratopes that can recognize protein motifs inaccessible to classical antibodies, VNARs have yet to be exploited for the development of SARS-CoV-2 therapeutics. Here, we detail the identification of a series of VNARs from a VNAR phage display library screened against the SARS-CoV-2 receptor binding domain (RBD). The ability of the VNARs to neutralize pseudotype and authentic live SARS-CoV-2 virus rivalled or exceeded that of full-length immunoglobulins and other single-domain antibodies. Crystallographic analysis of two VNARs found that they recognized separate epitopes on the RBD and had distinctly different mechanisms of virus neutralization unique to VNARs. Structural and biochemical data suggest that VNARs would be effective therapeutic agents against emerging SARS-CoV-2 mutants, including the Delta variant, and coronaviruses across multiple phylogenetic lineages. This study highlights the utility of VNARs as effective therapeutics against coronaviruses and may serve as a critical milestone for nearing a paradigm shift of the greater biologic landscape.

60 APPLIED LIFE SCIENCES↗

Metagenome-assembled genome extraction and analysis from microbiomes using KBase

Uncultivated Bacteria and Archaea account for the vast majority of species on Earth, but obtaining their genomes directly from the environment, using shotgun sequencing, has only become possible recently. In order to realize the hope of capturing Earth’s microbial genetic complement and to facilitate the investigation of the functional roles of specific lineages in a given ecosystem, technologies that accelerate the recovery of high-quality genomes are necessary. We present a series of analysis steps and data products for the extraction of high-quality metagenome-assembled genomes (MAGs) from microbiomes using the U.S. Department of Energy Systems Biology Knowledgebase (KBase) platform (http://www.kbase.us/). Overall, these steps take about a day to obtain extracted genomes when starting from smaller environmental shotgun read libraries, or up to about a week from larger libraries. In KBase, the process is end-to-end, allowing a user to go from the initial sequencing reads all the way through to MAGs, which can then be analyzed with other KBase capabilities such as phylogenetic placement, functional assignment, metabolic modeling, pangenome functional profiling, RNA-Seq and others. While portions of such capabilities are available individually from other resources, the combination of the intuitive usability, data interoperability and integration of tools in a freely available computational resource makes KBase a powerful platform for obtaining MAGs from microbiomes. While this workflow offers tools for each of the key steps in the genome extraction process, it also provides a scaffold that can be easily extended with additional MAG recovery and analysis tools, via the KBase software development kit (SDK).

59 BASIC BIOLOGICAL SCIENCES↗

Evolution and Antigenic Advancement of N2 Neuraminidase of Swine Influenza A Viruses Circulating in the United States following Two Separate Introductions from Human Seasonal Viruses

Two separate introductions of human seasonal N2 neuraminidase genes were sustained in U.S. swine since 1998 (N2-98) and 2002 (N2-02). Herein, we characterized the antigenic evolution of the N2 of swine influenza A virus (IAV) across 2 decades following each introduction. The N2-98 and N2-02 expanded in genetic diversity, with two statistically supported monophyletic clades within each lineage. To assess antigenic drift in swine N2 following the human-to-swine spillover events, we generated a panel of swine N2 antisera against representative N2 and quantified the antigenic distance between wild-type viruses using enzyme-linked lectin assay and antigenic cartography. The antigenic distance between swine and human N2 was smallest between human N2 circulating at the time of each introduction and the archetypal swine N2. However, sustained circulation and evolution in swine of the two N2 lineages resulted in significant antigenic drift, and the N2-98 and N2-02 swine N2 lineages were antigenically distinct. Although intralineage antigenic diversity was observed, the magnitude of antigenic drift did not consistently correlate with the observed genetic differences. These data represent the first quantification of the antigenic diversity of neuraminidase of IAV in swine and demonstrated significant antigenic drift from contemporary human seasonal strains as well as antigenic variation among N2 detected in swine. These data suggest that antigenic mismatch may occur between circulating swine IAV and vaccine strains. Consequently, consideration of the diversity of N2 in swine IAV for vaccine selection may likely result in more effective control and aid public health initiatives for pandemic preparedness.

59 BASIC BIOLOGICAL SCIENCES↗

A Survey of Bacterial Microcompartment Distribution in the Human Microbiome

Bacterial microcompartments (BMCs) are protein-based organelles that expand the metabolic potential of many bacteria by sequestering segments of enzymatic pathways in a selectively permeable protein shell. Sixty-eight different types/subtypes of BMCs have been bioinformatically identified based on the encapsulated enzymes and shell proteins encoded in genomic loci. BMCs are found across bacterial phyla. The organisms that contain them, rather than strictly correlating with specific lineages, tend to reflect the metabolic landscape of the environmental niches they occupy. From our recent comprehensive bioinformatic survey of BMCs found in genome sequence data, we find many in members of the human microbiome. Here we survey the distribution of BMCs in the different biotopes of the human body. Given their amenability to be horizontally transferred and bioengineered they hold promise as metabolic modules that could be used to probiotically alter microbiomes or treat dysbiosis.

59 BASIC BIOLOGICAL SCIENCES↗

Evolution of hematopoiesis: Three members of the PU.1 transcription factor family in a cartilaginous fish, Raja eglanteria

T lymphocytes and B lymphocytes are present in jawed vertebrates, including cartilaginous fishes, but not in jawless vertebrates or invertebrates. The origins of these lineages may be understood in terms of evolutionary changes in the structure and regulation of transcription factors that control lymphocyte development, such as PU.1. The identification and characterization of three members of the PU.1 family of transcription factors in a cartilaginous fish, Raja eglanteria, are described here. Two of these genes are orthologs of mammalian PU.1 and Spi-C, respectively, whereas the third gene, Spi-D, is a different family member. In addition, a PU.1-like gene has been identified in a jawless vertebrate, Petromyzon marinus (sea lamprey). Both DNA-binding and transactivation domains are highly conserved between mammalian and skate PU.1, in marked contrast to lamprey Spi, in which similarity is evident only in the DNA-binding domain. Phylogenetic analysis of sequence data suggests that the appearance of Spi-C may predate the divergence of the jawed and jawless vertebrates and that Spi-D arose before the divergence of the cartilaginous fish from the lineage leading to the mammals. The tissue-specific expression patterns of skate PU.1 and Spi-C suggest that these genes share regulatory as well as structural properties with their mammalian orthologs.

NASA Discipline Evolutionary Biology↗

Identification of hidden N4-like viruses and their interactions with hosts

The N4-like viruses, which were recently assigned to the novel viral family Schitoviridae in 2021, belong to a podoviral-like viral lineage and possess conserved genomic characteristics and a unique replication mechanism. Despite their significance, our understanding of N4-like viruses is primarily based on viral isolates. To address this knowledge gap, this study has established a comprehensive N4-like viral data sets comprising 342 high-quality N4-like viruses/proviruses (144 viral isolates, 158 uncultured viruses, and 40 integrated N4-like proviruses). These viruses were classified into 97 subfamilies (89 of which are newly identified), 148 genera (100 of which are newly identified), and 253 species (177 of which are newly identified). The study reveals that N4-like viruses inhibit the polar region, oligotrophic open oceans, and the human gut, where they infect various bacterial lineages, such as Alpha/Beta/Gamma/Epsilon-proteobacteria in the Proteobacteria phylum. Although N4-like viral endogenization appears to be prevalent in Proteobacteria, it has also been observed in Firmicutes. Additionally, the phylogenetic analysis has identified evolutionary divergence within the hallmark genes of N4-like viruses, indicating a complex origin of the different conserved parts of viral genomes. Moreover, 1,101 putative auxiliary metabolic genes (AMGs) were identified in the N4-like viral pan-proteome, which mainly participate in nucleotide and cofactor/vitamin metabolisms. Of these AMGs, 27 were found to be associated with virulence, suggesting their potential involvement in the spread of bacterial pathogenicity. The findings of this study are significant, as N4-like viruses represent a unique viral lineage with a distinct replication mechanism and a conserved core genome. This work has resulted in a comprehensive global map of the entire N4-like viral lineage, including information on their distribution in different biomes, evolutionary divergence, genomic diversity, and the potential for viral-mediated host metabolic reprogramming. As such, this work significantly contributes to our understanding of the ecological function and viral-host interactions of bacteriophages.

60 APPLIED LIFE SCIENCES↗

A genoscape–network model for conservation prioritization in a migratory bird

Migratory animals are declining worldwide and coordinated conservation efforts are needed to reverse current trends. We devised a novel genoscape-network model that combines genetic analyses with species distribution modeling and demographic data to overcome challenges with conceptualizing alternative risk factors in migratory species across their full annual cycle. We applied our method to the long distance, Neotropical migratory bird, Wilson's Warbler (Cardellina pusilla). Despite a lack of data from some wintering locations, we demonstrated how the results can be used to help prioritize conservation of breeding and wintering areas. For example, we showed that when genetic, demographic, and network modeling results were considered together it became clear that conservation recommendations will differ depending on whether the goal is to preserve unique genetic lineages or the largest number of birds per unit area. More specifically, if preservation of genetic lineages is the goal, then limited resources should be focused on preserving habitat in the California Sierra, Basin Rockies, or Coastal California, where the 3 most vulnerable genetic lineages breed, or in western Mexico, where 2 of the 3 most vulnerable lineages overwinter. Alternatively, if preservation of the largest number of individuals per unit area is the goal, then limited conservation dollars should be placed in the Pacific Northwest or Central America, where densities are estimated to be the highest. Overall, our results demonstrated the utility of adopting a genetically based network model for integrating multiple types of data across vast geographic scales and better inform conservation decision-making for migratory animals.

60 APPLIED LIFE SCIENCES↗

Loss of Cadherin-11 in pancreatic ductal adenocarcinoma alters tumor-immune microenvironment

Pancreatic ductal adenocarcinoma (PDAC) is one of the top five deadliest forms of cancer with very few treatment options. The 5-year survival rate for PDAC is 10% following diagnosis. Cadherin 11 (Cdh11), a cell-to-cell adhesion molecule, has been suggested to promote tumor growth and immunosuppression in PDAC, and Cdh11 inhibition significantly extended survival in mice with PDAC. However, the mechanisms by which Cdh11 deficiency influences PDAC progression and anti-tumor immune responses have yet to be fully elucidated. To investigate Cdh11-deficiency induced changes in PDAC tumor microenvironment (TME), we crossed p48-Cre; LSL-Kras G12D/+ ; LSLTrp53 R172H/+ (KPC) mice with Cdh11 +/- mice and performed single-cell RNA sequencing (scRNA-seq) of the non-immune (CD45 - ) and immune (CD45 + ) compartment of KPC tumor-bearing Cdh11 proficient (KPC-Cdh11 +/+ ) and Cdh11 deficient (KPC-Cdh11 +/- ) mice. Our analysis showed that Cdh11 is expressed primarily in cancer-associated fibroblasts (CAFs) and at low levels in epithelial cells undergoing epithelial-to-mesenchymal transition (EMT). Cdh11 deficiency altered the molecular profile of CAFs, leading to a decrease in the expression of myofibroblast markers such as Acta2 and Tagln and cytokines such as Il6, Il33 and Midkine (Mdk). We also observed a significant decrease in the presence of monocytes/macrophages and neutrophils in KPC-Cdh11 +/- tumors while the proportion of T cells was increased. Additionally, myeloid lineage cells from Cdh11-deficient tumors had reduced expression of immunosuppressive cytokines that have previously been shown to play a role in immune suppression. In summary, our data suggests that Cdh11 deficiency significantly alters the fibroblast and immune microenvironments and contributes to the reduction of immunosuppressive cytokines, leading to an increase in anti-tumor immunity and enhanced survival.

59 BASIC BIOLOGICAL SCIENCES↗

Benchmark datasets for SARS-CoV-2 surveillance bioinformatics

Severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2), the cause of coronavirus disease 2019 (COVID-19), has spread globally and is being surveilled with an international genome sequencing effort. Surveillance consists of sample acquisition, library preparation, and whole genome sequencing. This has necessitated a classification scheme detailing Variants of Concern (VOC) and Variants of Interest (VOI), and the rapid expansion of bioinformatics tools for sequence analysis. These bioinformatic tools are means for major actionable results: maintaining quality assurance and checks, defining population structure, performing genomic epidemiology, and inferring lineage to allow reliable and actionable identification and classification. Additionally, the pandemic has required public health laboratories to reach high throughput proficiency in sequencing library preparation and downstream data analysis rapidly. However, both processes can be limited by a lack of a standardized sequence dataset. We identified six SARS-CoV-2 sequence datasets from recent publications, public databases and internal resources. In addition, we created a method to mine public databases to identify representative genomes for these datasets. Using this novel method, we identified several genomes as either VOI/VOC representatives or non-VOI/VOC representatives. To describe each dataset, we utilized a previously published datasets format, which describes accession information and whole dataset information. Additionally, a script from the same publication has been enhanced to download and verify all data from this study.

60 APPLIED LIFE SCIENCES↗

IUE observations of central stars

IUE satellite data on sixty galactic planetary nebulae (PN) and three PNs in the Magellanic clouds are examined to establish a mass distribution among the central star types. An evolutionary lineage was determined for the observed central stars, based on UV magnitudes, demonstrating that central stars in optically thin nebulae have a narrow distribution around 0.58 solar mass, whereas stars in optically thick nebulae exhibited the highest masses of the sample, implying that highest mass stars in PN are the most difficult to detect. No definitive correlation was found between the mass of an object and its spectral type.

Heap, S. R.↗

Archaeal phylogeny: reexamination of the phylogenetic position of Archaeoglobus fulgidus in light of certain composition-induced artifacts

A major and too little recognized source of artifact in phylogenetic analysis of molecular sequence data is compositional difference among sequences. The problem becomes particularly acute when alignments contain ribosomal RNAs from both mesophilic and thermophilic species. Among prokaryotes the latter are considerably higher in G + C content than the former, which often results in artificial clustering of thermophilic lineages and their being placed artificially deep in phylogenetic trees. In this communication we review archaeal phylogeny in the light of this consideration, focusing in particular on the phylogenetic position of the sulfate reducing species Archaeoglobus fulgidus, using both 16S rRNA and 23S rRNA sequences. The analysis shows clearly that the previously reported deep branching of the A. fulgidus lineage (very near the base of the euryarchaeal side of the archaeal tree) is incorrect, and that the lineage actually groups with a previously recognized unit that comprises the Methanomicrobiales and extreme halophiles.

NASA Discipline Exobiology↗

Genomic variation across Chinook salmon populations reveals effects of a duplication on migration alleles and supports fine scale structure

Abstract The distribution of ecotypic variation in natural populations is influenced by neutral and adaptive evolutionary forces that are challenging to disentangle. This study provides a high‐resolution portrait of genomic variation in Chinook salmon ( Oncorhynchus tshawytscha ) with emphasis on a region of major effect for ecotypic variation in migration timing. With a filtered data set of ~13 million single nucleotide polymorphisms (SNPs) from low‐coverage whole genome resequencing of 53 populations (3566 barcoded individuals), we contrasted patterns of genomic structure within and among major lineages and examined the extent of a selective sweep at a major effect region underlying migration timing (GREB1L/ROCK1). Neutral variation provided support for fine‐scale structure of populations, while allele frequency variation in GREB1L/ROCK1 was highly correlated with mean return timing for early and late migrating populations within each of the lineages ( r 2 = .58–.95; p < .001). However, the extent of selection within the genomic region controlling migration timing was much narrower in one lineage (interior stream‐type) compared to the other two major lineages, which corresponded to the breadth of phenotypic variation in migration timing observed among lineages. Evidence of a duplicated block within GREB1L/ROCK1 may be responsible for reduced recombination in this portion of the genome and contributes to phenotypic variation within and across lineages. Lastly, SNP positions across GREB1L/ROCK1 were assessed for their utility in discriminating migration timing among lineages, and we recommend multiple markers nearest the duplication to provide highest accuracy in conservation applications such as those that aim to protect early migrating Chinook salmon. These results highlight the need to investigate variation throughout the genome and the effects of structural variants on ecologically relevant phenotypic variation in natural species.

Horn, Rebekah L.↗

Regulation of bacterial stringent response by an evolutionarily conserved ribosomal protein L11 methylation

Lysine and arginine methylation is an important regulator of enzyme activity and transcription in eukaryotes. However, little is known about this covalent modification in bacteria. In this work, we investigated the role of methylation in bacteria. By reanalyzing a large phyloproteomics data set from 48 bacterial strains representing six phyla, we found that almost a quarter of the bacterial proteome is methylated. Many of these methylated proteins are conserved across diverse bacterial lineages, including those involved in central carbon metabolism and translation. Among the proteins with the most conserved methylation sites is ribosomal protein L11 (bL11). bL11 methylation has been a mystery for five decades, as the deletion of its methyltransferase PrmA causes no cell growth defects. Comparative proteomics analysis combined with inorganic polyphosphate and guanosine tetra/pentaphosphate assays of the ΔprmA mutant in Escherichia coli revealed that bL11 methylation is important for stringent response signaling. In the stationary phase, we found that the ΔprmA mutant has impaired guanosine tetra/pentaphosphate production. This leads to a reduction in inorganic polyphosphate levels, accumulation of RNA and ribosomal proteins, and an abnormal polysome profile. Overall, our investigation demonstrates that the evolutionarily conserved bL11 methylation is important for stringent response signaling and ribosomal activity regulation and turnover.

59 BASIC BIOLOGICAL SCIENCES↗

Circularization of 23S rRNA but not 16S rRNA within archaeal ribosomes

Background Processing of archaeal 16S and 23S rRNAs is believed to involve excision of individual rRNAs from polycistronic precursors, circularization of excised rRNAs, and re-linearization before the incorporation into ribosomes. However, all the knowledge is derived from several isolated species, leaving open the possibility that different processes may occur in other archaeal groups. Results Here, we investigate rRNAs from diverse and mostly uncultivated archaea. Sequencing of total cellular RNA from eight phylum-level lineages indicates that archaeal circular 23S rRNA transcript abundances vastly exceed those of linear counterparts, and linear versions are often undetectable. As the majority of rRNAs derive from mature ribosomes, the data suggest that ribosomes contain circular 23S rRNAs. Thus, we directly sequence RNA extracted from isolated ribosomes of a model archaeon, Methanosarcina acetivorans, and confirm that the 23S rRNAs in the ribosomes are circular. Structural modeling places the 5′ and 3′ ends of the linear precursors of archaeal 23S rRNAs in close proximity to form a GNRA tetraloop (in which N is A, C, G, or U and R is A or G), consistent with their existence as circular molecules. We also confirm the existence of circular 16S rRNA intermediates in transcriptomes of most archaea, yet a circular form is not evident in some distinct archaeal groups, suggesting that certain archaea do not circularize 16S rRNA during processing. Conclusions Our findings uncover unexpected variations in the processing required to generate mature rRNAs and the conformation of functional molecules in archaeal ribosomes.

Archaea↗