Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Sequence Homology”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Homology and the optimization of DNA sequence data

Three methods of nucleotide character analysis are discussed. Their implications for molecular sequence homology and phylogenetic analysis are compared. The criterion of inter-data set congruence, both character based and topological, are applied to two data sets to elucidate and potentially discriminate among these parsimony-based ideas. c2001 The Willi Hennig Society.

Non-NASA Center

Orthologs, paralogs and genome comparisons

During the past decade, ancient gene duplications were recognized as one of the main forces in the generation of diverse gene families and the creation of new functional capabilities. New tools developed to search data banks for homologous sequences, and an increased availability of reliable three-dimensional structural information led to the recognition that proteins with diverse functions can belong to the same superfamily. Analyses of the evolution of these superfamilies promises to provide insights into early evolution but are complicated by several important evolutionary processes. Horizontal transfer of genes can lead to a vertical spread of innovations among organisms, therefore finding a certain property in some descendants of an ancestor does not guarantee that it was present in that ancestor. Complete or partial gene conversion between duplicated genes can yield phylogenetic trees with several, apparently independent gene duplications, suggesting an often surprising parallelism in the evolution of independent lineages. Additionally, the breakup of domains within a protein and the fusion of domains into multifunctional proteins makes the delineation of superfamilies a task that remains difficult to automate.

Non-NASA Center

Kinetic Induction of Oat Shoot Pulvinus Invertase mRNA by Gravistimulation and Partial cDNA Cloning by the Polymerase Chain Reaction

An asymmetric (top vs. bottom halves of pulvini) induction of invertase mRNA by gravistimulation was analyzed in oat shoot pulvini. Total RNA and poly(A)(+) RNA, isolated from oat pulvini, and two oli-gonucleotide primers, corresponding to two conserved amino acid sequences (NDPNG and WECPD) found in invertase from other species, were used for the polymerase chain reaction (PCR). A partial length cDNA (550 bp) was obtained and characterized. A 62% nucleotide sequence homology and 58% deduced amino acid sequence homology, as compared to beta-fructosidase of carrot cell wall, was found. Northern blot analysis showed that there was an obviously transient induction of invertase mRNA by gravistimulation in the oat pulvinus system. The mRNA was rapidly induced to a maximum level at 1 hour after gravistimulation treatment and gradually decreased afterwards. The mRNA level in the bottom half of the oat pulvinus was significantly higher than that in the top half of the pulvinus tissue. The kinetic induction of invertase mRNA was consistent with the transient accumulation of invertase activity during the graviresponse of the pulvinus. This indicates that the expression of the invertase gene(s) could be regulated by gravistimulation at the transcriptional level. Southern blot analysis showed that there were two to three genomic DNA fragments which hybridized with the partial-length invertase cDNA.

Wu, Liu-Lai

Preliminary crystallographic examination of a novel fungal lysozyme from Chalaropsis

The lysozyme from the fungus of the Chalaropsis species has been crystallized. This lysozyme displays no sequence homology with avian, phage, or mammalian lysozymes, however, preliminary studies indicate significant sequence homology with the bacterial lysozyme from Streptomyces. Both enzymes are unusual in possessing beta-1,4-N-acetylmuramidase and beta-1,4-N,6-O-diacetylmuramidase activity. The crystals grow from solutions of ammonium sulfate during growth periods from several months to a year. The space group is P2(1)2(1)2(1) with a = 34.0 A, b = 42.6 A, c = 122.1 A. Preliminary data indicate that there is 1 molecule/asymmetric unit.

Carter, Daniel C.

Bottom-Up Simulation, Reconstruction, and Quantification of Macromolecule Sequences from Experimental Polymerizations

Motivated by the canonical sequence–structure–function paradigm, tools to characterize chemical patterning in natural biomacromolecules, from proteins to nucleic acids, have grown exponentially in recent years. However, analogous strategies for synthetic macromolecules remain in nascent stages, complicated by sequence polydispersity and analytical limitations. To address this, we have developed a comprehensive and open-source Python package, PRISM (polymer rate insights and sequence modeling), an end-to-end workflow that provides a path from experimental kinetics measurements to quantitative and qualitative metrics for describing chemical patterning in stochastic polymers. First, a numerical integration strategy was constructed to simulate and fit experimental data from reversible addition–fragmentation chain transfer (RAFT) polymerization kinetics, enabling the facile estimation of relevant reactivity ratios. These ratios were then used in a mechanism-specific stochastic kinetic simulation strategy to simulate sequence ensembles corresponding to model systems spanning experimental copolymers, classes of statistical polymers (e.g., alternating, block, and gradient), and multiblock copolymers. Lastly, inspired by sequence homology metrics from bioinformatics, we introduce visualization strategies and quantitative metrics to facilitate comparisons of different sequence ensembles. As the sequence–structure–function paradigm becomes increasingly central in de novo design of synthetic macromolecules, this toolkit provides a first step toward accurate and representative sequence description and featurization.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

PRIME: An evaluation framework for protein representation inference and generalization in viral mutation space

Background Protein language models (PLMs) have revolutionized protein fitness prediction, yet their application to rapidly evolving viral pathogens is often confounded by extreme sequence homology. This homology leads to “data leakage” in standard random validation splits, yielding inflated performance metrics that fail to translate into real-world biosurveillance utility. Results We present Protein Representation Inference for Mutation Evaluation (PRIME), a framework that integrates domain-specific fine-tuning with a rigorous position-stratified validation protocol to evaluate viral threats. Using a dataset of 347,432 SARS-CoV-2 receptor binding domain (RBD) sequences, we demonstrate that while random training data split yields deceptive R 2 values (> 0.90), they fail to generalize to novel mutational sites. By benchmarking models up to 650 M parameters, we show that domain-specific fine-tuning of the ESM-C 600 M model with correctly stratified data provides an initial demonstration of predictive signal for binding affinity and expression at unseen mutational sites of binding affinity and expression on unseen sites (R 2 ~0.23), a significant advancement over base foundation models which exhibit no predictive power (R 2 <0). PRIME’s embedding-based clustering identified 3.03% of bat coronavirus sequences as candidates for further experimental prioritization based on their functional similarity to human-infective strains in embedding space, offering a perspective complementary to traditional phylogenetic methods. Conclusion PRIME establishes a new benchmark for the application of PLMs in pathogen surveillance. Our findings demonstrate that state-of-the-art models and fine-tuning, when paired with stratified validation, provide biologically meaningful insights into pathogen evolution and zoonotic risk.

59 BASIC BIOLOGICAL SCIENCES

Establishing homologies in protein sequences

Computer-based statistical techniques used to determine homologies between proteins occurring in different species are reviewed. The technique is based on comparison of two protein sequences, either by relating all segments of a given length in one sequence to all segments of the second or by finding the best alignment of the two sequences. Approaches discussed include selection using printed tabulations, identification of very similar sequences, and computer searches of a database. The use of the SEARCH, RELATE, and ALIGN programs (Dayhoff, 1979) is explained; sample data are presented in graphs, diagrams, and tables and the construction of scoring matrices is considered.

Dayhoff, M. O.

Identification of the prooxidant site of human ceruloplasmin: a model for oxidative damage by copper bound to protein surfaces

Free transition metal ions oxidize lipids and lipoproteins in vitro; however, recent evidence suggests that free metal ion-independent mechanisms are more likely in vivo. We have shown previously that human ceruloplasmin (Cp), a serum protein containing seven Cu atoms, induces low density lipoprotein oxidation in vitro and that the activity depends on the presence of a single, chelatable Cu atom. We here use biochemical and molecular approaches to determine the site responsible for Cp prooxidant activity. Experiments with the His-specific reagent diethylpyrocarbonate (DEPC) showed that one or more His residues was specifically required. Quantitative [14C]DEPC binding studies indicated the importance of a single His residue because only one was exposed upon removal of the prooxidant Cu. Plasmin digestion of [14C]DEPC-treated Cp (and N-terminal sequence analysis of the fragments) showed that the critical His was in a 17-kDa region containing four His residues in the second major sequence homology domain of Cp. A full length human Cp cDNA was modified by site-directed mutagenesis to give His-to-Ala substitutions at each of the four positions and was transfected into COS-7 cells, and low density lipoprotein oxidation was measured. The prooxidant site was localized to a region containing His426 because CpH426A almost completely lacked prooxidant activity whereas the other mutants expressed normal activity. These observations support the hypothesis that Cu bound at specific sites on protein surfaces can cause oxidative damage to macromolecules in their environment. Cp may serve as a model protein for understanding mechanisms of oxidant damage by copper-containing (or -binding) proteins such as Cu, Zn superoxide dismutase, and amyloid precursor protein.

NASA Discipline Regulatory Physiology

Regiochemical control of monolignol radical coupling: a new paradigm for lignin and lignan biosynthesis

BACKGROUND: Although the lignins and lignans, both monolignol-derived coupling products, account for nearly 30% of the organic carbon circulating in the biosphere, the biosynthetic mechanism of their formation has been poorly understood. The prevailing view has been that lignins and lignans are produced by random free-radical polymerization and coupling, respectively. This view is challenged, mechanistically, by the recent discovery of dirigent proteins that precisely determine both the regiochemical and stereoselective outcome of monolignol radical coupling. RESULTS: To understand further the regulation and control of monolignol coupling, leading to both lignan and lignin formation, we sought to clone the first genes encoding dirigent proteins from several species. The encoding genes, described here, have no sequence homology with any other protein of known function. When expressed in a heterologous system, the recombinant protein was able to confer strict regiochemical and stereochemical control on monolignol free-radical coupling. The expression in plants of dirigent proteins and proposed dirigent protein arrays in developing xylem and in other lignified tissues indicates roles for these proteins in both lignan formation and lignification. CONCLUSIONS: The first understanding of regiochemical and stereochemical control of monolignol coupling in lignan biosynthesis has been established via the participation of a new class of dirigent proteins. Immunological studies have also implicated the involvement of potential corresponding arrays of dirigent protein sites in controlling lignin biopolymer assembly.

NASA Discipline Plant Biology

Csa-19, a radiation-responsive human gene, identified by an unbiased two-gel cDNA library screening method in human cancer cells

A novel polymerase chain reaction (PCR)-based method was used to identify candidate genes whose expression is altered in cancer cells by ionizing radiation. Transcriptional induction of randomly selected genes in control versus irradiated human HL60 cells was compared. Among several complementary DNA (cDNA) clones recovered by this approach, one cDNA clone (CL68-5) was downregulated in X-irradiated HL60 cells but unaffected by 12-O-tetradecanoyl phorbol-13-acetate, forskolin, or cyclosporin-A. DNA sequencing of the CL68-5 cDNA revealed 100% nucleotide sequence homology to the reported human Csa-19 gene. Northern blot analysis of RNA from control and irradiated cells revealed the expression of a single 0.7-kilobase (kb) messenger RNA (mRNA) transcript. This 0.7-kb Csa-19 mRNA transcript was also expressed in a variety of human adult and corresponding fetal normal tissues. Moreover, when the effect of X- or fission neutron-irradiation on Csa-19 mRNA was compared in cultured human cells differing in p53 gene status (p53-/- versus p53+/+), downregulation of Csa-19 by X-rays or fission neutrons was similar in p53-wild type and p53-null cell lines. Our results provide the first known example of a radiation-responsive gene in human cancer cells whose expression is not associated with p53, adenylate cyclase or protein kinase C.

NASA Discipline Radiation Health

An alternative pocket for binding the N‐degrons by the UBR1 and UBR2 ubiquitin E3 ligases

The UBR family of ubiquitin ligases binds to N-termini of their targets (known as N-degron) to induce their ubiquitination and degradation via a conserved domain known as UBR-box. UBR1 and UBR2 share the highest sequence homology among the family, and substantial structural studies were previously performed for substrate binding by the UBR-boxes of UBR1 and UBR2. Here, we describe a new pocket in the UBR-boxes of UBR1 and UBR2 for binding the second residues of N-degrons through determining five co-crystal structures of the UBR-boxes with various N-degron peptides. Together with binding affinities measured by fluorescence polarization, we show that the two highly homologous UBR-boxes can interact with the second residue of an N-degron differently. In addition, the UBR-boxes undergo different conformational changes when binding N-degrons. Furthermore, we demonstrate that the sidechain of the third amino acid of an N-degron has no contribution to binding the UBR-boxes. These findings represent a new conceptual advancement for the UBR E3 ligases and the new insights described here can be leveraged for developing their selective ligands for research and potential therapies.

N-end rule

Evolutionary trajectory of transcription factors and selection of targets for metabolic engineering

Transcription factors (TFs) provide potentially powerful tools for plant metabolic engineering as they often control multiple genes in a metabolic pathway. However, selecting the best TF for a particular pathway has been challenging, and the selection often relies significantly on phylogenetic relationships. Here, we offer examples where evolutionary relationships have facilitated the selection of the suitable TFs, alongside situations where such relationships are misleading from the perspective of metabolic engineering. We argue that the evolutionary trajectory of a particular TF might be a better indicator than protein sequence homology alone in helping decide the best targets for plant metabolic engineering efforts. This article is part of the theme issue ‘The evolution of plant metabolism’.

Life Sciences & Biomedicine - Other Topics

Genetics of Flooding Tolerance in an F 2 Miscanthus sacchariflorus ssp. lutarioriparius × M. sinensis Population

Miscanthus is a warm-season, perennial grass cultivated as a feedstock for bioenergy and bioproducts. M. sacchariflorus ssp. lutarioriparius has high yield potential and is well-adapted to seasonal flooding, but little is known about the genetics of this adaptation. We conducted a quantitative trait locus (QTL) analysis on a population of 332 diploid Miscanthus ×giganteus (Mxg) F2s derived from an initial cross between diploid M. sacchariflorus ssp. lutarioriparius ‘PF30022’ and diploid M. sinensis ‘PMS-014’, followed by intermating 50 F 1 s. Using tanks in a greenhouse to assess the effects of partial submergence on actively growing plants, we compared an aerobic soil control to a 6-week flood treatment. The study's primary objectives were to (1) identify QTL for flooding tolerance in Miscanthus , (2) identify candidate genes and (3) compare ethylene response factors in Miscanthus with those in rice and Arabidopsis , sorghum and maize for binding site sequence homology and synteny, especially those associated with flooding tolerance. In total, 10 QTL and 66 candidate genes for partial submergence tolerance were identified (including many for ethylene signalling), a first report for Miscanthus . Notably, none of the Miscanthus candidates were orthologs of rice Sub1A, SK1 or SK2 , yet the ‘PF30022’ parent exhibited a snorkeling phenotype, indicating convergent evolution. This study will facilitate breeding of climate-resiliant Mxg.

abiotic stress tolerance

Molecular and structural characterization of a Bacillus cereus strain producing an anthrax-like capsule

Bacillus cereus is a ubiquitous Gram-positive, spore-forming, rod-shaped saprophytic bacterium, occasionally reported to cause food-borne illnesses. However, instances of B. cereus strains harboring anthrax toxin and capsule genes have elevated certain strains as formidable pathogens and biothreats. This study focuses on the genomic analysis and the structural characterization of capsular material produced by the virulent B. cereus PATH2418 strain, isolated from the wound of a traumatic open fracture patient. The genome was sequenced using Nanopore MinION sequencing, revealing a chromosome of 5,270,283 bp and three plasmids. One plasmid, pATH1, was found to encode an operon for the biosynthesis of a bacterial capsule. This operon had sequence homology to the Bacillus anthracis capBCADE operon, which encodes the poly-γ-D-glutamate (PDGA) capsule. The capsule production in B. cereus PATH2418 was influenced by temperature and CO 2 levels. Structural analysis of the capsular material using a combined approach of nuclear magnetic resonance (NMR) and high-performance liquid chromatography (HPLC) techniques confirmed the presence of a high-molecular-weight poly-γ-glutamate capsule, with an enantiomeric composition of approximately 67% D-glutamic acid and 33% L-glutamic acid, matching that of B. anthracis.

Bacillus cereus

A measure of the denseness of a phylogenetic network

An objective measure of phylogenetic denseness is developed to examine various phylogenetic criteria: alpha- and beta-hemoglobin, myoglobin, cytochrome c, and the parvalbumin family. Attention is given to the number of nucleotide replacements separating homologous sequences, and to the topology of the network (in other words, to the qualitative nature of the network as defined by how closely the studied species are related). Applications include quantitative comparisons of species origin, relation, and rates of evolution.

Holmquist, R.

Phytanyl-glycerol ethers and squalenes in the archaebacterium Methanobacterium thermoautotrophicum

Gas chromatographic and mass- and infrared-spectrometric techniques are used to assay the lipids of a thermophilic chemolithotroph, Methanobacterium thermoautotrophicum. Of the chloroform-soluble lipids, 79% are polar and 21% non-polar. Attention is given to the detection of squalene and hydrosqualene derivatives, which, coupled with 16S r-RNA sequence homologies, indicate that the extreme halophiles and the methanogens share a common ancestor.

Tornabene, T. G.

Molecular Basis of the Increase in Invertase Activity Elicited by Gravistimulation of Oat-Shoot Pulvini

An asymmetric (top vs. bottom) increase in invertase activity is elicited by gravistimulation in oatshoot pulvini starting within 3h after treatment. In order to analyze the regulation of invertase gene expression in this system, we examined the effect of gravistimulation on invertase mRNA induction. Total RNA and poly(A)(+)RNA, isolated from oat pulvini, and two oligonucleotide primers, corresponding to two conserved amino-acid sequences (NDPNG and WECPD) found in invertase from other species, were used for the Polymerase Chain Reaction (PCR). A partial-length cDNA (550 base pairs) was obtained and characterized. There was a 52 % deduced amino-acid sequence homology to that of carrot beta-fructosi- dase and a 48 % homology to that of tomato invertase. Northern blot analysis showed that there was an obvious transient accumulation of invertase mRNA elicited by gravistimulation of oat pulvini. The mRNA was rapidly induced to a maximum level at 1h following gravistimulation treatment and gradually decreased afterwards. The mRNA level in the bottom half of the oat pulvinus was significantly higher (five-fold) than that in the top half of the pulvinus tissue. The induction of invertase mRNA was consistent with the transient enhancement of invertase activity during the graviresponse of the pulvinus. These data indicate that the expression of the invertase gene(s) could be regulated by gravistimulation at the transcriptional and/or translational levels. Southern blot analysis showed that there were four genomic DNA fragments hybridized to the invertase cDNA. This suggests that an invertase gene family may exist in oat plants.

Wu, Liu-Lai

gyrB as a phylogenetic discriminator for members of the Bacillus anthracis-cereus-thuringiensis group

Bacillus anthracis, the causative agent of the human disease anthrax, Bacillus cereus, a food-borne pathogen capable of causing human illness, and Bacillus thuringiensis, a well-characterized insecticidal toxin producer, all cluster together within a very tight clade (B. cereus group) phylogenetically and are indistinguishable from one another via 16S rDNA sequence analysis. As new pathogens are continually emerging, it is imperative to devise a system capable of rapidly and accurately differentiating closely related, yet phenotypically distinct species. Although the gyrB gene has proven useful in discriminating closely related species, its sequence analysis has not yet been validated by DNA:DNA hybridization, the taxonomically accepted "gold standard". We phylogenetically characterized the gyrB sequences of various species and serotypes encompassed in the "B. cereus group," including lab strains and environmental isolates. Results were compared to those obtained from analyses of phenotypic characteristics, 16S rDNA sequence, DNA:DNA hybridization, and virulence factors. The gyrB gene proved more highly differential than 16S, while, at the same time, as analytical as costly and laborious DNA:DNA hybridization techniques in differentiating species within the B. cereus group.

Phylogeny