Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “protein sequence”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Codon2Vec v1.0

Background: Codon2Vec is an embedding neural network that predicts 'high' or 'low' gene expression directly from the protein-coding sequences. Embedding neural networks are commonly used for natural language processing (NLP) applications. Analogous to how an English sentence is a string of words, a gene can be thought of as a string of codons. Similar to how NLP neural networks model English sentences as a non-random sequence of words, we considered a coding sequence as a non-random non-overlapping array of codons (k-mers of length = 3). Value Proposition: - Codon2Vec achieved a high median AUC-ROC score of 83.8% when trained and applied to transcriptomic data from 300 fungal species - Unlike Codo2Vec, conventional methods predicting for expression based on codon usage rely on a priori knowledge of optimal codons or a set of reference genes. - Unlike Codon2vec, these methods do not account for the effect of codon order on gene expression. - Codon2Vec neural network bypasses the need for artisanal feature selection step that is necessary for traditional machine learning models.

Wint, Rhondene↗

Functional Diversification of Populus FLOWERING LOCUS D-LIKE3 Transcription Factor and Two Paralogs in Shoot Ontogeny, Flowering, and Vegetative Phenology

Both the evolution of tree taxa and whole-genome duplication (WGD) have occurred many times during angiosperm evolution. Transcription factors are preferentially retained following WGD suggesting that functional divergence of duplicates could contribute to traits distinctive to the tree growth habit. We used gain- and loss-of-function transgenics, photoperiod treatments, and circannual expression studies in adult trees to study the diversification of three Populus FLOWERING LOCUS D-LIKE (FDL) genes encoding bZIP transcription factors. Expression patterns and transgenic studies indicate that FDL2.2 promotes flowering and that FDL1 and FDL3 function in different vegetative phenophases. Study of dominant repressor FDL versions indicates that the FDL proteins are partially equivalent in their ability to alter shoot growth. Like its paralogs, FDL3 overexpression delays short day-induced growth cessation, but also induces distinct heterochronic shifts in shoot development—more rapid phytomer initiation and coordinated delay in both leaf expansion and the transition to secondary growth in long days, but not in short days. Our results indicate that both regulatory and protein coding sequence variation contributed to diversification of FDL paralogs that has led to a degree of specialization in multiple developmental processes important for trees and their local adaptation.

59 BASIC BIOLOGICAL SCIENCES↗

Differential transcriptomic alterations in nasal versus lung tissue of acrolein-exposed rats

Introduction: Acrolein is a significant component of anthropogenic and wildfire emissions, as well as cigarette smoke. Although acrolein primarily deposits in the upper respiratory tract upon inhalation, patterns of site-specific injury in nasal versus pulmonary tissues are not well characterized. This assessment is critical in the design of in vitro and in vivo studies performed for assessing health risk of irritant air pollutants. Methods: In this study, male and female Wistar-Kyoto rats were exposed nose-only to air or acrolein. Rats in the acrolein exposure group were exposed to incremental concentrations of acrolein (0, 0.1, 0.316, 1 ppm) for the first 30 min, followed by a 3.5 h exposure at 3.16 ppm. In the first cohort of male and female rats, nasal and bronchoalveolar lavage fluids were analyzed for markers of inflammation, and in a second cohort of males, nasal airway and left lung tissues were used for mRNA sequencing. Results: Protein leakage in nasal airways of acrolein-exposed rats was similar in both sexes; however, inflammatory cells and cytokine increases were more pronounced in males when compared to females. No consistent changes were noted in bronchoalveolar lavage fluid of males or females except for increases in total cells and IL-6. Acrolein-exposed male rats had 452 differentially expressed genes (DEGs) in nasal tissue versus only 95 in the lung. Pathway analysis of DEGs in the nose indicated acute phase response signaling, Nrf2-mediated oxidative stress, unfolded protein response, and other inflammatory pathways, whereas in the lung, xenobiotic metabolism pathways were changed. Genes associated with glucocorticoid and GPCR signaling were also changed in the nose but not in the lung. Discussion: These data provide insights into inhaled acrolein-mediated sex-specific injury/inflammation in the nasal and pulmonary airways. The transcriptional response in the nose reflects acrolein-induced acute oxidative and cytokine signaling changes, which might have implications for upper airway inflammatory disease susceptibility.

Alewel, Devin I.↗

Nanotechnology Based Materials and Devices for Health Care

This viewgraph presentation provides information on trends in NASA nanotechnology research and development, and future biotechnological applications for that nanotechnology. The presentation covers nanoelectronics, nanosensors, and nanomaterials, biomimetics, devices and materials for health care, carbon nanotubes, biosensors for astrobiology, solid-state nanopores for DNA sequencing, and protein nanotubes.

Srivastava, Deepaka↗

Technical advance: stringent control of transgene expression in Arabidopsis thaliana using the Top10 promoter system

We show that the tightly regulated tetracycline-sensitive Top10 promoter system (Weinmann et al. Plant J. 1994, 5, 559-569) is functional in Arabidopsis thaliana. A pure breeding A. thaliana line (JL-tTA/8) was generated which expressed a chimeric fusion of the tetracycline repressor and the activation domain of Herpes simplex virus (tTA), from a single transgenic locus. Plants from this line were crossed with transgenics carrying the ER-targeted green fluorescent protein coding sequence (mGFP5) under control of the Top10 promoter sequence. Progeny from this cross displayed ER-targeted GFP fluorescence throughout the plant, indicating that the tTA-Top10 promoter interaction was functional in A. thaliana. GFP expression was repressed by 100 ng ml-1 tetracycline, an order of magnitude lower than the concentration used previously to repress expression in Nicotiana tabacum. Moreover, the level of GFP expression was controlled by varying the concentration of tetracycline in the medium, allowing a titred regulation of transgenic activity that was previously unavailable in A. thaliana. The kinetics of GFP activity were determined following de-repression of the Top10:mGFP5 transgene, with a visible ER-targeted GFP signal appearing from 24 to 48 h after de-repression.

NASA Discipline Plant Biology↗

The regulation and regulatory role of collagenase in bone

Interstitial collagenase plays an important role in both the normal and pathological remodeling of collagenous extracellular matrices, including skeletal tissues. The enzyme is a member of the family of matrix metalloproteinases. Only one rodent interstitial collagenase has been found but there are two human enzymes, human collagenase-1 and -3, the latter being the homologue of the rat enzyme. In developing rat and mouse bone, collagenase is expressed by hypertrophic chondrocytes, osteoblasts, and osteocytes, a situation that is replicated in a fracture callus. Cultured osteoblasts derived from neonatal rat calvariae show greater amounts of collagenase transcripts late in differentiation. These levels can be regulated by parathyroid hormone (PTH), retinoic acid, and insulin-like growth factors, as well as the degree of matrix mineralization. Much of the work on collagenase in bone has been derived from studies on the rat osteosarcoma cell line, UMR 106-01. All bone-resorbing agents stimulate these cells to produce collagenase mRNA and protein, with PTH being the most potent stimulator. Determination of secreted levels of collagenase has been difficult because UMR cells, normal rat osteoblasts, and rat fibroblasts possess a scavenger receptor that removes the enzyme from the extracellular space, internalizes and degrades it, thus imposing another level of control. PTH can also regulate the abundance of the receptor as well as the expression and synthesis of the enzyme. Regulation of the collagenase gene by PTH appears to involve the cAMP pathway as well as a primary response gene, possibly Fos, which then contributes to induction of the collagenase gene. The rat collagenase gene contains an activator protein-1 sequence that is necessary for basal expression, but other promoter regions may also participate in PTH regulation. Thus, there are many levels of regulation of collagenase in bone perhaps constraining what would otherwise be a rampant enzyme.

NASA Discipline Musculoskeletal↗

Expression of the Acyl-Coenzyme A: Cholesterol Acyltransferase GFP Fusion Protein in Sf21 Insect Cells

The enzyme acyl-coenzyme A:cholesterol acyltransferase (ACAT) is an important contributor to the pathological expression of plaque leading to artherosclerosis n a major health problem. Adequate knowledge of the structure of this protein will enable pharmaceutical companies to design drugs specific to the enzyme. ACAT is a membrane protein located in the endoplasmic reticulum.t The protein has never been purified to homogeneity.T.Y. Chang's laboratory at Dartmouth College provided a 4-kb cDNA clone (K1) coding for a structural gene of the protein. We have modified the gene sequence and inserted the cDNA into the BioGreen His Baculovirus transfer vector. This was successfully expressed in Sf2l insect cells as a GFP-labeled ACAT protein. The advantage to this ACAT-GFP fusion protein (abbreviated GCAT) is that one can easily monitor its expression as a function of GFP excitation at 395 nm and emission at 509 nm. Moreover, the fusion protein GCAT can be detected on Western blots with the use of commercially available GFP antibodies. Antibodies against ACAT are not readily available. The presence of the 6xHis tag in the transfer vector facilitates purification of the recombinant protein since 6xHis fusion proteins bind with high affinity to Ni-NTA agarose. Obtaining highly pure protein in large quantities is essential for subsequent crystallization. The purified GCAT fusion protein can readily be cleaved into distinct GFP and ACAT proteins in the presence of thrombin. Thrombin digests the 6xHis tag linking the two protein sequences. Preliminary experiments have indicated that both GCAT and ACAT are expressed as functional proteins. The ultimate aim is to obtain large quantities of the ACAT protein in pure and functional form appropriate for protein crystal growth. Determining protein structure is the key to the design and development of effective drugs. X-ray analysis requires large homogeneous crystals that are difficult to obtain in the gravity environment of earth. Protein crystals grown in microgravity are often larger and have fewer defects than those grown on earth. The analysis of higher quality space-grown crystals will assist in structure-based drug design. We have successfully grown GCAT-infected Sf21 cells in both adhesion and suspension cultures. Expression levels of GCAT in cell lines such as Sf9 and High Five appear to be reduced. We intend to replicate GCAT expression in all three cell lines using the NASA rotating wall bioreactor which effectively duplicates a microgravity environment. The bioreactor itself could be launched to study the expression of the GFP and GCAT proteins in the actual microgravity environment achieved in orbit.

Mahtani, H. K.↗

TopPICR: A Companion R Package for Top-Down Proteomics Data Analysis

Top-down proteomics is the analysis of proteins in their intact form without proteolysis, thus preserving valuable information about post-translational modifications, isoforms, and proteolytic processing. However, it is still a developing field due to limitations in the instrumentation, difficulties with interpretation of complex mass spectra, and a lack of well-established quantification approaches. TopPIC is one of the popular tools for proteoform identification. Here we extended its capabilities into label-free proteoform quantification by developing a companion R package (TopPICR). Key steps in the TopPICR pipeline include filtering identifications, inferring a minimal set of protein accessions explaining the observed sequences, aligning retention times, recalibrating measured masses, clustering features across datasets, and finally compiling feature intensities using the match-between-runs approach. The output of the pipeline is an MSnSet object which makes downstream data analysis seamlessly compatible with packages from the Bioconductor project. It also provides the capability for visualizing proteoforms within the context of the parent protein sequence. The functionality of TopPICR is demonstrated on top-down LC-MS/MS datasets of 10 human-in-mouse xenografts of luminal and basal breast tumor samples.

59 BASIC BIOLOGICAL SCIENCES↗

Characterization and distribution of a maize cDNA encoding a peptide similar to the catalytic region of second messenger dependent protein kinases

Maize (Zea mays) roots respond to a variety of environmental stimuli which are perceived by a specialized group of cells, the root cap. We are studying the transduction of extracellular signals by roots, particularly the role of protein kinases. Protein phosphorylation by kinases is an important step in many eukaryotic signal transduction pathways. As a first phase of this research we have isolated a cDNA encoding a maize protein similar to fungal and animal protein kinases known to be involved in the transduction of extracellular signals. The deduced sequence of this cDNA encodes a polypeptide containing amino acids corresponding to 33 out of 34 invariant or nearly invariant sequence features characteristic of protein kinase catalytic domains. The maize cDNA gene product is more closely related to the branch of serine/threonine protein kinase catalytic domains composed of the cyclic-nucleotide- and calcium-phospholipid-dependent subfamilies than to other protein kinases. Sequence identity is 35% or more between the deduced maize polypeptide and all members of this branch. The high structural similarity strongly suggests that catalytic activity of the encoded maize protein kinase may be regulated by second messengers, like that of all members of this branch whose regulation has been characterized. Northern hybridization with the maize cDNA clone shows a single 2400 base transcript at roughly similar levels in maize coleoptiles, root meristems, and the zone of root elongation, but the transcript is less abundant in mature leaves. In situ hybridization confirms the presence of the transcript in all regions of primary maize root tissue.

NASA Program Space Biology↗

Inferring the palaeoenvironment of ancient bacteria on the basis of resurrected proteins

Features of the physical environment surrounding an ancestral organism can be inferred by reconstructing sequences of ancient proteins made by those organisms, resurrecting these proteins in the laboratory, and measuring their properties. Here, we resurrect candidate sequences for elongation factors of the Tu family (EF-Tu) found at ancient nodes in the bacterial evolutionary tree, and measure their activities as a function of temperature. The ancient EF-Tu proteins have temperature optima of 55-65 degrees C. This value seems to be robust with respect to uncertainties in the ancestral reconstruction. This suggests that the ancient bacteria that hosted these particular genes were thermophiles, and neither hyperthermophiles nor mesophiles. This conclusion can be compared and contrasted with inferences drawn from an analysis of the lengths of branches in trees joining proteins from contemporary bacteria, the distribution of thermophily in derived bacterial lineages, the inferred G + C content of ancient ribosomal RNA, and the geological record combined with assumptions concerning molecular clocks. The study illustrates the use of experimental palaeobiochemistry and assumptions about deep phylogenetic relationships between bacteria to explore the character of ancient life.

Bacteria/classification/metabolism↗

Structural Models and Sequence Alignment Results of the Rhodospirillum rubrum Proteome

This dataset contains the structural models for the primary transcripts of the Rhodospirillum rubrum proteome as well as sequence alignment results for a subset of the encoded proteins. For each protein, the five models inferred from AlphaFold 2 are provided. The largest pTM-scoring model for each protein was energy minimized; this minimized structure as well as its AlphaFold pickle output file are also provided. This set of structures represent an alternate source of models for the R. rubrum proteome to those available in the AlphaFold Protein Structure Database. For proteins that have been annotated as hypothetical, sequence alignment results from the HHblits and SAdLSA alignment methods are provided. These methods are often more capable to resolve sequence homology than other methods. Therefore, the results from both HHblits and SAdLSA are provided to identify possible homologs for these challenging proteins. Numerous sequence databases are utilized for these alignments. References AlphaFold v2 Multimer: https://doi.org/10.1101/2021.10.04.463034. References HHBlits: https://doi.org/10.1186/s12859-019-3019-7. References SAdLSA: https://doi.org/10.3389/fbinf.2021.689960.

59 BASIC BIOLOGICAL SCIENCES↗

NMPFamsDB: a database of novel protein families from microbial metagenomes and metatranscriptomes

Abstract The Novel Metagenome Protein Families Database (NMPFamsDB) is a database of metagenome- and metatranscriptome-derived protein families, whose members have no hits to proteins of reference genomes or Pfam domains. Each protein family is accompanied by multiple sequence alignments, Hidden Markov Models, taxonomic information, ecosystem and geolocation metadata, sequence and structure predictions, as well as 3D structure models predicted with AlphaFold2. In its current version, NMPFamsDB hosts over 100 000 protein families, each with at least 100 members. The reported protein families significantly expand (more than double) the number of known protein sequence clusters from reference genomes and reveal new insights into their habitat distribution, origins, functions and taxonomy. We expect NMPFamsDB to be a valuable resource for microbial proteome-wide analyses and for further discovery and characterization of novel functions. NMPFamsDB is publicly available in http://www.nmpfamsdb.org/ or https://bib.fleming.gr/NMPFamsDB.

59 BASIC BIOLOGICAL SCIENCES↗

Simultaneous enhancement of multiple functional properties using evolution-informed protein design

Abstract A major challenge in protein design is to augment existing functional proteins with multiple property enhancements. Altering several properties likely necessitates numerous primary sequence changes, and novel methods are needed to accurately predict combinations of mutations that maintain or enhance function. Models of sequence co-variation (e.g., EVcouplings), which leverage extensive information about various protein properties and activities from homologous protein sequences, have proven effective for many applications including structure determination and mutation effect prediction. We apply EVcouplings to computationally design variants of the model protein TEM-1 β -lactamase. Nearly all the 14 experimentally characterized designs were functional, including one with 84 mutations from the nearest natural homolog. The designs also had large increases in thermostability, increased activity on multiple substrates, and nearly identical structure to the wild type enzyme. This study highlights the efficacy of evolutionary models in guiding large sequence alterations to generate functional diversity for protein design applications.

59 BASIC BIOLOGICAL SCIENCES↗

The sequence, and its evolutionary implications, of a Thermococcus celer protein associated with transcription

Through random search, a gene from Thermococcus celer has been identified and sequenced that appears to encode a transcription-associated protein (110 amino acid residues). The sequence has clear homology to approximately the last half of an open reading frame reported previously for Sulfolobus acidocaldarius [Langer, D. & Zillig, W. (1993) Nucleic Acids Res. 21, 2251]. The protein translations of these two archaeal genes in turn are homologs of a small subunit found in eukaryotic RNA polymerase I (A12.2) and the counterpart of this from RNA polymerase II (B12.6). Homology is also seen with the eukaryotic transcription factor TFIIS, but it involves only the terminal 45 amino acids of the archaeal proteins. Evolutionary implications of these homologies are discussed.

Non-NASA Center↗