Engineering Papers⌕ Search

DOE OSTI · 1853183

Optimal Bayesian supervised domain adaptation for RNA sequencing data

Abstract

Abstract Motivation When learning to subtype complex disease based on next-generation sequencing data, the amount of available data is often limited. Recent works have tried to leverage data from other domains to design better predictors in the target domain of interest with varying degrees of success. But they are either limited to the cases requiring the outcome label correspondence across domains or cannot leverage the label information at all. Moreover, the existing methods cannot usually benefit from other information available a priori such as gene interaction networks. Results In this article, we develop a generative optimal Bayesian supervised domain adaptation (OBSDA) model that can integrate RNA sequencing (RNA-Seq) data from different domains along with their labels for improving prediction accuracy in the target domain. Our model can be applied in cases where different domains share the same labels or have different ones. OBSDA is based on a hierarchical Bayesian negative binomial model with parameter factorization, for which the optimal predictor can be derived by marginalization of likelihood over the posterior of the parameters. We first provide an efficient Gibbs sampler for parameter inference in OBSDA. Then, we leverage the gene-gene network prior information and construct an informed and flexible variational family to infer the posterior distributions of model parameters. Comprehensive experiments on real-world RNA-Seq data demonstrate the superior performance of OBSDA, in terms of accuracy in identifying cancer subtypes by utilizing data from different domains. Moreover, we show that by taking advantage of the prior network information we can further improve the performance. Availability and implementation The source code for implementations of OBSDA and SI-OBSDA are available at the following link. https://github.com/SHBLK/BSDA. Supplementary information Supplementary data are available at Bioinformatics online.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Boluki, Shahin, Qian, Xiaoning, Dougherty, Edward R., Gorodkin, Jan. 2021-04-05. Optimal Bayesian supervised domain adaptation for RNA sequencing data. https://doi.org/10.1093/bioinformatics%2Fbtab228

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related reports

Matrix Metalloproteinases as Candidate Antigenic Determinants for Anti‐Tumor Autoantibodies in Human Ovarian Cancer: A Post Hoc Analysis

Circulating antibodies in patients with cancer can facilitate the identification of accessible epitopes on autoantigens expressed by tumors. To identify previously unrecognized protein targets in ovarian cancer, we computationally assessed a heptapeptide consensus motif (VPELGHE, flanked by two cysteine residues yielding a cyclic nonapeptide under oxidizing conditions) previously discovered via phage display-based epitope mapping of autoantibodies in patients. Eight proteins associated with ovarian cancer encompass amino acid sequences similar to the consensus motif and were, therefore, considered as candidate native autoantigens. Among these candidate targets, however, matrix metalloproteinase 14 (MMP14) demonstrates gene expression that is both high and negatively correlated with survival in ovarian cancer patient cohorts. MMP14 protein levels are also stable in tumor versus non-tumor tissues. Moreover, the corresponding heptapeptide mimic in MMP14 occurs within an α-helical secondary structural element observed in its catalytic domain. These findings demonstrate that a subset of patient-derived autoantibodies may interact with a previously unknown antigenic epitope found in MMP14 and other MMPs, thereby providing opportunities for the development of new targeted agents.

Biochemistry & Molecular Biology↗

Structures of a synthetic antibody selected against and bound to the C‐terminal domain of Clostridium perfringens enterotoxin

Abstract Clostridium perfringensenterotoxin (CpE) causes cytotoxic gastrointestinal disease in mammalian epithelium by binding membrane protein receptors called claudins. Claudins direct the formation of cell/cell tight junctions through oligomerization and govern the transport of molecules between individual cells. CpE binds claudins through its C‐terminal domain (cCpE) and induces cytotoxicity through its N‐terminal domain. The non‐toxic cCpE is a useful tool to study claudins, tight junctions, and for translational applications, such as increasing the permeability of restrictive tissues like the blood–brain barrier or selective targeting of claudin overexpressing cancers. Conversely, there are no specialized molecular tools to study CpE or cCpE, or to modulate or inhibit their functions. We previously reported the development of synthetic antigen‐binding fragments (sFabs) that bind cCpE, and low‐resolution structures of them bound to claudin/cCpE complexes. Here, we determine high‐resolution structures of sFab COP‐2 bound to cCpE using X‐ray crystallography and cryogenic electron microscopy. The structures and biophysical findings provide the mechanism of COP‐2 binding to cCpE and the molecular determinants driving their interactions. These insights can advance the design of new antibody‐based tools from our COP‐2 scaffold to study or alter cCpE function and give rise to a “Trojan horse” strategy that exploits cCpE's tight junction barrier disrupting function to selectively deliver conjugated therapeutics through normally impermeable tissues.

Biochemistry & Molecular Biology↗

Functional Relevance of CASP16 Nucleic Acid Predictions as Evaluated by Structure Providers

ABSTRACT Accurate biomolecular structure prediction enables the prediction of mutational effects, the speculation of function based on predicted structural homology, the analysis of ligand binding modes, experimental model building, and many other applications. Such algorithms to predict essential functional and structural features remain out of reach for biomolecular complexes containing nucleic acids. Here, we report a quantitative and qualitative evaluation of nucleic acid structures for the CASP16 blind prediction challenge by 12 of the experimental groups who provided nucleic acid targets. Blind predictions accurately model secondary structure and some aspects of tertiary structure, including reasonable global folds for some complex RNAs; however, predictions often lack accuracy in the regions of highest functional importance. All models have inaccuracies in non‐canonical regions where, for example, the nucleic‐acid backbone bends, deviating from an A‐form helix geometry, or a base forms a non‐standard hydrogen bond (not a Watson‐Crick base pair). These bends and non‐canonical interactions are integral to forming functionally important regions such as RNA enzymatic active sites. Additionally, the modeling of conserved and functional interfaces between nucleic acids and ligands, proteins, or other nucleic acids remains poor. For some targets, the experimental structures may not represent the only structure the biomolecular complex occupies in solution or in its functional life cycle, posing a future challenge for the community.

Biochemistry & Molecular Biology↗