Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Supplementary Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Optimal Bayesian supervised domain adaptation for RNA sequencing data

Abstract Motivation When learning to subtype complex disease based on next-generation sequencing data, the amount of available data is often limited. Recent works have tried to leverage data from other domains to design better predictors in the target domain of interest with varying degrees of success. But they are either limited to the cases requiring the outcome label correspondence across domains or cannot leverage the label information at all. Moreover, the existing methods cannot usually benefit from other information available a priori such as gene interaction networks. Results In this article, we develop a generative optimal Bayesian supervised domain adaptation (OBSDA) model that can integrate RNA sequencing (RNA-Seq) data from different domains along with their labels for improving prediction accuracy in the target domain. Our model can be applied in cases where different domains share the same labels or have different ones. OBSDA is based on a hierarchical Bayesian negative binomial model with parameter factorization, for which the optimal predictor can be derived by marginalization of likelihood over the posterior of the parameters. We first provide an efficient Gibbs sampler for parameter inference in OBSDA. Then, we leverage the gene-gene network prior information and construct an informed and flexible variational family to infer the posterior distributions of model parameters. Comprehensive experiments on real-world RNA-Seq data demonstrate the superior performance of OBSDA, in terms of accuracy in identifying cancer subtypes by utilizing data from different domains. Moreover, we show that by taking advantage of the prior network information we can further improve the performance. Availability and implementation The source code for implementations of OBSDA and SI-OBSDA are available at the following link. https://github.com/SHBLK/BSDA. Supplementary information Supplementary data are available at Bioinformatics online.

Biochemistry & Molecular Biology↗

MOSAIC: a joint modeling methodology for combined circadian and non-circadian analysis of multi-omics data

Abstract Motivation Circadian rhythms are approximately 24-h endogenous cycles that control many biological functions. To identify these rhythms, biological samples are taken over circadian time and analyzed using a single omics type, such as transcriptomics or proteomics. By comparing data from these single omics approaches, it has been shown that transcriptional rhythms are not necessarily conserved at the protein level, implying extensive circadian post-transcriptional regulation. However, as proteomics methods are known to be noisier than transcriptomic methods, this suggests that previously identified arrhythmic proteins with rhythmic transcripts could have been missed due to noise and may not be due to post-transcriptional regulation. Results To determine if one can use information from less-noisy transcriptomic data to inform rhythms in more-noisy proteomic data, and thus more accurately identify rhythms in the proteome, we have created the Multi-Omics Selection with Amplitude Independent Criteria (MOSAIC) application. MOSAIC combines model selection and joint modeling of multiple omics types to recover significant circadian and non-circadian trends. Using both synthetic data and proteomic data from Neurospora crassa, we showed that MOSAIC accurately recovers circadian rhythms at higher rates in not only the proteome but the transcriptome as well, outperforming existing methods for rhythm identification. In addition, by quantifying non-circadian trends in addition to circadian trends in data, our methodology allowed for the recognition of the diversity of circadian regulation as compared to non-circadian regulation. Availability and implementation MOSAIC’s full interface is available at https://github.com/delosh653/MOSAIC. An R package for this functionality, mosaic.find, can be downloaded at https://CRAN.R-project.org/package=mosaic.find. Supplementary information Supplementary data are available at Bioinformatics online.

De los Santos, Hannah↗

Complete collision data set for electrons scattering on molecular hydrogen and its isotopologues: I. Fully vibrationally-resolved electronic excitation of H 2 $X^1Σ^+_g)$

Here, we present a comprehensive set of vibrationally-resolved cross sections for electron-impact electronic excitation of molecular hydrogen suitable for implementation in collisional-radiative models. The adiabatic-nuclei molecular convergent close-coupling method is used to calculate cross sections for excitation of all bound vibrational levels and dissociative excitation of the B 1 Σ u + , C 1 Π u , E F 1 Σ g + , B ′ 1 Σ u + , G K 1 Σ g + , I 1 Π g , J 1 Δ g , D 1 Π u , H 1 Σ g + , b 3 Σ u + , c 3 Π u , a 3 Σ g + , e 3 Σ u + , d 3 Π u , h 3 Σ g + , g 3 Σ g + , i 3 Π g , and j 3 Δ g electronic states from all –14 bound vibrational levels of the ground electronic ($\text{X}$ $^1&#x3A3^+_g$) state. The data set consists of cross sections from threshold to 500 eV for over 5000 transitions, representing all possible electronic and vibrational transitions between the state and the –3 singlet and triplet states (where refers to the united-atoms-limit principle quantum number). The cross sections are presented in graphical form and provided as both numerical values and analytic fit functions in supplementary data files.

74 ATOMIC AND MOLECULAR PHYSICS↗

pKPDB: a protein data bank extension database of p Ka and pI theoretical values

Abstract Summary pKa values of ionizable residues and isoelectric points of proteins provide valuable local and global insights about their structure and function. These properties can be estimated with reasonably good accuracy using Poisson–Boltzmann and Monte Carlo calculations at a considerable computational cost (from some minutes to several hours). pKPDB is a database of over 12 M theoretical pKa values calculated over 120k protein structures deposited in the Protein Data Bank. By providing precomputed pKa and pI values, users can retrieve results instantaneously for their protein(s) of interest while also saving countless hours and resources that would be spent on repeated calculations. Furthermore, there is an ever-growing imbalance between experimental pKa and pI values and the number of resolved structures. This database will complement the experimental and computational data already available and can also provide crucial information regarding buried residues that are under-represented in experimental measurements. Availability and implementation Gzipped csv files containing p Ka and isoelectric point values can be downloaded from https://pypka.org/pKPDB. To query a single PDB code please use the PypKa free server at https://pypka.org. The pKPDB source code can be found at https://github.com/mms-fcul/pKPDB. Supplementary information Supplementary data are available at Bioinformatics online.

Reis, Pedro B. P. S. (ORCID:0000000335636239)↗

EDGE COVID-19: a web platform to generate submission-ready genomes from SARS-CoV-2 sequencing efforts

Abstract Summary Genomics has become an essential technology for surveilling emerging infectious disease outbreaks. A range of technologies and strategies for pathogen genome enrichment and sequencing are being used by laboratories worldwide, together with different and sometimes ad hoc, analytical procedures for generating genome sequences. A fully integrated analytical process for raw sequence to consensus genome determination, suited to outbreaks such as the ongoing COVID-19 pandemic, is critical to provide a solid genomic basis for epidemiological analyses and well-informed decision making. We have developed a web-based platform and integrated bioinformatic workflows that help to provide consistent high-quality analysis of SARS-CoV-2 sequencing data generated with either the Illumina or Oxford Nanopore Technologies (ONT). Using an intuitive web-based interface, this workflow automates data quality control, SARS-CoV-2 reference-based genome variant and consensus calling, lineage determination and provides the ability to submit the consensus sequence and necessary metadata to GenBank, GISAID and INSDC raw data repositories. We tested workflow usability using real world data and validated the accuracy of variant and lineage analysis using several test datasets, and further performed detailed comparisons with results from the COVID-19 Galaxy Project workflow. Our analyses indicate that EC-19 workflows generate high-quality SARS-CoV-2 genomes. Finally, we share a perspective on patterns and impact observed with Illumina versus ONT technologies on workflow congruence and differences. Availability and implementation https://edge-covid19.edgebioinformatics.org, and https://github.com/LANL-Bioinformatics/EDGE/tree/SARS-CoV2. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

Poisson hurdle model-based method for clustering microbiome features

Abstract Motivation High-throughput sequencing technologies have greatly facilitated microbiome research and have generated a large volume of microbiome data with the potential to answer key questions regarding microbiome assembly, structure and function. Cluster analysis aims to group features that behave similarly across treatments, and such grouping helps to highlight the functional relationships among features and may provide biological insights into microbiome networks. However, clustering microbiome data are challenging due to the sparsity and high dimensionality. Results We propose a model-based clustering method based on Poisson hurdle models for sparse microbiome count data. We describe an expectation–maximization algorithm and a modified version using simulated annealing to conduct the cluster analysis. Moreover, we provide algorithms for initialization and choosing the number of clusters. Simulation results demonstrate that our proposed methods provide better clustering results than alternative methods under a variety of settings. We also apply the proposed method to a sorghum rhizosphere microbiome dataset that results in interesting biological findings. Availability and implementation R package is freely available for download at https://cran.r-project.org/package=PHclust. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

Reactome and the Gene Ontology: digital convergence of data resources

Abstract Motivation Gene Ontology Causal Activity Models (GO-CAMs) assemble individual associations of gene products with cellular components, molecular functions and biological processes into causally linked activity flow models. Pathway databases such as the Reactome Knowledgebase create detailed molecular process descriptions of reactions and assemble them, based on sharing of entities between individual reactions into pathway descriptions. Results To convert the rich content of Reactome into GO-CAMs, we have developed a software tool, Pathways2GO, to convert the entire set of normal human Reactome pathways into GO-CAMs. This conversion yields standard GO annotations from Reactome content and supports enhanced quality control for both Reactome and GO, yielding a nearly seamless conversion between these two resources for the bioinformatics community. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

Complete collision data set for electrons scattering on molecular hydrogen and its isotopologues: IV. Vibrationally-resolved ionization of the ground and excited electronic states

Here, we present a comprehensive set of vibrationally-resolved cross sections for electron-impact ionization of molecular hydrogen and its isotopologues (H 2 , D 2 , T 2 , HD, HT, and DT) in both the ground and excited electronic states. We apply the adiabatic-nuclei molecular convergent close-coupling (MCCC) method to calculate cross sections from threshold to 1000 eV for ionization of the ground and excited vibrational levels of the X 1 Σ$^{+}_{g}$, B 1 Σ$^{+}_{u}$, C 1 π u , EF 1 Σ$^{+}_{g}$, a 3 Σ$^{+}_{g}$, and c 3 π u electronic states, representing all states with united-atoms-limit principle quantum number n=1–2. The cross sections are presented in graphical form and provided as both numerical values and analytic fit functions in supplementary data files. The data can also be downloaded from the MCCC database at mccc-db.org.

74 ATOMIC AND MOLECULAR PHYSICS↗

An open-source high-content analysis workflow for CFTR function measurements using the forskolin-induced swelling assay

Abstract Motivation The forskolin-induced swelling (FIS) assay has become the preferential assay to predict the efficacy of approved and investigational CFTR-modulating drugs for individuals with cystic fibrosis (CF). Currently, no standardized quantification method of FIS data exists thereby hampering inter-laboratory reproducibility. Results We developed a complete open-source workflow for standardized high-content analysis of CFTR function measurements in intestinal organoids using raw microscopy images as input. The workflow includes tools for (i) file and metadata handling; (ii) image quantification and (iii) statistical analysis. Our workflow reproduced results generated by published proprietary analysis protocols and enables standardized CFTR function measurements in CF organoids. Availability and implementation All workflow components are open-source and freely available: the htmrenamer R package for file handling https://github.com/hmbotelho/htmrenamer; CellProfiler and ImageJ analysis scripts/pipelines https://github.com/hmbotelho/FIS_image_analysis; the Organoid Analyst application for statistical analysis https://github.com/hmbotelho/organoid_analyst; detailed usage instructions and a demonstration dataset https://github.com/hmbotelho/FIS_analysis. Distributed under GPL v3.0. Supplementary information Supplementary data are available at Bioinformatics online.

Hagemeijer, Marne C.↗

A catalogue of cataclysmic variables from 20 yr of the Sloan Digital Sky Survey with new classifications, periods, trends, and oddities

ABSTRACT We present a catalogue of 507 cataclysmic variables (CVs) observed in SDSS I to IV including 70 new classifications collated from multiple archival data sets. This represents the largest sample of CVs with high-quality and homogeneous optical spectroscopy. We have used this sample to derive unbiased space densities and period distributions for the major sub-types of CVs. We also report on some peculiar CVs, period bouncers and also CVs exhibiting large changes in accretion rates. We report 70 new CVs, 59 new periods, 178 unpublished spectra, and 262 new or updated classifications. From the SDSS spectroscopy, we also identified 18 systems incorrectly identified as CVs in the literature. We discuss the observed properties of 13 peculiar CVS, and we identify a small set of eight CVs that defy the standard classification scheme. We use this sample to investigate the distribution of different CV sub-types, and we estimate their individual space densities, as well as that of the entire CV population. The SDSS I to IV sample includes 14 period bounce CVs or candidates. We discuss the variability of CVs across the Hertzsprung–Russell diagram, highlighting selection biases of variability-based CV detection. Finally, we searched for, and found eight tertiary companions to the SDSS CVs. We anticipate that this catalogue and the extensive material included in the Supplementary Data will be useful for a range of observational population studies of CVs.

79 ASTRONOMY AND ASTROPHYSICS↗

A Statistician’s Overview of Physics-Informed Neural Networks for Spatio-Temporal Data

The recent success of deep neural network models with physical constraints (so-called, Physics-Informed Neural Networks, PINNs) has led to renewed interest in the incorporation of mechanistic information in predictive models. Statisticians and others have long been interested in this problem, which has led to several practical and innovative solutions dating back decades. In this overview, we focus on the problem of data-driven prediction and inference of dynamic spatio-temporal processes that include mechanistic information, such as would be available from partial differential equations, with a strong focus on the quantification of uncertainty associated with data, process, and parameters. Here, we give a brief review of several paradigms and focus our attention on Bayesian implementations given they naturally accommodate uncertainty quantification. We then show that it is straight-forward to include the Bayesian PINN (B-PINN) within the Bayesian hierarchical model (BHM) framework that has long been considered for modeling dynamic spatio-temporal processes. Such a BHM-PINN is illustrated via a simulation study in which a latent nonlinear Burgers’ equation PDE governs the dynamics of Poisson distributed spatio-temporal data. Supplementary materials for this article are available online, including a standardized description of the materials available for reproducing the work.

Bayesian↗

2D radial-azimuthal particle-in-cell benchmark for E × B discharges

In this paper we propose a representative simulation test-case of E × B discharges accounting for plasma wall interactions with the presence of both the electron cyclotron drift instability and the modified-two-stream-instability. Seven independently developed particle-in-cell (PIC) codes have simulated this benchmark case, with the same specified conditions. The characteristics of the different codes and computing times are given. Here, results show that both instabilities were captured in a similar fashion and good agreement between the different PIC codes is reported as main plasma parameters were closely related within a 5% interval. The number of macroparticles per cell was also varied and statistical convergence was reached. Detailed outputs are given in the supplementary data, to be used by other similar groups in the perspective of code verification.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Persistent minimal sequences of SARS-CoV-2

Abstract Motivation Severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) has caused more than 14 million cases and more than half million deaths. Given the absence of implemented therapies, new analysis, diagnosis and therapeutics are of great importance. Results Analysis of SARS-CoV-2 genomes from the current outbreak reveals the presence of short persistent DNA/RNA sequences that are absent from the human genome and transcriptome (PmRAWs). For the PmRAWs with length 12, only four exist at the same location in all SARS-CoV-2. At the gene level, we found one PmRAW of size 13 at the Spike glycoprotein coding sequence. This protein is fundamental for binding in human ACE2 and further use as an entry receptor to invade target cells. Applying protein structural prediction, we localized this PmRAW at the surface of the Spike protein, providing a potential targeted vector for diagnostics and therapeutics. In addition, we show a new pattern of relative absent words (RAWs), characterized by the progressive increase of GC content (Guanine and Cytosine) according to the decrease of RAWs length, contrarily to the virus and host genome distributions. New analysis shows the same property during the Ebola virus outbreak. At a computational level, we improved the alignment-free method to identify pathogen-specific signatures in balance with GC measures and removed previous size limitations. Availability and implementation https://github.com/cobilab/eagle. Supplementary information Supplementary data are available at Bioinformatics online.

Pratas, Diogo↗

A public website for the automated assessment and validation of SARS-CoV-2 diagnostic PCR assays

Abstract Summary Polymerase chain reaction-based assays are the current gold standard for detecting and diagnosing SARS-CoV-2. However, as SARS-CoV-2 mutates, we need to constantly assess whether existing PCR-based assays will continue to detect all known viral strains. To enable the continuous monitoring of SARS-CoV-2 assays, we have developed a web-based assay validation algorithm that checks existing PCR-based assays against the ever-expanding genome databases for SARS-CoV-2 using both thermodynamic and edit-distance metrics. The assay-screening results are displayed as a heatmap, showing the number of mismatches between each detection and each SARS-CoV-2 genome sequence. Using a mismatch threshold to define detection failure, assay performance is summarized with the true-positive rate (recall) to simplify assay comparisons. Availability and implementation The assay evaluation website and supporting software are Open Source and freely available at https://covid19.edgebioinformatics.org/#/assayValidation, https://github.com/jgans/thermonucleotide BLAST and https://github.com/LANL-Bioinformatics/assay_validation. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

VPF-Class: taxonomic assignment and host prediction of uncultivated viruses based on viral protein families

Abstract Motivation Two key steps in the analysis of uncultured viruses recovered from metagenomes are the taxonomic classification of the viral sequences and the identification of putative host(s). Both steps rely mainly on the assignment of viral proteins to orthologs in cultivated viruses. Viral Protein Families (VPFs) can be used for the robust identification of new viral sequences in large metagenomics datasets. Despite the importance of VPF information for viral discovery, VPFs have not yet been explored for determining viral taxonomy and host targets. Results In this work, we classified the set of VPFs from the IMG/VR database and developed VPF-Class. VPF-Class is a tool that automates the taxonomic classification and host prediction of viral contigs based on the assignment of their proteins to a set of classified VPFs. Applying VPF-Class on 731K uncultivated virus contigs from the IMG/VR database, we were able to classify 363K contigs at the genus level and predict the host of over 461K contigs. In the RefSeq database, VPF-class reported an accuracy of nearly 100% to classify dsDNA, ssDNA and retroviruses, at the genus level, considering a membership ratio and a confidence score of 0.2. The accuracy in host prediction was 86.4%, also at the genus level, considering a membership ratio of 0.3 and a confidence score of 0.5. And, in the prophages dataset, the accuracy in host prediction was 86% considering a membership ratio of 0.6 and a confidence score of 0.8. Moreover, from the Global Ocean Virome dataset, over 817K viral contigs out of 1 million were classified. Availability and implementation The implementation of VPF-Class can be downloaded from https://github.com/biocom-uib/vpf-tools. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

A variant selection framework for genome graphs

Abstract Motivation Variation graph representations are projected to either replace or supplement conventional single genome references due to their ability to capture population genetic diversity and reduce reference bias. Vast catalogues of genetic variants for many species now exist, and it is natural to ask which among these are crucial to circumvent reference bias during read mapping. Results In this work, we propose a novel mathematical framework for variant selection, by casting it in terms of minimizing variation graph size subject to preserving paths of length α with at most δ differences. This framework leads to a rich set of problems based on the types of variants [e.g. single nucleotide polymorphisms (SNPs), indels or structural variants (SVs)], and whether the goal is to minimize the number of positions at which variants are listed or to minimize the total number of variants listed. We classify the computational complexity of these problems and provide efficient algorithms along with their software implementation when feasible. We empirically evaluate the magnitude of graph reduction achieved in human chromosome variation graphs using multiple α and δ parameter values corresponding to short and long-read resequencing characteristics. When our algorithm is run with parameter settings amenable to long-read mapping (α = 10 kbp, δ = 1000), 99.99% SNPs and 73% SVs can be safely excluded from human chromosome 1 variation graph. The graph size reduction can benefit downstream pan-genome analysis. Availability and implementation https://github.com/AT-CG/VF. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

Quartet-based inference is statistically consistent under the unified duplication-loss-coalescence model

Abstract Motivation The classic multispecies coalescent (MSC) model provides the means for theoretical justification of incomplete lineage sorting-aware species tree inference methods. This has motivated an extensive body of work on phylogenetic methods that are statistically consistent under MSC. One such particularly popular method is ASTRAL, a quartet-based species tree inference method. Novel studies suggest that ASTRAL also performs well when given multi-locus gene trees in simulation studies. Further, Legried et al. recently demonstrated that ASTRAL is statistically consistent under the gene duplication and loss model (GDL). GDL is prevalent in evolutionary histories and is the first core process in the powerful duplication-loss-coalescence evolutionary model (DLCoal) by Rasmussen and Kellis. Results In this work, we prove that ASTRAL is statistically consistent under the general DLCoal model. Therefore, our result supports the empirical evidence from the simulation-based studies. More broadly, we prove that the quartet-based inference approach is statistically consistent under DLCoal. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

AutoCCS: automated collision cross-section calculation software for ion mobility spectrometry–mass spectrometry

Abstract Motivation Ion mobility spectrometry (IMS) separations are increasingly used in conjunction with mass spectrometry (MS) for separation and characterization of ionized molecular species. Information obtained from IMS measurements includes the ion’s collision cross section (CCS), which reflects its size and structure and constitutes a descriptor for distinguishing similar species in mixtures that cannot be separated using conventional approaches. Incorporating CCS into MS-based workflows can improve the specificity and confidence of molecular identification. At present, there is no automated, open-source pipeline for determining CCS of analyte ions in both targeted and untargeted fashion, and intensive user-assisted processing with vendor software and manual evaluation is often required. Results We present AutoCCS, an open-source software to rapidly determine CCS values from IMS-MS measurements. We conducted various IMS experiments in different formats to demonstrate the flexibility of AutoCCS for automated CCS calculation: (i) stepped-field methods for drift tube-based IMS (DTIMS), (ii) single-field methods for DTIMS (supporting two calibration methods: a standard and a new enhanced method) and (iii) linear calibration for Bruker timsTOF and non-linear calibration methods for traveling wave based-IMS in Waters Synapt and Structures for Lossless Ion Manipulations. We demonstrated that AutoCCS offers an accurate and reproducible determination of CCS for both standard and unknown analyte ions in various IMS-MS platforms, IMS-field methods, ionization modes and collision gases, without requiring manual processing. Availability and implementation https://github.com/PNNL-Comp-Mass-Spec/AutoCCS. Supplementary information Supplementary data are available at Bioinformatics online. Demo datasets are publicly available at MassIVE (Dataset ID: MSV000085979).

47 OTHER INSTRUMENTATION↗