Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “computational biology and bioinformatics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Energy metric prediction for double insertion mutants via the RoseNet deep learning framework

Studying the structural and functional implications of protein mutations is an important task in computational biology and bioinformatics. We leverage our previously proposed RoseNet neural network architecture to predict energy metrics of proteins with double amino acid insertions or deletions (InDels). We train models on previously generated benchmark datasets containing the exhaustive double InDel mutations for three proteins, as well as an additional three proteins for which ∼145k random mutants, each with two InDels, have been generated. We expand on our previous work by evaluating three additional proteins and analyzing domain features that impact the prediction capabilities of RoseNet. These features include InDels into secondary structures and the solvent accessible surface area (SASA) scores of the residues. We uncover further evidence to support that RoseNet has a higher proficiency of generalizing to unseen residue combinations than unseen insertion positions. We also observe that RoseNet produces higher-quality predictions when inserting into a β-sheet over an α-helix. Additionally, when the insertions fall in an area of high SASA, RoseNet often displays better performance than inserting into areas of low SASA.

59 BASIC BIOLOGICAL SCIENCES↗

KBase Educators Handbook

The KBase Educators Handbook is a community resource for educators teaching biology, computational biology, and bioinformatics using KBase. The Handbook includes supporting documentation on how to join and access community-developed resources for teaching with KBase, best practices, and guidelines on how to contribute to the KBase Educators Community.

59 BASIC BIOLOGICAL SCIENCES↗

The 'Biologically-Inspired Computing' Column

The field of Biology changed dramatically in 1953, with the determination by Francis Crick and James Dewey Watson of the double helix structure of DNA. This discovery changed Biology for ever, allowing the sequencing of the human genome, and the emergence of a "new Biology" focused on DNA, genes, proteins, data, and search. Computational Biology and Bioinformatics heavily rely on computing to facilitate research into life and development. Simultaneously, an understanding of the biology of living organisms indicates a parallel with computing systems: molecules in living cells interact, grow, and transform according to the "program" dictated by DNA. Moreover, paradigms of Computing are emerging based on modelling and developing computer-based systems exploiting ideas that are observed in nature. This includes building into computer systems self-management and self-governance mechanisms that are inspired by the human body's autonomic nervous system, modelling evolutionary systems analogous to colonies of ants or other insects, and developing highly-efficient and highly-complex distributed systems from large numbers of (often quite simple) largely homogeneous components to reflect the behaviour of flocks of birds, swarms of bees, herds of animals, or schools of fish. This new field of "Biologically-Inspired Computing", often known in other incarnations by other names, such as: Autonomic Computing, Pervasive Computing, Organic Computing, Biomimetics, and Artificial Life, amongst others, is poised at the intersection of Computer Science, Engineering, Mathematics, and the Life Sciences. Successes have been reported in the fields of drug discovery, data communications, computer animation, control and command, exploration systems for space, undersea, and harsh environments, to name but a few, and augur much promise for future progress.

Hinchey, Mike↗

On whole-genome demography of world’s ethnic groups and individual genomic identity

All current categorizations of human population, such as ethnicity, ancestry and race, are based on various selections and combinations of complex and dynamic common characteristics, that are mostly societal and cultural in nature, perceived by the members within or from outside of the categorized group. During the last decade, a massive amount of a new type of characteristics, that are exclusively genomic in nature, became available that allows us to analyze the inherited whole-genome demographics of extant human, especially in the fields such as human genetics, health sciences and medical practices (e.g., 1,2,3), where such health-related characteristics can be related to whole-genome-based categorization. Here we show the feasibility of deriving such whole-genome-based categorization. We observe that, within the available genomic data at present, (a) the study populations form about 14 genomic groups, each consisting of multiple ethnic groups; and (b), at an individual level, approximately 99.8%, on average, of the whole autosomal-genome contents are identical between any two individuals regardless of their genomic or ethnic groups.

59 BASIC BIOLOGICAL SCIENCES↗

Machine-learning-based dynamic-importance sampling for adaptive multiscale simulations

Multiscale simulations are a well-accepted way to bridge the length and time scales required for scientific studies with the solution accuracy achievable through available computational resources. Traditional approaches either solve a coarse model with selective refinement or coerce a detailed model into faster sampling, both of which have limitations. Here, we present a paradigm of adaptive, multiscale simulations that couple different scales using a dynamic-importance sampling approach. Our method uses machine learning to dynamically and exhaustively sample the phase space explored by a macro model using microscale simulations and enables an automatic feedback from the micro to the macro scale, leading to a self-healing multiscale simulation. As a result, our approach delivers macro length and time scales, but with the effective precision of the micro scale. Our approach is arbitrarily scalable as well as transferable to many different types of simulations. Overall, our method made possible a multiscale scientific campaign of unprecedented scale to understand the interactions of RAS proteins with a plasma membrane in the context of cancer research running over several days on Sierra, which is currently the second-most-powerful supercomputer in the world.

59 BASIC BIOLOGICAL SCIENCES↗

AI-accelerated protein-ligand docking for SARS-CoV-2 is 100-fold faster with no significant change in detection

Protein-ligand docking is a computational method for identifying drug leads. The method is capable of narrowing a vast library of compounds down to a tractable size for downstream simulation or experimental testing and is widely used in drug discovery. While there has been progress in accelerating scoring of compounds with artificial intelligence, few works have bridged these successes back to the virtual screening community in terms of utility and forward-looking development. We demonstrate the power of high-speed ML models by scoring 1 billion molecules in under a day (50 k predictions per GPU seconds). We showcase a workflow for docking utilizing surrogate AI-based models as a pre-filter to a standard docking workflow. Our workflow is ten times faster at screening a library of compounds than the standard technique, with an error rate less than 0.01% of detecting the underlying best scoring 0.1% of compounds. Our analysis of the speedup explains that another order of magnitude speedup must come from model accuracy rather than computing speed. In order to drive another order of magnitude of acceleration, we share a benchmark dataset consisting of 200 million 3D complex structures and 2D structure scores across a consistent set of 13 million “in-stock” molecules over 15 receptors, or binding sites, across the SARS-CoV-2 proteome. We believe this is strong evidence for the community to begin focusing on improving the accuracy of surrogate models to improve the ability to screen massive compound libraries 100 × or even 1000 × faster than current techniques and reduce missing top hits. The technique outlined aims to be a fast drop-in replacement for docking for screening billion-scale molecular libraries.

59 BASIC BIOLOGICAL SCIENCES↗

Structural basis of promiscuous substrate transport by Organic Cation Transporter 1

Organic Cation Transporter 1 (OCT1) plays a crucial role in hepatic metabolism by mediating the uptake of a range of metabolites and drugs. Genetic variations can alter the efficacy and safety of compounds transported by OCT1, such as those used for cardiovascular, oncological, and psychological indications. Despite its importance in drug pharmacokinetics, the substrate selectivity and underlying structural mechanisms of OCT1 remain poorly understood. Here, we present cryo-EM structures of full-length human OCT1 in the inward-open conformation, both ligand-free and drug-bound, indicating the basis for its broad substrate recognition. Comparison of our structures with those of outward-open OCTs provides molecular insight into the alternating access mechanism of OCTs. We observe that hydrophobic gates stabilize the inward-facing conformation, whereas charge neutralization in the binding pocket facilitates the release of cationic substrates. These findings provide a framework for understanding the structural basis of the promiscuity of drug binding and substrate translocation in OCT1.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Integrating chromatin conformation information in a self-supervised learning model improves metagenome binning

Metagenome binning is a key step, downstream of metagenome assembly, to group scaffolds by their genome of origin. Although accurate binning has been achieved on datasets containing multiple samples from the same community, the completeness of binning is often low in datasets with a small number of samples due to a lack of robust species co-abundance information. In this study, we exploited the chromatin conformation information obtained from Hi-C sequencing and developed a new reference-independent algorithm, Metagenome Binning with Abundance and Tetra-nucleotide frequencies—Long Range (metaBAT-LR), to improve the binning completeness of these datasets. This self-supervised algorithm builds a model from a set of high-quality genome bins to predict scaffold pairs that are likely to be derived from the same genome. Then, it applies these predictions to merge incomplete genome bins, as well as recruit unbinned scaffolds. We validated metaBAT-LR’s ability to bin-merge and recruit scaffolds on both synthetic and real-world metagenome datasets of varying complexity. Benchmarking against similar software tools suggests that metaBAT-LR uncovers unique bins that were missed by all other methods.

59 BASIC BIOLOGICAL SCIENCES↗

A multi-scale pipeline linking drug transcriptomics with pharmacokinetics predicts in vivo interactions of tuberculosis drugs

Tuberculosis (TB) is the deadliest infectious disease worldwide. The design of new treatments for TB is hindered by the large number of candidate drugs, drug combinations, dosing choices, and complex pharmaco-kinetics/dynamics (PK/PD). Here we study the interplay of these factors in designing combination therapies by linking a machine-learning model, INDIGO-MTB, which predicts in vitro drug interactions using drug transcriptomics, with a multi-scale model of drug PK/PD and pathogen-immune interactions called GranSim. We calculate an in vivo drug interaction score (iDIS) from dynamics of drug diffusion, spatial distribution, and activity within lesions against various pathogen sub-populations. The iDIS of drug regimens evaluated against non-replicating bacteria significantly correlates with efficacy metrics from clinical trials. Our approach identifies mechanisms that can amplify synergistic or mitigate antagonistic drug interactions in vivo by modulating the relative distribution of drugs. Our mechanistic framework enables efficient evaluation of in vivo drug interactions and optimization of combination therapies.

59 BASIC BIOLOGICAL SCIENCES↗

Bioinformatic Teaching Resources – For Educators, by Educators – Using KBase, a Free, User-Friendly, Open Source Platform

Over the past year, biology educators and staff at the U.S. Department of Energy Systems Biology Knowledgebase (KBase) initiated a collaborative effort to develop a curriculum for bioinformatics education. KBase is a free web-based platform where anyone can conduct sophisticated and reproducible bioinformatic analyses via a graphical user interface. Here, we demonstrate the utility of KBase as a platform for bioinformatics education, and present a set of modular, adaptable, and customizable instructional units for teaching concepts in Genomics, Metagenomics, Pangenomics, and Phylogenetics. Each module contains teaching resources, publicly available data, analysis tools, and Markdown capability, enabling instructors to modify the lesson as appropriate for their specific course. We present initial student survey data on the effectiveness of using KBase for teaching bioinformatic concepts, provide an example case study, and detail the utility of the platform from an instructor’s perspective. Even as in-person teaching returns, KBase will continue to work with instructors, supporting the development of new active learning curriculum modules. For anyone utilizing the platform, the growing KBase Educators Organization provides an educators network, accompanied by community-sourced guidelines, instructional templates, and peer support, for instructors wishing to use KBase within a classroom at any educational level–whether virtual or in-person.

59 BASIC BIOLOGICAL SCIENCES↗

Metagenomic features of bioburden serve as outcome indicators in combat extremity wounds

Abstract Battlefield injury management requires specialized care, and wound infection is a frequent complication. Challenges related to characterizing relevant pathogens further complicates treatment. Applying metagenomics to wounds offers a comprehensive path toward assessing microbial genomic fingerprints and could indicate prognostic variables for future decision support tools. Wound specimens from combat-injured U.S. service members, obtained during surgical debridements before delayed wound closure, were subjected to whole metagenome analysis and targeted enrichment of antimicrobial resistance genes. Results did not indicate a singular, common microbial metagenomic profile for wound failure, instead reflecting a complex microenvironment with varying bioburden diversity across outcomes. Genus-level Pseudomonas detection was associated with wound failure at all surgeries. A logistic regression model was fit to the presence and absence of antimicrobial resistance classes to assess associations with nosocomial pathogens. A. baumannii detection was associated with detection of genomic signatures for resistance to trimethoprim, aminoglycosides, bacitracin, and polymyxin. Machine learning classifiers were applied to identify wound and microbial variables associated with outcome. Feature importance rankings averaged across models indicated the variables with the largest effects on predicting wound outcome, including an increase in P. putida sequence reads. These results describe the microbial genomic determinants in combat wound bioburden and demonstrate metagenomic investigation as a comprehensive tool for providing information toward aiding treatment of combat-related injuries.

59 BASIC BIOLOGICAL SCIENCES↗

Nonlinear model of infection wavy oscillation of COVID-19 in Japan based on diffusion kinetics

The infectious propagation of SARS-CoV-2 is continuing worldwide, and specifically, Japan is facing severe circumstances. Medical resource maintenance and action limitations remain the central measures. An analysis of long-term follow-up reports in Japan shows that the infection number follows a unique wavy oscillation, increasing and decreasing over time. However, only a few studies explain the infection wavy oscillation. This study introduces a novel nonlinear mathematical model of the new infection wavy oscillation by applying the macromolecule diffusion theory. In this model, the diffusion coefficient that depends on population density gives nonlinearity in infection propagation. As a result, our model accurately simulated infection wavy oscillations, and the infection wavy oscillation frequency and amplitude were closely linked with the recovery rate of infected individuals. In conclusion, our model provides a novel nonlinear contact infection analysis framework.

60 APPLIED LIFE SCIENCES↗

Author Correction: Perspectives on ENCODE

In the original article, the authors Rizi Ai (Department of Chemistry and Biochemistry, University of California, San Diego, La Jolla, CA, USA) and Shantao Li (Program in Computational Biology and Bioinformatics, Yale University, New Haven, CT, USA) were mistakenly omitted from the ENCODE Project Consortium author list. The original Article has been corrected online.

99 GENERAL AND MISCELLANEOUS↗

Author Correction: Expanded encyclopaedias of DNA elements in the human and mouse genomes

In the version of this article initially published, two members of the ENCODE Project Consortium were missing from the author list. Rizi Ai (Department of Chemistry and Biochemistry, University of California, San Diego, La Jolla, CA, USA) and Shantao Li (Program in Computational Biology and Bioinformatics, Yale University, New Haven, CT, USA) are now included in the author list. These errors have been corrected in the online version of the article.

59 BASIC BIOLOGICAL SCIENCES↗

The performance of ensemble-based free energy protocols in computing binding affinities to ROS1 kinase

Optimization of binding affinities for compounds to their target protein is a primary objective in drug discovery. Herein we report on a collaborative study that evaluates a set of compounds binding to ROS1 kinase. We use ESMACS (enhanced sampling of molecular dynamics with approximation of continuum solvent) and TIES (thermodynamic integration with enhanced sampling) protocols to rank the binding free energies. The predicted binding free energies from ESMACS simulations show good correlations with experimental data for subsets of the compounds. Consistent binding free energy differences are generated for TIES and ESMACS. Although an unexplained overestimation exists, we obtain excellent statistical rankings across the set of compounds from the TIES protocol, with a Pearson correlation coefficient of 0.90 between calculated and experimental activities.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Network-medicine framework for studying disease trajectories in U.S. veterans

A better understanding of the sequential and temporal aspects in which diseases occur in patient’s lives is essential for developing improved intervention strategies that reduce burden and increase the quality of health services. Here we present a network-based framework to study disease relationships using Electronic Health Records from > 9 million patients in the United States Veterans Health Administration (VHA) system. We create the Temporal Disease Network, which maps the sequential aspects of disease co-occurrence among patients and demonstrate that network properties reflect clinical aspects of the respective diseases. We use the Temporal Disease Network to identify disease groups that reflect patterns of disease co-occurrence and the flow of patients among diagnoses. Finally, we define a strategy for the identification of trajectories that lead from one disease to another. The framework presented here has the potential to offer new insights for disease treatment and prevention in large health care systems.

60 APPLIED LIFE SCIENCES↗

Energy landscapes from cryo-EM snapshots: a benchmarking study

Abstract Biomolecules undergo continuous conformational motions, a subset of which are functionally relevant. Understanding, and ultimately controlling biomolecular function are predicated on the ability to map continuous conformational motions, and identify the functionally relevant conformational trajectories. For equilibrium and near-equilibrium processes, function proceeds along minimum-energy pathways on one or more energy landscapes, because higher-energy conformations are only weakly occupied. With the growing interest in identifying functional trajectories, the need for reliable mapping of energy landscapes has become paramount. In response, various data-analytical tools for determining structural variability are emerging. A key question concerns the veracity with which each data-analytical tool can extract functionally relevant conformational trajectories from a collection of single-particle cryo-EM snapshots. Using synthetic data as an independently known ground truth, we benchmark the ability of four leading algorithms to determine biomolecular energy landscapes and identify the functionally relevant conformational paths on these landscapes. Such benchmarking is essential for systematic progress toward atomic-level movies of continuous biomolecular function.

59 BASIC BIOLOGICAL SCIENCES↗

Information-incorporated gene network construction with FDR control

Abstract Motivation Large-scale gene expression studies allow gene network construction to uncover associations among genes. To study direct associations among genes, partial correlation-based networks are preferred over marginal correlations. However, FDR control for partial correlation-based network construction is not well-studied. In addition, currently available partial correlation-based methods cannot take existing biological knowledge to help network construction while controlling FDR. Results In this paper, we propose a method called Partial Correlation Graph with Information Incorporation (PCGII). PCGII estimates partial correlations between each pair of genes by regularized node-wise regression that can incorporate prior knowledge while controlling the effects of all other genes. It handles high-dimensional data where the number of genes can be much larger than the sample size and controls FDR at the same time. We compare PCGII with several existing approaches through extensive simulation studies and demonstrate that PCGII has better FDR control and higher power. We apply PCGII to a plant gene expression dataset where it recovers confirmed regulatory relationships and a hub node, as well as several direct associations that shed light on potential functional relationships in the system. We also introduce a method to supplement observed data with a pseudogene to apply PCGII when no prior information is available, which also allows checking FDR control and power for real data analysis. Availability and implementation R package is freely available for download at https://cran.r-project.org/package=PCGII.

59 BASIC BIOLOGICAL SCIENCES↗