CoVTransformer
The code is a transformer model to forecast SARS-CoV-2 lineage frequencies in the future.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
The code is a transformer model to forecast SARS-CoV-2 lineage frequencies in the future.
Code developed to fit and analyze viral load data for acute respiratory infections such as SARS-CoV-2. The code can also implement different treatments and generate in-silicon cohorts to examine the effects of treatments.
Protein language machine learning models built upon existing ESM-2 model developed by Evolutionary Scale (evolutionaryscale.ai) and an in-house protein language model based on the BERT model developed by Google. The code also includes model training scripts and saved checkpoints from our own training using publicly available SARS-CoV-2 protein sequences.
Broadly reactive antibodies that target sequence-diverse antigens are of interest for vaccine design and monoclonal antibody therapeutic development because they can protect against multiple strains of a virus and provide a barrier to evolution of escape mutants. Using LIBRA-seq (linking B cell receptor to antigen specificity through sequencing) data for the B cell repertoire of an individual chronically infected with human immunodeficiency virus type 1 (HIV-1), we identified a lineage of IgG3 antibodies predicted to bind to HIV-1 Envelope (Env) and influenza A Hemagglutinin (HA). Two lineage members, antibodies 2526 and 546, were confirmed to bind to a large panel of diverse antigens, including several strains of HIV-1 Env, influenza HA, coronavirus (CoV) spike, hepatitis C virus (HCV) E protein, Nipah virus (NiV) F protein, and Langya virus (LayV) F protein. We found that both antibodies bind to complex glycans on the antigenic surfaces. Antibody 2526 targets the stem region of influenza HA and the N-terminal domain (NTD) region of SARS-CoV-2 spike. A crystal structure of 2526 Fab bound to mannose revealed the presence of a glycan-binding pocket on the light chain. Antibody 2526 cross-reacted with antigens from multiple pathogens and displayed no signs of autoreactivity. These features distinguish antibody 2526 from previously described glycan-reactive antibodies. Further study of this antibody class may aid in the selection and engineering of broadly reactive antibody therapeutics and can inform the development of effective vaccines with exceptional breadth of pathogen coverage.
The LDRD ER “Building a Computational and Experimental Rapid Response Pipeline to Counter the Coronavirus Disease 2019 Outbreak and Emerging Biothreats” was conceived to address a need for rapid, scalable, evaluation of computationally designed therapeutic or prophylactic antibodies and vaccine antigens, two important classes of protein medical countermeasure (MCM). This was done in complement to a computationally driven LDRD 20ERD032 “Active Learning for Rapid Design of Vaccines and Antibodies.” Natural antibodies and antigens are often insufficiently broad or robust across different pathogens and their variants. Leveraging a collaboration of simulation driven machine learning, structural expertise, and high-throughput characterization of candidate antibodies, we successfully re-targeted three different anti-SARS-CoV-1 antibodies to neutralize SARS-CoV-2 in vitro. Our antibody design work reached its most important stage in rapid response to the emergence of the Omicron variant of concern (VOC) in late 2021. In a matter of weeks, we computationally designed derivative antibodies of COV2-2130, one of two antibodies from Vanderbilt that form the basis of the AstraZeneca Evusheld prophylactic drug product. This drug product suffers a serious loss of efficacy against Omicron BA.1 and BA.1.1, the first Omicron strains. Our designs were successful, including a pair of designs which provide potent neutralization of not only Omicron BA.1 and BA.1.1, but also the earlier Delta variant, and subsequent Omicron strains including BA.2, BA.4, BA.5, and BA.2.75, demonstrating that our multi-target design process can, by its nature, produce robust antibody designs that strictly improve over the parental antibody. These results, recognized by a 2022 Director’s Science and Technology award, have enabled the follow-on GUIDE program, to commence in FY23.
Virulence assessment of new, emerging, and engineered pathogens is critical to mounting an appropriate response to a biothreat agent. The capacity of the pathogen to colonize human and harm tissues must be characterized to understand pathogenicity pathways and optimize diagnosis and treatment of resulting disease. Respiratory pathogens are of interest because they can have high transmissibility rates, as observed with the SARS-CoV-2 virus, the causative agent of Covid-19. Current technologies are insufficient to assess threats due to their reliance on systems with only one cell type and on sequencing the pathogen. However, it is known that sequence is not an accurate predictor of function, and sequencing can be unreliable for newly emerged or engineered pathogens. An ideal system would consist of relevant epithelial cell types and an assay sensitive enough to detect changes in host responses that do not rely on DNA sequencing. We chose a system consisting of host lung epithelial cells that can be used to assess the virulence of unknown respiratory pathogens. We interrogated pathogens using this model and assess features of pathogenicity. Our objective is to leverage PNNLs strengths in tissue engineering and proteomics capabilities to build a multiple reaction monitoring (MRM) or parallel reaction monitoring (PRM) liquid chromatography-tandem mass spectrometry assay for human host cell proteins whose abundance is influenced by infection. These responses can were then assessed for relative virulence using pathogen agnostic signatures. When confronted with a pathogen, cells activate dedicated signaling pathways, typically through phosphorylation of regulatory proteins and downstream activation of host cell networks.
The major goal of this project is to develop machine learning (ML) methods to enable improved predictive power on real drug discovery for novel targets. More specifically, we plan to demonstrate the capability and effectiveness of ML tools utilizing unlabeled large-volume protein-ligand datasets. We also plan to demonstrate the capability and effectiveness of the developed methods by testing on a realistic drug discovery task to identify pan-coronavirus protease inhibitors such as SARS-CoV-2. While the overall goals and milestones remain consistent with the original proposal, certain technical details have been modified, which we will describe in this report.
The natural tendency of virus to mutate and the under-sampling of the environment makes it difficult to become aware of the emergence of new viral strains with pandemic potential. Being able to predict what mutations make a virus more infective might allow to spot such strains with minimal sampling and potentially allow to predict/prevent the next pandemic. The team attempted to mimic natural viral mutations and recombination through computational and experimental methods producing a variety of mutants of a SARS-COV-2 protein (receptor binding domain, RBD, of spike protein) responsible for viral entry in mammalian cells. The library of mutants was then interrogated for ability and lack-there-of to interact with the host cell receptor ushering viral entry, Angiotensin-converting enzyme 2 (ACE2). The negative and positive data set is intended to “teach the rules” of virus-host receptor interaction. Additionally, the positive clones were used to screen a set of antibody mutants designed to widen the breadth of viral mutants recognition, to demonstrate that this kind of libraries could also be a tool to produce antibody therapeutics impervious to viral mutation, even before a pandemic strain is discovered.
Our project established and demonstrated a transfer learning framework that enables prediction of antibody–antigen interactions across related viruses. The approach focused on three major activities: 1. Conserved region and epitope identification – We compared viral protein structures and sequences to identify shared receptor-binding domains and neutralizing epitope regions across variants and related viruses. These conserved features formed the foundation for discovering broadly functional antibodies. 2. Machine learning model development – We built neural network–based models that integrate epitope features with antibody sequence information. Instead of relying solely on structural or physical properties, the models learned transferable patterns that describe antibody binding potential across different viral families. 3. Transfer learning and validation – Using SARS-CoV-2 and Ebola as source systems, we successfully transferred learned epitope features to predict antibody interactions for SARS CoV-1 and Marburg virus. Iterative cycles of dataset generation, retraining, and evaluation improved generalization and predictive power, ensuring the framework can adapt to new threats.
The major goal of this project is to develop machine learning (ML) methods to enable improved predictive power on real drug discovery for novel targets. More specifically, we planned to demonstrate the capability and effectiveness of ML tools utilizing unlabeled large-volume protein-ligand datasets. We investigated multiple pre-training approaches for 3D protein-ligand structure-based foundation models, without relying on experimental binding data. We also addressed scenarios in which crystal structures are unavailable or binding data are limited. We also planned to develop a complete pipeline to screen novel compounds as well as to demonstrate the capability and effectiveness of the developed methods by testing on a realistic drug discovery task such as SARS-CoV-2. While the major goals and milestones remain consistent with the original proposal, certain technical details have been adjusted, based on the experimental results and related outcomes.
Vaccines have historically played a pivotal role in controlling epidemics. Effective vaccines for viruses causing significant human disease, e.g., Ebola, Lassa fever, or Crimean Congo hemorrhagic fever virus, would be invaluable to public health strategies and counter-measure development missions. Here, we propose coverage metrics to quantify vaccine-induced CD8 + T cell-mediated immune protection, as well as metrics to characterize immuno-dominant epitopes, in light of human genetic heterogeneity and viral evolution. Proof-of-principle of our approach and methods are demonstrated for Ebola virus, SARS-CoV-2, and Burkholderia pseudomallei (vaccine) proteins.
Recent advances in 3D structure-based deep learning approaches demonstrate improved accuracy in predicting protein-ligand binding affinity in drug discovery. These methods complement physics-based computational modeling such as molecular docking for virtual high-throughput screening. Despite recent advances and improved predictive performance, most methods in this category primarily rely on utilizing co-crystal complex structures and experimentally measured binding affinities as both input and output data for model training. Nevertheless, co-crystal complex structures are not readily available and the inaccurate predicted structures from molecular docking can degrade the accuracy of the machine learning methods. We introduce a novel structure-based inference method utilizing multiple molecular docking poses for each complex entity. Our proposed method employs multi-instance learning with an attention network to predict binding affinity from a collection of docking poses. We validate our method using multiple datasets, including PDBbind and compounds targeting the main protease of SARS-CoV-2. The results demonstrate that our method leveraging docking poses is competitive with other state-of-the-art inference models that depend on co-crystal structures. This method offers binding affinity prediction without requiring co-crystal structures, thereby increasing its applicability to protein targets lacking such data.
Decades of drug development research have explored a vast chemical space for highly active compounds. The exponential growth of virtual libraries enables easy access to billions of synthesizable molecules. Computational modeling, particularly molecular docking, utilizes physics-based calculations to prioritize molecules for synthesis and testing. Nevertheless, the molecular docking process often yields docking poses with favorable scores that prove to be inaccurate with experimental testing. To address these issues, several approaches using machine learning (ML) have been proposed to filter incorrect poses based on the crystal structures. However, most of the methods are limited by the availability of structure data. Here, we propose a new pose classification approach, PECAN2 (Pose Classification with 3D Atomic Network 2), without the need for crystal structures, based on a 3D atomic neural network with Point Cloud Network (PCN). The new approach uses the correlation between docking scores and experimental data to assign labels, instead of relying on the crystal structures. We validate the proposed classifier on multiple datasets including human mu, delta, and kappa opioid receptors and SARS-CoV-2 Mpro. Our results demonstrate that leveraging the correlation between docking scores and experimental data alone enhances molecular docking performance by filtering out false positives and false negatives.
In the absence of preventive therapies or effective treatment for most cases of coronavirus disease 2019 (COVID-19), governments worldwide have sought to minimize person-to-person severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) transmission through a variety of lock-down measures and social distancing policies. Extreme events like the COVID-19 pandemic present a tremendous opportunity to make quantitative connections between changes in anthropogenic forcing, social and economic activity, and the related Earth system response. In this comment, we examine air quality impacts associated with pandemic response measures in the Northeastern United States.
Saliva has been used as a source of biological markers for a wide spectrum of normal and disease states for a long time. It is a non-invasive, easily accessible, and self-collected body fluid that contains a variety of measurable biological substances. While mostly water, saliva also contains ions, carbohydrates, proteins and peptides, exfoliated cells, nucleic acids, and microorganisms. Saliva can reflect tissue levels of some natural substances and a large variety of molecules introduced for therapeutic use, emotional status; hormonal status, immunological status, neurological effects, and nutritional and metabolic status. It can also be used to monitor a variety of drugs including marijuana, cocaine, and alcohol. It is the most cost-effective approach for screening large population in community mass screening programs and for longitudinal sampling of hospitalized individuals aimed at monitoring viral load dynamics and treatment response. During the COVID-19 pandemic, scientific evidence emerged indicating that molecular tests performed on saliva have diagnostic sensitivity and specificity comparable to those observed with nasopharyngeal swabs for SARS-CoV-2 RNA detection. The presence of IgA and IgG antibodies at the mucosal level has been demonstrated to influence the progression of viral infection and the severity of clinical manifestation. As saliva contains both respiratory secretions and immunological components, it has wide applications, ranging from clinical diagnostics to post-vaccine disease burden and immunity surveillance.
Immunogenic compositions and methods of their use in eliciting immune responses to coronaviruses, such as SARS-CoV-2 are provided.
The present disclosure relates to an isolated or purified antibody, or a fragment thereof, having a binding domain that binds to a coronavirus (e.g., SARS-COV-2) or a portion thereof. In other embodiments, the antibody includes a binding domain that competes with binding to angiotensin converting enzyme 2 (ACE2) or a portion thereof. Methods of using such antibodies are also described herein, such as methods of treating or delaying the progression of a disease associated with a coronavirus.
Most virus infection assays have indirect readout such as virus number following entry (e.g., PCR, cell lysis). While effective, these technologies are labor‐intensive, require specialized environments (e.g., sterile or RNA‐free), and detect later‐stage viral events like lysis or cell death, lacking sensitivity to early fusion events. To address these limitations, we present biologically relevant 2D membrane materials, host‐cell‐derived supported lipid bilayers (hcd‐SLBs), integrated with organic microelectrode arrays (OMEAs) for detection of severe acute respiratory syndrome coronavirus 2 (SARS‐CoV‐2) fusion. By overexpressing angiotensin‐converting enzyme 2 (ACE2) receptors on the native membranes, the platform functions as a viral sensor capable of detecting virus pseudo particles (VPPs) through the late pathway. Additionally, hcd‐SLBs extracted from human lung epithelium expressing native ACE2 detect fusion events through the early pathway. The platform's utility as a drug‐screening tool is demonstrated by testing antibodies targeting either the ACE2 on the host membrane or the viral spike (S) proteins. To enhance the throughput, microfluidics are integrated for automation and OMEAs are incorporated within each channel, miniaturizing the testing units. This system supports high‐throughput data generation, automation, and scalability, providing an efficient platform for viral fusion detection that advances the study of pathogen‐host interactions and accelerates antiviral drug discovery.