Engineering Papers⌕ Search

Engineering topics

Li, Po-E

Publications and source records attributed to Li, Po-E.

Recommendations for Uniform Variant Calling of SARS-CoV-2 Genome Sequence across Bioinformatic Workflows

Genomic sequencing of clinical samples to identify emerging variants of SARS-CoV-2 has been a key public health tool for curbing the spread of the virus. As a result, an unprecedented number of SARS-CoV-2 genomes were sequenced during the COVID-19 pandemic, which allowed for rapid identification of genetic variants, enabling the timely design and testing of therapies and deployment of new vaccine formulations to combat the new variants. However, despite the technological advances of deep sequencing, the analysis of the raw sequence data generated globally is neither standardized nor consistent, leading to vastly disparate sequences that may impact identification of variants. Here, we show that for both Illumina and Oxford Nanopore sequencing platforms, downstream bioinformatic protocols used by industry, government, and academic groups resulted in different virus sequences from same sample. These bioinformatic workflows produced consensus genomes with differences in single nucleotide polymorphisms, inclusion and exclusion of insertions, and/or deletions, despite using the same raw sequence as input datasets. Here, we compared and characterized such discrepancies and propose a specific suite of parameters and protocols that should be adopted across the field. Consistent results from bioinformatic workflows are fundamental to SARS-CoV-2 and future pathogen surveillance efforts, including pandemic preparation, to allow for a data-driven and timely public health response.

60 APPLIED LIFE SCIENCES↗

Comparing variability in diagnosis of upper respiratory tract infections in patients using syndromic, next generation sequencing, and PCR-based methods

Early and accurate diagnosis of respiratory pathogens and associated outbreaks can allow for the control of spread, epidemiological modeling, targeted treatment, and decision making–as is evident with the current COVID-19 pandemic. Many respiratory infections share common symptoms, making them difficult to diagnose using only syndromic presentation. Yet, with delays in getting reference laboratory tests and limited availability and poor sensitivity of point-of-care tests, syndromic diagnosis is the most-relied upon method in clinical practice today. Here, we examine the variability in diagnostic identification of respiratory infections during the annual infection cycle in northern New Mexico, by comparing syndromic diagnostics with polymerase chain reaction (PCR) and sequencing-based methods, with the goal of assessing gaps in our current ability to identify respiratory pathogens. Of 97 individuals that presented with symptoms of respiratory infection, only 23 were positive for at least one RNA virus, as confirmed by sequencing. Whereas influenza virus (n = 7) was expected during this infection cycle, we also observed coronavirus (n = 7), respiratory syncytial virus (n = 8), parainfluenza virus (n = 4), and human metapneumovirus (n = 1) in individuals with respiratory infection symptoms. Four patients were coinfected with two viruses. In 21 individuals that tested positive using PCR, RNA sequencing completely matched in only 12 (57%) of these individuals. Few individuals (37.1%) were diagnosed to have an upper respiratory tract infection or viral syndrome by syndromic diagnostics, and the type of virus could only be distinguished in one patient. Thus, current syndromic diagnostic approaches fail to accurately identify respiratory pathogens associated with infection and are not suited to capture emerging threats in an accurate fashion. We conclude there is a critical and urgent need for layered agnostic diagnostics to track known and unknown pathogens at the point of care to control future outbreaks.

60 APPLIED LIFE SCIENCES↗

EDGE COVID-19: a web platform to generate submission-ready genomes from SARS-CoV-2 sequencing efforts

Abstract Summary Genomics has become an essential technology for surveilling emerging infectious disease outbreaks. A range of technologies and strategies for pathogen genome enrichment and sequencing are being used by laboratories worldwide, together with different and sometimes ad hoc, analytical procedures for generating genome sequences. A fully integrated analytical process for raw sequence to consensus genome determination, suited to outbreaks such as the ongoing COVID-19 pandemic, is critical to provide a solid genomic basis for epidemiological analyses and well-informed decision making. We have developed a web-based platform and integrated bioinformatic workflows that help to provide consistent high-quality analysis of SARS-CoV-2 sequencing data generated with either the Illumina or Oxford Nanopore Technologies (ONT). Using an intuitive web-based interface, this workflow automates data quality control, SARS-CoV-2 reference-based genome variant and consensus calling, lineage determination and provides the ability to submit the consensus sequence and necessary metadata to GenBank, GISAID and INSDC raw data repositories. We tested workflow usability using real world data and validated the accuracy of variant and lineage analysis using several test datasets, and further performed detailed comparisons with results from the COVID-19 Galaxy Project workflow. Our analyses indicate that EC-19 workflows generate high-quality SARS-CoV-2 genomes. Finally, we share a perspective on patterns and impact observed with Illumina versus ONT technologies on workflow congruence and differences. Availability and implementation https://edge-covid19.edgebioinformatics.org, and https://github.com/LANL-Bioinformatics/EDGE/tree/SARS-CoV2. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

Using Deep Mutational Data and Machine Learning to Guide Outbreak and Pandemic Response

A significant fraction of pathogens known to infect humans originate in non-human (zoonotic) hosts (Taylor, Latham, and Woolhouse 2001), and new and emerging pathogens continue to spill over into the human population more frequently at an alarming rate (e.g., SARS, MERS, Cholera, etc.). The recent outbreaks of Ebola virus in West Africa and the ongoing SARS-CoV-2 pandemic demonstrate the need for rapid and reliable assessments of viral phenotype information to help inform scientists and policy makers how best to control the spread of disease. Further understanding of the virus pathogenic evolutionary space and potential trajectory could guide appropriate control measures to limit the spread of a new virus throughout the local and global human population.

59 BASIC BIOLOGICAL SCIENCES↗

Challenges in Bioinformatics Workflows for Processing Microbiome Omics Data at Scale

The nascent field of microbiome science is transitioning from a descriptive approach of cataloging taxa and functions present in an environment to applying multi-omics methods to investigate microbiome dynamics and function. A large number of new tools and algorithms have been designed and used for very specific purposes on samples collected by individual investigators or groups. While these developments have been quite instructive, the ability to compare microbiome data generated by many groups of researchers is impeded by the lack of standardized application of bioinformatics methods. Additionally, there are few examples of broad bioinformatics workflows that can process metagenome, metatranscriptome, metaproteome and metabolomic data at scale, and no central hub that allows processing, or provides varied omics data that are findable, accessible, interoperable and reusable (FAIR). Here, we review some of the challenges that exist in analyzing omics data within the microbiome research sphere, and provide context on how the National Microbiome Data Collaborative has adopted a standardized and open access approach to address such challenges.

NMDC, Microbiome↗