Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “bioinformatics tool”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Discovery, characterization, and application of chromosomal integration sites for stable heterologous gene expression in Rhodotorula toruloides

Rhodotorula toruloides is a non-model, oleaginous yeast uniquely suited to produce acetyl-CoA-derived chemicals. However, the lack of well-characterized genomic integration sites has impeded the metabolic engineering of this organism. Here we report a set of computationally predicted and experimentally validated chromosomal integration sites in R. toruloides. We first implemented an in silico platform by integrating essential gene information and transcriptomic data to identify candidate sites that meet stringent criteria. We then conducted a full experimental characterization of these sites, assessing integration efficiency, gene expression levels, impact on cell growth, and long-term expression stability. Among the identified sites, 12 exhibited integration efficiencies of 50% or higher, making them sufficient for most metabolic engineering applications. Using selected high-efficiency sites, we achieved simultaneous double and triple integrations and efficiently integrated long functional pathways (up to 14.7 kb). Additionally, we developed a new inducible marker recycling system that allows multiple rounds of integration at our characterized sites. Here, we validated this system by performing five sequential rounds of GFP integration and three sequential rounds of MaFAR integration for fatty alcohol production, demonstrating, for the first time, precise gene copy number tuning in R. toruloides. These characterized integration sites should significantly advance metabolic engineering efforts and future genetic tool development in R. toruloides.

59 BASIC BIOLOGICAL SCIENCES↗

VPF-Class: taxonomic assignment and host prediction of uncultivated viruses based on viral protein families

Abstract Motivation Two key steps in the analysis of uncultured viruses recovered from metagenomes are the taxonomic classification of the viral sequences and the identification of putative host(s). Both steps rely mainly on the assignment of viral proteins to orthologs in cultivated viruses. Viral Protein Families (VPFs) can be used for the robust identification of new viral sequences in large metagenomics datasets. Despite the importance of VPF information for viral discovery, VPFs have not yet been explored for determining viral taxonomy and host targets. Results In this work, we classified the set of VPFs from the IMG/VR database and developed VPF-Class. VPF-Class is a tool that automates the taxonomic classification and host prediction of viral contigs based on the assignment of their proteins to a set of classified VPFs. Applying VPF-Class on 731K uncultivated virus contigs from the IMG/VR database, we were able to classify 363K contigs at the genus level and predict the host of over 461K contigs. In the RefSeq database, VPF-class reported an accuracy of nearly 100% to classify dsDNA, ssDNA and retroviruses, at the genus level, considering a membership ratio and a confidence score of 0.2. The accuracy in host prediction was 86.4%, also at the genus level, considering a membership ratio of 0.3 and a confidence score of 0.5. And, in the prophages dataset, the accuracy in host prediction was 86% considering a membership ratio of 0.6 and a confidence score of 0.8. Moreover, from the Global Ocean Virome dataset, over 817K viral contigs out of 1 million were classified. Availability and implementation The implementation of VPF-Class can be downloaded from https://github.com/biocom-uib/vpf-tools. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

GeneLab for High Schools: Data Mining for the Next Generation

Modern biological sciences have become increasingly based on molecular biology and high-throughput molecular techniques, such as genomics, transcriptomics, and proteomics. NASA Scientists and the NASA Space Biology Program have aimed to examine the fundamental building blocks of life (RNA, DNA and protein) in order to understand the response of living organisms to space and aid in fundamental research discoveries on Earth. In an effort to enable NASA funded science to be available to everyone, NASA has collected the data from omics studies and curated them in a data system called GeneLab. Whilst most college-level interns, academics and other scientists have had some interaction with omics data sets and analysis tools, high school students often have not. Therefore, the Space Biology Program is implementing a new Summer Program for high-school students that aims to inspire the next generation of scientists to learn about and get involved in space research using GeneLabs Data System. The program consists of three main components core learning modules, focused on developing students knowledge on the Space Biology Program and Space Biology research, Genelab and the data system, and previous research conducted on model organisms in space; networking and team work, enabling students to interact with guest lecturers from local universities and their fellow peers, and also enabling them to visit local universities and genomics centers around the Bay area; and finally an independent learning project, whereby students will be required to form small groups, analyze a dataset on the Genelab platform, generate a hypothesis and develop a research plan to test their hypothesis. This program will not only help inspire high-school students to become involved in space-based research but will also help them develop key critical thinking and bioinformatics skills required for most college degrees and furthermore, will enable them to establish networks with their peers and connections with university Professors that may help them achieve their educational goals.

genelab↗

Mycobacterium Phage Butters-Encoded Proteins Contribute to Host Defense against Viral Attack [plus supplemental information]

A diverse set of prophage-mediated mechanisms protecting bacterial hosts from infection has been recently uncovered within cluster N mycobacteriophages isolated on the host, Mycobacterium smegmatis mc 2 155. In that context, we unveil a novel defense mechanism in cluster N prophage Butters. By using bioinformatics analyses, phage plating efficiency experiments, microscopy, and immunoprecipitation assays, we show that Butters genes located in the central region of the genome play a key role in the defense against heterotypic viral attack. Our study suggests that a two-component system, articulated by interactions between protein products of genes 30 and 31, confers defense against heterotypic phage infection by PurpleHaze (cluster A/subcluster A3) or Alma (cluster A/subcluster A9) but is insufficient to confer defense against attack by the heterotypic phage Island3 (cluster I/subcluster I1). Therefore, based on heterotypic phage plating efficiencies on the Butters lysogen, additional prophage genes required for defense are implicated and further show specificity of prophage-encoded defense systems. IMPORTANCE: Many sequenced bacterial genomes, including those of pathogenic bacteria, contain prophages. Some prophages encode defense systems that protect their bacterial host against heterotypic viral attack. Understanding the mechanisms undergirding these defense systems is crucial to appreciate the scope of bacterial immunity against viral infections and will be critical for better implementation of phage therapy that would require evasion of these defenses. Furthermore, such knowledge of prophage-encoded defense mechanisms may be useful for developing novel genetic tools for engineering phage-resistant bacteria of industrial importance.

59 BASIC BIOLOGICAL SCIENCES↗

Expanding Repository Data Available For Sharing and Knowledge Discovery

Some of the hardest space biology and space health challenges require data-intensive, bioinformatic, meta-analytical, and computer-assisted research approaches. These challenges include examining interdisciplinary space life science research across experiments and across interacting spaceflight hazards (radiation, altered gravity, confinement, hostile-closed environments, distance-duration from Earth). The approaches to confront these challenges involve mining multiple datasets simultaneously from various hierarchical organizations of biological complexity, all while concurrently evaluating how experimental design factors affect endpoints of standard assays. To enable this field, it is essential that principal investigators (PIs) submit data in a structure so it can be maximally re-used. The purpose of the NASA Ames Life Sciences Data Archive (ALSDA) is to collect, curate, and make publicly available all non-human space-relevant biological data. ALSDA must also ensure data are open-access, and maximally findable, accessible, interoperable, and reusable (FAIR). The scope of ALSDA data collected and submitted by PIs include subject and study design metadata, assay metadata parameters, raw and processed assay data, assay imagery/video, and subject-experienced mission data telemetry (radiation, temperature, humidity, acoustics, vibrations, etc.). ALSDA recently integrated into a collaborative group of Open Science projects to facilitate a suite of new tools and workflows that will improve data submission, accessibility, and reusability by implementing digital data submission agreements, and adopting the data management system originally developed by NASA GeneLab. ALSDA intends to bring current biological repository data and all future collected data into this new scientific data reuse reality. This new suite of tools will enable ALSDA to deploy a science curation system using scientific assay configurations for the data submission portal. It will capture essential assay parameters according to established standards in each sub-field within biology. The submission portal expedites data collection by enhancing ease of PI data submission, providing a user interface and specificity for which data is to be submitted. Data submissions can be brought into cutting-edge informatic analysis portals to enable mining of physiological, behavioral, biochemical, and imaging datasets in conjunction with ‘omics-level datasets. As ALSDA datasets are submitted, curated, and published (e.g., micro-computed tomography, histology, pulse oximetry, serum metabolites, magnetic resonance imaging, intraocular pressure, novel object recognition, etc.), the merging together of spaceflight data along this multi-hierarchical complexity of biology will enable informatics and data-intensive approaches resulting in knowledge discoveries across missions, space hazards, and biological disciplines.

Biology↗

Expanding Repository Data Available For Sharing And Knowledge Discovery

Some of the hardest space biology and space health challenges require data-intensive, bioinformatic, meta-analytical, and computer-assisted research approaches. These challenges include examining interdisciplinary space life science research across experiments and across interacting spaceflight hazards (radiation, altered gravity, confinement, hostile-closed environments, distance-duration from Earth). The approaches to confront these challenges involve mining multiple datasets simultaneously from various hierarchical organizations of biological complexity, all while concurrently evaluating how experimental design factors affect endpoints of standard assays. To enable this field, it is essential that principal investigators (PIs) submit data in a structure so it can be maximally re-used. The purpose of the NASA Ames Life Sciences Data Archive (ALSDA) is to collect, curate, and make publicly available all non-human space-relevant biological data. ALSDA must also ensure data are open-access, and maximally findable, accessible, interoperable, and reusable (FAIR). The scope of ALSDA data collected and submitted by PIs include subject and study design metadata, assay metadata parameters, raw and processed assay data, assay imagery/video, and subject-experienced mission data telemetry (radiation, temperature, humidity, acoustics, vibrations, etc.). ALSDA recently integrated into a collaborative group of Open Science projects to facilitate a suite of new tools and workflows that will improve data submission, accessibility, and reusability by implementing digital data submission agreements, and adopting the data management system originally developed by NASA GeneLab. ALSDA intends to bring current biological repository data and all future collected data into this new scientific data reuse reality. This new suite of tools will enable ALSDA to deploy a science curation system using scientific assay configurations for the data submission portal. It will capture essential assay parameters according to established standards in each sub-field within biology. The submission portal expedites data collection by enhancing ease of PI data submission, providing a user interface and specificity for which data is to be submitted. Data submissions can be brought into cutting-edge informatic analysis portals to enable mining of physiological, behavioral, biochemical, and imaging datasets in conjunction with ‘omics-level datasets. As ALSDA datasets are submitted, curated, and published (e.g., micro-computed tomography, histology, pulse oximetry, serum metabolites, magnetic resonance imaging, intraocular pressure, novel object recognition, etc.), the merging together of spaceflight data along this multi-hierarchical complexity of biology will enable informatics and data-intensive approaches resulting in knowledge discoveries across missions, space hazards, and biological disciplines.

life science↗

Evaluating the factors influencing accuracy, interpretability, and reproducibility in the use of machine learning classifiers in biology to enable standardization

The complexity and variability of biological data has promoted the increased use of machine learning methods to understand processes and predict outcomes. These same features complicate reliable, reproducible, interpretable, and responsible use of such methods, resulting in questionable relevance of the derived. outcomes. Here we systematically explore challenges associated with applying machine learning to predict and understand biological processes using a well- characterized in vitro experimental system. We evaluated factors that vary while applying machine learning classifers: (1) type of biochemical signature (transcripts vs. proteins), (2) data curation methods (pre- and post-processing), and (3) choice of machine learning classifier. Using accuracy, generalizability, interpretability, and reproducibility as metrics, we found that the above factors significantly mod- ulate outcomes even within a simple model system. Our results caution against the unregulated use of machine learning methods in the biological sciences, and strongly advocate the need for data standards and validation tool-kits for such studies.

59 BASIC BIOLOGICAL SCIENCES↗

Single-cell chromatin accessibility and cis -regulatory element analyses in plants using the scPlantReg platform

Understanding gene regulation is fundamental to plant improvement, but the lack of plant-specific single-cell assay for transposase-accessible chromatin using sequencing (scATAC-seq) frameworks and cross-species databases has limited insights into cell-type-specific cellular regulation. Here we present ‘scPlantReg’, an integrated framework and database for plant scATAC-seq data. scPlantReg supports end-to-end analyses from raw data processing to biological interpretation and features ‘scATACtor’, a supervised machine-learning approach that outperforms existing tools for cell-type annotation. We applied scPlantReg to pearl millet to characterize cell-type-specific chromatin accessibility and identify validated activating and repressing accessible chromatin regions (ACRs), revealing WRKY transcription factors as potential regulators of xylem development. Furthermore, we reanalysed scATAC-seq datasets from 8 plant species, spanning 11 tissues and multiple developmental stages, enabling cross-species comparisons. Furthermore, these analyses uncovered conserved regulatory programmes, including AP2/EREBP-associated ACRs linked to cell wall development and cell-type-conserved TFs across grasses. Collectively, scPlantReg provides a general framework and resource for comparative regulatory analysis in plants.

Epigenomics↗

BRAKER3: Fully automated genome annotation using RNA-seq and protein evidence with GeneMark-ETP, AUGUSTUS, and TSEBRA

Gene prediction has remained an active area of bioinformatics research for a long time. Still, gene prediction in large eukaryotic genomes presents a challenge that must be addressed by new algorithms. The amount and significance of the evidence available from transcriptomes and proteomes vary across genomes, between genes, and even along a single gene. User-friendly and accurate annotation pipelines that can cope with such data heterogeneity are needed. The previously developed annotation pipelines BRAKER1 and BRAKER2 use RNA-seq or protein data, respectively, but not both. A further significant performance improvement integrating all three data types was made by the recently released GeneMark-ETP. We here present the BRAKER3 pipeline that builds on GeneMark-ETP and AUGUSTUS, and further improves accuracy using the TSEBRA combiner. BRAKER3 annotates protein-coding genes in eukaryotic genomes using both short-read RNA-seq and a large protein database, along with statistical models learned iteratively and specifically for the target genome. We benchmarked the new pipeline on genomes of 11 species under an assumed level of relatedness of the target species proteome to available proteomes. BRAKER3 outperforms BRAKER1 and BRAKER2. The average transcript-level F1-score is increased by about 20 percentage points on average, whereas the difference is most pronounced for species with large and complex genomes. BRAKER3 also outperforms other existing tools, MAKER2, Funannotate, and FINDER. The code of BRAKER3 is available on GitHub and as a ready-to-run Docker container for execution with Docker or Singularity. Overall, BRAKER3 is an accurate, easy-to-use tool for eukaryotic genome annotation.

59 BASIC BIOLOGICAL SCIENCES↗

Metagenomic Methods for Addressing NASA's Planetary Protection Policy Requirements on Future Missions: A Workshop Report

Molecular biology methods and technologies have advanced substantially over the past decade. These new molecular methods should be incorporated among the standard tools of planetary protection (PP) and could be validated for incorporation by 2026. To address the feasibility of applying modern molecular techniques to such an application, NASA conducted a technology workshop with private industry partners, academics, and government agency stakeholders, along with NASA staff and contractors. The technical discussions and presentations of the Multi-Mission Metagenomics Technology Development Workshop focused on modernizing and supplementing the current PP assays. The goals of the workshop were to assess the state of metagenomics and other advanced molecular techniques in the context of providing a validated framework to supplement the bacterial endospore-based NASA Standard Assay and to identify knowledge and technology gaps. In particular, workshop participants were tasked with discussing metagenomics as a stand-alone technology to provide rapid and comprehensive analysis of total nucleic acids and viable microorganisms on spacecraft surfaces, thereby allowing for the development of tailored and cost-effective microbial reduction plans for each hardware item on a spacecraft. Workshop participants recommended metagenomics approaches as the only data source that can adequately feed into quantitative microbial risk assessment models for evaluating the risk of forward (exploring extraterrestrial planet) and back (Earth harmful biological) contamination. Participants were unanimous that a metagenomics workflow, in tandem with rapid targeted quantitative (digital) PCR, represents a revolutionary advance over existing methods for the assessment of microbial bioburden on spacecraft surfaces. The workshop highlighted low biomass sampling, reagent contamination, and inconsistent bioinformatics data analysis as key areas for technology development. Finally, it was concluded that implementing metagenomics as an additional workflow for addressing concerns of NASA's robotic mission will represent a dramatic improvement in technology advancement for PP and will benefit future missions where mission success is affected by backward and forward contamination.

59 BASIC BIOLOGICAL SCIENCES↗

The endohyphal microbiome: current progress and challenges for scaling down integrative multi-omic microbiome research

Abstract As microbiome research has progressed, it has become clear that most, if not all, eukaryotic organisms are hosts to microbiomes composed of prokaryotes, other eukaryotes, and viruses. Fungi have only recently been considered holobionts with their own microbiomes, as filamentous fungi have been found to harbor bacteria (including cyanobacteria), mycoviruses, other fungi, and whole algal cells within their hyphae. Constituents of this complex endohyphal microbiome have been interrogated using multi-omic approaches. However, a lack of tools, techniques, and standardization for integrative multi-omics for small-scale microbiomes (e.g., intracellular microbiomes) has limited progress towards investigating and understanding the total diversity of the endohyphal microbiome and its functional impacts on fungal hosts. Understanding microbiome impacts on fungal hosts will advance explorations of how “microbiomes within microbiomes” affect broader microbial community dynamics and ecological functions. Progress to date as well as ongoing challenges of performing integrative multi-omics on the endohyphal microbiome is discussed herein. Addressing the challenges associated with the sample extraction, sample preparation, multi-omic data generation, and multi-omic data analysis and integration will help advance current knowledge of the endohyphal microbiome and provide a road map for shrinking microbiome investigations to smaller scales.

59 BASIC BIOLOGICAL SCIENCES↗

Visualizing group II intron dynamics between the first and second steps of splicing

Group II introns are ubiquitous self-splicing ribozymes and retrotransposable elements evolutionarily and chemically related to the eukaryotic spliceosome, with potential applications as gene-editing tools. Recent biochemical and structural data have captured the intron in multiple conformations at different stages of catalysis. Here, we employ enzymatic assays, X-ray crystallography, and molecular simulations to resolve the spatiotemporal location and function of conformational changes occurring between the first and the second step of splicing. We show that the first residue of the highly-conserved catalytic triad is protonated upon 5’-splice-site scission, promoting a reversible structural rearrangement of the active site (toggling). Protonation and active site dynamics induced by the first step of splicing facilitate the progression to the second step. Our insights into the mechanism of group II intron splicing parallels functional data on the spliceosome, thus reinforcing the notion that these evolutionarily-related molecular machines share the same enzymatic strategy.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

PDB-IHM: A System for Deposition, Curation, Validation, and Dissemination of Integrative Structures

Structures of many large biomolecular assemblies are now being determined using integrative approaches. In these approaches, information derived from multiple experimental and computational methods is combined to compute three-dimensional structures of multi-protein complexes and other macromolecular machines. A standalone prototype data resource for integrative structures called PDB-Dev was built, based on recommendations of the Integrative and Hybrid Methods (IHM) Task Force of the Worldwide Protein Data Bank (wwPDB). This effort included developing data standards and software tools for collecting, curating, validating, visualizing, archiving, and disseminating integrative structures that span diverse spatiotemporal scales and conformational states. Mechanisms have been created to validate integrative structures based on the experimental data underpinning them. Building upon this foundational framework, PDB-Dev has been further expanded to handle large dynamic macromolecular systems and integrative structures that combine, for example, experimental restraints with atomic coordinates computed by machine learning algorithms. Data standards and supporting tools have also been extended to capture information about biomolecular dynamics, such as conformational transitions and related kinetic data derived from biophysical methods. Recently, PDB-Dev was unified with the PDB archive and rebranded as PDB-IHM (pdb-ihm.org), further promoting FAIR (Findable, Accessible, Interoperable, and Reusable) principles of data stewardship for integrative structural biology.

IHMCIF↗

Transcript and blood-microbiome analysis towards a blood diagnostic tool for goats affected by Haemonchus contortus

The Alpine goat (Capra aegagrus hircus) is parasitized by the barber pole worm (Haemonchus contortus). Hematological parameters from transcript and metagenome analysis in the host are reflective of infestation. We explored comparisons between blood samples of control, infected, infected zoledronic acid-treated, and infected antibody (anti-γδ T cells) treated wethers under controlled conditions. Seven days post-inoculation (dpi), we identified 7,627 transcripts associated with the different treatment types. Microbiome measurements at 7 dpi revealed fewer raw read counts across all treatments and a less diverse microbial flora than at 21 dpi. This study identifies treatment specific transcripts and an increase in microflora abundance and diversity as wethers age. Further, F/B ratio reflect health, based on depression or elevation above thresholds defined by the baseline of non-infected controls. Forty Alpine wethers were studied where blood samples were collected from five goats in four treatment groups on 7 dpi and 21 dpi. Transcript and microbiome profiles were obtained using the Partek Flow (St. Louis, Missouri, USA) software suites pipelines. Inflammation comparisons were based on the Firmicutes/Bacteriodetes ratios that are calculated as well as the reduction of microbial diversity.

60 APPLIED LIFE SCIENCES↗

Association of Diet and Antimicrobial Resistance in Healthy U.S. Adults

Antimicrobial resistance (AMR) represents a significant source of morbidity and mortality worldwide, with expectations that AMR-associated consequences will continue to worsen throughout the coming decades. Since resistance to antibiotics is encoded in the microbiome, interventions aimed at altering the taxonomic composition of the gut might allow us to prophylactically engineer microbiomes that harbor fewer antibiotic resistant genes (ARGs). Diet is one method of intervention, and yet little is known about the association between diet and antimicrobial resistance. To address this knowledge gap, we examined diet using the food frequency questionnaire (FFQ; habitual diet) and 24-h dietary recalls (Automated Self-Administered 24-h [ASA24 ® ] tool) coupled with an analysis of the microbiome using shotgun metagenome sequencing in 290 healthy adult participants of the United States Department of Agriculture (USDA) Nutritional Phenotyping Study. We found that aminoglycosides were the most abundant and prevalent mechanism of AMR in these healthy adults and that aminoglycoside-O-phosphotransferases (aph3-dprime) correlated negatively with total calories and soluble fiber intake. Individuals in the lowest quartile of ARGs (low-ARG) consumed significantly more fiber in their diets than medium- and high-ARG individuals, which was concomitant with increased abundances of obligate anaerobes, especially from the family Clostridiaceae, in their gut microbiota. Finally, we applied machine learning to examine 387 dietary, physiological, and lifestyle features for associations with antimicrobial resistance, finding that increased phylogenetic diversity of diet was associated with low-ARG individuals. These data suggest diet may be a potential method for reducing the burden of AMR.

59 BASIC BIOLOGICAL SCIENCES↗

Cross-reactive immunogenicity of group A streptococcal vaccines designed using a recurrent neural network to identify conserved M protein linear epitopes

The M protein of group A streptococci (Strep A) is a major virulence determinant and protective antigen. The N-terminal sequence of the protein defines the more than 200 M types of Strep A and also contains epitopes that elicit opsonic antibodies, some of which cross-react with heterologous M types. Current efforts to develop broadly protective M protein-based vaccines are directed at identifying potential cross-protective epitopes located in the N-terminal regions of cluster-related M proteins for use as vaccine antigens. In this study, we have used a comprehensive approach using the recurrent neural network ABCpred and IEDB epitope conservancy analysis tools to predict 16 residue linear B-cell epitopes from 117 clinically relevant M types of Strep A (~88% of global Strep A infections). Furthermore, to examine the immunogenicity of these epitope-based vaccines, nine peptides that together shared ≥60% sequence identity with 37 heterologous M proteins were incorporated into two recombinant hybrid protein vaccines, in which the epitopes were repeated 2 or 3 times, respectively. The combined immune responses of immunized rabbits showed that the vaccines elicited significant levels of antibodies against all nine vaccine epitopes present in homologous N-terminal 1–50 amino acid synthetic M peptides, as well as cross-reactive antibodies against 16 of 37 heterologous M peptides predicted to contain similar epitopes. The epitope-specificity of the cross-reactive antibodies was confirmed by ELISA inhibition assays and functional opsonic activity was assayed in HL-60-based bactericidal assays. The results provide important information for the future design of broadly protective M protein-based Strep A vaccines.

60 APPLIED LIFE SCIENCES↗

Computational Basis for On-Demand Production of Diversified Therapeutic Phage Cocktails

New therapies are necessary to combat increasingly antibiotic-resistant bacterial pathogens. We have developed a technology platform of computational, molecular biology, and microbiology tools which together enable on-demand production of phages that target virtually any given bacterial isolate. Two complementary computational tools that identify and precisely map prophages and other integrative genetic elements in bacterial genomes are used to identify prophage-laden bacteria that are close relatives of the target strain. Phage genomes are engineered to disable lysogeny, through use of long amplicon PCR and Gibson assembly. Finally, the engineered phage genomes are introduced into host bacteria for phage production. As an initial demonstration, we used this approach to produce a phage cocktail against the opportunistic pathogen Pseudomonas aeruginosa PAO1. Two prophageladen P. aeruginosa strains closely related to PAO1 were identified, ATCC 39324 and ATCC 27853. Deep sequencing revealed that mitomycin C treatment of these strains induced seven phages that grow on P. aeruginosa PAO1. The most diverse five phages were engineered for nonlysogeny by deleting the integrase gene (int), which is readily identifiable and typically conveniently located at one end of the prophage. The Δint phages, individually and in cocktails, killed P. aeruginosa PAO1 in liquid culture as well as in a waxworm (Galleria mellonella) model of infection.

59 BASIC BIOLOGICAL SCIENCES↗

Leveraging Structured Biological Knowledge for Counterfactual Inference: A Case Study of Viral Pathogenesis

Counterfactual inference is a useful tool for comparing outcomes of interventions on complex systems. It requires us to represent the system in form of a structural causal model, complete with a causal diagram, probabilistic assumptions on exogenous variables, and functional assignments. Specifying such models can be extremely difficult in practice. The process requires substantial domain expertise, and does not scale easily to large systems, multiple systems, or novel system modifications. At the same time, many application domains, such as molecular biology, are rich in structured causal knowledge that is qualitative in nature. This manuscript proposes a general approach for querying a causal knowledge graph with a causal question and converting the qualitative result into a quantitative structural causal model that can learn from data to answer the question. Here, we demonstrate the feasibility, accuracy and versatility of this approach using two case studies in systems biology. The first demonstrates the appropriateness of the underlying assumptions and the accuracy of the results. The second demonstrates the versatility of the approach by querying a knowledge base for the molecular determinants of a severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2)-induced cytokine storm and performing counterfactual inference to predict the causal effect of medical countermeasures for severely ill COVID-19 patients.

60 APPLIED LIFE SCIENCES↗