Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “bioinformatics tool”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Targeted metagenomic assessment reflects critical colonization in battlefield injuries

Current diagnostics and clinical management strategies for combat wounds are based on decisions made by expert clinicians. However, even in the hands of experienced surgeons, wounds from combat injuries can exhibit failed healing and complications related to limitations in the rapid and comprehensive generation of diagnostic information. Previous studies have demonstrated the possible use of genomic sequencing approaches to detect microbial signatures involved in combat casualty care. While effective, whole metagenome sequencing is limited by the depth required to confidently detect all relevant signatures. To address this, we developed a targeted capture sequencing panel to detect microbial signatures relevant to wound healing. These targets include known microbial nosocomial pathogens, wound colonizers, and genes involved in virulence and antimicrobial resistance. A bioinformatics pipeline was built to identify genomic regions of interest and over 8,000 oligonucleotide probes were designed for capture. The panel was synthesized and validated using control reference genomes in human background and on wound-effluent samples from a cohort of combat-injured U.S. service members. Our panel was sensitive against wound-colonizing species, Acinetobacter baumannii and Pseudomonas aeruginosa, and was specific in detecting corresponding virulence and antimicrobial-resistance genes as well as other pathogenic species present in microflora mixtures. Random forest feature permutation confirmed the prevalence of Acinetobacter and Pseudomonas in critically colonized wounds and wounds that failed to heal, respectively. Our results demonstrate the capability of targeted sequencing tools and analysis platforms to profile and deliver information on pathogenic factors influencing wound progression, thereby guiding therapeutic intervention.

59 BASIC BIOLOGICAL SCIENCES↗

$\mathrm{CROPSR}$: an automated platform for complex genome-wide $\mathrm{CRISPR}$ g$\mathrm{RNA}$ design and validation

CRISPR/Cas9 technology has become an important tool to generate targeted, highly specific genome mutations. The technology has great potential for crop improvement, as crop genomes are tailored to optimize specific traits over generations of breeding. Many crops have highly complex and polyploid genomes, particularly those used for bioenergy or bioproducts. The majority of tools currently available for designing and evaluating gRNAs for CRISPR experiments were developed based on mammalian genomes that do not share the characteristics or design criteria for crop genomes. We have developed an open source tool for genome-wide design and evaluation of gRNA sequences for CRISPR experiments, CROPSR. The genome-wide approach provides a significant decrease in the time required to design a CRISPR experiment, including validation through PCR, at the expense of an overhead compute time required once per genome, at the first run. To better cater to the needs of crop geneticists, restrictions imposed by other packages on design and evaluation of gRNA sequences were lifted. A new machine learning model was developed to provide scores while avoiding situations in which the currently available tools sometimes failed to provide guides for repetitive, A/T-rich genomic regions. We show that our gRNA scoring model provides a significant increase in prediction accuracy over existing tools, even in non-crop genomes. CROPSR provides the scientific community with new methods and a new workflow for performing CRISPR/Cas9 knockout experiments. CROPSR reduces the challenges of working in crops, and helps speed gRNA sequence design, evaluation and validation. We hope that the new software will accelerate discovery and reduce the number of failed experiments.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

RCSB protein data Bank: Next‐generation advanced search for exploration of experimental structures and computed structure models

Abstract The Protein Data Bank (PDB), established in 1971, is the primary global, open‐access archive for experimentally determined 3D macromolecular structures (proteins, RNA, DNA). The research‐focused RCSB.org web‐portal provides access to these data alongside more than one million machine‐learning‐predicted structure models, greatly expanding the available structural landscape. Rapid growth of both experimental and computational structures has increased the need for powerful yet accessible search tools that serve a broad and diverse scientific community. Herein, we describe a redesigned RCSB Protein Data Bank RCSB.org Advanced Search capability that supports intuitive discovery of 3D structures through a unified interface. This interface integrates annotation‐, sequence‐, and 3D structure‐based searches, embeds an interactive 3D viewer, and incorporates curated biological knowledge, such as catalytic site definitions from Mechanism and Catalytic Site Atlas and ligand‐guided structural motifs, for constructing geometry‐driven queries. A new Chemical Search tool allows definition of chemical queries via an integrated drawing tool or standard identifiers, seamlessly combining them with annotation filters. By allowing query definition directly within spatial and chemical contexts, these search interfaces reduce the need for detailed knowledge of residue numbering, chain identifiers, or external cheminformatics software. This capability enables efficient exploration of structures, chemical diversity, and structure–function relationships across all life domains. The redesigned interfaces can be accessed directly at rcsb.org/search/advanced for Advanced Search and rcsb.org/search/chemical for Chemical Search.

Rose, Yana [Research Collaboratory for Structural ↗

SBbadger: biochemical reaction networks with definable degree distributions

Abstract Motivation An essential step in developing computational tools for the inference, optimization and simulation of biochemical reaction networks is gauging tool performance against earlier efforts using an appropriate set of benchmarks. General strategies for the assembly of benchmark models include collection from the literature, creation via subnetwork extraction and de novo generation. However, with respect to biochemical reaction networks, these approaches and their associated tools are either poorly suited to generate models that reflect the wide range of properties found in natural biochemical networks or to do so in numbers that enable rigorous statistical analysis. Results In this work, we present SBbadger, a python-based software tool for the generation of synthetic biochemical reaction or metabolic networks with user-defined degree distributions, multiple available kinetic formalisms and a host of other definable properties. SBbadger thus enables the creation of benchmark model sets that reflect properties of biological systems and generate the kinetics and model structures typically targeted by computational analysis and inference software. Here, we detail the computational and algorithmic workflow of SBbadger, demonstrate its performance under various settings, provide sample outputs and compare it to currently available biochemical reaction network generation software. Availability and implementation SBbadger is implemented in Python and is freely available at https://github.com/sys-bio/SBbadger and via PyPI at https://pypi.org/project/SBbadger/. Documentation can be found at https://SBbadger.readthedocs.io. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

Energy landscapes from cryo-EM snapshots: a benchmarking study

Abstract Biomolecules undergo continuous conformational motions, a subset of which are functionally relevant. Understanding, and ultimately controlling biomolecular function are predicated on the ability to map continuous conformational motions, and identify the functionally relevant conformational trajectories. For equilibrium and near-equilibrium processes, function proceeds along minimum-energy pathways on one or more energy landscapes, because higher-energy conformations are only weakly occupied. With the growing interest in identifying functional trajectories, the need for reliable mapping of energy landscapes has become paramount. In response, various data-analytical tools for determining structural variability are emerging. A key question concerns the veracity with which each data-analytical tool can extract functionally relevant conformational trajectories from a collection of single-particle cryo-EM snapshots. Using synthetic data as an independently known ground truth, we benchmark the ability of four leading algorithms to determine biomolecular energy landscapes and identify the functionally relevant conformational paths on these landscapes. Such benchmarking is essential for systematic progress toward atomic-level movies of continuous biomolecular function.

59 BASIC BIOLOGICAL SCIENCES↗

Using low-coverage whole genome sequencing (genome skimming) to delineate three introgressed species of buffalofish ( Ictiobus )

Consumption of buffalofish has been sporadically associated with Haff disease-like illnesses involving sudden onset muscle pain and weakness due to skeletal muscle rhabdomyolysis, but determination of precisely which species are associated with these illnesses has been impeded by a lack of species-specific DNA-based markers. Here, three closely related species of buffalofish native to the Mississippi River Basin (Ictiobus bubalus, Ictiobus cyprinellus and Ictiobus niger) that have previously proven genetically indistinguishable using both mitochondrial and nuclear single-locus sequencing were reliably discriminated using low-coverage whole genome sequencing (‘genome skimming’). Using 44 specimens representing the three species collected from the mid/upper (Missouri) and lower (Louisiana) regions of the species’ native ranges, the SISRS (Site Identification from Short Read Sequences) bioinformatics pipeline was adapted to (1) identify over 620Mbp of putatively homologous nuclear sequence data and (2) isolate over 140,000 single-nucleotide polymorphisms (SNPs) that supported accurate species delimitation, all without the use of a reference genome or annotation data. These sites were used to classify Ictiobus spp. samples with genome-skim data, along with a larger set (n = 67) where ultraconserved elements (UCEs) were sequenced. Analyses of whole mitochondrial data revealed more limited signal. Nearly all samples matched their purported species based on morphologic identification, but two Missouri samples morphologically identified as I. niger grouped with samples of I. bubalus, albeit with significant enrichment of I. niger SNPs. To our knowledge this is the first report of a DNA-based tool to reliably discriminate these three morphologically distinct species.

59 BASIC BIOLOGICAL SCIENCES↗

EcoPLOT: dynamic analysis of biogeochemical data

Motivation: We have created EcoPLOT (parameterized linkage of omics-driven technologies), a web-app for the dynamic, interactive analysis of biogeochemical datasets that combines state-of-the-art analysis tools to statistically and graphically explore environmental, geochemical and microbiome datasets. Using the iterative random forest, a machine learning algorithm, EcoPLOT allows for the de novo discovery of drivers which exhibit significant impact on plant, microbial or soil dynamics. Availability and implementation: EcoPLOT is built entirely within the R language. It can be accessed through any system where R is installed, including Windows, Mac and most Linux systems. EcoPLOT is free to use and can be accessed at https://github.com/cdsanchez18/EcoPLOT.

59 BASIC BIOLOGICAL SCIENCES↗

HDBind: encoding of molecular structure with hyperdimensional binary representations

Traditional methods for identifying “hit” molecules from a large collection of potential drug-like candidates rely on biophysical theory to compute approximations to the Gibbs free energy of the binding interaction between the drug and its protein target. These approaches have a significant limitation in that they require exceptional computing capabilities for even relatively small collections of molecules. Increasingly large and complex state-of-the-art deep learning approaches have gained popularity with the promise to improve the productivity of drug design, notorious for its numerous failures. However, as deep learning models increase in their size and complexity, their acceleration at the hardware level becomes more challenging. Hyperdimensional Computing (HDC) has recently gained attention in the computer hardware community due to its algorithmic simplicity relative to deep learning approaches. The HDC learning paradigm, which represents data with high-dimension binary vectors, allows the use of low-precision binary vector arithmetic to create models of the data that can be learned without the need for the gradient-based optimization required in many conventional machine learning and deep learning methods. This algorithmic simplicity allows for acceleration in hardware that has been previously demonstrated in a range of application areas (computer vision, bioinformatics, mass spectrometery, remote sensing, edge devices, etc.). To the best of our knowledge, our work is the first to consider HDC for the task of fast and efficient screening of modern drug-like compound libraries. We also propose the first HDC graph-based encoding methods for molecular data, demonstrating consistent and substantial improvement over previous work. We compare our approaches to alternative approaches on the well-studied MoleculeNet dataset and the recently proposed LIT-PCBA dataset derived from high quality PubChem assays. We demonstrate our methods on multiple target hardware platforms, including Graphics Processing Units (GPUs) and Field Programmable Gate Arrays (FPGAs), showing at least an order of magnitude improvement in energy efficiency versus even our smallest neural network baseline model with a single hidden layer. Our work thus motivates further investigation into molecular representation learning to develop ultra-efficient pre-screening tools. We make our code publicly available at https://github.com/LLNL/hdbind.

59 BASIC BIOLOGICAL SCIENCES↗

Evaluation of Correction Methods for NASA GeneLab Transcriptomic Datasets

Conducting space biology experiments aboard the International Space Station, particularly those utilizing complex model organisms like mice, is expensive and difficult due to limited crew availability, hardware, and space. As a result, sample numbers from these studies are low, reducing the statistical power of any one experiment. Aggregating spaceflight datasets serves as a method to increase sample numbers, allowing for novel insights through bioinformatic analysis of ‘omics data from merged datasets. However, aggregating datasets can introduce unwanted variation including 1) differences in sample handling, processing, and sequencing platforms between datasets (technical variation) as well as 2) differences in experimental design between datasets. In the present study, NASA GeneLab-hosted RNAseq datasets from mouse liver tissues were used to evaluate several statistical methods to correct for this unwanted variation through two approaches, reference-based and standard. The following correction algorithms were applied with (reference-based) and/or without (standard) considering Universal Mouse RNA Reference samples: ComBat and ComBat_seq from the SVA package, median polish, empirical Bayes, and ANOVA-based algorithms from the MBatch package, and negative binomial regression normalization in the DESeq2 package. For each approach, after the correction algorithm was applied, differential gene expression (DGE) analysis of flight and ground control samples was performed with the combined data. The robustness of each tool was evaluated using BatchQC to determine statistical differences between datasets before and after correction, Principal Component Analysis to evaluate global gene expression in samples before and after correction, and by comparing DGE analysis of individual datasets and combined datasets before and after correction. The results showed that the reference-based approach introduced several additional (and likely artificial) DEGs when compared with the respective standard approach. Of the methods tested, standard ComBat and DESeq2 were identified as the most robust correction methods for combining spaceflight mouse liver RNAseq datasets hosted on GeneLab.

GeneLab↗

Combining RNA-SEQ Datasets from NASA GENELAB: An Evaluation of Correction Methods

Background: Conducting space biology experiments aboard the International Space Station, particularly those utilizing complex model organisms like mice, is expensive and difficult due to limited crew availability, hardware, and space. As a result, sample numbers from these studies are low, reducing the statistical power of any one experiment. Aggregating spaceflight datasets serves as a method to increase sample numbers, allowing for novel insights through bioinformatic analysis of ‘omics data from merged datasets. However, aggregating datasets can introduce unwanted variation including 1) differences in sample handling, processing, and sequencing platforms between datasets (technical variation) as well as 2) differences in experimental design between datasets. Methods: In the present study, NASA GeneLab-hosted RNAseq datasets from mouse liver tissues were used to evaluate several statistical methods to correct for this unwanted variation through two approaches, reference-based and standard. The following correction algorithms were applied with (reference-based) and/or without (standard) considering Universal Mouse RNA Reference samples: ComBat and ComBat_seq from the SVA package, the median polish, empirical Bayes, and ANOVA-based algorithms from the MBatch package, and negative binomial regression normalization in the DESeq2 package. For each approach, after the correction algorithm was applied, differential gene expression (DGE) analysis of flight and ground control samples was performed with the combined data. The robustness of each tool was evaluated using BatchQC to determine statistical differences between datasets before and after correction, Principal Component Analysis to evaluate global gene expression in samples before and after correction, and by comparing DGE analysis of individual datasets and combined datasets before and after correction. Results: The results showed that the reference-based approach introduced several additional (and likely artificial) differentially expressed genes when compared with the respective standard approach. Conclusions: Of the methods tested, standard ComBat_seq and DESeq2 were identified as the most robust correction methods for combining spaceflight mouse liver RNAseq datasets hosted on GeneLab.

Finsam Samson↗

Bottom-Up Simulation, Reconstruction, and Quantification of Macromolecule Sequences from Experimental Polymerizations

Motivated by the canonical sequence–structure–function paradigm, tools to characterize chemical patterning in natural biomacromolecules, from proteins to nucleic acids, have grown exponentially in recent years. However, analogous strategies for synthetic macromolecules remain in nascent stages, complicated by sequence polydispersity and analytical limitations. To address this, we have developed a comprehensive and open-source Python package, PRISM (polymer rate insights and sequence modeling), an end-to-end workflow that provides a path from experimental kinetics measurements to quantitative and qualitative metrics for describing chemical patterning in stochastic polymers. First, a numerical integration strategy was constructed to simulate and fit experimental data from reversible addition–fragmentation chain transfer (RAFT) polymerization kinetics, enabling the facile estimation of relevant reactivity ratios. These ratios were then used in a mechanism-specific stochastic kinetic simulation strategy to simulate sequence ensembles corresponding to model systems spanning experimental copolymers, classes of statistical polymers (e.g., alternating, block, and gradient), and multiblock copolymers. Lastly, inspired by sequence homology metrics from bioinformatics, we introduce visualization strategies and quantitative metrics to facilitate comparisons of different sequence ensembles. As the sequence–structure–function paradigm becomes increasingly central in de novo design of synthetic macromolecules, this toolkit provides a first step toward accurate and representative sequence description and featurization.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Evaluation of Correction Methods for NASA GeneLab Transcriptomic Datasets

Conducting space biology experiments aboard the International Space Station, particularly those utilizing complex model organisms like mice, is expensive and difficult due to limited crew availability, hardware, and space. As a result, sample numbers from these studies are low, reducing the statistical power of any one experiment. Aggregating spaceflight datasets serves as a method to increase sample numbers, allowing for novel insights through bioinformatic analysis of ‘omics data from merged datasets. However, aggregating datasets can introduce unwanted variation including 1) differences in sample handling, processing, and sequencing platforms between datasets (technical variation) as well as 2) differences in experimental design between datasets such as sex or age of the model organism used. In the present study, NASA GeneLab-hosted RNAseq datasets from rodent liver tissues were used to evaluate several statistical methods to correct for this unwanted variation through two approaches, reference-based and standard. The following correction algorithms were applied with (reference-based) and/or without (standard) considering Universal Mouse RNA Reference samples: ComBat and ComBat_seq from the SVA package, median polish, empirical Bayes, and ANOVA-based algorithms from the MBatch package, and negative binomial regression normalization in the DESeq2 package. For each approach, after the correction algorithm was applied, differential gene expression (DGE) analysis of flight and ground control samples was performed with the combined data. The robustness of each tool was evaluated using BatchQC, to determine statistical differences between datasets before and after correction, Principal Component Analysis, to evaluate global gene expression in samples before and after correction, and by comparing DGE analysis of individual datasets and combined datasets before and after correction. The results showed that the reference-based approach introduced several additional (and likely artificial) DEGs when compared with the standard approach. Thus, the most robust standard correction will be implemented in the GeneLab Visualization 2.0 platform when datasets are combined.

GeneLab, RNA-seq, Batch Correction↗

Open-Source and FAIR Research Software for Proteomics

Scientific discovery relies on innovative software as much as experimental methods, especially in proteomics, where computational tools are essential for mass spectrometer setup, data analysis, and interpretation. Since the introduction of SEQUEST, proteomics software has grown into a complex ecosystem of algorithms, predictive models, and workflows, but the field faces challenges, including the increasing complexity of mass spectrometry data, limited reproducibility due to proprietary software, and difficulties integrating with other omics disciplines. Closed-source, platform-specific tools exacerbate these issues by restricting innovation, creating inefficiencies, and imposing hidden costs on the community. Open-source software (OSS), aligned with the FAIR Principles (Findable, Accessible, Interoperable, Reusable), offers a solution by promoting transparency, reproducibility, and community-driven development, which fosters collaboration and continuous improvement. In this manuscript, we explore the role of OSS in computational proteomics, its alignment with FAIR principles, and its potential to address challenges related to licensing, distribution, and standardization. Drawing on lessons from other omics fields, we present a vision for a future where OSS and FAIR principles underpin a transparent, accessible, and innovative proteomics community.

97 MATHEMATICS AND COMPUTING↗

kb_DRAM: annotation and metabolic profiling of genomes with DRAM in KBase

Microbial genome annotation is the process of identifying structural and functional elements in DNA sequences and subsequently attaching biological information to those elements. DRAM is a tool developed to annotate bacterial, archaeal, and viral genomes derived from pure cultures or metagenomes. DRAM goes beyond traditional annotation tools by distilling multiple gene annotations to genome level summaries of functional potential. Despite these benefits, a downside of DRAM is the requirement of large computational resources, which limits its accessibility. Further, it did not integrate with downstream metabolic modeling tools that require genome annotation. To alleviate these constraints, DRAM and the viral counterpart, DRAM-v, are now available and integrated with the freely accessible KBase cyberinfrastructure. With kb_DRAM users can generate DRAM annotations and functional summaries from microbial or viral genomes in a point-and-click interface, as well as generate genome-scale metabolic models from DRAM annotations.

59 BASIC BIOLOGICAL SCIENCES↗

Expanding standards in viromics: in silico evaluation of dsDNA viral genome identification, classification, and auxiliary metabolic gene curation

Viruses influence global patterns of microbial diversity and nutrient cycles. Though viral metagenomics (viromics), specifically targeting dsDNA viruses, has been critical for revealing viral roles across diverse ecosystems, its analyses differ in many ways from those used for microbes. To date, viromics benchmarking has covered read pre-processing, assembly, relative abundance, read mapping thresholds and diversity estimation, but other steps would benefit from benchmarking and standardization. Here we use in silico-generated datasets and an extensive literature survey to evaluate and highlight how dataset composition (i.e., viromes vs bulk metagenomes) and assembly fragmentation impact (i) viral contig identification tool, (ii) virus taxonomic classification, and (iii) identification and curation of auxiliary metabolic genes (AMGs). The in silico benchmarking of five commonly used virus identification tools show that gene-content-based tools consistently performed well for long (≥3 kbp) contigs, while k -mer- and blast-based tools were uniquely able to detect viruses from short (≤3 kbp) contigs. Notably, however, the performance increase of k -mer- and blast-based tools for short contigs was obtained at the cost of increased false positives (sometimes up to ~5% for virome and ~75% bulk samples), particularly when eukaryotic or mobile genetic element sequences were included in the test datasets. Furthermore, for viral classification, variously sized genome fragments were assessed using gene-sharing network analytics to quantify drop-offs in taxonomic assignments, which revealed correct assignations ranging from ~95% (whole genomes) down to ~80% (3 kbp sized genome fragments). A similar trend was also observed for other viral classification tools such as VPF-class, ViPTree and VIRIDIC, suggesting that caution is warranted when classifying short genome fragments and not full genomes. Finally, we highlight how fragmented assemblies can lead to erroneous identification of AMGs and outline a best-practices workflow to curate candidate AMGs in viral genomes assembled from metagenomes. Together, these benchmarking experiments and annotation guidelines should aid researchers seeking to best detect, classify, and characterize the myriad viruses ‘hidden’ in diverse sequence datasets.

59 BASIC BIOLOGICAL SCIENCES↗

NASA GeneLab Concept of Operations

NASA's GeneLab aims to greatly increase the number of scientists that are using data from space biology investigations on board ISS, emphasizing a systems biology approach to the science. When completed, GeneLab will provide the integrated software and hardware infrastructure, analytical tools and reference datasets for an assortment of model organisms. GeneLab will also provide an environment for scientists to collaborate thereby increasing the possibility for data to be reused for future experimentation. To maximize the value of data from life science experiments performed in space and to make the most advantageous use of the remaining ISS research window, GeneLab will apply an open access approach to conducting spaceflight experiments by generating, and sharing the datasets derived from these biological studies in space.Onboard the ISS, a wide variety of model organisms will be studied and returned to Earth for analysis. Laboratories on the ground will analyze these samples and provide genomic, transcriptomic, metabolomic and proteomic data. Upon receipt, NASA will conduct data quality control tasks and format raw data returned from the omics centers into standardized, annotated information sets that can be readily searched and linked to spaceflight metadata. Once prepared, the biological datasets, as well as any analysis completed, will be made public through the GeneLab Space Bioinformatics System webb as edportal. These efforts will support a collaborative research environment for spaceflight studies that will closely resemble environments created by the Department of Energy (DOE), National Center for Biotechnology Information (NCBI), and other institutions in additional areas of study, such as cancer and environmental biology. The results will allow for comparative analyses that will help scientists around the world take a major leap forward in understanding the effect of microgravity, radiation, and other aspects of the space environment on model organisms. These efforts will speed the process of scientific sharing, iteration, and discovery.

Space Life Science↗

Development of platforms for functional characterization and production of phenazines using a multi-chassis approach via CRAGE

Phenazines (Phzs), a family of chemicals with a phenazine backbone, are secondary metabolites with diverse properties such as antibacterial, anti-fungal, or anticancer activity. The core derivatives of phenazine, phenazine-1-carboxylic acid (PCA) and phenazine-1,6-dicarboxylic acid (PDC), are themselves precursors for various other derivatives. Recent advances in genome mining tools have enabled researchers to identify many biosynthetic gene clusters (BGCs) that might produce novel Phzs. Here, to characterize the function of these BGCs efficiently, we performed modular construct assembly and subsequent multi-chassis heterologous expression using chassis-independent recombinase-assisted genome engineering (CRAGE). CRAGE allowed rapid integration of a PCA BGC into 23 diverse γ-proteobacteria species and allowed us to identify top PCA producers. We then used the top five chassis hosts to express four partially refactored PDC BGCs. A few of these platforms produced high levels of PDC. Specifically, Xenorhabdus doucetiae and Pseudomonas simiae produced PDC at a titer of 293 mg/L and 373 mg/L, respectively, in minimal media. These titers are significantly higher than those previously reported. Furthermore, selectivity toward PDC production over PCA production was improved by up to 9-fold. The results show that these strains are promising chassis for production of PCA, PDC, and their derivatives, as well as for function characterization of Phz BGCs identified via bioinformatics mining.

54 ENVIRONMENTAL SCIENCES↗

scPlantAnnotate: an accurate and robust transformer-based model for plant cell type annotation

Accurate cell type annotation remains a major bottleneck in plant single-cell RNA sequencing (scRNA-seq), where existing tools are often adapted from animal studies and perform sub-optimally on plant data. The lack of plant-specific computational frameworks limits the construction of plant cell atlases and downstream biological discovery. We develop and evaluate scPlantAnnotate, a Transformer-based reference annotation framework tailored for plant scRNA-seq data, and benchmark it against state-of-the-art deep learning and conventional methods across multiple plant species. Species-specific scPlantAnnotate models were trained using curated datasets from Arabidopsis thaliana, Zea mays, Oryza sativa, and Glycine max. We compared scPlantAnnotate with leading baselines under both standard random-split evaluation and a more stringent leave-one-dataset-out setting, which tests robustness to completely unseen datasets and tissue types. scPlantAnnotate consistently outperforms existing approaches across all four species under random-split evaluation. In the leave-one-dataset-out setting for A. thaliana, where performance drops markedly for all methods due to strong batch effects and dataset heterogeneity, scPlantAnnotate nonetheless achieves the highest Accuracy, Macro-F1, Balanced Accuracy, and Macro-AUROC on average and ranks first on most held-out datasets. These results demonstrate improved robustness to dataset shifts, a critical yet underexplored challenge in plant scRNA-seq analysis. A freely accessible web server enables users to annotate their own datasets using pretrained models. scPlantAnnotate provides a plant-specific, Transformer-based framework for single-cell annotation that delivers state-of-the-art performance and enhanced robustness to unseen datasets. By addressing limitations of existing tools and enabling scalable reference-based annotation, scPlantAnnotate supports the development of comprehensive plant cell atlases and facilitates broader use of single-cell genomics in plant biology.

Bioinformatics↗