Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Sequence Annotation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

In vivo mapping of mutagenesis sensitivity of human enhancers

Distant-acting enhancers are central to human development1. However, our limited understanding of their functional sequence features prevents the interpretation of enhancer mutations in disease2. Here we determined the functional sensitivity to mutagenesis of human developmental enhancers in vivo. Focusing on seven enhancers that are active in the developing brain, heart, limb and face, we created over 1,700 transgenic mice for over 260 mutagenized enhancer alleles. Systematic mutation of 12-base-pair blocks collectively altered each sequence feature in each enhancer at least once. We show that 69% of all blocks are required for normal in vivo activity, with mutations more commonly resulting in loss (60%) than in gain (9%) of function. Using predictive modelling, we annotated critical nucleotides at the base-pair resolution. The vast majority of motifs predicted by these machine learning models (88%) coincided with changes in in vivo function, and the models showed considerable sensitivity, identifying 59% of all functional blocks. Taken together, our results reveal that human enhancers contain a high density of sequence features that are required for their normal in vivo function and provide a rich resource for further exploration of human enhancer logic.

Kosicki, Michael↗

The reference genome for the northeastern Pacific bull kelp, Nereocystis luetkeana

Bull kelp, Nereocystis luetkeana, is a northeastern Pacific kelp with broad distribution from Alaska to central California. Its population declines have caused severe concerns in northern California, the Salish Sea in Washington, and recently in some populations in Oregon. Despite bull kelp's accumulated ecological and physiological studies, an assembled and annotated genomic reference was still unavailable. Here, we report the complete and annotated genome of Nereocystis luetkeana, produced by the California Conservation Genomics Project (CCGP), which aims to reveal genomic diversity patterns across California by sequencing the complete genomes of approximately 150 carefully selected species. The genome was assembled into 1562 scaffolds with 449.82 Mb, 80x of coverage and 22 952 gene models. BUSCO assembly showed a completeness score of 72% for the stramenopiles gene set. The mitochondria and chloroplast genome sequences have 37 Kb and 131 Mb, respectively. The orthology analysis between 10 Phaeophycean genomes showed 1065 expanded and 286 unique orthogroups for this species. Pairwise comparisons showed 542 orthogroups present only in N. luetkeana and M. pyrifera, another large-body kelp. The enrichment analysis of these orthogroups showed important functions related to central metabolism and signaling due to ATPases enrichment in these two species. This genome assembly will provide an essential resource for the ecology, evolution, conservation, and breeding of bull kelp.

California Conservation Genomics Project—CCGP↗

6051R & 6051S Assembly and Annotation

We report the draft genomes of two morphologically distinct variants of Bacillus subtilis ATCC 6051 [NCBI3610]. The two isolates exhibit differences in not only morphology but also their genetics, despite identical 16S rRNA sequences. Investigating the genetic differences of colony morphology variation in this model organism can provide valuable insights.

59 BASIC BIOLOGICAL SCIENCES↗

A practical approach to using the Genomic Standards Consortium MIxS reporting standard for comparative genomics and metagenomics

Comparative analysis of (meta)genomes necessitates aggregation, integration, and synthesis of well-annotated data using standards. The Genomic Standards Consortium (GSC) collaborates with the research community to develop and maintain the Minimal Information about any (x) Sequence (MIxS) reporting standard for genomic data. To facilitate use of the GSC’s MIxS reporting standard, we provide a description of the structure and terminology, how to navigate ontologies for required terms in MIxS, and demonstrate practical usage through a soil metagenome example.

standards, metadata, genome, metagenome, schema, v↗

Optimizing inference of segmentation on high-resolution images in MLExchange

MLExchange is a machine learning (ML) operations platform providing web user-interfaces (UIs) for data visualization and analysis pipelines at synchrotron facilities. Among these UIs is the segmentation app which helps synchrotron users utilize ML algorithms to automatically segment high-resolution scientific images with minimal manual annotation effort. In this work, we share code optimizations that significantly speed up the segmentation inference workflow of large data in short time. By optimizing the sequence of CPU-GPU data transfers and introducing CPU parallelization to key operations, we improve the per-device, per-image frame computational efficiency and observe close to 3×$$\times$$ speedup over the original segmentation inference workflow run time when utilizing a single GPU. Further adaptations enabling multi-GPU inference yield more than 40×$$\times$$ speedup with 100 GPUs compared to the optimized single GPU inference workflow. This acceleration of the segmentation inference workflow will provide MLExchange users with easy access to segmentation results with little wait time.

Lu, Shizhao↗

A global soil plasmidome resource unveils functional and ecological roles of plasmids in soil microbiomes

Plasmids play significant roles in microbial adaptation to ecosystems, yet their dynamics remain poorly understood due to identification challenges. We present the Global Soil Plasmidome Resource (GSPR), a comprehensive dataset of 98,728 plasmid sequences amassed from 6860 terrestrial microbial communities and isolates. We explore this resource through various computational approaches, including phylogenetic diversity analysis, host prediction, and extensive functional annotation, to understand the contribution of plasmids to the genetic and functional diversity in soil, correlating these findings with sample type, as well as the soil habitat they were retrieved from. Our analysis reveals insights into plasmid-encoded functions such as effector modules, quorum sensing, and stress resistance, which may contribute to their persistence and microbial adaptation in soil. Furthermore, CRISPR analysis suggests a prevalent role of these elements related to intra-plasmid competition. By contrasting plasmids from cultivated and uncultivated organisms, we identify important functions that expand existing knowledge of plasmid roles in these habitats. This study represents a notable step forward in elucidating plasmid diversity and function within soil microbiomes and establishes a foundational framework for exploring their roles in natural environments.

Fiamenghi, Mateus B↗

Signatures of Selection for Resistance/Tolerance to Perkinsus olseni in Grooved Carpet Shell Clam ( Ruditapes decussatus ) Using a Population Genomics Approach

ABSTRACT The grooved carpet shell clam ( Ruditapes decussatus ) is a bivalve of high commercial value distributed throughout the European coast. Its production has suffered a decline caused by different factors, especially by the parasite Perkinsus olsenii . Improving production of R . decussatus requires genomic resources to ascertain the genetic factors underlying resistance/tolerance to P. olseni i . In this study, the first reference genome of R . decussatus was assembled through long‐ and short‐read sequencing (1677 contigs; 1.386 Mb) and further scaffolded at chromosome level with Hi‐C (19 superscaffolds; 95.4% of assembly). Repetitive elements were identified (32%) and masked for annotation of 38,276 coding‐ and 13,056 non‐coding genes. This genome was used as a reference to develop a 2bRAD‐Seq 13,438 SNP panel for a genomic screening on six shellfish beds distributed across the Atlantic Ocean and Mediterranean Sea. Beds were selected by perkinsosis prevalence and the infection level was individually evaluated in all the samples. Genetic diversity was significantly higher in the Mediterranean than in the Atlantic region. The main genetic breakage was detected between those regions (F ST = 0.224), being the Mediterranean more heterogeneous than the Atlantic. Several loci under divergent selection (394 outliers; 261 genomic windows) were detected across shellfish beds. Samples were also inspected to detect signals of selection for resistance/tolerance to P. olseni i by using infection‐level and population‐genomics approaches, and 90 common divergent outliers for resistance/tolerance to perkinsosis were identified and used for gene mining. Candidate genes and markers identified provide invaluable information for controlling perkinsosis and for improving production of the grooved carpet shell clam.

Sambade, Inés M. [Department of Zoology, Genetics ↗

High-quality draft genome sequence of Thermobifida halotolerans DSM 44931

Here, we report the genome sequence of Thermobifida halotolerans DSM 44931, a bacterium that was originally isolated from a salt mine in the Yunnan Province of China. This genome was sequenced using Pacific Biosciences sequencing technology and was assembled into 2 contigs in 2 scaffolds. It has a total length of 5,506,851 bp and a GC content of 71.16%. Functional annotation of this genome provides further metabolic insight into this species.

actinomycete↗

Addressing the dynamic nature of reference data: a new nucleotide database for robust metagenomic classification

Accurate metagenomic classification relies on comprehensive, up-to-date, and validated reference databases. While the NCBI BLAST Nucleotide (nt) database, encompassing a vast collection of sequences from all domains of life, represents an invaluable resource, its massive size—currently exceeding 10 12 nucleotides—and exponential growth pose significant challenges for researchers seeking to maintain current nt-based indices for metagenomic classification. Recognizing that no current nt-based indices exist for the widely used Centrifuge classifier, and the last public version currently available was released in 2018, we addressed this critical gap by leveraging advanced high-performance computing resources. We present new Centrifuge-compatible nt databases, meticulously constructed using a novel pipeline incorporating different quality control measures, including reference decontamination and filtering. These measures demonstrably reduce spurious classifications, as shown through our reanalysis of published metagenomic data where Plasmodium annotations were dramatically reduced using our decontaminated database, highlighting how database quality can significantly impact research conclusions. Through temporal comparisons, we also reveal how our approach minimizes inconsistencies in taxonomic assignments stemming from asynchronous updates between public sequence and taxonomy databases. These discrepancies are particularly evident in taxa such as Listeria monocytogenes and Naegleria fowleri, where classification accuracy varied significantly across database versions. These new databases, made available as pre-built Centrifuge indexes, respond to the need for an open, robust, nt-based pipeline for taxonomic classification in metagenomics. Applications such as environmental metagenomics, forensics, and clinical metagenomics, which require comprehensive taxonomic coverage, will benefit from this resource. Our work highlights the importance of treating reference databases as dynamic entities, subject to ongoing quality control and validation akin to software development best practices. This approach is crucial for ensuring accuracy and reliability of metagenomic analysis, especially as databases continue to expand in size and complexity.

59 BASIC BIOLOGICAL SCIENCES↗

AlgaeOrtho, a bioinformatics tool for processing ortholog inference results in algae

Introduction: Microalgae constitute a prominent feedstock for producing biofuels and biochemicals by virtue of their prolific reproduction, high bioproduct accumulation, and the ability to grow in brackish and saline water. However, naturally occurring wild type algal strains are rarely optimal for industrial use; therefore, bioengineering of algae is necessary to generate superior performing strains that can address production challenges in industrial settings, particularly the bioenergy and bioproduct sectors. One of the crucial steps in this process is deciding on a bioengineering target: namely, which gene/protein to differentially express. These targets are often orthologs which are defined as genes/proteins originating from a common ancestor in divergent species. Although bioinformatics tools for the identification of protein orthologs already exist, processing the output from such tools is nontrivial, especially for a researcher with little or no bioinformatics experience. Methods: The present study introduces AlgaeOrtho, a user-friendly tool that builds upon the SonicParanoid orthology inference tool (based on an algorithm that identifies potential protein orthologs based on amino acid sequences) and the PhycoCosm database from JGI (Joint Genome Institute) to help researchers identify orthologs of their proteins of interest in multiple diverse algal species. Results: The output of this application includes a table of the putative orthologs of their protein of interest, a heatmap showing sequence similarity (%), and an unrooted tree of the putative protein orthologs. Notably, the tool would be instrumental in identifying novel bioengineering targets in different algal strains, including targets in not-fully annotated algal species, since it does not depend on existing protein annotations. We tested AlgaeOrtho using three case studies, for which orthologs of proteins relevant to bioengineering targets, were identified from diverse algal species, demonstrating its ease of use and utility for bioengineering researchers. Discussion: This tool is unique in the protein ortholog identification space as it can visualize putative orthologs, as desired by the user, across several algal species.

09 BIOMASS FUELS↗

Human Host Cellular Response to HCoV-229E Infection Proteomics (ACS-JM-DP2)

The purpose of this experiment was to evaluate the human host cellular response to wild-type Human coronavirus strain 229E (HCoV-229E) infection. Sample data was obtained for mock and infected immortalized human lung epithelial cells (A549) (MOI 5) nuclear extracts, immortalized human lung fibroblasts cells (MRC5) (MOI5) nuclear extracts, and primary human airway epithelial (HAE) (MOI 3) cells from lung tissue and processed for proteome analysis. Processed datasets are openly accessible from the download button and contain secondary processed proteomic results files and supporting metadata materials. Experimental proteomics samples were prepared using Limited Proteolysis (LiP) methods for Label-free quantification (LFQ) and global proteomic evaluation. Sample data was acquired using a Q-Exactive HF-X mass spectrometer and was processed and compiled using MaxQuant software (v.1.6.17.0). Processed proteomic data downloads include a sample naming key, processed MaxQuant results/parameters, and protein annotated relative abundance files. See corresponding primary data accessions below and Viral Experiment LiP Analysis source code supporting data transparency and reuse. Experimental transcriptomics samples were collected in parallel and processed for RNA sequencing (RNA-Seq) as summarized under ACS-DP1 (https://data.pnnl.gov/group/nodes/dataset/34069).

59 BASIC BIOLOGICAL SCIENCES↗

MVP: a modular viromics pipeline to identify, filter, cluster, annotate, and bin viruses from metagenomes

While numerous computational frameworks and workflows are available for recovering prokaryote and eukaryote genomes from metagenome data, only a limited number of pipelines are designed specifically for viromics analysis. With many viromics tools developed in the last few years alone, it can be challenging for scientists with limited bioinformatics experience to easily recover, evaluate quality, annotate genes, dereplicate, assign taxonomy, and calculate relative abundance and coverage of viral genomes using state-of-the-art methods and standards. Here, we describe Modular Viromics Pipeline (MVP) v.1.0, a user-friendly pipeline written in Python and providing a simple framework to perform standard viromics analyses. MVP combines multiple tools to enable viral genome identification, characterization of genome quality, filtering, clustering, taxonomic and functional annotation, genome binning, and comprehensive summaries of results that can be used for downstream ecological analyses. Overall, MVP provides a standardized and reproducible pipeline for both extensive and robust characterization of viruses from large-scale sequencing data including metagenomes, metatranscriptomes, viromes, and isolate genomes. As a typical use case, we show how the entire MVP pipeline can be applied to a set of 20 metagenomes from wetland sediments using only 10 modules executed via command lines, leading to the identification of 11,656 viral contigs and 8,145 viral operational taxonomic units (vOTUs) displaying a clear beta-diversity pattern. Further, acting as a dynamic wrapper, MVP is designed to continuously incorporate updates and integrate new tools, ensuring its ongoing relevance in the rapidly evolving field of viromics. MVP is available at https://gitlab.com/ccoclet/mvp and as versioned packages in PyPi and Conda.

59 BASIC BIOLOGICAL SCIENCES↗

Coupling Metabolic Source Isotopic Pair Labeling and Genome Wide Association for Metabolite and Gene Annotation in Plants (Final Technical Report)

In this project, we applied our labeling pipeline to Arabidopsis and sorghum by feeding tissues with isotopically labeled versions of commercially available amino acids to identify all metabolite features that incorporate the label. In sorghum, we fed five accessions, sampled across the diversity of sorghum, to identify the precursor-of-origin for metabolites that vary between accessions as well as those that may be missing from a single reference genotype. This provided us with precursor-of-origin annotation for thousands of unknown metabolites. We then used GWA to map genes responsible for the synthesis of precursor-of-origin classified metabolites. For sorghum leaf and root ducible metabolites, we performed untargeted metabolomics on leaf and root tissues from 300 diverse genotyped sorghum inbred lines. The amino acid precursor-of-origin metabolite library were then used to identify the corresponding metabolites in the GWA data sets and to identify novel gene-metabolite associations. Finally, we utilized existing and newly generated sequenced EMS mutants of sorghum to validate the predicted gene-metabolite relationships that our labelling analysis identified. In parallel, we conducted similar feeding experiments in Arabidopsis to categorize metabolites based on precursor-of-origin, identify those that vary across our existing Arabidopsis metabolite GWA dataset, and identify genes required for the synthesis of each metabolite. To provide an independent test of gene annotation and pathway involvement, we tested the GWA gene-metabolite associations in Arabidopsis by analyzing the metabolic phenotypes of gene knockouts. Genes of particular interest from both sorghum and Arabidopsis were studied in detail by directly measuring the activity of the corresponding enzymes following heterologous expression. In summary, this work classified as-yet-unknown amino acid-derived metabolites and identified genes involved in their production generated through “omics” technologies. This information was used to validate gene function and identify new metabolism in Arabidopsis and sorghum.

09 BIOMASS FUELS↗

The small protein SbtC is a functional component of the CO 2 concentrating mechanism in Synechocystis sp. PCC 6803

Oxygenic phototrophs fix CO 2 via the enzyme ribulose-1,5-bisphosphate carboxylase/oxygenase (RubisCO), which shows relatively low CO 2 affinity and specificity. To circumvent low and fluctuating CO 2 concentrations in aquatic systems, cyanobacteria and algae have evolved sophisticated inorganic carbon (Ci) concentrating mechanisms (CCMs). Bicarbonate transporters such as SbtA play a crucial role in the cyanobacterial CCM and hence display multiple layers of tight regulation. Control of sbtA gene expression and corresponding transporter activity involves the PII-like protein SbtB, whose gene is frequently co-transcribed with sbtA. A previously non-annotated gene located upstream of the sbtAB operon in the model Synechocystis sp. PCC 6803 encodes the small protein SbtC, composed of 80 amino acids. Presence of SbtC was confirmed by immunoblotting of the sbtC-coding sequence fused to a Flag-tag. Similar to sbtAB , transcription of the sbtC locus is induced by low CO 2 availability; however, it is controlled independently. Mutation of the sbtC locus in a wild-type background produced only a mild phenotype, even under low CO 2 , but impaired diurnal growth resembled that of the mutant ΔsbtB . Biochemical analysis indicated a trimeric SbtABC complex in the membrane. Bicarbonate leakage from cells was strongly elevated when either sbtB or sbtC was deleted from recombinant Synechocystis strains harboring only SbtA as single Ci uptake system. Here, our results provide evidence that SbtC contributes to the formation of the SbtAB complex, thereby regulating bicarbonate exchange at the cytoplasmic membrane. Well-conserved SbtC-like proteins encoded in the neighborhood of sbtAB exist in many cyanobacterial genomes, pointing toward an important role in the cyanobacterial CCM.

Walke, Peter [Univ. of Rostock (Germany)] (ORCID:0↗

Building a FAIR data ecosystem for incorporating single-cell transcriptomics data into agricultural genome to phenome research

Introduction The agriculture genomics community has numerous data submission standards available, but the standards for describing and storing single-cell (SC, e.g., scRNA- seq) data are comparatively underdeveloped. Methods To bridge this gap, we leveraged recent advancements in human genomics infrastructure, such as the integration of the Human Cell Atlas Data Portal with Terra, a secure, scalable, open-source platform for biomedical researchers to access data, run analysis tools, and collaborate. In parallel, the Single Cell Expression Atlas at EMBL-EBI offers a comprehensive data ingestion portal for high-throughput sequencing datasets, including plants, protists, and animals (including humans). Developing data tools connecting these resources would offer significant advantages to the agricultural genomics community. The FAANG data portal at EMBL-EBI emphasizes delivering rich metadata and highly accurate and reliable annotation of farmed animals but is not computationally linked to either of these resources. Results Herein, we describe a pilot-scale project that determines whether the current FAANG metadata standards for livestock can be used to ingest scRNA-seq datasets into Terra in a manner consistent with HCA Data Portal standards. Importantly, rich scRNA-seq metadata can now be brokered through the FAANG data portal using a semi-automated process, thereby avoiding the need for substantial expert curation. We have further extended the functionality of this tool so that validated and ingested SC files within the HCA Data Portal are transferred to Terra for further analysis. In addition, we verified data ingestion into Terra, hosted on Azure, and demonstrated the use of a workflow to analyze the first ingested porcine scRNA-seq dataset. Additionally, we have also developed prototype tools to visualize the output of scRNA-seq analyses on genome browsers to compare gene expression patterns across tissues and cell populations. This JBrowse tool now features distinct tracks, showcasing PBMC scRNA-seq alongside two bulk RNA-seq experiments. Discussion We intend to further build upon these existing tools to construct a scientist-friendly data resource and analytical ecosystem based on Findable, Accessible, Interoperable, and Reusable (FAIR) SC principles to facilitate SC-level genomic analysis through data ingestion, storage, retrieval, re-use, visualization, and comparative annotation across agricultural species.

Genetics & Heredity↗

RolyPoly (rp) v0.1.0

The Rolypoly pipeline is designed to process raw RNA-seq data and identify potential RNA viral sequences. It is split into several self contained steps: 1. input data filtering and QC, 2. Genome assembly and refinement, 3. Assembly filtering, 4. Mapping to known RNA viral genomes, 5. Searching for RNA viral marker genes. 6. Genome functional and structural annotation. 6. Report preparation and potential downstream analysis The last module, may include taxonomic assignment, host range estimation, and phenotypic prediction. There are many similar software, but they focus on human related viruses, and lack the downstream applications or differ in their sensitivity. The initial user base are non-computational microbial ecologists who wish to better understand the potential RNA viruses in their own generated samples.

Neri, Uri↗

Database of virus genomes from ultra-deep sequencing of wastewater

Researchers at University of Missouri have conducted ultra-deep RNA sequencing of viral concentrates from wastewater (1 billion Illumina reads per sample). The resulting dataset spans 321 samples collected weekly from 11 cities between 2023-2025. As part of a tri-lab collaboration, scientists at LLNL and LANL cleaned, assembled, and annotated this metagenomic data, identifying nearly 200,000 viral genomes. Careful data curation resulted in a database containing 21,015 high-quality, near-complete viral genomes from wastewater. This database contains viruses predicted to infect a range of hosts including bacteria (most common viruses), plants (most abundant viruses), and vertebrates (rarest viruses). There are also numerous novel viruses that could not be well identified and whose host(s) are unknown. Just 7% of all genomes in the wastewater virus database had genus-level matches in the public NCBI database, and 17% matched to a recently created metagenomic virus database at that level (metaVR). The database will provide baseline information about viruses in wastewater that may be used to additional identify novel viruses during ongoing monitoring

Allen, Jonathan [Lawrence Livermore National Labor↗