Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Biological databases”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

106 records · Page 6

Warming is Associated With More Encoded Antimicrobial Resistance Genes and Transcriptions Within Five Drug Classes in Soil Bacteria: A Case Study and Synthesis

ABSTRACT The effect of warming on anti‐microbial resistance (AMR) genes in the environment has critical implications for public health but is little studied. We collected published soil bacterial genomes from the BV‐BRC database and tested the correlation between reported optimal growth temperature and the number of encoded AMR genes. Furthermore, we tested the relationship between temperature and AMR gene transcription in a natural ecosystem by analysing soil transcriptomes from a warming manipulation experiment in an Alaskan boreal forest. We hypothesised that there is a positive relationship between warming and AMR prevalence in gene content in bacterial genomes and transcriptomic sequences, and that this effect would vary by drug class. Regarding the bacterial genomes, we found a positive relationship between the fraction of encoded AMR genes and the reported optimal temperature of soil bacteria. The drug classes tetracycline and lincosamide/macrolide/streptogramin had the strongest positive relationship with reported optimal temperature. For the case study in a natural ecosystem, we found 61 significantly upregulated AMR gene‐associated transcripts spanning eight drug classes in warmed plots. In the Alaskan soil samples, we found that warming elicited the strongest positive effect on transcripts targeting lincosamide/streptogramin, beta‐lactam and phenicol/quinolone antibiotics. Overall, higher temperatures were linked to AMR gene prevalence.

Hacopian, Melanie T. [Department of Ecology and Ev↗

The secondary metabolism collaboratory: a database and web discussion portal for secondary metabolite biosynthetic gene clusters

Secondary metabolites are small molecules produced by all corners of life, often with specialized bioactive functions with clinical and environmental relevance. Secondary metabolite biosynthetic gene clusters (BGCs) can often be identified within DNA sequences by various sequence similarity tools, but determining the exact functions of genes in the pathway and predicting their chemical products can often only be done by careful, manual comparative analysis. To facilitate this, we report the first release of the secondary metabolism collaboratory (SMC), which aims to provide a comprehensive, tool-agnostic repository of BGC sequence data drawn from all publicly available and user-submitted bacterial and archaeal genome and contig sources. On the website, users are provided a searchable catalog of putative BGCs identified from each source, along with visualizations of gene and domain annotations derived from multiple sequence analysis tools. SMC’s data is also available through publicly-accessible application programming interface (API) endpoints to facilitate programmatic access. Users are encouraged to share their findings (and search for others’) through comment posts on BGC and source pages. At the time of writing, SMC is the largest repository of BGC information, holding 13.1M BGC regions from 1.3M source sequences and growing, and can be found at https://smc.jgi.doe.gov.

59 BASIC BIOLOGICAL SCIENCES↗

Metabolic interactions underpinning high methane fluxes across terrestrial freshwater wetlands

Current estimates of wetland contributions to the global methane budget carry high uncertainty, particularly in accurately predicting emissions from high methane-emitting wetlands. Microorganisms drive methane cycling, but little is known about their conservation across wetlands. To address this, we integrate 16S rRNA amplicon datasets, metagenomes, metatranscriptomes, and annual methane flux data across 9 wetlands, creating the Multi-Omics for Understanding Climate Change (MUCC) v2.0.0 database. This resource is used to link microbiome composition to function and methane emissions, focusing on methane-cycling microbes and the networks driving carbon decomposition. We identify eight methane-cycling genera shared across wetlands and show wetland-specific metabolic interactions in marshes, revealing low connections between methanogens and methanotrophs in high-emitting wetlands. Methanoregula emerged as a hub methanogen across networks and is a strong predictor of methane flux. In these wetlands it also displays the functional potential for methylotrophic methanogenesis, highlighting the importance of this pathway in these ecosystems. Collectively, our findings illuminate trends between microbial decomposition networks and methane flux while providing an extensive publicly available database to advance future wetland research.

54 ENVIRONMENTAL SCIENCES↗

AlgaeOrtho, a bioinformatics tool for processing ortholog inference results in algae

Introduction: Microalgae constitute a prominent feedstock for producing biofuels and biochemicals by virtue of their prolific reproduction, high bioproduct accumulation, and the ability to grow in brackish and saline water. However, naturally occurring wild type algal strains are rarely optimal for industrial use; therefore, bioengineering of algae is necessary to generate superior performing strains that can address production challenges in industrial settings, particularly the bioenergy and bioproduct sectors. One of the crucial steps in this process is deciding on a bioengineering target: namely, which gene/protein to differentially express. These targets are often orthologs which are defined as genes/proteins originating from a common ancestor in divergent species. Although bioinformatics tools for the identification of protein orthologs already exist, processing the output from such tools is nontrivial, especially for a researcher with little or no bioinformatics experience. Methods: The present study introduces AlgaeOrtho, a user-friendly tool that builds upon the SonicParanoid orthology inference tool (based on an algorithm that identifies potential protein orthologs based on amino acid sequences) and the PhycoCosm database from JGI (Joint Genome Institute) to help researchers identify orthologs of their proteins of interest in multiple diverse algal species. Results: The output of this application includes a table of the putative orthologs of their protein of interest, a heatmap showing sequence similarity (%), and an unrooted tree of the putative protein orthologs. Notably, the tool would be instrumental in identifying novel bioengineering targets in different algal strains, including targets in not-fully annotated algal species, since it does not depend on existing protein annotations. We tested AlgaeOrtho using three case studies, for which orthologs of proteins relevant to bioengineering targets, were identified from diverse algal species, demonstrating its ease of use and utility for bioengineering researchers. Discussion: This tool is unique in the protein ortholog identification space as it can visualize putative orthologs, as desired by the user, across several algal species.

09 BIOMASS FUELS↗

PAVC: The foundation for a Pan-Arctic Vegetation Cover database

Field-measured Arctic vegetation cover data is essential for creating accurate, high-quality vegetation structure and composition maps. Extrapolating field data into high-resolution cover maps provides detailed, function-specific information for use in Earth System Models, vegetation classifications, and monitoring vegetation change over time and space. However, field campaigns that collect plant cover vary substantially in scope, method, and purpose, which makes them difficult to unify across data stores, and they are often not designed to meet remote sensing needs. In this work, we synthesized and harmonized field-based fractional cover data from various data stores to create a high-quality, consistent repository schema for remote sensing-based vegetation cover mapping applications. We developed a reproducible workflow for synthesizing visual estimate and point-intercept fractional cover data. The resultant Pan-Arctic Vegetation Cover (PAVC) database contains synthesized fractional cover at both the species and plant functional type levels. The latter includes absolute foliar cover for deciduous shrubs and trees, evergreen shrubs and trees, forbs, graminoids, lichen, bryophytes, and “other” vegetation, as well as absolute cover for litter and top cover for water and bare ground.

Steckler, Morgan R. [Oak Ridge National Laboratory↗

The impact of curation errors in the PDBBind Database on machine learning predictions of protein–protein binding affinity

The PDBBind database has been widely utilized for the computational prediction of protein–protein binding affinities. While the accuracy of the PDBBind-curated equilibrium dissociation constants (K D ) has been reported for the protein–ligand subset of the PDBBind database, the curation accuracy has not been reported for the protein–protein subset. Here, we present a detailed manual analysis for the subset of PDBBind records with PubMed Central Open Access primary publications and find that ~19% of these records had K D values that were not supported by their primary publications. The impact of these putative curation errors on the machine learning-based prediction of K D from experimental protein–protein 3D structures was evaluated and correcting the curation errors improved the Pearson correlation coefficient between measured and random forest-predicted log 10 (K D ) values by ~8 percentage points. This finding underscores the importance of dataset accuracy for computational modelling and highlights the need for more stringent curation processes when extracting information from the scientific literature.

59 BASIC BIOLOGICAL SCIENCES↗

Produced Water DNA Database (PW-DNA): Utilizing KBase to generate an environmental specific curated molecular database

The deep subsurface is estimated to host the majority of Earth’s microbial biomass yet remains one of the most challenging environments to access and study. One common approach to investigate these microbial communities is through the analysis of produced water from subsurface reservoirs, where researchers can assess water and gas chemistry along with molecular (DNA/RNA) sequence data. Advances in high-throughput sequencing have greatly expanded our understanding of these environments and their biotechnological potential. However, further progress requires large-scale, integrative meta-analyses across diverse datasets. To address this need, we developed the Produced Water-DNA (PW-DNA) Database, a curated, publicly available resource that consolidates microbial DNA/RNA sequences, geochemical data, and relevant metadata from in situ hydrocarbon environments such as coal beds, oil reservoirs, and natural gas systems. The PW-DNA database delivers three core benefits to the research community: (1) it improves data sharing by linking environmental microbial datasets with corresponding geochemical parameters, enabling more robust filtering and analysis; (2) it connects with complementary research databases to promote broader dissemination and interoperability; and (3) it supports technological innovation by serving as a resource for identifying microbial trends and exploring genetic potential. While individual studies have highlighted basin-specific microbial communities and functional redundancy in biogeochemical cycling, a comprehensive, system-wide perspective is needed to better understand connectivity and novelty across subsurface ecosystems. By designing the PW-DNA in the KBase platform, we provide a reproducible, visual framework for integrating large-scale genomic and geochemical data, enabling researchers to perform more informed analyses and experimental design. Ultimately, this resource enhances the ability to identify, characterize, and interpret microbial functions across diverse subsurface environments, thereby accelerating discovery in subsurface microbiology and biotechnology.

59 BASIC BIOLOGICAL SCIENCES↗

A functional microbiome catalogue crowdsourced from North American rivers

Predicting elemental cycles and maintaining water quality under increasing anthropogenic influence requires knowledge of the spatial drivers of river microbiomes. However, understanding of the core microbial processes governing river biogeochemistry is hindered by a lack of genome-resolved functional insights and sampling across multiple rivers. Here we used a community science effort to accelerate the sampling, sequencing and genome-resolved analyses of river microbiomes to create the Genome Resolved Open Watersheds database (GROWdb). GROWdb profiles the identity, distribution, function and expression of microbial genomes across river surface waters covering 90% of United States watersheds. Specifically, GROWdb encompasses microbial lineages from 27 phyla, including novel members from 10 families and 128 genera, and defines the core river microbiome at the genome level. GROWdb analyses coupled to extensive geospatial information reveals local and regional drivers of microbial community structuring, while also presenting foundational hypotheses about ecosystem function. Building on the previously conceived River Continuum Concept, we layer on microbial functional trait expression, which suggests that the structure and function of river microbiomes is predictable. We make GROWdb available through various collaborative cyberinfrastructures, so that it can be widely accessed across disciplines for watershed predictive modelling and microbiome-based management practices.

59 BASIC BIOLOGICAL SCIENCES↗

A metagenomic perspective on the microbial prokaryotic genome census

Following 30 years of sequencing, we assessed the phylogenetic diversity (PD) of >1.5 million microbial genomes in public databases, including metagenome-assembled genomes (MAGs) of uncultivated microbes. As compared to the vast diversity uncovered by metagenomic sequences, cultivated taxa account for a modest portion of the overall diversity, 9.73% in bacteria and 6.55% in archaea, while MAGs contribute 48.54% and 57.05%, respectively. Therefore, a substantial fraction of bacterial (41.73%) and archaeal PD (36.39%) still lacks any genomic representation. This unrepresented diversity manifests primarily at lower taxonomic ranks, exemplified by 134,966 species identified in 18,087 metagenomic samples. Our study exposes diversity hotspots in freshwater, marine subsurface, sediment, soil, and other environments, whereas human samples yielded minimal novelty within the context of existing datasets. These results offer a roadmap for future genome recovery efforts, delineating uncaptured taxa in underexplored environments and underscoring the necessity for renewed isolation and sequencing.

59 BASIC BIOLOGICAL SCIENCES↗

Fractionation of Filamentous Algae from Mixed Biofilms

Filamentous algae, which grow in long, hair-like filaments within biofilms, play a crucial role in wastewater treatment due to their ability to produce significant biomass and their resistance to predation compared to traditional microalgal treatments. These algae can effectively uptake and utilize pollutants, particularly excessive nitrogen (ammonia, nitrate, nitrite) and phosphorus (phosphate), making filamentous algae valuable for wastewater treatment, as well as bioethanol and biodiesel production due to high lipid productions. However, each algal species possesses different capacities, necessitating a thorough genetic identification and understanding of each community. A major challenge in accurately assessing these communities is the lack of coverage in large sequencing databases which can lead to misrepresentation of the true composition and abundance of organisms and overall sequencing bias. To address this, I evaluated chemical and physical techniques for separating filamentous algae from mixed biofilms to achieve clean genetic sequencing results. I employed pH washing (0.001M HCl, 0.001M HCl, DiH2O, 0.0001M HCl, 0.001M HCl) for chemical treatment, followed by physical separation through centrifugation (5000rpm, 6500rpm) or filtration (2mm, 250um, 75um). The most successful method was deionized water washing, which yielded clear differences across stacked filters; the 2mm filtrate showed high levels of filamentous algae, with microalgae eluting in the 75um filtrate or remaining within agglutinations of algae larger filters. Base washing eluted the highest concentrations of microalgae, with larger filter sizes retaining more filamentous algae, indicating the breakdown of extracellular polymeric substances (EPS). Our downstream plans include sending the high-throughput next-generation sequencing to confirm the purity and ratios of filamentous and non-filamentous algae, as well as bacteria present, thereby validating the success of our treatments. Potential applications include creating community-based fractions for analysis, refining current sequencing data with clearer isolations, and generating designer biofilms to enhance our understanding of community interactions.

59 BASIC BIOLOGICAL SCIENCES↗

Estimating irrigation water use from remotely sensed evapotranspiration data: Accuracy and uncertainties at field, water right, and regional scales

Irrigated agriculture is the dominant user of water globally, but most water withdrawals are not monitored or reported. As a result, it is largely unknown when, where, and how much water is used for irrigation. Here, we evaluated the ability of remotely sensed evapotranspiration (ET) data, integrated with other datasets, to calculate irrigation water withdrawals and applications in an intensively irrigated portion of the United States. We compared irrigation calculations based on an ensemble of satellite-driven ET models from OpenET with reported groundwater withdrawals from hundreds of farmer irrigation application records and a statewide flowmeter database at three spatial scales (field, water right group, and management area). At the field scale, we found that ET-based calculations of irrigation agreed best with reported irrigation when the OpenET ensemble mean was aggregated to the growing season timescale (bias = 1.6–4.9%, R 2 = 0.53–0.74), and agreement between calculated and reported irrigation was better for multi-year averages than for individual years. At the water right group scale, linking pumping wells to specific irrigated fields was the primary source of uncertainty. At the management area scale, calculated irrigation exhibited similar temporal patterns as flowmeter data but tended to be positively biased with more interannual variability. Disagreement between calculated and reported irrigation was strongly correlated with annual precipitation, and calculated and reported irrigation agreed more closely after statistically adjusting for annual precipitation. The selection of an ET model was also an important consideration, as variability across ET models was larger than the potential impacts of conservation measures employed in the region. From these results, we suggest key practices for working with ET-based irrigation data that include accurately accounting for changes in soil moisture, deep percolation, and runoff; careful verification of irrigated area and well-field linkages; and conducting application-specific evaluations of uncertainty.

59 BASIC BIOLOGICAL SCIENCES↗

Annotation of DOM metabolomes with an ultrahigh resolution mass spectrometry molecular formula library

Current approaches to analyzing metabolomic data often rely on matching MS/MS fragmentation data to sparse libraries or databases. This approach results in limited identification of features, often with less than 10% of the dataset being annotated. A complementary approach is to assign molecular formula to features based on accurate mass measurements, but the platforms commonly used for metabolomics do not have the needed accuracy or resolving power to do this robustly, particularly for larger molecules. Using our newly modified analysis tool, CoreMS, we generated a library of molecular formula from pooled samples analyzed with LC-21T FT-ICR MS. This library successfully annotated approximately 53.2% of features identified from the exometabolome of marine diatom Phaeodactylum tricornutum – a nearly ten-fold increase over the 5.9% annotation rate achieved using a conventional MS/MS library matching approach. Using this FT-ICR MS library approach, we were able to differentiate differences in the exometabolome of P. tricornutum in iron replete and iron limited conditions, with 668 metabolites being differentially expressed (p < 0.05, 2 x intensity difference) under these conditions. The traditional MS/MS fragmentation-based annotation approach only annotated 61 of these metabolites, while our novel pipeline annotated 450 metabolites and revealed 12 metabolites that were significantly more abundant under low iron conditions. Our results demonstrate the utility of ultrahigh resolution mass spectrometry for generating more comprehensive and confident molecular annotations.

21T-FTICR-MS, CoreMS↗

Genomic analysis of Klebsiella aerogenes circulating in New Mexico

Klebsiella aerogenes is an opportunistic pathogen and a growing cause of healthcare-associated infections, characterized by multidrug resistance and the emergence of global high-risk clones. However, regional genomic surveillance data remain limited. Here, we sought to characterize the population structure, transmission dynamics and resistance mechanisms of clinical K. aerogenes in Albuquerque, New Mexico. We sequenced 177 clinical isolates collected between 2021 and 2023. We also developed a novel, species-specific PopPUNK database to facilitate rapid, high-resolution typing. The New Mexico K. aerogenes population was diverse but dominated by two global pandemic lineages, ST93 (47.5%) and ST4 (7.9%), which were significantly enriched for the virulence factors yersiniabactin and colibactin. Genomic evidence for recent local transmission was rare, with only four putative transmission pairs identified. The resistome was characterized by intrinsic and adaptive mutations. Nearly all isolates possessed gyrA mutations associated with decreased fluoroquinolone susceptibility. Mutations in the AmpC regulator AmpD and the outer membrane porin Omp36 were common, particularly within the dominant ST93 lineage. These mutations have been associated with increased AmpC-mediated carbapenem resistance. Our findings underscore the critical importance of genomic surveillance to monitor the transmission and evolution of adaptive resistance.

59 BASIC BIOLOGICAL SCIENCES↗

NEAR: Neural Embeddings for Amino acid Relationships

Protein language models (PLMs) have recently demonstrated potential to supplant classical protein database search methods based on sequence alignment, but are slower than common alignment-based tools and appear to be prone to a high rate of false labeling. Here, we present NEAR, a method based on neural representation learning that is designed to improve both speed and accuracy of search for likely homologs in a large protein sequence database. NEAR’s ResNet embedding model is trained using contrastive learning guided by trusted sequence alignments. It computes per-residue embeddings for target and query protein sequences, and identifies alignment candidates with a pipeline consisting of residue-level k-NN search and a simple neighbor aggregation scheme. Tests on a benchmark consisting of trusted remote homologs and randomly shuffled decoy sequences reveal that NEAR substantially improves accuracy relative to state-of-the-art PLMs, with lower memory requirements and faster embedding and search speed. While these results suggest that the NEAR model may be useful for standalone homology detection with increased sensitivity over standard alignment-based methods, in this manuscript we focus on a more straightforward analysis of the model’s value as a high-speed pre-filter for sensitive annotation. In that context, NEAR is at least 5x faster than the pre-filter currently used in the widely-used profile hidden Markov model (pHMM) search tool HMMER3, and also outperforms the pre-filter used in our fast pHMM tool, nail.

59 BASIC BIOLOGICAL SCIENCES↗

Post-composing ontology terms for efficient phenotyping in plant breeding

Abstract Ontologies are widely used in databases to standardize data, improving data quality, integration, and ease of comparison. Within ontologies tailored to diverse use cases, post-composing user-defined terms reconciles the demands for standardization on the one hand and flexibility on the other. In many instances of Breedbase, a digital ecosystem for plant breeding designed for genomic selection, the goal is to capture phenotypic data using highly curated and rigorous crop ontologies, while adapting to the specific requirements of plant breeders to record data quickly and efficiently. For example, post-composing enables users to tailor ontology terms to suit specific and granular use cases such as repeated measurements on different plant parts and special sample preparation techniques. To achieve this, we have implemented a post-composing tool based on orthogonal ontologies providing users with the ability to introduce additional levels of phenotyping granularity tailored to unique experimental designs. Post-composed terms are designed to be reused by all breeding programs within a Breedbase instance but are not exported to the crop reference ontologies. Breedbase users can post-compose terms across various categories, such as plant anatomy, treatments, temporal events, and breeding cycles, and, as a result, generate highly specific terms for more accurate phenotyping.

Mathematical & Computational Biology↗

Signatures of Mollicutes-related endobacteria in publicly available Mucoromycota genomes

ABSTRACT Mucoromycota fungi and their Mollicutes-related endobacteria (MRE) are an ideal system for studying bacterial–fungal interactions and evolution due to the long-term and intimate nature of their interactions. However, methods for detecting MRE face specific challenges due to the poor representation of MRE in sequencing databases coupled with the high sequence divergence of their genomes, making traditional similarity searches unreliable. This has precluded estimations on the diversity of MRE associated with Mucoromycota. To determine the prevalence of previously undetected MRE in fungal genome sequences, we scanned 389 Mucoromycota genome assemblies available from the National Center for Biotechnology Information for the presence of MRE sequences using publicly available tools to map contigs from fungal assemblies to publicly available MRE genomes. We demonstrate a higher diversity of MRE genomes than previously described in Mucoromycota and a lack of cophylogeny between MRE and the majority of their fungal hosts. This supports the late invasion hypothesis regarding MRE acquisition across most of the examined fungal families. In contrast with other Mucoromycota lineages, MRE from the Gigasporaceae displayed some degree of cophylogeny with their hosts, which may indicate that horizontal transmission is restricted between members of this family or that transmission is strictly vertical. These results underscore the need for a refined process to capture sequencing data from potential fungal endosymbionts to discern their evolution and transmission. Screens of fungal genomes for MRE can help improve the quality of fungal genome assemblies while identifying new MRE lineages to further test hypotheses on their origin and evolution. IMPORTANCE Mollicutes-related endobacteria (MRE) are obligate intracellular bacteria found within Mucoromycota fungi. Despite their frequent detection, MRE roles in host functioning are still unknown. Comparative genomic investigations can improve our understanding of the impact of MRE on their fungal hosts by identifying similarities and differences in MRE genome evolution. However, MRE genomes have only been assembled from a small fraction of Mucoromycota hosts. Here, we demonstrate that MRE can be present yet undetected in publicly available Mucoromycota genome assemblies. We use these newfound sequences to assess the broader diversity of MRE and their phylogenetic relationships with respect to their hosts. We demonstrate that publicly available tools can be used to extract novel MRE sequences from assembled fungal genomes leading to insights on MRE evolution. This work contributes to a greater understanding of the fungal microbiome, which is crucial to improving knowledge on the dynamics and impacts of fungi in microbial ecosystems.

59 BASIC BIOLOGICAL SCIENCES↗