Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Dataset Annotations”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Relevancy Ranking of Satellite Dataset Search Results

As the Variety of Earth science datasets increases, science researchers find it more challenging to discover and select the datasets that best fit their needs. The most common way of search providers to address this problem is to rank the datasets returned for a query by their likely relevance to the user. Large web page search engines typically use text matching supplemented with reverse link counts, semantic annotations and user intent modeling. However, this produces uneven results when applied to dataset metadata records simply externalized as a web page. Fortunately, data and search provides have decades of experience in serving data user communities, allowing them to form heuristics that leverage the structure in the metadata together with knowledge about the user community. Some of these heuristics include specific ways of matching the user input to the essential measurements in the dataset and determining overlaps of time range and spatial areas. Heuristics based on the novelty of the datasets can prioritize later, better versions of data over similar predecessors. And knowledge of how different user types and communities use data can be brought to bear in cases where characteristics of the user (discipline, expertise) or their intent (applications, research) can be divined. The Earth Observing System Data and Information System has begun implementing some of these heuristics in the relevancy algorithm of its Common Metadata Repository search engine.

science data management↗

Pickaxe: a Python library for the prediction of novel metabolic reactions

Abstract Background Biochemical reaction prediction tools leverage enzymatic promiscuity rules to generate reaction networks containing novel compounds and reactions. The resulting reaction networks can be used for multiple applications such as designing novel biosynthetic pathways and annotating untargeted metabolomics data. It is vital for these tools to provide a robust, user-friendly method to generate networks for a given application. However, existing tools lack the flexibility to easily generate networks that are tailor-fit for a user’s application due to lack of exhaustive reaction rules, restriction to pre-computed networks, and difficulty in using the software due to lack of documentation. Results Here we present Pickaxe, an open-source, flexible software that provides a user-friendly method to generate novel reaction networks. This software iteratively applies reaction rules to a set of metabolites to generate novel reactions. Users can select rules from the prepackaged JN1224min ruleset, derived from MetaCyc, or define their own custom rules. Additionally, filters are provided which allow for the pruning of a network on-the-fly based on compound and reaction properties. The filters include chemical similarity to target molecules, metabolomics, thermodynamics, and reaction feasibility filters. Example applications are given to highlight the capabilities of Pickaxe: the expansion of common biological databases with novel reactions, the generation of industrially useful chemicals from a yeast metabolome database, and the annotation of untargeted metabolomics peaks from an E. coli dataset. Conclusion Pickaxe predicts novel metabolic reactions and compounds, which can be used for a variety of applications. This software is open-source and available as part of the MINE Database python package ( https://pypi.org/project/minedatabase/ ) or on GitHub ( https://github.com/tyo-nu/MINE-Database ). Documentation and examples can be found on Read the Docs ( https://mine-database.readthedocs.io/en/latest/ ). Through its documentation, pre-packaged features, and customizable nature, Pickaxe allows users to generate novel reaction networks tailored to their application.

59 BASIC BIOLOGICAL SCIENCES↗

YeastWGD2025

Supplementary data for Discovery of additional ancient genome duplications in yeasts wgd_syn / - directory containing wgd syn output for all contiguous genomes [dataset] Tree - phylogeny [dataset]Duplications - duplication table from OrthoFinder output KOannotations - KEGG annotations used for enrichment analysis IPRannotations - InterPro annotations used for enrichment analysis DipodascalesOrthogroups - formatted orthogroup assignments for Dipodascales genes.fa and .gff3 files for each new genome assembly are also provided, those these are not required to replicate the analysis

Genomics↗

EI_MS_ML

The unambiguous identification of compounds from their electron ionization mass (EI-MS) spectra remains a significant unsolved problem in the field of metabolomics and analytical chemistry as a whole. Typically EI-MS spectra are compared using various mathematical operations that convert the spectral similarity or differences into a distance-like metric that roughly approximates the similarity of any two spectra. A commonly used metric for this is the cosine similarity metric which has values close to one for very similar spectra and a value of zero for very dissimilar spectra; however, no metric is perfect. Due to the prevalence of structurally-similar compounds such as isomers and the prevalence of certain fragmentation patterns across structurally-dissimilar compounds, the unambiguous assignment of EI-MS spectra compounds remains difficult. Frequently, querying an observed EI-MS spectrum against a large database such as the NIST17 library yields multiple possible assignments requiring the end user to distinguish between multiple high scoring hits, or multiple low scoring hits while keeping in mind that the correct hit may not be in the database at all. Although techniques such as orthogonal information from techniques such as chromatography can greatly aid in unambiguous assignment, this also requires more complicated experimental designs and access to more complicated analytical instrumentation. Substructures can be trivially detected and represented as strings using a previously published technique called node coloring from a known chemical structure. However, for experimentally-derived EI-MS spectra this information must be derived from the spectra itself (i.e., because we do not know what compound it represents). To achieve this, the software uses techniques from the field of machine learning and a large training dataset of EI-MS spectra corresponding to known structures annotated with substructure strings, to build models that can predict the presence of a given chemical substructure from an EI-MS spectrum directly.If these predictions are of high-quality (i.e., are unlikely to be false positives), the presence of one or more predicted substructures can be used to constrain the number of possible hits for a query spectrum. Mathematically, this restriction could be expressed in many forms, but the most straight-forward implementation is to weight the cosine similarity of a query spectrum and a plausible database match with a Tanimoto-like coefficient based on the ratio of the number of substructures predicted to the number of substructures present in the potential database hit. Determining which combination of models best reduces assignment ambiguity will be achieved using a combination of manual curation and optimization techniques such as genetic algorithms. This software will perform all the steps necessary to construct said models from a training dataset and evaluate them using a holdout dataset. Various statistical analyses can be performed to determine if this approach does decrease assignment ambiguity. For example, if this approach works, on average, the rank-order of the correct assignment for the holdout set of EI-MS spectra should decrease and the weighted cosine similarities for most of the possible matches in the database should be better than the unweighted cosine similarities. Furthermore, this same pipeline can be used on real experimental data to generate less ambiguous assignments.

Mitchell, Joshua↗

efam: an e xpanded, metaproteome-supported HMM profile database of viral protein fam ilies

Viruses infect, reprogram and kill microbes, leading to profound ecosystem consequences, from elemental cycling in oceans and soils to microbiome-modulated diseases in plants and animals. Although metagenomic datasets are increasingly available, identifying viruses in them is challenging due to poor representation and annotation of viral sequences in databases. Here, we establish efam, an expanded collection of Hidden Markov Model (HMM) profiles that represent viral protein families conservatively identified from the Global Ocean Virome 2.0 dataset. This resulted in 240 311 HMM profiles, each with at least 2 protein sequences, making efam >7-fold larger than the next largest, pan-ecosystem viral HMM profile database. Adjusting the criteria for viral contig confidence from ‘conservative’ to ‘eXtremely Conservative’ resulted in 37 841 HMM profiles in our efam-XC database. To assess the value of this resource, we integrated efam-XC into VirSorter viral discovery software to discover viruses from less-studied, ecologically distinct oxygen minimum zone (OMZ) marine habitats. This expanded database led to an increase in viruses recovered from every tested OMZ virome by ~24% on average (up to ~42%) and especially improved the recovery of often-missed shorter contigs (<5 kb). Additionally, to help elucidate lesser-known viral protein functions, we annotated the profiles using multiple databases from the DRAM pipeline and virion-associated metaproteomic data, which doubled the number of annotations obtainable by standard, single-database annotation approaches. Together, these marine resources (efam and efam-XC) are provided as searchable, compressed HMM databases that will be updated bi-annually to help maximize viral sequence discovery and study from any ecosystem.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A clinically and genomically annotated nerve sheath tumor biospecimen repository

Nerve sheath tumors occur as a heterogeneous group of neoplasms in patients with neurofibromatosis type 1 (NF1). The malignant form represents the most common cause of death in people with NF1, and even when benign, these tumors can result in significant disfigurement, neurologic dysfunction, and a range of profound symptoms. Lack of human tissue across the peripheral nerve tumors common in NF1 has been a major limitation in the development of new therapies. To address this unmet need, we have created an annotated collection of patient tumor samples, patient-derived cell lines, and patient-derived xenografts, and carried out high-throughput genomic and transcriptomic characterization to serve as a resource for further biologic and preclinical therapeutic studies. In this work, we release genomic and transcriptomic datasets comprised of 55 tumor samples derived from 23 individuals, complete with clinical annotation. All data are publicly available through the NF Data Portal and at http://synapse.org/jhubiobank.

60 APPLIED LIFE SCIENCES↗

Building kinetic models for metabolic engineering

Kinetic formalisms of metabolism link metabolic fluxes to enzyme levels, metabolite concentrations and their allosteric regulatory interactions. Though they require the identification of physiologically relevant values for numerous parameters, kinetic formalisms uniquely establish a mechanistic link across heterogeneous omics datasets and provide an overarching vantage point to effectively inform metabolic engineering strategies. Advances in computational power, gene annotation coverage, and formalism standardization have led to significant progress over the past few years. However, careful interpretation of model predictions, limited metabolic flux datasets, and assessment of parameter sensitivity remain as challenges. In this study we highlight fundamental considerations which influence model quality and prediction, advances in methodologies, and success stories of deploying kinetic models to guide metabolic engineering.

59 BASIC BIOLOGICAL SCIENCES↗

Deep active learning for classifying cancer pathology reports

Abstract Background Automated text classification has many important applications in the clinical setting; however, obtaining labelled data for training machine learning and deep learning models is often difficult and expensive. Active learning techniques may mitigate this challenge by reducing the amount of labelled data required to effectively train a model. In this study, we analyze the effectiveness of 11 active learning algorithms on classifying subsite and histology from cancer pathology reports using a Convolutional Neural Network as the text classification model. Results We compare the performance of each active learning strategy using two differently sized datasets and two different classification tasks. Our results show that on all tasks and dataset sizes, all active learning strategies except diversity-sampling strategies outperformed random sampling, i.e., no active learning. On our large dataset (15K initial labelled samples, adding 15K additional labelled samples each iteration of active learning), there was no clear winner between the different active learning strategies. On our small dataset (1K initial labelled samples, adding 1K additional labelled samples each iteration of active learning), marginal and ratio uncertainty sampling performed better than all other active learning techniques. We found that compared to random sampling, active learning strongly helps performance on rare classes by focusing on underrepresented classes. Conclusions Active learning can save annotation cost by helping human annotators efficiently and intelligently select which samples to label. Our results show that a dataset constructed using effective active learning techniques requires less than half the amount of labelled data to achieve the same performance as a dataset constructed using random sampling.

59 BASIC BIOLOGICAL SCIENCES↗

Advancing Tassel Detection and Counting: Annotation and Algorithms

Tassel counts provide valuable information related to flowering and yield prediction in maize, but are expensive and time-consuming to acquire via traditional manual approaches. High-resolution RGB imagery acquired by unmanned aerial vehicles (UAVs), coupled with advanced machine learning approaches, including deep learning (DL), provides a new capability for monitoring flowering. In this article, three state-of-the-art DL techniques, CenterNet based on point annotation, task-aware spatial disentanglement (TSD), and detecting objects with recursive feature pyramids and switchable atrous convolution (DetectoRS) based on bounding box annotation, are modified to improve their performance for this application and evaluated for tassel detection relative to Tasselnetv2+. The dataset for the experiments is comprised of RGB images of maize tassels from plant breeding experiments, which vary in size, complexity, and overlap. Results show that the point annotations are more accurate and simpler to acquire than the bounding boxes, and bounding box-based approaches are more sensitive to the size of the bounding boxes and background than point-based approaches. Overall, CenterNet has high accuracy in comparison to the other techniques, but DetectoRS can better detect early-stage tassels. The results for these experiments were more robust than Tasselnetv2+, which is sensitive to the number of tassels in the image.

54 ENVIRONMENTAL SCIENCES↗

deadtrees.earth — An open-access and interactive database for centimeter-scale aerial imagery to uncover global tree mortality dynamics

Excessive tree mortality is a global concern and remains poorly understood as it is a complex phenomenon. We lack global and temporally continuous coverage on tree mortality data. Ground-based observations on tree mortality, e.g., derived from national inventories, are very sparse, and may not be standardized or spatially explicit. Earth observation data, combined with supervised machine learning, offer a promising approach to map overstory tree mortality in a consistent manner over space and time. However, global-scale machine learning requires broad training data covering a wide range of environmental settings and forest types. Low altitude observation platforms (e.g., drones or airplanes) provide a cost-effective source of training data by capturing high-resolution orthophotos of overstory tree mortality events at centimeter-scale resolution. Here, we introduce deadtrees.earth, an open-access platform hosting more than two thousand centimeter-resolution orthophotos, covering more than 1,000,000 ha, of which more than 58,000 ha are manually annotated with live/dead tree classifications. This community-sourced and rigorously curated dataset can serve as a comprehensive reference dataset to uncover tree mortality patterns from local to global scales using space-based Earth observation data and machine learning models. This will provide the basis to attribute tree mortality patterns to environmental changes or project tree mortality dynamics to the future. The open nature of deadtrees.earth, together with its curation of high-quality, spatially representative, and ecologically diverse data will continuously increase our capacity to uncover and understand tree mortality dynamics.

Citizen science↗

The value of human data annotation for machine learning based anomaly detection in environmental systems

Anomaly detection is the process of identifying unexpected data samples in datasets. Automated anomaly detection is either performed using supervised machine learning models, which require a labelled dataset for their calibration, or unsupervised models, which do not require labels. While academic research has produced a vast array of tools and machine learning models for automated anomaly detection, the research community focused on environmental systems still lacks a comparative analysis that is simultaneously comprehensive, objective, and systematic. This knowledge gap is addressed for the first time in this study, where 15 different supervised and unsupervised anomaly detection models are evaluated on 5 different environmental datasets from engineered and natural aquatic systems. To this end, anomaly detection performance, labelling efforts, as well as the impact of model and algorithm tuning are taken into account. As a result, our analysis reveals the relative strengths and weaknesses of the different approaches in an objective manner without bias for any particular paradigm in machine learning. Most importantly, our results show that expert-based data annotation is extremely valuable for anomaly detection based on machine learning.

54 ENVIRONMENTAL SCIENCES↗

Barcoded overexpression screens in gut Bacteroidales identify genes with roles in carbon utilization and stress resistance

Abstract A mechanistic understanding of host-microbe interactions in the gut microbiome is hindered by poorly annotated bacterial genomes. While functional genomics can generate large gene-to-phenotype datasets to accelerate functional discovery, their applications to study gut anaerobes have been limited. For instance, most gain-of-function screens of gut-derived genes have been performed in Escherichia coli and assayed in a small number of conditions. To address these challenges, we develop Barcoded Overexpression BActerial shotgun library sequencing (Boba-seq). We demonstrate the power of this approach by assaying genes from diverse gut Bacteroidales overexpressed in Bacteroides thetaiotaomicron . From hundreds of experiments, we identify new functions and phenotypes for 29 genes important for carbohydrate metabolism or tolerance to antibiotics or bile salts. Highlights include the discovery of a d -glucosamine kinase, a raffinose transporter, and several routes that increase tolerance to ceftriaxone and bile salts through lipid biosynthesis. This approach can be readily applied to develop screens in other strains and additional phenotypic assays.

59 BASIC BIOLOGICAL SCIENCES↗

Gaia: An AI-enabled genomic context–aware platform for protein sequence annotation

Protein sequence similarity search is fundamental to biology research, but current methods are typically not able to consider crucial genomic context information indicative of protein function, especially in microbial systems. Here, we present Gaia (Genomic AI Annotator), a sequence annotation platform that enables rapid, context-aware protein sequence search across genomic datasets. Gaia leverages gLM2, a mixed-modality genomic language model trained on both amino acid sequences and their genomic neighborhoods to generate embeddings that integrate sequence-structure-context information. This approach allows for the identification of functionally and/or evolutionarily related genes that are found in conserved genomic contexts, which may be missed by traditional sequence- or structure-based search alone. Gaia enables real-time search of a curated database comprising more than 85 million protein clusters from 131,744 microbial genomes. We compare the homolog retrieval performance of Gaia search against other embedding and alignment-based approaches. We provide Gaia as a web-based, freely available tool.

Jha, Nishant↗

Linked Open Data in the Global Change Information System (GCIS)

The U.S. Global Change Research Program (http://globalchange.gov) coordinates and integrates federal research on changes in the global environment and their implications for society. The USGCRP is developing a Global Change Information System (GCIS) that will centralize access to data and information related to global change across the U.S. federal government. The first implementation will focus on the 2013 National Climate Assessment (NCA) . (http://assessment.globalchange.gov) The NCA integrates, evaluates, and interprets the findings of the USGCRP; analyzes the effects of global change on the natural environment, agriculture, energy production and use, land and water resources, transportation, human health and welfare, human social systems, and biological diversity; and analyzes current trends in global change, both human-induced and natural, and projects major trends for the subsequent 25 to 100 years. The NCA has received over 500 distinct technical inputs to the process, many of which are reports distilling and synthesizing even more information, coming from thousands of individuals around the federal, state and local governments, academic institutions and non-governmental organizations. The GCIS will present a web-based version of the NCA including annotations linking the findings and content of the NCA with the scientific research, datasets, models, observations, etc. that led to its conclusions. It will use semantic tagging and a linked data approach, assigning globally unique, persistent, resolvable identifiers to all of the related entities and capturing and presenting the relationships between them, both internally and referencing out to other linked data sources and back to agency data centers. The developing W3C PROV Data Model and ontology will be used to capture the provenance trail and present it in both human readable web pages and machine readable formats such as RDF and SPARQL. This will improve visibility into the assessment process, increase understanding and reproducibility, and ultimately increase credibility and trust of the resulting report. Building on the foundation of the NCA, longer term plans for the GCIS include extending these capabilities throughout the U.S. Global Change Research Program, centralizing access to global change data and information across the thirteen agencies that comprise the program.

Tilmes, Curt A.↗

NASA GeneLab Concept of Operations

NASA's GeneLab aims to greatly increase the number of scientists that are using data from space biology investigations on board ISS, emphasizing a systems biology approach to the science. When completed, GeneLab will provide the integrated software and hardware infrastructure, analytical tools and reference datasets for an assortment of model organisms. GeneLab will also provide an environment for scientists to collaborate thereby increasing the possibility for data to be reused for future experimentation. To maximize the value of data from life science experiments performed in space and to make the most advantageous use of the remaining ISS research window, GeneLab will apply an open access approach to conducting spaceflight experiments by generating, and sharing the datasets derived from these biological studies in space.Onboard the ISS, a wide variety of model organisms will be studied and returned to Earth for analysis. Laboratories on the ground will analyze these samples and provide genomic, transcriptomic, metabolomic and proteomic data. Upon receipt, NASA will conduct data quality control tasks and format raw data returned from the omics centers into standardized, annotated information sets that can be readily searched and linked to spaceflight metadata. Once prepared, the biological datasets, as well as any analysis completed, will be made public through the GeneLab Space Bioinformatics System webb as edportal. These efforts will support a collaborative research environment for spaceflight studies that will closely resemble environments created by the Department of Energy (DOE), National Center for Biotechnology Information (NCBI), and other institutions in additional areas of study, such as cancer and environmental biology. The results will allow for comparative analyses that will help scientists around the world take a major leap forward in understanding the effect of microgravity, radiation, and other aspects of the space environment on model organisms. These efforts will speed the process of scientific sharing, iteration, and discovery.

Space Life Science↗

Ontology-Enriched Specifications Enabling Findable, Accessible, Interoperable, and Reusable Marine Metagenomic Datasets in Cyberinfrastructure Systems

Marine microbial ecology requires the systematic comparison of biogeochemical and sequence data to analyze environmental influences on the distribution and variability of microbial communities. With ever-increasing quantities of metagenomic data, there is a growing need to make datasets Findable, Accessible, Interoperable, and Reusable (FAIR) across diverse ecosystems. FAIR data is essential to developing analytical frameworks that integrate microbiological, genomic, ecological, oceanographic, and computational methods. Although community standards defining the minimal metadata required to accompany sequence data exist, they haven’t been consistently used across projects, precluding interoperability. Moreover, these data are not machine-actionable or discoverable by cyberinfrastructure systems. By making ‘omic and physicochemical datasets FAIR to machine systems, we can enable sequence data discovery and reuse based on machine-readable descriptions of environments or physicochemical gradients. In this work, we developed a novel technical specification for dataset encapsulation for the FAIR reuse of marine metagenomic and physicochemical datasets within cyberinfrastructure systems. This includes using Frictionless Data Packages enriched with terminology from environmental and life-science ontologies to annotate measured variables, their units, and the measurement devices used. This approach was implemented in Planet Microbe, a cyberinfrastructure platform and marine metagenomic web-portal. Here, we discuss the data properties built into the specification to make global ocean datasets FAIR within the Planet Microbe portal. We additionally discuss the selection of, and contributions to marine-science ontologies used within the specification. Finally, we use the system to discover data by which to answer various biological questions about environments, physicochemical gradients, and microbial communities in meta-analyses. This work represents a future direction in marine metagenomic research by proposing a specification for FAIR dataset encapsulation that, if adopted within cyberinfrastructure systems, would automate the discovery, exchange, and re-use of data needed to answer broader reaching questions than originally intended.

59 BASIC BIOLOGICAL SCIENCES↗

Overcoming small minirhizotron datasets using transfer learning

Minirhizotron technology is widely used to study root growth and development. Yet, standard approaches for tracing roots in minirhiztron imagery is extremely tedious and time consuming. Machine learning approaches can help to automate this task. However, lack of enough annotated training data is a major limitation for the application of machine learning methods. Transfer learning is a useful technique to help with training when available datasets are limited. In this paper, we investigated the effect of pre-trained features from the massives-cale, irrelevant ImageNet dataset and a relatively moderate-scale, but relevant peanut root dataset on switchgrass root imagery segmentation applications. We compiled two minirhizotron image datasets to accomplish this study: one with 17,550 peanut root images and another with 28 switchgrass root images. Both datasets were paired with manually labeled ground truth masks. Deep neural networks based on the U-net architecture were used with different pre-trained features as initialization for automated, precise pixel-wise root segmentation in minirhizotron imagery. We observed that features pre-trained on a closely related but relatively moderate size dataset like our peanut dataset were more effective than features pre-trained on the large but unrelated ImageNet dataset. Here, we achieved high quality segmentation on peanut root dataset with 99.04% accuracy at the pixel-level and overcame errors in human-labeled ground truth masks. By applying transfer learning technique on limited switchgrass dataset with features pre-trained on peanut dataset, we obtained 99% segmentation accuracy in switchgrass imagery using only 21 images for training (fine tuning). Furthermore, the peanut pre-trained features can help the model converge faster and have much more stable performance.

59 BASIC BIOLOGICAL SCIENCES↗