Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Dataset Annotations”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Characterizing Families of Spectral Similarity Scores and Their Use Cases for Gas Chromatography–Mass Spectrometry Small Molecule Identification

Metabolomics provides a unique snapshot into the world of small molecules and the complex biological processes that govern the human, animal, plant, and environmental ecosystems encapsulated by the One Health modeling framework. However, this “molecular snapshot” is only as informative as the number of metabolites confidently identified within it. The spectral similarity (SS) score is traditionally used to identify compound(s) in mass spectrometry approaches to metabolomics, where spectra are matched to reference libraries of candidate spectra. Unfortunately, there is little consensus on which of the dozens of available SS metrics should be used. This lack of standard SS score creates analytic uncertainty and potentially leads to issues in reproducibility, especially as these data are integrated across other domains. In this work, we use metabolomic spectral similarity as a case study to showcase the challenges in consistency within just one piece of the One Health framework that must be addressed to enable data science approaches for One Health problems. Here, using a large cohort of datasets comprising both standard and complex datasets with expert-verified truth annotations, we evaluated the effectiveness of 66 similarity metrics to delineate between correct matches (true positives) and incorrect matches (true negatives). We additionally characterize the families of these metrics to make informed recommendations for their use. Our results indicate that specific families of metrics (the Inner Product, Correlative, and Intersection families of scores) tend to perform better than others, with no single similarity metric performing optimally for all queried spectra. This work and its findings provide an empirically-based resource for researchers to use in their selection of similarity metrics for GC-MS identification, increasing scientific reproducibility through taking steps towards standardizing identification workflows.

59 BASIC BIOLOGICAL SCIENCES↗

BigNeuron: a resource to benchmark and predict performance of algorithms for automated tracing of neurons in light microscopy datasets

BigNeuron is an open community bench-testing platform with the goal of setting open standards for accurate and fast automatic neuron tracing. We gathered a diverse set of image volumes across several species that is representative of the data obtained in many neuroscience laboratories interested in neuron tracing. Here, we report generated gold standard manual annotations for a subset of the available imaging datasets and quantified tracing quality for 35 automatic tracing algorithms. The goal of generating such a hand-curated diverse dataset is to advance the development of tracing algorithms and enable generalizable benchmarking. Together with image quality features, we pooled the data in an interactive web application that enables users and developers to perform principal component analysis, t-distributed stochastic neighbor embedding, correlation and clustering, visualization of imaging and tracing data, and benchmarking of automatic tracing algorithms in user-defined data subsets. The image quality metrics explain most of the variance in the data, followed by neuromorphological features related to neuron size. Furthermore, we observed that diverse algorithms can provide complementary information to obtain accurate results and developed a method to iteratively combine methods and generate consensus reconstructions. The consensus trees obtained provide estimates of the neuron structure ground truth that typically outperform single algorithms in noisy datasets. However, specific algorithms may outperform the consensus tree strategy in specific imaging conditions. Finally, to aid users in predicting the most accurate automatic tracing results without manual annotations for comparison, we used support vector machine regression to predict reconstruction quality given an image volume and a set of automatic tracings.

97 MATHEMATICS AND COMPUTING↗

Profile Images and Annotations for Vehicle Re-identification Algorithms (PRIMAVERA)

This dataset contains 636,246 profile images of vehicles representing 13,963 unique vehicles. The data was collected by a set of roadside sensors over the course of three years. Each time a vehicle passed by one of the sensors, a series of images was collected. The images were processed to detect and localize each vehicle, and a license plate reader collocated with the sensor was used to provide a unique ID for the vehicle. Actual license plate numbers have been obfuscated by replacing with an arbitrary numerical ID for each vehicle. After localizing the vehicle in each image, the original RGB image was rotated, scaled, and shifted to produce a new RGB image of size 234x234 pixels such that the outermost two wheels are located at predetermined pixel locations in the image. In this way, all vehicle images are aligned to one another. This registration process occasionally results in a portion of certain vehicles being cutoff at the edges of the image. The dataset has been partitioned into two sets called training and validation. The two partitions no common vehicles, i.e., a vehicle present in one partition is guaranteed not to be present in the other. In this way, an algorithm can be validated against a set of new vehicles that were not seen during the training process. The training set contains 543,926 images from 64,440 vehicle passes representing 11,918 unique vehicles, while the validation set contains 92,320 images from 10,991 vehicle passes representing 2,045 unique vehicles. Vehicle images are organized by directories corresponding to unique vehicles. The file naming scheme is as follows: veh_{vehID}_tr_{passID}_{frameID}_{elevation}_{timeofday}.jpg where {vehID} is the vehicle ID (unique across the entire dataset), {passID} is an identifier for each tracked vehicle pass (unique across the entire dataset), {frameID} is the index of the frame within the given vehicle pass starting at 0, {elevation} is a two-letter string indicating whether the sensor was elevated (el) or at ground-level (gl), and {timeofday} is a two-letter string indicating whether the image was captured during daytime (dt) or nighttime (nt).

image↗

A new paradigm in electron microscopy: Automated microstructure analysis utilizing a dynamic segmentation convolutional neutral network

Over the past half century, the transmission electron microscope enabled insight into the fundamental arrangements and structures of materials. State-of-the-art electron microscopes can acquire large image datasets across multiple imaging modalities. However, the manual annotation process for feature or defect quantification may not be feasible with the modern microscope. Convolutional neural networks emerged to characterize individual microstructural features from an image in a cost-effective, consistent manner. However, many of these neural network approaches rely on thousands to hundreds of thousands of manual annotations of each feature type across hundreds of images to train the network for adequate performance. This work focused on the development and application of a pixel-wise defect detection machine-learning dynamic segmentation convolutional neural network with associated automated acquisition and postprocessing to identify microstructural features rapidly and quantitatively from a small initial dataset incorporating multiple imaging modes. The approach was demonstrated for characterization of superalloy 718 from both single image acquisition on multiple detectors to in-situ evolution captured with a single detector on a standard desktop computer to demonstrate the low barrier to entry required for widespread adoption. Pixel-by-pixel class identification was excellent with strong identification of chemically distinct phases, structurally distinct phases, and defect structures, thus demonstrating the new paradigm of machine learning-assisted characterization.

36 MATERIALS SCIENCE↗

Metabolite discovery through global annotation of untargeted metabolomics data

Liquid chromatography–high-resolution mass spectrometry (LC-MS)-based metabolomics aims to identify and quantify all metabolites, but most LC-MS peaks remain unidentified. Here we present a global network optimization approach, NetID, to annotate untargeted LC-MS metabolomics data. The approach aims to generate, for all experimentally observed ion peaks, annotations that match the measured masses, retention times and (when available) tandem mass spectrometry fragmentation patterns. Peaks are connected based on mass differences reflecting adduction, fragmentation, isotopes, or feasible biochemical transformations. Global optimization generates a single network linking most observed ion peaks, enhances peak assignment accuracy, and produces chemically informative peak–peak relationships, including for peaks lacking tandem mass spectrometry spectra. Applying this approach to yeast and mouse data, we identified five previously unrecognized metabolites (thiamine derivatives and N-glucosyl-taurine). Isotope tracer studies indicate active flux through these metabolites. Furthermore, NetID applies existing metabolomic knowledge and global optimization to substantially improve annotation coverage and accuracy in untargeted metabolomics datasets, facilitating metabolite discovery.

59 BASIC BIOLOGICAL SCIENCES↗

Are Phosphatidic Acids Ubiquitous in Mammalian Tissues or Overemphasized in Mass Spectrometry Imaging Applications?

Abstract Mass spectrometry imaging (MSI) is an invaluable tool for the spatial visualization of molecules in vivo. However, the question of whether observed annotations are endogenous or artificial (i. e., from in‐source fragmentation) is critical and has been largely unexplored in multimodal MSI. In matrix‐assisted laser desorption/ionization (MALDI)‐MSI datasets from researchers worldwide, PAs were found to represent up to 18 % of annotations in rat brain. Rat brain was additionally imaged here using nanospray desorption electrospray ionization (nano‐DESI), a softer ionization strategy. No PAs observed with MALDI were present in the nano‐DESI dataset. Further investigation strongly indicated lipid fragmentation to PAs for MALDI‐MSI, but not with nano‐DESI‐MSI. We finally extend this observation to the MALDI‐MSI analyses of human tissues, showing that PA annotations comprised up to 16 % of annotations. Therefore, this study shows that MSI annotations should be carefully interrogated, as in‐source fragmentation or modification of lipids may contribute substantially to false annotations and incorrect biological interpretations.

Vandergrift, Gregory W.↗

RadioGalaxyNET: Dataset and novel computer vision algorithms for the detection of extended radio galaxies and infrared hosts

Abstract Creating radio galaxy catalogues from next-generation deep surveys requires automated identification of associated components of extended sources and their corresponding infrared hosts. In this paper, we introduce RadioGalaxyNET, a multimodal dataset, and a suite of novel computer vision algorithms designed to automate the detection and localization of multi-component extended radio galaxies and their corresponding infrared hosts. The dataset comprises 4 155 instances of galaxies in 2 800 images with both radio and infrared channels. Each instance provides information about the extended radio galaxy class, its corresponding bounding box encompassing all components, the pixel-level segmentation mask, and the keypoint position of its corresponding infrared host galaxy. RadioGalaxyNET is the first dataset to include images from the highly sensitive Australian Square Kilometre Array Pathfinder (ASKAP) radio telescope, corresponding infrared images, and instance-level annotations for galaxy detection. We benchmark several object detection algorithms on the dataset and propose a novel multimodal approach to simultaneously detect radio galaxies and the positions of infrared hosts.

Astronomy & Astrophysics↗

Implementation of FAIR principles in the IPCC: the WGI AR6 Atlas repository

The Sixth Assessment Report (AR6) of the Intergovernmental Panel on Climate Change (IPCC) has adopted the FAIR Guiding Principles. We present the Atlas chapter of Working Group I (WGI) as a test case. We describe the application of the FAIR principles in the Atlas, the challenges faced during its implementation, and those that remain for the future. We introduce the open source repository resulting from this process, including coding (e.g., annotated Jupyter notebooks), data provenance, and some aggregated datasets used in some figures in the Atlas chapter and its interactive companion (the Interactive Atlas), open to scrutiny by the scientific community and the general public. We describe the informal pilot review conducted on this repository to gather recommendations that led to significant improvements. Finally, a working example illustrates the re-use of the repository resources to produce customized regional information, extending the Interactive Atlas products and running the code interactively in a web browser using Jupyter notebooks.

54 ENVIRONMENTAL SCIENCES↗

Attention-based Aspect Reasoning for Knowledge Base Question Answering on Clinical Notes

Question Answering (QA) in clinical notes has gained a lot of attention in the past few years. Existing machine reading comprehension approaches in clinical domain can only handle questions about a single block of clinical texts and fail to retrieve information about different patients and clinical notes. To handle more complex questions, we aim at creating knowledge base from clinical notes to link different patients and clinical notes, and performing knowledge base question answering (KBQA). Based on the expert annotations in n2c2, we first created the ClinicalKBQA dataset that includes 8,952 QA pairs and covers questions about seven medical topics through 322 question templates. Then, we proposed an attention-based aspect reasoning (AAR) method for KBQA and investigated the impact of different aspects of answers (e.g., entity, type, path, and context) for prediction. The AAR method achieves better performance due to the well-designed encoder and attention mechanism. In the experiments, we find that both aspects, type and path, enable the model to identify answers satisfying the general conditions and produce lower precision and higher recall. On the other hand, the aspects, entity and context, limit the answers by node-specific information and lead to higher precision and lower recall.

Wang, Ping↗

MassIVE MSV000087434 - Raw data files for NetID manuscript

This dataset is associated with this manuscript: Metabolite discovery through global annotation of untargeted metabolomics data. Li Chen, Wenyun Lu, Lin Wang, Xi Xing, Xin Teng, Xianfeng Zeng, Antonio D Muscarella, Yihui Shen, Alexis Cowan, Melanie R McReynolds, Brandon Kennedy, Ashley M Lato, Shawn R Campagna, Mona Singh, Joshua Rabinowitz. The dataset has six main folders: (1) Yeast-MS1 (2) Yeast-labeling (3) Liver-MS1 (4) Liver-targeted-MS2 (separate folders for mzxml files and matched csv files used as the inclusion list to set up PRM targeted-MS2 runs) (5) Mouse-labeling (6) Glucosyl_taurine-liver-vs-synthesis. Please refer to README-raw-data-NetID-20210512.doc for more details. [doi:10.25345/C5WV53] [dataset license: CC0 1.0 Universal (CC0 1.0)]

full scan MS1↗

Variation in forest root image annotation by experts, novices, and AI

Abstract Background The manual study of root dynamics using images requires huge investments of time and resources and is prone to previously poorly quantified annotator bias. Artificial intelligence (AI) image-processing tools have been successful in overcoming limitations of manual annotation in homogeneous soils, but their efficiency and accuracy is yet to be widely tested on less homogenous, non-agricultural soil profiles, e.g., that of forests, from which data on root dynamics are key to understanding the carbon cycle. Here, we quantify variance in root length measured by human annotators with varying experience levels. We evaluate the application of a convolutional neural network (CNN) model, trained on a software accessible to researchers without a machine learning background, on a heterogeneous minirhizotron image dataset taken in a multispecies, mature, deciduous temperate forest. Results Less experienced annotators consistently identified more root length than experienced annotators. Root length annotation also varied between experienced annotators. The CNN root length results were neither precise nor accurate, taking ~ 10% of the time but significantly overestimating root length compared to expert manual annotation ( p = 0.01). The CNN net root length change results were closer to manual ( p = 0.08) but there remained substantial variation. Conclusions Manual root length annotation is contingent on the individual annotator. The only accessible CNN model cannot yet produce root data of sufficient accuracy and precision for ecological applications when applied to a complex, heterogeneous forest image dataset. A continuing evaluation and development of accessible CNNs for natural ecosystems is required.

Handy, Grace↗

Identification of mobile genetic elements with geNomad

Identifying and characterizing mobile genetic elements in sequencing data is essential for understanding their diversity, ecology, biotechnological applications and impact on public health. Here we introduce geNomad, a classification and annotation framework that combines information from gene content and a deep neural network to identify sequences of plasmids and viruses. geNomad uses a dataset of more than 200,000 marker protein profiles to provide functional gene annotation and taxonomic assignment of viral genomes. Using a conditional random field model, geNomad also detects proviruses integrated into host genomes with high precision. In benchmarks, geNomad achieved high classification performance for diverse plasmids and viruses (Matthews correlation coefficient of 77.8% and 95.3%, respectively), substantially outperforming other tools. Leveraging geNomad’s speed and scalability, we processed over 2.7 trillion base pairs of sequencing data, leading to the discovery of millions of viruses and plasmids that are available through the IMG/VR and IMG/PR databases. geNomad is available at https://portal.nersc.gov/genomad.

59 BASIC BIOLOGICAL SCIENCES↗

Pickaxe: a Python library for the prediction of novel metabolic reactions

Abstract Background Biochemical reaction prediction tools leverage enzymatic promiscuity rules to generate reaction networks containing novel compounds and reactions. The resulting reaction networks can be used for multiple applications such as designing novel biosynthetic pathways and annotating untargeted metabolomics data. It is vital for these tools to provide a robust, user-friendly method to generate networks for a given application. However, existing tools lack the flexibility to easily generate networks that are tailor-fit for a user’s application due to lack of exhaustive reaction rules, restriction to pre-computed networks, and difficulty in using the software due to lack of documentation. Results Here we present Pickaxe, an open-source, flexible software that provides a user-friendly method to generate novel reaction networks. This software iteratively applies reaction rules to a set of metabolites to generate novel reactions. Users can select rules from the prepackaged JN1224min ruleset, derived from MetaCyc, or define their own custom rules. Additionally, filters are provided which allow for the pruning of a network on-the-fly based on compound and reaction properties. The filters include chemical similarity to target molecules, metabolomics, thermodynamics, and reaction feasibility filters. Example applications are given to highlight the capabilities of Pickaxe: the expansion of common biological databases with novel reactions, the generation of industrially useful chemicals from a yeast metabolome database, and the annotation of untargeted metabolomics peaks from an E. coli dataset. Conclusion Pickaxe predicts novel metabolic reactions and compounds, which can be used for a variety of applications. This software is open-source and available as part of the MINE Database python package ( https://pypi.org/project/minedatabase/ ) or on GitHub ( https://github.com/tyo-nu/MINE-Database ). Documentation and examples can be found on Read the Docs ( https://mine-database.readthedocs.io/en/latest/ ). Through its documentation, pre-packaged features, and customizable nature, Pickaxe allows users to generate novel reaction networks tailored to their application.

59 BASIC BIOLOGICAL SCIENCES↗

YeastWGD2025

Supplementary data for Discovery of additional ancient genome duplications in yeasts wgd_syn / - directory containing wgd syn output for all contiguous genomes [dataset] Tree - phylogeny [dataset]Duplications - duplication table from OrthoFinder output KOannotations - KEGG annotations used for enrichment analysis IPRannotations - InterPro annotations used for enrichment analysis DipodascalesOrthogroups - formatted orthogroup assignments for Dipodascales genes.fa and .gff3 files for each new genome assembly are also provided, those these are not required to replicate the analysis

Genomics↗

EI_MS_ML

The unambiguous identification of compounds from their electron ionization mass (EI-MS) spectra remains a significant unsolved problem in the field of metabolomics and analytical chemistry as a whole. Typically EI-MS spectra are compared using various mathematical operations that convert the spectral similarity or differences into a distance-like metric that roughly approximates the similarity of any two spectra. A commonly used metric for this is the cosine similarity metric which has values close to one for very similar spectra and a value of zero for very dissimilar spectra; however, no metric is perfect. Due to the prevalence of structurally-similar compounds such as isomers and the prevalence of certain fragmentation patterns across structurally-dissimilar compounds, the unambiguous assignment of EI-MS spectra compounds remains difficult. Frequently, querying an observed EI-MS spectrum against a large database such as the NIST17 library yields multiple possible assignments requiring the end user to distinguish between multiple high scoring hits, or multiple low scoring hits while keeping in mind that the correct hit may not be in the database at all. Although techniques such as orthogonal information from techniques such as chromatography can greatly aid in unambiguous assignment, this also requires more complicated experimental designs and access to more complicated analytical instrumentation. Substructures can be trivially detected and represented as strings using a previously published technique called node coloring from a known chemical structure. However, for experimentally-derived EI-MS spectra this information must be derived from the spectra itself (i.e., because we do not know what compound it represents). To achieve this, the software uses techniques from the field of machine learning and a large training dataset of EI-MS spectra corresponding to known structures annotated with substructure strings, to build models that can predict the presence of a given chemical substructure from an EI-MS spectrum directly.If these predictions are of high-quality (i.e., are unlikely to be false positives), the presence of one or more predicted substructures can be used to constrain the number of possible hits for a query spectrum. Mathematically, this restriction could be expressed in many forms, but the most straight-forward implementation is to weight the cosine similarity of a query spectrum and a plausible database match with a Tanimoto-like coefficient based on the ratio of the number of substructures predicted to the number of substructures present in the potential database hit. Determining which combination of models best reduces assignment ambiguity will be achieved using a combination of manual curation and optimization techniques such as genetic algorithms. This software will perform all the steps necessary to construct said models from a training dataset and evaluate them using a holdout dataset. Various statistical analyses can be performed to determine if this approach does decrease assignment ambiguity. For example, if this approach works, on average, the rank-order of the correct assignment for the holdout set of EI-MS spectra should decrease and the weighted cosine similarities for most of the possible matches in the database should be better than the unweighted cosine similarities. Furthermore, this same pipeline can be used on real experimental data to generate less ambiguous assignments.

Mitchell, Joshua↗

efam: an e xpanded, metaproteome-supported HMM profile database of viral protein fam ilies

Viruses infect, reprogram and kill microbes, leading to profound ecosystem consequences, from elemental cycling in oceans and soils to microbiome-modulated diseases in plants and animals. Although metagenomic datasets are increasingly available, identifying viruses in them is challenging due to poor representation and annotation of viral sequences in databases. Here, we establish efam, an expanded collection of Hidden Markov Model (HMM) profiles that represent viral protein families conservatively identified from the Global Ocean Virome 2.0 dataset. This resulted in 240 311 HMM profiles, each with at least 2 protein sequences, making efam >7-fold larger than the next largest, pan-ecosystem viral HMM profile database. Adjusting the criteria for viral contig confidence from ‘conservative’ to ‘eXtremely Conservative’ resulted in 37 841 HMM profiles in our efam-XC database. To assess the value of this resource, we integrated efam-XC into VirSorter viral discovery software to discover viruses from less-studied, ecologically distinct oxygen minimum zone (OMZ) marine habitats. This expanded database led to an increase in viruses recovered from every tested OMZ virome by ~24% on average (up to ~42%) and especially improved the recovery of often-missed shorter contigs (<5 kb). Additionally, to help elucidate lesser-known viral protein functions, we annotated the profiles using multiple databases from the DRAM pipeline and virion-associated metaproteomic data, which doubled the number of annotations obtainable by standard, single-database annotation approaches. Together, these marine resources (efam and efam-XC) are provided as searchable, compressed HMM databases that will be updated bi-annually to help maximize viral sequence discovery and study from any ecosystem.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A clinically and genomically annotated nerve sheath tumor biospecimen repository

Nerve sheath tumors occur as a heterogeneous group of neoplasms in patients with neurofibromatosis type 1 (NF1). The malignant form represents the most common cause of death in people with NF1, and even when benign, these tumors can result in significant disfigurement, neurologic dysfunction, and a range of profound symptoms. Lack of human tissue across the peripheral nerve tumors common in NF1 has been a major limitation in the development of new therapies. To address this unmet need, we have created an annotated collection of patient tumor samples, patient-derived cell lines, and patient-derived xenografts, and carried out high-throughput genomic and transcriptomic characterization to serve as a resource for further biologic and preclinical therapeutic studies. In this work, we release genomic and transcriptomic datasets comprised of 55 tumor samples derived from 23 individuals, complete with clinical annotation. All data are publicly available through the NF Data Portal and at http://synapse.org/jhubiobank.

60 APPLIED LIFE SCIENCES↗