Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “annotations”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

A Publicly Available, Annotated Dataset for Naturalistic Driving Study and Computer Vision Algorithm Development

Oak Ridge National Laboratory developed and implemented a data collection effort to create a dataset for use in evaluating and testing algorithms for analyzing driver behavior under controlled settings for support of the Federal Highway Administration’s Exploratory Advanced Research Program. This collection is called the ORNL Naturalistic Driving Study Sample (ONDSS). The dataset is designed to emulate aspects of the Second Strategic Highway Research Project (SHRP2), which contained a massive naturalistic driving study (NDS) with over 3000 drivers between 2010 and 2013 using their personal vehicles, with over 4300 person-years of data collected [HANKEY].

42 ENGINEERING↗

U.S. Nuclear Declaratory Policy 2021: the Renewed Debate about Sole Purpose and No-First-Use. Annotated Bibliography

Declaratory policy and public statements about the potential use of nuclear weapons serve many important roles. They provide an assessment of the security environment, and inform the public debate. These statements also enhance deterrence messages and signals towards adversaries, and reassure allies and partners. On the global level, U.S. declaratory policy has the potential to shape international trends and norms, influence nuclear proliferation, and it may also affect the policy decisions of other nuclear possessors. As the Biden administration reviews the elements of U.S. nuclear declaratory policy, the issue of sole purpose and no-first-use is likely to resurface. Previous administrations have examined these policies in multiple rounds of review, and they decided that the time was not right for such declarations. This literature review was prepared to inform the debate by collecting some of the most prominent articles on the topic that highlight the potential risks and benefits of these policies.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

China and Multi-Domain Strategic Stability (Annotated Bibliography)

Brief summaries of literature relating to four areas: China's Approach to Multi-Domain Complexity, China's Approach to Multi-Domain Strategic Stability, A Cooperative Management Approach, and Risk Mitigation in the Absence of Cooperative Approaches.

97 MATHEMATICS AND COMPUTING↗

Multi-Species Complex and Standard Metabolomic Samples with Verified Truth Annotations Dataset

This dataset contains 4523251 (~6.35 GB) metabolite-spectra matches following identification with CoreMS. Data were manually curated as true positives, true negatives, or unknowns. Calculations for spectral similarity scores were carried out with two methods for a total of ~12.7 GB (6.35 * 2) of data. They are all .tsv files, though can easily be changed to .txt. The file types are: * human cerebrospinal fluid (CSF), human blood plasma human urine: already published here https://www.nature.com/articles/s41597-021-00894-y, • purchased FAMES standards • fungi species (A. niger, A. nidulans, T. reesei) • soil crust

59 BASIC BIOLOGICAL SCIENCES↗

6051R & 6051S Assembly and Annotation

We report the draft genomes of two morphologically distinct variants of Bacillus subtilis ATCC 6051 [NCBI3610]. The two isolates exhibit differences in not only morphology but also their genetics, despite identical 16S rRNA sequences. Investigating the genetic differences of colony morphology variation in this model organism can provide valuable insights.

59 BASIC BIOLOGICAL SCIENCES↗

Original images, ground truth annotations, and precited masks from U-NET for a fungal species x nitrogen experiment

This document describes the “MicroVision-MV003” dataset. This folder contains three subfolders: “originals”, “predicted_masks”, and “gt_masks”. The “originals” folder contains 720 original images acquired using two imaging modalities: 360 images from overhead imaging and 360 images from transmission (raw) imaging. Naming scheme:YYYY-MM-dd__plate_strain_nitrogen-level_replication_experiment_imaging-modalityYYYY-MM-dd: 2024-12-20, 2024-12-21, 2024-12-22,2024-12-23,2024-12-24,2024-12-25,2024-12-26,2024-12-27,2024-12-28,2024-12-29,2024-12-30,2024-12-31.plate: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30.strain: FG, LE.nitrogen-level: N-1, N-10, N-100.replication: 1, 2, 3, 4, 5.experiment: MV.003.imaging-modality: overhead, transmission (raw). For example, the image2024-12-20__1_FG_N-10_3_MV.003_overheadwas taken on 2024-12-20. The plate number is 1, the fungal strain is FG, the nitrogen treatment is N-10, and the replication number is 3 for experiment MV.003, acquired using the overhead imaging modality. The “predicted_masks” folder contains masks generated by a trained U-Net model for the transmission imaging modality.The “gt_masks” folder contains 326 human hand-traced masks that serve as ground truth.

09 BIOMASS FUELS↗

Metabolite, annotation, and gene integration system and method

Disclosed herein are systems and methods for associating metabolites with genes. In some embodiments, after potential metabolites are identified based on spectroscopy data of the content of an organism, possible reactions capable of producing the potential metabolites are determined. The possible reactions are compared to gene sequences in a database, and an association score for the likelihood that a gene sequence is related to the potential metabolites is calculated.

Erbilgin, Onur↗

Finding the missing pieces: filling gaps that impede the translation of omics data into models

High-throughput omics technologies such as DNA sequencing have made the sequencing and computational assembly of microbial genomes recovered from the environment relatively routine. Computational inference of the protein products encoded by these genomes, and the associated biochemical functions, should enable the accurate prediction and modeling of microbial metabolism, organismal interactions, and ecosystem processes. However, a lack of scalable, probabilistic protein annotation tools limits the full potential of modeling for understanding the metabolism and biogeochemical cycles of microbial communities. Our approach to improve inference of protein annotations and metabolic models relied on learning from and emulating expert manual curation, leveraging software engineering and data science best practices to scale up the throughput and accuracy of annotations and metabolic model construction, building software to objectively evaluate different annotation strategies, and more closely linking the protein annotation and metabolic model inference process. Outcomes of this research include several improved or new computational tools, including DRAM (Distilled and Refined Annotation of Metabolism) for annotating microbial genomes with protein function and metabolic traits, CAMPER (Curated Annotations for Microbial Polyphenol Enzymes and Reactions) for annotating key polyphenol metabolisms, EC-Bench for comprehensive and unbiased benchmarking of annotation tools, and several apps available via the DOE Systems Biology Knowledgebase (KBase) for building genome-scale metabolic models. We demonstrate that these tools allow us to scalably annotate and understand thousands of genomes for microbial communities from a variety of systems and test cases, including rivers, thawing permafrost, and gut microbiomes. All of these computational tools are available as open-source software, with most broadly and easily accessible to the scientific community via KBase apps.

59 BASIC BIOLOGICAL SCIENCES↗

Learning from Crowds by Modeling Common Confusions

Crowdsourcing provides a practical way to obtain large amounts of labeled data at a low cost. However, the annotation quality of annotators varies considerably, which imposes new challenges in learning a high-quality model from the crowdsourced annotations. In this work, we provide a new perspective to decompose annotation noise into common noise and individual noise and differentiate the source of confusion based on instance difficulty and annotator expertise on a per-instance-annotator basis. We realize this new crowdsourcing model by an end-to-end learning solution with two types of noise adaptation layers: one is shared across annotators to capture their commonly shared confusions, and the other one is pertaining to each annotator to realize individual confusion. To recognize the source of noise in each annotation, we use an auxiliary network to choose from the two noise adaptation layers with respect to both instances and annotators. Extensive experiments on both synthesized and real-world benchmarks demonstrate the effectiveness of our proposed common noise adaptation solution.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

The Gene Ontology resource: enriching a GOld mine

The Gene Ontology Consortium (GOC) provides the most comprehensive resource currently available for computable knowledge regarding the functions of genes and gene products. Here, we report the advances of the consortium over the past two years. The new GO-CAM annotation framework was notably improved, and we formalized the model with a computational schema to check and validate the rapidly increasing repository of 2838 GO-CAMs. In addition, we describe the impacts of several collaborations to refine GO and report a 10% increase in the number of GO annotations, a 25% increase in annotated gene products, and over 9,400 new scientific articles annotated. As the project matures, we continue our efforts to review older annotations in light of newer findings, and, to maintain consistency with other ontologies. As a result, 20,000 annotations derived from experimental data were reviewed, corresponding to 2.5% of experimental GO annotations. The website (http://geneontology.org) was redesigned for quick access to documentation, downloads and tools. To maintain an accurate resource and support traceability and reproducibility, we have made available a historical archive covering the past 15 years of GO data with a consistent format and file structure for both the ontology and annotations.

59 BASIC BIOLOGICAL SCIENCES↗