Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Dataset Annotations”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Integration of Simulated and Real Distributed Acoustic Sensor Measurements to Develop AI Enhanced Intelligent Sensing Operations

Distributed acoustic sensors (DAS) have shown remarkable success in monitoring critical civil, energy and transportation assets over long distances and under harsh environmental conditions. Integration of artificial intelligence (AI) technologies with DAS can extend their operational capabilities from conventional tasks (e.g., vibration frequency detection) to more advanced intelligent tasks like anomaly detection. However, these AI technologies often require large labeled/ annotated DAS measurements datasets of the assets in both normal and anomalous operating conditions. This data acquisition task can be both time consuming and costly. This work develops a generative adversarial network based unsupervised domain adaptation framework to build intelligent DAS operational capabilities.

Venketeswaran, Abhishek↗

Expanded genetic variation (SNP) and phenomics (image based) dataset for Populus trichocarpa

The image dataset consists of 11,791 images representing 1,219 genotypes of Populus trichocarpa undergoing in planta regeneration. Genotypes were imaged with a median of four weekly timepoints and a median of two replicates each. A representative and diverse subset of 249 images was annotated using the IDEAS annotation interface (ideas.eecs.oregonstate.edu) and these annotated images were used to train a deep semantic segmentation model (PSPNet), which was deployed for inference over the entire dataset. Annotated classes include specific stages of regeneration (callus and shoot) in addition to unregenerated plant material and background. Statistics of relative tissue area were extracted and used for downstream genetic association mapping in a genome-wide association study. The SNP dataset consists of over 40 million single-nucleotide polymorphisms across 1,323 wild accessions of Populus trichocarpa

09 BIOMASS FUELS↗

Data for "Genetics of flooding tolerance in an F2 Miscanthus sacchariflorus ssp. lutarioriparius × M. sinensis population"

This dataset contains all data and supplementary materials from "Genetics of flooding tolerance in an F2 Miscanthus sacchariflorus ssp. lutarioriparius × M. sinensis population". 1. The dataset S1 table contains the raw phenotypic data collected during the experiment. 2. The dataset S2 table contains the LSmean values for the 24 traits studied. 3. The dataset S3 table contains the TASSEL GBSv2 map, marker information, and genotype data used for mapping. 4. The dataset S4 table contains information on candidate genes found in each of the QTL intervals. 5. The dataset S5 table contains the GO annotations and KEGG enrichment analyses for those candidate genes. 6. The dataset S6 table contains information on the sequences used to classify AP2 ERF transcription factors. 7. The dataset S7 table contains information on AP2 ERF orthologs between Miscanthus and rice based on synteny. 8. Supplementary file 1 contains the ANOVA results using the raw phenotypic data collected from protocol "A". 9. Supplementary file 2 contains the ANOVA results using the raw phenotypic data collected from protocol "B". 10. Supplementary file 3 contains notes on the comparison of SNP calling methods. 11. Supplementary file 4 is a script for analyzing candidate genes found in QTL intervals.

Miscanthus, flood, partial submergence, complete s↗

PyCMG-based Simulation of Volumetric Concrete Microstructure

Concrete is a complex, heterogeneous material with a microstructure composed of aggregates, cement paste, and pores spanning multiple length scales. Understanding this microstructure is critical for advancing the performance, durability, and modeling of concrete-based systems. While experimental imaging such as X-ray computed tomography (XCT) provides valuable insights, generating large datasets with detailed ground truth annotations is both costly and labor-intensive due to challenges in segmenting similar phases, such as aggregates and cement paste, that often share similar attenuation properties. To address this, we developed a pipeline to simulate realistic 3D concrete microstructures using the open-source Python package PyCMG. This simulation effort focuses on generating high-fidelity, annotated microstructures that can serve as training or benchmarking datasets for image analysis, segmentation algorithms, and machine learning models, particularly in scenarios where experimental data is scarce.

Ziabari, Amir [Oak Ridge National Laboratory; ORNL↗

GOLEM: GOld standard for Learning and Evaluation of Motifs

Motifs are distinctive, recurring, widely used idiom-like words or phrases, often originating from folklore, whose meaning is anchored in a narrative and have a significance as communicative devices across a wide range of media, including news, literature, and propaganda. Many motifs concisely imply a large constellation of culturally relevant information, and their broad usage suggests their cognitive importance as touchstones of cultural knowledge. As such, their detection is a step towards culturally aware natural language processing. We present GOLEM (GOld standard for Learning and Evaluation of Motifs) a dataset of English news articles, opinion pieces, and broadcast transcripts annotated for motific information. The dataset identifies 25,737 motif candidates across 34 motif types drawn from three cultural or national groups: Jewish, Irish, and Puerto Rican. The dataset contains 2,024,141 words split into 25,737 text snippets drawn from 8,073 articles. Each motif candidate is labeled according to a scheme which identifies the type of usage (motific, referential, eponymic, or unrelated), resulting in 1,743 actual motific instances in the data. Annotation was performed by individuals identifying as members of each group and achieved a Fleiss’ kappa (?) of > 0.55. In addition to the data, we demonstrate that classification of the candidate type is a challenging task for Large Language Models (LLMs) using a few-shot approach; recent models such as T5, FLAN-T5, GPT-2, and Llama 2 (7B) achieved a performance of 41% accuracy at best, where the majority class accuracy is 41% and the average chance accuracy is 27%. These data will support development of new models and approaches for detecting (and reasoning about) motific information in text.

motif, culture, natural language, artificial intel↗

AMVOS: Additive Manufacturing Video Object Segmentation Dataset

This dataset provides labeled video frames from four additive manufacturing (AM) processes for video object segmentation (VOS) tasks. It contains 90 video segments comprising 900 individually annotated frames across five AM datasets: laser hot-wire directed energy deposition (LHW-DED), tungsten inert gas wire arc additive manufacturing (TIG-WAAM), plasma arc welding (PAW), visible-light polymer extrusion (visPolymer), and near-infrared polymer extrusion (irPolymer). Each video segment consists of 10 contiguous frames with corresponding pixel-level object instance annotations. Depending on the process, two of four object classes are labeled per frame: Melt Pool, Feed Wire, Nozzle, or Material. Raw frames are provided as .jpg files and annotations as palettized .png files. The dataset follows the directory structure of established VOS benchmarks (DAVIS, YouTube-VOS, MOSE), enabling direct integration into VOS model training and evaluation pipelines for foundation model fine-tuning, domain adaptation, or zero-shot performance benchmarking. Data was collected at Oak Ridge National Laboratory's Manufacturing Demonstration Facility.

Wetzel, Jon [ORNL]↗

A Step-by-Step Protocol from METASPACE to Biological Interpretation

Mass spectrometry imaging (MSI) represents an exceptional tool for exploring complex biological systems spatially at the molecular level. However, due to its multidimensional nature and large-scale data output, it presents considerable challenges when it comes to extracting meaningful biological insights. Recent advancements, such as the METASPACE platform, have enabled researchers to efficiently process, annotate, and interpret MSI datasets by leveraging machine learning and cloud-based infrastructure. In this tutorial, we present a detailed and user-friendly R-pipeline designed to help METASPACE users navigate untargeted metabolomic annotations and transform them into practical insights about their biological systems. By combining METASPACE annotations with rapid R-based screening, this workflow not only streamlined the analytical process but also enhanced the understanding of spatial molecular distribution, especially for complex systems. Here, this easy-to-follow approach has the potential for applications in diagnostics, drug discovery, environmental and ecological processes, and more. We envision this pipeline to be particularly useful for newcomers to the field of MSI and

Moreno Pedraza, Abigail↗

Missing microbial eukaryotes and misleading meta-omic conclusions

Meta-omics is commonly used for large-scale analyses of microbial eukaryotes, including species or taxonomic group distribution mapping, gene catalog construction, and inference on the functional roles and activities of microbial eukaryotes in situ. Here, we explore the potential pitfalls of common approaches to taxonomic annotation of protistan meta-omic datasets. We re-analyze three environmental datasets at three levels of taxonomic hierarchy in order to illustrate the crucial importance of database completeness and curation in enabling accurate environmental interpretation. We show that taxonomic membership of sequence clusters estimates community composition more accurately than returning exact sequence labels, and overlap between clusters can address database shortcomings. Clustering approaches can be applied to diverse environments while continuing to exploit the wealth of annotation data collated in databases, and selecting and evaluating these databases is a critical part of correctly annotating protistan taxonomy in environmental datasets. We argue that ongoing curation of genetic resources is crucial in accurately annotating protists in in situ meta-omic datasets. Moreover, we propose that precise taxonomic annotation of meta-omic data is a clustering problem rather than a feasible alignment problem.

59 BASIC BIOLOGICAL SCIENCES↗

The Chlamydomonas Genome Project, version 6: reference assemblies for mating type plus and minus strains reveal extensive structural mutation in the laboratory

Five versions of the Chlamydomonas reinhardtii reference genome have been produced over the last two decades. Here we present version 6, bringing significant advances in assembly quality and structural annotations. PacBio-based chromosome-level assemblies for two laboratory strains, CC-503 and CC-4532, provide resources for the plus and minus mating type alleles. We corrected major misassemblies in previous versions and validated our assemblies via linkage analyses. Contiguity increased over ten-fold and >80% of filled gaps are within genes. We used Iso-Seq and deep RNA-seq datasets to improve structural annotations, and updated gene symbols and textual annotation of functionally characterized genes via extensive manual curation. We discovered that the cell wall-less classical reference strain CC-503 exhibits genomic instability potentially caused by deletion of the helicase RECQ3, with major structural mutations identified that affect >100 genes. We therefore present the CC-4532 assembly as the primary reference, although this strain also carries unique structural mutations and is experiencing rapid proliferation of a Gypsy retrotransposon. We expect all laboratory strains to harbor gene-disrupting mutations, which should be considered when interpreting and comparing experimental results. Collectively, the resources presented here herald a new era of Chlamydomonas genomics and will provide the foundation for continued research in this important reference organism.

59 BASIC BIOLOGICAL SCIENCES↗

From Text to Maps: LLM-Driven Extraction and Geotagging of Epidemiological Data

Epidemiological datasets are essential for public health analysis and decision-making, yet they remain scarce and often difficult to compile due to inconsistent data formats, language barriers, and evolving political boundaries. Traditional methods of creating such datasets involve extensive manual effort and are prone to errors in accurate location extraction. To address these challenges, we propose utilizing large language models (LLMs) to automate the extraction and geotagging of epidemiological data from textual documents. Our approach significantly reduces the manual effort required, limiting human intervention to validating a subset of records against text snippets and verifying the geotagging reasoning, as opposed to reviewing multiple entire documents manually to extract, clean, and geotag. Additionally, the LLMs identify information often overlooked by human annotators, further enhancing the dataset’s completeness. Our findings demonstrate that LLMs can be effectively used to semi-automate the extraction and geotagging of epidemiological data, offering several key advantages: (1) comprehensive information extraction with minimal risk of missing critical details; (2) minimal human intervention; (3) higher-resolution data with more precise geotagging; and (4) significantly reduced resource demands compared to traditional methods.

Harrod, Karly↗

Layer-wise Imaging Dataset from Powder Bed Additive Manufacturing Processes for Machine Learning Applications (Peregrine v2021-03)

This dataset contains layer-wise powder bed images from three different powder bed printing technologies – laser powder bed fusion, electron beam powder bed fusion, and binder jetting. This dataset was collected and annotated using the internally-developed Peregrine software tool and is designed primarily to facilitate research into anomaly defect detection using image segmentation or similar techniques. A total of 20 layers are provided for each printing technology, with each layer of data consisting of one or more calibrated images and an annotation file containing pixel-wise ground truth labels. The ground truths were labeled by domain experts, typically printer technicians. Data in this release were collected at Oak Ridge National Laboratory between 2016 and 2020 and were compiled in March 2021.

36 MATERIALS SCIENCE↗

Text Mining for Process–Structure–Properties Relationships in Metals

With the advent of large language models (LLMs), the vast unstructured text within millions of academic papers is increasingly accessible for materials discovery—although significant challenges remain. While LLMs offer promising few- and zero-shot learning capabilities, particularly valuable in the materials domain where expert annotations are scarce, general-purpose LLMs often fail to address key materials-specific queries without further adaptation. To bridge this gap, fine-tuning LLMs on human-labeled data is essential for effective structured knowledge extraction (Liu in The Importance of Human-Labeled Data in the Era of LLMs, 2023). Here, in this study, we introduce a novel annotation schema designed to extract generic process–structure–properties relationships from scientific literature. We demonstrate the utility of this approach using a dataset of 128 abstracts, with annotations drawn from two distinct domains: high-temperature materials (Domain I) and uncertainty quantification in simulating materials microstructure (Domain II). Initially, we developed a conditional random field (CRF) model based on MatBERT—a domain-specific BERT variant—and evaluated its performance on Domain I. Subsequently, we compared this model with a fine-tuned LLM (GPT-4o from OpenAI) under identical conditions. Our results indicate that fine-tuning LLMs can significantly improve entity extraction performance over the BERT-CRF baseline on Domain I. However, when additional examples from Domain II were incorporated, the performance of the BERT-CRF model became comparable to that of the GPT-4o model. These findings underscore the potential of our schema for structured knowledge extraction and highlight the complementary strengths of both modeling approaches.

Materials science↗

Hpc Natural Language Understanding (nlu) Dataset

This code provides natural language annotation and labels for multiple activities in high performance. The purpose is to enable machine learning for natural language processing algorithms with high performance computing systems.

Biggs, BrandonS↗

Structural models and functional annotations for the Sphagnum divinum proteome

This dataset contains the structural models for the primary transcripts of the Sphagnum divinum proteome. Additionally, for a subset of these proteins, sequence and structural alignment results are provided. This dataset represents the most thorough structural study of a Sphagnum species, also known as peat mosses, by providing three-dimensional atomic resolution structures of the majority of the encoded proteins as well as structural alignment results used in the application of annotating the proteome. References (DOI) AlphaFold v2 Monomer: https://doi.org/10.1038/s41586-021-03819-2. References (DOI) US-align2: https://doi.org/10.1038/s41592-022-01585-1

59 BASIC BIOLOGICAL SCIENCES↗

Structure-aware annotation of leucine-rich repeat domains

Protein domain annotation is typically done by predictive models such as HMMs trained on sequence motifs. However, sequence-based annotation methods are prone to error, particularly in calling domain boundaries and motifs within them. These methods are limited by a lack of structural information accessible to the model. With the advent of deep learning-based protein structure prediction, existing sequenced-based domain annotation methods can be improved by taking into account the geometry of protein structures. We develop dimensionality reduction methods to annotate repeat units of the Leucine Rich Repeat solenoid domain. The methods are able to correct mistakes made by existing machine learning-based annotation tools and enable the automated detection of hairpin loops and structural anomalies in the solenoid. The methods are applied to 127 predicted structures of LRR-containing intracellular innate immune proteins in the model plant Arabidopsis thaliana and validated against a benchmark dataset of 172 manually-annotated LRR domains.

Xu, Boyan↗

EXSCLAIM!: Harnessing materials science literature for self-labeled microscopy datasets

This work introduces the EXSCLAIM! toolkit for the automatic extraction, separation, and caption-based natural language annotation of images from scientific literature. EXSCLAIM! is used to show how rule-based natural language processing and image recognition can be leveraged to construct an electron microscopy data set containing thousands of keyword-annotated nanostructure images. Moreover, it is demonstrated how a combination of statistical topic modeling and semantic word similarity comparisons can be used to increase the number and variety of keyword annotations on top of the standard annotations from EXSCLAIM! With large-scale imaging datasets constructed from scientific literature, users are well positioned to train neural networks for classification and recognition tasks specific to microscopy-tasks often otherwise inhibited by a lack of sufficient annotated training data.

36 MATERIALS SCIENCE↗

Comprehensive SNP Data for 1,323 GWAS Population in Populus trichocarpa and Combined Annotation Files for P. trichocarpa v3.0 and v3.1

The VCF dataset includes genetic variations found in 1,323 Populus trichocarpa genotypes, providing valuable information for scientists studying plant genetics. Researchers have generated this dataset using whole-genome DNA short-read sequencing on the Illumina Genome Analyzer, HiSeq 2000, and HiSeq 2500 platforms. This sequencing effort ensured a minimum expected sequencing depth of 15×. The dataset comprises more than 9.7 million single nucleotide polymorphisms (SNPs) and indel variants. The combined annotation files are derived from P. trichocarpa v3.0 and v3.1. We merged these files to create a comprehensive annotation file used for GWAS analysis. In total, 38,830 genes overlapped between the two versions. For overlapping genes, we defined the start as the smaller and the end as the larger among the two versions to increase the likelihood of locating candidate genetic loci. Additionally, we included 2,505 unique genes from v3.0 and 4,120 unique genes from v3.1, resulting in a total of 45,455 genes in the updated annotation file.

09 BIOMASS FUELS↗

Phenotypically anchored transcriptomics across diverse agrichemicals reveals conserved pathways and unique gene expression signatures in zebrafish

Agrichemicals such as herbicides, fungicides, insecticides, and biocides are widely used in agriculture, yet some are associated with adverse effects in humans and the environment. While many of these chemicals have been extensively studied in vitro and are included in the EPA’s ToxCast program, comprehensive in vivo comparisons using RNA sequencing across structurally diverse agrichemicals, in a single screening platform, are lacking. In this study, we examined structurally diverse agrichemicals found in the U.S. Environmental Protection Agency’s (EPA) Toxcast Phase I and II library by statically exposing early life stage zebrafish at 6 h post fertilization (hpf) until 120 hpf at concentrations ranging from 0.25 to 100 µM. Morphological outcomes were assessed at 120 hpf across 10 endpoints, including yolk sac edema, craniofacial malformations, and axis abnormalities. Chemicals that produced robust concentration-response relationships were selected for transcriptomic profiling. For transcriptomic analysis, zebrafish were statically exposed to each chemical and sampled at 48 hpf, prior to the onset of morphological effects observed at 120 hpf. Differential expression analysis identified between 0 and 4,538 differentially expressed genes (DEGs) per chemical, with no clear correlation to morphological severity. Both DEG and co-expression network analyses revealed chemical-specific expression patterns that converged on shared biological pathways, including neurodevelopment and cytoskeletal organization. Key regulatory genes such as mylpfa and krt4 were identified within co-expression modules, suggesting their potential role in conserved toxicity mechanisms. Semantic similarity analysis of enriched gene ontology (GO) terms, when compared to existing datasets, highlighted gaps in the annotation of neurodevelopmental processes, indicating that some in vivo effects may not be fully captured by current curated resources. The results provide new insights into the modes of action of diverse agrichemicals and establish a framework for understanding how agrichemical structure relates to biological function in a vertebrate model.

agrichemical↗