Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “annotations”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

The value of human data annotation for machine learning based anomaly detection in environmental systems

Anomaly detection is the process of identifying unexpected data samples in datasets. Automated anomaly detection is either performed using supervised machine learning models, which require a labelled dataset for their calibration, or unsupervised models, which do not require labels. While academic research has produced a vast array of tools and machine learning models for automated anomaly detection, the research community focused on environmental systems still lacks a comparative analysis that is simultaneously comprehensive, objective, and systematic. This knowledge gap is addressed for the first time in this study, where 15 different supervised and unsupervised anomaly detection models are evaluated on 5 different environmental datasets from engineered and natural aquatic systems. To this end, anomaly detection performance, labelling efforts, as well as the impact of model and algorithm tuning are taken into account. As a result, our analysis reveals the relative strengths and weaknesses of the different approaches in an objective manner without bias for any particular paradigm in machine learning. Most importantly, our results show that expert-based data annotation is extremely valuable for anomaly detection based on machine learning.

54 ENVIRONMENTAL SCIENCES↗

Extracting Material Property Measurements from Scientific Literature with Limited Annotations

Extracting material property data from scientific text is pivotal for advancing data-driven research in chemistry and materials science; however, the extensive annotation effort required to produce training data for named entity recognition (NER) models for this task often makes it a barrier to extracting specialized data sets. Here, in this work, we present a comparative study of the conventional, supervised NER methodology to alternative few-shot learning architectures and large language model (LLM)-based approaches that mitigate the need to label large training data sets. We find that the best-performing LLM (GPT-4o) not only excels in directly extracting relevant material properties based on limited examples but also enhances supervised learning through data augmentation. We supplement our findings with error and data quality assessments to provide a nuanced understanding of factors that impact property measurement extraction.

36 MATERIALS SCIENCE↗

PeakDecoder enables machine learning-based metabolite annotation and accurate profiling in multidimensional mass spectrometry measurements

Multidimensional measurements using state-of-the-art separations and mass spectrometry provide advantages in untargeted metabolomics analyses for studying biological and environmental bio-chemical processes. However, the lack of rapid analytical methods and robust algorithms for these heterogeneous data has limited its application. Here, we develop and evaluate a sensitive and high-throughput analytical and computational workflow to enable accurate metabolite profiling. Our workflow combines liquid chromatography, ion mobility spectrometry and data-independent acquisition mass spectrometry with PeakDecoder, a machine learning-based algorithm that learns to distinguish true co-elution and co-mobility from raw data and calculates metabolite identification error rates. We apply PeakDecoder for metabolite profiling of various engineered strains of Aspergillus pseudoterreus, Aspergillus niger, Pseudomonas putida and Rhodosporidium toruloides. Results, validated manually and against selected reaction monitoring and gas-chromatography platforms, show that 2683 features could be confidently annotated and quantified across 116 microbial sample runs using a library built from 64 standards.

59 BASIC BIOLOGICAL SCIENCES↗

ContScout: sensitive detection and removal of contamination from annotated genomes

Contamination of genomes is an increasingly recognized problem affecting several downstream applications, from comparative evolutionary genomics to metagenomics. Here we introduce ContScout, a precise tool for eliminating foreign sequences from annotated genomes. It achieves high specificity and sensitivity on synthetic benchmark data even when the contaminant is a closely related species, outperforms competing tools, and can distinguish horizontal gene transfer from contamination. A screen of 844 eukaryotic genomes for contamination identified bacteria as the most common source, followed by fungi and plants. Furthermore, we show that contaminants in ancestral genome reconstructions lead to erroneous early origins of genes and inflate gene loss rates, leading to a false notion of complex ancestral genomes. Taken together, we offer here a tool for sensitive removal of foreign proteins, identify and remove contaminants from diverse eukaryotic genomes and evaluate their impact on phylogenomic analyses.

59 BASIC BIOLOGICAL SCIENCES↗

Enhancing tandem mass spectrometry-based metabolite annotation with online chemical labeling

Abstract Metabolite identification in non-targeted mass spectrometry-based metabolomics remains a major challenge due to limited spectral library coverage and difficulties in predicting metabolite fragmentation patterns. Here, we introduce Multiplexed Chemical Metabolomics (MCheM), which employs orthogonal post-column derivatization reactions integrated into a unified mass spectrometry data framework. MCheM generates orthogonal structural information that substantially improves metabolite annotation through in silico spectrum matching and open-modification searches, offering a powerful new toolbox for the structure elucidation of unknown metabolites at scale.

Science & Technology - Other Topics↗

Automated annotation of scientific texts for ML-based keyphrase extraction and validation

Advanced omics technologies and facilities generate a wealth of valuable data daily; however, the data often lack the essential metadata required for researchers to find, curate, and search them effectively. The lack of metadata poses a significant challenge in the utilization of these data sets. Machine learning (ML)–based metadata extraction techniques have emerged as a potentially viable approach to automatically annotating scientific data sets with the metadata necessary for enabling effective search. Text labeling, usually performed manually, plays a crucial role in validating machine-extracted metadata. However, manual labeling is time-consuming and not always feasible; thus, there is a need to develop automated text labeling techniques in order to accelerate the process of scientific innovation. This need is particularly urgent in fields such as environmental genomics and microbiome science, which have historically received less attention in terms of metadata curation and creation of gold-standard text mining data sets. In this paper, we present two novel automated text labeling approaches for the validation of ML-generated metadata for unlabeled texts, with specific applications in environmental genomics. Our techniques show the potential of two new ways to leverage existing information that is only available for select documents within a corpus to validate ML models, which can then be used to describe the remaining documents in the corpus. The first technique exploits relationships between different types of data sources related to the same research study, such as publications and proposals. The second technique takes advantage of domain-specific controlled vocabularies or ontologies. In this paper, we detail applying these approaches in the context of environmental genomics research for ML-generated metadata validation. Our results show that the proposed label assignment approaches can generate both generic and highly specific text labels for the unlabeled texts, with up to 44% of the labels matching with those suggested by a ML keyword extraction algorithm.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Chromosome assembled and annotated genome sequence of Aspergillus flavus NRRL 3357

Abstract Aspergillus flavus is an opportunistic pathogen of crops, including peanuts and maize, and is the second leading cause of aspergillosis in immunocompromised patients. A. flavus is also a major producer of the mycotoxin, aflatoxin, a potent carcinogen, which results in significant crop losses annually. The A. flavus isolate NRRL 3357 was originally isolated from peanut and has been used as a model organism for understanding the regulation and production of secondary metabolites, such as aflatoxin. A draft genome of NRRL 3357 was previously constructed, enabling the development of molecular tools and for understanding population biology of this particular species. Here, we describe an updated, near complete, telomere-to-telomere assembly and re-annotation of the eight chromosomes of A. flavus NRRL 3357 genome, accomplished via long-read PacBio and Oxford Nanopore technologies combined with Illumina short-read sequencing. A total of 13,715 protein-coding genes were predicted. Using RNA-seq data, a significant improvement was achieved in predicted 5’ and 3’ untranslated regions, which were incorporated into the new gene models.

59 BASIC BIOLOGICAL SCIENCES↗

IMG/PR: a database of plasmids from genomes and metagenomes with rich annotations and metadata

Plasmids are mobile genetic elements found in many clades of Archaea and Bacteria. They drive horizontal gene transfer, impacting ecological and evolutionary processes within microbial communities, and hold substantial importance in human health and biotechnology. To support plasmid research and provide scientists with data of an unprecedented diversity of plasmid sequences, we introduce the IMG/PR database, a new resource encompassing 699 973 plasmid sequences derived from genomes, metagenomes and metatranscriptomes. IMG/PR is the first database to provide data of plasmid that were systematically identified from diverse microbiome samples. IMG/PR plasmids are associated with rich metadata that includes geographical and ecosystem information, host taxonomy, similarity to other plasmids, functional annotation, presence of genes involved in conjugation and antibiotic resistance. The database offers diverse methods for exploring its extensive plasmid collection, enabling users to navigate plasmids through metadata-centric queries, plasmid comparisons and BLAST searches. The web interface for IMG/PR is accessible at https://img.jgi.doe.gov/pr. Plasmid metadata and sequences can be downloaded from https://genome.jgi.doe.gov/portal/IMG_PR.

59 BASIC BIOLOGICAL SCIENCES↗

Annotated Genome Sequence of the High-Biomass-Producing Yellow-Green Alga Tribonema minus

Here, we report the annotated genome sequence for a heterokont alga from the class Xanthophyceae. This high-biomass-producing strain, Tribonema minus UTEX B 3156, was isolated from a wastewater treatment plant in California. It is stable in outdoor raceway ponds and is a promising industrial feedstock for biofuels and bioproducts.

59 BASIC BIOLOGICAL SCIENCES↗

Dataset: Breaking the barrier of human-annotated training data for machine-learning-aided plant research using aerial imagery

This dataset supports the implementation described in the manuscript "Breaking the Barrier of Human-Annotated Training Data for Machine-Learning-Aided Biological Research Using Aerial Imagery." It comprises UAV aerial imagery used to execute the code available at https://github.com/pixelvar79/GAN-Flowering-Detection-paper. For detailed information on dataset usage and instructions for implementing the code to reproduce the study, please refer to the GitHub repository.

generative and adversarial learning↗

Review of Literature and Utility Commission Proceedings Relevant to Integrated System Planning: Annotated Bibliography Prepared to Support the Washington Utilities and Transportation Commission

In 2024 the Washington State Legislature passed the Decarbonization Act for Large Combination Utilities (Engrossed Substitute House Bill 1589 – the Act). The Act requires a large combination electric and gas utility to conduct integrated system planning supporting electrification and gas system decarbonization, including a reduction in the gas rate base. The utility is required to submit the first integrated system plan (ISP) by January 1, 2027. The requirements are to be developed and adopted by the Washington Utilities and Transportation Commission (UTC) by July 1, 2025. In October 2024 Pacific Northwest National Laboratory (PNNL) and Lawrence Berkeley National Laboratory (LBNL) began providing technical assistance to the UTC to support the ISP rulemaking. PNNL and LBNL have prepared this annotated bibliography of research and reports, and state examples of coordinated gas and electric planning, future of gas, and future of heat proceedings in other U.S. States and one Canadian Province. The items described here have been selected by the authors for their potential relevance to the UTC’s integrated system planning rule discussions.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Co-Located Wave Energy Converter (WEC) and Aquaculture System Annotated Bibliography

This annotated bibliography includes references that could aid in the design of a co-located WEC and aquaculture system off the coast of Guam. The breadth of this work covers multiple co-location archetypes such as: 1. WEC seawater desalination system a) Nearshore and deepwater WEC deployment b) Onshore and offshore aquaculture 2. WEC powering an offshore aquaculture platform a) Nearshore or deepwater WEC deployment 3. WEC powering an onshore aquaculture system a) Nearshore WEC deployment 4. Wave powered seawater pump a) Nearshore WEC deployment b) Onshore aquaculture There are two archetypes that may be of immediate interest to the community in Guam are to service the existing Fadian Hatchery (Mangilao) and to support freshwater aquaculture activities. First, the seawater pump at the hatchery that fills the facility’s seawater storage unit is broken. A nearshore seawater pumping WEC could be a solution to this issue. In addition, due to the high frequency of typhoons/extreme conditions and the island’s bathymetry, the likelihood of community support for an offshore aquaculture platform or WEC deployed in deepwater is low. Proactive and resilient solutions not just for power, but for freshwater are of interest as well to support any freshwater aquaculture activities. Therefore, a nearshore WEC desalination system is another archetype to consider.

16 TIDAL AND WAVE POWER↗

Introduction to Digital Image Correlation (DIC) with annotated bibliography

Digital Image Correlation (DIC) is “a non-contact means of measuring motion and deformation using digital images of the object of interest”, capable of full-field measurements over large areas. Getting started in DIC can be difficult, since there are many different techniques, and the body of literature on DIC is extensive. IEEE Xplore alone lists nearly 11,000 references, including magazines, conference proceedings, journal articles, books, etc. A Google search returns around 145 million results. The intent of this annotated bibliography is to provide an entry point for beginners and to collect references that may be useful to more advanced practitioners.

42 ENGINEERING↗

Transporter annotations are holding up progress in metabolic modeling

Mechanistic, constraint-based models of microbial isolates or communities are a staple in the metabolic analysis toolbox, but predictions about microbe-microbe and microbe-environment interactions are only as good as the accuracy of transporter annotations. A number of hurdles stand in the way of comprehensive functional assignments for membrane transporters. These include general or non-specific substrate assignments, ambiguity in the localization, directionality and reversibility of a transporter, and the many-to-many mapping of substrates, transporters and genes. In this perspective, we summarize progress in both experimental and computational approaches used to determine the function of transporters and consider paths forward that integrate both. Investment in accurate, high-throughput functional characterization is needed to train the next-generation of predictive tools toward genome-scale metabolic network reconstructions that better predict phenotypes and interactions. More reliable predictions in this domain will benefit fields ranging from personalized medicine to metabolic engineering to microbial ecology.

Casey, John↗

A Preferences Corpus and Annotation Scheme for Human-Guided Alignment of Time-Series GPTs

The process of time-series forecasting such as predicting trajectories of silicon content in blast furnaces is a difficult task. Most time-series approaches today focus on scalar-type MSE loss optimization. This optimization approach, while widely common, could benefit from the use of human expert or process-level preferences. In this paper, we introduce a novel alignment and fine-tuning approach that involves learning from a corpus of preferred and dis-preferred time-series prediction trajectories. Our contributions include (1) a preference annotation pipeline for time-series forecasts, (2) the application of Score-based Preference Optimization (SPO) to train decoder-only transformers from preferences, and (3) results showing improvements in forecast quality. The approach is validated on both proprietary blast furnace data and the UCI Appliances Energy dataset. The proposed preference corpus and training strategy offer a new option for fine-tuning sequence models in industrial settings.

DPO↗