Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “annotations”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

Filling gaps in bacterial catabolic pathways with computation and high-throughput genetics

To discover novel catabolic enzymes and transporters, we combined high-throughput genetic data from 29 bacteria with an automated tool to find gaps in their catabolic pathways. GapMind for carbon sources automatically annotates the uptake and catabolism of 62 compounds in bacterial and archaeal genomes. For the compounds that are utilized by the 29 bacteria, we systematically examined the gaps in GapMind’s predicted pathways, and we used the mutant fitness data to find additional genes that were involved in their utilization. We identified novel pathways or enzymes for the utilization of glucosamine, citrulline, myo-inositol, lactose, and phenylacetate, and we annotated 299 diverged enzymes and transporters. We also curated 125 proteins from published reports. For the 29 bacteria with genetic data, GapMind finds high-confidence paths for 85% of utilized carbon sources. In diverse bacteria and archaea, 38% of utilized carbon sources have high-confidence paths, which was improved from 27% by incorporating the fitness-based annotations and our curation. GapMind for carbon sources is available as a web server ( http://papers.genomics.lbl.gov/carbon ) and takes just 30 seconds for the typical genome.

59 BASIC BIOLOGICAL SCIENCES↗

Individual Wave Detection and Tracking within a Rotating Detonation Engine through Computer Vision Object Detection applied to High-Speed Images

Known for their simplistic design and continuous detonation, rotating detonation engines (RDEs) constitute a majority of current pressure gain combustion (PGC) research efforts. Experimental RDE operation times have been continuously extended through the use of rig cooling techniques. As the window of observable behavior is expanded, and as the technology matures toward eventual integration within gas turbines, monitoring techniques must evolve to better match industrial diagnostics. High-speed image analysis techniques prove useful to capture and evaluate the unsteady detonation behavior within the RDE. Traditional image analysis techniques, however, require extensive processing times which prohibit simultaneous monitoring. To better address this problem, a computer vision object detection methodology is proposed to quickly detect individual detonation waves within a single down-axis image. Detonation waves are detected in individual images by the implemented computer vision method You Only Look Once (YOLO) object detection network. In order to detect detonation waves, the network must first be trained using RDE images of interest, for which each required phase of network development is outlined. Detection of waves is improved through proper treatment of the collected image set, variation of Intersection over Union (IoU) and confidence thresholding, and through a parametric study of annotation dimensions. Each detected wave is described by its location and rotational direction, and locations are tracked to calculate wave velocity across each frame, leading to a timestep resolution of 20 µs. Wave velocities are also calculated through a series of frames, leading to a suitable average velocity estimation using as few as 10 frames. Uncertainty analysis accounting for variation in camera framerate, pixel width and annotation centroid locations estimates a total uncertainty of ±4.3% for velocity calculations, using the smallest annotation boxes. This method offers great reductions in processing times, as a step toward real-time monitoring of detonation waves within an RDE. Improving on previous studies, this technique is impartial to wave modes not included in the original training set and calculates wave velocities independent of high-speed pressure data. The ability to isolate waves within predicted bounding boxes will likely facilitate analysis of pixel intensity variation as an estimation of wave strength in future work.

Johnson, Kristyn↗

Dataset for the Danczak et al., 2025 manuscript about bacterial-fungal interactions

We generated genome-resolved multiomics data from a series of metagenomic and metatranscriptomic sequencing. Specifically, we acquired, functionally annotated, and taxonomically classified both bacterial and eukaryotic metagenome assembled genomes (MAGs). For bacterial MAGs, we assembled eukaryotic float metagenomic sequencing data from JGI using MEGAHIT, binned and refined MAGs using MetaWRAP and dRep, functionally annotated MAGs using eggNOG mapper, and assigned taxonomy using GTDB-tk. For eukaryotic MAGs, we first identified potentially eukaryotic contigs from a coassembly of eukaryotic float metagenomic sequencing data from JGI using EukRep and Whokaryote, binned MAGs using MetaBAT2, functionally annotated MAGs using eggNOG mapper, and assigned taxonomy using Eukulele. Bulk metatranscriptomic reads were mapped to bacterial MAGs and polyA-metatranscriptomic read were mapped to eukaryotic MAGs using bbmap.

Danczak, Robert E. [Pacific Northwest National Lab↗

Hyaloscypha finlandica Metabolome Repository

This repository provides the curated data tables, manuscript figure and table exports, dependency records, and workflow scripts supporting an integrated comparative genomics and untargeted LC-MS/MS metabolomics analysis of Hyaloscypha finlandica strain PMI 746, a root-associated dark septate endophyte of poplar. The repository includes genome-mining summaries from antiSMASH, FunBGCeX, BGC-Prophet, and BiG-SCAPE; processed metabolomics inputs; metabolite annotation evidence; statistical outputs; and publication-facing figures and tables. Raw LC-MS/MS spectra, full genome/protein downloads, and large generated tool outputs are referenced through public archive/accession records and are not stored in Git.

59 BASIC BIOLOGICAL SCIENCES↗

Genome-scale model development and genomic sequencing of the oleaginous clade Lipomyces

The Lipomyces clade contains oleaginous yeast species with advantageous metabolic features for biochemical and biofuel production. Limited knowledge about the metabolic networks of the species and limited tools for genetic engineering have led to a relatively small amount of research on the microbes. Here, a genome-scale metabolic model (GSM) of Lipomyces starkeyi NRRL Y-11557 was built using orthologous protein mappings to model yeast species. Phenotypic growth assays were used to validate the GSM (66% accuracy) and indicated that NRRL Y-11557 utilized diverse carbohydrates but had more limited catabolism of organic acids. The final GSM contained 2,193 reactions, 1,909 metabolites, and 996 genes and was thus named iLst996. The model contained 96 of the annotated carbohydrate-active enzymes. iLst996 predicted a flux distribution in line with oleaginous yeast measurements and was utilized to predict theoretical lipid yields. Twenty-five other yeasts in the Lipomyces clade were then genome sequenced and annotated. Sixteen of the Lipomyces species had orthologs for more than 97% of the iLst996 genes, demonstrating the usefulness of iLst996 as a broad GSM for Lipomyces metabolism. Pathways that diverged from iLst996 mainly revolved around alternate carbon metabolism, with ortholog groups excluding NRRL Y-11557 annotated to be involved in transport, glycerolipid, and starch metabolism, among others. Overall, this study provides a useful modeling tool and data for analyzing and understanding Lipomyces species metabolism and will assist further engineering efforts in Lipomyces .

59 BASIC BIOLOGICAL SCIENCES↗

NCBI’s Virus Discovery Codeathon: Building “FIVE” —The Federated Index of Viral Experiments API Index

Viruses represent important test cases for data federation due to their genome size and the rapid increase in sequence data in publicly available databases. However, some consequences of previously decentralized (unfederated) data are lack of consensus or comparisons between feature annotations. Unifying or displaying alternative annotations should be a priority both for communities with robust entry representation and for nascent communities with burgeoning data sources. To this end, during this three-day continuation of the Virus Hunting Toolkit codeathon series (VHT-2), a new integrated and federated viral index was elaborated. This Federated Index of Viral Experiments (FIVE) integrates pre-existing and novel functional and taxonomy annotations and virus–host pairings. Variability in the context of viral genomic diversity is often overlooked in virus databases. As a proof-of-concept, FIVE was the first attempt to include viral genome variation for HIV, the most well-studied human pathogen, through viral genome diversity graphs. As per the publication of this manuscript, FIVE is the first implementation of a virus-specific federated index of such scope. FIVE is coded in BigQuery for optimal access of large quantities of data and is publicly accessible. Many projects of database or index federation fail to provide easier alternatives to access or query information. To this end, a Python API query system was developed to enhance the accessibility of FIVE.

59 BASIC BIOLOGICAL SCIENCES↗

AMVOS: Additive Manufacturing Video Object Segmentation Dataset

This dataset provides labeled video frames from four additive manufacturing (AM) processes for video object segmentation (VOS) tasks. It contains 90 video segments comprising 900 individually annotated frames across five AM datasets: laser hot-wire directed energy deposition (LHW-DED), tungsten inert gas wire arc additive manufacturing (TIG-WAAM), plasma arc welding (PAW), visible-light polymer extrusion (visPolymer), and near-infrared polymer extrusion (irPolymer). Each video segment consists of 10 contiguous frames with corresponding pixel-level object instance annotations. Depending on the process, two of four object classes are labeled per frame: Melt Pool, Feed Wire, Nozzle, or Material. Raw frames are provided as .jpg files and annotations as palettized .png files. The dataset follows the directory structure of established VOS benchmarks (DAVIS, YouTube-VOS, MOSE), enabling direct integration into VOS model training and evaluation pipelines for foundation model fine-tuning, domain adaptation, or zero-shot performance benchmarking. Data was collected at Oak Ridge National Laboratory's Manufacturing Demonstration Facility.

Wetzel, Jon [ORNL]↗

A Tool for Automatic Data Distribution for CFD Applications on Structured Grids

Development of HPF versions of NPB and ARC3D has shown that HPF provides an efficient, concise way to express parallelism and to organize data traffic. The use of HPF, as noted in the papers, requires an intimate knowledge of the applications and a detailed analysis of data affinity, data movement, and data granularity. To simplify and accelerate the task of developing HPF versions of existing CFD applications we have designed and implemented ADAPT (Automatic Data Alignment and Placement Tool). ADAPT analyzes a CFD application working on a single structured grid and generates HPF TEMPLATE, (RE)DISTRIBUTION, ALIGNMENT, and INDEPENDENT directives. The directives can be generated on the nest level, subroutine level, application level, or on the application interface level. ADAPT annotates an existing CFD FORTRAN application, performing computations on single or multiple grids. On each grid the application is considered as a sequence of operators, each applied to a set of variables defined in a particular grid domain. ADAPT automatically detects implicit operators (i.e., having data dependences) and explicit operators (without data dependences). For parallelization of an explicit operator ADAPT creates a template for the operator domain, aligns arrays used in the operator with the template, distributes the template, and declares the loops over the distributed dimensions as INDEPENDENT. For parallelization of an implicit operator, the distribution of the operator's domain should be consistent with the operator's dependences. Any dependence between sections distributed on different processors would preclude parallelization if the compiler does not have an ability to pipeline computations. If a data distribution is "orthogonal" to the dependences of an implicit operator, then the loop which implements the operator can be declared as INDEPENDENT. ADAPT starts with an analysis of array index expressions of the loop nests. For each pair of arrays referenced in an assignment statement, it generates an arc in the alignment graph and annotates it with an affinity relation. The template, alignment, and distribution directives for a particular loop nest are then derived from a transitive closure of the affinity relation. A compromise of data distributions in different nests and subroutines is achieved by merging annotated alignment graphs for adjacent nests/stibroutine calls in the nest/call graph of the application in the process called distribution lifting. ADAPT has been implemented as a C++ program running in conjunction with a parallelization tool called CAPTools. ADAPT uses the parse tree, interprocedural analysis and application database generated by CAPTools. It also uses the Directed Graph class, initially implemented in p2d2 (parallel debugger oi distributed programs), and some other classes supporting symbolic computations. ADAPT uses data distribution techniques described. ADAPT was tested with ARC3D and the FT benchmark and has demonstrated a code performance within a factor of 1.5 of handwritten versions.

Frumkin, Michael↗

A Software Architecture for Intelligent Synthesis Environments

The NASA's Intelligent Synthesis Environment (ISE) program is a grand attempt to develop a system to transform the way complex artifacts are engineered. This paper discusses a "middleware" architecture for enabling the development of ISE. Desirable elements of such an Intelligent Synthesis Architecture (ISA) include remote invocation; plug-and-play applications; scripting of applications; management of design artifacts, tools, and artifact and tool attributes; common system services; system management; and systematic enforcement of policies. This paper argues that the ISA extend conventional distributed object technology (DOT) such as CORBA and Product Data Managers with flexible repositories of product and tool annotations and "plug-and-play" mechanisms for inserting "ility" or orthogonal concerns into the system. I describe the Object Infrastructure Framework, an Aspect Oriented Programming (AOP) environment for developing distributed systems that provides utility insertion and enables consistent annotation maintenance. This technology can be used to enforce policies such as maintaining the annotations of artifacts, particularly the provenance and access control rules of artifacts-, performing automatic datatype transformations between representations; supplying alternative servers of the same service; reporting on the status of jobs and the system; conveying privileges throughout an application; supporting long-lived transactions; maintaining version consistency; and providing software redundancy and mobility.

Filman, Robert E.↗

AIRS Maps from Space Processing Software

This software package processes Atmospheric Infrared Sounder (AIRS) Level 2 swath standard product geophysical parameters, and generates global, colorized, annotated maps. It automatically generates daily and multi-day averaged colorized and annotated maps of various AIRS Level 2 swath geophysical parameters. It also generates AIRS input data sets for Eyes on Earth, Puffer-sphere, and Magic Planet. This program is tailored to AIRS Level 2 data products. It re-projects data into 1/4-degree grids that can be combined and averaged for any number of days. The software scales and colorizes global grids utilizing AIRS-specific color tables, and annotates images with title and color bar. This software can be tailored for use with other swath data products for the purposes of visualization.

Thompson, Charles K.↗

Individual Wave Detection and Tracking within a Rotating Detonation Engine through Computer Vision Object Detection applied to High-Speed Images

Known for their simplistic design and continuous detonation, rotating detonation engines (RDEs) constitute a majority of current pressure gain combustion (PGC) research efforts. Experimental RDE operation times have been continuously extended through the use of rig cooling techniques. As the window of observable behavior is expanded, and as the technology matures toward eventual integration within gas turbines, monitoring techniques must evolve to better match industrial diagnostics. High-speed image analysis techniques prove useful to capture and evaluate the unsteady detonation behavior within the RDE. Traditional image analysis techniques, however, require extensive processing times which prohibit simultaneous monitoring. To better address this problem, a computer vision object detection methodology is proposed to quickly detect individual detonation waves within a single down-axis image. Detonation waves are detected in individual images by the implemented computer vision method You Only Look Once (YOLO) object detection network. In order to detect detonation waves, the network must first be trained using RDE images of interest, for which each required phase of network development is outlined. Detection of waves is improved through proper treatment of the collected image set, variation of Intersection over Union (IoU) and confidence thresholding, and through a parametric study of annotation dimensions. Each detected wave is described by its location and rotational direction, and locations are tracked to calculate wave velocity across each frame, leading to a timestep resolution of 20 µs. Wave velocities are also calculated through a series of frames, leading to a suitable average velocity estimation using as few as 10 frames. Uncertainty analysis accounting for variation in camera framerate, pixel width and annotation centroid locations estimates a total uncertainty of ±4.3% for velocity calculations, using the smallest annotation boxes. This method offers great reductions in processing times, as a step toward real-time monitoring of detonation waves within an RDE. Improving on previous studies, this technique is impartial to wave modes not included in the original training set and calculates wave velocities independent of high-speed pressure data. The ability to isolate waves within predicted bounding boxes will likely facilitate analysis of pixel intensity variation as an estimation of wave strength in future work.

Johnson, Kristyn↗

A traffic accident dataset for Chattanooga, Tennessee

This publication presents an annotated accident dataset which fuses traffic data from radar detection sensors, weather condition data, and light condition data with traffic accident data (as illustrated in Fig. 1) in a format that is easy to process using machine learning tools, databases, or data workflows. The purpose of this data is to analyze, predict, and detect traffic patterns when accidents occur. Each file contains a timeseries of traffic speeds, flows, and occupancies at the sensor nearest to the accident, as well as 5 neighboring sensors upstream and downstream. It also contains information about the accident type, date, and time. In addition to the accident data, we provide baseline data for typical traffic patterns during a given time of day. Overall, the dataset contains 6 months of annotated traffic data from November 2020 to April 2021. During this timeframe, and 361 accidents occurred in the monitored area around Chattanooga, Tennessee. This dataset served as the basis for a study on topology-aware automated accident detection for a companion publication [1].

97 MATHEMATICS AND COMPUTING↗

Identification and preliminary characterization of conserved uncharacterized proteins from Chlamydomonas reinhardtii , Arabidopsis thaliana , and Setaria viridis

Abstract The rapid accumulation of sequenced plant genomes in the past decade has outpaced the still difficult problem of genome‐wide protein‐coding gene annotation. A substantial fraction of protein‐coding genes in all plant genomes are poorly annotated or unannotated and remain functionally uncharacterized. We identified unannotated proteins in three model organisms representing distinct branches of the green lineage (Viridiplantae): Arabidopsis thaliana (eudicot), Setaria viridis (monocot), and Chlamydomonas reinhardtii (Chlorophyte alga). Using similarity searching, we identified a subset of unannotated proteins that were conserved between these species and defined them as Deep Green proteins. Bioinformatic, genomic, and structural predictions were performed to begin classifying Deep Green genes and proteins. Compared to whole proteomes for each species, the Deep Green set was enriched for proteins with predicted chloroplast targeting signals predictive of photosynthetic or plastid functions, a result that was consistent with enrichment for daylight phase diurnal expression patterning. Structural predictions using AlphaFold and comparisons to known structures showed that a significant proportion of Deep Green proteins may possess novel folds. Though only available for three organisms, the Deep Green genes and proteins provide a starting resource of high‐value targets for further investigation of potentially new protein structures and functions conserved across the green lineage.

59 BASIC BIOLOGICAL SCIENCES↗

A Comment on “Deep Proteogenomics of a Photosynthetic Cyanobacterium”

Proteomic researchers strive to achieve complete annotation of protein-coding DNA sequences to provide a foundational context for their relevant biological data. A recent deep proteogenomic study using a photosynthetic cyanobacterium Synechocystis sp. PCC 6803 by Spät et al. proposed 64 refined open reading frames (ORFs). By searching LC-MS/MS data from affinity chromatography-isolated protein complexes, our laboratory identified that six of these high-abundance ORFs possess Nterminal initiation start sites that differ than those proposed in the alternative models. Our findings are supported by highly confident MS2 data, phylogenetic analysis, chemical labeling, and established data from two independent research groups. Based on these highquality experimental identifications, we subsequently propose a standardized strategy and set of criteria for future deep proteogenomic efforts to ensure accurate and stringent proteogenomic annotation.

cyanobacteria↗

High-throughput protein characterization by complementation using DNA barcoded fragment libraries

Abstract Our ability to predict, control, or design biological function is fundamentally limited by poorly annotated gene function. This can be particularly challenging in non-model systems. Accordingly, there is motivation for new high-throughput methods for accurate functional annotation. Here, we used co mplementation of aux otrophs and DNA barcode seq uencing (Coaux-Seq) to enable high-throughput characterization of protein function. Fragment libraries from eleven genetically diverse bacteria were tested in twenty different auxotrophic strains of Escherichia coli to identify genes that complement missing biochemical activity. We recovered 41% of expected hits, with effectiveness ranging per source genome, and observed success even with distant E. coli relatives like Bacillus subtilis and Bacteroides thetaiotaomicron . Coaux-Seq provided the first experimental validation for 53 proteins, of which 11 are less than 40% identical to an experimentally characterized protein. Among the unexpected function identified was a sulfate uptake transporter, an O-succinylhomoserine sulfhydrylase for methionine synthesis, and an aminotransferase. We also identified instances of cross-feeding wherein protein overexpression and nearby non-auxotrophic strains enabled growth. Altogether, Coaux-Seq’s utility is demonstrated, with future applications in ecology, health, and engineering.

59 BASIC BIOLOGICAL SCIENCES↗

DOE JGI Metagenome Workflow

The DOE Joint Genome Institute (JGI) Metagenome Workflow performs metagenome data processing, including assembly; structural, functional, and taxonomic annotation; and binning of metagenomic data sets that are subsequently included into the Integrated Microbial Genomes and Microbiomes (IMG/M) (I.-M. A. Chen, K. Chu, K. Palaniappan, A. Ratner, et al., Nucleic Acids Res, 49:D751–D763, 2021, https://doi.org/10.1093/nar/gkaa939) comparative analysis system and provided for download via the JGI data portal (https://genome.jgi.doe.gov/portal/). This workflow scales to run on thousands of metagenome samples per year, which can vary by the complexity of microbial communities and sequencing depth. Here, we describe the different tools, databases, and parameters used at different steps of the workflow to help with the interpretation of metagenome data available in IMG and to enable researchers to apply this workflow to their own data. We use 20 publicly available sediment metagenomes to illustrate the computing requirements for the different steps and highlight the typical results of data processing. The workflow modules for read filtering and metagenome assembly are available as a workflow description language (WDL) file (https://code.jgi.doe.gov/BFoster/jgi_meta_wdl). The workflow modules for annotation and binning are provided as a service to the user community at https://img.jgi.doe.gov/submit and require filling out the project and associated metadata descriptions in the Genomes OnLine Database (GOLD) (S. Mukherjee, D. Stamatis, J. Bertsch, G. Ovchinnikova, et al., Nucleic Acids Res, 49:D723–D733, 2021, https://doi.org/10.1093/nar/gkaa983).

59 BASIC BIOLOGICAL SCIENCES↗

Structural Models and Sequence Alignment Results of the Desulfovibrio vulgaris Proteome

This dataset contains the structural models for the primary transcripts of the Desulfovibrio vulgaris proteome as well as sequence alignment results for a subset of the encoded proteins. For each protein, the five models inferred from AlphaFold 2 are provided. The largest pTM-scoring model for each protein was energy minimized; this minimized structure as well as its AlphaFold pickle output file are also provided. This set of structures represent an alternate source of models for the D. vulgaris proteome to those available in the AlphaFold Protein Structure Database (AFDB). This is a bit more complicated since the proteins reporting in the AFDB originate from an outdated form of the D. vulgaris sequence. The different versions of the D. vulgaris gene annotation are collected in the Chronology subdirectory; further consideration of these changes on the structural space of the proteome are currently underway. For proteins that have been annotated as hypothetical, sequence alignment results from the HHblits and SAdLSA alignment methods are provided. These methods are often more capable to resolve sequence homology than other methods. Therefore, the results from both HHblits and SAdLSA are provided to identify possible homologs for these challenging proteins. Numerous sequence databases are utilized for these alignments. References AlphaFold v2 Multimer: https://doi.org/10.1101/2021.10.04.463034. References HHblits: hhtps://doi.org/10.1186/s12859-019-3019-7. References SAdLSA: hhtps://doi.org/10.3389/fbinf.2021.689960.

59 BASIC BIOLOGICAL SCIENCES↗

SFA-VirOmics

PNNL's Soil Microbiome Science Focus Area (SFA) is focused on understanding the basic biology underpinning how interactions among various soil microbial community members, across trophic levels, lead to the emergence of community functions. Moisture, in particular, drives microbial interactions and influences everything from cell function to substrate fate within soils. The group predicts this results in repeatable, predictable phenotypes. The sum of these phenotypes comprises the “soil metaphenome”. Understanding how the soil metaphenome shifts in response to moisture will provide a basis for modeling and predicting these shifts in reaction network responses. Visit the PNNL Soil Microbiome Science Focus Area Program homepage for more information. PNNL’s Soil Microbiome SFA Virome Dataset Annotation page is an extension to the PNNL Soil Microbiome SFA repository on DataHub allows for exploring and downloading integrated experimental omics dataset annotations, associated experimental metadata, pre- and post-processed viral data files, and other associated materials directly related to experimental viral project data.

59 BASIC BIOLOGICAL SCIENCES↗