Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “annotations”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22

Hyperspectral segmentation of plants in fabricated ecosystems

Hyperspectral imaging provides a powerful tool for analyzing above-ground plant characteristics in fabricated ecosystems, offering rich spectral information across diverse wavelengths. This study presents an efficient workflow for hyperspectral data segmentation and subsequent data analytics, minimizing the need for user annotation through the use of ensembles of sparse mixed scale convolution neural networks. The segmentation process leverages the diversity of ensembles to achieve high accuracy with minimal labeled data, reducing labor-intensive annotation efforts. To further enhance robustness, we incorporate image alignment techniques to address spatial variability in the dataset. Downstream analysis focuses on using the segmented data for processing spectral data, enabling monitoring of plant health. This approach provides a scalable solution for spectral segmentation, and facilitates actionable insights into plant conditions in complex, controlled environments. Our results demonstrate the utility of combining advanced machine learning techniques with hyperspectral analytics for high-throughput plant monitoring.

Zwart, Petrus H.↗

FeGenie: a comprehensive tool for the identification of iron genes and iron gene neighborhoods in genomes and metagenome assemblies

Iron is a micronutrient for nearly all life on Earth. It can be used as an electron donor and electron acceptor by iron-oxidizing and iron-reducing microorganisms and is used in a variety of biological processes, including photosynthesis and respiration. While it is the fourth most abundant metal in the Earth’s crust, iron is often limiting for growth in oxic environments because it is readily oxidized and precipitated. Much of our understanding of how microorganisms compete for and utilize iron is based on laboratory experiments. However, the advent of next-generation sequencing and surge in publicly available sequence data has made it possible to probe the structure and function of microbial communities in the environment. To bridge the gap between our understanding of iron acquisition, iron redox cycling, iron storage, and magnetosome formation in model microorganisms and the plethora of sequence data available from environmental studies, we have created a comprehensive database of hidden Markov models (HMMs) based on genes related to iron acquisition, storage, and reduction/oxidation in Bacteria and Archaea. Along with this database, we present FeGenie, a bioinformatics tool that accepts genome and metagenome assemblies as input and uses our comprehensive HMM database to annotate provided datasets with respect to iron-related genes and gene neighborhood. An important contribution of this tool is the efficient identification of genes involved in iron oxidation and dissimilatory iron reduction, which have been largely overlooked by standard annotation pipelines. We validated FeGenie against a selected set of 28 isolate genomes and showcase its utility in exploring iron genes present in 27 metagenomes, 4 isolate genomes from human oral biofilms, and 17 genomes from candidate organisms, including members of the candidate phyla radiation. We show that FeGenie accurately identifies iron genes in isolates. Furthermore, analysis of metagenomes using FeGenie demonstrates that the iron gene repertoire and abundance of each environment is correlated with iron richness. While this tool will not replace the reliability of culture-dependent analyses of microbial physiology, it provides reliable predictions derived from the most up-to-date genetic markers. FeGenie’s database will be maintained and continually updated as new genes are discovered.

59 BASIC BIOLOGICAL SCIENCES↗

The Transcriptional Response of Soil Bacteria to Long-Term Warming and Short-Term Seasonal Fluctuations in a Terrestrial Forest

Terrestrial ecosystems are an important carbon store, and this carbon is vulnerable to microbial degradation with climate warming. After 30 years of experimental warming, carbon stocks in a temperate mixed deciduous forest were observed to be reduced by 30% in the heated plots relative to the controls. In addition, soil respiration was seasonal, as was the warming treatment effect. We therefore hypothesized that long-term warming will have higher expressions of genes related to carbohydrate and lipid metabolism due to increased utilization of recalcitrant carbon pools compared to controls. Because of the seasonal effect of soil respiration and the warming treatment, we further hypothesized that these patterns will be seasonal. We used RNA sequencing to show how the microbial community responds to long-term warming (~30 years) in Harvard Forest, MA. Total RNA was extracted from mineral and organic soil types from two treatment plots (+5°C heated and ambient control), at two time points (June and October) and sequenced using Illumina NextSeq technology. Treatment had a larger effect size on KEGG annotated transcripts than on CAZymes, while soil types more strongly affected CAZymes than KEGG annotated transcripts, though effect sizes overall were small. Although, warming showed a small effect on overall CAZymes expression, several carbohydrate-associated enzymes showed increased expression in heated soils (~68% of all differentially expressed transcripts). Further, exploratory analysis using an unconstrained method showed increased abundances of enzymes related to polysaccharide and lipid metabolism and decomposition in heated soils. Compared to long-term warming, we detected a relatively small effect of seasonal variation on community gene expression. Together, these results indicate that the higher carbohydrate degrading potential of bacteria in heated plots can possibly accelerate a self-reinforcing carbon cycle-temperature feedback in a warming climate.

54 ENVIRONMENTAL SCIENCES↗

AlgaeOrtho, a bioinformatics tool for processing ortholog inference results in algae

Introduction: Microalgae constitute a prominent feedstock for producing biofuels and biochemicals by virtue of their prolific reproduction, high bioproduct accumulation, and the ability to grow in brackish and saline water. However, naturally occurring wild type algal strains are rarely optimal for industrial use; therefore, bioengineering of algae is necessary to generate superior performing strains that can address production challenges in industrial settings, particularly the bioenergy and bioproduct sectors. One of the crucial steps in this process is deciding on a bioengineering target: namely, which gene/protein to differentially express. These targets are often orthologs which are defined as genes/proteins originating from a common ancestor in divergent species. Although bioinformatics tools for the identification of protein orthologs already exist, processing the output from such tools is nontrivial, especially for a researcher with little or no bioinformatics experience. Methods: The present study introduces AlgaeOrtho, a user-friendly tool that builds upon the SonicParanoid orthology inference tool (based on an algorithm that identifies potential protein orthologs based on amino acid sequences) and the PhycoCosm database from JGI (Joint Genome Institute) to help researchers identify orthologs of their proteins of interest in multiple diverse algal species. Results: The output of this application includes a table of the putative orthologs of their protein of interest, a heatmap showing sequence similarity (%), and an unrooted tree of the putative protein orthologs. Notably, the tool would be instrumental in identifying novel bioengineering targets in different algal strains, including targets in not-fully annotated algal species, since it does not depend on existing protein annotations. We tested AlgaeOrtho using three case studies, for which orthologs of proteins relevant to bioengineering targets, were identified from diverse algal species, demonstrating its ease of use and utility for bioengineering researchers. Discussion: This tool is unique in the protein ortholog identification space as it can visualize putative orthologs, as desired by the user, across several algal species.

09 BIOMASS FUELS↗

A chromosome-level genome assembly of the Chinese cork oak (Quercus variabilis)

Quercus variabilis (Fagaceae) is an ecologically and economically important deciduous broadleaved tree species native to and widespread in East Asia. It is a valuable woody species and an indicator of local forest health, and occupies a dominant position in forest ecosystems in East Asia. However, genomic resources from Q. variabilis are still lacking. Here, we present a high-quality Q. variabilis genome generated by PacBio HiFi and Hi-C sequencing. The assembled genome size is 787 Mb, with a contig N50 of 26.04 Mb and scaffold N50 of 64.86 Mb, comprising 12 pseudo-chromosomes. The repetitive sequences constitute 67.6% of the genome, of which the majority are long terminal repeats, accounting for 46.62% of the genome. We used ab initio , RNA sequence-based and homology-based predictions to identify protein-coding genes. A total of 32,466 protein-coding genes were identified, of which 95.11% could be functionally annotated. Evolutionary analysis showed that Q. variabilis was more closely related to Q. suber than to Q. lobata or Q. robur. We found no evidence for species-specific whole genome duplications in Quercus after the species had diverged. This study provides the first genome assembly and the first gene annotation data for Q. variabilis. These resources will inform the design of further breeding strategies, and will be valuable in the study of genome editing and comparative genomics in oak species.

Han, Biao↗

Survey of Thirteen Novel Pseudomonas putida Bacteriophages

Bacteriophages have been widely investigated as a promising treatment of food, medical equipment, and humans colonized by antibiotic-resistant bacteria. Phages pose particular interest in combating those bacteria which form biofilms, such as the medically important human pathogen Pseudomonas aeruginosa and several plant pathogens, including P. syringae . In an undergraduate lab course, P. putida was used as the host to isolate novel anti-pseudomonal bacteriophages. Environmental samples of soil and water were collected, and purified phage isolates were obtained. After Illumina sequencing, genomes of these phages were assembled de novo and annotated. Assembled genomes were compared with known genomes in the literature and GenBank to identify taxonomic relations and to refine their functional annotations. The thirteen phages described are sipho-, myo-, and podoviruses in several families of Caudoviricetes , spanning several novel genera, with genomes ranging from 40,000 to 96,000 bp. One phage (DDSR119) is unique and is the first reported P. putida siphovirus. The remaining 12 can be clustered into four distinct groups. Six are highly related to each other and to previously described Autotranscriptaviridae phages: Waldo5, PlaquesPlease, and Laces98 all belong to the Waldovirus genus, whereas Stalingrad, Bosely, and Stamos belong to the Troedvirus genus. Zuri was previously classified as the founding member of a new genus Zurivirus within the family Schitoviridae . Ebordelon and Holyagarpour each represent different species within Zurivirus , whereas Meara is a more distantly related member of the Schitoviridae . Dolphis and Jeremy are similar enough to form a genus but have only a few distant relatives among sequenced phages and are notable for being temperate. We identified the lysis cassettes in all 13 phages, compared tail spike structures, and found auxiliary metabolic genes in several. Studies like these, which isolate and characterize infectious virions, enable the identification of novel proteins and molecular systems and also provide the raw materials for further study, evaluation, and manipulation of phage proteins and their hosts.

Pseudomonas putida↗

Natural Product Gene Clusters in the Filamentous Nostocales Cyanobacterium HT-58-2

Cyanobacteria are known as rich repositories of natural products. One cyanobacterial-microbial consortium (isolate HT-58-2) is known to produce two fundamentally new classes of natural products: the tetrapyrrole pigments tolyporphins A–R, and the diterpenoid compounds tolypodiol, 6-deoxytolypodiol, and 11-hydroxytolypodiol. The genome (7.85 Mbp) of the Nostocales cyanobacterium HT-58-2 was annotated previously for tetrapyrrole biosynthesis genes, which led to the identification of a putative biosynthetic gene cluster (BGC) for tolyporphins. Here, bioinformatics tools have been employed to annotate the genome more broadly in an effort to identify pathways for the biosynthesis of tolypodiols as well as other natural products. A putative BGC (15 genes) for tolypodiols has been identified. Four BGCs have been identified for the biosynthesis of other natural products. Two BGCs related to nitrogen fixation may be relevant, given the association of nitrogen stress with production of tolyporphins. The results point to the rich biosynthetic capacity of the HT-58-2 cyanobacterium beyond the production of tolyporphins and tolypodiols.

60 APPLIED LIFE SCIENCES↗

A Genome-Scale Metabolic Model of Anabaena 33047 to Guide Genetic Modifications to Overproduce Nylon Monomers

Nitrogen fixing-cyanobacteria can significantly improve the economic feasibility of cyanobacterial production processes by eliminating the requirement for reduced nitrogen. Anabaena sp. ATCC 33047 is a marine, heterocyst forming, nitrogen fixing cyanobacteria with a very short doubling time of 3.8 h. We developed a comprehensive genome-scale metabolic (GSM) model, iAnC892, for this organism using annotations and content obtained from multiple databases. iAnC892 describes both the vegetative and heterocyst cell types found in the filaments of Anabaena sp. ATCC 33047. iAnC892 includes 953 unique reactions and accounts for the annotation of 892 genes. Comparison of iAnC892 reaction content with the GSM of Anabaena sp. PCC 7120 revealed that there are 109 reactions including uptake hydrogenase, pyruvate decarboxylase, and pyruvate-formate lyase unique to iAnC892. iAnC892 enabled the analysis of energy production pathways in the heterocyst by allowing the cell specific deactivation of light dependent electron transport chain and glucose-6-phosphate metabolizing pathways. The analysis revealed the importance of light dependent electron transport in generating ATP and NADPH at the required ratio for optimal N2 fixation. When used alongside the strain design algorithm, OptForce, iAnC892 recapitulated several of the experimentally successful genetic intervention strategies that over produced valerolactam and caprolactam precursors.

59 BASIC BIOLOGICAL SCIENCES↗

Bottleneck Detection in Modular Construction Factories Using Computer Vision

The construction industry is increasingly adopting off-site and modular construction methods due to the advantages offered in terms of safety, quality, and productivity for construction projects. Despite the advantages promised by this method of construction, modular construction factories still rely on manually-intensive work, which can lead to highly variable cycle times. As a result, these factories experience bottlenecks in production that can reduce productivity and cause delays to modular integrated construction projects. To remedy this effect, computer vision-based methods have been proposed to monitor the progress of work in modular construction factories. However, these methods fail to account for changes in the appearance of the modular units during production, they are difficult to adapt to other stations and factories, and they require a significant amount of annotation effort. Due to these drawbacks, this paper proposes a computer vision-based progress monitoring method that is easy to adapt to different stations and factories and relies only on two image annotations per station. In doing so, the Scale-invariant feature transform (SIFT) method is used to identify the presence of modular units at workstations, and the Mask R-CNN deep learning-based method is used to identify active workstations. This information was synthesized using a near real-time data-driven bottleneck identification method suited for assembly lines in modular construction factories. This framework was successfully validated using 420 h of surveillance videos of a production line in a modular construction factory in the U.S., providing 96% accuracy in identifying the occupancy of the workstations and an F-1 Score of 89% in identifying the state of each station on the production line. The extracted active and inactive durations were successfully used via a data-driven bottleneck detection method to detect bottleneck stations inside a modular construction factory. The implementation of this method in factories can lead to continuous and comprehensive monitoring of the production line and prevent delays by timely identification of bottlenecks.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Acoustic Rocket Signatures Collected by Smartphones

Rockets generate complex acoustic signatures that can be detected over a thousand kilometers from their source. While many far-field acoustic rocket signatures have been collected and released to the public, very few signatures collected at distances less than 100 km are available. This work presents a curated and annotated dataset of acoustic signatures of 243 rocket launches collected by a network of smartphones stationed at distances between 10 and 70 km from the launch sites, resulting in 1089 individual recordings. Due to the frequency dependence of atmospheric attenuation and the relatively short propagation distances, higher-frequency features not preserved in most publicly available data are observed. The signals are time-aligned to allow for different segments of the signal (ignition, launch, trajectory, chronology) to be more easily examined and compared. Initial analysis of the features of these rocket launch stages is performed, observed features are compared to those found in the existing literature, and comparisons between signals from launches of different rocket types are made. The dataset is annotated and made available to the public to aid future analysis of the characteristics and source mechanisms of rocket acoustics as well as applications such as rocket detection and classification models.

33 ADVANCED PROPULSION SYSTEMS↗

PRMI: A Dataset of Minirhizotron Images for Diverse Plant Root Study

Understanding a plant's root system architecture (RSA) is crucial for a variety of plant science problem domains including sustainability and climate adaptation. Minirhizotron (MR) technology is a widely-used approach for phenotyping RSA non-destructively by capturing root imagery over time. Precisely segmenting roots from the soil in MR imagery is a critical step in studying RSA features. In this paper, we introduce a large-scale dataset of plant root images captured by MR technology. In total, there are over 72K RGB root images across six different species including cotton, papaya, peanut, sesame, sunflower, and switchgrass in the dataset. The images span a variety of conditions including varied root age, root structures, soil types, and depths under the soil surface. All of the images have been annotated with weak image-level labels indicating whether each image contains roots or not. The image-level labels can be used to support weakly supervised learning in plant root segmentation tasks. In addition, 63K images have been manually annotated to generate pixel-level binary masks indicating whether each pixel corresponds to root or not. These pixel-level binary masks can be used as ground truth for supervised learning in semantic segmentation tasks. By introducing this dataset, we aim to facilitate the automatic segmentation of roots and the research of RSA with deep learning and other image analysis algorithms.

Xu, Weihuang↗

RU Net for Automatic Characterization of TRISO Fuel Cross Sections

TRistructural ISOtropic (TRISO) particle fuel is a type of nuclear fuel known for its high-temperature and high-burnup performance. Each sub-millimeter diameter TRISO particle consists of uranium-oxycarbide (UCO) or UO2 fuel kernel, coated with buffer, inner pyrolytic carbon (IPyC), silicon carbide (SiC), and outer pyrolytic carbon (OPyC) layers. The SiC layer acts as the main containment barrier for the TRISO particle to retain the fission products, while the IPyC and OPyC layers provide additional barriers to the release of fission products, especially fission gases. During irradiation, phenomena like kernel swelling, buffer densification, and IPyC fracture may impact fuel performance. Post-irradiation microscopy on entire compact cross sections or samples of individual particles deconsolidated from compacts is often used to identify these irradiation-induced changes in morphology. However, each fuel compact generally contains thousands of TRISO particles. To get statistical information on these phenomena, it is cumbersome work if done manually. For example, to get information about swelling/densification behaviors of different layers or kernels after irradiation, researchers previously manually measured the perimeter of each TRISO layer in hundreds of particles after four rounds of iterative grinding and polishing encompassing more than 2000 cross-section images for a total of four fuel compacts. To attempt to reduce the subjectivity inherent in that process and accelerate data analysis, we conducted a study on the automatic TRISO layer segmentation on cross-sectional microscopic images using Convolutional Neural Networks (CNNs). CNNs are a class of machine learning algorithms specifically designed for processing structured grid data that have gained popularity in recent years due to their remarkable performance in various computer vision tasks, including image classification, object detection, and image segmentation. In this research, we have generated the large irradiated TRISO layer dataset with more than 2000 cross-section TRISO microscopic images and the corresponding annotated images. Based on these annotated images, we have employed different CNNs for automatic segmentation of different TRISO layers. These include RU-Net (developed in this study), as well as three existing architectures: U-Net, Residual Network (ResNet), and Attention U-Net. The preliminary results show that the model based on RU-Net has the best performance in terms of intersection-over-union (IoU). Through the aid of these CNN models, we can expedite the analysis of TRISO particle cross-sections, significantly reducing the manual labor involved and improving the objectivity of the segmentation results.

Convolutional Neural Networks↗

YeastWGD2025

Supplementary data for Discovery of additional ancient genome duplications in yeasts wgd_syn / - directory containing wgd syn output for all contiguous genomes [dataset] Tree - phylogeny [dataset]Duplications - duplication table from OrthoFinder output KOannotations - KEGG annotations used for enrichment analysis IPRannotations - InterPro annotations used for enrichment analysis DipodascalesOrthogroups - formatted orthogroup assignments for Dipodascales genes.fa and .gff3 files for each new genome assembly are also provided, those these are not required to replicate the analysis

Genomics↗

Geologic interpretation of Apollo 6 stereophotography from Baja California to west Texas

Excellent space photography of parts of the southwestern United States and northwestern Mexico was obtained during the unmanned Apollo 6 spaceflight. Two features of this photography made it useful for geologic interpretations: its vertical stereocoverage and its exposure under a relatively low angle of solar illumination through an unusually cloud-free and clear atmosphere. The structural patterns, which were topographically enhanced by the longer shadows, were annotated on the photographs, in order to analyze their trends with respect to the continental tectonic framework, and to attempt to correlate the pattern with known copper or other base metal deposits. The annotated fracture patterns showed the regional trends and their distribution. The area studied was a 100- to 105-mile swath of terrain covering a total land area of approximately 60,000 square statute miles. The coverage began from a point centered on Punta Colnett on the Pacific coast of Baja California and extended to the Sacramento Mountains of New Mexico and west Texas.

Gawarecki, S. J.↗

ERTS data user investigation to develop a multistage forest sampling inventory system

The author has identified the following significant results. A system to provide precision annotation of predetermined forest inventory sampling units on the ERTS-1 MSS images was developed. In addition, an annotation system for high altitude U2 photographs was completed. MSS bulk image accuracy is good enough to allow the use of one square mile sampling units. IMANCO image analyzer interpretation work for small scale images demonstrated the need for much additional analyses. Continuing image interpretation work for the next reporting period is concentrated on manual image interpretation work as well as digital interpretation system development using the computer compatible tapes.

Langley, P. G.↗

Applicability of ERTS-1 to Montana geology

The author has identified the following significant results. A detailed band 7 ERTS-1 lineament map covering western Montana and northern Idaho has been prepared and is being evaluated by direct comparison with geologic maps, by statistical plots of lineaments and known faults, and by field checking. Lineament patterns apparent in the Idaho and Boulder batholiths do not correspond to any known geologic structures. A band 5 mosaic of Montana and adjacent areas has been laid and a lineament annotation prepared for comparison with the band 7 map. All work to date indicates that ERTS-1 imagery is very useful for revealing patterns of high-angle faults, though much less useful for mapping rock units and patterns of low-angle faults. Large-scale mosaics of U-2 photographs of three test sites have been prepared for annotation and comparison with ERTS-1 maps. Mapping of Quaternary deposits in the Glacial Lake Missoula basin using U-2 color infrared transparencies has been successful resulting in the discovery of some deposits not previously mapped. Detailed work has been done for Test Site 354 D using ERTS-1 imagery; criteria for recognition of several rock types have been found. Photogeologic mapping for southeastern Montana suggest Wasatch deposits where none shown of geologic map.

Weidman, R. M.↗

Physiologic responses to water immersion in man: A compendium of research

A total of 221 reports published through December 1973 in the area of physiologic responses to water immersion in man were summarized. The author's abstract or summary was used whenever possible. Otherwise, a detailed annotation was provided under the subheadings: (1) purpose, (2) procedures and methods, (3) results, and (4) conclusions. The annotations are in alphabetical order by first author; author and subject indexes are included. Additional references are provided in the selected bibliography.

Kollias, J.↗

Toward autonomous driving: The CMU Navlab. II - Architecture and systems

A description is given of EDDIE, the architecture for the Navlab mobile robot which provides a toolkit for building specific systems quickly and easily. Included in the discussion are the annotated maps used by EDDIE and the Navlab's road-following system, called the Autonomous Mail Vehicle, which was built using EDDIE and its annotated maps as a basis. The contributions of the Navlab project and the lessons learned from it are examined.

Thorpe, Charles↗