Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “annotations”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Visualizing and analyzing 3D biomolecular structures using Mol* at RCSB.org: Influenza A H5N1 virus proteome case study

The easiest and often most useful way to work with experimentally determined or computationally predicted structures of biomolecules is by viewing their three-dimensional (3D) shapes using a molecular visualization tool. Mol* was collaboratively developed by RCSB Protein Data Bank (RCSB PDB, RCSB.org) and Protein Data Bank in Europe (PDBe, PDBe.org) as an open-source, web-based, 3D visualization software suite for examination and analyses of biostructures. It is capable of displaying atomic coordinates and related experimental data of biomolecular structures together with a variety of annotations, facilitating basic and applied research, training, education, and information dissemination. Across RCSB.org, the RCSB PDB research-focused web portal, Mol* has been implemented to support single-mouse-click atomic-level visualization of biomolecules (e.g., proteins, nucleic acids, carbohydrates) with bound cofactors, small-molecule ligands, ions, water molecules, or other macromolecules. RCSB.org Mol* can seamlessly display 3D structures from various sources, allowing structure interrogation, superimposition, and comparison. Using influenza A H5N1 virus as a topical case study of an important pathogen, we exemplify how Mol* has been embedded within various RCSB.org tools—allowing users to view polymer sequence and structure-based annotations integrated from trusted bioinformatics data resources, assess patterns and trends in groups of structures, and view structures of any size and compositional complexity. In addition to being linked to every experimentally determined biostructure and Computed Structure Model made available at RCSB.org, Standalone Mol* is freely available for visualizing any atomic-level or multi-scale biostructure at rcsb.org/3d-view.

3D biostructure↗

RCSB protein data Bank: Next‐generation advanced search for exploration of experimental structures and computed structure models

Abstract The Protein Data Bank (PDB), established in 1971, is the primary global, open‐access archive for experimentally determined 3D macromolecular structures (proteins, RNA, DNA). The research‐focused RCSB.org web‐portal provides access to these data alongside more than one million machine‐learning‐predicted structure models, greatly expanding the available structural landscape. Rapid growth of both experimental and computational structures has increased the need for powerful yet accessible search tools that serve a broad and diverse scientific community. Herein, we describe a redesigned RCSB Protein Data Bank RCSB.org Advanced Search capability that supports intuitive discovery of 3D structures through a unified interface. This interface integrates annotation‐, sequence‐, and 3D structure‐based searches, embeds an interactive 3D viewer, and incorporates curated biological knowledge, such as catalytic site definitions from Mechanism and Catalytic Site Atlas and ligand‐guided structural motifs, for constructing geometry‐driven queries. A new Chemical Search tool allows definition of chemical queries via an integrated drawing tool or standard identifiers, seamlessly combining them with annotation filters. By allowing query definition directly within spatial and chemical contexts, these search interfaces reduce the need for detailed knowledge of residue numbering, chain identifiers, or external cheminformatics software. This capability enables efficient exploration of structures, chemical diversity, and structure–function relationships across all life domains. The redesigned interfaces can be accessed directly at rcsb.org/search/advanced for Advanced Search and rcsb.org/search/chemical for Chemical Search.

Rose, Yana [Research Collaboratory for Structural ↗

Genomic variation within the maize stiff-stalk heterotic germplasm pool

The stiff-stalk heterotic group in Maize (Zea mays L.) is an important source of inbreds used in U.S. commercial hybrid production. Founder inbreds B14, B37, B73, and, to a lesser extent, B84, are found in the pedigrees of a majority of commercial seed parent inbred lines. We created high-quality genome assemblies of B84 and four expired Plant Variety Protection (ex-PVP) lines LH145 representing B14, NKH8431 of mixed descent, PHB47 representing B37, and PHJ40, which is a Pioneer Hi-Bred International (PHI) early stiff-stalk type. Sequence was generated using long-read sequencing achieving highly contiguous assemblies of 2.13-2.18 Gbp with N50 scaffold lengths >200 Mbp. Inbred-specific gene annotations were generated using a core five-tissue gene expression atlas, whereas transposable element (TE) annotation was conducted using de novo and homology-directed methodologies. Compared with the reference inbred B73, synteny analyses revealed extensive collinearity across the five stiff-stalk genomes, although unique components of the maize pangenome were detected. Comparison of this set of stiff-stalk inbreds with the original Iowa Stiff Stalk Synthetic breeding population revealed that these inbreds represent only a proportion of variation in the original stiff-stalk pool and there are highly conserved haplotypes in released public and ex-Plant Variety Protection inbreds. Despite the reduction in variation from the original stiff-stalk population, substantial genetic and genomic variation was identified supporting the potential for continued breeding success in this pool. The assemblies described here represent stiff-stalk inbreds that have historical and commercial relevance and provide further insight into the emerging maize pangenome.

59 BASIC BIOLOGICAL SCIENCES↗

Systematic identification of transcriptional activation domains from non-transcription factor proteins in plants and yeast

Transcription factors can promote gene expression through activation domains. Whole-genome screens have systematically mapped activation domains in transcription factors but not in non-transcription factor proteins (e.g., chromatin regulators and coactivators). To fill this knowledge gap, we employed the activation domain predictor PADDLE to analyze the proteomes of Arabidopsis thaliana and Saccharomyces cerevisiae. We screened 18,000 predicted activation domains from >800 non-transcription factor genes in both species, confirming that 89% of candidate proteins contain active fragments. Our work enables the annotation of hundreds of nuclear proteins as putative coactivators, many of which have never been ascribed any function in plants. Analysis of peptide sequence compositions reveals how the distribution of key amino acids dictates activity. Finally, we validated short, "universal" activation domains with comparable performance to state-of-the-art activation domains used for genome engineering. Our approach enables the genome-wide discovery and annotation of activation domains that can function across diverse eukaryotes.

59 BASIC BIOLOGICAL SCIENCES↗

PlantSegNet: 3D point cloud instance segmentation of nearby plant organs with identical semantics

In this study, we introduce PlantSegNet, a novel neural network model for instance segmentation of nearby objects with similar geometric structures. Our work addresses the challenges of instance segmentation of plant point clouds, including the difficulty of annotating and labeling point clouds, the loss of local structural information in neural network components, and the generation of large numbers of incorrect small clusters due to poor choices of the loss function. One of the key contributions of our approach is a digital twin of sorghum, i.e., a procedural sorghum model, which was used to generate point clouds of sorghum fields. This allowed us to create a large-scale, annotated, synthetic dataset of sorghum plants that we used to train our PlantSegNet model. We demonstrated the effectiveness of our method in segmenting instances of sorghum leaves grown in outdoor field settings. To the best of our knowledge, this is the first study to address this specific instance segmentation problem for plants grown in such a setting. We compared our proposed method with other state-of-the-art methods for indoor settings, including SGPN and TreePartNet, on both synthetic and real data. Furthermore, our results show that PlantSegNet outperforms these methods regarding accuracy, robustness, and efficiency.

97 MATHEMATICS AND COMPUTING↗

MjCyc: Rediscovering the pathway-genome landscape of the first sequenced archaeon, Methanocaldococcus (Methanococcus) jannaschii

The genome of Methanocaldococcus (Methanococcus) jannaschii DSM 2661 was the first Archaeal genome to be sequenced in 1996. Subsequent sequence-based annotation cycles led to its first metabolic reconstruction in 2005. Leveraging new experimental results and function assignments, we have now re-annotated M. jannaschii, creating an updated resource with novel information and testable predictions in a pathway-genome database available at BioCyc.org. This reannotation effort has resulted in 652 function assignments with enzyme roles, accounting for a third of the total protein-coding entries for this genome. The updated resource includes 883 reactions, 540 enzymes, and 142 individual pathways. Despite notable progress in computational genomics, more than a third of the genome remains functionally uncharacterized. The publicly available MjCyc pathway-genome database holds great potential for the wider community to conduct research on the biology of methanogenic Archaea.

59 BASIC BIOLOGICAL SCIENCES↗

RU-net for automatic characterization of TRISO fuel cross sections

During irradiation, phenomena such as kernel swelling and buffer densification may impact the performance of tristructural isotropic (TRISO) particle fuel. Post-irradiation microscopy is often used to identify these irradiation-induced morphologic changes. However, each fuel compact generally contains thousands of TRISO particles. Manually performing the work to get statistical information on these phenomena is cumbersome and subjective. Here, to reduce the subjectivity inherent in that process and to accelerate data analysis, we used convolutional neural networks (CNNs) to automatically segment cross-sectional images of microscopic TRISO layers. CNNs are a class of machine-learning algorithms specifically designed for processing structured grid data. They have gained popularity in recent years due to their remarkable performance in various computer vision tasks, including image classification, object detection, and image segmentation. In this research, we generated a large irradiated TRISO layer dataset with more than 2,000 microscopic images of cross-sectional TRISO particles and the corresponding annotated images. Based on these annotated images, we used different CNNs to automatically segment different TRISO layers. These CNNs include RU-Net (developed in this study), as well as three existing architectures: U-Net, Residual Network (ResNet), and Attention U-Net. The preliminary results show that the model based on RU-Net performs best in terms of Intersection over Union (IoU). Using CNN models, we can expedite the analysis of TRISO particle cross sections, significantly reducing the manual labor involved and improving the objectivity of the segmentation results.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

A new paradigm in electron microscopy: Automated microstructure analysis utilizing a dynamic segmentation convolutional neutral network

Over the past half century, the transmission electron microscope enabled insight into the fundamental arrangements and structures of materials. State-of-the-art electron microscopes can acquire large image datasets across multiple imaging modalities. However, the manual annotation process for feature or defect quantification may not be feasible with the modern microscope. Convolutional neural networks emerged to characterize individual microstructural features from an image in a cost-effective, consistent manner. However, many of these neural network approaches rely on thousands to hundreds of thousands of manual annotations of each feature type across hundreds of images to train the network for adequate performance. This work focused on the development and application of a pixel-wise defect detection machine-learning dynamic segmentation convolutional neural network with associated automated acquisition and postprocessing to identify microstructural features rapidly and quantitatively from a small initial dataset incorporating multiple imaging modes. The approach was demonstrated for characterization of superalloy 718 from both single image acquisition on multiple detectors to in-situ evolution captured with a single detector on a standard desktop computer to demonstrate the low barrier to entry required for widespread adoption. Pixel-by-pixel class identification was excellent with strong identification of chemically distinct phases, structurally distinct phases, and defect structures, thus demonstrating the new paradigm of machine learning-assisted characterization.

36 MATERIALS SCIENCE↗

Omics-guided metabolic pathway discovery in plants: Resources, approaches, and opportunities

Plants produce a vast array of metabolites, the biosynthetic routes of which remain largely undetermined. Genome-scale enzyme and pathway annotations and omics technologies have revolutionized research to decrypt plant metabolism and produced a growing list of functionally characterized metabolic genes and pathways. However, what is known is still a tiny fraction of the metabolic capacity harbored by plants. Here, in this work, we review plant enzyme and pathway annotation resources and cutting-edge omics approaches to guide discovery and characterization of plant metabolic pathways. We also discuss strategies for improving enzyme function prediction by integrating protein 3D structure information and single cell omics. This review aims to serve as a primer for plant biologists to leverage omics datasets to facilitate understanding and engineering plant metabolism.

59 BASIC BIOLOGICAL SCIENCES↗

Improved Characterization of Soil Organic Matter by Integrating FT-ICR MS, Liquid Chromatography Tandem Mass Spectrometry, and Molecular Networking: A Case Study of Root Litter Decay under Drought Conditions

Understanding of how soil organic matter (SOM) chemistry is altered in a changing climate has advanced considerably; however, most SOM components remain unidentified, impeding the ability to characterize a major fraction of organic matter and predict what types of molecules, and from which sources, will persist in soil. Here we present a novel approach to better characterize SOM extracts by integrating information from three types of analyses, and we deploy this method to characterize decaying root-detritus soil microcosms subjected to either drought or normal conditions. To observe broad differences in composition, we employed direct infusion Fourier-transform ion cyclotron resonance mass spectrometry (DI-FT-ICR MS). We complemented this with liquid chromatography tandem mass spectrometry (LC-MS/MS) to identify components by library matching. Since libraries contain only a small fraction of SOM components, we also used fragment spectral cosine similarity scores to relate unknowns and library matches through molecular networks. This integrated approach allowed us to corroborate DI-FT-ICR MS molecular formulas using library matches, which included fungal metabolites and related polyphenolic compounds. We also inferred structures of unknowns from molecular networks and improved LC-MS/MS annotation rates from ~5 to 35% by considering DI-FT-ICR MS molecular formula assignments. Under drought conditions, we found greater relative amounts of lignin-like vs condensed aromatic polyphenol formulas and lower average nominal oxidation state of carbon, suggesting reduced decomposition of SOM and/or microbes under stress. Our integrated approach provides a framework for enhanced annotation of SOM components that is more comprehensive than performing individual data analyses in parallel.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Untargeted Spatial Metabolomics and Spatial Proteomics on the Same Tissue Section

An increasing number of spatial multiomic workflows have been recently developed. Some of these approaches have leveraged initial mass spectrometry imaging (MSI)-based spatial metabolomics to inform region of interest (ROI) selection for downstream spatial proteomics. However, these workflows have been limited by varied substrate requirements between modalities or have required analyzing serial sections (i.e., one section per modality). To mitigate these issues, we present a novel multiomic workflow that uses desorption electrospray ionization (DESI)-MSI to identify representative spatial metabolite patterns on-tissue prior to spatial proteomic analyses on the same tissue section. Further, this workflow is demonstrated here with a model mammalian tissue (coronal rat brain section) mounted on a polyethylene naphthalate-membrane slide. Initial DESI-MSI resulted in 160 annotations (SwissLipids) within to the METASPACE platform (≤20% false discovery rate). A segmentation map from the annotated ion images informed downstream ROI selection for spatial proteomics characterization from the same sample. The unspecific substrate requirements and minimal sample disruption inherent to DESI-MSI allowed for an optimized, downstream spatial proteomics assay, resulting in 3888 ± 240 to 4717 ± 48 proteins being confidently directed per ROI (200 µm x 200 µm). Finally, we demonstrate the integration of multiomic information, where we found ceramide localization to be correlated with SMPD3 abundance (ceramide synthesis protein), and we also utilized protein abundance to resolve metabolite isomeric ambiguity. Overall, the integration of DESI-MSI into the multiomic workflow allows for complementary spatial and molecular-level information to be achieved from optimized implementations of each MS assay inherent to the workflow itself.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

PubChemLite Plus Collision Cross Section (CCS) Values for Enhanced Interpretation of Nontarget Environmental Data

Finding relevant chemicals in the vast (known) chemical space is a major challenge for environmental and exposomics studies leveraging nontarget high resolution mass spectrometry (NT-HRMS) methods. Chemical databases now contain hundreds of millions of chemicals, yet many are not relevant. This article details an extensive collaborative, open science effort to provide a dynamic collection of chemicals for environmental, metabolomics, and exposomics research, along with supporting information about their relevance to assist researchers in the interpretation of candidate hits. The PubChemLite for Exposomics collection is compiled from ten annotation categories within PubChem, enhanced with patent, literature and annotation counts, predicted partition coefficient (logP) values, as well as predicted collision cross section (CCS) values using CCSbase. Monthly versions are archived on Zenodo under a CC-BY license, supporting reproducible research, and a new interface has been developed, including historical trends of patent and literature data, for researchers to browse the collection. This article details how PubChemLite can support researchers in environmental and exposomics studies, describes efforts to increase the availability of experimental CCS values, and explores known limitations and potential for future developments. The data and code behind these efforts are openly available.

PubChem↗

Seasonal Patterns in Mobile Colistin Resistance Gene Variants in Wastewater Bioaerosols and Liquid Sludge

To address the growing threat of global antimicrobial resistance, a one-health approach is needed to understand the complex socioecological cycling of antibiotic resistance genes. In this study, a metagenomics approach using DNA shotgun sequencing, metagenome assembly, and antibiotic resistance gene (ARG) annotation was used to examine seasonal patterns in the abundance of mobile colistin resistance (mcr) gene variants in bioaerosols and liquid sludge in three wastewater treatment plants (WWTPs). ARGs represented 0.2–0.8 and 0.1–0.2% of the bioaerosol and liquid sludge metagenomes, respectively, while mcr genes represented 0–0.3 and 0–0.5% of the identified ARGs in bioaerosol and liquid sludge metagenomes. Seven of the ten known mcr variants were detected in wastewater bioaerosol and liquid samples, with mcr-5 and mcr-8 being the most prevalent across all seasons and sites. Furthermore, additional functional and taxonomic annotation of mcr-containing metagenomic contigs showed that mcr genes were often located on contigs with other co-occurring ARGs and mobile genetic elements and may be harbored by opportunistic human pathogens and other bacterial taxa not previously associated with mcr genes. Atmospheric dispersion modeling showed that mcr-containing bioaerosols can be transported kilometers away from the WWTPs, resulting in the possible dissemination of these ARGs into surrounding environments and communities.

Antimicrobial agents↗

Metabolomics Analysis of Bacterial Pathogen Burkholderia thailandensis and Mammalian Host Cells in Co-culture

The Tier 1 HHS/USDA Select Agent Burkholderia pseudomallei is a bacterial pathogen that is highly virulent when introduced into the respiratory tract and intrinsically resistant to many antibiotics. Transcriptomic- and proteomic-based methodologies have been used to investigate mechanisms of virulence employed by B. pseudomallei and Burkholderia thailandensis, a convenient surrogate; however, analysis of the pathogen and host metabolomes during infection is lacking. Changes in the metabolites produced can be a result of altered gene expression and/or post-transcriptional processes. Thus, metabolomics complements transcriptomics and proteomics by providing a chemical readout of a biological phenotype, which serves as a snapshot of an organism’s physiological state. However, the poor signal from bacterial metabolites in the context of infection poses a challenge in their detection and robust annotation. In this work, we coupled mammalian cell culture-based metabolomics with feature-based molecular networking of mono- and co-cultures to annotate the pathogen’s secondary metabolome during infection of mammalian cells. These methods enabled us to identify several key secondary metabolites produced by B. thailandensis during infection of airway epithelial and macrophage cell lines. Additionally, the use of in silico approaches provided insights into shifts in host biochemical pathways relevant to defense against infection. Using chemical class enrichment analysis, for example, we identified changes in a number of host-derived compounds including immune lipids such as prostaglandins, which were detected exclusively upon pathogen challenge. Taken together, our findings indicate that co-culture of B. thailandensis with mammalian cells alters the metabolome of both pathogen and host and provides a new dimension of information for in-depth analysis of the host–pathogen interactions underlying Burkholderia infection.

60 APPLIED LIFE SCIENCES↗

Deep Learning on Multimodal Chemical and Whole Slide Imaging Data for Predicting Prostate Cancer Directly from Tissue Images

Prostate cancer is one of the most common cancers globally and is the second most common cancer in the male population in the US. Here we develop a study based on correlating the hematoxylin and eosin (H&E)-stained biopsy data with MALDI mass-spectrometric imaging data of the corresponding tissue to determine the cancerous regions and their unique chemical signatures and variations of the predicted regions with original pathological annotations. We obtain features from high-resolution optical micrographs of whole slide H&E stained data through deep learning and spatially register them with mass spectrometry imaging (MSI) data to correlate the chemical signature with the tissue anatomy of the data. We then use the learned correlation to predict prostate cancer from observed H&E images using trained coregistered MSI data. This multimodal approach can predict cancerous regions with ~80% accuracy, which indicates a correlation between optical H&E features and chemical information found in MSI. Further, we show that such paired multimodal data can be used for training feature extraction networks on H&E data which bypasses the need to acquire expensive MSI data and eliminates the need for manual annotation saving valuable time. Two chemical biomarkers were also found to be predicting the ground truth cancerous regions. This study shows promise in generating improved patient treatment trajectories by predicting prostate cancer directly from readily available H&E-stained biopsy images aided by coregistered MSI data.

60 APPLIED LIFE SCIENCES↗

Event‐Based Training in Label‐Limited Regimes

Abstract The distribution of attributes assigned using data on independent sensors for a specific source, for example, magnitude, can be richly descriptive for final event characterization and associated uncertainty. Attribute distributions can also provide powerful context for event characterization in the absence of comprehensive annotation. This work develops a way to leverage distributional information across a set of sensors in the absence of comprehensive annotation as a domain‐informed regularization term applied during gradient‐based learning. The regularization term is the basis of event‐based training which I show can be a powerful semi‐supervised learning (SSL) approach. I first use a simple feed forward neural network and a toy data set to outline how data set structure interacts with the assumptions inherent to many semi‐supervised learning approaches. I then demonstrate the effectiveness of event‐based training using a deep convolutional neural network for seismic event classification in Utah, which increases SSL accuracy from 92% to 97% on event classification with a limited number of training labels.

Linville, Lisa M.↗

Guided construction of single cell reference for human and mouse lung

Accurate cell type identification is a key and rate-limiting step in single-cell data analysis. Single-cell references with comprehensive cell types, reproducible and functionally validated cell identities, and common nomenclatures are much needed by the research community for automated cell type annotation, data integration, and data sharing. Here, we develop a computational pipeline utilizing the LungMAP CellCards as a dictionary to consolidate single-cell transcriptomic datasets of 104 human lungs and 17 mouse lung samples to construct LungMAP single-cell reference (CellRef) for both normal human and mouse lungs. CellRefs define 48 human and 40 mouse lung cell types catalogued from diverse anatomic locations and developmental time points. We demonstrate the accuracy and stability of LungMAP CellRefs and their utility for automated cell type annotation of both normal and diseased lungs using multiple independent methods and testing data. We develop user-friendly web interfaces for easy access and maximal utilization of the LungMAP CellRefs.

59 BASIC BIOLOGICAL SCIENCES↗

An expanded registry of candidate cis -regulatory elements

Mammalian genomes contain millions of regulatory elements that control the complex patterns of gene expression. Previously, the ENCODE consortium mapped biochemical signals across hundreds of cell types and tissues and integrated these data to develop a registry containing 0.9 million human and 300,000 mouse candidate cis-regulatory elements (cCREs) annotated with potential functions. Here we have expanded the registry to include 2.37 million human and 967,000 mouse cCREs, leveraging new ENCODE datasets and enhanced computational methods. This expanded registry covers hundreds of unique cell and tissue types, providing a comprehensive understanding of gene regulation. Functional characterization data from assays such as STARR-seq, massively parallel reporter assay, CRISPR perturbation and transgenic mouse assays have profiled more than 90% of human cCREs, revealing complex regulatory functions. We identified thousands of novel silencer cCREs and demonstrated their dual enhancer and silencer roles in different cellular contexts. Integrating the registry with other ENCODE annotations facilitates genetic variation interpretation and trait-associated gene identification, exemplified by the identification of KLF1 as a novel causal gene for red blood cell traits. This expanded registry is a valuable resource for studying the regulatory genome and its impact on health and disease.

Moore, Jill E. [Univ. of Massachusetts, Worchester↗