Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “annotations”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25

Structural Models and Sequence Alignment Results of the Rhodospirillum rubrum Proteome

This dataset contains the structural models for the primary transcripts of the Rhodospirillum rubrum proteome as well as sequence alignment results for a subset of the encoded proteins. For each protein, the five models inferred from AlphaFold 2 are provided. The largest pTM-scoring model for each protein was energy minimized; this minimized structure as well as its AlphaFold pickle output file are also provided. This set of structures represent an alternate source of models for the R. rubrum proteome to those available in the AlphaFold Protein Structure Database. For proteins that have been annotated as hypothetical, sequence alignment results from the HHblits and SAdLSA alignment methods are provided. These methods are often more capable to resolve sequence homology than other methods. Therefore, the results from both HHblits and SAdLSA are provided to identify possible homologs for these challenging proteins. Numerous sequence databases are utilized for these alignments. References AlphaFold v2 Multimer: https://doi.org/10.1101/2021.10.04.463034. References HHBlits: https://doi.org/10.1186/s12859-019-3019-7. References SAdLSA: https://doi.org/10.3389/fbinf.2021.689960.

59 BASIC BIOLOGICAL SCIENCES↗

RCSB Protein Data Bank: visualizing groups of experimentally determined PDB structures alongside computed structure models of proteins

Recent advances in Artificial Intelligence and Machine Learning (e.g., AlphaFold, RosettaFold, and ESMFold) enable prediction of three-dimensional (3D) protein structures from amino acid sequences alone at accuracies comparable to lower-resolution experimental methods. These tools have been employed to predict structures across entire proteomes and the results of large-scale metagenomic sequence studies, yielding an exponential increase in available biomolecular 3D structural information. Given the enormous volume of this newly computed biostructure data, there is an urgent need for robust tools to manage, search, cluster, and visualize large collections of structures. Equally important is the capability to efficiently summarize and visualize metadata, biological/biochemical annotations, and structural features, particularly when working with vast numbers of protein structures of both experimental origin from the Protein Data Bank (PDB) and computationally-predicted models. Moreover, researchers require advanced visualization techniques that support interactive exploration of multiple sequences and structural alignments. This paper introduces a suite of tools provided on the RCSB PDB research-focused web portal RCSB. org, tailor-made for efficient management, search, organization, and visualization of this burgeoning corpus of 3D macromolecular structure data.

3D visualization↗

Utilizing Amino Acid Composition and Entropy of Potential Open Reading Frames to Identify Protein-Coding Genes

One of the main steps in gene-finding in prokaryotes is determining which open reading frames encode for a protein, and which occur by chance alone. There are many different methods to differentiate the two; the most prevalent approach is using shared homology with a database of known genes. This method presents many pitfalls, most notably the catch that you only find genes that you have seen before. The four most popular prokaryotic gene-prediction programs (GeneMark, Glimmer, Prodigal, Phanotate) all use a protein-coding training model to predict protein-coding genes, with the latter three allowing for the training model to be created ab initio from the input genome. Different methods are available for creating the training model, and to increase the accuracy of such tools, we present here GOODORFS, a method for identifying protein-coding genes within a set of all possible open reading frames (ORFS). Our workflow begins with taking the amino acid frequencies of each ORF, calculating an entropy density profile (EDP), using KMeans to cluster the EDPs, and then selecting the cluster with the lowest variation as the coding ORFs. To test the efficacy of our method, we ran GOODORFS on 14,179 annotated phage genomes, and compared our results to the initial training-set creation step of four other similar methods (Glimmer, MED2, PHANOTATE, Prodigal). We found that GOODORFS was the most accurate (0.94) and had the best F1-score (0.85), while Glimmer had the highest precision (0.92) and PHANOTATE had the highest recall (0.96).

59 BASIC BIOLOGICAL SCIENCES↗

Characterization of the thermophilic xylanase Fsa02490Xyn from the hyperthermophile Fervidibacter sacchari belonging to glycoside hydrolase family 10

Fervidibacter sacchari is an aerobic hyperthermophile belonging to the phylum Armatimonadota that degrades a variety of polysaccharides. Its genome encodes 117 enzymes with one or more annotated glycoside hydrolase (GH) domain, but the roles of these putative GHs in polysaccharide catabolism are poorly defined. Here, we describe one F. sacchari enzyme encoding a GH10 domain, Fsa02490Xyn, that was previously shown to be active on Miscanthus, oat β-glucan, and beech-wood xylan, with optimal activity at 90-100 °C. We show that Fsa02490Xyn is also active on birch-wood xylan and gellan gum. The pH range on beech-wood xylan was 4.5 to 9.5 (pHopt 7.0-8.0). Fsa024940Xyn had a Km of 2.375 mm, Vmax of 1250 μm·min-1, and kcat/Km of 1.259 × 104 s-1·m-1 when using a para-nitrophenyl-?-xylobioside assay. A phylogenetic analysis of GH10 family enzymes revealed a large clade of enzymes from diverse members of the class Fervidibacteria, including Fsa02490Xyn and a second enzyme from F. sacchari, with apparent horizontal gene transfer within Fervidibacteria and between Fervidibacteria and thermophilic Bacillota. This study establishes Fsa02490Xyn as a hyperthermophilic GH10 enzyme with endo-β-1,4-xylanase activity and identifies a large clade of homologous GH10 enzymes within the class Fervidibacteria. Impact statement The depolymerization of xylan at high temperatures is important because this process limits the degradation of polysaccharides in nature and the synthesis of biofuels from plant wastes. Our study is also important because F. sacchari is one of only a few cultivated members of the Armatimonadota, which are polysaccharide-degradation specialists.

Armatimonadota↗

Automated Bacterial Identification and Morphological Feature Analysis in Low‐Dose Cryo‐EM Using YOLOv11

Bacteria rapidly adapt to environmental cues through morphological and ultrastructural changes that correlate with physiology and behavior. Cryogenic transmission electron microscopy (cryo‐TEM) can capture these phenotypic changes in near‐native, vitrified states, but manual analysis of low‐dose micrographs is labor intensive and limits throughput. Here, we present an end‐to‐end workflow that combines low‐dose cryo‐TEM imaging with a YOLOv11‐based instance‐segmentation model to automatically identify bacteria and quantify key structural features directly from the micrographs. This workflow enables (i) robust bacterial localization and counting from low‐magnification atlas/montage images, (ii) automated measurements of cell‐envelope (outer–inner membrane) thickness and anisotropy from higher‐magnification views, and (iii) detection and quantification of bacteria–flagella interactions, including overlap length and curvature metrics for interacting versus noninteracting flagella. Using Pantoea sp. YR343 grown under distinct media conditions, we show that the automated measurements agree with manual annotations while substantially reducing analysis time. Together, these tools provide a practical framework for scalable bacterial identification and quantitative phenotyping in low‐dose cryo‐TEM datasets and establish a foundation for extending cryo‐TEM image analysis toward higher‐throughput studies of microbial heterogeneity and biointerfaces.

YOLOv11↗

Oleaginous Yeast Biology Elucidated With Comparative Transcriptomics

ABSTRACT Extremophilic yeasts have favorable metabolic and tolerance traits for biomanufacturing‐ like lipid biosynthesis, flavinogenesis, and halotolerance – yet the connection between these favorable phenotypes and strain genotype is not well understood. To this end, this study compares the phenotypes and gene expression patterns of biotechnologically relevant yeasts Yarrowia lipolytica , Debaryomyces hansenii , and Debaryomyces subglobosus grown under nitrogen starvation, iron starvation, and salt stress. To analyze the large data set across species and conditions, two approaches were used: a “network‐first” approach where a generalized metabolic network serves as a scaffold for mapping genes and a “cluster‐first” approach where unsupervised machine learning co‐expression analysis clusters genes. Both approaches provide insight into strain behavior. The network‐first approach corroborates that Yarrowia upregulates lipid biosynthesis during nitrogen starvation and provides new evidence that riboflavin overproduction in Debaryomyces yeasts is overflow metabolism that is routed to flavin cofactor production under salt stress. The cluster‐first approach does not rely on annotation; therefore, the coexpression analysis can identify known and novel genes involved in stress responses, mainly transcription factors and transporters. Therefore, this work links the genotype to the phenotype of biotechnologically relevant yeasts and demonstrates the utility of complementary computational approaches to gain insight from transcriptomics data across species and conditions.

Weintraub, Sarah J. [Department of Bioinformatics ↗

Artificial intelligence to unlock real-world evidence in clinical oncology: A primer on recent advances

Purpose: Real world evidence is crucial to understanding the diffusion of new oncologic therapies, monitoring cancer outcomes, and detecting unexpected toxicities. In practice, real world evidence is challenging to collect rapidly and comprehensively, often requiring expensive and time-consuming manual case-finding and annotation of clinical text. In this Review, we summarise recent developments in the use of artificial intelligence to collect and analyze real world evidence in oncology. Methods: We performed a narrative review of the major current trends and recent literature in artificial intelligence applications in oncology. Results: Artificial intelligence (AI) approaches are increasingly used to efficiently phenotype patients and tumors at large scale. These tools also may provide novel biological insights and improve risk prediction through multimodal integration of radiographic, pathological, and genomic datasets. Custom language processing pipelines and large language models hold great promise for clinical prediction and phenotyping. Conclusions: Despite rapid advances, continued progress in computation, generalizability, interpretability, and reliability as well as prospective validation are needed to integrate AI approaches into routine clinical care and real-time monitoring of novel therapies.

60 APPLIED LIFE SCIENCES↗

Perimatrix of middle ear cholesteatoma: A granulation tissue with a specific transcriptomic signature

Objectives/Hypothesis: To establish comprehensive transcriptomic profiles of cholesteatoma perimatrix tissue and granulation tissue from chronic otitis media (COM) that did not develop cholesteatoma, which can indicate molecular pathways involved in the cholesteatoma perimatrix pathology and invasiveness. Study Design: Retrospective Case Series. Methods: Transcriptome data were obtained from cholesteatoma perimatrix tissue and COM granulation tissue by an Illumina iScan microarray. Differentially expressed genes (DEGs) were subsequently analyzed using both bioinformatical functional annotation and network analysis. Expression of candidate genes (MMP9 and LCN2) was validated by quantitative reverse transcription-polymerase chain reaction (qRT-PCR) on a larger group of samples. Results: Analysis of the transcriptome led to the identification of 169 differentially expressed genes between investigated tissues. Bioinformatic analysis suggested that most significant biological processes involving DEGs were previously described in cholesteatoma pathology. Network analysis identified ERBB2, TFAP2A, and TP63 as major hubs of the DEGs molecular network. Furthermore, it was observed that the cellular component most significantly enriched in DEGs was extracellular space containing 47 DEGs. Using qRT-PCR, it was confirmed that mRNA levels of the major extracellular hub (MMP9) are increased, whereas its interacting molecule (LCN2) mRNA levels were decreased in cholesteatoma perimatrix tissue compared to COM granulation tissue. Conclusions: The current study approach offers an overall look at molecular mechanisms that describe the cholesteatoma entity by focusing exclusively on the perimatrix processes in comparison to COM granulation tissue. The observed differences in gene expression between cholesteatoma perimatrix and COM granulation tissue could suggest novel markers potentially influenced by the perimatrix–matrix molecular interplay, which is not present in COM without cholesteatoma. Level of Evidence: NA. Laryngoscope, 130:E220–E227, 2020. © 2019 The American Laryngological, Rhinological and Otological Society, Inc.

60 APPLIED LIFE SCIENCES↗

An optimized ChIP-Seq framework for profiling histone modifications in Chromochloris zofingiensis

The eukaryotic green alga Chromochloris zofingiensis is a reference organism for studying carbon partitioning and a promising candidate for the production of biofuel precursors. Recent transcriptome profiling transformed our understanding of its biology and generally algal biology, but epigenetic regulation remains understudied and represents a fundamental gap in our understanding of algal gene expression. Chromatin immunoprecipitation followed by deep sequencing (ChIP-Seq) is a powerful tool for the discovery of such mechanisms, by identifying genome-wide histone modification patterns and transcription factor-binding sites alike. Here, we established a ChIP-Seq framework for Chr. zofingiensis yielding over 20 million high-quality reads per sample. The most critical steps in a ChIP experiment were optimized, including DNA shearing to obtain an average DNA fragment size of 250 bp and assessment of the recommended formaldehyde concentration for optimal DNA-protein cross-linking. We used this ChIP-Seq framework to generate a genome-wide map of the H3K4me3 distribution pattern and to integrate these data with matching RNA-Seq data. In line with observations from other organisms, H3K4me3 marks predominantly transcription start sites of genes. Our H3K4me3 ChIP-Seq data will pave the way for improved genome structural annotation in the emerging reference alga Chr. zofingiensis.

59 BASIC BIOLOGICAL SCIENCES↗

First Plant Cell Atlas symposium report

The Plant Cell Atlas (PCA) community hosted a virtual symposium on December 9 and 10, 2021 on single cell and spatial omics technologies. The conference gathered almost 500 academic, industry, and government leaders to identify the needs and directions of the PCA community and to explore how establishing a data synthesis center would address these needs and accelerate progress. This report details the presentations and discussions focused on the possibility of a data synthesis center for a PCA and the expected impacts of such a center on advancing science and technology globally. Community discussions focused on topics such as data analysis tools and annotation standards; computational expertise and cyber-infrastructure; modes of community organization and engagement; methods for ensuring a broad reach in the PCA community; recruitment, training, and nurturing of new talent; and the overall impact of the PCA initiative. These targeted discussions facilitated dialogue among the participants to gauge whether PCA might be a vehicle for formulating a data synthesis center. The conversations also explored how online tools can be leveraged to help broaden the reach of the PCA (i.e., online contests, virtual networking, and social media stakeholder engagement) and decrease costs of conducting research (e.g., virtual REU opportunities). Major recommendations for the future of the PCA included establishing standards, creating dashboards for easy and intuitive access to data, and engaging with a broad community of stakeholders. The discussions also identified the following as being essential to the PCA's success: identifying homologous cell-type markers and their biocuration, publishing datasets and computational pipelines, utilizing online tools for communication (such as Slack), and user-friendly data visualization and data sharing. In conclusion, the development of a data synthesis center will help the PCA community achieve these goals by providing a centralized repository for existing and new data, a platform for sharing tools, and new analytical approaches through collaborative, multidisciplinary efforts. A data synthesis center will help the PCA reach milestones, such as community-supported data evaluation metrics, accelerating plant research necessary for human and environmental health.

59 BASIC BIOLOGICAL SCIENCES↗

Ca X ML: Chemistry‐informed machine learning explains mutual changes between protein conformations and calcium ions in calcium‐binding proteins using structural and topological features

Proteins' flexibility is a feature in communicating changes in cell signaling instigated by binding with secondary messengers, such as calcium ions, associated with the coordination of muscle contraction, neurotransmitter release, and gene expression. When binding with the disordered parts of a protein, calcium ions must balance their charge states with the shape of calcium-binding proteins and their versatile pool of partners depending on the circumstances they transmit. Accurately determining the ionic charges of those ions is essential for understanding their role in such processes. However, it is unclear whether the limited experimental data available can be effectively used to train models to accurately predict the charges of calcium-binding protein variants. Here, we developed a chemistry-informed, machine-learning algorithm that implements a game theoretic approach to explain the output of a machine-learning model without the prerequisite of an excessively large database for high-performance prediction of atomic charges. We used the ab initio electronic structure data representing calcium ions and the structures of the disordered segments of calcium-binding peptides with surrounding water molecules to train several explainable models. Network theory was used to extract the topological features of atomic interactions in the structurally complex data dictated by the coordination chemistry of a calcium ion, a potent indicator of its charge state in protein. Our design created a computational tool of Ca X ML, which provided a framework of explainable machine learning model to annotate ionic charges of calcium ions in calcium-binding proteins in response to the chemical changes in an environment. Our framework will provide new insights into protein design for engineering functionality based on the limited size of scientific data in a genome space.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Assembly and comparative genome analysis of a Patagonian Aureobasidium pullulans isolate reveals unexpected intraspecific variation

Aureobasidium pullulans is a yeast-like fungus with remarkable phenotypic plasticity widely studied for its importance for the pharmaceutical and food industries. So far, genomic studies with strains from all over the world suggest they constitute a genetically unstructured population, with no association by habitat. However, the mechanisms by which this genome supports so many phenotypic permutations are still poorly understood. Recent works have shown the importance of sequencing yeast genomes from extreme environments to increase the repertoire of phenotypic diversity of unconventional yeasts. In this study, we present the genomic draft of A. pullulans strain from a Patagonian yeast diversity hotspot, re-evaluate its taxonomic classification based on taxogenomic approaches, and annotate its genome with high-depth transcriptomic data. Here, our analysis suggests this isolate could be considered a novel variant at an early stage of the speciation process. The discovery of divergent strains in a genomically homogeneous group, such as A. pullulans, can be valuable in understanding the evolution of the species. The identification and characterization of new variants will not only allow finding unique traits of biotechnological importance, but also optimize the choice of strains whose phenotypes will be characterized, providing new elements to explore questions about plasticity and adaptation.

59 BASIC BIOLOGICAL SCIENCES↗

A practical approach to using the Genomic Standards Consortium MIxS reporting standard for comparative genomics and metagenomics

Comparative analysis of (meta)genomes necessitates aggregation, integration, and synthesis of well-annotated data using standards. The Genomic Standards Consortium (GSC) collaborates with the research community to develop and maintain the Minimal Information about any (x) Sequence (MIxS) reporting standard for genomic data. To facilitate use of the GSC’s MIxS reporting standard, we provide a description of the structure and terminology, how to navigate ontologies for required terms in MIxS, and demonstrate practical usage through a soil metagenome example.

standards, metadata, genome, metagenome, schema, v↗

Privacy-Preserving Knowledge Transfer with Bootstrap Aggregation of Teacher Ensembles

There is a need to transfer knowledge among institutions and organizations to save effort in annotation and labeling or in enhancing task performance. However, knowledge transfer is difficult because of restrictions that are in place to ensure data security and privacy. Institutions are not allowed to exchange data or perform any activity that may expose personal information. With the leverage of a differential privacy algorithm in a high-performance computing environment, we propose a new training protocol, Bootstrap Aggregation of Teacher Ensembles (BATE), which is applicable to various types of machine learning models. The BATE algorithm is based on and provides enhancements to the PATE algorithm, maintaining competitive task performance scores on complex datasets with underrepresented class labels.We conducted a proof-of-the-concept study of the information extraction from cancer pathology report data from four cancer registries and performed comparisons between four scenarios: no collaboration, no privacy-preserving collaboration, the PATE algorithm, and the proposed BATE algorithm. The results showed that the BATE algorithm maintained competitive macro-averaged F1 scores, demonstrating that the suggested algorithm is an effective yet privacy-preserving method for machine learning and deep learning solutions.

Yoon, Hong-Jun↗

The Kokkos OpenMPTarget Backend: Implementation and Lessons Learned

As the supercomputing landscape diversifies, solutions such as Kokkos to write vendor agnostic applications and libraries have risen in popularity. Kokkos provides a programming model designed for performance portability, which allows developers to write a single source implementation that can run efficiently on various architectures. At its heart, Kokkos maps parallel algorithms to architecture and vendor specific backends written in lower level programming models such as CUDA and HIP. Another approach to writing vendor agnostic parallel code is using OpenMP’s directives based approach, which lets developers annotate code to express parallelism. It is implemented at the compiler level and is supported by all major high performance computing vendors, as well as the primary Open Source toolchains GNU and LLVM. Since its inception, Kokkos has used OpenMP to parallelize on CPU architectures. In this paper, we explore leveraging OpenMP for a GPU backend and discuss the challenges we encountered when mapping the Kokkos APIs and semantics to OpenMP target constructs. As an exemplar workload we chose a simple conjugate gradient solver for sparse matrices. We find that performance on NVIDIA and AMD GPUs varies widely based on details of the implementation strategy and the chosen compiler. Furthermore, the performance of the OpenMP implementations decreases with increasing complexity of the investigated algorithms.

Gayatri, Rahulkumar↗

Automatic Digitization and Orientation of Scanned Mesh Data for Floor Plan and 3D Model Generation

This paper describes a novel approach for generating accurate floor plans and 3D models of building interiors using scanned mesh data. Unlike previous methods, which begin with a high resolution point cloud from a laser range-finder, our approach begins with triangle mesh data, as from a Microsoft HoloLens. It generates two types of floor plans, a “pen-and-ink” style that preserves details and a drafting-style that reduces clutter. It processes the 3D model for use in applications by aligning it with coordinate axes, annotating important objects, dividing it into stories, and removing the ceiling. Its performance is evaluated on commercial and residential buildings, with experiments to assess quality and dimensional accuracy. Our approach demonstrates promising potential for automatic digitization and orientation of scanned mesh data, enabling floor plan and 3D model generation in various applications such as navigation, interior design, furniture placement, facilities management, building construction, and HVAC design.

Sharma, Ritesh↗

Domain-Specific Type-Safe APIs for Hierarchical Scientific Data with Modern C++

General-purpose library application programming interfaces (APIs) for self-describing hierarchical scientific data storage, such as the HDF5 and NetCDF libraries, are traditionally of runtime nature. Runtime errors for entry existence and data types are typically caught later in the development process of higher-level application-specific APIs. In this paper, we propose exploiting modern C++ metaprogramming features to add compile-time type-safety to improve the interaction with a well-defined metadata-rich scientific schema in domain-specific hierarchical datasets. We tackle two aspects of common use: (i) direct data access, (ii) flexible “in-memory” index models for efficient search and data processing. The proposed APIs use C++17’s template type auto deduction features, C++11’s enum class for type-safety and C-style preprocessor macros for generative templated code. We showcase the pros and cons of our initial work on the standard NeXus schema used for annotating and storing experimental neutron scattering data at several facilities around the world on top of HDF5. Extendable compile-time type-safe APIs are a desirable feature that could be indexed by any modern integrated development environment (IDE). Hence, such APIs can help ease the learning curve for domain scientists using a less error-prone software interaction to enhance the findability of their data without resorting to a domain-specific language (DSL).

Godoy, William↗