Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “annotations”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 613 records · Page 34

MIBiG 4.0: advancing biosynthetic gene cluster curation through global collaboration

Specialized or secondary metabolites are small molecules of biological origin, often showing potent biological activities with applications in agriculture, engineering and medicine. Usually, the biosynthesis of these natural products is governed by sets of co-regulated and physically clustered genes known as biosynthetic gene clusters (BGCs). To share information about BGCs in a standardized and machine-readable way, the Minimum Information about a Biosynthetic Gene cluster (MIBiG) data standard and repository was initiated in 2015. Since its conception, MIBiG has been regularly updated to expand data coverage and remain up to date with innovations in natural product research. Here, we describe MIBiG version 4.0, an extensive update to the data repository and the underlying data standard. In a massive community annotation effort, 267 contributors performed 8304 edits, creating 557 new entries and modifying 590 existing entries, resulting in a new total of 3059 curated entries in MIBiG. Particular attention was paid to ensuring high data quality, with automated data validation using a newly developed custom submission portal prototype, paired with a novel peer-reviewing model. MIBiG 4.0 also takes steps towards a rolling release model and a broader involvement of the scientific community. MIBiG 4.0 is accessible online at https://mibig.secondarymetabolites.org/.

59 BASIC BIOLOGICAL SCIENCES↗

BioPortal: an open community resource for sharing, searching, and utilizing biomedical ontologies

Abstract BioPortal (https://bioportal.bioontology.org) is the world’s most comprehensive repository of biomedical ontologies. It provides infrastructure for finding, sharing, searching, and utilizing biomedical ontologies. Launched in 2005, BioPortal now includes 1549 ontologies (1182 of them public). Its open, freely accessible website enables anyone (i) to browse the ontology library, (ii) to search for terms across ontologies, (iii) to browse mappings between terms, (iv) to see popularity ratings and recommendations on which ontologies are most relevant to their use cases, (v) to annotate text with ontology terms, (vi) to submit an ontology, and (vii) to request ontology changes. The library of ontologies can be accessed programmatically via a REST application programming interface (API). Recent enhancements include a BioPortal knowledge graph that integrates knowledge from multiple ontologies; a unified data model for interoperability with other knowledge sources; ontology popularity ratings and recommendations for relevant ontologies; and the ability to request ontology changes via a simple user interface that automatically converts user change requests to GitHub Pull Requests that specify the edits that will be made to the ontology upon approval.

Vendetti, Jennifer↗

metagRoot: a comprehensive database of protein families associated with plant root microbiomes

The plant root microbiome is vital in plant health, nutrient uptake, and environmental resilience. To explore and harness this diversity, we present metagRoot, a specialized and enriched database focused on the protein families of the plant root microbiome. MetagRoot integrates metagenomic, metatranscriptomic, and reference genome-derived protein data to characterize 71 091 enriched protein families, each containing at least 100 sequences. These families are annotated with multiple sequence alignments, CRISPR elements, hidden Markov models, taxonomic and functional classifications, ecosystem and geolocation metadata, and predicted 3D structures using AlphaFold2. MetagRoot is a powerful tool for decoding the molecular landscape of root-associated microbial communities and advancing microbiome-informed agricultural practices by enriching protein family information with ecological and structural context. The database is available at https://pavlopoulos-lab.org/metagroot/ or https://www.metagroot.org.

Chasapi, Maria N↗

Metabolic Source Isotopic Pair Labeling and Genome-Wide Association Are Complementary Tools for the Identification of Metabolite-Gene Associations in Plants

The optimal extraction of information from untargeted metabolomics analyses is a continuing challenge. Here, we describe an approach that combines stable isotope labeling, liquid chromatography– mass spectrometry (LC–MS), and a computational pipeline to automatically identify metabolites produced from a selected metabolic precursor. Here, we identified the subset of the soluble metabolome generated from phenylalanine (Phe) in Arabidopsis thaliana, which we refer to as the Phe-derived metabolome (FDM) In addition to identifying Phe-derived metabolites present in a single wild-type reference accession, the FDM was established in nine enzymatic and regulatory mutants in the phenylpropanoid pathway. To identify genes associated with variation in Phe-derived metabolites in Arabidopsis, MS features collected by untargeted metabolite profiling of an Arabidopsis diversity panel were retrospectively annotated to the FDM and natural genetic variants responsible for differences in accumulation of FDM features were identified by genome-wide association. Large differences in Phe-derived metabolite accumulation and presence/absence variation of abundant metabolites were observed in the nine mutants as well as between accessions from the diversity panel. Many Phe-derived metabolites that accumulated in mutants also accumulated in non-Col-0 accessions and was associated to genes with known or suspected functions in the phenylpropanoid pathway as well as genes with no known functions. Overall, we show that cataloguing a biochemical pathway’s products through isotopic labeling across genetic variants can substantially contribute to the identification of metabolites and genes associated with their biosynthesis.

09 BIOMASS FUELS↗

Targeted genetic manipulation and yeast-like evolutionary genomics in the green alga Auxenochlorella

Auxenochlorella spp. are diploid oleaginous green algae whose streamlined genomes can be readily manipulated by homologous recombination, making them highly amenable to discovery research and bioengineering. Vegetatively diploid organisms experience specific evolutionary phenomena, including allodiploid hybridization, mitotic recombination, loss-of-heterozygosity, and aneuploidy; however, studies of these forces have largely focused on yeasts. Here, we present a telomere-to-telomere phased diploid genome assembly of Auxenochlorella UTEX 250-A (haploid length 22 Mb) and introduce a genetic toolkit for site-specific manipulation of the nuclear genome in multiple strains, featuring several selectable markers, inducible promoters, and fluorescent reporters for protein localization. UTEX 250-A is an allodiploid hybrid of Auxenochlorella protothecoides and Auxenochlorella symbiontica, two species differentiated by extensive chromosomal rearrangements. UTEX 250-A haplotypes are a mosaic of each parental species following mitotic recombination, and two chromosomes are trisomic. Loss-of-heterozygosity events are pervasive across Auxenochlorella and can evolve rapidly in the laboratory. High-quality structural annotation yielded ∼7,500 genes per haplotype. Auxenochlorella have experienced gene family loss and reduction, including core photosynthesis genes, and exhibit periodic adenine and cytosine methylation at promoters and gene bodies, respectively. Approximately 10% of genes, especially those involved in DNA repair and sex, overlap antisense long noncoding RNAs, which may participate in a regulatory mechanism. We demonstrate the utility of Auxenochlorella for fundamental research by knockout of a chlorophyll biosynthesis enzyme, and confirm one trisomy by allele-specific transformation. These results demonstrate the generality of several evolutionary forces associated with vegetative diploidy and provide a foundation for the use of Auxenochlorella as a reference organism.

CHL27↗

Hydroxycinnamaldehyde-derived benzofuran components in lignins

Abstract Lignin is an abundant polymer in plant secondary cell walls. Prototypical lignins derive from the polymerization of monolignols (hydroxycinnamyl alcohols), mainly coniferyl and sinapyl alcohol, via combinatorial radical coupling reactions and primarily via the endwise coupling of a monomer with the phenolic end of the growing polymer. Hydroxycinnamaldehyde units have long been recognized as minor components of lignins. In plants deficient in cinnamyl alcohol dehydrogenase, the last enzyme in the monolignol biosynthesis pathway that reduces hydroxycinnamaldehydes to monolignols, chain-incorporated aldehyde unit levels are elevated. The nature and relative levels of aldehyde components in lignins can be determined from their distinct and dispersed correlations in 2D 1H–13C-correlated nuclear magnetic resonance (NMR) spectra. We recently became aware of aldehyde NMR peaks, well resolved from others, that had been overlooked. NMR of isolated low-molecular-weight oligomers from biomimetic radical coupling reactions involving coniferaldehyde revealed that the correlation peaks belonged to hydroxycinnamaldehyde-derived benzofuran moieties. Coniferaldehyde 8-5-coupling initially produces the expected phenylcoumaran structures, but the derived phenolic radicals undergo preferential disproportionation rather than radical coupling to extend the growing polymer. As a result, the hydroxycinnamaldehyde-derived phenylcoumaran units are difficult to detect in lignins, but the benzofurans are now readily observed by their distinct and dispersed correlations in the aldehyde region of NMR spectra from any lignin or monolignol dehydrogenation polymer. Hydroxycinnamaldehydes that are coupled to coniferaldehyde can be distinguished from those coupled with a generic guaiacyl end-unit. These benzofuran peaks may now be annotated and reported and their structural ramifications further studied.

59 BASIC BIOLOGICAL SCIENCES↗

The small protein SbtC is a functional component of the CO 2 concentrating mechanism in Synechocystis sp. PCC 6803

Oxygenic phototrophs fix CO 2 via the enzyme ribulose-1,5-bisphosphate carboxylase/oxygenase (RubisCO), which shows relatively low CO 2 affinity and specificity. To circumvent low and fluctuating CO 2 concentrations in aquatic systems, cyanobacteria and algae have evolved sophisticated inorganic carbon (Ci) concentrating mechanisms (CCMs). Bicarbonate transporters such as SbtA play a crucial role in the cyanobacterial CCM and hence display multiple layers of tight regulation. Control of sbtA gene expression and corresponding transporter activity involves the PII-like protein SbtB, whose gene is frequently co-transcribed with sbtA. A previously non-annotated gene located upstream of the sbtAB operon in the model Synechocystis sp. PCC 6803 encodes the small protein SbtC, composed of 80 amino acids. Presence of SbtC was confirmed by immunoblotting of the sbtC-coding sequence fused to a Flag-tag. Similar to sbtAB , transcription of the sbtC locus is induced by low CO 2 availability; however, it is controlled independently. Mutation of the sbtC locus in a wild-type background produced only a mild phenotype, even under low CO 2 , but impaired diurnal growth resembled that of the mutant ΔsbtB . Biochemical analysis indicated a trimeric SbtABC complex in the membrane. Bicarbonate leakage from cells was strongly elevated when either sbtB or sbtC was deleted from recombinant Synechocystis strains harboring only SbtA as single Ci uptake system. Here, our results provide evidence that SbtC contributes to the formation of the SbtAB complex, thereby regulating bicarbonate exchange at the cytoplasmic membrane. Well-conserved SbtC-like proteins encoded in the neighborhood of sbtAB exist in many cyanobacterial genomes, pointing toward an important role in the cyanobacterial CCM.

Walke, Peter [Univ. of Rostock (Germany)] (ORCID:0↗

Reekeekee- and roodoodooviruses, two different Microviridae clades constituted by the smallest DNA phages

Small circular single-stranded DNA viruses of the Microviridae family are both prevalent and diverse in all ecosystems. They usually harbor a genome between 4.3 and 6.3 kb, with a microvirus recently isolated from a marine Alphaproteobacteria being the smallest known genome of a DNA phage (4.248 kb). A subfamily, Amoyvirinae, has been proposed to classify this virus and other related small Alphaproteobacteria-infecting phages. Here, we report the discovery, in meta-omics data sets from various aquatic ecosystems, of sixteen complete microvirus genomes significantly smaller (2.991–3.692 kb) than known ones. Phylogenetic analysis reveals that these sixteen genomes represent two related, yet distinct and diverse, novel groups of microviruses—amoyviruses being their closest known relatives. We propose that these small microviruses are members of two tentatively named subfamilies Reekeekeevirinae and Roodoodoovirinae. As known microvirus genomes encode many overlapping and overprinted genes that are not identified by gene prediction software, we developed a new methodology to identify all genes based on protein conservation, amino acid composition, and selection pressure estimations. Surprisingly, only four to five genes could be identified per genome, with the number of overprinted genes lower than that in phiX174. These small genomes thus tend to have both a lower number of genes and a shorter length for each gene, leaving no place for variable gene regions that could harbor overprinted genes. Even more surprisingly, these two Microviridae groups had specific and different gene content, and major differences in their conserved protein sequences, highlighting that these two related groups of small genome microviruses use very different strategies to fulfill their lifecycle with such a small number of genes. The discovery of these genomes and the detailed prediction and annotation of their genome content expand our understanding of ssDNA phages in nature and are further evidence that these viruses have explored a wide range of possibilities during their long evolution.

59 BASIC BIOLOGICAL SCIENCES↗

Genome Assembly and Transcriptome of Colletotrichum sublineola CsGL1, a New Resource to Study Anthracnose Disease in Sorghum

Colletotrichum species are globally distributed and well known as members of a destructive phytopathogenic genus, causing the anthracnose disease in a wide variety of crops and fruits. Colletotrichum sublineola is the causal agent of the anthracnose disease in sorghum, causing losses of up to 50% in yield. Here, we used PacBio sequencing combined with RNA-seq to generate a chromosome-level assembly and annotation of the Colletotrichum sublineola strain CsGL1. [Formula: see text] Copyright © 2021 The Author(s). This is an open access article distributed under the CC BY-NC-ND 4.0 International license .

Biochemistry & Molecular Biology↗

Community-Driven Metadata Standards for Agricultural Microbiome Research

Accelerating the pace of microbiome science to enhance crop productivity and agroecosystem health will require transdisciplinary studies, comparisons among datasets, and synthetic analyses of research from diverse crop management contexts. However, despite the widespread availability of crop-associated microbiome data, variation in field sampling and laboratory processing methodologies, as well as metadata collection and reporting, significantly constrains the potential for integrative and comparative analyses. Here we discuss the need for agriculture-specific metadata standards for microbiome research, and propose a list of “required” and “desirable” metadata categories and ontologies essential to be included in a future minimum information metadata standards checklist for describing agricultural microbiome studies. We begin by briefly reviewing existing metadata standards relevant to agricultural microbiome research, and describe ongoing efforts to enhance the potential for integration of data across research studies. Our goal is not to delineate a fixed list of metadata requirements. Instead, we hope to advance the field by providing a starting point for discussion, and inspire researchers to adopt standardized procedures for collecting and reporting consistent and well-annotated metadata for agricultural microbiome research.

59 BASIC BIOLOGICAL SCIENCES↗

PDB‐101: Molecular Explorations through Biology and Medicine

PDB‐101 is an online portal for teachers, students, and the general public to promote exploration of the structural biology of proteins and nucleic acids ( pdb101.rcsb.org ). Learning about the diverse shapes and functions of these biological macromolecules helps to understand all aspects of biomedicine and agriculture, from protein synthesis to health and disease to biological energy. Why PDB‐101? Researchers around the world are studying these molecules at the atomic level. These 3D structures are freely available at the Protein Data Bank (PDB), the central storehouse of biomolecular structures. This website builds introductory materials to help beginners get started in the basics of biomolecular structure and function (“101”, as in an entry level course) as well as resources for extended learning. Since 2011, PDB‐101 has been developed by the RCSB PDB , a global resource for the advancement of research and education in biology and medicine. Along with our Worldwide PDB collaborators, RCSB PDB curates, annotates, and makes publicly available the PDB data deposited by scientists around the globe. The RCSB PDB then provides a window to these data through a rich online resource with powerful searching, reporting, and visualization tools for researchers. This information is then streamlined for students and teachers at PDB‐101. Features include the ongoing Molecule of the Month series, educational materials such as paper models, posters, molecular animations, educational curricula and more. The section “Guide to Understanding PDB Data” is a primer for detailed PDB‐specific information: PDB Data, Visualizing Structures, Reading Coordinate Files, scientific methods for structure determination, and more. PDB‐101 also runs annual Video Challenges for high school students. Participants create short videos that tell molecular stories that connect structural biology and medicine. Previous topics have included HIV/AIDS, diabetes, and antimicrobial resistance. The 2022 challenge will focus on Molecular Mechanisms of Cancer. PDB‐101 activities are evaluated using user surveys, feedback from in‐person activities, and website analytics. In 2020, PDB‐101 hosted >850,000 users and >2.6 million page views.

Zardecki, Christine↗

Quantifying uncertainties and correlations in the nuclear-matter equation of state

We perform statistically rigorous uncertainty quantification (UQ) for chiral effective field theory (χ EFT) applied to infinite nuclear matter up to twice nuclear saturation density. The equation of state (EOS) is based on high-order many-body perturbation theory calculations with nucleon-nucleon and three-nucleon interactions up to fourth order in the χ EFT expansion. From these calculations our newly developed Bayesian machine-learning approach extracts the size and smoothness properties of the correlated EFT truncation error. Furthermore, we then propose a novel extension that uses multitask machine learning to reveal correlations between the EOS at different proton fractions. The inferred in-medium χ EFT breakdown scale in pure neutron matter and symmetric nuclear matter is consistent with that from free-space nucleon-nucleon scattering. These significant advances allow us to provide posterior distributions for the nuclear saturation point and propagate theoretical uncertainties to derived quantities: the pressure and incompressibility of symmetric nuclear matter, the nuclear symmetry energy, and its derivative. Our results, which are validated by statistical diagnostics, demonstrate that an understanding of truncation-error correlations between different densities and different observables is crucial for reliable UQ. The methods developed here are publicly available as annotated Jupyter notebooks.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

ROI-Finder : machine learning to guide region-of-interest scanning for X-ray fluorescence microscopy

The microscopy research at the Bionanoprobe (currently at beamline 9-ID and later 2-ID after APS-U) of Argonne National Laboratory focuses on applying synchrotron X-ray fluorescence (XRF) techniques to obtain trace elemental mappings of cryogenic biological samples to gain insights about their role in critical biological activities. The elemental mappings and the morphological aspects of the biological samples, in this instance, the bacterium Escherichia coli ( E. Coli ), also serve as label-free biological fingerprints to identify E. coli cells that have been treated differently. The key limitations of achieving good identification performance are the extraction of cells from raw XRF measurements via binary conversion, definition of features, noise floor and proportion of cells treated differently in the measurement. Automating cell extraction from raw XRF measurements across different types of chemical treatment and the implementation of machine-learning models to distinguish cells from the background and their differing treatments are described. Principal components are calculated from domain knowledge specific features and clustered to distinguish healthy and poisoned cells from the background without manual annotation. The cells are ranked via fuzzy clustering to recommend regions of interest for automated experimentation. The effects of dwell time and the amount of data required on the usability of the software are also discussed.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

RCSB Protein Data Bank: supporting research and education worldwide through explorations of experimentally determined and computationally predicted atomic level 3D biostructures

The Protein Data Bank (PDB) was established as the first open-access digital data resource in biology and medicine in 1971 with seven X-ray crystal structures of proteins. Today, the PDB houses >210 000 experimentally determined, atomic level, 3D structures of proteins and nucleic acids as well as their complexes with one another and small molecules ( e.g. approved drugs, enzyme cofactors). These data provide insights into fundamental biology, biomedicine, bioenergy and biotechnology. They proved particularly important for understanding the SARS-CoV-2 global pandemic. The US-funded Research Collaboratory for Structural Bioinformatics Protein Data Bank (RCSB PDB) and other members of the Worldwide Protein Data Bank (wwPDB) partnership jointly manage the PDB archive and support >60 000 `data depositors' (structural biologists) around the world. wwPDB ensures the quality and integrity of the data in the ever-expanding PDB archive and supports global open access without limitations on data usage. The RCSB PDB research-focused web portal at https://www.rcsb.org/ (RCSB.org) supports millions of users worldwide, representing a broad range of expertise and interests. In addition to retrieving 3D structure data, PDB `data consumers' access comparative data and external annotations, such as information about disease-causing point mutations and genetic variations. RCSB.org also provides access to >1 000 000 computed structure models (CSMs) generated using artificial intelligence/machine-learning methods. To avoid doubt, the provenance and reliability of experimentally determined PDB structures and CSMs are identified. Related training materials are available to support users in their RCSB.org explorations.

59 BASIC BIOLOGICAL SCIENCES↗

QLiG: Query Like a Graph For Subgraph Matching

A graph is a natural and flexible modeling approach to represent entities and relationships between them in real-world. A Knowledge Graphs (KG) is a specialized graph with formal and structured representation of facts, relationships, annotated with semantic descriptions. Subgraph matching is one of the fundamental graph problems to identify relationships, interactions and activities of interest within a large graph. A query specification is a collection of abstract components, operations, and constraints to express a pattern. The specification can be implemented in different ways based on underlying data model. Various graph query specifications have been developed over the years and have led to the development of different open-sourced and vendor-specific query languages. Such specification are modeled as an extension of relational algebra used to develop relational query languages such as SQL. Such relational concepts do not inherently support graph queries. There is a need to represent graph queries in terms on graph-based components to expedite query construction by non-database experts. We present a graph-based query approach QLiG (pronounced cleeg), to perform subgraph matching in Labeled Property Graph. We present the query specifications, salient features, and a use case to show functional examples.

Purohit, Sumit↗

Indicator-directed Dynamic Power Management for Iterative Workloads on GPU-Accelerated Systems

Modern high-performance and warehouse computing centers show strong interest in minimizing system power consumption while satisfying customers’ quality of service (QoS). Dynamic voltage and frequency scaling (DVFS) is effective for achieving this goal. Nevertheless, automating the process online and making it transparent to users must address three major challenges: (1) Complexity — today’s hardware components (e.g., CPUs, GPUs, memory, network, etc.) can be configured in several or dozens of frequency/voltage states for satisfying divergent system demands. Given their combination and the emergence of heterogeneity, searching the optimal configuration in the design space online can be timing consuming. (2) QoS guarantee — user-defined objectives such as power constraint and performance target must be monitored, predicted and ensured at the best effort. (3) Adaptability — various known and unknown workloads run on systems. Workloads characteristics should be quickly determined and configurations dynamically adjusted in accord with workloads and QoS. In this work, we focus on applications exhibiting an interesting feature – iterative or periodic, which is common among conventional HPC and emerging machine learning workloads. We propose an online dynamic power-performance (ODPP) management framework to dynamically adjust GPU DVFS configurations to meet performance and power objectives and constraints, without any code annotation or intrusion. Particularly, ODPP extracts the performance and power indicators for applications from their resources utilization profiles in a short episode. It further automatically constructs an accurate model that infers from the indicators how the application's performance and power vary with GPU core and memory frequencies. Aided with the model, for both seen and unseen applications, ODPP can quickly determine the most appropriate DVFS configuration for their execution. We evaluate ODPP on an NVIDIA GPU using multiple exascale computing (ECP) and deep learning applications.

Zou, Pengfei↗

Expanding on the BRIAR Dataset: A Comprehensive Whole Body Biometric Recognition Resource at Extreme Distances and Real-World Scenarios (Collections 1-4)

The state-of-the-art in biometric recognition algorithms and operational systems has advanced quickly in recent years providing high accuracy and robustness in more challenging collection environments and consumer applications. However, the technology still suffers greatly when applied to non-conventional settings such as those seen when performing identification at extreme distances or from elevated cameras on buildings or mounted to UAVs. This paper summarizes an extension to the largest dataset currently focused on addressing these operational challenges, and describes its composition as well as methodologies of collection, curation, and annotation.

Cornett, David [ORNL] (ORCID:0000000222910860)↗

DOC-DICAM: Domain Aware One Class Defect Identification in Composite Aerostructure Material

Fiber-reinforced composites are a common material used in the design of aircraft structures due to their good tensile strength and resistance to compression. During the manufacturing process, these structures are thoroughly inspected for flaws and defects to ensure structural integrity during commercial use. Non-destructive testing (NDT) is a collection of inspection methods that allow inspectors to evaluate material without altering it. Due to the high safety standards in aerospace manufacturing, the NDT process is done manually and can be a significant bottleneck in the development workflow. In this paper, we develop an AI-based assistance tool to drastically reduce inspection time. Typical AI workflows require large amounts of annotated data, but defects rarely occur resulting in strong class imbalance. To overcome this, we formulate the problem of defect identification as an anomaly detection task in which our primary focus is learning non-defect characteristics. To do this, we develop a multi-task self-supervised learning framework that embeds problem specific domain knowledge into the deep learning model. We verify our method using fuselage data generated in a production environment. As a result, we show that our method can effectively identify defects and requires minimal training and inference time.

anomaly detection↗