Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Sequence Annotation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Identification of transcribed sequences in Arabidopsis thaliana by using high-resolution genome tiling arrays

Using a maskless photolithography method, we produced DNA oligonucleotide microarrays with probe sequences tiled throughout the genome of the plant Arabidopsis thaliana. RNA expression was determined for the complete nuclear, mitochondrial, and chloroplast genomes by tiling 5 million 36-mer probes. These probes were hybridized to labeled mRNA isolated from liquid grown T87 cells, an undifferentiated Arabidopsis cell culture line. Transcripts were detected from at least 60% of the nearly 26,330 annotated genes, which included 151 predicted genes that were not identified previously by a similar genome-wide hybridization study on four different cell lines. In comparison with previously published results with 25-mer tiling arrays produced by chromium masking-based photolithography technique, 36-mer oligonucleotide probes were found to be more useful in identifying intron-exon boundaries. Using two-dimensional HPLC tandem mass spectrometry, a small-scale proteomic analysis was performed with the same cells. A large amount of strongly hybridizing RNA was found in regions "antisense" to known genes. Similarity of antisense activities between the 25-mer and 36-mer data sets suggests that it is a reproducible and inherent property of the experiments. Transcription activities were also detected for many of the intergenic regions and the small RNAs, including tRNA, small nuclear RNA, small nucleolar RNA, and microRNA. Expression of tRNAs correlates with genome-wide amino acid usage.

Arabidopsis/genetics

Genome collection processing for “Conserved upper thermal limits and small safety margins in soil copiotrophic bacteria”

We extracted the genomic DNA of 400 randomly selected isolates using a Quick-DNA Microprep Kit (Zymo Research D3020) according to the manufacturer’s protocol. We then submitted the extracted gDNA samples for short-read Illumina sequencing (200 Mbp) at SeqCoast Genomics (Portsmouth, NH, USA). After preprocessing the sequences using Trimmommatic (Bolger et al. 2014), we assembled the genomes using SPADES (Bankevich et al. 2012) and checked the quality of each assembly using QUAST (Gurevich et al. 2013). We processed the genome assemblies using a KBase (v1.4.0) pipeline (Allen et al. 2017; Arkin et al. 2018). Briefly, we used DRAM (v0.1.2) with default settings to annotate the genome assemblies. We then evaluated genome quality and possible contamination levels using CheckM (v1.0.18) (Parks et al. 2015) and retained genomes with completeness above 98% and contamination below 5% (n = 354), following the authors' guidelines. We then obtained taxonomic assignments for all remaining isolates using the Genome Taxonomy Database tool GTDB-Tk (v2.3.2, database version r214) (Chaumeil et al. 2019). We constructed a phylogenetic tree using the tool SpeciesTree (v2.2.0). We then trimmed the tree (using Trim SpeciesTree to GenomeSet- v1.4.0), retaining only tips within our collection with measured thermal performance.

59 BASIC BIOLOGICAL SCIENCES

Tissue Photolithography

Tissue lithography will enable physicians and researchers to obtain macromolecules with high purity (greater than 90 percent) from desired cells in conventionally processed, clinical tissues by simply annotating the desired cells on a computer screen. After identifying the desired cells, a suitable lithography mask will be generated to protect the contents of the desired cells while allowing destruction of all undesired cells by irradiation with ultraviolet light. The DNA from the protected cells can be used in a number of downstream applications including DNA sequencing. The purity (i.e., macromolecules isolated form specific cell types) of such specimens will greatly enhance the value and information of downstream applications. In this method, the specific cells are isolated on a microscope slide using photolithography, which will be faster, more specific, and less expensive than current methods. It relies on the fact that many biological molecules such as DNA are photosensitive and can be destroyed by ultraviolet irradiation. Therefore, it is possible to protect the contents of desired cells, yet destroy undesired cells. This approach leverages the technologies of the microelectronics industry, which can make features smaller than 1 micrometer with photolithography. A variety of ways has been created to achieve identification of the desired cell, and also to designate the other cells for destruction. This can be accomplished through chrome masks, direct laser writing, and also active masking using dynamic arrays. Image recognition is envisioned as one method for identifying cell nuclei and cell membranes. The pathologist can identify the cells of interest using a microscopic computerized image of the slide, and appropriate custom software. In one of the approaches described in this work, the software converts the selection into a digital mask that can be fed into a direct laser writer, e.g. the Heidelberg DWL66. Such a machine uses a metalized glass plate (with chrome metallization) on which there is a thin layer of photoresist. The laser transfers the digital mask onto the photoresist by direct writing, with typical best resolution of 2 micrometers. The plate is then developed to remove the exposed photoresist, which leaves the exposed areas susceptible to chemical chrome etch. The etch removes the unprotected chrome. The rest of the photoresist is then removed, by either ultraviolet organic solvent or over-development. The remaining chrome pattern is quickly oxidized by atmospheric exposure (typically within 30 seconds). The ready chrome mask is now applied to the tissue slide and aligned manually, or using automatic software and pre-designed alignment marks. The slide plate sandwich is then exposed to UV to destroy the DNA of the unwanted cells. The slide and plate are separated and the slide is processed in a standard way to prepare for polymerase chain reaction (PCR) and potential identification of cancer sequences.

Wade, Lawrence A.

Plant sulfate transporter protein sequences for phylogenetic analysis

Sulfur is an essential macronutrient that supports plant growth, development, and responses to environmental stress. Sulfate is the predominant inorganic form of sulfur in soils, and its uptake by roots and translocation to shoots are facilitated by the sulfate transporter (SULTR) family of proteins. Although the first plant SULTR gene was identified nearly three decades ago, several subfamily members, particularly those in the expansive and angiosperm-specific SULTR3 group, remain poorly characterized. To support comprehensive phylogenetic and sequence-based analyses, we compiled a curated dataset of 262 SULTR protein sequences from 22 plant species spanning the evolutionary breadth of land plants. This collection includes representatives from two basal lineages, two early-divergent angiosperms, six monocots, and ten dicots. All sequences were extracted from genome assemblies available in Phytozome v13 (Joint Genome Institute) and manually curated, with cross-referencing to additional databases such as NCBI when needed. This dataset provides a valuable resource for reconstructing the evolutionary history of the SULTR family, with particular emphasis on the diversification of SULTR3 transporters in flowering plants. This resource may also support functional annotation, comparative genomics, and structural modeling of sulfate transport proteins.

CBI

High-quality Acinetobacter genomes recovered from combat wounds via metagenomic sequencing resemble cultured isolate genomes

The ability to accurately characterize wound pathogens is critical to informing clinical decisions for wound infections with complex treatment requirements. Acinetobacter baumannii is an impactful nosocomial pathogen in combat wounds and civilian hospital-acquired infections. An informed understanding of the phylogenetics and epidemiology of A. baumannii infections in military and civilian environments could guide approaches that improve antibiotic treatment regimens for both military and civilian patients. Whole-genome data for bacterial strains can be difficult to obtain due to challenges in culturing isolates from preserved military specimens. Metagenomic sequencing and assembly create opportunities for genomic analysis of pathogens directly from clinical specimens. The ability to perform comparative analyses between metagenome-derived genomes and culture-derived genomes would support a range of comparative bacterial genomic studies. Wound tissue biopsy and effluent samples from combat injuries were subjected to metagenomic sequencing and assembly. In total, 42 microbial metagenome-assembled genomes (MAGs) were obtained directly from metagenomic sequence data, 36 of which were designated “high” quality. Thirty of these genomes corresponded to Acinetobacter, with 29 mapping specifically to A. baumannii. Other observed genera included Bordetella, Citrobacter, Escherichia, and Pseudomonas. Single-copy and multi-copy orthologs were identified across Acinetobacter MAGs and publicly available isolate genomes derived from military and civilian sources. Both MAG and military isolate genomes were annotated with antimicrobial resistance data, and MAG genomes were statistically comparable to genomes obtained from isolates. Our results highlight the potential of de novo metagenome assembly for enabling high-resolution characterization directly from clinical specimens, thereby improving diagnostic precision, guiding antimicrobial stewardship, and enhancing understanding of pathogen evolution across diverse healthcare and battlefield environments.

Acinetobacter baumannii

A Preferences Corpus and Annotation Scheme for Human-Guided Alignment of Time-Series GPTs

The process of time-series forecasting such as predicting trajectories of silicon content in blast furnaces is a difficult task. Most time-series approaches today focus on scalar-type MSE loss optimization. This optimization approach, while widely common, could benefit from the use of human expert or process-level preferences. In this paper, we introduce a novel alignment and fine-tuning approach that involves learning from a corpus of preferred and dis-preferred time-series prediction trajectories. Our contributions include (1) a preference annotation pipeline for time-series forecasts, (2) the application of Score-based Preference Optimization (SPO) to train decoder-only transformers from preferences, and (3) results showing improvements in forecast quality. The approach is validated on both proprietary blast furnace data and the UCI Appliances Energy dataset. The proposed preference corpus and training strategy offer a new option for fine-tuning sequence models in industrial settings.

DPO

Single-cell chromatin accessibility and cis -regulatory element analyses in plants using the scPlantReg platform

Understanding gene regulation is fundamental to plant improvement, but the lack of plant-specific single-cell assay for transposase-accessible chromatin using sequencing (scATAC-seq) frameworks and cross-species databases has limited insights into cell-type-specific cellular regulation. Here we present ‘scPlantReg’, an integrated framework and database for plant scATAC-seq data. scPlantReg supports end-to-end analyses from raw data processing to biological interpretation and features ‘scATACtor’, a supervised machine-learning approach that outperforms existing tools for cell-type annotation. We applied scPlantReg to pearl millet to characterize cell-type-specific chromatin accessibility and identify validated activating and repressing accessible chromatin regions (ACRs), revealing WRKY transcription factors as potential regulators of xylem development. Furthermore, we reanalysed scATAC-seq datasets from 8 plant species, spanning 11 tissues and multiple developmental stages, enabling cross-species comparisons. Furthermore, these analyses uncovered conserved regulatory programmes, including AP2/EREBP-associated ACRs linked to cell wall development and cell-type-conserved TFs across grasses. Collectively, scPlantReg provides a general framework and resource for comparative regulatory analysis in plants.

Epigenomics

NEAR: Neural Embeddings for Amino acid Relationships

Protein language models (PLMs) have recently demonstrated potential to supplant classical protein database search methods based on sequence alignment, but are slower than common alignment-based tools and appear to be prone to a high rate of false labeling. Here, we present NEAR, a method based on neural representation learning that is designed to improve both speed and accuracy of search for likely homologs in a large protein sequence database. NEAR’s ResNet embedding model is trained using contrastive learning guided by trusted sequence alignments. It computes per-residue embeddings for target and query protein sequences, and identifies alignment candidates with a pipeline consisting of residue-level k-NN search and a simple neighbor aggregation scheme. Tests on a benchmark consisting of trusted remote homologs and randomly shuffled decoy sequences reveal that NEAR substantially improves accuracy relative to state-of-the-art PLMs, with lower memory requirements and faster embedding and search speed. While these results suggest that the NEAR model may be useful for standalone homology detection with increased sensitivity over standard alignment-based methods, in this manuscript we focus on a more straightforward analysis of the model’s value as a high-speed pre-filter for sensitive annotation. In that context, NEAR is at least 5x faster than the pre-filter currently used in the widely-used profile hidden Markov model (pHMM) search tool HMMER3, and also outperforms the pre-filter used in our fast pHMM tool, nail.

59 BASIC BIOLOGICAL SCIENCES

Genomic and transcriptomic characterization of carbohydrate-active enzymes in the anaerobic fungus Neocallimastix cameroonii var. constans

Anaerobic gut fungi effectively degrade lignocellulose in the guts of large herbivores, but there remain a limited number of isolated, publicly available, and sequenced strains that impede our understanding of the role of anaerobic fungi within microbial communities. We isolated and characterized a new fungal isolate, Neocallimastix cameroonii var. constans, providing a transcriptomic and genomic understanding of its ability to degrade diverse carbohydrates. This anaerobic fungal strain was stably cultivated for multiple years in vitro among members of an initial enrichment microbial community derived from goat feces, and it demonstrated the ability to pair with other microbial members, namely, archaeal methanogens to produce methane from lignocellulose. Genomic analysis revealed a higher number of predicted carbohydrate-active enzymes encoded in the N. cameroonii var. constans genome compared to most other sequenced anaerobic fungi. The carbohydrate-active enzyme profile for this isolate contained 660 glycoside hydrolases, 160 carbohydrate esterases, 194 glycosyltransferases, and 85 polysaccharide lyases. Differential gene expression analysis showed the upregulation of thousands of genes (including predicted carbohydrate-active enzymes) when N. cameroonii var. constans was grown on lignocellulose (reed canary grass) compared to less complex substrates, such as cellulose (filter paper), cellobiose, and glucose. AlphaFold was used to predict functions of transcriptionally active yet poorly annotated genes, revealing feruloyl esterases that likely play an important role in lignocellulose degradation by anaerobic fungi. The combination of this strain's genomic and transcriptomic characterization, omics-informed structural prediction, and robustness in microbial co-culture make it a well-suited platform to conduct future investigations into bioprocessing and enzyme discovery.

CAZymes

Integrative Modeling and Analysis of Fungal Central Carbon Metabolism

Over a thousand fungal genomes have been sequenced, yet manually curated genome-scale metabolic models (GEMs) are available for only a limited number of species. Moreover, these models have often been developed independently, leading to inconsistencies in namespaces, compartment definitions, and pathway representations that hinder comparative analysis, the systematic reuse of prior curation efforts, and the integration of consolidated metabolic knowledge. Here, we present the Consolidated Fungal Core Metabolism Model (CFCMM), constructed by integrating thirteen published fungal models spanning Ascomycota, Mucoromycota, and both Crabtree-positive and Crabtree-negative yeasts. We harmonized metabolites and reactions into a non-redundant shared ModelSEED ontological space, standardized compartmentalization, and refined gene–protein–reaction (GPR) rules. Using pathway-level visualization and systematic gap detection, we further improved the integrated network through literature-guided curation to correct stoichiometry, stereospecificity, and pathway architecture. Orthologous protein family reconstruction and functional annotation workflows were used to validate and inform GPR associations, with particular emphasis on ambiguous enzyme superfamilies and membrane-associated components. Using the resulting CFCMM, we built high-quality central carbon core models for each fungus and performed flux balance analysis to quantify ATP-yield variation under aerobic and anaerobic conditions, explicitly evaluating scenarios driven by differences in electron transport chain (ETC) composition. Simulations reproduced the expected fermentative yield of approximately 2 mmol ATP per mmol glucose under anaerobic conditions and separated the thirteen fungi into two bioenergetic groups under aerobic respiration based on Complex I status, with predicted yields of approximately 30 versus 22 mmol ATP per mmol glucose. Forcing flux through the alternative oxidase bypass further reduced ATP yields to approximately 12 and 4 mmol ATP per mmol glucose in Complex I-containing and Complex I-lacking fungi, respectively. Collectively, this work provides a manually curated, ModelSEED-consistent, and extensible fungal core metabolic template, deployed in DOE KBase as a resource for automated reconstruction of central carbon core models from any sequenced fungal genome. In addition, the CFCMM provides modular components for developing GEMs with more accurate energy predictions and enables robust comparative analyses of fungal bioenergetics and core metabolic diversity

59 BASIC BIOLOGICAL SCIENCES

Beyond microbial abundance: metadata integration enhances disease prediction in human microbiome studies

Multiple studies have highlighted the interaction of the human microbiome with physiological systems such as the gut, immune, liver, and skin, via key axes. Advances in sequencing technologies and high-performance computing have enabled the analysis of large-scale metagenomic data, facilitating the use of machine learning to predict disease likelihood from microbiome profiles. However, challenges such as compositionality, high dimensionality, sparsity, and limited sample sizes have hindered the development of actionable models. One strategy to improve these models is by incorporating key metadata from both the human host and sample collection/processing protocols. This remains challenging due to sparsity and inconsistency in metadata annotation and availability. In this paper, we introduce a machine learning-based pipeline for predicting human disease states by integrating host and protocol metadata with microbiome abundance profiles from 68 different studies, processed through a consistent pipeline. Our findings indicate that metadata can enhance machine learning predictions, particularly at higher taxonomic ranks like Kingdom and Phylum, though this effect diminishes at lower ranks. Our study leverages a large collection of microbiome datasets comprising 11,208 samples, therefore enhancing the robustness and statistical confidence of our findings. This work is a critical step toward utilizing microbiome and metadata for predicting diseases such as gastrointestinal infections, diabetes, cancer, and neurological disorders.

Mathematics and Computing

The landscape of regulatory element evolution in a C4 perennial grass

Gene regulatory evolution is a well-known source of phenotypic diversity and adaptive evolution. Although cis-regulatory elements (CREs) play a vital role in gene expression evolution, the molecular evolution of CREs remains mostly unknown due to the difficulty in identifying and characterizing these functional elements. Comparative genomic analyses of noncoding DNA can be leveraged to identify conserved noncoding sequences (CNS), many of which may harbor functional CREs conserved by purifying selection. However, purely computational inference of CREs from putative CNS can be erroneous due to the complex genomic architecture in plants. One promising experimental approach to identify CREs is by profiling accessible chromatin regions (ACRs) that are often associated with the location of CREs. In this study, we use comparative genomics along with the profiling of ACRs to study the molecular evolution of putative functional noncoding regulatory regions in Panicoid grasses. We identified sets of CNS that varied in relationship to the degree of evolutionary divergence among the studied taxa, including identifying core-Panicoid-CNS. We augmented this analysis by profiling ACRs in Panicum hallii ecotypes using ATAC-seq. ACRs had low SNP density at the summit, harbored a high frequency of core-Panicoid-CNS, and were enriched with expression QTL. These data help to annotate the P. hallii genome for putative functional elements and suggest that a large proportion of these ACRs are evolving under purifying selection. Turnover in CNS and ACR between ecotypes of P. hallii identifies a small set of putatively divergent CREs that may underlie differences in gene regulation between genotypes from inland and coastal habitats. In summary, we profiled ACRs in Panicoid grasses and integrated this data with our putative CNS prediction framework, which provides unique insight into patterns of polymorphism and divergence in CREs in C4 perennial grasses.

59 BASIC BIOLOGICAL SCIENCES

Characterization of Multiple Trichloroethene, cis-Dichloroethene and 1,1-Dichloroethene Degrading Propanotrophic Communities

Aerobic cometabolism offers a viable strategy for the remediation of chlorinated solvent plumes at oxic sites where anaerobic approaches are limited. In this study, propane-enriched mixed cultures (derived from agricultural soils and an impacted site sediment) which previously degraded 1,4-dioxane, were evaluated for their capacity to also degrade trichloroethene (TCE), cis-1,2-dichloroethene (cDCE), and 1,1-dichloroethene (1,1-DCE) over successive transfers. Sustained biodegradation of TCE and cDCE was observed across multiple enrichments, and cultures enriched on one compound generally degraded the other. In contrast, 1,1-DCE biodegradation was restricted to a subset of cultures and removal times increased over transfers. Further, 1,1-DCE removal was absent at elevated concentrations, both trends consistent with inhibitory or toxic effects. Whole genome sequencing analyses revealed pronounced substrate-dependent selection of microbial communities, with cDCE-degrading cultures being dominated by Mycobacterium and Mycolicibacterium, whereas TCE-degrading cultures were dominated by Rhodococcus. Rhodococcus metagenome-assembled genomes (MAGs) in the TCE degrading cultures classified as R. opacus or R. wratislaviensis. 1,1-DCE degrading cultures were dominated by Pseudonocardia, although the associated MAGs contained a truncated propane monooxygenase alpha subunit. Functional gene analysis identified both group 5 (prmABCD) and putative group 6 propane monooxygenases. The following KBase narratives contain the quality controlled reads, MAGs (fasta assemblies) and the prokka annotations for each assembly TCE Site 1A and B Propanotrophic MAGs (https://narrative.kbase.us/narrative/254918) TCE Soil 2A and B Propanotrophic MAGs (https://narrative.kbase.us/narrative/254919) TCE Soil T3 A and B Propanotrophic MAGs (https://narrative.kbase.us/narrative/254920) TCE Soil T4 A and B Propanotrophic MAGs (https://narrative.kbase.us/narrative/254921) cDCE Site 1A 1B Propanotrophic MAGs (https://narrative.kbase.us/narrative/254915) cDCE Soils T2 A and B Propanotrophic MAGs (https://narrative.kbase.us/narrative/254927) cDCE Soil T3 A and B Propanotrophic MAGs (https://narrative.kbase.us/narrative/254928) cDCE Soil 4A and B Propanotrophic MAGs (https://narrative.kbase.us/narrative/254942) 1,1-DCE T2 T3 Propanotrophic MAGs (https://narrative.kbase.us/narrative/254903)

59 BASIC BIOLOGICAL SCIENCES

Geology of the One Earth Energy Site

The One Earth Energy site is one of two sites in the Illinois Storage Corridor (ISC) project. The objectives of the ISC project is to accelerate commercial deployment of carbon capture utilization and storage at two individual sites and receive approvals for Underground Injection Control (UIC) Class VI permits for construction at each site. At the One Earth Energy site, an extensive data collection program was undertaken, which included the drilling of a test well (One Earth Energy #1 [OEE #1]), four 2D seismic lines, and a small 3D seismic survey. The OEE #1 well was drilled in 2022 and acquired extensive core, log, and testing data to characterize the subsurface geology of the site. Coring was focused on the storage interval, the Mt. Simon Sandstone, and the confining interval, the Eau Claire Formation. The core and log data were used to evaluate the sedimentology and sequence stratigraphy, as well as to develop the conceptual geologic model. This report includes the geological summaries of the Mt. Simon Sandstone and the Eau Claire Formation. The extensive analysis of the log data is included in the petrophysical section, showing ranges of porosity, estimated pore size, and the mineral content of selected zones in the well. The separate petrographic technical report entitled “Petrographic and Advanced Geologic Characterization Report on One Earth Energy #1 (API# 1211325373)”, report number DOE-UIUC-0031892-04, details thin section point-counting analysis that includes mineralogical and pore space analysis, including grain size analysis, annotated thin section photomicrographs, scanning electron microscopy (SEM) with energy dispersive X-ray spectroscopy (EDS), and statistics of grain size analysis on Mt. Simon thin sections from OEE #1. The final OEE #1 well data to be included in this geology report is the routine core analysis of both whole core plugs and rotary sidewall core plugs. In addition to the OEE #1 well, four 2D seismic lines and a small 3D survey were acquired as part of the overall subsurface geological characterization. This geology report references the seismic interpretation report, entitled “One Earth Energy Site Seismic Interpretation Task 5.0”, report number DOE-UIUC-0031892-07. This report details the stratigraphic and structural interpretation of the 2D and 3D seismic data acquired at the One Earth Energy site. The 2D seismic data was acquired in 2019 and 2021, and the 3D survey was acquired in 2022. The objectives of the seismic programs were to contribute to the subsurface characterization of the Mt. Simon-Eau Claire Storage Complex by evaluating the continuity of potential storage reservoirs and containment intervals across the project area, and to determine if any geologic features are present that would increase containment risk to the proposed carbon storage project.

09 BIOMASS FUELS

Relics of interspecific hybridization retained in the genome of a drought-adapted peanut cultivar

Peanut (Arachis hypogaea L.) is a globally important oil and food crop frequently grown in arid, semi-arid, or dryland environments. Improving drought tolerance is a key goal for peanut crop improvement efforts. Here, we present the genome assembly and gene model annotation for “Line8,” a peanut genotype bred from drought-tolerant cultivars. Our assembly and annotation are the most contiguous and complete peanut genome resources currently available. The high contiguity of the Line8 assembly allowed us to explore structural variation both between peanut genotypes and subgenomes. We detect several large inversions between Line8 and other peanut genome assemblies, and there is a trend for the inversions between more genetically diverged genotypes to have higher gene content. We also relate patterns of subgenome exchange to structural variation between Line8 homeologous chromosomes. Unexpectedly, we discover that Line8 harbors an introgression from A.cardenasii, a diploid peanut relative and important donor of disease resistance alleles to peanut breeding populations. The fully resolved sequences of both haplotypes in this introgression provide the first in situ characterization of A.cardenasii candidate alleles that can be leveraged for future targeted improvement efforts. The completeness of our genome will support peanut biotechnology and broader research into the evolution of hybridization and polyploidy.

60 APPLIED LIFE SCIENCES

The Conservation of Structure and Mechanism of Catalytic Action in a Family of Thiamin Pyrophosphate (TPP)-dependent Enzymes

Thiamin pyrophosphate (TPP)-dependent enzymes are a divergent family of TPP and metal ion binding proteins that perform a wide range of functions with the common decarboxylation steps of a -(O=)C-C(OH)- fragment of alpha-ketoacids and alpha- hydroxyaldehydes. To determine how structure and catalytic action are conserved in the context of large sequence differences existing within this family of enzymes, we have carried out an analysis of TPP-dependent enzymes of known structures. The common structure of TPP-dependent enzymes is formed at the interface of four alpha/beta domains from at least two subunits, which provide for two metal and TPP-binding sites. Residues around these catalytic sites are conserved for functional purpose, while those further away from TPP are conserved for structural reasons. Together they provide a network of contacts required for flip-flop catalytic action within TPP-dependent enzymes. Thus our analysis defines a TPP-action motif that is proposed for annotating TPP-dependent enzymes for advancing functional proteomics.

Dominiak, P.

Specialization Restricts the Evolutionary Paths Available to Yeast Sugar Transporters

Functional innovation at the protein level is a key source of evolutionary novelties. The constraints on functional innovations are likely to be highly specific in different proteins, which are shaped by their unique histories and the extent of global epistasis that arises from their structures and biochemistries. These contextual nuances in the sequence–function relationship have implications both for a basic understanding of the evolutionary process and for engineering proteins with desirable properties. Here, we have investigated the molecular basis of novel function in a model member of an ancient, conserved, and biotechnologically relevant protein family. These Major Facilitator Superfamily sugar porters are a functionally diverse group of proteins that are thought to be highly plastic and evolvable. By dissecting a recent evolutionary innovation in an α-glucoside transporter from the yeast Saccharomyces eubayanus, we show that the ability to transport a novel substrate requires high-order interactions between many protein regions and numerous specific residues proximal to the transport channel. To reconcile the functional diversity of this family with the constrained evolution of this model protein, we generated new, state-of-the-art genome annotations for 332 Saccharomycotina yeast species spanning ~400 My of evolution. By integrating phylogenetic and phenotypic analyses across these species, we show that the model yeast α-glucoside transporters likely evolved from a multifunctional ancestor and became subfunctionalized. The accumulation of additive and epistatic substitutions likely entrenched this subfunction, which made the simultaneous acquisition of multiple interacting substitutions the only reasonably accessible path to novelty.

59 BASIC BIOLOGICAL SCIENCES

General Biology 2 Sugar Beet Lab - SU2025 - Final

This module uses the Department of Energy Systems Biology Knowledgebase (KBase) platform to explore topics such as genome assembly, metagenomics, and phylogenomics. Here students will use sequences from DNA that they collected to compare the metagenomes of microbial communities from the rhizosphere of plants grown in fertilized vs. unfertilized soils. Using these data students will evaluate the impact of fertilizer on these communities and how these microbial communities influence soil health and plant growth.

59 BASIC BIOLOGICAL SCIENCES