Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “annotations”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29

Opportunities and Challenges for Machine Learning-Assisted Enzyme Engineering

Enzymes can be engineered at the level of their amino acid sequences to optimize key properties such as expression, stability, substrate range, and catalytic efficiency or even to unlock new catalytic activities not found in nature. Because the search space of possible proteins is vast, enzyme engineering usually involves discovering an enzyme starting point that has some level of the desired activity followed by directed evolution to improve its “fitness” for a desired application. Recently, machine learning (ML) has emerged as a powerful tool to complement this empirical process. ML models can contribute to (1) starting point discovery by functional annotation of known protein sequences or generating novel protein sequences with desired functions and (2) navigating protein fitness landscapes for fitness optimization by learning mappings between protein sequences and their associated fitness values. In this Outlook, we explain how ML complements enzyme engineering and discuss its future potential to unlock improved engineering outcomes.

60 APPLIED LIFE SCIENCES↗

An Improved CRISPR Interference Tool to Engineer Rhodococcus opacus

Rhodococcus opacus is a non-model bacterium that is well suited for valorizing lignin. Despite recent advances in our systems-level understanding of its versatile metabolism, studies of its gene functions at a single gene level are still lagging. Elucidating gene functions in non-model organisms is challenging due to limited genetic engineering tools that are convenient to use. To address this issue, we developed a simple gene repression system based on CRISPR interference (CRISPRi). This gene repression system uses a T 7 RNA polymerase system to express a small guide RNA, demonstrating improved repression compared to the previously demonstrated CRISPRi system (i.e., the maximum repression efficiency improved from 58% to 85%). Additionally, our cloning strategy allows for building multiple CRISPRi plasmids in parallel without any PCR step, facilitating the engineering of this GC-rich organism. Using the improved CRISPRi system, we confirmed the annotated roles of four metabolic pathway genes, which had been identified by our previous transcriptomic analysis to be related to the consumption of benzoate, vanillate, catechol, and acetate. Furthermore, we showed our tool’s utility by demonstrating the inducible accumulation of muconate that is a precursor of adipic acid, an important monomer for nylon production. While the maximum muconate yield obtained using our tool was 30% of the yield obtained using gene knockout, our tool showed its inducibility and partial repressibility. In conclusion, our CRISPRi tool will be useful to facilitate functional studies of this non-model organism and engineer this promising microbial chassis for lignin valorization.

09 BIOMASS FUELS↗

The Factors Governing Metal Dependence of an Emergent Superfamily of Bimetallic Oxygenases

Metalloenzyme superfamilies are typically defined by their protein scaffolds and active sites. Owing to the high tunability of protein structures, members of a single superfamily can catalyze diverse reactions with the same metallocofactor. Some superfamilies, such as amidohydrolase-related dinuclear oxygenases (AROs), display further versatility by utilizing multiple metallocofactors. We have shown that certain AROs catalyze monooxygenation reactions with diiron, dimanganese, and/or mixed manganese−iron cofactors, but the molecular factors governing the selection of a particular cofactor remain unknown, and the extent of this superfamily in biology is unclear. Here, we report bioinformatic analyses that expand the ARO superfamily to approximately 17,000 unique UniProt sequences, far exceeding the number of previously characterized enzymes. Through the integration of structural, spectroscopic, and thermodynamic analyses of representative proteins with a bioinformatic pipeline that identifies key secondary- and tertiary-sphere residues, we can predict in silico the metal preference for the majority of reported ARO sequences. These annotations were validated via the characterization of multiple new AROs, including ones implicated in key oxidative steps of natural product biosyntheses. This study establishes the key structure−function relationships governing metal preferences in AROs and highlights their vastly underappreciated role in myriad biological processes.

Liu, Chang [University of California, Berkeley, CA↗

Storage Conditions of Human Kidney Tissue Sections Affect Spatial Lipidomics Analysis Reproducibility

Lipids often are labile, unstable, and tend to degrade overtime, so it is of the upmost importance to study these molecules in their most native state. In this study, we sought to understand the optimal storage conditions for spatial lipidomic analysis of human kidney tissue sections. Specifically, we evaluated human kidney tissue sections on several different days throughout the span of a week using our established protocol for elucidating lipids using high mass resolution matrix-assisted laser desorption/ionization mass spectrometry imaging (MALDI-MSI). We studied kidney tissue sections stored under five different conditions: open stored at –80 °C, vacuumed sealed and stored at –80 °C, with matrix preapplied before storage at –80 °C, under a nitrogen atmosphere and stored at –80 °C, and at room temperature in a desiccator. Results were compared to data obtained from kidney tissue sections that were prepared and analyzed immediately after cryosectioning. Data was processed using METASPACE. After a week of storage, the sections stored at room temperature showed the largest amount of lipid degradation, while sections stored under nitrogen and at –80 °C retained the greatest number of overlapping annotations in relation to freshly cut tissue. Overall, we found that molecular degradation of the tissue sections was unavoidable over time, regardless of storage conditions, but storing tissue sections in an inert gas at low temperatures can curtail molecular degradation within tissue sections.

60 APPLIED LIFE SCIENCES↗

Evaluation of a Reference-Free Collision Cross Section Calibration Strategy for Proteomics Using SLIM-Based High-Resolution Ion Mobility Spectrometry–Mass Spectrometry

Ion mobility spectrometry (IMS) is a gas-phase analytical technique that separates ions with different sizes and shapes and is compatible with mass spectrometry (MS) to provide an additional separation dimension. The rapid nature of the IMS separation combined with the high sensitivity of MS-based detection and the ability to derive structural information on analytes in the form of the property collision cross section (CCS) makes IMS particularly well-suited for characterizing complex samples in -omics applications. In such applications, the quality of CCS from IMS measurements is critical to confident annotation of the detected components in the complex -omics samples. However, most IMS instrumentation in mainstream use requires calibration to calculate CCS from measured arrival times, with the most notable exception being drift tube IMS measurements using multifield methods. The strategy for calibrating CCS values, particularly selection of appropriate calibrants, has important implications for CCS accuracy, reproducibility, and transferability between laboratories. The conventional approach to CCS calibration involves explicitly defining calibrants ahead of data acquisition and crucially relies upon availability of reference CCS values. In this work, we present a novel reference-free approach to CCS calibration which leverages trends among putatively identified features and computational CCS prediction to conduct calibrations post-data acquisition and without relying on explicitly defined calibrants. We demonstrated the utility of this reference-free CCS calibration strategy for proteomics application using high-resolution structures for lossless ion manipulations (SLIM)-based IMS-MS. In conclusion, we first validated the accuracy of CCS values using a set of synthetic peptides and then demonstrated using a complex peptide sample from cell lysate.

59 BASIC BIOLOGICAL SCIENCES↗

Autotrophic and mixotrophic metabolism of an anammox bacterium revealed by in vivo 13 C and 2 H metabolic network mapping

Anaerobic ammonium-oxidizing (anammox) bacteria mediate a key step in the biogeochemical nitrogen cycle and have been applied worldwide for the energy-efficient removal of nitrogen from wastewater. However, outside their core energy metabolism, little is known about the metabolic networks driving anammox bacterial anabolism and use of different carbon and energy substrates beyond genome-based predictions. Here, we experimentally resolved the central carbon metabolism of the anammox bacterium Candidatus ‘Kuenenia stuttgartiensis’ using time-series 13 C and 2 H isotope tracing, metabolomics, and isotopically nonstationary metabolic flux analysis. Our findings confirm predicted metabolic pathways used for CO 2 fixation, central metabolism, and amino acid biosynthesis in K. stuttgartiensis , and reveal several instances where genomic predictions are not supported by in vivo metabolic fluxes. This includes the use of the oxidative branch of an incomplete tricarboxylic acid cycle for alpha-ketoglutarate biosynthesis, despite the genome not having an annotated citrate synthase. We also demonstrate that K. stuttgartiensis is able to directly assimilate extracellular formate via the Wood–Ljungdahl pathway instead of oxidizing it completely to CO 2 followed by reassimilation. In contrast, our data suggest that K. stuttgartiensis is not capable of using acetate as a carbon or energy source in situ and that acetate oxidation occurred via the metabolic activity of a low-abundance microorganism in the bioreactor’s side population. Together, these findings provide a foundation for understanding the carbon metabolism of anammox bacteria at a systems-level and will inform future studies aimed at elucidating factors governing their function and niche differentiation in natural and engineered ecosystems.

54 ENVIRONMENTAL SCIENCES↗

A Guanidine-Degrading Enzyme Controls Genomic Stability of Ethylene-Producing Cyanobacteria

Recent studies have revealed the prevalence and biological significance of guanidine metabolism in nature. However, the metabolic pathways used by microbes to degrade guanidine or mitigate its toxicity have not been widely studied. Here, via comparative proteomics and subsequent experimental validation, we demonstrate that Sll1077, previously annotated as an agmatinase enzyme in the model cyanobacterium Synechocystis sp. PCC 6803, is more likely a guanidinase as it can break down guanidine rather than agmatine into urea and ammonium. The model cyanobacterium Synechococcus elongatus PCC 7942 strain engineered to express the bacterial ethylene-forming enzyme (EFE) exhibits unstable ethylene production due to toxicity and genomic instability induced by accumulation of the EFE-byproduct guanidine. Co-expression of EFE and Sll1077 significantly enhances genomic stability and enables the resulting strain to achieve sustained high-level ethylene production. These findings expand our knowledge of natural guanidine degradation pathways and demonstrate their biotechnological application to support ethylene bioproduction.

59 BASIC BIOLOGICAL SCIENCES↗

Seasonal activities of the phyllosphere microbiome of perennial crops

Understanding the interactions between plants and microorganisms can inform microbiome management to enhance crop productivity and resilience to stress. Here, we apply a genome-centric approach to identify ecologically important leaf microbiome members on replicated plots of field-grown switchgrass and miscanthus, and to quantify their activities over two growing seasons for switchgrass. We use metagenome and metatranscriptome sequencing and curate 40 medium- and high-quality metagenome-assembled-genomes (MAGs). We find that classes represented by these MAGs (Actinomycetia, Alpha- and Gamma- Proteobacteria, and Bacteroidota) are active in the late season, and upregulate transcripts for short-chain dehydrogenase, molybdopterin oxidoreductase, and polyketide cyclase. Stress-associated pathways are expressed for most MAGs, suggesting engagement with the host environment. We also detect seasonally activated biosynthetic pathways for terpenes and various non-ribosomal peptide pathways that are poorly annotated. Our findings support that leaf-associated bacterial populations are seasonally dynamic and responsive to host cues.

59 BASIC BIOLOGICAL SCIENCES↗

Genome analyses reveal population structure and a purple stigma color gene candidate in finger millet

Finger millet is a key food security crop widely grown in eastern Africa, India and Nepal. Long considered a ‘poor man’s crop’, finger millet has regained attention over the past decade for its climate resilience and the nutritional qualities of its grain. To bring finger millet breeding into the 21 st century, here we present the assembly and annotation of a chromosome-scale reference genome. We show that this ~1.3 million years old allotetraploid has a high level of homoeologous gene retention and lacks subgenome dominance. Population structure is mainly driven by the differential presence of large wild segments in the pericentromeric regions of several chromosomes. Trait mapping, followed by variant analysis of gene candidates, reveals that loss of purple coloration of anthers and stigma is associated with loss-of-function mutations in the finger millet orthologs of the maize R1/B1 and Arabidopsis GL3/EGL3 anthocyanin regulatory genes. Proanthocyanidin production in seed is not affected by these gene knockouts.

59 BASIC BIOLOGICAL SCIENCES↗

Central transcriptional regulator controls photosynthetic growth and carbon storage in response to high light

Carbon capture and biochemical storage are some of the primary drivers of photosynthetic yield and productivity. To elucidate the mechanisms governing carbon allocation, we designed a photosynthetic light response test system for genetic and metabolic carbon assimilation tracking, using microalgae as simplified plant models. The systems biology mapping of high light-responsive photophysiology and carbon utilization dynamics between two variants of the same Picochlorum celeri species, TG1 and TG2 elucidated metabolic bottlenecks and transport rates of intermediates using instationary 13 C-fluxomics. Simultaneous global gene expression dynamics showed 73% of the annotated genes responding within one hour, elucidating a singular, diel-responsive transcription factor, closely related to the CCA1/LHY clock genes in plants, with significantly altered expression in TG2. Transgenic P. celeri TG1 cells expressing the TG2 CCA1/LHY gene, showed 15% increase in growth rates and 25% increase in storage carbohydrate content, supporting a coordinating regulatory function for a single transcription factor.

09 BIOMASS FUELS↗

Barcoded overexpression screens in gut Bacteroidales identify genes with roles in carbon utilization and stress resistance

Abstract A mechanistic understanding of host-microbe interactions in the gut microbiome is hindered by poorly annotated bacterial genomes. While functional genomics can generate large gene-to-phenotype datasets to accelerate functional discovery, their applications to study gut anaerobes have been limited. For instance, most gain-of-function screens of gut-derived genes have been performed in Escherichia coli and assayed in a small number of conditions. To address these challenges, we develop Barcoded Overexpression BActerial shotgun library sequencing (Boba-seq). We demonstrate the power of this approach by assaying genes from diverse gut Bacteroidales overexpressed in Bacteroides thetaiotaomicron . From hundreds of experiments, we identify new functions and phenotypes for 29 genes important for carbohydrate metabolism or tolerance to antibiotics or bile salts. Highlights include the discovery of a d -glucosamine kinase, a raffinose transporter, and several routes that increase tolerance to ceftriaxone and bile salts through lipid biosynthesis. This approach can be readily applied to develop screens in other strains and additional phenotypic assays.

59 BASIC BIOLOGICAL SCIENCES↗

Structural diversity and clustering of bacterial flagellar outer domains

Supercoiled flagellar filaments function as mechanical propellers within the bacterial flagellum complex, playing a crucial role in motility. Flagellin, the building block of the filament, features a conserved inner D0/D1 core domain across different bacterial species. In contrast, approximately half of the flagellins possess additional, highly divergent outer domain(s), suggesting varied functional potential. In this study, we report atomic structures of flagellar filaments from three distinct bacterial species: Cupriavidus gilardii , Stenotrophomonas maltophilia , and Geovibrio thiophilus . Our findings reveal that the flagella from the facultative anaerobic G. thiophilus possesses a significantly more negatively charged surface, potentially enabling adhesion to positively charged minerals. Furthermore, we analyze all AlphaFold predicted structures for annotated bacterial flagellins, categorizing the flagellin outer domains into 682 structural clusters. This classification provides insights into the prevalence and experimental verification of these outer domains. Remarkably, two of the flagellar structures reported herein belong to a distinct cluster, indicating additional opportunities on the study of the functional diversity of flagellar outer domains. Our findings underscore the complexity of bacterial flagellins and open up possibilities for future studies into their varied roles beyond motility.

Science & Technology - Other Topics↗

Functional protein mining with conformal guarantees

Molecular structure prediction and homology detection offer promising paths to discovering protein function and evolutionary relationships. However, current approaches lack statistical reliability assurances, limiting their practical utility for selecting proteins for further experimental and in-silico characterization. To address this challenge, we introduce a statistically principled approach to protein search leveraging principles from conformal prediction, offering a framework that ensures statistical guarantees with user-specified risk and provides calibrated probabilities (rather than raw ML scores) for any protein search model. Our method (1) lets users select many biologically-relevant loss metrics (i.e. false discovery rate) and assigns reliable functional probabilities for annotating genes of unknown function; (2) achieves state-of-the-art performance in enzyme classification without training new models; and (3) robustly and rapidly pre-filters proteins for computationally intensive structural alignment algorithms. Our framework enhances the reliability of protein homology detection and enables the discovery of uncharacterized proteins with likely desirable functional properties.

59 BASIC BIOLOGICAL SCIENCES↗

A global soil plasmidome resource unveils functional and ecological roles of plasmids in soil microbiomes

Plasmids play significant roles in microbial adaptation to ecosystems, yet their dynamics remain poorly understood due to identification challenges. We present the Global Soil Plasmidome Resource (GSPR), a comprehensive dataset of 98,728 plasmid sequences amassed from 6860 terrestrial microbial communities and isolates. We explore this resource through various computational approaches, including phylogenetic diversity analysis, host prediction, and extensive functional annotation, to understand the contribution of plasmids to the genetic and functional diversity in soil, correlating these findings with sample type, as well as the soil habitat they were retrieved from. Our analysis reveals insights into plasmid-encoded functions such as effector modules, quorum sensing, and stress resistance, which may contribute to their persistence and microbial adaptation in soil. Furthermore, CRISPR analysis suggests a prevalent role of these elements related to intra-plasmid competition. By contrasting plasmids from cultivated and uncultivated organisms, we identify important functions that expand existing knowledge of plasmid roles in these habitats. This study represents a notable step forward in elucidating plasmid diversity and function within soil microbiomes and establishes a foundational framework for exploring their roles in natural environments.

Fiamenghi, Mateus B↗

Synthetic data-driven deep learning for label-free autonomous atomic force microscopy

Atomic force microscopy (AFM) is a widely used tool for nanoscale characterization across materials science, energy research, and biology. However, its adoption in high-throughput materials discovery and statistically driven studies remains limited by a strong dependence on expert operator input and by the scarcity of annotated experimental AFM datasets needed to enable data-driven automation. Here, we introduce SimuScan, a synthetic-data–driven framework that enables reliable AFM feature identification, segmentation, and targeted imaging without requiring large manually labeled experimental datasets. SimuScan generates tunable, high-fidelity synthetic AFM images of defined morphologies while incorporating realistic experimental artifacts, including tip–sample convolution, noise, flattening distortions, and surface debris. These datasets are shown to support scalable, label-free training of modern deep learning models for AFM analysis. When integrated into data-driven AFM workflows, SimuScan-trained models can locate and analyze nanoscale structures across large datasets and guide targeted follow-up imaging. We validate this approach on nanostructured surfaces, DNA assemblies, and bacterial cells, demonstrating robust generalization across diverse sample types with minimal operator intervention. More broadly, this work establishes a general strategy for generating explicitly conditioned, task-relevant synthetic data to improve the reliability of downstream models in autonomous microscopy.

Millan-Solsona, Ruben [Oak Ridge National Laborato↗

Dynamic genome evolution in a model fern

The large size and complexity of most fern genomes have hampered efforts to elucidate fundamental aspects of fern biology and land plant evolution through genome-enabled research. Here we present a chromosomal genome assembly and associated methylome, transcriptome and metabolome analyses for the model fern species Ceratopteris richardii. The assembly reveals a history of remarkably dynamic genome evolution including rapid changes in genome content and structure following the most recent whole-genome duplication approximately 60 million years ago. These changes include massive gene loss, rampant tandem duplications and multiple horizontal gene transfers from bacteria, contributing to the diversification of defence-related gene families. The insertion of transposable elements into introns has led to the large size of the Ceratopteris genome and to exceptionally long genes relative to other plants. Gene family analyses indicate that genes directing seed development were co-opted from those controlling the development of fern sporangia, providing insights into seed plant evolution. Our findings and annotated genome assembly extend the utility of Ceratopteris as a model for investigating and teaching plant biology.

59 BASIC BIOLOGICAL SCIENCES↗

Single-cell chromatin accessibility and cis -regulatory element analyses in plants using the scPlantReg platform

Understanding gene regulation is fundamental to plant improvement, but the lack of plant-specific single-cell assay for transposase-accessible chromatin using sequencing (scATAC-seq) frameworks and cross-species databases has limited insights into cell-type-specific cellular regulation. Here we present ‘scPlantReg’, an integrated framework and database for plant scATAC-seq data. scPlantReg supports end-to-end analyses from raw data processing to biological interpretation and features ‘scATACtor’, a supervised machine-learning approach that outperforms existing tools for cell-type annotation. We applied scPlantReg to pearl millet to characterize cell-type-specific chromatin accessibility and identify validated activating and repressing accessible chromatin regions (ACRs), revealing WRKY transcription factors as potential regulators of xylem development. Furthermore, we reanalysed scATAC-seq datasets from 8 plant species, spanning 11 tissues and multiple developmental stages, enabling cross-species comparisons. Furthermore, these analyses uncovered conserved regulatory programmes, including AP2/EREBP-associated ACRs linked to cell wall development and cell-type-conserved TFs across grasses. Collectively, scPlantReg provides a general framework and resource for comparative regulatory analysis in plants.

Epigenomics↗

Leveraging unlabeled SEM datasets with self-supervised learning for enhanced particle segmentation

Scanning Electron Microscopes (SEMs) are widely used in experimental science laboratories, often requiring cumbersome and repetitive user analysis. Automating SEM image analysis processes is highly desirable to address this challenge. In particle sample analysis, Machine Learning (ML) has emerged as the most effective approach for particle segmentation. However, the time-intensive process of manually annotating thousands of SEM images limits the applicability of supervised learning approaches. Self-Supervised Learning (SSL) offers a promising alternative by enabling knowledge extraction from raw, unlabeled data. This study presents a framework for evaluating SSL techniques in SEM image analysis, focusing on novel methods leveraging the ConvNeXtV2 architecture for particle detection. A dataset comprising 25,000 SEM images is curated to benchmark these proposed SSL methods. The results demonstrate that ConvNeXtV2 models, with varying parameter counts, consistently outperform other techniques in particle detection across different length scales, achieving up to a 34% reduction in relative error compared to established SSL methods. Furthermore, an ablation study explores the relationship between dataset size and SSL performance, providing actionable insights for practitioners regarding model selection and resource efficiency. This research advances the integration of SSL into autonomous analysis pipelines and supports its application in accelerating materials science discovery.

Rettenberger, Luca↗