Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Sequence Function Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

A family portrait of lanmodulin selectivity for enhanced rare-earth separations

Proteins offer a molecular design space to create bespoke ligands for the separation of critical metals like rare earth elements (REs). However, data-intensive approaches to tune metalloprotein selectivity are constrained by the low-throughput nature of existing characterization methods. Here we invented an assay called ‘SpyTag-Catcher Immobilization of Lanmodulin for Assaying Metal-Binding Selectivity’ (SpyCI-LAMBS) to measure metalloprotein selectivity en masse. This 96-format workflow was used to study the selectivity of 621 lanmodulin (LanM) orthologs for 15 REs, revealing eight distinct selectivity profiles based on sequence-to-function analyses. We discovered >200 LanMs with stronger selectivity against low-value LaIII relative to the prototypical LanM. This includes a LanM that can perform a challenging one-stage separation of PrIII from LaIII with up to >99.9 mol% purity and 83% yield. SpyCI-LAMBS is a powerful tool that can rapidly collect high-fidelity selectivity data to inform metal ion separations and machine-learning-assisted metalloprotein design.

59 BASIC BIOLOGICAL SCIENCES

GenomeFace v1.0

GenomeFace is meta-genome binning software. Metagenomic binning, the process of grouping DNA sequences into taxonomic units, is critical for understanding the functions, interactions, and evolutionary dynamics of microbial communities. We propose a deep learning approach to binning using two neural networks, one based on composition and another on environmental abundance, dynamically weighting the contribution of each based on characteristics of the input data. Trained on over 43,000 prokaryotic genomes, our network for composition-based binning is inspired by metric learning techniques used for facial recognition. Using a task-specific, multi-GPU accelerated algorithm to cluster the embeddings produced by our network, our binner leverages marker genes observed to be universally present in nearly all taxa to grade and select optimal clusters of sequences from a hierarchy of candidates. We evaluate our approach on four simulated datasets with known ground truth. Our linear time integration of marker genes recovers more near complete genomes than state of the art but computationally infeasible solutions using them, while being over an order of magnitude faster. Finally, we demonstrate the scalability and acuity of our approach by testing it on three of the largest metagenome assemblies ever performed. Compared to other binners, we produced 47%-183% more near complete genomes. From these datasets, we find over the genomes of over 3000 new candidate species which have never been previously cataloged, representing a potential 4% expansion of the known bacterial tree of life.

Lettich, Richard [Lawrence Berkeley National Labor

Data and scripts associated with a manuscript analyzing ELM-FATES parameter sensitivity under pre-fire and postfire scenarios using machine learning

NOTE: The manuscript associated with this data package is currently in review. The data may be revised based on reviewer feedback. Upon manuscript acceptance, this data package will be updated with the final dataset and additional metadata. This data package is associated with the manuscript “Fire Severity-Dependent Shifts in Vegetation Parameter Sensitivity: A Pre- and Post-Fire Analysis Using ELM-FATES and Explainable AI” submitted to Journal of Advances in Modeling Earth Systems (Zahura et al. 2026). The study examines vegetation physiological parameters controlling pre-fire and post-fire vegetation dynamics. To support this analysis, 73 vegetation parameters in Functionally Assembled Terrestrial Ecosystem Simulator (FATES) (Fisher et al., 2018) , which is coupled with E3SM (Energy Exascale Earth System Model) land model (ELM, ELM-FATES), were perturbed using a Sobol sequence to generate 1,024 ensemble members for two plant functional types: needleleaf evergreen extratropical trees (NEET) and C3 grass. Simulations were conducted for the pre-fire period (2016) and post-fire period (2018–2023). Burn severity was represented by modifying the Nesterov index in FATES to 75,000, 150,000, and 300,000 for low, moderate, and high severity, respectively. A no-fire scenario was also included. Simulations were performed for 16 grid cells in the American River Watershed across different burn severities and plant functional types. XGBoost (eXtreme Gradient Boosting) models were trained using the parameter ensembles and ELM-FATES-simulated outputs, including leaf area index (LAI), gross primary productivity (GPP), aboveground biomass, vegetation evaporation, transpiration, and soil evaporation. Models were trained separately for each year and burn severity, followed by SHAP (SHapley Additive exPlanations) analysis to identify changes in dominant parameters after fire disturbance. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. The data package contains the ELM-FATES simulation data. The scripts and data related to the analysis will be added later. The inputs and outputs from ELM-FATES are inside the “FATES” folder. “FATES_domain_surface” contains the domain and surface netcdfs that were used to run ELM-FATES in the study area. “FATES_parameters” contains the 1024 ensembles that were generated using Sobol sequence. “FATES_outputs” folder contains ELM-FATES simulated variables. All files are .csv and .nc (NetCDF).

Aboveground biomass

Functional characterization of glycosyltransferases in duckweed to enable predictive biology

Glycosyltransferases (GTs) catalyze the formation of glycosidic linkages to produce almost all complex carbohydrates. This project used a multi-disciplinary, high-throughput (HTP) biochemical and computational biology approach focused on duckweed as a model energy crop, to study carbohydrate metabolic processes. To achieve this, developed and carried out out high-throughput (HTP) functional characterization of plant glycosyltransferases (GTs) role of enzymatic microenvironments be assessed through a combined proteomic and computational biology approach, and the combined data was used to populate deep-learning frameworks to predict plant GT function. Functional validation achieved through this research is being used to assign gene function and study plant processes at the systems level to efficiently link the genome sequence with gene function. Together, the combined approaches used within this study provide a foundation for how computational prediction, in combination with high-throughput functional validation, can be used to study plant processes at the systems level and translate knowledge gained to efficiently link genome sequence with gene function in a species agnostic manner.

09 BIOMASS FUELS

Identification and characterization of a skin microbiome on Caenorhabditis elegans suggests environmental microbes confer cuticle protection

ABSTRACT In the wild, C. elegans are emersed in environments teeming with a veritable menagerie of microorganisms. The C. elegans cuticular surface serves as a barrier and first point of contact with their microbial environments. In this study, we identify microbes from C. elegans natural habitats that associate with its cuticle, constituting a simple “skin microbiome.” We rear our animals on a modified CeMbio, mCeMbio, a consortium of ecologically relevant microbes. We first combine standard microbiological methods with an adapted micro skin-swabbing tool to describe the skin-resident bacteria on the C. elegans surface. Furthermore, we conduct 16S rRNA gene sequencing studies to identify relative shifts in the proportion of mCeMbio bacteria upon surface-sterilization, implying distinct skin- and gut-microbiomes. We find that some strains of bacteria, including Enterobacter sp. JUb101 , are primarily found on the nematode skin, while others like Stenotrophomonas indicatrix JUb19 and Ochrobactrum vermis MYb71 are predominantly found in the animal’s gut. Finally, we show that this skin microbiome promotes host cuticle integrity in harsh environments. Together, we identify a skin microbiome for the well-studied nematode model and propose its value in conferring host fitness advantages in naturalized contexts. IMPORTANCE The genetic model organism C. elegans has recently emerged as a tool for understanding host–microbiome interactions. Nearly all of these studies either focus on pathogenic or gut-resident microbes. Little is known about the existence of native, nonpathogenic skin microbes or their function. We demonstrate that members of a modified C. elegans model microbiome, mCeMbio, can adhere to the animal's cuticle and confer protection from noxious environments. We combine a novel micro-swab tool, the first 16S microbial sequencing data from relatively unperturbed C. elegans , and physiological assays to demonstrate microbially mediated protection of the skin. This work serves as a foundation to explore wild C. elegans skin microbiomes and use C. elegans as a model for skin research.

16S RNA

RolyPoly (rp) v0.1.0

The Rolypoly pipeline is designed to process raw RNA-seq data and identify potential RNA viral sequences. It is split into several self contained steps: 1. input data filtering and QC, 2. Genome assembly and refinement, 3. Assembly filtering, 4. Mapping to known RNA viral genomes, 5. Searching for RNA viral marker genes. 6. Genome functional and structural annotation. 6. Report preparation and potential downstream analysis The last module, may include taxonomic assignment, host range estimation, and phenotypic prediction. There are many similar software, but they focus on human related viruses, and lack the downstream applications or differ in their sensitivity. The initial user base are non-computational microbial ecologists who wish to better understand the potential RNA viruses in their own generated samples.

Neri, Uri

Protocol for applying a network-enabled gene discovery pipeline to non-model plant species

Identifying upstream regulators of key genes is essential for understanding gene regulatory mechanisms and translating these insights into functional targets. Here, we present a protocol for applying the network-enabled gene discovery pipeline (NEEDLE) to non-model plant species. We describe steps for environment setup, data preparation, computational analysis, expected outputs, and parameter considerations. NEEDLE integrates RNA sequencing (RNA-seq) processing, weighted gene co-expression analysis (WGCNA), Gene Network Inference with Ensemble of trees (GENIE3), and promoter conservation analysis to prioritize candidate transcriptional regulators.

Plant Sciences

Single-cell proteomics of Arabidopsis leaf mesophyll reveals dynamic protein responses to water-deficit stress

Background The application of single-cell omics tools to biological systems can provide unique insights into diverse cellular populations and their heterogeneous responses to internal and external perturbations. Thus far, most single-cell studies in plant systems have been limited to RNA-sequencing approaches, which only provide indirect readouts of cellular functions. Results Here, we present a single-cell proteomics workflow for plant cells that integrates tape-sandwich protoplasting, piezoelectric cell sorting, nanoPOTS sample preparation, and ion mobility-based MS data acquisition method for label-free single-cell proteomics analysis of Arabidopsis leaf mesophyll cells. From a single leaf protoplast, over 3,000 proteins were quantified with high precision. The workflow is demonstrated to identify stress associated changes in protein abundance by analyzing 117 protoplasts from well-watered and water-deficit stressed plants. Additionally, we describe a new approach for constructing covarying protein networks at the single-cell level and demonstrate how single-cell protein covariation analysis can reveal previously unrecognized protein functions while also capturing stress-induced changes in protein–protein dynamics. Conclusions The label-free scProteomic approach presented here represents a significant advance through the demonstration of a facile protoplast isolation method combined with deep and precise proteomic coverage of Arabidopsis leaf mesophyll cell types. We believe this study will serve as an informative reference to future plant scProteomic investigations.

Arabidopsis

Signatures of Mollicutes-related endobacteria in publicly available Mucoromycota genomes

ABSTRACT Mucoromycota fungi and their Mollicutes-related endobacteria (MRE) are an ideal system for studying bacterial–fungal interactions and evolution due to the long-term and intimate nature of their interactions. However, methods for detecting MRE face specific challenges due to the poor representation of MRE in sequencing databases coupled with the high sequence divergence of their genomes, making traditional similarity searches unreliable. This has precluded estimations on the diversity of MRE associated with Mucoromycota. To determine the prevalence of previously undetected MRE in fungal genome sequences, we scanned 389 Mucoromycota genome assemblies available from the National Center for Biotechnology Information for the presence of MRE sequences using publicly available tools to map contigs from fungal assemblies to publicly available MRE genomes. We demonstrate a higher diversity of MRE genomes than previously described in Mucoromycota and a lack of cophylogeny between MRE and the majority of their fungal hosts. This supports the late invasion hypothesis regarding MRE acquisition across most of the examined fungal families. In contrast with other Mucoromycota lineages, MRE from the Gigasporaceae displayed some degree of cophylogeny with their hosts, which may indicate that horizontal transmission is restricted between members of this family or that transmission is strictly vertical. These results underscore the need for a refined process to capture sequencing data from potential fungal endosymbionts to discern their evolution and transmission. Screens of fungal genomes for MRE can help improve the quality of fungal genome assemblies while identifying new MRE lineages to further test hypotheses on their origin and evolution. IMPORTANCE Mollicutes-related endobacteria (MRE) are obligate intracellular bacteria found within Mucoromycota fungi. Despite their frequent detection, MRE roles in host functioning are still unknown. Comparative genomic investigations can improve our understanding of the impact of MRE on their fungal hosts by identifying similarities and differences in MRE genome evolution. However, MRE genomes have only been assembled from a small fraction of Mucoromycota hosts. Here, we demonstrate that MRE can be present yet undetected in publicly available Mucoromycota genome assemblies. We use these newfound sequences to assess the broader diversity of MRE and their phylogenetic relationships with respect to their hosts. We demonstrate that publicly available tools can be used to extract novel MRE sequences from assembled fungal genomes leading to insights on MRE evolution. This work contributes to a greater understanding of the fungal microbiome, which is crucial to improving knowledge on the dynamics and impacts of fungi in microbial ecosystems.

59 BASIC BIOLOGICAL SCIENCES

Dynamic Temporal Graph Sequence Data for Resilience-Oriented Distribution Network Reconfiguration

This dataset comprises temporal dynamic graph sequences generated from power grid simulations focused on grid reconfiguration to enhance resilience. The simulations model failure propagation under varying conditions, with nodes assigned distinct failure probabilities. For each time step, the dataset captures the evolution of node states (functional or failed) and features critical to grid operations, such as pv_output, load_profile, load_dispatch, dg_output, loss, and voltage. Node types include sources, normal loads, and nodes with specific equipment like PVs, micro turbines, or shunt capacitors. The dataset is structured to support the training of dynamic graph neural networks, facilitating research on node feature prediction and edge dynamics under failure scenarios. Three distinct configurations are included, providing a robust foundation for modeling power grid resilience.

29 ENERGY PLANNING, POLICY, AND ECONOMY

Bottom-Up Simulation, Reconstruction, and Quantification of Macromolecule Sequences from Experimental Polymerizations

Motivated by the canonical sequence–structure–function paradigm, tools to characterize chemical patterning in natural biomacromolecules, from proteins to nucleic acids, have grown exponentially in recent years. However, analogous strategies for synthetic macromolecules remain in nascent stages, complicated by sequence polydispersity and analytical limitations. To address this, we have developed a comprehensive and open-source Python package, PRISM (polymer rate insights and sequence modeling), an end-to-end workflow that provides a path from experimental kinetics measurements to quantitative and qualitative metrics for describing chemical patterning in stochastic polymers. First, a numerical integration strategy was constructed to simulate and fit experimental data from reversible addition–fragmentation chain transfer (RAFT) polymerization kinetics, enabling the facile estimation of relevant reactivity ratios. These ratios were then used in a mechanism-specific stochastic kinetic simulation strategy to simulate sequence ensembles corresponding to model systems spanning experimental copolymers, classes of statistical polymers (e.g., alternating, block, and gradient), and multiblock copolymers. Lastly, inspired by sequence homology metrics from bioinformatics, we introduce visualization strategies and quantitative metrics to facilitate comparisons of different sequence ensembles. As the sequence–structure–function paradigm becomes increasingly central in de novo design of synthetic macromolecules, this toolkit provides a first step toward accurate and representative sequence description and featurization.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

A combinatorially complete epistatic fitness landscape in an enzyme active site

Protein engineering often targets binding pockets or active sites which are enriched in epistasis—nonadditive interactions between amino acid substitutions—and where the combined effects of multiple single substitutions are difficult to predict. Few existing sequence-fitness datasets capture epistasis at large scale, especially for enzyme catalysis, limiting the development and assessment of model-guided enzyme engineering approaches. We present here a combinatorially complete, 160,000-variant fitness landscape across four residues in the active site of an enzyme. Assaying the native reaction of a thermostable β-subunit of tryptophan synthase (TrpB) in a nonnative environment yielded a landscape characterized by significant epistasis and many local optima. These effects prevent simulated directed evolution approaches from efficiently reaching the global optimum. There is nonetheless wide variability in the effectiveness of different directed evolution approaches, which together provide experimental benchmarks for computational and machine learning workflows. The most-fit TrpB variants contain a substitution that is nearly absent in natural TrpB sequences—a result that conservation-based predictions would not capture. Thus, although fitness prediction using evolutionary data can enrich in more-active variants, these approaches struggle to identify and differentiate among the most-active variants, even for this near-native function. Overall, this work presents a large-scale testing ground for model-guided enzyme engineering and suggests that efficient navigation of epistatic fitness landscapes can be improved by advances in both machine learning and physical modeling.

biocatalysis

Data for Comparison of Genotyping Assays for Detection of Targeted CRISPR/Cas Mutagenesis in Highly Polyploid Sugarcane

Sugarcane ( Saccharum spp.) is an important biofuel feedstock and a leading source of global table sugar. Saccharum hybrid cultivars are highly polyploid (2n = 100–130), containing large numbers of functionally redundant hom(e)ologs in their genomes. Genome editing with sequence-specific nucleases holds tremendous promise for sugarcane breeding. However, identification of plants with the desired level of co-editing within a pool of primary transformants can be difficult. While DNA sequencing provides direct evidence of targeted mutagenesis, it is cost-prohibitive as a primary screening method in sugarcane and most other methods of identifying mutant lines have not been optimized for use in highly polyploid species. In this study, non-sequencing methods of mutant screening, including capillary electrophoresis (CE), Cas9 RNP assay, and high-resolution melt analysis (HRMA), were compared to assess their potential for CRISPR/Cas9-mediated mutant screening in sugarcane. These assays were used to analyze sugarcane lines containing mutations at one or more of six sgRNA target sites. All three methods distinguished edited lines from wild type, with co-mutation frequencies ranging from 2% to 100%. Cas9 RNP assays were able to identify mutant sugarcane lines with as low as 3.2% co-mutation frequency, and samples could be scored based on undigested band intensity. CE was highlighted as the most comprehensive assay, delivering precise information on both mutagenesis frequency and indel size to a 1 bp resolution across all six targets. This represents an economical and comprehensive alternative to sequencing-based genotyping methods which could be applied in other polyploid species.

Genomics

Microbial spies and bloggers: programming cells to convert environmental information into discernible signals

Microbes regulate their dynamic behaviors using the chemical and physical characteristics of their environment. The ability of microbes to continuously convert this physicochemical information into biochemical information and to use organic matter in the environment as a power source makes these organisms attractive as chassis for building sensors. However, most biosensors have severe limitations when considering applications in hard-to-image settings like soils, sediments, and wastewater. Emerging technologies at the interface of biomolecular design, microbiome engineering, and synthetic biology offer new tools to program cells and communities as biosensors for these settings. Here, in this review, we describe innovations in biosensor outputs that are enabling new applications in complex environments, including reporters that are read out using electrochemical, gas chromatography, hyperspectral imaging, and next-generation sequencing methods. We also discuss computational advances that are accelerating the diversification of sensing components by mining metagenomics data for new transcriptional regulators and by designing allosteric protein switches that directly regulate reporter outputs using analytes. We highlight emerging opportunities for programming undomesticated microbes in communities to function as distributed sensors in the environment. Finally, we discuss the need for responsible biosensor development and to modernize regulatory frameworks to support evidence-based assessment of environmental biosensors.

analyte

The secondary metabolism collaboratory: a database and web discussion portal for secondary metabolite biosynthetic gene clusters

Secondary metabolites are small molecules produced by all corners of life, often with specialized bioactive functions with clinical and environmental relevance. Secondary metabolite biosynthetic gene clusters (BGCs) can often be identified within DNA sequences by various sequence similarity tools, but determining the exact functions of genes in the pathway and predicting their chemical products can often only be done by careful, manual comparative analysis. To facilitate this, we report the first release of the secondary metabolism collaboratory (SMC), which aims to provide a comprehensive, tool-agnostic repository of BGC sequence data drawn from all publicly available and user-submitted bacterial and archaeal genome and contig sources. On the website, users are provided a searchable catalog of putative BGCs identified from each source, along with visualizations of gene and domain annotations derived from multiple sequence analysis tools. SMC’s data is also available through publicly-accessible application programming interface (API) endpoints to facilitate programmatic access. Users are encouraged to share their findings (and search for others’) through comment posts on BGC and source pages. At the time of writing, SMC is the largest repository of BGC information, holding 13.1M BGC regions from 1.3M source sequences and growing, and can be found at https://smc.jgi.doe.gov.

59 BASIC BIOLOGICAL SCIENCES

favela3/Maize.N-cycle.Function

Supplemental sequence processing and R statistical analysis for publication which compares the microbiome of 27 Zea cultivars: 12 Inbred maize genotypes, 9 hybrids, and 6 wild teosinte. The project contains amplicon data for various genes: 16S rRNA, ITS, bacterial amoA, Archeal amoA, nirS, nirK, and nosZ. In addition to functional potential assay data, and N2O flux.

Favela, Alonso

GenomeDepot v1.0

GenomeDepot is a web-based platform for annotation, management, and comparative analysis of microbial genomic sequences and associated data including ortholog families, protein domains, operons, regulatory interactions, strain taxonomy, and sample metadata. GenomeDepot supports rapid creation of web-sites for user-defined genome collections that include bioinformatic tools for interactive genome browsing, BLAST search, annotation search, comparative genomic neighborhood visualization, and sequence download. Gene function annotations are generated by a customizable annotation pipeline. The pipeline runs annotation tools in Conda environments and can be easily extended with additional user-specified tools.

Kazakov, Alexey [Lawrence Berkeley National Labora

Discovery of additional ancient genome duplications in yeasts

Whole-genome duplication (WGD) has had profound macroevolutionary impacts on diverse lineages, preceding adaptive radiations in vertebrates, teleost fish, and angiosperms. In contrast to the many known ancient WGDs in animals, and especially plants, we are aware of evidence for only four WGDs in fungi. The oldest of these occurred ∼100 million years ago (mya) and is shared by ∼60 extant Saccharomycetales species, including the baker’s yeast Saccharomyces cerevisiae. Notably, this is the only known ancient WGD event in the yeast subphylum Saccharomycotina. The dearth of ancient WGD events in fungi remains a mystery. Some studies have suggested that fungal lineages that experience chromosome and genome duplication quickly go extinct, leaving no trace in the genomic record, while others contend that the lack of known WGDs is due to an absence of data. Under the second hypothesis, additional sampling and deeper sequencing of fungal genomes should lead to the discovery of more WGD events. Coupling hundreds of recently published genomes from nearly every described Saccharomycotina species, with three additional long-read assemblies, we discovered three novel WGD events. Although the functions of retained duplicate genes originating from these events are broad, they bear similarities to the well-known WGD that occurred in the Saccharomycetales. In conclusion, our results suggest that WGD may be a more common evolutionary force in fungi than previously believed.

convergent evolution