Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Sequence Analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Unique Structural Features Relate to Evolutionary Adaptation of Cytochrome P450 in the Abyssal Zone

Cytochromes P450 (CYPs) form one of the largest enzyme superfamilies, with similar structural folds yet biological functions varying from synthesis of physiologically essential compounds to metabolism of myriad xenobiotics. Sterol 14α-demethylases (CYP51s) represent a very special P450 family, regarded as a possible evolutionary progenitor for all currently existing P450s. In metazoans CYP51 is critical for the biosynthesis of sterols including cholesterol. Here we determined the crystal structures of ligand-free CYP51s from the abyssal fish Coryphaenoides armatus and human-. Comparative sequence–structure–function analysis revealed specific structural elements that imply elevated conformational flexibility, uncovering a molecular basis for faster catalytic rates, lower substrate selectivity, and intrinsic resistance to inhibition. In addition, the C. armatus structure displayed a large-scale repositioning of structural segments that, in vivo, are immersed in the endoplasmic reticulum membrane and border the substrate entrance (the FG arm, >20 Å, and the β4 hairpin, >15 Å). The structural distinction of C. armatus CYP51, which is the first structurally characterized deep sea P450, suggests stronger involvement of the membrane environment in regulation of the enzyme function. We interpret this as a co-adaptation of the membrane protein structure with membrane lipid composition during evolutionary incursion to life in the deep sea.

Biochemistry & Molecular Biology↗

The secondary metabolism collaboratory: a database and web discussion portal for secondary metabolite biosynthetic gene clusters

Secondary metabolites are small molecules produced by all corners of life, often with specialized bioactive functions with clinical and environmental relevance. Secondary metabolite biosynthetic gene clusters (BGCs) can often be identified within DNA sequences by various sequence similarity tools, but determining the exact functions of genes in the pathway and predicting their chemical products can often only be done by careful, manual comparative analysis. To facilitate this, we report the first release of the secondary metabolism collaboratory (SMC), which aims to provide a comprehensive, tool-agnostic repository of BGC sequence data drawn from all publicly available and user-submitted bacterial and archaeal genome and contig sources. On the website, users are provided a searchable catalog of putative BGCs identified from each source, along with visualizations of gene and domain annotations derived from multiple sequence analysis tools. SMC’s data is also available through publicly-accessible application programming interface (API) endpoints to facilitate programmatic access. Users are encouraged to share their findings (and search for others’) through comment posts on BGC and source pages. At the time of writing, SMC is the largest repository of BGC information, holding 13.1M BGC regions from 1.3M source sequences and growing, and can be found at https://smc.jgi.doe.gov.

59 BASIC BIOLOGICAL SCIENCES↗

Host analysis-guided selection and targeted engineering (HASTE) of Lipomyces tetrasporus for the conversion of CO2-derived feedstocks

Efficient and cost-competitive bioproduction calls for utilizing CO2-derived feedstocks, such as products from electro-reduction of CO2 and hydrolysate from lignocellulosic biomass. However, efficiently using all their carbon components, including acetate, glucose, and xylose, remains a challenge. Here, we characterize Lipomyces tetrasporus, a novel, robust yeast strain capable of effectively assimilating these carbon sources. We used an integrated systems biology approach combining ¹³C metabolic flux analysis, dynamic labeling experiments, and RNA sequencing. We conducted the first metabolic flux analysis for glucose, xylose, and acetate catabolism in this species. Dynamic labeling revealed a highly active TCA cycle during acetate metabolism, evidenced by rapid citrate and malate accumulation. The strain demonstrated strong NADH/NADPH production and acetyl-CoA synthase activity. Using insights and gene targets from this analysis, we engineered L. tetrasporus for malate production. The engineered strain produced 7.5 g/L malic acid (0.25 g/g yield) in shake flasks with glucose-acetate media and 28.8 g/L malic acid at a yield of 0.20 g/g in fed-batch mode with corn-stover hydrolysate. Together, these insights and rational strain engineering establish L. tetrasporus as a versatile, Crabtree-negative platform that is an energy-CO2-bioproduction nexus for channeling CO2 carbon into value-added bioproducts.

Xiao, Zhengyang↗

Genomic factors shaping codon usage across the Saccharomycotina subphylum

Codon usage bias, or the unequal use of synonymous codons, is observed across genes, genomes, and between species. It has been implicated in many cellular functions, such as translation dynamics and transcript stability, but can also be shaped by neutral forces. We characterized codon usage across 1,154 strains from 1,051 species from the fungal subphylum Saccharomycotina to gain insight into the biases, molecular mechanisms, evolution, and genomic features contributing to codon usage patterns. We found a general preference for A/T-ending codons and correlations between codon usage bias, GC content, and tRNA-ome size. Codon usage bias is distinct between the 12 orders to such a degree that yeasts can be classified with an accuracy >90% using a machine learning algorithm. We also characterized the degree to which codon usage bias is impacted by translational selection. We found it was influenced by a combination of features, including the number of coding sequences, BUSCO count, and genome length. Our analysis also revealed an extreme bias in codon usage in the Saccharomycodales associated with a lack of predicted arginine tRNAs that decode CGN codons, leaving only the AGN codons to encode arginine. Analysis of Saccharomycodales gene expression, tRNA sequences, and codon evolution suggests that avoidance of the CGN codons is associated with a decline in arginine tRNA function. Consistent with previous findings, codon usage bias within the Saccharomycotina is shaped by genomic features and GC bias. However, we find cases of extreme codon usage preference and avoidance along yeast lineages, suggesting additional forces may be shaping the evolution of specific codons.

59 BASIC BIOLOGICAL SCIENCES↗

Hybridization capture sequencing for Vibrio spp. and associated virulence factors

ABSTRACT Proliferation ofVibriospp. in aquatic ecosystems is associated with climate change and, concomitantly, increased incidence of vibriosis. They are autochthonous to aquatic environments globally, but traditional metagenomic methods for detecting and typing pathogenicVibriospp. are challenged by their presence in relatively low abundance and ability to persist in a viable but nonculturable state. In the study reported here, hybridization capture sequencing (HCS) was employed to profile low-abundanceVibriospp. in environmental samples. The HCS panel targeted a family of molecular chaperones (CPN60) specific to 69Vibriospp. and 162Vibrio-specific virulence factors. This approach was evaluated in parallel with traditional whole-community shotgun sequencing in a metagenomic analysis of water and oyster samples collected from the Chesapeake Bay. In addition,Vibrio parahaemolyticusandVibrio vulnificusstrains isolated from the samples were subjected to whole-genome sequencing to determine the genetic characteristics of pathogenicVibriospp. circulating in an aquatic environment. HCS, employed to determine the incidence and characterization of specificVibriospp., yielded significantly greater metagenomic insight, notably a variety of otherVibriospp., including detection ofVibrio cholerae,Vibrio fluvialis, andVibrio aestuarianus, in addition toVibrio parahaemolyticusandVibrio vulnificus, and also important virulence factors not detectable using traditional molecular methods. Thus, pathogenicVibriospp. in aquatic ecosystems may be far more common than currently understood. It is concluded that environmental surveillance should include HCS, a valuable tool for the detection and characterization of pathogenic agents in aquatic ecosystems, notably vibrios. IMPORTANCE The increasing prevalence of pathogenicVibriospp. in aquatic ecosystems, driven by climate change, is closely linked to a rise in cholera and vibriosis cases, emphasizing the need for improved environmental surveillance. Vibrios are naturally occurring in aquatic environments globally, but traditional metagenomic methods for detecting and typing pathogenicVibriospp. are challenged by their presence in relatively low abundance and ability to persist in a viable but nonculturable state. In the study reported here, hybridization capture sequencing was employed to profile low-abundanceVibriospp. in metagenomic samples, namely water and oysters collected from the Chesapeake Bay. This approach was evaluated in parallel with traditional whole-community shotgun sequencing and whole-genome sequencing ofVibrio parahaemolyticusandVibrio vulnificusstrains isolated from the samples. Results suggest pathogenicVibriospp. in aquatic ecosystems may be far more common than currently understood, when multiple methods are considered for environmental surveillance.

Microbiology↗

Metagenome-assembled genomes provide insight into the metabolic potential during early production of Hydraulic Fracturing Test Site 2 in the Delaware Basin

Demand for natural gas continues to climb in the United States, having reached a record monthly high of 104.9 billion cubic feet per day (Bcf/d) in November 2023. Hydraulic fracturing, a technique used to extract natural gas and oil from deep underground reservoirs, involves injecting large volumes of fluid, proppant, and chemical additives into shale units. This is followed by a “shut-in” period, during which the fracture fluid remains pressurized in the well for several weeks. The microbial processes that occur within the reservoir during this shut-in period are not well understood; yet, these reactions may significantly impact the structural integrity and overall recovery of oil and gas from the well. To shed light on this critical phase, we conducted an analysis of both pre-shut-in material alongside production fluid collected throughout the initial production phase at the Hydraulic Fracturing Test Site 2 (HFTS 2) located in the prolific Wolfcamp formation within the Permian Delaware Basin of west Texas, USA. Specifically, we aimed to assess the microbial ecology and functional potential of the microbial community during this crucial time frame. Prior analysis of 16S rRNA sequencing data through the first 35 days of production revealed a strong selection for a Clostridia species corresponding to a significant decrease in microbial diversity. Here, we performed a metagenomic analysis of produced water sampled on Day 33 of production. This analysis yielded three high-quality metagenome-assembled genomes (MAGs), one of which was a Clostridia draft genome closely related to the recently classified Petromonas tenebris. This draft genome likely represents the dominant Clostridia species observed in our 16S rRNA profile. Annotation of the MAGs revealed the presence of genes involved in critical metabolic processes, including thiosulfate reduction, mixed acid fermentation, and biofilm formation. These findings suggest that this microbial community has the potential to contribute to well souring, biocorrosion, and biofouling within the reservoir. Our research provides unique insights into the early stages of production in one of the most prolific unconventional plays in the United States, with important implications for well management and energy recovery.

natural gas↗

A Chemoselective and Stereodivergent Platform of Heme‐Nitrene Transferases to Access Chiral Aryl‐β‐Amino Esters and An Investigation of the Sequence‐Activity Landscape

Engineered biocatalysts can utilize nitrene precursors to access enantioenriched amination products, yet they have not been applied to produce valuable, enantiomerically enriched noncanonical β-amino esters. Current approaches to synthesizing β-amino acids rely on pre-oxidized precursors and multistep synthetic approaches involving various protecting groups. We engineered a platform of heme enzymes for stereoselective C–H bond amination of readily available carboxylic ester derivatives to install primary amines. A directed evolution campaign coupled with sequencing of over 1000 variants enabled us to develop engineered variants that use either O-pivaloylhydroxylamine triflic acid (PONT) or hydroxylamine hydrochloride (H 2 NOH∙HCl) as aminating reagents. An analysis of the resulting sequence–activity dataset revealed additional improvements that could be made to the final variant, highlighting the utility of sequencing data to guide future steps in directed evolution campaigns. Furthermore, the evolved nitrene transferases expand the scope of accessible chiral β-amino acid building blocks for peptidomimetic applications and provide new starting points for the design and synthesis of enantioenriched β-amino acid motifs.

amino ester building blocks↗

Rhizosphere Microbiome Diversity Potentially Supports Robust Nature of Field Pennycress ( Thlaspi arvense L.) in Dryland Cropping Systems of Eastern Washington

ABSTRACT Field pennycress ( Thlaspi arvense L.) is an annual in the Brassicaceae family and is currently being developed as an oilseed intermediate crop suitable for renewable biodiesel and jet fuel. It displays many desirable characteristics for this role including cold tolerance, a rapid life cycle, and a seed fatty acid profile conducive to bioenergy generation. These traits make field pennycress favorable for winter oilseed cultivation in the inland Pacific Northwest (iPNW). Simultaneously, intermediate crops are an increasingly recognized component of both agronomic sustainability and soil health management. Intermediate crops enhance soil microbial diversity, which benefits both soil and plant health. To understand the impact of field pennycress on soil microbial diversity, two natural accessions and seven experimental accessions were grown at three sites in Eastern Washington. Aboveground biomass and rhizosphere soil were then collected. Soil genomic DNA was extracted from rhizosphere samples and used to generate an amplicon library for bacterial (16S) and fungal (ITS) rRNA sequences. The resulting libraries were analyzed in QIIME2, which revealed that not only did the fad2 deficient line from the Spring32‐10 background have significantly increased aboveground biomass production compared to other pennycress genotypes, but also displayed significantly higher β‐diversity in the rhizosphere community specifically at the site experiencing the driest conditions. ANCOM analysis showed that multiple sequences similar to beneficial plant and soil health enhancing organisms such as Trichoderma spirale , Pseudomonas spp., and Methylobacterium goesingense were found to be enriched in the microbiome of the fad2 Spring32‐10 background also at that site. To add additional context to rhizosphere community data, root exudates from two pennycress genotypes were captured in magenta boxes and analyzed using HPLC. Future work will expand our understanding of the mechanisms by which field pennycress creates diversity in the rhizosphere, thus expanding our ability to cultivate this crop in the iPNW.

54 ENVIRONMENTAL SCIENCES↗

Enhanced Resistance Pines for Improved Renewable Biofuel and Chemical Production (Technical Report)

We completed phenotyping constitutive and inducible oleoresin flow across two seasons, constitutive resin canal number and density and wood terpene content in our ADEPT2 and CCLONES populations. We completed genetic association between 19 oleoresin phenotypes and a total of 523,192 SNP markers from ADEPT2 and 13,883 SNP markers in CCLONES using four mixed linear models. A total of 293 significant SNPs (FDR = 0.20) were identified. We used the MENTOR tool to mine mechanistic connections from a multiplex network constructed from poplar multi-omic data to construct a conceptual model for a subset of these significant SNPs. Our model contains 6 transcriptional regulators in addition to 3 monoterpene synthases. To generate more lines of evidence for these significant SNPs, we completed a time course RNAseq experiment after inducing vascular zone cells to differentiate into new resin canals with a methyl jasmonate treatment, a single nuclei RNAseq that identified differentiating resin canal epithelial cells and are completing analysis for a QTL study in a hybrid pine population. The time course identified 4634 significantly down and 1890 significantly up regulated transcripts after treatment with methyl jasmonate, an inducer of new resin canal formation in the vascular cambial meristem. To analyze this large set of differentially regulated genes, we created a predictive expression network and analyzed it with random walk restart using 6 seed genes coding for transcription factors regulating xylem differentiation in poplar. Of the top ranked 200 transcripts, 119 transcripts were significant differentially expressed supporting these transcripts as potential candidates regulating resin canal formation. Analysis of single nuclei sequencing of shoot tips that contain differentiating resin canals, identified 10 clusters. One cluster was highly enriched in transcripts coding for 9 of the enzymes in the MEP pathway 3 prenyl synthetases, and 3 monoterpene synthases strongly suggesting that this cluster represents resin canal epithelial cells. We are mining the additional transcripts to create a trajectory analysis. In summary, we have identified > 10 novel genes that are strongly supported candidates for further analysis in breeding lines and for genetic engineering over- and under- expressing lines to increase wood terpene content to improve resistance to insect and fungal pathogens while simultaneously increasing terpene supplies for renewable chemicals and biofuels.

59 BASIC BIOLOGICAL SCIENCES↗

High-Power Clock Laser Spectrally Tailored for High-Fidelity Quantum State Engineering

Highly frequency-stable lasers are ubiquitous tools for optical-frequency metrology, precision interferometry, and quantum information science. While making a universally applicable laser is unrealistic, spectral noise can be tailored for specific applications. Here we report a high-power 698-nm clock laser with a maximum output of 4W and minimized frequency noise up to a few kHz Fourier frequency, together with long-term instability of 3.5 × 10 −17 at one to thousands of seconds. The laser-frequency noise is precisely characterized with atom-based spectral analysis that employs a pulse sequence designed to suppress sensitivity to intensity noise. This method provides universally applicable tunability of the spectral response and analysis of quantum sensors over a wide frequency range. With the optimized laser system characterized by this technique, we achieve an average single-qubit Clifford gate fidelity of up to 𝐹$^2_1$ = 0.999⁢64⁢(3) when simultaneously driving 3000 optical qubits with a homogeneous Rabi frequency ranging from 10 Hz to 1 kHz. This result represents the highest single optical-qubit-gate fidelity for a large number of atoms.

atomic gases↗

Pressure–Temperature–Magnetic Field Phase Diagram of Multiferroic (NH 4 ) 2 FeCl 5 ·H 2 O

We combined synchrotron-based infrared absorbance and Raman scattering spectroscopies with diamond anvil cell techniques and a symmetry analysis to explore the properties of multiferroic (NH 4 ) 2 FeCl 5 ·H 2 O under extreme pressure–temperature conditions. Compression-induced splitting of the Fe–Cl stretching, Cl–Fe–Cl and Cl–Fe–O bending, and NH 4 + librational modes defines two structural phase transitions, and a group–subgroup analysis reveals space group sequences that vary depending upon proximity to the unexpectedly wide order–disorder transition. Here, we bring these findings together with prior high-field work to develop the pressure–temperature–magnetic field phase diagram uncovering competing polar, chiral, and magnetic phases in this system.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Multimodal framework for the joint analysis of single-cell RNA and T cell receptor sequencing data predicts T cell response to cancer immunotherapy

T cell states are prognostic in different cancer types. Recent technologies enable joint profiling of T cell RNA and T cell receptor (TCR) sequences at single-cell resolution. Here we present the TCR-RNA Integrating Model (TRIM), a multi-modal variational autoencoder framework that integrates RNA-TCR data and predicts T cell clonality and transcriptional states. TRIM learns a shared representation of the data conditioned on patient, tissue source, and treatment timepoint. We applied TRIM to three independent datasets that included T cells collected before and after checkpoint inhibitor treatment, sourced either from blood and tumor biopsies in patients with head and neck squamous cell carcinoma and colorectal cancer, or from tumor and adjacent tissue in a pan-cancer dataset. In all settings, TRIM accurately predicted intra-tumor T cell clonal expansion and transcriptional status based on T cells from blood or normal tissue before treatment, demonstrating its utility in modeling multimodal T cell data and predicting T cell response to treatment and disease progression.

60 APPLIED LIFE SCIENCES↗

Deficiency in transmitter release triggers homeostatic transcriptional changes that increase presynaptic excitability

Weakening of synaptic transmission at theDrosophilalarval neuromuscular junction triggers two forms of homeostatic compensation, one that increases the probability of glutamate release per action potential (P r ) and another that increases motoneuron (MN) activity. We investigated the molecular changes in MNs that underlie the increase in MN activity. RNA sequencing (RNA-seq) analysis on MNs whose glutamate release is weakened by knockdown of components of the MN transmitter release machinery reveals a reduction in expression of a group of genes that encode potassium channels and their positive modulators. These results identify a mechanism of compensation for weakened synaptic transmission by MNs, which engages a transcriptional program in those cells to increase firing and, thereby, ensure sufficient locomotory drive.

Science & Technology - Other Topics↗

GenomeDepot: data management system for microbial comparative genomics

Summary GenomeDepot is an open-source web-based platform for annotation, management, and comparative analysis of microbial genomic sequences and associated data including ortholog families, protein domains, operons, regulatory interactions, strain taxonomy, and sample metadata. GenomeDepot supports rapid creation of websites for user-defined genome collections that include bioinformatic tools for interactive genome browsing, Basic Local Alignment Search Tool (BLAST) search, annotation search, comparative genomic neighborhood visualization, and sequence download. Gene function annotations are generated by a customizable annotation pipeline. The pipeline runs annotation tools in Conda environments and can be easily extended with additional user-specified tools. Availability and implementation GenomeDepot is open source and distributed under the GNU General Public License via GitHub (https://github.com/aekazakov/genome-depot). GenomeDepot is implemented in Python and was tested in Ubuntu Linux. Full installation instructions and documentation are available at https://aekazakov.github.io/genome-depot/. GenomeDepot demo server is freely accessible at https://iseq.lbl.gov/demogd/.

Kazakov, Alexey [Lawrence Berkeley National Labora↗

CAHS: Context-Aware Homology Search

Protein homology search is foundational to bioinformatics: it supports annotation transfer, structure/function inference, and evolutionary analysis over rapidly expanding sequence repositories (e.g., UniProtKB). Profile hidden Markov models (pHMMs), as implemented in HMMER, remain the most widely trusted approach because they provide statistically calibrated E-values; however, their gap behavior is fixed once a profile is trained, despite biological evidence that insertion/deletion tolerance varies across flexible loops and intrinsically disordered regions. We present CAHS (Context-Aware Homology Search), a lightweight query-time adapter for pHMM search that incorporates learned and biologically motivated signals without changing HMMER's downstream search pipeline or its calibrated E-value reporting. Given a query sequence, CAHS computes per-residue representations from a protein language model and a disorder predictor, maps these to profile coordinates, and modulates only match-state transition rows (gap-open and gap-extension probabilities) while preserving Plan7 constraints. We comprehensively evaluate CAHS across six structurally diverse protein families and multi-domain architectures against a 570k-sequence target corpus. CAHS expands detection capability, retrieving thousands of additional remote homologs at relaxed thresholds by maintaining alignment quality through flexible regions. For multi-domain proteins, context-aware modulation resolves 94% of fragmented alignments. Crucially, CAHS preserves hit-set invariance at stringent operating points (E<10-10), demonstrating increased statistical confidence without inflating false positives. Furthermore, sharper statistical distinction between homologs and background noise during early filter stages yields up to a 3.87× acceleration in end-to-end wall-clock time on high-performance computing clusters. Overall, CAHS illustrates a practical AI-for-science design pattern: augmenting a trusted probabilistic model with query-specific learned signals to improve interpretable, reproducible inference in data-rich biology.

Bhattaram, Swethasree [Georgia Institute of Techno↗

GenomeDepot v1.0

GenomeDepot is a web-based platform for annotation, management, and comparative analysis of microbial genomic sequences and associated data including ortholog families, protein domains, operons, regulatory interactions, strain taxonomy, and sample metadata. GenomeDepot supports rapid creation of web-sites for user-defined genome collections that include bioinformatic tools for interactive genome browsing, BLAST search, annotation search, comparative genomic neighborhood visualization, and sequence download. Gene function annotations are generated by a customizable annotation pipeline. The pipeline runs annotation tools in Conda environments and can be easily extended with additional user-specified tools.

Kazakov, Alexey [Lawrence Berkeley National Labora↗

From soil to sequence: filling the critical gap in genome-resolved metagenomics is essential to the future of soil microbial ecology

Abstract Soil microbiomes are heterogeneous, complex microbial communities. Metagenomic analysis is generating vast amounts of data, creating immense challenges in sequence assembly and analysis. Although advances in technology have resulted in the ability to easily collect large amounts of sequence data, soil samples containing thousands of unique taxa are often poorly characterized. These challenges reduce the usefulness of genome-resolved metagenomic (GRM) analysis seen in other fields of microbiology, such as the creation of high quality metagenomic assembled genomes and the adoption of genome scale modeling approaches. The absence of these resources restricts the scale of future research, limiting hypothesis generation and the predictive modeling of microbial communities. Creating publicly available databases of soil MAGs, similar to databases produced for other microbiomes, has the potential to transform scientific insights about soil microbiomes without requiring the computational resources and domain expertise for assembly and binning.

59 BASIC BIOLOGICAL SCIENCES↗

Enzyme Engineering Database (EnzEngDB): a platform for sharing and interpreting sequence–function relationships across protein engineering campaigns

The discovery and engineering of new enzymes is important across the bioeconomy, with diverse applications from foods to pharmaceuticals, sensors to agriculture. However, enzyme engineering, in particular machine learning-guided engineering, is hampered by a lack of data. Currently there exists no database designed to capture and interpret datasets created in this domain, nor are there easy analysis and visualisation tools. We developed the Enzyme Engineering Database to provide a centralized resource and an online analysis tool to consolidate sequence-function data from enzyme engineering campaigns, thereby making three contributions: (i) a database into which researchers can deposit public data, (ii) visualisation and analysis tools for protein engineers to analyse their own data or compare enzyme variants to other engineering campaigns, and (iii) a gold-standard dataset for benchmarking automated extraction along with the first large language model extraction pipeline specific for enzyme engineering campaigns. The Enzyme Engineering Database is accessible at http://enzengdb.org/.

Long, Yueming [California Institute of Technology ↗