Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “sequencing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Constructing a High‐Resolution Aftershock Catalog for the 2017 Mw 8.2 Tehuantepec Earthquake Sequence Using a Machine Learning–Based Workflow

The 8 September 2017 Mw 8.2 Tehuantepec earthquake was the largest instrumentally recorded normal‐faulting earthquake in Mexico. The mainshock occurred offshore within the Tehuantepec seismic gap, generating >30,000 aftershocks in the following year. We applied an open‐source, machine learning (ML)–assisted workflow to construct a high‐resolution aftershock catalog using data from temporary and permanent seismic networks in southern Mexico. The workflow integrates PhaseNet for phase detection; GaMMA for phase association; and VELEST, HypoInverse, and HypoDD for velocity modeling and relocation. We processed seven months of continuous waveform data from 29 broadband stations, including a temporary rapid‐response deployment that improved station coverage of the offshore rupture zone. To evaluate performance, we compared our results against analyst‐reviewed picks and event locations from the Servicio Sismológico Nacional catalog. The resulting catalog contains 11,374 relocated earthquakes and represents the most comprehensive published dataset for this sequence, incorporating the first full use of the temporary network. Relocated hypocenters show improved depth control and align well with the Slab2.0 subduction geometry, revealing clearer separation between offshore slab events and onshore crustal seismicity. This study demonstrates that combining ML‐based detection with established methods provides a scalable and reproducible approach for constructing high‐quality earthquake catalogs in tectonically complex environments and offers practical guidance for adapting similar workflows to other earthquake sequences.

Garcia, Marc [The University of Texas at El Paso, ↗

KBase Narrative - Complete genome sequence of a novel Microbacterium sp. strain Clip185.

We have isolated a new species of Microbacterium, an Actinobacterium. We have temporarily named this bacterium as Microbacterium sp. strain Clip185 (hereafter called strain Clip185) from a contaminated Tris-Acetate-Phosphate (TAP) medium culture plate of a green micro-alga Chlamydomonas reinhardtii strain LMJ.RY0402.185141 (a Chlamydomonas Library project CLiP strain). We sequenced the whole genome of strain Clip185 using the PacBio Sequel II Continuous Long Read technology and have submitted it to NCBI along with the SRA and PacBio methylation motif data. Additionally, we have submitted the PacBio methylome to REBASE, Ref#35996. We present the whole genome sequence of this new Microbacterium species that offers insights into its coding and non-coding genes and its nearest taxonomic neighbors.

Mitra, Mautusi↗

Complete genome sequence of Sphingobium yanoikuyae strain CC4533

We have isolated a new strain of Sphingobium yanoikuyae , which belongs to the class Alphaproteobacteria, order Sphingomonadales, and family Sphingomonadaceae. This carotenoid-producing strain is capable of degrading xenobiotics and is tolerant to toxic levels of six heavy metals. We have designated the newly isolated strain of S. yanoikuyae as S. yanoikuyae strain CC4533 (hereafter called strain CC4533) because it was isolated from a contaminated Tris-Acetate-Phosphate (TAP) medium culture plate of a green micro-alga Chlamydomonas reinhardtii wild type strain CC4533. We sequenced the whole genome of strain CC4533 using the PacBio Sequel II Continuous Long Read technology and have submitted it to NCBI along with the SRA and PacBio methylation motif data. Additionally, we have submitted the PacBio methylome to REBASE, Ref#35996. We present the whole genome sequence of S. yanoikuyae strain CC4533 that offers insights into its coding and non-coding genes and its nearest taxonomic neighbors.

59 BASIC BIOLOGICAL SCIENCES↗

Next-generation sequencing dataset of genome-scale CRISPRi in Synechococcus sp. PCC 7002 across seven conditions

A 33,298-member sgRNA library developed for Synechococus sp. PCC 7002 was screened with two replicates across seven growth conditions and sequenced with Illumina NextSeq (paired end, 2x150 bp) for a total of ~700M reads. The original plasmid library and the library after transformation into a dCas9-containing and dCas9-absent strain were also sequenced as a reference for initial sgRNA abundance.

genome- wide screens environmental acclimation spe↗

Section-level genome sequencing and comparative genomics of Aspergillus sections Cavernicolus and Usti.

The genus Aspergillus is diverse, including species of industrial importance, human pathogens, plant pests, and model organisms. Aspergillus includes species from sections Usti and Cavernicolus, which until recently were joined in section Usti, but have now been proposed to be non-monophyletic and were split by section Nidulantes, Aenei and Raperi. To learn more about these sections, we have sequenced the genomes of 13 Aspergillus species from section Cavernicolus (A. cavernicola, A. californicus, and A. egyptiacus), section Usti (A. carlsbadensis, A. germanicus, A. granulosus, A. heterothallicus, A. insuetus, A. keveii, A. lucknowensis, A. pseudodeflectus and A. pseudoustus), and section Nidulantes (A. quadrilineatus, previously A. tetrazonus). We compared these genomes with 16 additional species from Aspergillus to explore their genetic diversity, based on their genome content, repeat-induced point mutations (RIPs), transposable elements, carbohydrate-active enzyme (CAZyme) profile, growth on plant polysaccharides, and secondary metabolite gene clusters (SMGCs). All analyses support the split of section Usti and provide additional insights: Analyses of genes found only in single species show that these constitute genes which appear to be involved in adaptation to new carbon sources, regulation to fit new niches, and bioactive compounds for competitive advantages, suggesting that these support species differentiation in Aspergillus species. Sections Usti and Cavernicolus have mainly unique SMGCs. Section Usti contains very large and information-rich genomes, an expansion partially driven by CAZymes, as section Usti contains the most CAZyme-rich species seen in genus Aspergillus. Section Usti is clearly an underutilized source of plant biomass degraders and shows great potential as industrial enzyme producers. Citation: Nybo JL, Vesth TC, Theobald S, Frisvad JC, Larsen TO, Kjaerboelling I, Rothschild-Mancinelli K, Lyhne EK, Barry K, Clum A, Yoshinaga Y, Ledsgaard L, Daum C, Lipzen A, Kuo A, Riley R, Mondo S, LaButti K, Haridas S, Pangalinan J, Salamov AA, Simmons BA, Magnuson JK, Chen J, Drula E, Henrissat B, Wiebenga A, Lubbers RJM, Müller A, dos Santos Gomes AC, Mäkelä MR, Stajich JE, Grigoriev IV, Mortensen UH, de Vries RP, Baker SE, Andersen MR (2025). Section-level genome sequencing and comparative genomics of Aspergillus sections Cavernicolus and Usti. Studies in Mycology 111: 101-114. doi: 10.3114/sim.2025.111.03.

59 BASIC BIOLOGICAL SCIENCES↗

Sequence, structure prediction, and epitope analysis of the polymorphic membrane protein family in Chlamydia trachomatis

The polymorphic membrane proteins (Pmps) are a family of autotransporters that play an important role in infection, adhesion and immunity in Chlamydia trachomatis. Here we show that the characteristic GGA(I,L,V) and FxxN tetrapeptide repeats fit into a larger repeat sequence, which correspond to the coils of a large beta-helical domain in high quality structure predictions. Analysis of the protein using structure prediction algorithms provided novel insight to the chlamydial Pmp family of proteins. While the tetrapeptide motifs themselves are predicted to play a structural role in folding and close stacking of the beta-helical backbone of the passenger domain, we found many of the interesting features of Pmps are localized to the side loops jutting out from the beta helix including protease cleavage, host cell adhesion, and B-cell epitopes; while T-cell epitopes are predominantly found in the beta-helix itself. This analysis more accurately defines the Pmp family of Chlamydia and may better inform rational vaccine design and functional studies.

59 BASIC BIOLOGICAL SCIENCES↗

Long-term effects of cycle time and volume exchange ratio on poly(3-hydroxybutyrate-co-3-hydroxyvalerate) production from food waste digestate by Haloferax mediterranei cultivated in sequencing batch reactors for 450 days

Food waste digestate was fed into a sequencing batch reactor (SBR) for Haloferax mediterranei (HM) to produce poly(3-hydroxybutyrate-co-3-hydroxyvalerate) (PHBV). This SBR was operated uninterruptedly for 450 days to test its stability, during which the cycle time and volume exchange ratio were varied to understand their impacts on the PHBV fermentation performance under ranged organic loading rates (OLR). Results showed that 1) PHBV productivity was proportional to OLR of food waste digestate; 2) substrate and product inhibitions were two limiting factors constraining substrate utilization and PHBV yields; 3) PHBV titer was dependent on the hydraulic retention time of the SBR while a volume exchange ratio lower than 0.5 is unfavorable due to the product inhibitor accumulation. Furthermore, this study for the first time demonstrated that the long-term stability of food waste-fed PHBV production by HM and revealed that inhibition effects could be barriers in SBR limiting the full-scale application of the technology.

09 BIOMASS FUELS↗

Single cell RNA sequencing reveals shifts in cell maturity and function of endogenous and infiltrating cell types in response to acute intervertebral disc injury

Intervertebral disc (IVD) degeneration contributes to disabling back pain. Degeneration can be initiated by injury and progressively leads to an irreversible loss of cells and function. IVD function restoration through cell replacement therapies have had limited success due to knowledge gaps in the critical cell populations important for repair. Here, in this study, we used single cell RNA sequencing to identify the transcriptional changes of IVD resident and infiltrating cell populations from Control and Injured coccygeal IVDs extracted from 12-week-old female C57BL/6J mice 7 days post injury. Clustering, gene ontology, and pseudotime trajectory analyses determined transcriptomic divergences with injury, flow cytometry identified they types of infiltrating immune cells, and immunofluorescence was utilized to define mesenchymal stem cell (MSC) localization. We identified 11 distinct clusters that included IVD, immune, vascular cells, and MSCs. Differential gene expression analysis determined that Outer Annulus Fibrosus, Neutrophils, Saa2-High MSCs, Macrophages, and Krt18 + Nucleus Pulposus (NP) cells were the major drivers of transcriptomic differences between Control and Injured cells. Gene ontology revealed that the most upregulated biological pathways were angiogenesis and T cell-related while wound healing and ECM regulation were downregulated. Pseudotime trajectory analyses revealed that IVD injury directed cells towards increased differentiation in all clusters, except for Krt18 + NP cells which remained in a less mature cell state. Saa2-High and Grem1-High MSCs populations shifted towards more differentiated IVD cells profiles with injury and localized distinctly within the IVD. This study revealed novel MSC populations with the potential to be leveraged for future IVD repair studies.

Cartilage↗

Frameshifting Stimulatory Sequence Induces Large Structural Change of Ribosomal Proteins When Bound to E. coli Ribosomes

Biological macromolecular machines occupy a continuum of structural conformations to perform cellular tasks. Mapping this conformational space provides an insight into its functionality. While the cryo-electron microscopy resolution revolution has expanded our ability to characterize the conformational continuums, there are obstacles in structurally characterizing regions of high flexibility. These technical barriers have impeded characterization of flexible ribosomal proteins when the ribosome is interacting with mRNA stem-loop structures such as a frameshifting stimulatory sequence (FSS). Small-angle neutron/X-ray scattering and electron microscopy were used to study ribosomal samples and compared structural differences between a ribosome that is bound to an FSS stem-loop compared to a ribosome bound to linear mRNA. This comparison shows that a large protein stalk elongates by 22% when the 70S interacts with an mRNA stem-loop. Finally, our results suggest that ribosomal proteins have extensive flexibility and may influence important ribosomal mechanisms, such as those that involve FSS.

36 MATERIALS SCIENCE↗

Optimized Substrate Positioning Enables Switches in the C–H Cleavage Site and Reaction Outcome in the Hydroxylation–Epoxidation Sequence Catalyzed by Hyoscyamine 6β-Hydroxylase

Hyoscyamine 6β-hydroxylase (H6H) is an Fe(II)- and 2-oxoglutarate-dependent (Fe/2OG) oxygenase that catalyzes the last two steps in the biosynthesis of scopolamine, a prolifically administered anti-nausea drug. After its namesake first reaction, H6H couples the newly installed C6-bonded oxygen to C7 to form the epoxide of scopolamine. Oxoiron(IV) (ferryl) intermediates initiate both reactions by cleaving C–H bonds, but it remains unclear how the enzyme switches target site and promotes (C6)O–C7 coupling in preference to C7 hydroxylation in the second step. In one possible epoxidation mechanism, the C6 oxygen would – analogously to mechanisms proposed for the Fe/2OG halogenases and, in the preceding paper, N-acetylnorloline synthase (LolO) – coordinate as alkoxide to the C7–H-cleaving ferryl intermediate to enable alkoxyl coupling to the ensuing C7 radical. Here we provide structural and kinetic evidence that H6H instead exploits the distinct spatial dependencies of competitive C–H-cleavage (C6 vs C7) and C–O-coupling (oxygen rebound vs cyclization) steps to promote the two-step sequence without substrate coordination or repositioning for the epoxidation step. Structural comparisons of ferryl-mimicking vanadyl complexes of wild-type H6H and a variant that preferentially hydroxylates C7 of 6-hydroxyhyoscyamine suggest that only a modest (~ 10°) shift in the Fe–O–H(C7) approach angle is sufficient to change the outcome. Finally, the observation that, in wild-type H6H, 2 H 2 O solvent also increases the C7-hydroxylation:epoxidation ratio by ~ 8-fold implies that the latter outcome requires cleavage of the alcohol O-H bond, which, unlike in the LolO oxacyclization, is not accomplished in advance of C–H cleavage.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Seismic Features Predict Ground Motions During Repeating Caldera Collapse Sequence

Abstract Applying machine learning to continuous acoustic emissions, signals previously deemed noise, from laboratory faults and slowly slipping subduction‐zone faults, demonstrates hidden signatures are emitted that describe physical details, including fault displacement and friction. However, no evidence currently exists to demonstrate that similar hidden signals occur during seismogenic stick‐slip on earthquake faults—the damaging earthquakes of most societal interest. We show that continuous seismic emissions emitted during the 2018 multi‐month caldera collapse sequence at the Kı̄lauea volcano in Hawai'i contain hidden signatures characterizing the earthquake cycle. Multi‐spectral data features extracted from 30 s intervals of the continuous seismic emission are used to train a gradient boosted tree regression model to predict the GNSS‐derived contemporaneous surface displacement and time‐to‐failure of the upcoming collapse event. This striking result suggests that at least some faults emit such signals and provide a potential path to characterizing the instantaneous and future behavior of earthquake faults.

58 GEOSCIENCES↗

Multimodal framework for the joint analysis of single-cell RNA and T cell receptor sequencing data predicts T cell response to cancer immunotherapy

T cell states are prognostic in different cancer types. Recent technologies enable joint profiling of T cell RNA and T cell receptor (TCR) sequences at single-cell resolution. Here we present the TCR-RNA Integrating Model (TRIM), a multi-modal variational autoencoder framework that integrates RNA-TCR data and predicts T cell clonality and transcriptional states. TRIM learns a shared representation of the data conditioned on patient, tissue source, and treatment timepoint. We applied TRIM to three independent datasets that included T cells collected before and after checkpoint inhibitor treatment, sourced either from blood and tumor biopsies in patients with head and neck squamous cell carcinoma and colorectal cancer, or from tumor and adjacent tissue in a pan-cancer dataset. In all settings, TRIM accurately predicted intra-tumor T cell clonal expansion and transcriptional status based on T cells from blood or normal tissue before treatment, demonstrating its utility in modeling multimodal T cell data and predicting T cell response to treatment and disease progression.

60 APPLIED LIFE SCIENCES↗

Dissecting neurofilament tail sequence-phosphorylation-structure relationships with multicomponent reconstituted protein brushes

Neurofilaments (NFs) are multisubunit, bottlebrush-shaped intermediate filaments abundant in the axonal cytoskeleton. Each NF subunit contains a long intrinsically disordered tail domain, which protrudes from the NF core to form a “brush” surrounding each NF. Precisely how the tails’ variable charge patterns and repetitive phosphorylation sites mediate their conformation within the brush remains an open question in axonal biology. We address this problem by grafting recombinant NF tail protein constructs NF-Light, -Medium, and -Heavy (NFL, NFM, and NFH) to surfaces, yielding protein brushes of defined stoichiometry that can be phosphorylated in vitro. Atomic force microscopy measurements reveal that brush height depends on composition monotonically but not always linearly for binary NFL:NFM or NFL:NFH systems, and that NFM-based brushes are highly extended, while brushes incorporating the much larger NFH are surprisingly compact even after multisite phosphorylation. Complementary self-consistent field theory (SCFT) predicts multilayer brush morphologies for NFM and phosphorylated NFH brushes. Further experiments and SCFT analysis with designed mutants reveal that N-terminal negative charges in the NFH tail repel phosphorylated residues to generate the multilayer morphology, while the C-terminal charge-neutral region contributes to multilayer brush morphology but not total brush height. Charge-shuffled NFM variants show that charge segregation promotes brush collapse near physiological ionic strengths. Collectively, this study supports a role for NFM in establishing a dynamic range for NF brush conformation, lending insight into previous in vitro and in vivo findings. More broadly, this work establishes a platform for dissecting contributions of disordered protein sequence to conformation at interfaces.

Science & Technology - Other Topics↗

Sequence modeling of higher-order wave modes of quasi-circular, spinning, non-precessing binary black hole mergers

Higher-order gravitational wave modes from quasi-circular, spinning, non-precessing binary-black-hole (BBH) mergers encode rich information about the nonlinear dynamics of strong-field gravity. We present a transformer-based sequence-completion surrogate that, given an early-inspiral segment, forecasts the subsequent late inspiral, merger, and ringdown. The intended applications are (i) patching or completing expensive or interrupted numerical-relativity (NR) simulations and (ii) providing late-time cross-checks and rapid hybridization studies. The training set is built from the NRHybSur3dq8 surrogate, which provides spherical-harmonic modes up to $\ell$ ≤ 4 (excluding (4, 0) and (4,±1), and including (5, 5)) for mass ratios q ≤ 8, dimensionless spin components s$^{z}_{1,2}$ ϵ[–0.8, 0.8], and inclination angles θ ϵ [0, π]. Waveforms are supplied on the interval t ϵ [–5000M, –100 M) and the model autoregressively generates the plus and cross polarizations (h + , h x ) on t ϵ [–100 M, 130M]. Training on the Delta supercomputer with 16 NVIDIA A100 GPUs required ~15 h on more than 14 million hybrid waveforms. Evaluation on a held-out test set of 840,000 samples yields mean and median overlaps of 0.996 and 0.997, respectively, with respect to the surrogate ground truth.

black-hole merger↗

Microbiome Comparison and Pathogen Identification for Three Migrating Passerines Captured During Spring Season in Jordan Using 16S rRNA Sequencing

Jordan is located on an important spot along the Mediterranean and Black Sea Flyway. Hundreds of migratory bird species have been identified stopping over in Jordan during spring and autumn migratory seasons. Compared to mammals and economically important birds, the microbiomes of wild bird species are severely understudied. Gut microbial composition is a valuable source of information that reflects food preferences, foraging behavior, and the risk of pathogen transmission to humans and other animals. In this study, we assessed the microbiome composition of three species of migrating passerines (willow warblers, lesser whitethroats, and common reed warblers) captured during the spring migration stopover in Jordan in 2023. A total of 59 fecal samples were selected evenly from the three species and subjected to 16S sequencing and microbiome analysis. Our objectives were to determine the diversity of bacteria in these three species, assess the amount of intra- and inter-specific variation, and detect pathogenic genera and species that could pose health risks to humans, domestic animals, and wildlife. Bacteria mainly belonged to the phyla Proteobacteria (62%), Actinobacteriota (18%), Firmicutes (13%), Cyanobacteria (5%), and Bacteroidota (1%). The results reveal that lesser whitethroats had the greatest variation in bacterial genus richness, Shannon diversity, and microbial composition compared to willow warblers and common reed warblers. The three bird species harbored several pathogenic genera and species, including Campylobacter, Enterococcus, Escherichia-Shigella, Mycoplasma, Rickettsia, Clostridium perfringens, and Vibrio cholerae. We suggest further investigation to understand the relationship between migratory behavior and their gut microbiome. We advocate for the use of advanced molecular techniques to characterize the pathogens found in migratory birds that might have public and environmental health impacts in addition to economic loss.

59 BASIC BIOLOGICAL SCIENCES↗

Enzyme Engineering Database (EnzEngDB): a platform for sharing and interpreting sequence–function relationships across protein engineering campaigns

The discovery and engineering of new enzymes is important across the bioeconomy, with diverse applications from foods to pharmaceuticals, sensors to agriculture. However, enzyme engineering, in particular machine learning-guided engineering, is hampered by a lack of data. Currently there exists no database designed to capture and interpret datasets created in this domain, nor are there easy analysis and visualisation tools. We developed the Enzyme Engineering Database to provide a centralized resource and an online analysis tool to consolidate sequence-function data from enzyme engineering campaigns, thereby making three contributions: (i) a database into which researchers can deposit public data, (ii) visualisation and analysis tools for protein engineers to analyse their own data or compare enzyme variants to other engineering campaigns, and (iii) a gold-standard dataset for benchmarking automated extraction along with the first large language model extraction pipeline specific for enzyme engineering campaigns. The Enzyme Engineering Database is accessible at http://enzengdb.org/.

Long, Yueming [California Institute of Technology ↗

Improving precision and accuracy of genetic mapping with genotyping‐by‐sequencing data in outcrossing species

Abstract Genotyping‐by‐sequencing (GBS) is a widely used strategy for obtaining large numbers of genetic markers in model and non‐model organisms. In crop plants, GBS‐derived marker datasets are frequently used to perform quantitative trait locus (QTL) mapping. In some plant species, however, high heterozygosity and complex genome structure mean that researchers must use care in handling GBS data to conduct QTL mapping most effectively. Such outbred crops include most of the perennial grass and tree species used for bioenergy. To identify strategies for increasing accuracy and precision of QTL mapping using GBS data in outbred crops, we conducted an empirical study of SNP‐calling and genetic map‐building pipeline parameters in a Miscanthus sinensis population, and a complementary simulation study to estimate the relationship between genome‐wide error rate, read depth, and marker number. The bioenergy grass Miscanthus is an obligate outcrossing species with a recent (diploidized) whole‐genome duplication. For the study of empirical M. sinensis data, we compared two SNP‐calling methods (one non‐reference‐based and one reference‐based), a series of depth filters (12×, 20×, 30×, and 40×) and two map‐construction methods (i.e., marker ordering: linkage‐only and order‐corrected based on a reference genome). We found that correcting the order of markers on a linkage map by using a high‐quality reference genome improved QTL precision (shorter confidence intervals). For typical GBS datasets of between 1000 and 5000 markers to build a genetic map for biparental populations, a depth filter set at 30× to 40× applied to outbred populations provided a genome‐wide genotype‐calling error rate of less than 1%, improved accuracy of QTL point estimates and minimized type I errors for identifying QTL. Based on these results, we recommend using a reference genome to correct the marker order of genetic maps and a robust genotype depth filter to improve QTL mapping for outbred crops.

59 BASIC BIOLOGICAL SCIENCES↗