Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Base Sequence”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Lanthanide binding peptide surfactants at air–aqueous interfaces for interfacial separation of rare earth elements

Rare earth elements (REEs) are critical materials to modern technologies. They are obtained by selective separation from mining feedstocks consisting of mixtures of their trivalent cation. We are developing an all-aqueous, bioinspired, interfacial separation using peptides as amphiphilic molecular extractants. Lanthanide binding tags (LBTs) are amphiphilic peptide sequences based on the EF-hand metal binding loops of calcium-binding proteins which complex selectively REEs. We study LBTs optimized for coordination to Tb 3+ using luminescence spectroscopy, surface tensiometry, X-ray reflectivity, and X-ray fluorescence near total reflection, and find that these LBTs capture Tb 3+ in bulk and adsorb the complex to the interface. Molecular dynamics show that the binding pocket remains intact upon adsorption. We find that, if the net negative charge on the peptide results in a negatively charged complex, excess cations are recruited to the interface by nonselective Coulombic interactions that compromise selective REE capture. If, however, the net negative charge on the peptide is −3, resulting in a neutral complex, a 1:1 surface ratio of cation to peptide is achieved. Surface adsorption of the neutral peptide complexes from an equimolar mixture of Tb 3+ and La 3+ demonstrates a switchable platform dictated by bulk and interfacial effects. The adsorption layer becomes enriched in the favored Tb 3+ when the bulk peptide is saturated, but selective to La 3+ for undersaturation due to a higher surface activity of the La 3+ complex.

Ortuno Macias, Luis E. (ORCID:0000000284342192)

Association between optically identified galaxy clusters and the underlying dark matter halos

Clusters of galaxies trace massive dark matter halos in the Universe, but they can include multiple halos projected along lines of sight. Here, we study the halos contributing to clusters using the Cardinal simulation, which mimics the Dark Energy Survey data. We use the red-sequence-based cluster finding algorithm redMaPPer as a case study. For each cluster, we identify the halos hosting its member galaxies, and we define the main halo as the one contributing the most to the cluster's richness ($λ$, the estimated number of member galaxies). At $z=0.3$, for clusters with $λ> 60$, the main halo typically contributes to $92\%$ of the richness, and this fraction drops to $67\%$ for $λ\approx 20$. Defining "clean" clusters as those with $\geq50\%$ of the richness contributed by the main halo, we find that $100\%$ of the $λ> 60$ clusters are clean, while $73\%$ of the $λ\approx 20$ clusters are clean. Three halos can usually account for more than $80\%$ of the richness of a cluster. The main halos associated with redMaPPer clusters have a completeness ranging from $98\%$ at virial mass $10^{14.6}~h^{-1}M_{\odot}$ to $64\%$ at $10^{14}~h^{-1}M_{\odot}$. In addition, we compare the inferred cluster centers with true halo centers, finding that $30\%$ of the clusters are miscentered with a mean offset $40\%$ of the cluster radii, in agreement with recent X-ray studies. These systematics worsen as redshift increases, but we expect that upcoming surveys extending to longer wavelengths will improve the cluster finding at high redshifts. Our results affirm the robustness of the redMaPPer algorithm and provide a framework for benchmarking other cluster-finding strategies.

79 ASTRONOMY AND ASTROPHYSICS

SPADES

Sequence-based Pathogen-Agnostic Diagnostics/Detection Solution

Li, Po-E [Los Alamos National Laboratory]

Status on Genetic Resistance to Rice Blast Disease in the Post-Genomic Era

Rice blast, caused by Magnaporthe oryzae, is a major threat to global rice production, necessitating the development of resistant cultivars through genetic improvement. Breakthroughs in rice genomics, including the complete genome sequencing of japonica and indica subspecies and the availability of various sequence-based molecular markers, have greatly advanced the genetic analysis of blast resistance. To date, approximately 122 blast-resistance genes have been identified, with 39 of these genes cloned and molecularly characterized. The application of these findings in marker-assisted selection (MAS) has significantly improved rice breeding, allowing for the efficient integration of multiple resistance genes into elite cultivars, enhancing both the durability and spectrum of resistance. Pangenomic studies, along with AI-driven tools like AlphaFold2, RoseTTAFold, and AlphaFold3, have further accelerated the identification and functional characterization of resistance genes, expediting the breeding process. Future rice blast disease management will depend on leveraging these advanced genomic and computational technologies. Emphasis should be placed on enhancing computational tools for the large-scale screening of resistance genes and utilizing gene editing technologies such as CRISPR-Cas9 for functional validation and targeted resistance enhancement and deployment. These approaches will be crucial for advancing rice blast resistance, ensuring food security, and promoting agricultural sustainability.

Pedrozo, Rodrigo

Xanthos-Lake Model Source Code

This repository contains the source code for Xanthos-Lake, a lake-modeling extension of the Xanthos framework that introduces a coupled lake component comprising the Xanthos-Lake Snow and Ice Model (xLSIM) and the Xanthos-Lake Water Balance Model (xLWBM). xLSIM is a basin-aware machine-learning model for lake snow, ice, and thermal conditions. It predicts monthly lake ice thickness, snow depth, snow-cover fraction, mixing-layer temperature, and lake ice fraction from meteorological forcing and lake surface-area information. It uses sequence-based deep-learning architectures, including Transformer and hybrid Long Short-Term Memory–Transformer (LSTM–Transformer) models, together with seasonal encoding, multi-lake learning, physical masking, and basin-level cryospheric and non-cryospheric classification. The training workflow uses Ray for scalable execution and includes optional Ray Tune hyperparameter optimization. Model predictions, observations, diagnostics, and feature-importance outputs are written in NetCDF. xLWBM is the water-balance component of the new lake framework. It simulates monthly lake storage, surface area, evaporation, inflow, outflow, and lake–groundwater exchange. It combines physical water-balance equations with calibrated bathymetric relationships, weir-based outlet flow, modified Penman open-water evaporation, groundwater head relaxation, Penman–Monteith snow and ice sublimation, and snow, ice, and thermal conditions supplied by xLSIM. The model calibrates lake parameters against satellite-derived surface-area data, using evaporation-based calibration where surface-area data are unavailable, and supports small, medium, and large lake classes. For large lakes, xLWBM is integrated with the managed-routing workflow so that lake storage and outflow interact directly with downstream river routing and reservoir operations. Together, xLSIM and xLWBM provide Xanthos with a coupled lake-modeling capability. xLSIM supplies the snow, ice, and thermal conditions that affect lake evaporation and snow- and ice-related water exchanges, while xLWBM translates those conditions into dynamic lake storage, surface area, evaporation, and discharge. In return, xLWBM supplies evolving lake surface area to xLSIM. This coupling enables Xanthos to represent lakes as active hydrologic components within basin-scale water-availability and routing simulations.

Machine Learning

Beyond sequence similarity: toward function-based screening of nucleic acid synthesis

Synthetic nucleic acids are a key input to modern biotechnology, yet they represent dual-use materials that require robust screening to mitigate biosecurity risks. The prevailing screening paradigm, which identifies sequences of concern (SoCs) through sequence similarity to controlled pathogens and toxins, may not fully capture risks posed by AI tools that can decouple biomolecular function from reliance on known sequences. Rapidly advancing biodesign capabilities enable the generation of genes and proteins that might evade sequence-based detection. We highlight the critical need for function-based screening approaches that can detect sequences capable of hazardous biological functions, regardless of similarity to known SoCs. We examine the feasibility of function-based screening with an initial focus on proteins, arguing that, while protein sequence space is vast, biologically functional proteins are significantly constrained by biophysical and biochemical requirements that can be learned and modeled. We propose a concrete implementation framework organized along a continuum of complexity, starting with toxins as the most tractable targets before expanding to more complex pathogenic functions. We then discuss open challenges and describe a research and development strategy to address them.

59 BASIC BIOLOGICAL SCIENCES

An FPGA-based hardware accelerator supporting sensitive sequence homology filtering with profile hidden Markov models

Abstract Background Sequence alignment lies at the heart of genome sequence annotation. While the BLAST suite of alignment tools has long held an important role in alignment-based sequence database search, greater sensitivity is achieved through the use of profile hidden Markov models (pHMMs). Here, we describe an FPGA hardware accelerator, called HAVAC, that targets a key bottleneck step (SSV) in the analysis pipeline of the popular pHMM alignment tool, HMMER. Results The HAVAC kernel calculates the SSV matrix at 1739 GCUPS on a $$\sim$$ ∼ $3000 Xilinx Alveo U50 FPGA accelerator card, $$\sim$$ ∼ 227× faster than the optimized SSV implementation in nhmmer . Accounting for PCI-e data transfer data processing, HAVAC is 65× faster than nhmmer’s SSV with one thread and 35× faster than nhmmer with four threads, and uses $$\sim$$ ∼ 31% the energy of a traditional high end Intel CPU. Conclusions HAVAC demonstrates the potential offered by FPGA hardware accelerators to produce dramatic speed gains in sequence annotation and related bioinformatics applications. Because these computations are performed on a co-processor, the host CPU remains free to simultaneously compute other aspects of the analysis pipeline.

59 BASIC BIOLOGICAL SCIENCES

Nanopore Activity Assays for Detection of Biomarker Protease Activity: Design and Testing of Substrates for Both Nanopore Sequencing and PCR-Based Detection Methods

The work performed in this project has demonstrated the ability to construct proteolytic enzyme substrates that are PCR and sequencing-readable reporter molecules. Specifically, the goal was to detect those reporter molecules via PCR and Oxford Nanopore Technologies MinION sequencing methods following exposure to the biomarker protease thrombin. The assay development focused on binding the constructed peptide-oligonucleotide chimera to immobilized streptavidin. The action of thrombin on the peptide portion of the molecule released the oligonucleotide for detection. Detection of protease activity was demonstrated in a concentration-dependent manner using MALDI-MS, RT-PCR and DNA sequencing. Additional steps to remove background release of reporter molecules during the assay was used to improve the difference in detected oligonucleotide reporter following protease activity. Additional steps in assay development will be to (1) test the assay in an appropriate matrix, (2) investigate detection using additional DNA sequencing platforms and (3) demonstrate multiplexed detection of multiple protease markers in a single reaction.

59 BASIC BIOLOGICAL SCIENCES

NEAR: Neural Embeddings for Amino acid Relationships

Protein language models (PLMs) have recently demonstrated potential to supplant classical protein database search methods based on sequence alignment, but are slower than common alignment-based tools and appear to be prone to a high rate of false labeling. Here, we present NEAR, a method based on neural representation learning that is designed to improve both speed and accuracy of search for likely homologs in a large protein sequence database. NEAR’s ResNet embedding model is trained using contrastive learning guided by trusted sequence alignments. It computes per-residue embeddings for target and query protein sequences, and identifies alignment candidates with a pipeline consisting of residue-level k-NN search and a simple neighbor aggregation scheme. Tests on a benchmark consisting of trusted remote homologs and randomly shuffled decoy sequences reveal that NEAR substantially improves accuracy relative to state-of-the-art PLMs, with lower memory requirements and faster embedding and search speed. While these results suggest that the NEAR model may be useful for standalone homology detection with increased sensitivity over standard alignment-based methods, in this manuscript we focus on a more straightforward analysis of the model’s value as a high-speed pre-filter for sensitive annotation. In that context, NEAR is at least 5x faster than the pre-filter currently used in the widely-used profile hidden Markov model (pHMM) search tool HMMER3, and also outperforms the pre-filter used in our fast pHMM tool, nail.

59 BASIC BIOLOGICAL SCIENCES

Role of Ribosomal Protein bS1 in Orthogonal mRNA Start Codon Selection

In many bacteria, the location of the mRNA start codon is determined by a short ribosome binding site sequence that base pairs with the 3'-end of 16S rRNA (rRNA) in the 30S subunit. Many groups have changed these short sequences, termed the Shine-Dalgarno (SD) sequence in the mRNA and the anti-Shine-Dalgarno (ASD) sequence in 16S rRNA, to create "orthogonal" ribosomes to enable the synthesis of orthogonal polymers in the presence of the endogenous translation machinery. However, orthogonal ribosomes are prone to SD-independent translation. Ribosomal protein bS1, which binds to the 30S ribosomal subunit, is thought to promote translation initiation by shuttling the mRNA to the ribosome. Thus, a better understanding of how the SD and bS1 contribute to start codon selection could help efforts to improve the orthogonality of ribosomes. Here, we engineered the Escherichia coli ribosome to prevent binding of bS1 to the 30S subunit and separate the activity of bS1 binding to the ribosome from the role of the mRNA SD sequence in start codon selection. We find that ribosomes lacking bS1 are slightly less active than wild-type ribosomes in vitro. Furthermore, orthogonal 30S subunits lacking bS1 do not have an improved orthogonality. Our findings suggest that mRNA features outside the SD sequence and independent of binding of bS1 to the ribosome likely contribute to start codon selection and the lack of orthogonality of present orthogonal ribosomes.

59 BASIC BIOLOGICAL SCIENCES

Short-period post-common envelope binaries with Balmer emission from SDSS and LAMOST based on ZTF photometric data

ABSTRACT We present here 55 short-period post-common envelope binaries (PCEBs) containing a hot white dwarf (WD) and a low-mass main sequence (MS). Based on the photometric data from Zwicky Transient Facility survey data Release 19 (ZTF DR19), the light curves are analysed for about 200 WDMS binaries with emission line(s) identified from the Sloan Digital Sky Survey (SDSS) or the Large Sky Area Multi-Object Fibre Spectroscopic Telescope (LAMOST) spectra, in which 55 WDMS binaries are found to exhibit variability in their luminosities with a short period and are thus short-period binaries (i.e. PCEBs). In addition, it is found that the orbital periods of these PCEBs locate in a range from 2.2643 to 81.1526 h. However, only six short-period PCEBs are newly discovered and the orbital periods of 19 PCEBs are improved in this work. Meanwhile, it is found that three objects are newly discovered eclipsing PCEBs, and a object (i.e. SDSS J1541) might be the short-period PCEB with a late M-type star or a brown dwarf companion based on the analysis of its spectral energy distribution. At last, the mechanism(s) being responsible for the emission features in the spectra of these PCEBs are discussed, the emission features arising in their optical spectra might be caused by the stellar activity or an irradiated component owing to a hot WD companion because most of them contain a WD with an effective temperature higher than $\sim$10 000 K.

Li, Lifang

Sequence modeling of higher-order wave modes of quasi-circular, spinning, non-precessing binary black hole mergers

Higher-order gravitational wave modes from quasi-circular, spinning, non-precessing binary-black-hole (BBH) mergers encode rich information about the nonlinear dynamics of strong-field gravity. We present a transformer-based sequence-completion surrogate that, given an early-inspiral segment, forecasts the subsequent late inspiral, merger, and ringdown. The intended applications are (i) patching or completing expensive or interrupted numerical-relativity (NR) simulations and (ii) providing late-time cross-checks and rapid hybridization studies. The training set is built from the NRHybSur3dq8 surrogate, which provides spherical-harmonic modes up to $\ell$ ≤ 4 (excluding (4, 0) and (4,±1), and including (5, 5)) for mass ratios q ≤ 8, dimensionless spin components s$^{z}_{1,2}$ ϵ[–0.8, 0.8], and inclination angles θ ϵ [0, π]. Waveforms are supplied on the interval t ϵ [–5000M, –100 M) and the model autoregressively generates the plus and cross polarizations (h + , h x ) on t ϵ [–100 M, 130M]. Training on the Delta supercomputer with 16 NVIDIA A100 GPUs required ~15 h on more than 14 million hybrid waveforms. Evaluation on a held-out test set of 840,000 samples yields mean and median overlaps of 0.996 and 0.997, respectively, with respect to the surrogate ground truth.

black-hole merger

Constructing a High‐Resolution Aftershock Catalog for the 2017 Mw 8.2 Tehuantepec Earthquake Sequence Using a Machine Learning–Based Workflow

The 8 September 2017 Mw 8.2 Tehuantepec earthquake was the largest instrumentally recorded normal‐faulting earthquake in Mexico. The mainshock occurred offshore within the Tehuantepec seismic gap, generating >30,000 aftershocks in the following year. We applied an open‐source, machine learning (ML)–assisted workflow to construct a high‐resolution aftershock catalog using data from temporary and permanent seismic networks in southern Mexico. The workflow integrates PhaseNet for phase detection; GaMMA for phase association; and VELEST, HypoInverse, and HypoDD for velocity modeling and relocation. We processed seven months of continuous waveform data from 29 broadband stations, including a temporary rapid‐response deployment that improved station coverage of the offshore rupture zone. To evaluate performance, we compared our results against analyst‐reviewed picks and event locations from the Servicio Sismológico Nacional catalog. The resulting catalog contains 11,374 relocated earthquakes and represents the most comprehensive published dataset for this sequence, incorporating the first full use of the temporary network. Relocated hypocenters show improved depth control and align well with the Slab2.0 subduction geometry, revealing clearer separation between offshore slab events and onshore crustal seismicity. This study demonstrates that combining ML‐based detection with established methods provides a scalable and reproducible approach for constructing high‐quality earthquake catalogs in tectonically complex environments and offers practical guidance for adapting similar workflows to other earthquake sequences.

Garcia, Marc [The University of Texas at El Paso,

Genomic and morphological characterization of Knufia obscura isolated from the Mars 2020 spacecraft assembly facility

Members of the family Trichomeriaceae, belonging to the Chaetothyriales order and the Ascomycota phylum, are known for their capability to inhabit hostile environments characterized by extreme temperatures, oligotrophic conditions, drought, or presence of toxic compounds. The genus Knufia encompasses many polyextremophilic species. In this report, the genomic and morphological features of the strain FJI-L2-BK-P2 presented, which was isolated from the Mars 2020 mission spacecraft assembly facility located at the Jet Propulsion Laboratory in Pasadena, California. The identification is based on sequence alignment for marker genes, multi-locus sequence analysis, and whole genome sequence phylogeny. The morphological features were studied using a diverse range of microscopic techniques (bright field, phase contrast, differential interference contrast and scanning electron microscopy). The phylogenetic marker genes of the strain FJI-L2-BK-P2 exhibited highest similarities with type strain of Knufia obscura (CBS 148926 T ) that was isolated from the gas tank of a car in Italy. To validate the species identity, whole genomes of both strains (FJI-L2-BK-P2 and CBS 148926 T ) were sequenced, annotated, and strain FJI-L2-BK-P2 was confirmed as K. obscura. The morphological analysis and description of the genomic characteristics of K. obscura FJI-L2-BK-P2 may contribute to refining the taxonomy of Knufia species. Key morphological features are reported in this K. obscura strain, resembling microsclerotia and chlamydospore-like propagules. These features known to be characteristic features in black fungi which could potentially facilitate their adaptation to harsh environments.

59 BASIC BIOLOGICAL SCIENCES

Understanding the structural mechanics of ligated DNA crystals via molecular dynamics simulation

DNA self-assembly is a highly programmable method to construct arbitrary architectures based on sequence complementarity. Among various constructs, DNA crystals are macroscopic crystalline materials formed by assembling motifs via sticky end association. Due to their high structural integrity and size ranging from tens to hundreds of micrometers, DNA crystals offer unique opportunities to study the structural properties and deformation behaviors of DNA assemblies. For example, enzymatic ligation of sticky ends can selectively seal nicks resulting in more robust structures with enhanced mechanical properties. However, the research efforts have been mostly on experiments involving different motif designs, structural optimization, or new synthesis methods, while their mechanics are not yet fully understood. The complex properties of DNA crystals are difficult to study via experiments alone, and numerical simulation can complement and aid the experiments. The coarse-grained molecular dynamics (MD) simulation is a powerful tool that can probe the mechanics of DNA assemblies. Here, we investigate DNA crystals made of four different motif lengths with various ligation patterns (full ligation, major directions, connectors, and in-plane) using oxDNA, an open-source, coarse-grained MD platform. We found that several distinct deformation stages emerge in response to mechanical loading and that the number and the location of ligated nucleotides can significantly modulate structural behaviors. These findings should be useful for predicting crystal properties and thus improving the design.

DNA crystal

Nanopore Readable Activity Probes for Ribosomal Inactivating Protein (RIP) Toxins

Ribosome inactivating proteins (RIPs) such as ricin and abrin depurinate an adenine base in the sarcin/ricin loop in the large ribosomal subunit, leading to inhibtion of protein synthesis and cell death. Here, we demonstrate that RIP toxin activity can be detected via nanopore-based DNA sequencing using synthetic oligonucleotide substrates. This is achieved by monitoring the mismatch proportion at the canonical target sequences incorporated into the synthetic substrate and determining the sequence length distribution throughout the entire substrate sequence. The mismatch proportion increases and sequence length distribution decreases with increasing toxin concentration for both ricin and abrin in buffer as well as in more complex backgrounds such as saliva and nasal secretions.

Turner, Matthew W [Pacific Northwest National Labo

Depth-resolved sagebrush root metabolomics, rhizosphere microbial communities, and geochemistry at the East River Watershed

This data set consists of results from soil nutrient profile, untargeted metabolomics, mass spec imaging, and amplicon sequencing. Data for soil nutrient profile includes common cations (Ca, Mg, Na, and K etc.) extracted from 3 digesting steps – ammonia acetate (for exchangeable cations), nitric acid (for acid dissolved fraction), and hydrofluoric acid/perchloric acid (HF/HClO4) for whole soil digestion. It also includes concentration of organic carbon, inorganic nitrogen (ammonia and nitrate) and phosphorus (Bray-1 P and nitric acid extract), and total nitrogen and phosphorus. Data for untargeted metabolomics includes metabolomic profile for root exudate/tissues and soil extracts from depths at surface soil to saprolite, that were measured using gas chromatography – mass spectrometry (GC-MS), and liquid chromatography – tandem mass spectrometry (LC-MS/MS). Data for mass spec imaging includes spatial distribution of metabolites that were detected and annotated with Fourier transformation ion cyclotron resonance mass spectrometer (FTICR-MS). Data for amplicon sequencing includes the base paired 16S and ITS ribosomal RNA sequences from Miseq Illumina sequencing. All samples were collected from 2 sampling campaign October 2022 and June 2023. Collectively, these datasets enable a mechanistic evaluation of how nutrient acquisition, especially nitrogen and phosphorus, differs between shallow roots operating in soil and deep roots functioning within the fractured bedrock zone. All files are provided as comma-separated values (CSV) fies (.csv) and (GZIP) file (.gz). The compressed .gz FASTQ files can be read directly in R using the dada2 package as part of the amplicon sequence analysis workflow. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. This research was performed on a project award 60563 (https://dx.doi.org/10.46936/expl.proj.2022.60563/60008727) from the Environmental Molecular Sciences Laboratory, a DOE Office of Science User Facility sponsored by the Biological and Environmental Research program under Contract No. DE-AC05-76RL01830.

EARTH SCIENCE > AGRICULTURE > SOILS > CARBON

Impact of Inverter-Based Resources on Grid Protection: A Review of Negative-Sequence Current Generation

The increasing integration of inverter-based resources (IBRs) in power grids poses challenges to traditional protection systems, primarily due to their different fault current signatures compared to conventional synchronous generators. Unlike synchronous generators whose fault response is dictated by their physical design, IBRs exhibit a wide range of fault characteristics due to manufacturer-specific control algorithms and settings. This dependence on proprietary control schemes makes modeling IBR behavior during faults significantly more complex, especially considering the rapid evolution of inverter technology and the diverse control strategies employed. While much research has focused on the positive-sequence current injections of IBRs during symmetrical faults, the understanding of negative-sequence current generation during non-symmetrical faults remains limited. This report provides an overview of current research on IBRs' negative-sequence current generation during unbalanced faults and its impact on protection schemes based on negative-sequence components. It covers both type III wind turbines and full-size converter-based IBRs. Additionally, this report reviews strategies for grid-forming controlled inverters to generate negative-sequence current during unbalanced faults, in addition to grid-following controlled ones.

24 POWER TRANSMISSION AND DISTRIBUTION