Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Direct Sequence”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

One-Pot Self-Assembly of Sequence-Controlled Mesoporous Heterostructures via Structure-Directing Agents

Multimaterial heterostructures have led to characteristics surpassing the individual components. Nature controls the architecture and placement of multiple materials through biomineralization of nanoparticles (NPs); however, synthetic heterostructure formation remains limited and generally departs from the elegance of self-assembly. Here, in this study, a class of block polymer structure-directing agents (SDAs) are developed containing repeat units capable of persistent (covalent) NP interactions that enable the direct fabrication of nanoscale porous heterostructures, where a single material is localized at the pore surface as a continuous layer. This SDA binding motif (design rule 1) enables sequence-controlled heterostructures, where the composition profile and interfaces correspond to the synthetic addition order. This approach is generalized with 5 material sequences using an SDA with only persistent SDA-NP interactions (“P-NP 1 –NP 2 ”; NP i = TiO 2 , Nb 2 O 5 , ZrO 2 ). Expanding these polymer SDA design guidelines, it is shown that the combination of both persistent and dynamic (noncovalent) SDA-NP interactions (“PD-NP 1 –NP 2 ”) improves the production of uniform interconnected porosity (design rule 2). The resulting competitive binding between two segments of the SDA (P- vs D-) requires additional time for the first NP type (NP 1 ) to reach and covalently attach to the SDA (design rule 3). The combination of these three design rules enables the direct self-assembly of heterostructures that localize a single material at the pore surface while preserving continuous porosity.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Crabtree/Warburg-like aerobic xylose fermentation by engineered Saccharomyces cerevisiae

Bottlenecks in the efficient conversion of xylose into cost-effective biofuels have limited the widespread use of plant lignocellulose as a renewable feedstock. The yeast Saccharomyces cerevisiae ferments glucose into ethanol with such high metabolic flux that it ferments high concentrations of glucose aerobically, a trait called the Crabtree/Warburg Effect. In contrast to glucose, most engineered S. cerevisiae strains do not ferment xylose at economically viable rates and yields, and they require respiration to achieve sufficient xylose metabolic flux and energy return for growth aerobically. Here, we evolved respiration-deficient S. cerevisiae strains that can grow on and ferment xylose to ethanol aerobically, a trait analogous to the Crabtree/Warburg Effect for glucose. Through genome sequence comparisons and directed engineering, we determined that duplications of genes encoding engineered xylose metabolism enzymes, as well as TKL1, a gene encoding a transketolase in the pentose phosphate pathway, were the causative genetic changes for the evolved phenotype. Reengineered duplications of these enzymes, in combination with deletion mutations in HOG1, ISU1, GRE3, and IRA2, increased the rates of aerobic and anaerobic xylose fermentation. Importantly, we found that these genetic modifications function in another genetic background and increase the rate and yield of xylose-to-ethanol conversion in industrially relevant switchgrass hydrolysate, indicating that these specific genetic modifications may enable the sustainable production of industrial biofuels from yeast. We propose a model for how key regulatory mutations prime yeast for aerobic xylose fermentation by lowering the threshold for overflow metabolism, allowing mutations to increase xylose flux and to redirect it into fermentation products.

59 BASIC BIOLOGICAL SCIENCES↗

Machine learning prediction of enzyme optimum pH

The relationship between pH and enzyme catalytic activity, especially the optimal pH (pH opt ) at which enzymes function, is critical for biotechnological applications. Hence, computational methods to predict pH opt will enhance enzyme discovery and design by facilitating accurate identification of enzymes that function optimally at specific pH levels, and by elucidating sequence-function relationships. Here, in this study, we proposed and evaluated various machine learning methods for predicting pH opt , conducting extensive hyperparameter optimization and training over 11,000 model instances. Our results demonstrate that models utilizing language model embeddings markedly outperform other methods in predicting pHopt. We present EpHod, the best-performing model, to predict pHopt, making it publicly available to researchers. From sequence data, EpHod directly learns structural and biophysical features that relate to pH opt , including proximity of residues to the catalytic centre and the accessibility of solvent molecules. Overall, EpHod presents a promising advancement in pH opt prediction and will potentially speed up the development of enzyme technologies.

97 MATHEMATICS AND COMPUTING↗

A curated benchmark for cofolding models on kinase conformational states

Abstract Protein kinases are critical drug targets, requiring therapeutics that can modulate their active and inactive conformational states. While cofolding models can generate global folds directly from kinase sequences and ligand SMILES strings, these models have not yet been tested on their ability to recover ligand-induced-fit conformational states of the kinase proteins. Here, we introduce KinConfBench, a curated benchmark of 2225 high-quality human kinase chains to evaluate the ability of four state-of-the-art cofolding models—Boltz-2, Chai-1, Protenix, and RoseTTAFold-All-Atom—to recover both canonical and rare conformational states. We show that geometric success metrics of a ligand pose in the active site do not correlate strongly with the correct kinase conformational state, motivating a new set of dynamical benchmarks for assessing cofolding models. While all four cofolding models achieve ~60–80% prediction accuracy for kinase conformational classification, they exhibit severe mode collapse when performing multiple inferences, show negligible structural diversity in sampling induced-fit motions, and display a prevalent “apo-drift” in which most cofolding models predominantly predict the kinase to be in its ligand-free state. Our results highlight that capturing ligand-induced protein conformational diversity, not just geometric fit, is critical for next-generation structure-based drug discovery.

Sun, Kunyang↗

Metabolic Modification of Sphingobium lignivorans SYK-6 for Lignin Valorization Through the Discovery of an Unusual Transcriptional Repressor of Lignin-Derived Dimer Catabolism

Sphingobium lignivorans SYK-6 catabolizes guaiacylglycerol-..beta..-guaiacyl ether (GGE, a ..beta..-O-4-type dimer) and 1,2-diguaiacylpropane-1,3-diol (DGPD, a ..beta..-1-type dimer) derived from lignin. Recently, SLG_35860 containing TetR- and MarR-type transcriptional regulator motifs was suggested to be involved in the regulation of GGE and DGPD catabolism. Here we investigated the role of SLG_35860 in the transcriptional regulation of GGE and DGPD catabolism genes. SLG_35860 designated ligS repressed 11 genes involved in GGE and DGPD catabolism. LigS binds directly to specific sequences in the promoter region of each gene. The MarR domain was shown to be involved in these bindings; however, GGE, DGPD, and their metabolites did not function as effectors of LigS. We discovered unidentified compound(s) in the black liquor of oxygen-soda anthraquinone pulping of Japanese cedar that SYK-6 cannot metabolize and that acted as effector(s). Therefore, LigS constantly represses the transcription of the GGE and DGPD catabolism genes to low levels. Based on these findings, we examined the productivity of a polymer building block, 2-pyrone-4,6-dicarboxylic acid (PDC), from GGE, DGPD, and a GGE metabolite using an engineered ligS mutant. The rates of PDC production from each compound by this strain were 1.5-6.0 times higher than those of a PDC-producing strain carrying ligS.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

The C2 domain augments Ras GTPase-activating protein catalytic activity

Regulation of Ras GTPases by GTPase-activating proteins (GAPs) is essential for their normal signaling. Nine of the ten GAPs for Ras contain a C2 domain immediately proximal to their canonical GAP domain, and in RasGAP (p120GAP, p120RasGAP;RASA1) mutation of this domain is associated with vascular malformations in humans. Here, we show that the C2 domain of RasGAP is required for full catalytic activity toward Ras. Analyses of the RasGAP C2-GAP crystal structure, AlphaFold models, and sequence conservation reveal direct C2 domain interaction with the Ras allosteric lobe. This is achieved by an evolutionarily conserved surface centered around RasGAP residue R707, point mutation of which impairs the catalytic advantage conferred by the C2 domain in vitro. In mice,R707Cmutation phenocopies the vascular and signaling defects resulting from constitutive disruption of theRASA1gene. In SynGAP, mutation of the equivalent conserved C2 domain surface impairs catalytic activity. Our results indicate that the C2 domain is required to achieve full catalytic activity of GAPs for Ras.

Science & Technology - Other Topics↗

V-HAMSTeR v1.0.0

V-HAMSTeR is a bioinformatics software tool designed to predict the hosts of viruses directly from genomic sequences. It can be used by researchers to predict animal, prokaryotic, plant, protist or fungal viral hosts including viruses that may be fragmented or discovered in environmental metagenomic datasets. Features & Uses: The software employs a novel dual-stream deep learning architecture that dynamically fuses implicit sequence embeddings from a genomic foundation model with 13 explicit, handcrafted biological features (e.g., coding density and strand switch rates). To ensure maximum reliability, V=HAMSTeR deploys a 5-fold deep ensemble calibrated via Joint Temperature Scaling, providing users with statistically rigorous confidence probabilities. It also features an automated sequence chunking and mean-pooling module to seamlessly process variable-length contigs. Advantages Over Similar Technologies: Existing tools (e.g., IPEV, RNAVirHost) typically rely on either basic k-mers or isolated neural networks. V-HAMSTeR's hybrid architecture captures both broad genomic context and specific biological motifs that standalone foundation models often miss. Furthermore, unlike competitor tools that struggle with incomplete data or exhibit extreme overconfidence, V-HAMSTeR is explicitly benchmarked and mathematically calibrated for fragmented assemblies (1kb–10kb). This makes it uniquely robust, accurate, and trustworthy for the messy reality of real-world environmental viromics.

Grigson, Susie [Lawrence Berkeley National Laborat↗

Selective Whole-Genome Amplification as a Tool to Enrich Specimens with Low Treponema pallidum Genomic DNA Copies for Whole-Genome Sequencing

Downstream next-generation sequencing (NGS) of the syphilis spirochete Treponema pallidum subspecies pallidum (T. pallidum) is hindered by low bacterial loads and the overwhelming presence of background metagenomic DNA in clinical specimens. In this study, we investigated selective whole-genome amplification (SWGA) utilizing multiple displacement amplification (MDA) in conjunction with custom oligonucleotides with an increased specificity for the T. pallidum genome and the capture and removal of 5'-C-phosphate-G-3' (CpG) methylated host DNA using the NEBNext Microbiome DNA enrichment kit followed by MDA with the REPLI-g single cell kit as enrichment methods to improve the yields of T. pallidum DNA in isolates and lesion specimens from syphilis patients. Sequencing was performed using the Illumina MiSeq v2 500 cycle or NovaSeq 6000 SP platform. These two enrichment methods led to 93 to 98% genome coverage at 5 reads/site in 5 clinical specimens from the United States and rabbit-propagated isolates, containing >14 T. pallidum genomic copies/μL of sample for SWGA and >129 genomic copies/μL for CpG methylation capture with MDA. Variant analysis using sequencing data derived from SWGA-enriched specimens showed that all 5 clinical strains had the A2058G mutation associated with azithromycin resistance. SWGA is a robust method that allows direct whole-genome sequencing (WGS) of specimens containing very low numbers of T. pallidum, which has been challenging until now.

59 BASIC BIOLOGICAL SCIENCES↗

Convergence acceleration of Monte Carlo many-body perturbation methods by direct sampling

In the Monte Carlo many-body perturbation (MC-MP) method, the conventional correlation-correction formula, which is a long sum of products of low-dimensional integrals, is first recast into a short sum of high-dimensional integrals over electron-pair and imaginary-time coordinates. These high-dimensional integrals are then evaluated by the Monte Carlo method with random coordinates generated by the Metropolis–Hasting algorithm according to a suitable distribution. The latter algorithm, while advantageous in its ability to sample nearly any distribution, introduces autocorrelation in sampled coordinates, which in turn increases the statistical uncertainty of the integrals and thus the computational cost. It also involves wasteful rejected moves and an initial “burn-in” step as well as displays hysteresis. Here, an algorithm is proposed that directly produces a random sequence of electron-pair coordinates for the same distribution used in the MC-MP method, which is free from autocorrelation, rejected moves, a burn-in step, or hysteresis. Furthermore, this direct-sampling algorithm is shown to accelerate second- (MC-MP2) and third-order Monte Carlo many-body perturbation (MC-MP3) calculations by up to 222% and 38%, respectively.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Evaluation of the Impact of Concentration and Extraction Methods on the Targeted Sequencing of Human Viruses from Wastewater

Sequencing human viruses in wastewater is challenging due to their low abundance compared to the total microbial background. This study compared the impact of four virus concentration/extraction methods (Innovaprep, Nanotrap, Promega, and Solids extraction) on probe-capture enrichment for human viruses followed by sequencing. Different concentration/extraction methods yielded distinct virus profiles. Innovaprep ultrafiltration (following solids removal) had the highest sequencing sensitivity and richness, resulting in the successful assembly of several near-complete human virus genomes. However, it was less sensitive in detecting SARS-CoV-2 by digital polymerase chain reaction (dPCR) compared to Promega and Nanotrap. Across all preparation methods, astroviruses and polyomaviruses were the most highly abundant human viruses, and SARS-CoV-2 was rare. These findings suggest that sequencing success can be increased using methods that reduce nontarget nucleic acids in the extract, though the absolute concentration of total extracted nucleic acid, as indicated by Qubit, and targeted viruses, as indicated by dPCR, may not be directly related to targeted sequencing performance. Further, using broadly targeted sequencing panels may capture viral diversity but risks losing signals for specific low-abundance viruses. Overall, this study highlights the importance of aligning wet lab and bioinformatic methods with specific goals when employing probe-capture enrichment for human virus sequencing from wastewater.

59 BASIC BIOLOGICAL SCIENCES↗

A Chemoselective and Stereodivergent Platform of Heme‐Nitrene Transferases to Access Chiral Aryl‐β‐Amino Esters and An Investigation of the Sequence‐Activity Landscape

Engineered biocatalysts can utilize nitrene precursors to access enantioenriched amination products, yet they have not been applied to produce valuable, enantiomerically enriched noncanonical β-amino esters. Current approaches to synthesizing β-amino acids rely on pre-oxidized precursors and multistep synthetic approaches involving various protecting groups. We engineered a platform of heme enzymes for stereoselective C–H bond amination of readily available carboxylic ester derivatives to install primary amines. A directed evolution campaign coupled with sequencing of over 1000 variants enabled us to develop engineered variants that use either O-pivaloylhydroxylamine triflic acid (PONT) or hydroxylamine hydrochloride (H 2 NOH∙HCl) as aminating reagents. An analysis of the resulting sequence–activity dataset revealed additional improvements that could be made to the final variant, highlighting the utility of sequencing data to guide future steps in directed evolution campaigns. Furthermore, the evolved nitrene transferases expand the scope of accessible chiral β-amino acid building blocks for peptidomimetic applications and provide new starting points for the design and synthesis of enantioenriched β-amino acid motifs.

amino ester building blocks↗

Biases in genome reconstruction from metagenomic data

Background Advances in sequencing, assembly, and assortment of contigs into species-specific bins has enabled the reconstruction of genomes from metagenomic data (MAGs). Though a powerful technique, it is difficult to determine whether assembly and binning techniques are accurate when applied to environmental metagenomes due to a lack of complete reference genome sequences against which to check the resulting MAGs. Methods We compared MAGs derived from an enrichment culture containing ~20 organisms to complete genome sequences of 10 organisms isolated from the enrichment culture. Factors commonly considered in binning software—nucleotide composition and sequence repetitiveness—were calculated for both the correctly binned and not-binned regions. This direct comparison revealed biases in sequence characteristics and gene content in the not-binned regions. Additionally, the composition of three public data sets representing MAGs reconstructed from the Tara Oceans metagenomic data was compared to a set of representative genomes available through NCBI RefSeq to verify that the biases identified were observable in more complex data sets and using three contemporary binning software packages. Results Repeat sequences were frequently not binned in the genome reconstruction processes, as were sequence regions with variant nucleotide composition. Genes encoded on the not-binned regions were strongly biased towards ribosomal RNAs, transfer RNAs, mobile element functions and genes of unknown function. Our results support genome reconstruction as a robust process and suggest that reconstructions determined to be >90% complete are likely to effectively represent organismal function; however, population-level genotypic heterogeneity in natural populations, such as uneven distribution of plasmids, can lead to incorrect inferences.

54 ENVIRONMENTAL SCIENCES↗

Genome dependent Cas9/gRNA search time underlies sequence dependent gRNA activity

Abstract CRISPR-Cas9 is a powerful DNA editing tool. A gRNA directs Cas9 to cleave any DNA sequence with a PAM. However, some gRNA sequences mediate cleavage at higher efficiencies than others. To understand this, numerous studies have screened large gRNA libraries and developed algorithms to predict gRNA sequence dependent activity. These algorithms do not predict other datasets as well as their training dataset and do not predict well between species. Here, to better understand these discrepancies, we retrospectively examine sequence features that impact gRNA activity in 44 published data sets. We find strong evidence that gRNA sequence dependent activity is largely influenced by the ability of the Cas9/gRNA complex to find the target site rather than activity at the target site and that this drives sequence dependent differences in gRNA activity between different species. This understanding will help guide future work to understand Cas9 activity as well as efforts to identify optimal gRNAs and improve Cas9 variants.

59 BASIC BIOLOGICAL SCIENCES↗

Differential laboratory passaging of SARS-CoV-2 viral stocks impacts the in vitro assessment of neutralizing antibodies

Viral populations in natural infections can have a high degree of sequence diversity, which can directly impact immune escape. However, antibody potency is often tested in vitro with a relatively clonal viral populations, such as laboratory virus or pseudotyped virus stocks, which may not accurately represent the genetic diversity of circulating viral genotypes. This can affect the validity of viral phenotype assays, such as antibody neutralization assays. To address this issue, we tested whether recombinant virus carrying SARS-CoV-2 spike (VSV-SARS-CoV-2-S) stocks could be made more genetically diverse by passage, and if a stock passaged under selective pressure was more capable of escaping monoclonal antibody (mAb) neutralization than unpassaged stock or than viral stock passaged without selective pressures. We passaged VSV-SARS-CoV-2-S four times concurrently in three cell lines and then six times with or without polyclonal antiserum selection pressure. All three of the monoclonal antibodies tested neutralized the viral population present in the unpassaged stock. The viral inoculum derived from serial passage without antiserum selection pressure was neutralized by two of the three mAbs. However, the viral inoculum derived from serial passage under antiserum selection pressure escaped neutralization by all three mAbs. Deep sequencing revealed the rapid acquisition of multiple mutations associated with antibody escape in the VSV-SARS-CoV-2-S that had been passaged in the presence of antiserum, including key mutations present in currently circulating Omicron subvariants. These data indicate that viral stock that was generated under polyclonal antiserum selection pressure better reflects the natural environment of the circulating virus and may yield more biologically relevant outcomes in phenotypic assays. Thus, mAb assessment assays that utilize a more genetically diverse, biologically relevant, virus stock may yield data that are relevant for prediction of mAb efficacy and for enhancing biosurveillance.

60 APPLIED LIFE SCIENCES↗

EHR-BERT: A BERT-based model for effective anomaly detection in electronic health records

Objective: Physicians and clinicians rely on data contained in electronic health records (EHRs), as recorded by health information technology (HIT), to make informed decisions about their patients. The reliability of HIT systems in this regard is critical to patient safety. Consequently, better tools are needed to monitor the performance of HIT systems for potential hazards that could compromise the collected EHRs, which in turn could affect patient safety. In this paper, we propose a new framework for detecting anomalies in EHRs using sequence of clinical events. This new framework, EHR-Bidirectional Encoder Representations from Transformers (BERT), is motivated by the gaps in the existing deep-learning related methods, including high false negatives, sub-optimal accuracy, higher computational cost, and the risk of information loss. EHR-BERT is an innovative framework rooted in the BERT architecture, meticulously tailored to navigate the hurdles in the contemporary BERT method; thus, enhancing anomaly detection in EHRs for healthcare applications.Methods: The EHR-BERT framework was designed using the Sequential Masked Token Prediction (SMTP) method. This approach treats EHRs as natural language sentences and iteratively masks input tokens during both training and prediction stages. This method facilitates the learning of EHR sequence patterns in both directions for each event and identifies anomalies based on deviations from the normal execution models trained on EHR sequences.Results: Extensive experiments on large EHR datasets across various medical domains demonstrate that EHR-BERT markedly improves upon existing models. It significantly reduces the number of false positives and enhances the detection rate, thus bolstering the reliability of anomaly detection in electronic health records. This improvement is attributed to the model’s ability to minimize information loss and maximize data utilization effectively.Conclusion: EHR-BERT showcases immense potential in decreasing medical errors related to anomalous clinical events, positioning itself as an indispensable asset for enhancing patient safety and the overall standard of healthcare services. The framework effectively overcomes the drawbacks of earlier models, making it a promising solution for healthcare professionals to ensure the reliability and quality of health data.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

A Multiplexed Quantitative Analysis of Germline Single Amino Acid Variants by Targeted Proteomics in Nondepleted Human Plasma

Single amino acid variants (SAAVs) in protein sequences are often a direct result of single-nucleotide polymorphisms (SNPs). Certain germline SAAVs have shown biological relevance in different disease conditions but lack precise quantification in circulation, which could hinder functional investigations and progress in biomarker development. Here, we have developed a multiplexed liquid chromatography-selected reaction monitoring (LC-SRM) assay that monitors 5 wild-type and variant peptide pairs (Complement Factor B: CFB-R32Q/R32W, Clusterin: CLU-N317H, Fetuin B: FETUB-K360R, and Kininogen: KNG1-L212P) in nondepleted human plasma. The assay was optimized for imprecision, linearity, stability, and calibration assessments with CVs of under 20%. The wild-type and variant peptide pairs were characterized in a set of healthy individual plasma samples. These target identifications were also validated by SNP genotyping with more than 99% accuracy. For all protein targets, we observed significantly lower concentrations of WT species in the presence variant peptides. In CFB, the concentration of R32Q was significantly lower than its counterpart R32W variant and WT species. Furthermore, our results distinguished phenotypes of homozygosity and heterozygosity of the SAAV presence through direct concentration level characterization. These findings provide some insights into how SAAVs affect quantitative assessments of target peptides. The assay demonstrates a platform for proteogenomic analyses with potential applications in both research and clinical settings.

genetics↗