Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Homology modelling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Topological Signatures of Adversaries in Multimodal Alignments

Topological Data Analysis for Adversarial Detection (LANL O4937) - Detects adversarial examples in vision-language models using persistent homology and two-sample testing. Combines TDA features from CLIP embeddings with statistical methods (ME, SCF, SAMMD, C2ST) for robust detection across ImageNet, CIFAR-10/100.

Bhattarai, Manish↗

Bacterial hemophilin homologs and their specific type eleven secretor proteins have conserved roles in heme capture and are diversifying as a family

Cellular life relies on enzymes that require metals, which must be acquired from extracellular sources. Bacteria utilize surface and secreted proteins to acquire such valuable nutrients from their environment. These include the cargo proteins of the type eleven secretion system (T11SS), which have been connected to host specificity, metal homeostasis, and nutritional immunity evasion. This Sec-dependent, Gram-negative secretion system is encoded by organisms throughout the phylum Proteobacteria, including human pathogens Neisseria meningitidis, Proteus mirabilis, Acinetobacter baumannii, and Haemophilus influenzae. Experimentally verified T11SS-dependent cargo include transferrin-binding protein B (TbpB), the hemophilin homologs heme receptor protein C (HrpC), hemophilin A (HphA), the immune evasion protein factor-H binding protein (fHbp), and the host symbiosis factor nematode intestinal localization protein C (NilC). Here, we examined the specificity of T11SS systems for their cognate cargo proteins using taxonomically distributed homolog pairs of T11SS and hemophilin cargo and explored the ligand binding ability of those hemophilin cargo homologs. In vivo expression in Escherichia coli of hemophilin homologs revealed that each is secreted in a specific manner by its cognate T11SS protein. Sequence analysis and structural modeling suggest that all hemophilin homologs share an N-terminal ligand-binding domain with the same topology as the ligand-binding domains of the Haemophilus haemolyticus heme binding protein (Hpl) and HphA. We term this signature feature of this group of proteins the hemophilin ligand-binding domain. Network analysis of hemophilin homologs revealed five subclusters and representatives from four of these showed variable heme-binding activities, which, combined with sequence-structure variation, suggests that hemophilins are diversifying in function.

59 BASIC BIOLOGICAL SCIENCES↗

Quantitative SANS and multi-model analysis of spacer-dependent micellization of urea-based gemini surfactants

The micellization behavior of urea-based cationic gemini surfactants was investigated using small-angle neutron scattering (SANS) with multi-model form factor analysis. A homologous series of surfactants with urea group included in the hydrophobic tail and polymethylene spacers consisting of two to ten methylene units was analyzed using three form factor models: a core–shell ellipsoid and two variants of homogeneous ellipsoids. The results from all models show a consistent trend of the micelle structures, confirming that the spacer length critically influences micellar geometry, aggregation number, and hydration. The surfactant with four CH 2 groups in the spacer formed the largest micelles with the highest aggregation number, while longer spacers led to progressively smaller, more compact aggregates. The shell hydration—quantified as the volume fraction of heavy water within the hydrophilic region—decreased systematically with increasing spacer length due to enhanced hydrophobicity of the headgroup-spacer region. Intermicellar interactions, modeled as screened Coulomb interaction using the rescaled mean spherical approximation (RMSA), revealed the strongest electrostatic repulsion for the case of four methylene groups in the spacer, corresponding to the highest micellar charge and largest interparticle spacing. The observed spacer-dependent trends were robust across all modeling approaches, demonstrating that the spacer length serves as a key structural determinant of self-assembly in this type of urea-based gemini systems. These findings provide insight into the design of gemini surfactants with tailored aggregation behavior for applications in drug delivery, nanostructure templating, and solubilization technologies.

Core–shell ellipsoid model↗

Harnessing Machine Learning and Data Fusion for Accurate Undocumented Well Identification in Satellite Images

This study utilizes satellite data to detect undocumented oil and gas wells, which pose significant environmental concerns, including greenhouse gas emissions. Three key findings emerge from the study. Firstly, the problem of imbalanced data is addressed by recommending oversampling techniques like Rotation–GaussianBlur–Solarization data augmentation (RGS), the Synthetic Minority Over-Sampling Technique (SMOTE), or ADASYN (an extension of SMOTE) over undersampling techniques. The performance of borderline SMOTE is less effective than that of the rest of the oversampling techniques, as its performance relies heavily on the quality and distribution of data near the decision boundary. Secondly, incorporating pre-trained models trained on large-scale datasets enhances the models’ generalization ability, with models trained on one county’s dataset demonstrating high overall accuracy, recall, and F1 scores that can be extended to other areas. This transferability of models allows for wider application. Lastly, including persistent homology (PH) as an additional input improves performance for in-distribution testing but may affect the model’s generalization for out-of-distribution testing. A careful consideration of PH’s impact on overall performance and generalizability is recommended. Overall, this study provides a robust approach to identifying undocumented oil and gas wells, contributing to the acceleration of a net-zero economy and supporting environmental sustainability efforts.

SMOTE↗

Modeling Protein–Protein and Protein–Ligand Interactions by the ClusPro Team in CASP16

ABSTRACT In the CASP16 experiment, our team employed hybrid computational strategies to predict both protein–protein and protein–ligand complex structures. For protein–protein docking, we combined physics‐based sampling—using ClusPro FFT docking and molecular dynamics—with AlphaFold (AF)‐based sampling, followed by AF‐based refinement. Our method produced numerous high‐accuracy complex models, including cases where AF alone failed, underscoring the critical role of physics‐based sampling alongside deep learning‐based refinement. For protein–ligand docking, we integrated the ClusPro LigTBM template‐based approach with a machine learning‐based confidence model for rescoring. The method preserves conserved interaction fragments derived from homologous complexes, followed by local resampling using physics‐based sampling and a diffusion model. Our template‐based strategy achieved a mean lDDT‐PLI of 0.69 across 233 targets, which was highly competitive. These results demonstrate that combining physics‐based modeling with AI‐driven refinement can significantly enhance the accuracy of both protein–protein and protein–ligand structure predictions.

Ashizawa, Ryota [Department of Applied Mathematics↗

Structural characterization of ligand binding and pH-specific enzymatic activity of mouse Acidic Mammalian Chitinase

Chitin is an abundant biopolymer and pathogen-associated molecular pattern that stimulates a host innate immune response. Mammals express chitin-binding and chitin-degrading proteins to remove chitin from the body. One of these proteins, Acidic Mammalian Chitinase (AMCase), is an enzyme known for its ability to function under acidic conditions in the stomach but is also active in tissues with more neutral pHs, such as the lung. Here, we used a combination of biochemical, structural, and computational modeling approaches to examine how the mouse homolog (mAMCase) can act in both acidic and neutral environments. We measured kinetic properties of mAMCase activity across a broad pH range, quantifying its unusual dual activity optima at pH 2 and 7. We also solved high-resolution crystal structures of mAMCase in complex with oligomeric GlcNAcn, the building block of chitin, where we identified extensive conformational ligand heterogeneity. Leveraging these data, we conducted molecular dynamics simulations that suggest how a key catalytic residue could be protonated via distinct mechanisms in each of the two environmental pH ranges. These results integrate structural, biochemical, and computational approaches to deliver a more complete understanding of the catalytic mechanism governing mAMCase activity at different pH. Engineering proteins with tunable pH optima may provide new opportunities to develop improved enzyme variants, including AMCase, for therapeutic purposes in chitin degradation.

59 BASIC BIOLOGICAL SCIENCES↗

NEAR: Neural Embeddings for Amino acid Relationships

Protein language models (PLMs) have recently demonstrated potential to supplant classical protein database search methods based on sequence alignment, but are slower than common alignment-based tools and appear to be prone to a high rate of false labeling. Here, we present NEAR, a method based on neural representation learning that is designed to improve both speed and accuracy of search for likely homologs in a large protein sequence database. NEAR’s ResNet embedding model is trained using contrastive learning guided by trusted sequence alignments. It computes per-residue embeddings for target and query protein sequences, and identifies alignment candidates with a pipeline consisting of residue-level k-NN search and a simple neighbor aggregation scheme. Tests on a benchmark consisting of trusted remote homologs and randomly shuffled decoy sequences reveal that NEAR substantially improves accuracy relative to state-of-the-art PLMs, with lower memory requirements and faster embedding and search speed. While these results suggest that the NEAR model may be useful for standalone homology detection with increased sensitivity over standard alignment-based methods, in this manuscript we focus on a more straightforward analysis of the model’s value as a high-speed pre-filter for sensitive annotation. In that context, NEAR is at least 5x faster than the pre-filter currently used in the widely-used profile hidden Markov model (pHMM) search tool HMMER3, and also outperforms the pre-filter used in our fast pHMM tool, nail.

59 BASIC BIOLOGICAL SCIENCES↗

The crystal structure of bacteriophage λ RexA provides novel insights into the DNA binding properties of Rex-like phage exclusion proteins

Abstract RexA and RexB function as an exclusion system that prevents bacteriophage T4rII mutants from growing on Escherichia coli λ phage lysogens. Recent data established that RexA is a non-specific DNA binding protein that can act independently of RexB to bias the λ bistable switch toward the lytic state, preventing conversion back to lysogeny. The molecular interactions underlying these activities are unknown, owing in part to a dearth of structural information. Here, we present the 2.05-Å crystal structure of the λ RexA dimer, which reveals a two-domain architecture with unexpected structural homology to the recombination-associated protein RdgC. Modelling suggests that our structure adopts a closed conformation and would require significant domain rearrangements to facilitate DNA binding. Mutagenesis coupled with electromobility shift assays, limited proteolysis, and double electron–electron spin resonance spectroscopy support a DNA-dependent conformational change. In vivo phenotypes of RexA mutants suggest that DNA binding is not a strict requirement for phage exclusion but may directly contribute to modulation of the bistable switch. We further demonstrate that RexA homologs from other temperate phages also dimerize and bind DNA in vitro. Collectively, these findings advance our mechanistic understanding of Rex functions and provide new evolutionary insights into different aspects of phage biology.

Biochemistry & Molecular Biology↗

Pathways to Electrochemical Ironmaking at Scale Via the Direct Reduction of Fe 2 O 3

Electrochemical ironmaking can provide an energy efficient, zero-emissions alternative to traditional methods of ironmaking, but the scalability of low-temperature electrochemical cells may be constrained by reactor throughput and the availability of acceptable feedstocks. Electrodes directly converting solid iron-oxide particles to metal circumvent traditional mass-transport limitations but are sensitive to both the particle size and nanoscale morphology of reactants. Furthermore, the effect of these properties on reactor throughput has not been systematically studied at model electrowinning surfaces. Here, we have used size-controlled, homologous α-Fe 2 O 3 particles to study how the nanoscale morphology of oxides influences the obtainable current density toward Fe metal and integrated these results in a technoeconomic model for alkaline iron electrowinning systems. Micron-scale α-Fe 2 O 3 with nanoscale porosity can be used to form Fe at current densities commensurate with industrial water electrolysis (>0.6 A cm –2 ) in the absence of external convection, providing a path to cost-competitive and scalable ironmaking using electrochemistry.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Functional Relevance of CASP16 Nucleic Acid Predictions as Evaluated by Structure Providers

ABSTRACT Accurate biomolecular structure prediction enables the prediction of mutational effects, the speculation of function based on predicted structural homology, the analysis of ligand binding modes, experimental model building, and many other applications. Such algorithms to predict essential functional and structural features remain out of reach for biomolecular complexes containing nucleic acids. Here, we report a quantitative and qualitative evaluation of nucleic acid structures for the CASP16 blind prediction challenge by 12 of the experimental groups who provided nucleic acid targets. Blind predictions accurately model secondary structure and some aspects of tertiary structure, including reasonable global folds for some complex RNAs; however, predictions often lack accuracy in the regions of highest functional importance. All models have inaccuracies in non‐canonical regions where, for example, the nucleic‐acid backbone bends, deviating from an A‐form helix geometry, or a base forms a non‐standard hydrogen bond (not a Watson‐Crick base pair). These bends and non‐canonical interactions are integral to forming functionally important regions such as RNA enzymatic active sites. Additionally, the modeling of conserved and functional interfaces between nucleic acids and ligands, proteins, or other nucleic acids remains poor. For some targets, the experimental structures may not represent the only structure the biomolecular complex occupies in solution or in its functional life cycle, posing a future challenge for the community.

Biochemistry & Molecular Biology↗

Partial spectral flow in the D1D5 CFT

The two-dimensional 𝒩 = 4 superconformal algebra has a free field realization with four bosons and four fermions. There is an automorphism of the algebra called spectral flow. Under spectral flow, the four fermions are transformed together. In this paper, we study partial spectral flow where only two of the four fermions are transformed. Partial spectral flow is applied to the D1D5 CFT where a marginal deformation moves the CFT away from the free point. The partial spectral flow is broken by the deformation. We show that this effect can be studied due to a transformation of the deformation which is well-defined under partial spectral flow. As a result in the spectrum, we demonstrate how to compute the second-order energy lift of a D1D5P state through its partial spectral flowed state. We find that D1D5P states related by partial spectral flow do not have the same lift in general.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

The carbon-concentrating mechanism of the extremophilic red microalga Cyanidioschyzon merolae

Abstract Cyanidioschyzon merolae is an extremophilic red microalga which grows in low-pH, high-temperature environments. The basis of C. merolae ’s environmental resilience is not fully characterized, including whether this alga uses a carbon-concentrating mechanism (CCM). To determine if C. merolae uses a CCM, we measured CO 2 uptake parameters using an open-path infra-red gas analyzer and compared them to values expected in the absence of a CCM. These measurements and analysis indicated that C. merolae had the gas-exchange characteristics of a CCM-operating organism: low CO 2 compensation point, high affinity for external CO 2 , and minimized rubisco oxygenation. The biomass δ 13 C of C. merolae was also consistent with a CCM. The apparent presence of a CCM in C. merolae suggests the use of an unusual mechanism for carbon concentration, as C. merolae is thought to lack a pyrenoid and gas-exchange measurements indicated that C. merolae primarily takes up inorganic carbon as carbon dioxide, rather than bicarbonate. We use homology to known CCM components to propose a model of a pH-gradient-based CCM, and we discuss how this CCM can be further investigated.

60 APPLIED LIFE SCIENCES↗

EC-Bench: A Benchmark for Enzyme Commission Number Prediction

Enzymes are proteins that catalyze specific biochemical reactions in cells. Enzyme Commission (EC) numbers are used to annotate enzymes in a four-level hierarchy that classifies enzymes based on the specific chemical reactions they catalyze. Accurate EC number prediction is essential for understanding enzyme functions. Despite the availability of numerous methods for predicting EC numbers from protein sequences, there is no unified framework for evaluating and studying such methods systematically. This gap limits the ability of the community to identify the most effective approaches for enzyme annotation. We introduce EC-Bench, a benchmark for EC number prediction, consisting of 1) an initial representative set of existing methods (including homology-based, deep learning, contrastive learning, and language model methods), 2) existing and novel accuracy and efficiency performance metrics, and 3) selected datasets to allow for comprehensive comparative study. EC-Bench is open-source and provides a framework for researchers to not only compare among existing methods objectively under uniform conditions, but also to introduce and effectively evaluate performance of new methods in a comparative framework. To demonstrate the utility of EC-Bench, we perform extensive experimentation to compare the existing EC number prediction methods and establish their advantages and disadvantages in a variety of prediction tasks, namely “exact EC number prediction”, “EC number completion” and (partial or additional) “EC number recommendation”. We find wide variation in the performance of different methods, but also subtle but potentially useful differences in the performance of different methods across tasks and for different parts of the EC hierarchy.

59 BASIC BIOLOGICAL SCIENCES↗

The essential Rhodobacter sphaeroides CenKR two-component system regulates cell division and envelope biosynthesis

Bacterial two-component systems (TCSs) often function through the detection of an extracytoplasmic stimulus and the transduction of a signal by a transmembrane sensory histidine kinase. This kinase then initiates a series of reversible phosphorylation modifications to regulate the activity of a cognate, cytoplasmic response regulator as a transcription factor. Several TCSs have been implicated in the regulation of cell cycle dynamics, cell envelope integrity, or cell wall development in Escherichia coli and other well-studied Gram-negative model organisms. However, many α-proteobacteria lack homologs to these regulators, so an understanding of how α-proteobacteria orchestrate extracytoplasmic events is lacking. In this work we identify an essential TCS, CenKR ( C ell en velope K inase and R egulator), in the α-proteobacterium Rhodobacter sphaeroides and show that modulation of its activity results in major morphological changes. Using genetic and biochemical approaches, we dissect the requirements for the phosphotransfer event between CenK and CenR, use this information to manipulate the activity of this TCS in vivo , and identify genes that are directly and indirectly controlled by CenKR in Rb . sphaeroides . Combining ChIP-seq and RNA-seq, we show that the CenKR TCS plays a direct role in maintenance of the cell envelope, regulates the expression of subunits of the Tol-Pal outer membrane division complex, and indirectly modulates the expression of peptidoglycan biosynthetic genes. CenKR represents the first TCS reported to directly control the expression of Tol-Pal machinery genes in Gram-negative bacteria, and we predict that homologs of this TCS serve a similar function in other closely related organisms. We propose that Rb . sphaeroides genes of unknown function that are directly regulated by CenKR play unknown roles in cell envelope biosynthesis, assembly, and/or remodeling in this and other α-proteobacteria.

Lakey, Bryan D. (ORCID:0000000332795242)↗

Understanding Electric Vehicle Range and Charging Needs: Interactions Between Ambient Temperature, Commute Patterns, and State-of-Charge Usage

Electric vehicle (EV) performance can vary substantially under real-world operating conditions, particularly due to ambient temperature effects on energy consumption, battery behavior, and thermal management requirements. This study quantifies how weather conditions, daily driving patterns, and State-of-Charge (SOC) usage strategies jointly influence EV driving range, charging frequency, and overall energy efficiency. A detailed and experimentally validated Autonomie vehicle model is developed, integrating a powertrain, a mono-zonal cabin model, and a battery electro-thermal model. Three battery sizes (200-, 300-, and 400-mile homologated ranges) are assessed across five commute profiles (20–200 miles) and six ambient temperatures (−18 °C to 50 °C), including scenarios with and without preconditioning. Results show that extreme temperatures could significantly decrease the maximum achievable range by up to 55% in cold conditions (−18 °C) and 40% in hot conditions (50 °C), relative to moderate conditions. Larger battery packs retain a greater fraction of their nominal range under thermal stress, while smaller packs experience sharper relative penalties due to the higher contribution of thermal loads to total energy demand. The analysis further demonstrates that limiting operation to partial SOC windows (e.g., 80–20%), a common real-world practice, significantly reduces achievable range and increases charging frequency, particularly in cold weather. Thermal preconditioning while plugged in is shown to mitigate these effects for short trips, reducing energy consumption by up to 31% in hot conditions and 7% in cold conditions. The findings demonstrate how climate, SOC usage behavior, and thermal management jointly shape the practical driving capability of EVs, highlighting the importance of efficient thermal management and realistic user charging strategies for ensuring reliable EV operation across diverse climatic scenarios.

33 ADVANCED PROPULSION SYSTEMS↗

Simultaneous enhancement of multiple functional properties using evolution-informed protein design

Abstract A major challenge in protein design is to augment existing functional proteins with multiple property enhancements. Altering several properties likely necessitates numerous primary sequence changes, and novel methods are needed to accurately predict combinations of mutations that maintain or enhance function. Models of sequence co-variation (e.g., EVcouplings), which leverage extensive information about various protein properties and activities from homologous protein sequences, have proven effective for many applications including structure determination and mutation effect prediction. We apply EVcouplings to computationally design variants of the model protein TEM-1 β -lactamase. Nearly all the 14 experimentally characterized designs were functional, including one with 84 mutations from the nearest natural homolog. The designs also had large increases in thermostability, increased activity on multiple substrates, and nearly identical structure to the wild type enzyme. This study highlights the efficacy of evolutionary models in guiding large sequence alterations to generate functional diversity for protein design applications.

59 BASIC BIOLOGICAL SCIENCES↗

Generative $β$-hairpin design using a residue-based physicochemical property landscape

De novo peptide design is a new frontier that has broad application potential in the biological and biomedical fields. Most existing models for de novo peptide design are largely based on sequence homology that can be restricted based on evolutionarily derived protein sequences and lack the physicochemical context essential in protein folding. Generative machine learning for de novo peptide design is a promising way to synthesize theoretical data that are based on, but unique from, the observable universe. In this study, we created and tested a custom peptide generative adversarial network intended to design peptide sequences that can fold into the -hairpin secondary structure. This deep neural network model is designed to establish a preliminary foundation of the generative approach based on physicochemical and conformational properties of 20 canonical amino acids, for example, hydrophobicity and residue volume, using extant structure-specific sequence data from the PDB. The beta generative adversarial network model robustly distinguishes secondary structures of hairpin from α helix and intrinsically disordered peptides with an accuracy of up to 96% and generates artificial -hairpin peptide sequences with minimum sequence identities around 31% and 50% when compared against the current NCBI PDB and nonredundant databases, respectively. These results highlight the potential of generative models specifically anchored by physicochemical and conformational property features of amino acids to expand the sequence-to-structure landscape of proteins beyond evolutionary limits.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Immunization of cows with HIV envelope trimers generates broadly neutralizing antibodies to the V2-apex from the ultralong CDRH3 repertoire

The generation of broadly neutralizing antibodies (bnAbs) to conserved epitopes on HIV Envelope (Env) is one of the cornerstones of HIV vaccine research. The animal models commonly used for HIV do not reliably produce a potent broadly neutralizing serum antibody response, with the exception of cows. Cows have previously produced a CD4 binding site response by homologous prime and boosting with a native-like Env trimer. In small animal models, other engineered immunogens were shown to focus antibody responses to the bnAb V2-apex region of Env. Here, we immunized two groups of cows (n = 4) with two regimens of V2-apex focusing Env immunogens to investigate whether antibody responses could be generated to the V2-apex on Env. Group 1 was immunized with chimpanzee simian immunodeficiency virus (SIV)-Env trimer that shares its V2-apex with HIV, followed by immunization with C108, a V2-apex focusing immunogen, and finally boosted with a cross-clade native-like trimer cocktail. Group 2 was immunized with HIV C108 Env trimer followed by the same HIV trimer cocktail as Group 1. Longitudinal serum analysis showed that one cow in each group developed serum neutralizing antibody responses to the V2-apex. Eight and 11 bnAbs were isolated from Group 1 and Group 2 cows, respectively, and showed moderate breadth and potency. Potent and broad responses in this study developed much later than previous cow immunizations that elicited CD4bs bnAbs responses and required several different immunogens. All isolated bnAbs were derived from the ultralong CDRH3 repertoire. The finding that cow antibodies can target more than one broadly neutralizing epitope on the HIV surface reveals the generality of elongated structures for the recognition of highly glycosylated proteins. The exclusive isolation of ultralong CDRH3 bnAbs, despite only comprising a small percent of the cow repertoire, suggests these antibodies outcompete the long and short CDRH3 antibodies during the bnAb response.

Microbiology↗