Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Homology modelling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Summary of Research Report

Ten papers, published in various publications, on buckling, and the effects of imperfections on various structures are presented. These papers are: (1) Buckling mode localization in elastic plates due to misplacement in the stiffner location; (2) On vibrational imperfection sensitivity on Augusti's model structure in the vicinity of a non-linear static state; (3) Imperfection sensitivity due to elastic moduli in the Roorda Koiter frame; (4) Buckling mode localization in a multi-span periodic structure with a disorder in a single span; (5) Prediction of natural frequency and buckling load variability due to uncertainty in material properties by convex modeling; (6) Derivation of multi-dimensional ellipsoidal convex model for experimental data; (7) Passive control of buckling deformation via Anderson localization phenomenon; (8)Effect of the thickness and initial im perfection on buckling on composite cylindrical shells: asymptotic analysis and numerical results by BOSOR4 and PANDA2; (9) Worst case estimation of homology design by convex analysis; (10) Buckling of structures with uncertain imperfections - Personal perspective.

Elishakoff, Isaac↗

Cell wall biology of the moss Physcomitrium patens

Abstract The moss Physcomitrium (previously Physcomitrella) patens is a non-vascular plant belonging to the bryophytes that has been used as a model species to study the evolution of plant cell wall structure and biosynthesis. Here, we present an updated review of the cell wall biology of P. patens. Immunocytochemical and structural studies have shown that the cell walls of P. patens mainly contain cellulose, hemicelluloses (xyloglucan, xylan, glucomannan, and arabinoglucan), pectin, and glycoproteins, and their abundance varies among different cell types and at different plant developmental stages. Genetic and biochemical analyses have revealed that a number of genes involved in cell wall biosynthesis are functionally conserved between P. patens and vascular plants, indicating that the common ancestor of mosses and vascular plants had already acquired most of the biosynthetic machinery to make various cell wall polymers. Although P. patens does not synthesize lignin, homologs of the phenylpropanoid biosynthetic pathway genes exist in P. patens and they play an essential role in the production of caffeate derivatives for cuticle formation. Further genetic and biochemical dissection of cell wall biosynthetic genes in P. patens promises to provide additional insights into the evolutionary history of plant cell wall structure and biosynthesis.

Plant Sciences↗

Large language models generate functional protein sequences across diverse families

Deep-learning language models have shown promise in various biotechnological applications, including protein design and engineering. Here, in this paper, we describe ProGen, a language model that can generate protein sequences with a predictable function across large protein families, akin to generating grammatically and semantically correct natural language sentences on diverse topics. The model was trained on 280 million protein sequences from >19,000 families and is augmented with control tags specifying protein properties. ProGen can be further fine-tuned to curated sequences and tags to improve controllable generation performance of proteins from families with sufficient homologous samples. Artificial proteins fine-tuned to five distinct lysozyme families showed similar catalytic efficiencies as natural lysozymes, with sequence identity to natural proteins as low as 31.4%. ProGen is readily adapted to diverse protein families, as we demonstrate with chorismate mutase and malate dehydrogenase.

59 BASIC BIOLOGICAL SCIENCES↗

IMITATION SWITCH is required for normal chromatin structure and gene repression in PRC2 target domains

Significance Polycomb Repressive Complex 2 (PRC2) methylates histones to regulate multicellular development, maintenance of stem cell identity, X-chromosome inactivation, and other important processes. Given these essential roles, there is significant interest in identifying components that function with PRC2 to establish and maintain transcriptionally repressive heterochromatin. Here we document an unexpected new role for a well-studied and conserved chromatin remodeling factor, ISWI. We found that the Neurospora ISWI homolog is required for normal facultative heterochromatin structure and gene repression at PRC2 target regions, and we defined requirements for ATP-dependent catalytic activity and accessory regulatory proteins. These findings provide mechanistic insights into the formation and function of facultative heterochromatin in a model eukaryote.

Kamei, Masayuki↗

Structural Models and Sequence Alignment Results of the Rhodospirillum rubrum Proteome

This dataset contains the structural models for the primary transcripts of the Rhodospirillum rubrum proteome as well as sequence alignment results for a subset of the encoded proteins. For each protein, the five models inferred from AlphaFold 2 are provided. The largest pTM-scoring model for each protein was energy minimized; this minimized structure as well as its AlphaFold pickle output file are also provided. This set of structures represent an alternate source of models for the R. rubrum proteome to those available in the AlphaFold Protein Structure Database. For proteins that have been annotated as hypothetical, sequence alignment results from the HHblits and SAdLSA alignment methods are provided. These methods are often more capable to resolve sequence homology than other methods. Therefore, the results from both HHblits and SAdLSA are provided to identify possible homologs for these challenging proteins. Numerous sequence databases are utilized for these alignments. References AlphaFold v2 Multimer: https://doi.org/10.1101/2021.10.04.463034. References HHBlits: https://doi.org/10.1186/s12859-019-3019-7. References SAdLSA: https://doi.org/10.3389/fbinf.2021.689960.

59 BASIC BIOLOGICAL SCIENCES↗

TLife-LSTM: Forecasting Future COVID-19 Progression with Topological Signatures of Atmospheric Conditions

Understanding the impact of atmospheric conditions on SARS-CoV2 is critical to model COVID-19 dynamics and sheds a light on the future spread around the world. Furthermore, geographic distri- butions of expected clinical severity of COVID-19 may be closely linked to prior history of respiratory diseases and changes in humidity, tem- perature, and air quality. In this context, we postulate that by tracking topological features of atmospheric conditions over time, we can provide a quanti?able structural distribution of atmospheric changes that are likely to be related to COVID-19 dynamics. As such, we apply the machinery of persistence homology on time series of graphs to extract topological signatures and to follow geographical changes in relative humidity and temperature. We develop an integrative machine learning framework named Topological Lifespan LSTM (TLife-LSTM) and test its predictive capabilities on forecasting the dynamics of SARS-CoV2 cases. We validate our framework using the number of con?rmed cases and hospitalization rates recorded in the states of Washington and California in the USA. Our results demonstrate the predictive potential of TLife-LSTM in forecasting the dynamics of COVID-19 and modeling its complex spatio-temporal spread dynamics.

Gel, Yulia R.↗

Molecular basis and biological relevance of bacterial and plant pinoresinol/lariciresinol reductase specificities

A bacterial pinoresinol/lariciresinol reductase (PLR) homolog named NrPinZ was obtained from a Novosphingobium rhizosphaerae sp. LY bacterial strain, with NrPinZ being part of its 5-step biochemical system catabolizing pinoresinol into coniferyl aldehyde and vanillin. Recombinant NrPinZ reduces racemic 8–8′ furanofuran lignans [(±)-pinoresinols, medioresinols, and syringaresinols] with similar overall catalytic efficiencies. In those reductions, only one of the two furan ring systems is reduced. Two other bacterial PLR homologs, NaPinZ and SlPinZ, from N. aromaticivorans F199 and Sphingobium lignivorans SYK-6, respectively, had comparable substrate versatilities and catalytic efficacies. Plant PLR homologs, by comparison, are either enantiospecific, enantioselective, or variants thereof, being able to reduce either one or both furan rings. For example, a recombinant enantioselective PLR (PLR_Tp2) from western red cedar (Thuja plicata) preferentially reduces both (+)-pinoresinol furan rings to afford (−)-secoisolariciresinol. BoltZ-2 modeling of NrPinZ and PLR_Tp2, together with substrate docking of (+)- and (−)-pinoresinols, medioresinols, and syringaresinols, was very instructive. The NrPinZ active site P1/P2 sub-pockets allow for both racemic forms to be catabolized. Conversely, the smaller P1 pocket in PLR_Tp2 preferentially positions (+)-pinoresinol for downstream metabolism into (−)-secoisolariciresinol, thereby providing a biochemical explanation for the different stereochemical outcomes. NrPinZ, NaPinZ, and SlPinZ, catalyzing substrate versatile catabolism of both racemic forms, may have important ramifications for gymnosperm and angiosperm lignin and lignan biodegradation, including its evolutionary significance and potential in enzyme engineering.

Boltz-2 molecular modeling↗

Utilizing Amino Acid Composition and Entropy of Potential Open Reading Frames to Identify Protein-Coding Genes

One of the main steps in gene-finding in prokaryotes is determining which open reading frames encode for a protein, and which occur by chance alone. There are many different methods to differentiate the two; the most prevalent approach is using shared homology with a database of known genes. This method presents many pitfalls, most notably the catch that you only find genes that you have seen before. The four most popular prokaryotic gene-prediction programs (GeneMark, Glimmer, Prodigal, Phanotate) all use a protein-coding training model to predict protein-coding genes, with the latter three allowing for the training model to be created ab initio from the input genome. Different methods are available for creating the training model, and to increase the accuracy of such tools, we present here GOODORFS, a method for identifying protein-coding genes within a set of all possible open reading frames (ORFS). Our workflow begins with taking the amino acid frequencies of each ORF, calculating an entropy density profile (EDP), using KMeans to cluster the EDPs, and then selecting the cluster with the lowest variation as the coding ORFs. To test the efficacy of our method, we ran GOODORFS on 14,179 annotated phage genomes, and compared our results to the initial training-set creation step of four other similar methods (Glimmer, MED2, PHANOTATE, Prodigal). We found that GOODORFS was the most accurate (0.94) and had the best F1-score (0.85), while Glimmer had the highest precision (0.92) and PHANOTATE had the highest recall (0.96).

59 BASIC BIOLOGICAL SCIENCES↗

An intron within the 16S ribosomal RNA gene of the archaeon Pyrobaculum aerophilum

The 16S rRNA genes of Pyrobaculum aerophilum and Pyrobaculum islandicum were amplified by the polymerase chain reaction, and the resulting products were sequenced directly. The two organisms are closely related by this measure (over 98% similar). However, they differ in that the (lone) 16S rRNA gene of Pyrobaculum aerophilum contains a 713-bp intron not seen in the corresponding gene of Pyrobaculum islandicum. To our knowledge, this is the only intron so far reported in the small subunit rRNA gene of a prokaryote. Upon excision the intron is circularized. A secondary structure model of the intron-containing rRNA suggests a splicing mechanism of the same type as that invoked for the tRNA introns of the Archaea and Eucarya and 23S rRNAs of the Archaea. The intron contains an open reading frame whose protein translation shows no certain homology with any known protein sequence.

NASA Discipline Exobiology↗

On the correlation between the stress exponent for creep determined by nanoindentation and the mechanism of action enabling stress relief in indium

Instrumented indentation performed at room temperature with a Berkovich and 10 μm radius sphere has been used to measure the stress exponent for creep before and after the strain burst observed in well-annealed, high-purity indium. Before the strain burst, the measured values are successfully rationalized using a new model based on stress directed diffusional flow along the interface between the indenter tip and test specimen. After the strain burst, the measured stress exponents are found to be representative of dislocation glide and climb assisted glide. Here these results are compared and contrasted to the previous experimental investigations and modeling efforts of Feng et al., Lucas et al., and Li et al. Collectively, the experimental observations and rationalization presented here provide significant new insight into the mechanisms of action that control the competition for stress relief in small, constrained volumes of crystalline metals subjected to high homologous temperatures.

36 MATERIALS SCIENCE↗

AQuaRef: machine learning accelerated quantum refinement of protein structures

Cryo-EM and X-ray crystallography provide crucial experimental data for obtaining atomic-detail models of biomacromolecules. Refining these models relies on library-based stereochemical data, which, in addition to being limited to known chemical entities, do not include meaningful noncovalent interactions. Quantum mechanical (QM) calculations could alleviate these issues but are too expensive for large molecules. Here we present a novel AI-enabled Quantum Refinement (AQuaRef) based on AIMNet2 machine learned interatomic potential (MLIP) mimicking QM at substantially lower computational costs. By refining 41 cryo-EM and 30 X-ray structures, we show that this approach yields atomic models with superior geometric quality compared to standard techniques, while maintaining an equal or better fit to experimental data. Notably, AQuaRef aids in determining proton positions, as illustrated in the challenging case of short hydrogen bonds in the parkinsonism-associated human protein DJ-1 and its bacterial homolog YajL.

Zubatyuk, Roman [Carnegie Mellon University, Pitts↗

De novo design of small beta barrel proteins

Small beta barrel proteins are attractive targets for computational design because of their considerable functional diversity despite their very small size (<70 amino acids). However, there are considerable challenges to designing such structures, and there has been little success thus far. Because of the small size, the hydrophobic core stabilizing the fold is necessarily very small, and the conformational strain of barrel closure can oppose folding; also intermolecular aggregation through free beta strand edges can compete with proper monomer folding. Here, we explore the de novo design of small beta barrel topologies using both Rosetta energy–based methods and deep learning approaches to design four small beta barrel folds: Src homology 3 (SH3) and oligonucleotide/oligosaccharide-binding (OB) topologies found in nature and five and six up-and-down-stranded barrels rarely if ever seen in nature. Both approaches yielded successful designs with high thermal stability and experimentally determined structures with less than 2.4 Å rmsd from the designed models. Using deep learning for backbone generation and Rosetta for sequence design yielded higher design success rates and increased structural diversity than Rosetta alone. The ability to design a large and structurally diverse set of small beta barrel proteins greatly increases the protein shape space available for designing binders to protein targets of interest.

59 BASIC BIOLOGICAL SCIENCES↗

Multi-Domain Routing in Delay Tolerant Networks

The goal of Delay Tolerant Networking (DTN) is to provide the missing ingredient for the ever-growing collection of communicating nodes in our solar system to become a Solar System Internet (SSI). Great strides have been made in modeling particular types of DTNs, such as schedule- or discovery-based. Now, analogously to the Internet, these smaller DTNs can be considered routing domains which must be stitched together to form the overall SSI. In this paper, we propose a framework for cross-domain routing in DTNs as well as methodologies for detecting these sub-domains. Example time-varying networks are given to demonstrate the techniques proposed. A basic component is the mathematical theory of sheaves, which unifies the underlying model of DTN routing algorithms, by giving rise to routing sheaves – these can be defined for the dynamic and scheduled networks as noted above, and can also be used to define the interfaces between these domains in order to route across them. An immediate application would be routing across discovery-based networks connected by scheduled networks. These DTN subdomains remain elusive, however, and need to become well-defined and properly sized for tractable computability. In particular, a balance must be determined between areas that are too large (i.e. large matrix computations) versus areas that are too small (i.e. “many” single-noded domains). Moreover, the connections between the domains should, at least locally, be chosen to optimize data flow and connectivity: we address this in three ways. First, tools from persistent homology are given to understand underlying structures, reminiscent of hierarchies in the Internet Protocol (IP) addressing. Second, we construct a notion of temporal graph curvature based on network geometry to analyze flows induced by dynamical processes on these networks. Finally, Schrodinger Bridges, a tool arising from statistical physics, are proposed as a method of constructing flows on time-evolving networks with desirable properties such as speed, robustness, and load sensitivity. We construct an approach to temporal hypergraphs to simultaneously model unicast, multicast, and broadcast, using the language of scheme theory, and then consider DTN network coding as a way to achieve network-level computation and organization. The paper concludes with a discussion and ideas for future work.

Alan Hylton↗

Insight into the autoproteolysis mechanism of the RsgI9 anti‐σ factor from Clostridium thermocellum

Abstract Clostridium thermocellum is a potential microbial platform to convert abundant plant biomass to biofuels and other renewable chemicals. It efficiently degrades lignocellulosic biomass using a surface displayed cellulosome, a megadalton sized multienzyme containing complex. The enzymatic composition and architecture of the cellulosome is controlled by several transmembrane biomass‐sensing RsgI‐type anti‐σ factors. Recent studies suggest that these factors transduce signals from the cell surface via a conserved RsgI extracellular (CRE) domain (also called a periplasmic domain) that undergoes autoproteolysis through an incompletely understood mechanism. Here we report the structure of the autoproteolyzed CRE domain from the C. thermocellum RsgI9 anti‐σ factor, revealing that the cleaved fragments forming this domain associate to form a stable α/β/α sandwich fold. Based on AlphaFold2 modeling, molecular dynamics simulations, and tandem mass spectrometry, we propose that a conserved Asn‐Pro bond in RsgI9 autoproteolyzes via a succinimide intermediate whose formation is promoted by a conserved hydrogen bond network holding the scissile peptide bond in a strained conformation. As other RsgI anti‐σ factors share sequence homology to RsgI9, they likely autoproteolyze through a similar mechanism.

Takayesu, Allen↗

Heterologous Expression of Cryptomaldamide in a Cyanobacterial Host

Filamentous marine cyanobacteria make a variety of bioactive molecules that are produced by polyketide synthases, nonribosomal peptide synthetases, and hybrid pathways that are encoded by large biosynthetic gene clusters. These cyanobacterial natural products represent potential drug leads; however, thorough pharmacological investigations have been impeded by the limited quantity of compound that is typically available from the native organisms. Additionally, investigations of the biosynthetic gene clusters and enzymatic pathways have been difficult due to the inability to conduct genetic manipulations in the native producers. Here we report a set of genetic tools for the heterologous expression of biosynthetic gene clusters in the cyanobacteria Synechococcus elongatus PCC 7942 and Anabaena (Nostoc) PCC 7120. To facilitate the transfer of gene clusters in both strains, we engineered a strain of Anabaena that contains S. elongatus homologous sequences for chromosomal recombination at a neutral site and devised a CRISPR-based strategy to efficiently obtain segregated double recombinant clones of Anabaena. These genetic tools were used to express the large 28.7 kb cryptomaldamide biosynthetic gene cluster from the marine cyanobacterium Moorena (Moorea) producens JHB in both model strains. S. elongatus did not produce cryptomaldamide; however, high-titer production of cryptomaldamide was obtained in Anabaena. Furthermore, the methods developed in this study will facilitate the heterologous expression of biosynthetic gene clusters isolated from marine cyanobacteria and complex metagenomic samples.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Novel Microbial Routes to Synthesize Industrially Significant Precursor Compounds

Ethylene is the most widely employed organic precursor compound in industry. The potential to impact ethylene formation via recently discovered microbial processes is tenable using plentiful CO2 feedstocks. The overall long-term objective of this project was to develop an industrially compatible microbial process to synthesize ethylene in high yields. The key objective of this project was to fully define and initially characterized a recently discovered and genetically regulated anaerobic pathway to produce high levels of ethylene called the Dihydroxyacetone Phosphate - Ethylene Pathway in phototrophic bacteria. This was addressed through the following specific aims: 1. Fully probe the catalytic potential of all enzymes of the DHAP ethylene pathway and determine the regulatory mechanism of DHAP-ethylene pathway gene expression. 2. Discover effective and active ethylene enzymes encoded in cultured and uncultured organisms from anoxic environments. 3.Model the thermodynamics and kinetics of ethylene synthetic pathways to guide engineering efforts in integrating best performing DHAP-ethylene pathway enzymes into model bacteria chassis for enhance ethylene yields. Through this project we discovered the initially missing genetic and enzyme component of the DHAP-ethylene pathway that directly synthesized ethylene and other important industrial compounds like methane and ethane from specific substrates. We uncovered and partially characterized a nitrogenase-like reductase that functions in DHAP-ethylene pathway specifically and in methionine synthesis in general. This nitrogenase-like system is called the Methylthio-Alkane Reductase (MAR) for its ability to cleave volatile organic sulfur compounds into methanethiol (CH3-SH) for methionine synthesis and a hydrocarbon byproduct. Key to the DHAP-ethylene pathway, MAR is the essential enzyme that cleaves 2-methylthioethanol (CH3-S-CH2-CH2-OH) into ethylene. Coordinately, we uncovered that the MAR genes and genes associated with conversion of methanethiol (CH3-SH) to methionine are under genetic control of a LysR Type Transcriptional Regulator called SalR, whose activity is dependent upon the amount of sulfate available to the cell. When sulfate as the preferred sulfur source for cell growth drops below 200 micromolar, SalR become active for expressing the MAR and methionine biosynthesis genes to enable the cell to grow from volatile organic sulfur compounds and make ethylene. Metabolic thermos-kinetic modeling revealed that these MAR reactions for ethylene and other hydrocarbon production are highly thermodynamically favorable and are one of the largest driving forces for ethylene production by the DHAP-ethylene pathway for high ethylene yields. Modeling also indicated that a key aldolase and to a lesser extent an isomerase of the DHAP-ethylene pathway for production of the ethylene precursor, 2-methylthioethanol, also would increase ethylene yields. Through metagenomic mining and gene synthesis by the JGI DNA synthesis program, over 500 aldolase and isomerase homologs were synthesized and screened. From this, variants were uncovered with substantially higher activity that increased ethylene yields 5-fold via the aldolase reaction and 1.5-fold via the isomerase reaction. Each of these elements that increase ethylene production were integrated together via plasmid under appropriate gene promoter elements in the phototrophic bacterium, Rhodospirillum rubrum, resulting in at least 3 orders of magnitude increase in ethylene yield from carbon dioxide feedstock.

10 SYNTHETIC FUELS↗

Hours-long Near-UV/Optical Emission from Mildly Relativistic Outflows in Black Hole–Neutron Star Mergers

The ongoing LIGO–Virgo–KAGRA observing run O4 provides an opportunity to discover new multimessenger events, including binary neutron star (BNS) mergers such as GW170817 and the highly anticipated first detection of a multimessenger black hole–neutron star (BH–NS) merger. While BNS mergers were predicted to exhibit early optical emission from mildly relativistic outflows, it has remained uncertain whether the BH–NS merger ejecta provides the conditions for similar signals to emerge. We present the first modeling of early near-ultraviolet/optical emission from mildly relativistic outflows in BH–NS mergers. Adopting optimal binary properties, a mass ratio of q = 2, and a rapidly rotating BH, we utilize numerical relativity and general relativistic magnetohydrodynamic (GRMHD) simulations to follow the binary's evolution from premerger to homologous expansion. We use an M1 neutrino transport GRMHD simulation to self-consistently estimate the opacity distribution in the outflows and find a bright near-ultraviolet/optical signal that emerges due to jet-powered cocoon cooling emission, outshining the kilonova emission at early time. The signal peaks at an absolute magnitude of ~–15 a few hours after the merger, longer than previous estimates, which did not consider the first principles–based jet launching. By late 2024, the Rubin Observatory will have the capability to track the entire signal evolution or detect its peak up to distances of ≳1 Gpc. In 2026, ULTRASAT will conduct all-sky surveys within minutes, detecting some of these events within ~200 Mpc. The BH–NS mergers with higher mass ratios or lower BH spins would produce shorter and fainter signals.

79 ASTRONOMY AND ASTROPHYSICS↗

Topological network analysis of patient similarity for precision management of acute blood pressure in spinal cord injury

Background: Predicting neurological recovery after spinal cord injury (SCI) is challenging. Using topological data analysis, we have previously shown that mean arterial pressure (MAP) during SCI surgery predicts long-term functional recovery in rodent models, motivating the present multicenter study in patients. Methods: Intra-operative monitoring records and neurological outcome data were extracted (n = 118 patients). We built a similarity network of patients from a low-dimensional space embedded using a non-linear algorithm, Isomap, and ensured topological extraction using persistent homology metrics. Confirmatory analysis was conducted through regression methods. Results: Network analysis suggested that time outside of an optimum MAP range (hypotension or hypertension) during surgery was associated with lower likelihood of neurological recovery at hospital discharge. Logistic and LASSO (least absolute shrinkage and selection operator) regression confirmed these findings, revealing an optimal MAP range of 76–[104-117] mmHg associated with neurological recovery. Conclusions: We show that deviation from this optimal MAP range during SCI surgery predicts lower probability of neurological recovery and suggest new targets for therapeutic intervention. Funding: NIH/NINDS: R01NS088475 (ARF); R01NS122888 (ARF); UH3NS106899 (ARF); Department of Veterans Affairs: 1I01RX002245 (ARF), I01RX002787 (ARF); Wings for Life Foundation (ATE, ARF); Craig H. Neilsen Foundation (ARF); and DOD: SC150198 (MSB); SC190233 (MSB); DOE: DE-AC02-05CH11231 (DM).

59 BASIC BIOLOGICAL SCIENCES↗