Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “protein complex structure prediction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

161 records · Page 9

Programming Amphiphilic Peptoid Oligomers for Hierarchical Assembly and Inorganic Crystallization

Natural organisms make a wide variety of exquisitely complex, nano-, micro-, and macroscale structured materials in an energy-efficient and highly reproducible manner. During these processes, the information-carrying biomolecules (e.g., proteins, peptides, and carbohydrates) enable (1) hierarchical organization to assemble scaffold materials and execute high-level functions and (2) exquisite control over inorganic materials synthesis, generating biominerals whose properties are optimized for their functions. Inspired by nature, significant efforts have been devoted to developing functional materials that can rival those natural molecules by mimicking in vivo functions using engineered proteins, peptides, DNAs, sequence-defined synthetic molecules (e.g., peptoids), and other biomimetic polymers. Among them, peptoids, a new type of synthetic mimetics of peptides and proteins, have received particular attention because they combine the merits of both synthetic polymers (e.g., high chemical stability and efficient synthesis) and biomolecules (e.g., sequence programmability and biocompatibility). The lack of both chirality and hydrogen bonds in their backbone results in a highly designable peptoid-based system with reduced structural complexity and side chain-chemistry-dominated properties. Here in this Account, we present our recent efforts in this field by programming amphiphilic peptoid sequences for (1) the controlled self-assembly into different hierarchically structured nanomaterials with favorable properties and (2) manipulating inorganic (nano)crystal nucleation, growth, and assembly into superstructures. First, we designed a series of amphiphilic peptoids with controlled side chain chemistries that self-assembled into 1D highly stiff and dynamic nanotubes, 2D membrane-mimetic nanosheets, hexagonally patterned nanoribbons, and 3D nanoflowers. These crystalline nanostructures exhibited sequence-dependent properties and showed promise for different applications. The corresponding peptoid self-assembly pathways and mechanisms were also investigated by leveraging in situ atomic force microscopy studies and molecular dynamics simulations, which showed precise sequence dependency. Second, inspired by peptide- and protein-controlled formation of hierarchical inorganic nanostructures in nature, we developed peptoid-based biomimetic approaches for controlled synthesis of inorganic materials (e.g., noble metals and calcite), in which we took advantage of the substantial side chain chemistry of peptoids and investigated the relationship between the peptoid sequences and the morphology and growth kinetics of inorganic materials. For example, to overcome the challenges (e.g., complexity of protein- and peptide-folding, poor thermal and chemical stabilities) facing the area of protein- and peptide-controlled synthesis of inorganic materials, we recently reported the design of sequence-defined peptoids for controlled synthesis of highly branched plasmonic gold particles. Moreover, we developed a rule of thumb for designing peptoids that predictively enabled the morphological evolution from spherical to coral-shaped gold nanoparticles (NPs). With this Account, we hope to stimulate the research interest of chemists and materials scientists and promote the predictive synthesis of functional and robust materials through the design of sequence-defined synthetic molecules.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Mechanical coupling in the nitrogenase complex

The enzyme nitrogenase reduces dinitrogen to ammonia utilizing electrons, protons, and energy obtained from the hydrolysis of ATP. Mo-dependent nitrogenase is a symmetric dimer, with each half comprising an ATP-dependent reductase, termed the Fe Protein, and a catalytic protein, known as the MoFe protein, which hosts the electron transfer P-cluster and the active-site metal cofactor (FeMo-co). A series of synchronized events for the electron transfer have been characterized experimentally, in which electron delivery is coupled to nucleotide hydrolysis and regulated by an intricate allosteric network. We report a graph theory analysis of the mechanical coupling in the nitrogenase complex as a key step to understanding the dynamics of allosteric regulation of nitrogen reduction. This analysis shows that regions near the active sites undergo large-scale, large-amplitude correlated motions that enable communications within each half and between the two halves of the complex. Computational predictions of mechanically regions were validated against an analysis of the solution phase dynamics of the nitrogenase complex via hydrogen-deuterium exchange. These regions include the P-loops and the switch regions in the Fe proteins, the loop containing the residue β-188Ser adjacent to the P-cluster in the MoFe protein, and the residues near the protein-protein interface. In particular, it is found that: (i) within each Fe protein, the switch regions I and II are coupled to the [4Fe-4S] cluster; (ii) within each half of the complex, the switch regions I and II are coupled to the loop containing β-188Ser; (iii) between the two halves of the complex, the regions near the nucleotide binding pockets of the two Fe proteins (in particular the P-loops, located over 130 Å apart) are also mechanically coupled. Notably, we found that residues next to the P-cluster (in particular the loop containing β-188Ser) are important for communication between the two halves.

59 BASIC BIOLOGICAL SCIENCES↗

Metaproteomics reveals insights into microbial structure, interactions, and dynamic regulation in defined communities as they respond to environmental disturbance

Abstract Background Microbe-microbe interactions between members of the plant rhizosphere are important but remain poorly understood. A more comprehensive understanding of the molecular mechanisms used by microbes to cooperate, compete, and persist has been challenging because of the complexity of natural ecosystems and the limited control over environmental factors. One strategy to address this challenge relies on studying complexity in a progressive manner, by first building a detailed understanding of relatively simple subsets of the community and then achieving high predictive power through combining different building blocks (e.g., hosts, community members) for different environments. Herein, we coupled this reductionist approach with high-resolution mass spectrometry-based metaproteomics to study molecular mechanisms driving community assembly, adaptation, and functionality for a defined community of ten taxonomically diverse bacterial members of Populus deltoides rhizosphere co-cultured either in a complex or defined medium. Results Metaproteomics showed this defined community assembled into distinct microbiomes based on growth media that eventually exhibit composition and functional stability over time. The community grown in two different media showed variation in composition, yet both were dominated by only a few microbial strains. Proteome-wide interrogation provided detailed insights into the functional behavior of each dominant member as they adjust to changing community compositions and environments. The emergence and persistence of select microbes in these communities were driven by specialization in strategies including motility, antibiotic production, altered metabolism, and dormancy. Protein-level interrogation identified post-translational modifications that provided additional insights into regulatory mechanisms influencing microbial adaptation in the changing environments. Conclusions This study provides high-resolution proteome-level insights into our understanding of microbe-microbe interactions and highlights specialized biological processes carried out by specific members of assembled microbiomes to compete and persist in changing environmental conditions. Emergent properties observed in these lower complexity communities can then be re-evaluated as more complex systems are studied and, when a particular property becomes less relevant, higher-order interactions can be identified.

59 BASIC BIOLOGICAL SCIENCES↗

Dynamics-Based Peptide–MHC Binding Optimization by a Convolutional Variational Autoencoder: A Use-Case Model for CASTELO

An unsolved challenge in the development of antigen-specific immunotherapies is determining the optimal antigens to target. Comprehension of antigen–major histocompatibility complex (MHC) binding is paramount toward achieving this goal. Here, we apply CASTELO, a combined machine learning-molecular dynamics (ML-MD) approach, to identify per-residue antigen binding contributions and then design novel antigens of increased MHC-II binding affinity for a type 1 diabetes-implicated system. We build upon a small-molecule lead optimization algorithm by training a convolutional variational autoencoder (CVAE) on MD trajectories of 48 different systems across four antigens and four HLA serotypes. We develop several new machine learning metrics including a structure-based anchor residue classification model as well as cluster comparison scores. ML-MD predictions agree well with experimental binding results and free energy perturbation-predicted binding affinities. Moreover, ML-MD metrics are independent of traditional MD stability metrics such as contact area and root-mean-square fluctuations (RMSF), which do not reflect binding affinity data. Finally, our work supports the role of structure-based deep learning techniques in antigen-specific immunotherapy design.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Identification of Small-Molecule Inhibitors of Fibroblast Growth Factor 23 Signaling via In Silico Hot Spot Prediction and Molecular Docking to α-Klotho

Fibroblast growth factor 23 (FGF23) is a therapeutic target for treating hereditary and acquired hypophosphatemic disorders, such as X-linked hypophosphatemic (XLH) rickets and tumor-induced osteomalacia (TIO), respectively. FGF23-induced hypophosphatemia is mediated by signaling through a ternary complex formed by FGF23, the FGF receptor (FGFR), and α-Klotho. Currently, disorders of excess FGF23 are treated with an FGF23-blocking antibody, burosumab. Small-molecule drugs that disrupt protein/protein interactions necessary for the ternary complex formation offer an alternative to disrupting FGF23 signaling. In this study, the FGF23:α-Klotho interface was targeted to identify small-molecule protein/protein interaction inhibitors since it was computationally predicted to have a large fraction of hot spots and two druggable residues on α-Klotho. We further identified Tyr433 on the KL1 domain of α-Klotho as a promising hot spot and α-Klotho as an appropriate drug-binding target at this interface. Subsequently, we performed in silico docking of ~5.5 million compounds from the ZINC database to the interface region of α-Klotho from the ternary crystal structure. Following docking, 24 and 20 compounds were in the final list based on the lowest binding free energies to α-Klotho and the largest number of contacts with Tyr433, respectively. Herein, five compounds were assessed experimentally by their FGF23-mediated extracellular signal-regulated kinase (ERK) activities in vitro, and two of these reduced activities significantly. Both these compounds were predicted to have favorable binding affinities to α-Klotho but not have a large number of contacts with the hot spot Tyr433. ZINC12409120 was found experimentally to disrupt FGF23:α-Klotho interaction to reduce FGF23-mediated ERK activities by 70% and have a half maximal inhibitory concentration (IC50) of 5.0 ± 0.23 μM. Molecular dynamics (MD) simulations of the ZINC12409120:α-Klotho complex starting from in silico docking poses reveal that the ligand exhibits contacts with residues on the KL1 domain, the KL1–KL2 linker, and the KL2 domain of α-Klotho simultaneously, thereby possibly disrupting the regular function of α-Klotho and impeding FGF23:α-Klotho interaction. ZINC12409120 is a candidate for lead optimization.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Enhancing Biopreparedness through a Model System to Understand the Molecular Mechanisms that Lead to Pathogenesis and Disease Transmission: NW-BRaVE

The science of biopreparedness to counter biological threats hinges on understanding the fundamental principles and molecular mechanisms that lead to pathogenesis and disease transmission. Our vision to address this challenge is to create a powerful and user-friendly platform to elucidate the fundamental principles of how molecular interactions drive pathogen-host relationships and host shifts. We will enable groundbreaking discoveries by integrating a wide range of structural, genomics, proteomics, and other advanced omics measurements, along with evolutionary and artificial intelligence predictions. To make sure the system is applicable to real-world problems, we will develop it in the context of a tractable model system, the small, abundant, and accessible photosynthetic cyanobacteria and their constantly co-adapting viral pathogens, cyanophages. This model will maintain the system’s applicability to real-world problems and techniques, but the overall focus will be on elucidating general principles of detecting, assessing, and surveilling molecular interaction, adaptation, and coevolution that are system agnostic and therefore extensible to other viral-host interactions. Our overall objectives are to (1) identify the molecular complexes that comprise the cyanobacteria redox macromolecular subsystem and how they dynamically change with bacteriophage infection in situ, using cryo-electron tomography; (2) profile regulatory changes during infection using proteomics, multiomics, and experimental validation, and integrate the data with in situ structures; (3) use genomics and metagenomics to determine environmental and population factors across time scales that impact the interactions between marine cyanobacteria and their cyanophage parasites, predicting the evolutionary origins of in situ structural and functional interactions, convergence and coevolution; and (4) develop a data integration and transformation platform that facilitates the integration of in situ, proteomic, and evolutionary measurements of molecular interactions to surveil diverse hosts and parasites in various environmental contexts. These objectives address Focus Area 2 Reveal Molecular Interactions Across Biological Scales for Design of Targeted Interventions. Our powerful and user-friendly platform will enhance connections between the often-siloed fields of structure, molecular phenotype, and evolutionary genomics that are key to biopreparedness, but in need of integration (Figure 1). We will build an integrated navigation tool to facilitate the effective use of globally distributed experimental data for integrated analysis and predictive modeling. The project will develop, implement, and test a platform to assess host-pathogen molecular interactions, adaptation to hosts and host shifts, and coevolution between hosts and pathogens, successfully impacting the research community by revolutionizing abilities to study any host-pathogen interaction, encourage diverse community contributions, and gain fundamental insights into how proteins adapt to new contexts. This ability will be critical for designing early interventions to address future threats. We will build surveillance training capability, aiming for a fair and equitable response to future pandemics and biothreats.

59 BASIC BIOLOGICAL SCIENCES↗

Cryo2StructData: A Large Labeled Cryo-EM Density Map Dataset for AI-based Modeling of Protein Structures

The advent of single-particle cryo-electron microscopy (cryo-EM) has brought forth a new era of structural biology, enabling the routine determination of large biological molecules and their complexes at atomic resolution. The high-resolution structures of biological macromolecules and their complexes significantly expedite biomedical research and drug discovery. However, automatically and accurately building atomic models from high-resolution cryo-EM density maps is still time-consuming and challenging when template-based models are unavailable. Artificial intelligence (AI) methods such as deep learning trained on limited amount of labeled cryo-EM density maps generate inaccurate atomic models. To address this issue, we created a dataset called Cryo2StructData consisting of 7,600 preprocessed cryo-EM density maps whose voxels are labelled according to their corresponding known atomic structures for training and testing AI methods to build atomic models from cryo-EM density maps. Cryo2StructData is larger than existing, publicly available datasets for training AI methods to build atomic protein structures from cryo-EM density maps. We trained and tested deep learning models on Cryo2StructData to validate its quality showing that it is ready for being used to train and test AI methods for building atomic models.

59 BASIC BIOLOGICAL SCIENCES↗

Comparative Analysis of TCR and TCR-pMHC Complex Structure Prediction Tools

The rapid development of computational approaches for predicting the structures of T cell receptors (TCRs) and TCR-peptide-major histocompatibility (TCR-pMHC) complexes, accelerated by AI breakthroughs such as AlphaFold, has made it feasible to calculate these structures with increasing accuracy. Although these tools show great potential, their relative accuracy and limitations remain unclear due to the lack of standardized benchmarks. Here, we systematically evaluate seven tools for predicting isolated TCR structures together with six tools for predicting TCR-pMHC complex structures. The methods include homology-based approaches, general prediction tools using AlphaFold, TCR-specific tools derived from AlphaFold2, and the newly developed tFold-TCR model. The evaluation uses a post-training data set comprising 40 αβ TCRs and 27 TCR-pMHC complexes (21 Class I and 6 Class II). Model accuracy is assessed at global, local, and interface levels using a variety of metrics. We find that each tool offers distinct advantages in various aspects of its predictions. AlphaFold2, AlphaFold3, and tFold-TCR excel in overall accuracy of TCR structure prediction, and TCRmodel2 and AlphaFold2 perform well in overall accuracy of TCR-pMHC structure prediction. However, TCR-specific tools derived from AlphaFold2 show lower accuracy in the framework region than both homology-based methods and general-purpose tools such as AlphaFold, and challenges remain for all in modeling CDR3 loops, docking orientations, TCR-peptide interfaces, and Class II MHC-peptide interfaces. Furthermore, these findings will guide researchers in selecting appropriate tools, emphasize the importance of using multiple evaluation metrics to assess model performance, and offer suggestions for improving TCR and TCR-pMHC structure prediction tools.

Chemical structure↗

Modeling the Activity of Single Genes

The central dogma of molecular biology states that information is stored in DNA, transcribed to messenger RNA (mRNA) and then translated into proteins. This picture is significantly augmentated when we consider the action of certain proteins in regulating transcription. These transcription factors provide a feedback pathway by which genes can regulate one another's expression as mRNA and then as protein. To review: DNA, RNA and proteins have different functions. DNA is the molecular storehouse of genetic information. When cells divide, the DNA is replicated, so that each daughter cell maintains the same genetic information as the mother cell. RNA acts as a go-between from DNA to proteins. Only a single copy of DNA is present, but multiple copies of the same piece of RNA may be present, allowing cells to make huge amounts of protein. In eukaryotes (organisms with a nucleus), DNA is found in the nucleus only. RNA is copied in the nucleus then translocates(moves) outside the nucleus, where it is transcribed into proteins. Along the way, the RNA may be spliced, i.e., may have pieces cut out. RNA then attaches to ribosomes and is translated to proteins. Proteins are the machinery of the cell other than DNA and RNA, all the complex molecules of the cell are proteins. Proteins are specialized machines, each of which fulfills its own task, which may be transporting oxygen, catalyzing reactions, or responding to extracellular signals, just to name a few. One of the more interesting functions a protein may have is binding directly or indirectly to DNA to perform transcriptional regulation, thus forming a closed feedback loop of gene regulation. The structure of DNA and the central dogma were understood in the 50s; in the early 80s it became possible to make arbitrary modifications to DNA and use cellular machinery to transcribe and translate the resulting genes; more recently, genomes (i.e., the complete DNA sequence) of many organisms have been sequenced. This large-scale sequencing began with simple organisms, viruses and bacteria, progressed to eukaryotes such as yeast, and more recently (1998) progressed to a multi-cellular animal, the nematode Caenorhabditis elegans. Sequencers have now moved on to the fruit fly Drosophila melanogaster, whose sequence is slated for completion by the end of 1999. The human genome project is expected to determine the complete sequence of all 3 billion bases of human DNA within the next five years. In the wake of genome-scale sequencing, further instrumentation is being developed to assay gene expression and function on a comparably large scale. Much of the work in computational biology focuses on computational tools used in sequencing, finding genes that are related to a particular gene, finding which parts of the DNA code for proteins and which do not, understanding what proteins will be formed from a given length of DNA, predicting how the proteins will fold from a one-dimensional structure into a three dimensional structure, and so on. Much less computational work has been done regarding the function of proteins. One reason for this is that different proteins function very differently, and so work on protein function is very specific to certain classes of proteins. There are, for example, proteins such enzymes that catalyze various intracellular reactions, receptors that respond to extracellular signals and ion channels that regulate the flow of charged particles into and out of the cell. In this chapter, we will consider a particular class of proteins called transcription factors(TFs), which are responsible for regulating when a certain gene is expressed in a certain cell, which cells it is express in, and how much is expressed. Understanding these processes will involve developing a deeper understanding of transcription, translation, and the cellular processes that control those processes. All of these elements fall under the aegis of gene regulation or more narrowly transcriptional regulation. Some of the key questions in gene regulation are: What genes are expressed in a certain cell at a certain time? How does gene expression differ from cell to cell in a multicellular organism? Which proteins act as transcription factors, i.e., are important in regulating gene expression? From questions like these, we hope to understand which genes are important for various macroscopic processes. Nearly all of the cells of a multicellular organism contain the same DNA. Yet this same genetic information yields a large number of different cell types. The fundamental difference between a neuron and a liver cell, for example, is which genes are expressed. Thus understanding gene regulation is an important step in understanding development. Furthermore, understanding the usual genes that are expressed in cells may give important clues about various diseases. Some diseases, such as sickle cell anemia and cystic fibrosis, are caused by defects in single, non-regulatory genes; others, such as certain cancers, are caused when the cellular control circuitry malfunctions - an understanding of these diseases will involve pathways of multiple interacting gene products. There are numerous challenges in the area of understanding and modeling gene regulation. First and foremost, biologists would like to develop a deeper understanding of the processes involved, including which genes and families of genes are important, how they interact, etc. From a computation point of view, there has been embarrassingly little work done. In this chapter there are many areas in which we can phrase meaningful, non-trivial computational questions, but questions that have not been addressed. Some of these are purely computational (what is a good algorithm for dealing with a model of type X) and others are more mathematical (given a system with certain characteristics, what sort of model can one use? How does one find biochemical parameters from system-level behavior using as few experiments as possible?). In addition to biological and algorithmic problems, there is also the ever-present issue of theoretical biology - what general principles can be derived from these systems, what can one do with models other than just simulate time-courses, what can be deduced about a class of systems without knowing all the details? The fundamental challenge to computationalists and theorists is to add value to the biology - to use models, modeling techniques and algorithms to understand the biology in new ways.

Mjolsness, Eric↗

Computing the Relative Affinity of Chlorophylls a and b to Light-Harvesting Complex II

In plants and algae, the primary antenna protein bound to photosystem II is light-harvesting complex II (LHCII), a pigment–protein complex that binds eight chlorophyll (Chl) a molecules and six Chl b molecules. Chl a and Chl b differ only in that Chl a has a methyl group (–CH 3 ) on one of its pyrrole rings, while Chl b has a formyl group (–CHO) at that position. This blue-shifts the Chl b absorbance relative to Chl a . It is not known how the protein selectively binds the right Chl type at each site. Knowing the selection criteria would allow the design of light-harvesting complexes that bind different Chl types, modifying an organism to utilize the light of different wavelengths. The difference in the binding affinity of Chl a and Chl b in pea and spinach LHCII was calculated using multiconformation continuum electrostatics and free energy perturbation. Both methods have identified some Chl sites where the bound Chl type ( a or b ) has a significantly higher affinity, especially when the protein provides a hydrogen bond for the Chl b formyl group. However, the Chl a sites often have little calculated preference for one Chl type, so they are predicted to bind a mixture of Chl a and b . The electron density of the spinach LHCII was reanalyzed, which, however, confirmed that there is negligible Chl b in the Chl a -binding sites. Finally, it is suggested that the protein chooses the correct Chl type during folding, segregating the preferred Chl to the correct binding site.

chemical calculations↗

Improved accuracy and transferability of molecular-orbital-based machine learning: Organics, transition-metal complexes, non-covalent interactions, and transition states

We report molecular-orbital-based machine learning (MOB-ML) provides a general framework for the prediction of accurate correlation energies at the cost of obtaining molecular orbitals. The application of Nesbet’s theorem makes it possible to recast a typical extrapolation task, training on correlation energies for small molecules and predicting correlation energies for large molecules, into an interpolation task based on the properties of orbital pairs. We demonstrate the importance of preserving physical constraints, including invariance conditions and size consistency, when generating the input for the machine learning model. Numerical improvements are demonstrated for different datasets covering total and relative energies for thermally accessible organic and transition-metal containing molecules, non-covalent interactions, and transition-state energies. MOB-ML requires training data from only 1% of the QM7b-T dataset (i.e., only 70 organic molecules with seven and fewer heavy atoms) to predict the total energy of the remaining 99% of this dataset with sub-kcal/mol accuracy. This MOB-ML model is significantly more accurate than other methods when transferred to a dataset comprising of 13 heavy atom molecules, exhibiting no loss of accuracy on a size intensive (i.e., per-electron) basis. It is shown that MOB-ML also works well for extrapolating to transition-state structures, predicting the barrier region for malonaldehyde intramolecular proton-transfer to within 0.35 kcal/mol when only trained on reactant/product-like structures. Finally, the use of the Gaussian process variance enables an active learning strategy for extending the MOB-ML model to new regions of chemical space with minimal effort. We demonstrate this active learning strategy by extending a QM7b-T model to describe non-covalent interactions in the protein backbone–backbone interaction dataset to an accuracy of 0.28 kcal/mol.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Tutorial: Machine-Learning-Based CREASE-2D Analysis of 2D SAXS Profiles to Characterize Anisotropic Nanostructures in Soft Materials

We present a tutorial to guide users on how to extend the Computational Reverse Engineering Analysis of Scattering Experiments-2D (CREASE-2D) framework to interpret their experimental two-dimensional small-angle scattering (SAS) data from soft materials (e.g., polymers, peptide amphiphiles, biomolecular fibrils). Unlike most traditional SAS analysis approaches, which typically rely on azimuthally averaged onedimensional (1D) profiles, CREASE-2D utilizes the complete 2D scattering profile to reveal information about anisotropy in the structure. In past applications, CREASE has provided insights into complex structural features, including the cross-sectional shapes of assembled nanostructures and dispersity in these features, which are difficult to discern with existing analytical models. While (1D- ) CREASE has been applied to SANS and SAXS data, this tutorial shares the steps for implementing CREASE-2D using an example of a dipeptide solution system, for which we have SAXS data. We present details for these steps involved in using CREASE-2D to interpret SAXS profiles: how to preprocess SAXS data, define relevant structural features, generate three-dimensional real-space structures for specific values of these features, train a machine learning (ML) surrogate model to predict scattering profiles for given structural features, and optimize these features using genetic algorithms (GA). Then, we use these steps to interpret complex 2DSAXS data collected from dipeptide solutions that, in microscopy images, exhibit nanoscale structures that could be elliptical tubes/ flat tapes/cylinders or a combination of these cross sections. Open-source codes, computational hardware, and software requirements, as well as the strengths and limitations of this protocol, are also presented. We expect researchers working with (soft) biomaterials, peptide amphiphiles, amphiphilic polymer solutions, polymer nanocomposites, and blends of particles/polymers will find this CREASE-2D method and this tutorial of use.

CREASE↗

The role of CenKR in the coordination of Rhodobacter sphaeroides cell elongation and division

ABSTRACT Cell elongation and division are essential aspects of the bacterial life cycle that must be coordinated for viability and replication. The impact of misregulation of these processes is not well understood as these systems are often not amenable to traditional genetic manipulation. Recently, we reported on the CenKR two-component system (TCS) in the Gram-negative bacterium Rhodobacter sphaeroides that is genetically tractable, widely conserved in α-proteobacteria, and directly regulates the expression of components crucial for cell elongation and division, including genes encoding subunit of the Tol-Pal complex. In this work, we show that overexpression of cenK results in cell filamentation and chaining. Using cryo-electron microscopy (cryo-EM) and cryo-electron tomography (cryo-ET), we generated high-resolution two-dimensional (2D) images and three-dimensional (3D) volumes of the cell envelope and division septum of wild-type cells and a cenK overexpression strain finding that these morphological changes stem from defects in outer membrane (OM) and peptidoglycan (PG) constriction. By monitoring the localization of Pal, PG biosynthesis, and the bacterial cytoskeletal proteins MreB and FtsZ, we developed a model for how increased CenKR activity leads to changes in cell elongation and division. This model predicts that increased CenKR activity decreases the mobility of Pal, delaying OM constriction, and ultimately disrupting the midcell positioning of MreB and FtsZ and interfering with the spatial regulation of PG synthesis and remodeling. IMPORTANCE By coordinating cell elongation and division, bacteria maintain their shape, support critical envelope functions, and orchestrate division. Regulatory and assembly systems have been implicated in these processes in some well-studied Gram-negative bacteria. However, we lack information on these processes and their conservation across the bacterial phylogeny. In R. sphaeroides and other α-proteobacteria, CenKR is an essential two-component system (TCS) that regulates the expression of genes known or predicted to function in cell envelope biosynthesis, elongation, and/or division. Here, we leverage unique features of CenKR to understand how increasing its activity impacts cell elongation/division and use antibiotics to identify how modulating the activity of this TCS leads to changes in cell morphology. Our results provide new insight into how CenKR activity controls the structure and function of the bacterial envelope, the localization of cell elongation and division machinery, and cellular processes in organisms with importance in health, host-microbe interactions, and biotechnology.

59 BASIC BIOLOGICAL SCIENCES↗

Mechanism and dynamics of fatty acid photodecarboxylase

Photoenzymes are rare biocatalysts driven by absorption of a photon at each catalytic cycle; they inspire development of artificial photoenzymes with valuable activities. Fatty acid photodecarboxylase (FAP) is a natural photoenzyme that has potential applications in the bio-based production of hydrocarbons, yet its mechanism is far from fully understood. RATIONALE To elucidate the mechanism of FAP, we studied the wild-type (WT) enzyme from Chlorella variabilis (CvFAP) and variants with altered active-site residues using a wealth of techniques, including static and time-resolved crystallography and spectroscopy, as well as biochemical and computational approaches. RESULTS A 1.8-Å-resolution CvFAP x-ray crystal structure revealed a dense hydrogen-bonding network positioning the fatty acid carboxyl group in the vicinity of the flavin adenine dinucleotide (FAD) cofactor. Structures solved from free electron laser and low-dose synchrotron x-ray crystal data further highlighted an unusual bent shape of the oxidized flavin chromophore, and showed that the bending angle (14°) did not change upon photon absorption (step 1) or throughout the photocycle. Calculations showed that bending substantially affected the energy levels of the flavin. Structural and spectroscopic analysis of WT and mutant proteins targeting two conserved active-site residues, R451 and C432, demonstrated that both residues were crucial for proper positioning of the substrate and water molecules and for oxidation of the fatty acid carboxylate by 1 FAD* (~300 ps in WT FAP) to form FAD ∙– (step 2). Time-resolved infrared spectroscopy demonstrated that decarboxylation occured quasi-instantaneously upon this forward electron transfer, consistent with barrierless bond cleavage predicted by quantum chemistry calculations and with snapshots obtained by time-resolved crystallography. Transient absorption spectroscopy in H 2 O and D 2 O buffers indicated that back electron transfer from FAD ∙– was coupled to and limited by transfer of an exchangeable proton or hydrogen atom (step 3). Unexpectedly, concomitant with FAD ∙– reoxidation (to a red-shifted form FAD RS ) in 100 ns, most of the CO 2 product was converted, most likely into bicarbonate (as inferred from FTIR spectra of the cryotrapped FAD RS intermediate). Calculations indicated that this catalytic transformation involved an active-site water molecule. Cryo-Fourier transform infrared spectroscopy studies suggested that bicarbonate formation (step 4) was preceded by deprotonation of an arginine residue (step 3). At room temperature, the remaining CO 2 left the protein in 1.5 μs (step 4'). The observation of residual electron density close to C432 in electron density maps derived from time-resolved and cryocrystallography data suggests that this residue may play a role in stabilizing CO 2 and/or bicarbonate. Three routes for alkane formation were identified by quantum chemistry calculations; the one shown in the figure is favored by the ensemble of experimental data. CONCLUSION Finally, we provide a detailed and comprehensive characterization of light-driven hydrocarbon formation by FAP, which uses a remarkably complex mechanism including unique catalytic steps. We anticipate that our results will help to expand the green chemistry toolkit.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Structural characterization and regulatory element analysis of the heart isoform of cytochrome c oxidase VIa

In order to investigate the mechanism(s) governing the striated muscle-specific expression of cytochrome c oxidase VIaH we have characterized the murine gene and analyzed its transcriptional regulatory elements in skeletal myogenic cell lines. The gene is single copy, spans 689 base pairs (bp), and is comprised of three exons. The 5'-ends of transcripts from the gene are heterogeneous, but the most abundant transcript includes a 5'-untranslated region of 30 nucleotides. When fused to the luciferase reporter gene, the 3.5-kilobase 5'-flanking region of the gene directed the expression of the heterologous protein selectively in differentiated Sol8 cells and transgenic mice, recapitulating the pattern of expression of the endogenous gene. Deletion analysis identified a 300-bp fragment sufficient to direct the myotube-specific expression of luciferase in Sol8 cells. The region lacks an apparent TATA element, and sequence motifs predicted to bind NRF-1, NRF-2, ox-box, or PPAR factors known to regulate other nuclear genes encoding mitochondrial proteins are not evident. Mutational analysis, however, identified two cis-elements necessary for the high level expression of the reporter protein: a MEF2 consensus element at -90 to -81 bp and an E-box element at -147 to -142 bp. Additional E-box motifs at closely located positions were mutated without loss of transcriptional activity. The dependence of transcriptional activation of cytochrome c oxidase VIaH on cis-elements similar to those found in contractile protein genes suggests that the striated muscle-specific expression is coregulated by mechanisms that control the lineage-specific expression of several contractile and cytosolic proteins.

Non-NASA Center↗

Prediction of protein assemblies, the next frontier: The CASP14-CAPRI experiment

Here we present the results for CAPRI Round 50, the fourth joint CASP-CAPRI protein assembly prediction challenge. The Round comprised a total of twelve targets, including six dimers, three trimers, and three higher-order oligomers. Four of these were easy targets, for which good structural templates were available either for the full assembly, or for the main interfaces (of the higher-order oligomers). Eight were difficult targets for which only distantly related templates were found for the individual subunits. Twenty-five CAPRI groups including eight automatic servers submitted ~1250 models per target. Twenty groups including six servers participated in the CAPRI scoring challenge submitted ~190 models per target. The accuracy of the predicted models was evaluated using the classical CAPRI criteria. The prediction performance was measured by a weighted scoring scheme that takes into account the number of models of acceptable quality or higher submitted by each group as part of their five top-ranking models. Compared to the previous CASP-CAPRI challenge, top performing groups submitted such models for a larger fraction (70–75%) of the targets in this Round, but fewer of these models were of high accuracy. Scorer groups achieved stronger performance with more groups submitting correct models for 70–80% of the targets or achieving high accuracy predictions. Servers performed less well in general, except for the MDOCKPP and LZERD servers, who performed on par with human groups. In addition to these results, major advances in methodology are discussed, providing an informative overview of where the prediction of protein assemblies currently stands.

protein complexes↗

Trends in lignin modification: a comprehensive analysis of the effects of genetic manipulations/mutations on lignification and vascular integrity

A comprehensive assessment of lignin configuration in transgenic and mutant plants is long overdue. This review thus undertook the systematic analysis of trends manifested through genetic and mutational manipulations of the various steps associated with monolignol biosynthesis; this included consideration of the downstream effects on organized lignin assembly in the various cell types, on vascular function/integrity, and on plant growth and development. As previously noted for dirigent protein (homologs), distinct and sophisticated monolignol forming metabolic networks were operative in various cell types, tissues and organs, and form the cell-specific guaiacyl (G) and guaiacyl-syringyl (G-S) enriched lignin biopolymers, respectively. Regardless of cell type undergoing lignification, carbon allocation to the different monolignol pools is apparently determined by a combination of phenylalanine availability and cinnamate-4-hydroxylase/"p-coumarate-3-hydroxylase" (C4H/C3H) activities, as revealed by transcriptional and metabolic profiling. Downregulation of either phenylalanine ammonia lyase or cinnamate-4-hydroxylase thus predictably results in reduced lignin levels and impaired vascular integrity, as well as affecting related (phenylpropanoid-dependent) metabolism. Depletion of C3H activity also results in reduced lignin deposition, albeit with the latter being derived only from hydroxyphenyl (H) units, due to both the guaiacyl (G) and syringyl (S) pathways being blocked. Apparently the cells affected are unable to compensate for reduced G/S levels by increasing the amounts of H-components. The downstream metabolic networks for G-lignin enriched formation in both angiosperms and gymnosperms utilize specific cinnamoyl CoA O-methyltransferase (CCOMT), 4-coumarate:CoA ligase (4CL), cinnamoyl CoA reductase (CCR) and cinnamyl alcohol dehydrogenase (CAD) isoforms: however, these steps neither affect carbon allocation nor H/G designations, this being determined by C4H/C3H activities. Such enzymes thus fulfill subsidiary processing roles, with all (except CCOMT) apparently being bifunctional for both H and G substrates. Their severe downregulation does, however, predictably result in impaired monolignol biosynthesis, reduced lignin deposition/vascular integrity, (upstream) metabolite build-up and/or shunt pathway metabolism. There was no evidence for an alternative acid/ester O-methyltransferase (AEOMT) being involved in lignin biosynthesis.The G/S lignin pathway networks are operative in specific cell types in angiosperms and employ two additional biosynthetic steps to afford the corresponding S components, i.e. through introduction of an hydroxyl group at C-5 and its subsequent O-methylation. [These enzymes were originally classified as ferulate-5-hydroxylase (F5H) and caffeate O-methyltransferase (COMT), respectively.] As before, neither step has apparently any role in carbon allocation to the pathway; hence their individual downregulation/manipulation, respectively, gives either a G enriched lignin or formation of the well-known S-deficient bm3 "lignin" mutant, with cell walls of impaired vascular integrity. In the latter case, COMT downregulation/mutation apparently results in utilization of the isoelectronic 5-hydroxyconiferyl alcohol species albeit in an unsuccessful attempt to form G-S lignin proper. However, there is apparently no effect on overall G content, thereby indicating that deposition of both G and S moieties in the G/S lignin forming cells are kept spatially, and presumably temporally, fully separate. Downregulation/mutation of further downstream steps in the G/S network [i.e. utilizing 4CL, CCR and CAD isoforms] gives predictable effects in terms of their subsidiary processing roles: while severe downregulation of 4CL gave phenotypes with impaired vascular integrity due to reduced monolignol supply, there was no evidence in support of increased growth and/or enhanced cellulose biosynthesis. CCR and CAD downregulation/mutations also established that a depletion in monolignol supply reduced both lignin contents supply reduced both lignin contents and vascular integrity, with a concomitant shift towards (upstream) metabolite build-up and/or shunting.The extraordinary claims of involvement of surrogate monomers (2-methoxybenzaldehyde, feruloyl tyramine, vanillic acid, etc.) in lignification were fully disproven and put to rest, with the investigators themselves having largely retracted former claims. Furthermore analysis of the well-known bm1 mutation, a presumed CAD disrupted system, apparently revealed that both G and S lignin components were reduced. This seems to imply that there is no monolignol specific dehydrogenase, such as the recently described sinapyl alcohol dehydrogenase (SAD) for sinapyl alcohol formation. Nevertheless, different CAD isoforms of differing homology seem to be operative in different lignifying cell types, thereby giving the G-enriched and G/S-enriched lignin biopolymers, respectively. For the G-lignin forming network, however, the CAD isoform is apparently catalytically less efficient with all three monolignols than that additionally associated with the corresponding G/S lignin forming network(s), which can more efficiently use all three monolignols. However, since CAD does not determine either H, G, or S designation, it again serves in a subsidiary role-albeit using different isoforms for different cell wall developmental and cell wall type responses.The results from this analysis contrasts further with speculations of some early investigators, who had viewed lignin assembly as resulting from non-specific oxidative coupling of monolignols and subsequent random polymerization. At that time, though, the study of the complex biological (biochemical) process of lignin assembly had begun without any of the (bio)chemical tools to either address or answer the questions posed as to how its formation might actually occur. Today, by contrast, there is growing recognition of both sophisticated and differential control of monolignol biosynthetic networks in different cell types, which serve to underscore the fact that complexity of assembly need not be confused any further with random formation. Moreover, this analysis revealed another factor which continues to cloud interpretations of lignin downregulation/mutational analyses, namely the serious technical problems associated with all aspects of lignin characterization, whether for lignin quantification, isolation of lignin-enriched preparations and/or in determining monomeric compositions. For example, in the latter analyses, some 50-90% of the lignin components still cannot be detected using current methodologies, e.g. by thioacidolysis cleavage and nitrobenzene oxidative cleavage. This deficiency in lignin characterization thus represents one of the major hurdles remaining in delineating how lignin assembly (in distinct cell types) and their configuration actually occurs.

Review, Academic↗