Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “protein domain boundary”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Multi-head attention-based U-Nets for predicting protein domain boundaries using 1D sequence features and 2D distance maps

Abstract The information about the domain architecture of proteins is useful for studying protein structure and function. However, accurate prediction of protein domain boundaries (i.e., sequence regions separating two domains) from sequence remains a significant challenge. In this work, we develop a deep learning method based on multi-head U-Nets (called DistDom) to predict protein domain boundaries utilizing 1D sequence features and predicted 2D inter-residue distance map as input. The 1D features contain the evolutionary and physicochemical information of protein sequences, whereas the 2D distance map includes the structural information of proteins that was rarely used in domain boundary prediction before. The 1D and 2D features are processed by the 1D and 2D U-Nets respectively to generate hidden features. The hidden features are then used by the multi-head attention to predict the probability of each residue of a protein being in a domain boundary, leveraging both local and global information in the features. The residue-level domain boundary predictions can be used to classify proteins as single-domain or multi-domain proteins. It classifies the CASP14 single-domain and multi-domain targets at the accuracy of 75.9%, 13.28% more accurate than the state-of-the-art method. Tested on the CASP14 multi-domain protein targets with expert annotated domain boundaries, the average per-target F1 measure score of the domain boundary prediction by DistDom is 0.263, 29.56% higher than the state-of-the-art method.

59 BASIC BIOLOGICAL SCIENCES↗

Modeling SARS-CoV-2 proteins in the CASP-commons experiment

Critical Assessment of Structure Prediction (CASP) is an organization aimed at advancing the state of the art in computing protein structure from sequence. In the spring of 2020, CASP launched a community project to compute the structures of the most structurally challenging proteins coded for in the SARS-CoV-2 genome. Forty-seven research groups submitted over 3000 three-dimensional models and 700 sets of accuracy estimates on 10 proteins. The resulting models were released to the public. CASP community members also worked together to provide estimates of local and global accuracy and identify structure-based domain boundaries for some proteins. Subsequently, two of these structures (ORF3a and ORF8) have been solved experimentally, allowing assessment of both model quality and the accuracy estimates. Models from the AlphaFold2 group were found to have good agreement with the experimental structures, with main chain GDT_TS accuracy scores ranging from 63 (a correct topology) to 87 (competitive with experiment).

59 BASIC BIOLOGICAL SCIENCES↗

Amantadine interactions with phase separated lipid membranes

Amantadine, a small amphilphic organic compound that consists of an adamantane backbone and an amino group, was first recognized as an antiviral in 1963 and received approval for prophylaxis against the type A influenza virus in 1976. Since then, it has also been used to treat Parkinson’s disease-related dyskinesia and is being considered as a treatment for corona viruses. Since amantadine usually targets membrane-bound proteins, its interactions with the membrane are also thought to be important. Biological membranes are now widely understood to be laterally heterogeneous and certain proteins are known to preferentially co-localize within specific lipid domains. Does amantadine, therefore, preferentially localize in certain lipid composition domains? To address this question, here, we studied amantadine’s interactions with phase separating membranes composed of cholesterol, DSPC (1,2-distearoyl-sn-glycero-3-phosphocholine), POPC (1-palmitoyl-2-oleoyl-glycero-3-phosphocholine), and DOPC (1,2-dioleoyl-sn-glycero-3-phosphocholine), as well as single-phase DPhPC (1,2-diphytanoyl-sn-glycero-3-phos-phocholine) membranes. From Langmuir trough and differential scanning calorimetry (DSC) measurements, we determined, respectively, that amantadine preferentially binds to disordered lipids, such as POPC, and lowers the phase transition temperature of POPC/DSPC/cholesterol mixtures, implying that amantadine increases membrane disorder. Further, using droplet interface bilayers (DIBs), we observed that amantadine disrupts DPhPC membranes, consistent with its disordering properties. Finally, we carried out molecular dynamics (MD) simulations on POPC/DSPC/cholesterol membranes with varying amounts of amantadine. Consistent with experiment, MD simulations showed that amantadine prefers to associate with disordered POPC-rich domains, domain boundaries, and lipid glycerol backbones. Since different proteins co-localize with different lipid domains, our results have possible implications as to which classes of proteins may be better targets for amantadine.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Structure-aware annotation of leucine-rich repeat domains

Protein domain annotation is typically done by predictive models such as HMMs trained on sequence motifs. However, sequence-based annotation methods are prone to error, particularly in calling domain boundaries and motifs within them. These methods are limited by a lack of structural information accessible to the model. With the advent of deep learning-based protein structure prediction, existing sequenced-based domain annotation methods can be improved by taking into account the geometry of protein structures. We develop dimensionality reduction methods to annotate repeat units of the Leucine Rich Repeat solenoid domain. The methods are able to correct mistakes made by existing machine learning-based annotation tools and enable the automated detection of hairpin loops and structural anomalies in the solenoid. The methods are applied to 127 predicted structures of LRR-containing intracellular innate immune proteins in the model plant Arabidopsis thaliana and validated against a benchmark dataset of 172 manually-annotated LRR domains.

Xu, Boyan↗

Biosensor Guided Polyketide Synthases Engineering for Optimization of Domain Exchange Boundaries

Type I modular polyketide synthases (PKSs) are multi-domain enzymes functioning like assembly lines. Many engineering attempts have been made for the last three decades to replace, delete and insert new functional domains into PKSs to produce novel molecules. However, inserting heterologous domains often destabilize PKSs, causing loss of activity and protein misfolding. To address this challenge, here we develop a fluorescence-based solubility biosensor that can quickly identify engineered PKSs variants with minimal structural disruptions. Using this biosensor, we screen a library of acyltransferase (AT)-exchanged PKS hybrids with randomly assigned domain boundaries, and we identify variants that maintain wild type production levels. We then probe each position in the AT linker region to determine how domain boundaries influence structural integrity and identify a set of optimized domain boundaries. Overall, we have successfully developed an experimentally validated, high-throughput method for making hybrid PKSs that produce novel molecules.

59 BASIC BIOLOGICAL SCIENCES↗

High-Throughput Directed Evolution of Marine Microalgae and Phototrophic Consortia for Improved Biomass Yields (Final Report)

Primary project achievements include using selective pressures (O 2 , light, temperature) and developing culturing regimes for the diatom Nitzschia inconspicua str. hildebrandi to attain enrichments with an ~90% increase in areal biomass productivity relative to the parental strain under pond-mimicking conditions with high O 2 stress in laboratory bioreactors. The resulting strain (GAI-337) was tested further for dilution time, culture density, CO 2 supplementation, pH, temperature, and dissolved O 2 concentration under outdoor pond-mimicking conditions to improve areal productivities. These experiments yielded an optimum harvest and dilution time just after sunset, ~0.45 g AFDW L -1 initial culture density for maximal productivities, no requirement for CO 2 supplementation or pH control, maximal performance under a diel temperature curve going from 24 °C at night to 36 °C during the day, and benefits from some O 2 removal from the culture by bubbling with air. Using pond-mimicking laboratory bioreactors, N. inconspicua GAI-337 achieved ~42 g AFDW m -2 d -1 . Nutrient limitation experiments resulted in a biomass composition that equated to ~160 Gallons of Gasoline Equivalent energy per ton AFDW, highlighting the potential of GAI-337 as a promising renewable fuel feedstock strain. Genome resequencing has revealed genome alterations potentially contributing to the improved growth of GAI-337 in the laboratory. Based on the comparative analyses of the GAI-337 and GAI-229 (reference) strains, we identified 144 single nucleotide substitutions that resulted in amino acid change, 7 single nucleotide substitutions that resulted in protein truncation; 5 deletions; and 1 frameshift mutation. From the mutations that potentially affect expression of functionally annotated genes, particular interest was noted for an interferon-induced 6-16 family protein that may be involved in the host immune response against microbe invasion; the chaperone protein DnaK, which may function to protect the folding of proteins within the cell; and SPRY domain protein that is found in many eukaryotic proteins important in cell signaling pathways. Transcriptome analysis revealed over 1000 genes with increased transcript levels. Many of these and many of the genes with mutations are not yet functionally annotated and an increased bioinformatics effort is necessary to more completely analyze the Nitzschia inconspicua genome. Adaptive laboratory evolution (ALE) was performed for over 300 days using consecutive 0.5°C temperature increases in a constant temperature incubator to attain greater thermal tolerance in Nitzschia inconspicua. The adapted strain was able to grow at a constant temperature of 37.5°C; whereas this constant temperature was lethal to the parental control, which had an upper temperature boundary of 35.5°C prior to adaptive evolution. Several high-temperature clonal isolates were obtained from the evolved population following ALE, and increased temperature tolerance was observed in clonal adapted cultures. The final temperature adaptation was maintained through cryopreservation and was observed in multiple clonal isolates, including multiple clonal isolates with significantly increased cell size, indicating the potential occurrence of a sexual cycle during the clonal isolation process. A survey of Nannochloropsis strains was conducted for tolerances to high pH and high bicarbonate media. Nannochloropsis granulata showed promising growth in diel bioreactors and was successfully grown at the GAI Kauai farm site in long-term growth campaigns. Co-culturing using Nitzschia inconspicua, Nannochloropsis and a cyanobacterium were assembled in the laboratory to determine if productivity synergies could be attained. Although all strains grew well in the laboratory high-bicarbonate media individually, the cyanobacterium quickly outgrew the other strains in the laboratory consortium pushing the co-culture away from a diverse (and potentially synergistic assemblage) phototroph culture towards a monoculture dominated by the cyanobacterium. Several outdoor growth campaigns were conducted, with productivities ranging between 10-20 g/m 2 /d of biomass. The best performing strain in the laboratory (GAI-337) did not outperform reference strains at the Kauai farm under the conditions used. Addition growth campaigns are necessary under conditions that result in higher biomass (>20 g/m 2 /d) and that attain higher O 2 levels are likely necessary. Initial data indicate that the thermally adapted strain did slightly better than the control strain at higher temperatures; however, additional campaigns are necessary to establish statistical significance. In summary, Nitzschia inconspicua is able to attain exemplary biomass and lipid yields in the laboratory bioreactors. Strain evolution to both O 2 and temperature resulted in targeted strain improvements. Additional outdoor campaigns are necessary to determine if laboratory improvements translate to the field.

09 BIOMASS FUELS↗

A multi-step nucleation process determines the kinetics of prion-like domain phase separation

Compartmentalization by liquid-liquid phase separation (LLPS) has emerged as a ubiquitous mechanism underlying the organization of biomolecules in space and time. Here, we combine rapid-mixing time-resolved small-angle X-ray scattering (SAXS) approaches to characterize the assembly kinetics of a prototypical prion-like domain with equilibrium techniques that characterize its phase boundaries and the size distribution of clusters prior to phase separation. We find two kinetic regimes on the micro- to millisecond timescale that are distinguished by the size distribution of clusters. At the nanoscale, small complexes are formed with low affinity. After initial unfavorable complex assembly, additional monomers are added with higher affinity. At the mesoscale, assembly resembles classical homogeneous nucleation. Careful multi-pronged characterization is required for the understanding of condensate assembly mechanisms and will promote understanding of how the kinetics of biological phase separation is encoded in biomolecules.

59 BASIC BIOLOGICAL SCIENCES↗

Tailoring Hierarchical Structure and Rare Earth Affinity of Compositionally Identical Polymers via Sequence Control

Macromolecule sequence, structure, and function are inherently intertwined. While well-established relationships exist in proteins, they are more challenging to define for synthetic polymer nanoparticles due to their molecular weight, sequence, and conformational dispersities. Furthermore, to explore the impact of sequence on nanoparticle structure, we synthesized a set of 16 compositionally identical, sequence-controlled polymers with distinct monomer patterning of dimethyl acrylamide and a bioinspired, structure-driving di(phenylalanine) acrylamide (FF). Sequence control was achieved through multiblock polymerizations, yielding unique ensembles of polymer sequences which were simulated by kinetic Monte Carlo simulations. Systematic analysis of the global (tertiary- and quaternary-like) structure in this amphiphilic copolymer series revealed the effect of multiple sequence descriptors: the number of domains, the hydropathy of terminal domains, and the patchiness (density) of FF within a domain, each of which impacted both chain collapse and the distribution of single- and multichain assemblies. Furthermore, both the conformational freedom of chain segments and local-scale, β-sheet-like interactions were sensitive to the patchiness of FF. To connect sequence, structure, and target function, we evaluated an additional series of nine sequence-controlled copolymers as sequestrants for rare earth elements (REEs) by incorporating a functional acrylic acid monomer into select polymer scaffolds. We identified key sequence variables that influence the binding affinity, capacity, and selectivity of the polymers for REEs. Collectively, these results highlight the potential of and boundaries of sequence control via multiblock polymerizations to drive primary sequence ensembles hierarchical structures, and ultimately the functionality of compositionally identical polymeric materials.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗