Engineering Papers⌕ Search

Engineering topics

Kozakov, Dima

Publications and source records attributed to Kozakov, Dima.

Protein folds vs. protein folding: Differing questions, different challenges

We report protein fold prediction using deep-learning artificial intelligence (AI) has transformed the field of protein structure prediction. By combining physical and geometric constraints—and especially patterns extracted from the Protein Data Bank —these machine learning algorithms can predict protein structures at or near atomic resolution and do so in seconds. Today, these computational methods have now solved more than 200 million protein structures, which are accessible from the AlphaFold Protein Structure Database. This accomplishment seems all the more remarkable because few thought it possible or saw it coming. Deservedly, deep-learning AI was named Science magazine’s 2021 “breakthrough of the year”. Clearly, deep-learning AI represents a major advance in protein fold prediction.

54 ENVIRONMENTAL SCIENCES↗

A simple technique to classify diffraction data from dynamic proteins according to individual polymorphs

One often observes small but measurable differences in the diffraction data measured from different crystals of a single protein. These differences might reflect structural differences in the protein and may reveal the natural dynamism of the molecule in solution. Partitioning these mixed-state data into single-state clusters is a critical step that could extract information about the dynamic behavior of proteins from hundreds or thousands of single-crystal data sets. Mixed-state data can be obtained deliberately (through intentional perturbation) or inadvertently (while attempting to measure highly redundant single-crystal data). To the extent that different states adopt different molecular structures, one expects to observe differences in the crystals; each of the polystates will create a polymorph of the crystals. After mixed-state diffraction data have been measured, deliberately or inadvertently, the challenge is to sort the data into clusters that may represent relevant biological polystates. Here, this problem is addressed using a simple multi-factor clustering approach that classifies each data set using independent observables, thereby assigning each data set to the correct location in conformational space. This procedure is illustrated using two independent observables, unit-cell parameters and intensities, to cluster mixed-state data from chymotrypsinogen (ChTg) crystals. It is observed that the data populate an arc of the reaction trajectory as ChTg is converted into chymotrypsin.

36 MATERIALS SCIENCE↗