DeepMind and EMBL-EBI publish the AlphaFold protein database

DeepMind and EMBL’s European Bioinformatics Institute opened the AlphaFold Protein Structure Database on 22 July 2021, with more than 350,000 predicted structures covering the roughly 20,000 proteins of the human proteome and 20 other organisms used in research.

Why it mattered The predictions were free to search from the day they were published, so researchers could use AlphaFold’s output on drug targets and basic biology without running the model themselves.

The CASP14 result eight months earlier had shown that AlphaFold 2 could predict the shape of a protein from its sequence to something near laboratory accuracy. It had not put a single structure into a biologist’s hands. The predictions sat inside DeepMind.

On 22 July 2021 DeepMind and EMBL’s European Bioinformatics Institute opened the AlphaFold Protein Structure Database. It held more than 350,000 predicted structures, searchable and free, covering the roughly 20,000 proteins of the human proteome and twenty other organisms chosen for their use in research, among them E. coli and the parasite that causes malaria. The partners said they intended to extend coverage toward almost every sequenced protein known to science.

The choice of partner mattered as much as the data. EMBL-EBI is a public institute that already ran databases molecular biologists consulted daily, so the predictions arrived in a place researchers were already looking, under conventions they already followed, rather than as a product to be licensed or a file attached to a paper.

Free access from the first day changed who could use the work. A group studying a neglected disease, or one without the computing budget to run the model itself, could look up a structure in the week the database opened and start from there.

What the database holds is still prediction. A computed structure is a hypothesis about a shape, not a measurement of one, and the entries carry no experimental authority of their own. The value is in what it costs to obtain a plausible answer, which fell from months of laboratory work to a database query for a large share of proteins.