PDB Statistics: Growth in Number of Unique Protein Sequences in Released PDB Structures (Cumulative) at Identity 95%

This chart shows the annual and cumulative numbers of protein sequences in released PDB structures. The chart can be viewed for a few different levels of sequence identity since the beginning of the PDB archive. The cumulative bars represent the growth in unique protein sequences (number of polymeric entities) across history. The yearly bars (dark blue) tell how many new protein sequences were added in a certain year.

Note: The total number of sequence clusters in the statistics table is linked to the sequence cluster group search result page. There is a default precision threshold in calculating the numbers for performance balance. So the statistics count may have a slight discrepancy compared to the actual non-redundant group search result when the result count approaches or goes above 10,000. The group search result page provides an accurate count. The statistics page provides the trend.

Chart is currently loading

Sequence cluster level:

YearNumber of New Protein SequencesTotal Number of Protein Sequences
19761313
19771023
1978326
1979632
1980436
19811046
19821864
19831175
19841186
19851298
19869107
198711118
198825143
198946189
199052241
199157298
199266364
1993232596
19944621,058
19953451,403
19964061,809
19975632,372
19987573,129
19998954,024
200010035,027
200110466,073
200211097,182
200315578,739
2004211610,855
2005233813,193
2006263915,832
2007296318,795
2008275621,551
2009281524,366
2010287227,238
2011263229,870
2012289132,761
2013308535,846
2014380939,655
2015313942,794
2016372346,517
2017395950,476
2018376754,243
2019401358,256
2020496963,225
2021457267,797
2022550473,301
2023521178,512
2024555084,062
2025482288,884