Phát hiện hiệu quả và tự động mối quan hệ cấu trúc quy mô lớn trong protein bằng cách sử dụng bộ căn chỉnh linh hoạt

BMC Bioinformatics - Tập 17 - Trang 1-13 - 2016
Fernando I. Gutiérrez1,2, Felipe Rodriguez-Valenzuela1, Ignacio L. Ibarra1,3, Damien P. Devos2,3, Francisco Melo1
1Departamento de Genética Molecular y Microbiología, Facultad de Ciencias Biológicas, Pontificia Universidad Católica de Chile, Santiago, Chile
2Centre for Organismal Studies (COS), Heidelberg University, Heidelberg, Germany
3Centro Andaluz de Biología del Desarrollo (CABD), Universidad Pablo de Olavide, Sevilla, Spain

Tóm tắt

Số lượng cấu trúc protein ba chiều đã biết đang gia tăng nhanh chóng. Do đó, nhu cầu tìm kiếm cấu trúc nhanh chóng trên các cơ sở dữ liệu đầy đủ mà không làm mất đi độ chính xác đáng kể đang ngày càng trở nên cấp thiết. Gần đây, TopSearch, một phương pháp siêu nhanh để tìm các mối quan hệ cấu trúc vững chắc giữa cấu trúc truy vấn và cơ sở dữ liệu Protein Data Bank (PDB) hoàn chỉnh, ở cấp độ nhiều chuỗi, đã được phát hành. Tuy nhiên, hiện chưa có các bộ căn chỉnh cấu trúc linh hoạt chính xác tương đương để thực hiện các tìm kiếm hiệu quả trên toàn bộ cơ sở dữ liệu của các protein nhiều miền. Việc có sẵn một công cụ như vậy là rất quan trọng cho việc thúc đẩy phát hiện sinh học một cách bền vững. Trong bài báo này, chúng tôi báo cáo về sự phát triển của một phương pháp mới cho việc so sánh nhanh chóng và linh hoạt các chuỗi cấu trúc protein. Phương pháp này dựa trên việc tính toán các ma trận 2D chứa mô tả về sự sắp xếp ba chiều của các yếu tố cấu trúc thứ cấp (các góc và khoảng cách). Việc so sánh liên quan đến việc khớp một tập hợp các tiểu cấu trúc thông qua một thuật toán lập trình động hai bước lồng ghép. Các đặc điểm độc đáo của phương pháp mới này là việc tích hợp và cân bằng những yếu tố sau: 1) tốc độ, 2) độ chính xác và 3) căn chỉnh cấu trúc linh hoạt toàn cục và bán toàn cục bằng cách tích hợp khớp tiểu cấu trúc địa phương. Việc so sánh và khớp với độ chính xác cạnh tranh của một cấu trúc truy vấn có kích thước vừa phải (250-aa) với cơ sở dữ liệu PDB hoàn chỉnh (216,322 chuỗi protein) mất khoảng 8 phút khi sử dụng máy tính để bàn trung bình. Phương pháp này nhanh hơn ít nhất 2-3 bậc so với các công cụ khác đã được thử nghiệm với độ chính xác tương tự. Chúng tôi xác nhận hiệu suất của phương pháp này cho việc phân loại kiểu gập và siêu gia đình trong một bộ điểm chuẩn lớn của các cấu trúc protein. Cuối cùng, chúng tôi cung cấp một loạt các ví dụ để minh họa tính hữu ích của phương pháp này và ứng dụng của nó trong việc phát hiện sinh học. Phương pháp này có khả năng phát hiện khớp cấu trúc một phần, dịch chuyển thể rắn, thay đổi hình dạng và chịu đựng sự biến đổi cấu trúc đáng kể do các sự chèn, xóa và sự khác biệt chuỗi gây ra, cũng như sự hội tụ cấu trúc của các protein không liên quan.

Từ khóa

#protein #cấu trúc ba chiều #căn chỉnh linh hoạt #Protein Data Bank #phát hiện sinh học #phương pháp so sánh protein

Tài liệu tham khảo

Erickson HP. Atomic structures of tubulin and FtsZ. Trends Cell Biol. 1998;8(4):133–7. van den Ent F, Amos LA, LoÈwe J. Prokaryotic origin of the actin cytoskeleton. Nature. 2001;413(6851):39–44. Hasegawa H, Holm L. Advances and pitfalls of protein structural alignment. Curr Opin Struct Biol. 2009;19(3):341–8. Holm L, Sander C. Dali: a network tool for protein structure comparison. Trends Biochem Sci. 1995;20(11):478–80. Gerstein M, Levitt M. Using iterative dynamic programming to obtain accurate pairwise and multiple alignments of protein structures. Proc Int Conf Intell Syst Mol Biol. 1996;4:59–67. Sippl MJ, Wiederstein M. Detection of spatial correlations in protein structures and molecular complexes. Structure (London, England : 1993). 2012;20(4):718–28. Ortiz AR, Strauss CEM, Olmea O. MAMMOTH (matching molecular models obtained from theory): an automated method for model comparison. Protein Sci. 2002;11(11):2606–21. Shindyalov IN, Bourne PE. Protein structure alignment by incremental combinatorial extension (CE) of the optimal path. Protein Eng. 1998;11(9):739–47. Konagurthu AS, Whisstock JC, Stuckey PJ, Lesk AM. MUSTANG: a multiple structural alignment algorithm. Proteins: Struct, Funct, Bioinf. 2006;64(3):559–74. Ye Y, Godzik A. Flexible structure alignment by chaining aligned fragment pairs allowing twists. Bioinformatics. 2003;19 suppl 2:ii246–55. Zhang Y, Skolnick J. TM-align: a protein structure alignment algorithm based on the TM-score. Nucleic Acids Res. 2005;33(7):2302–9. Gibrat J-F, Madej T, Bryant SH. Surprising similarities in structure comparison. Curr Opin Struct Biol. 1996;6(3):377–85. Orengo CA, Taylor WR. SSAP: sequential structure alignment program for protein structure comparison. Computer methods for macromolecular sequence analysis. 1996. Guerler A, Knapp EW. Novel protein folds and their nonsequential structural analogs. Protein Sci. 2008;17(8):1374–82. Stivala A, Wirth A, Stuckey PJ. Tableau-based protein substructure search using quadratic programming. BMC bioinformatics. 2009;10:153. Schwede T, Peitsch MC. Computational structural biology: Methods and applications. 1st ed. Singapore: World Scientific; 2008. Wiederstein M, Gruber M, Frank K, Melo F, Sippl MJ. Structure-based characterization of multiprotein complexes. Structure. 2014;22(7):1063–70. Brohawn SG, Leksa NC, Spear ED, Rajashankar KR, Schwartz TU. Structural evidence for common ancestry of the nuclear pore complex and vesicle coats. Science. 2008;322(5906):1369–73. Lesk AM. Systematic representation of protein folding patterns. J Mol Graph. 1995;13(3):159–64. Konagurthu AS, Stuckey PJ, Lesk AM. Structural search and retrieval using a tableau representation of protein folding patterns. Bioinformatics (Oxford, England). 2008;24(5):645–51. Konagurthu AS, Lesk AM. Structure description and identification using the tableau representation of protein folding patterns. Methods in molecular biology (Clifton, NJ). 2013;932:51–9. Kamat AP, Lesk AM. Contact patterns between helices and strands of sheet define protein folding patterns. Proteins. 2007;66(4):869–76. Murzin AG, Brenner SE, Hubbard T, Chothia C. SCOP: a structural classification of proteins database for the investigation of sequences and structures. J Mol Biol. 1995;247(4):536–40. Chen K, Ruan J, Kurgan L. Prediction of three dimensional structure of calmodulin. Protein J. 2006;25(1):57–70. Shatsky M, Nussinov R, Wolfson HJ. Flexible protein alignment and hinge detection. Proteins: Struct, Funct, Bioinf. 2002;48(2):242–56. Devos D, Dokudovskaya S, Alber F, Williams R, Chait BT, Sali A, et al. Components of coated vesicles and nuclear pore complexes share a common molecular architecture. PLoS Biol. 2004;2(12):e380. Field MC, Sali A, Rout MP. Evolution: On a bender--BARs, ESCRTs, COPs, and finally getting your coat. J Cell Biol. 2011;193(6):963–72. Frishman D, Argos P. Knowledge‐based protein secondary structure assignment. Proteins: Struct, Funct, Bioinf. 1995;23(4):566–79. Mizuguchi K, Deane CM, Blundell TL, Overington JP. HOMSTRAD: a database of protein structure alignments for homologous families. Protein Sci. 1998;7(11):2469–71. Slater AW, Castellanos JI, Sippl MJ, Melo F. Towards the development of standardized methods for comparison, ranking and evaluation of structure alignments. Bioinformatics (Oxford, England). 2013;29(1):47–53. Fox NK, Brenner SE, Chandonia JM. SCOPe: Structural Classification of Proteins--extended, integrating SCOP and ASTRAL data and classification of new structures. Nucleic Acids Res. 2014;42(Database issue):D304–309. Kabsch W, Sander C. Dictionary of protein secondary structure: pattern recognition of hydrogen-bonded and geometrical features. Biopolymers. 1983;22(12):2577–637. Jung J, Lee B. Protein structure alignment using environmental profiles. Protein Eng. 2000;13(8):535–43. Carpentier M, Brouillet S, Pothier J. YAKUSA: a fast structural database scanning method. Proteins. 2005;61(1):137–51. Kolodny R, Koehl P, Levitt M. Comprehensive evaluation of protein structure alignment methods: scoring by geometric measures. J Mol Biol. 2005;346(4):1173–88. Wall ME, Rechtsteiner A, Rocha LM. Singular value decomposition and principal component analysis. In: A practical approach to microarray data analysis. Springer. 2003: 91–109. Sung W-K. Algorithms in bioinformatics: A practical introduction: CRC Press; 2009. Broken Sound Parkway, NW Suite 300, Boca Raton, FL, 33487. USA. Smith TF, Waterman MS. Identification of common molecular subsequences. J Mol Biol. 1981;147(1):195–7. Sippl MJ. On distance and similarity in fold space. Bioinformatics (Oxford, England). 2008;24(6):872–3. Kabsch W. A discussion of the solution for the best rotation to relate two sets of vectors. Acta Crystallogr A. 1978;34(5):827–8. Kearsley SK. On the orthogonal transformation used for structural comparisons. Acta Crystallogr A. 1989;45(2):208–10. Wolda H. Similarity indices, sample size and diversity. Oecologia. 1981;50(3):296–302. Fawcett T. ROC graphs: Notes and practical considerations for researchers. Mach Learn. 2004;31:1–38. Vergara IA, Norambuena T, Ferrada E, Slater AW, Melo F. StAR: a simple tool for the statistical comparison of ROC curves. BMC bioinformatics. 2008;9:265.