clusterMaker: một plugin phân cụm đa thuật toán cho Cytoscape

BMC Bioinformatics - Tập 12 - Trang 1-14 - 2011
John H Morris1, Leonard Apeltsin1, Aaron M Newman2, Jan Baumbach3, Tobias Wittkop4, Gang Su5,6, Gary D Bader7,8, Thomas E Ferrin1,9
1Department of Pharmaceutical Chemistry, University of California, San Francisco, San Francisco, USA
2Institute for Stem Cell Biology and Regenerative Medicine, Stanford University School of Medicine, Stanford, USA
3Max Planck Institute for Informatics, Saarbrücken, Germany
4Buck Institute for Age Research, Novato, USA
5Bioinformatics Program, University of Michigan, Ann Arbor, USA
6National Center for Integrative Biomedical Informatics, University of Michigan, Ann Arbor, USA
7The Donnelly Centre, University of Toronto, Toronto, Canada
8Department of Molecular Genetics, University of Toronto, Ontario, Canada
9Department of Bioengineering and Therapeutic Sciences, University of California San Francisco, San Francisco, USA

Tóm tắt

Trong kỷ nguyên hậu hệ gen, sự gia tăng nhanh chóng của dữ liệu quy mô lớn đòi hỏi các công cụ tính toán có khả năng tích hợp các loại dữ liệu đa dạng và hỗ trợ nhận diện các mẫu có ý nghĩa sinh học trong đó. Ví dụ, các tập dữ liệu tương tác protein-protein đã được phân nhóm để xác định các phức hợp ổn định, nhưng các nhà khoa học vẫn thiếu các công cụ dễ tiếp cận để hỗ trợ phân tích kết hợp của nhiều tập dữ liệu từ các loại thí nghiệm khác nhau. Tại đây, chúng tôi giới thiệu clusterMaker, một plugin Cytoscape thực hiện một số thuật toán phân cụm và cung cấp các chế độ xem mạng lưới, cây phân nhánh và bản đồ nhiệt của các kết quả. Mạng Cytoscape được liên kết với tất cả các chế độ xem khác, vì vậy một lựa chọn ở một chế độ xem sẽ được phản ánh ngay lập tức ở các chế độ xem khác. clusterMaker là plugin Cytoscape đầu tiên thực hiện một loạt các thuật toán và hình ảnh phân cụm đa dạng như vậy, bao gồm cả các triển khai duy nhất của phân cụm phân cấp, hình ảnh cây phân nhánh cộng với bản đồ nhiệt (chế độ xem cây), k-means, k-medoid, SCPS, AutoSOME, và MCL gốc (Java). Các kết quả được trình bày dưới dạng ba kịch bản sử dụng: phân tích dữ liệu biểu hiện protein sử dụng một interactome chuột đã được công bố gần đây và một tập dữ liệu vi mạch chuột của gần một trăm loại tế bào/mô khác nhau; xác định các phức hợp protein trong nấm men Saccharomyces cerevisiae; và phân tích phân cụm của siêu họ enzyme chelate oxy vicinal (VOC). Đối với kịch bản một, chúng tôi khám phá các interactome chuột giàu chức năng đặc trưng cho các kiểu hình tế bào cụ thể và áp dụng phân cụm mờ. Đối với kịch bản hai, chúng tôi khám phá chi tiết phức hợp prefoldin sử dụng cả các cụm tương tác vật lý và di truyền. Đối với kịch bản ba, chúng tôi khám phá khả năng chú thích một protein như là methylmalonyl-CoA epimerase trong siêu họ VOC. Các tệp phiên Cytoscape cho cả ba kịch bản được cung cấp trong phần Tệp bổ sung. Plugin Cytoscape clusterMaker cung cấp một số thuật toán và hình ảnh phân cụm có thể được sử dụng độc lập hoặc kết hợp để phân tích và hình ảnh hóa các tập dữ liệu sinh học, và để xác nhận hoặc tạo ra các giả thuyết về chức năng sinh học. Một số hình ảnh và thuật toán này chỉ có sẵn cho người dùng Cytoscape thông qua plugin clusterMaker. clusterMaker có sẵn thông qua trình quản lý plugin Cytoscape.

Từ khóa

#Cytoscape #phân cụm #thuật toán #sinh học #tương tác protein

Tài liệu tham khảo

Eisen MB, Spellman PT, Brown PO, Botstein D: Cluster analysis and display of genome-wide expression patterns. Proc Natl Acad Sci USA 1998, 95(25):14863–14868. 10.1073/pnas.95.25.14863 Schuldiner M, Collins SR, Thompson NJ, Denic V, Bhamidipati A, Punna T, Ihmels J, Andrews B, Boone C, Greenblatt JF, et al.: Exploration of the function and organization of the yeast early secretory pathway through an epistatic miniarray profile. Cell 2005, 123(3):507–519. 10.1016/j.cell.2005.08.031 Costanzo M, Baryshnikova A, Bellay J, Kim Y, Spear ED, Sevier CS, Ding H, Koh JL, Toufighi K, Mostafavi S, et al.: The genetic landscape of a cell. Science 2010, 327(5964):425–431. 10.1126/science.1180823 Bader GD, Hogue CW: An automated method for finding molecular complexes in large protein interaction networks. BMC Bioinformatics 2003, 4: 2. 10.1186/1471-2105-4-2 King AD, Przulj N, Jurisica I: Protein complex prediction via cost-based clustering. Bioinformatics 2004, 20(17):3013–3020. 10.1093/bioinformatics/bth351 Blatt M, Wiseman S, Domany E: Super-paramagnetic clustering of data. Physical Review Leters 1996., 76: Enright AJ, Van Dongen S, Ouzounis CA: An efficient algorithm for large-scale detection of protein families. Nucleic Acids Res 2002, 30(7):1575–1584. 10.1093/nar/30.7.1575 van Dongen S: Graph Clustering by Flow Simulation. University of Utrecht; 2000. Rives AW, Galitski T: Modular organization of cellular networks. Proceedings of the National Academy of Sciences of the United States of America 2003, 100(3):1128–1133. 10.1073/pnas.0237338100 Wu CH, Nikolskaya A, Huang H, Yeh LS, Natale DA, Vinayaka CR, Hu ZZ, Mazumder R, Kumar S, Kourtesis P, et al.: PIRSF: family classification system at the Protein Information Resource. Nucleic acids research 2004, (32 Database):D112–114. Lee BJ, Shin MS, Oh YJ, Oh HS, Ryu KH: Identification of protein functions using a machine-learning approach based on sequence-derived properties. Proteome Sci 2009, 7: 27. 10.1186/1477-5956-7-27 Qiu JD, Luo SH, Huang JH, Liang RP: Using support vector machines to distinguish enzymes: approached by incorporating wavelet transform. J Theor Biol 2009, 256(4):625–631. 10.1016/j.jtbi.2008.10.026 Zhu F, Han LY, Chen X, Lin HH, Ong S, Xie B, Zhang HL, Chen YZ: Homology-free prediction of functional class of proteins and peptides by support vector machines. Curr Protein Pept Sci 2008, 9(1):70–95. 10.2174/138920308783565697 Kriventseva EV, Biswas M, Apweiler R: Clustering and analysis of protein families. Curr Opin Struct Biol 2001, 11(3):334–339. 10.1016/S0959-440X(00)00211-6 Apweiler R, Biswas M, Fleischmann W, Kanapin A, Karavidopoulou Y, Kersey P, Kriventseva EV, Mittard V, Mulder N, Phan I, et al.: Proteome Analysis Database: online application of InterPro and CluSTr for the functional classification of proteins in whole genomes. Nucleic acids research 2001, 29(1):44–48. 10.1093/nar/29.1.44 Li W, Jaroszewski L, Godzik A: Clustering of highly homologous sequences to reduce the size of large protein databases. Bioinformatics 2001, 17(3):282–283. 10.1093/bioinformatics/17.3.282 Li W, Jaroszewski L, Godzik A: Sequence clustering strategies improve remote homology recognitions while reducing search times. Protein Eng 2002, 15(8):643–649. 10.1093/protein/15.8.643 Yona G, Linial N, Linial M: ProtoMap: automatic classification of protein sequences and hierarchy of protein families. Nucleic acids research 2000, 28(1):49–55. 10.1093/nar/28.1.49 Sasson O, Vaaknin A, Fleischer H, Portugaly E, Bilu Y, Linial N, Linial M: ProtoNet: hierarchical classification of the protein space. Nucleic acids research 2003, 31(1):348–352. 10.1093/nar/gkg096 Kriventseva EV, Fleischmann W, Zdobnov EM, Apweiler R: CluSTr: a database of clusters of SWISS-PROT+TrEMBL proteins. Nucleic acids research 2001, 29(1):33–36. 10.1093/nar/29.1.33 Krause A, Haas SA, Coward E, Vingron M: SYSTERS, GeneNest, SpliceNest: exploring sequence space from genome to protein. Nucleic acids research 2002, 30(1):299–300. 10.1093/nar/30.1.299 Enright AJ, Ouzounis CA: GeneRAGE: a robust algorithm for sequence clustering and domain detection. Bioinformatics 2000, 16(5):451–457. 10.1093/bioinformatics/16.5.451 Abascal F, Valencia A: Clustering of proximal sequence space for the identification of protein families. Bioinformatics 2002, 18(7):908–921. 10.1093/bioinformatics/18.7.908 Nepusz T, Sasidharan R, Paccanaro A: SCPS: a fast implementation of a spectral method for detecting protein families on a genome-wide scale. BMC Bioinformatics 2010, 11: 120. 10.1186/1471-2105-11-120 Wittkop T, Emig D, Lange S, Rahmann S, Albrecht M, Morris JH, Bocker S, Stoye J, Baumbach J: Partitioning biological data with transitivity clustering. Nat Methods 2010, 7(6):419–420. 10.1038/nmeth0610-419 Wittkop T, Baumbach J, Lobo FP, Rahmann S: Large scale clustering of protein sequences with FORCE - A layout based heuristic for weighted cluster editing. BMC Bioinformatics 2007, 8: 396. 10.1186/1471-2105-8-396 Frey BJ, Dueck D: Clustering by passing messages between data points. Science 2007, 315(5814):972–976. 10.1126/science.1136800 van Dongen S: A cluster algorithm for graphs. Amsterdam: National Research Institue in the Netherlands; 2000. Wittkop T, Emig D, Truss A, Albrecht M, Bocker S, Baumbach J: Comprehensive cluster analysis with Transitivity Clustering. Nature protocols 2011, 6(3):285–295. 10.1038/nprot.2010.197 Shannon P, Markiel A, Ozier O, Baliga NS, Wang JT, Ramage D, Amin N, Schwikowski B, Ideker T: Cytoscape: a software environment for integrated models of biomolecular interaction networks. Genome Res 2003, 13(11):2498–2504. 10.1101/gr.1239303 Cline MS, Smoot M, Cerami E, Kuchinsky A, Landys N, Workman C, Christmas R, Avila-Campilo I, Creech M, Gross B, et al.: Integration of biological networks and gene expression data using Cytoscape. Nat Protoc 2007, 2(10):2366–2382. 10.1038/nprot.2007.324 Newman ME, Girvan M: Finding and evaluating community structure in networks. Phys Rev E Stat Nonlin Soft Matter Phys 2004, 69(2 Pt 2):026113. Su G, Kuchinsky A, Morris JH, States DJ, Meng F: GLay: community structure analysis of biological networks. Bioinformatics 2010, 26(24):3135–3137. 10.1093/bioinformatics/btq596 Newman AM, Cooper JB: AutoSOME: a clustering method for identifying gene expression modules without prior knowledge of cluster number. BMC Bioinformatics 2010, 11: 117. 10.1186/1471-2105-11-117 Krogan NJ, Cagney G, Yu H, Zhong G, Guo X, Ignatchenko A, Li J, Pu S, Datta N, Tikuisis AP, et al.: Global landscape of protein complexes in the yeast Saccharomyces cerevisiae. Nature 2006, 440(7084):637–643. 10.1038/nature04670 Vlasblom J, Wodak SJ: Markov clustering versus affinity propagation for the partitioning of protein interaction graphs. BMC Bioinformatics 2009, 10: 99. 10.1186/1471-2105-10-99 Yang F, Zhu Q-X, Tang D-M, Zhao M-Y: Clustering Protein Sequences Using Affinity Propagation Based on an Improved Similarity Measure. Evolutionary Bioinformatics 2010, 2009: 137. 1812-EBO-Clustering-Protein-Sequences-Using-Affinity-Propagation-Based-on-an-Im.pdf 1812-EBO-Clustering-Protein-Sequences-Using-Affinity-Propagation-Based-on-an-Im.pdf Rousseeuw PJ: Silhouettes: A graphical aid to the interpretation and validation of cluster analysis. Journal of Computational and Applied Mathematics 1987, 20: 53–65. 10.1016/0377-0427(87)90125-7 Apeltsin L, Morris JH, Babbitt PC, Ferrin TE: Improving the quality of protein similarity network clustering algorithms using the network edge weight distribution. Bioinformatics 2011, 27(3):326–333. 10.1093/bioinformatics/btq655 Saldanha AJ: Java Treeview--extensible visualization of microarray data. Bioinformatics 2004, 20(17):3246–3248. 10.1093/bioinformatics/bth349 Huang da W, Sherman BT, Lempicki RA: Systematic and integrative analysis of large gene lists using DAVID bioinformatics resources. Nature protocols 2009, 4(1):44–57. Maere S, Heymans K, Kuiper M: BiNGO: a Cytoscape plugin to assess overrepresentation of gene ontology categories in biological networks. Bioinformatics 2005, 21(16):3448–3449. 10.1093/bioinformatics/bti551 Ideker T, Ozier O, Schwikowski B, Siegel AF: Discovering regulatory and signalling circuits in molecular interaction networks. Bioinformatics 2002, 18(Suppl 1):S233–240. 10.1093/bioinformatics/18.suppl_1.S233 Giancarlo R, Scaturro D, Utro F: Computational cluster validation for microarray data analysis: experimental assessment of Clest, Consensus Clustering, Figure of Merit, Gap Statistics and Model Explorer. BMC Bioinformatics 2008, 9: 462. 10.1186/1471-2105-9-462 Li X, Cai H, Xu J, Ying S, Zhang Y: A mouse protein interactome through combined literature mining with multiple sources of interaction evidence. Amino Acids 2010, 38(4):1237–1252. 10.1007/s00726-009-0335-7 Lattin JE, Schroder K, Su AI, Walker JR, Zhang J, Wiltshire T, Saijo K, Glass CK, Hume DA, Kellie S, et al.: Expression analysis of G Protein-Coupled Receptors in mouse macrophages. Immunome Res 2008, 4: 5. 10.1186/1745-7580-4-5 Ito T, Tashiro K, Muta S, Ozawa R, Chiba T, Nishizawa M, Yamamoto K, Kuhara S, Sakaki Y: Toward a protein-protein interaction map of the budding yeast: A comprehensive system to examine two-hybrid interactions in all possible combinations between the yeast proteins. Proc Natl Acad Sci USA 2000, 97(3):1143–1147. 10.1073/pnas.97.3.1143 Uetz P, Giot L, Cagney G, Mansfield TA, Judson RS, Knight JR, Lockshon D, Narayan V, Srinivasan M, Pochart P, et al.: A comprehensive analysis of protein-protein interactions in Saccharomyces cerevisiae. Nature 2000, 403(6770):623–627. 10.1038/35001009 Johnsson N: A split-ubiquitin-based assay detects the influence of mutations on the conformational stability of the p53 DNA binding domain in vivo. FEBS Lett 2002, 531(2):259–264. 10.1016/S0014-5793(02)03533-0 Ho Y, Gruhler A, Heilbut A, Bader GD, Moore L, Adams SL, Millar A, Taylor P, Bennett K, Boutilier K, et al.: Systematic identification of protein complexes in Saccharomyces cerevisiae by mass spectrometry. Nature 2002, 415(6868):180–183. 10.1038/415180a Gavin AC, Aloy P, Grandi P, Krause R, Boesche M, Marzioch M, Rau C, Jensen LJ, Bastuck S, Dumpelfeld B, et al.: Proteome survey reveals modularity of the yeast cell machinery. Nature 2006, 440(7084):631–636. 10.1038/nature04532 Collins SR, Kemmeren P, Zhao XC, Greenblatt JF, Spencer F, Holstege FC, Weissman JS, Krogan NJ: Toward a comprehensive atlas of the physical interactome of Saccharomyces cerevisiae. Mol Cell Proteomics 2007, 6(3):439–450. Collins SR, Miller KM, Maas NL, Roguev A, Fillingham J, Chu CS, Schuldiner M, Gebbia M, Recht J, Shales M, et al.: Functional dissection of protein complexes involved in yeast chromosome biology using a genetic interaction map. Nature 2007, 446(7137):806–810. 10.1038/nature05649 Wilmes GM, Bergkessel M, Bandyopadhyay S, Shales M, Braberg H, Cagney G, Collins SR, Whitworth GB, Kress TL, Weissman JS, et al.: A genetic interaction map of RNA-processing factors reveals links between Sem1/Dss1-containing complexes and mRNA export and splicing. Mol Cell 2008, 32(5):735–746. 10.1016/j.molcel.2008.11.012 Jaroszewski L, Li Z, Krishna SS, Bakolitsa C, Wooley J, Deacon AM, Wilson IA, Godzik A: Exploration of uncharted regions of the protein universe. PLoS biology 2009, 7(9):e1000205. 10.1371/journal.pbio.1000205 Altschul SF, Gish W, Miller W, Myers EW, Lipman DJ: Basic local alignment search tool. Journal of molecular biology 1990, 215(3):403–410. Pegg SC, Brown SD, Ojha S, Seffernick J, Meng EC, Morris JH, Chang PJ, Huang CC, Ferrin TE, Babbitt PC: Leveraging enzyme structure-function relationships for functional inference and experimental design: the structure-function linkage database. Biochemistry 2006, 45(8):2545–2555. 10.1021/bi052101l Armstrong RN: Mechanistic diversity in a metalloenzyme superfamily. Biochemistry 2000, 39(45):13625–13632. 10.1021/bi001814v Babbitt PC: Exploring the VOC superfamily. In Edited by: Apeltsin L. 2011. Saeed AI, Sharov V, White J, Li J, Liang W, Bhagabati N, Braisted J, Klapa M, Currier T, Thiagarajan M, et al.: TM4: a free, open-source system for microarray data management and analysis. Biotechniques 2003, 34(2):374–378. Saeed AI, Bhagabati NK, Braisted JC, Liang W, Sharov V, Howe EA, Li J, Thiagarajan M, White JA, Quackenbush J: TM4 microarray software suite. Methods in enzymology 2006, 411: 134–193. J van der Laan M, Pollard KS: A new algorithm for hybrid hierarchical clustering with visualization and the bootstrap. Journal of Statistical Planning and Inference 2003, 117(2):275–303. 10.1016/S0378-3758(02)00388-9 Heyer LJ, Kruglyak S, Yooseph S: Exploring expression data: identification and analysis of coexpressed genes. Genome research 1999, 9(11):1106–1115. 10.1101/gr.9.11.1106 Bezdek JC: Pattern Recognition with Fuzzy Objective Function Algorithms. Kluwer Academic Publishers; 1981. Pavlopoulos GA, Moschopoulos CN, Hooper SD, Schneider R, Kossida S: jClust: a clustering and visualization toolbox. Bioinformatics 2009, 25(15):1994–1996. 10.1093/bioinformatics/btp330