Khung tối ưu hóa cho việc xác định không giám sát các biến thể số lượng bản sao hiếm từ dữ liệu SNP array

Genome Biology - Tập 10 - Trang 1-18 - 2009
Gökhan Yavaş1, Mehmet Koyutürk1,2, Meral Özsoyoğlu1, Meetha P Gould3, Thomas LaFramboise2,3,4
1Department of Electrical Engineering and Computer Science, Case Western Reserve University, Cleveland, USA
2Center for Proteomics and Bioinformatics, Case Western Reserve University, Cleveland, USA
3Department of Genetics, Case Western Reserve University, Cleveland, USA
4Genomic Medicine Institute, Lerner Research Institute, Cleveland Clinic Foundation, Cleveland, USA

Tóm tắt

Các biến thể số lượng bản sao (CNVs) có vai trò trong các bệnh lý của con người, và các microarray DNA là công cụ quan trọng để xác định chúng. Trong bài báo này, chúng tôi trình bày việc xác định CNV như một bài toán tối ưu hóa hàm mục tiêu. Chúng tôi áp dụng phương pháp của mình lên dữ liệu từ hàng trăm mẫu và chứng minh khả năng phát hiện CNVs với mức độ nhạy cao mà không hy sinh độ đặc hiệu. Hiệu suất của nó so sánh thuận lợi với các phương pháp hiện có và tiết lộ các tổn thất và lợi ích chưa từng được báo cáo trước đó.

Từ khóa

#biến thể số lượng bản sao #CNVs #microarray DNA #tối ưu hóa #phát hiện CNVs #phát hiện không giám sát

Tài liệu tham khảo

International HapMap Consortium: A haplotype map of the human genome. Nature. 2005, 437: 1299-1320. 10.1038/4371241a. Affymetrix: Genome-Wide Human SNP Array 6.0 Data Sheet. 2007, Santa Clara, California: Affymetrix Illumina: Human1M-duo Beadchip Data Sheet. 2007, San Diego, CA: Illumina Feuk L, Carson AR, Scherer SW: Structural variation in the human genome. Nat Rev Genet. 2006, 7: 85-97. 10.1038/nrg1767. Rovelet-Lecrux A, Hannequin D, Raux G, Le Meur N, Laquerrière A, Vital A, Dumanchin C, Feuillette S, Brice A, Vercelletto M, Dubas F, Frebourg T, Campion D: APP locus duplication causes autosomal dominant early-onset Alzheimer disease with cerebral amyloid angiopathy. Nat Genet. 2006, 38: 24-26. 10.1038/ng1718. Fellermann K, Stange DE, Schaeffeler E, Schmalzl H, Wehkamp J, Bevins CL, Reinisch W, Teml A, Schwab M, Lichter P, Radlwimmer B, Stange EF: A chromosome 8 gene-cluster polymorphism with low human beta-defensin 2 gene copy number predisposes to Crohn disease of the colon. Am J Hum Genet. 2006, 79: 439-448. 10.1086/505915. Sebat J, Lakshmi B, Malhotra D, Troge J, Lese-Martin C, Walsh T, Yamrom B, Yoon S, Krasnitz A, Kendall J, Leotta A, Pai D, Zhang R, Lee YH, Hicks J, Spence SJ, Lee AT, Puura K, Lehtimäki T, Ledbetter D, Gregersen PK, Bregman J, Sutcliffe JS, Jobanputra V, Chung W, Warburton D, King MC, Skuse D, Geschwind DH, Gilliam TC, et al: Strong association of de novo copy number mutations with autism. Science. 2007, 316: 445-449. 10.1126/science.1138659. Xu B, Roos JL, Levy S, van Rensburg EJ, Gogos JA, Karayiorgou M: Strong association of de novo copy number mutations with sporadic schizophrenia. Nat Genet. 2008, 40: 880-885. 10.1038/ng.162. Zhao X, Li C, Paez JG, Chin K, Jänne PA, Chen TH, Girard L, Minna J, Christiani D, Leo C, Gray JW, Sellers WR, Meyerson M: An integrated view of copy number and allelic alterations in the cancer genome using single nucleotide polymorphism arrays. Cancer Res. 2004, 64: 3060-3071. 10.1158/0008-5472.CAN-03-3308. Peiffer DA, Le JM, Steemers FJ, Chang W, Jenniges T, Garcia F, Haden K, Li J, Shaw CA, Belmont J, Cheung SW, Shen RM, Barker DL, Gunderson KL: High-resolution genomic profiling of chromosomal aberrations using Infinium whole-genome genotyping. Genome Res. 2006, 16: 1136-1148. 10.1101/gr.5402306. Gunderson KL, Steemers FJ, Lee G, Mendoza LG, Chee MS: A genome-wide scalable SNP genotyping assay using microarray technology. Nat Genet. 2005, 37: 549-554. 10.1038/ng1547. Lindblad-Toh K, Tanenbaum DM, Daly MJ, Winchester E, Lui WO, Villapakkam A, Stanton SE, Larsson C, Hudson TJ, Johnson BE, Lander ES, Meyerson M: Loss-of-heterozygosity analysis of small-cell lung carcinomas using single-nucleotide polymorphism arrays. Nat Biotechnol. 2000, 18: 1001-1005. 10.1038/79269. Bolstad BM, Irizarry RA, Astrand M, Speed TP: A comparison of normalization methods for high density oligonucleotide array data based on variance and bias. Bioinformatics. 2003, 19: 185-193. 10.1093/bioinformatics/19.2.185. Lin M, Wei LJ, Sellers WR, Lieberfarb M, Wong WH, Li C: dChipSNP: significance curve and clustering of SNP-array-based loss-of-heterozygosity data. Bioinformatics. 2004, 20: 1233-1240. 10.1093/bioinformatics/bth069. LaFramboise T, Harrington D, Weir BA: PLASQ: a generalized linear model-based procedure to determine allelic dosage in cancer cells from SNP array data. Biostatistics. 2007, 8: 323-336. 10.1093/biostatistics/kxl012. Bengtsson H, Irizarry R, Carvalho B, Speed TP: Estimation and assessment of raw copy numbers at the single locus level. Bioinformatics. 2008, 24: 759-767. 10.1093/bioinformatics/btn016. Korn JM, Kuruvilla FG, McCarroll SA, Wysoker A, Nemesh J, Cawley S, Hubbell E, Veitch J, Collins PJ, Darvishi K, Lee C, Nizzari MM, Gabriel SB, Purcell S, Daly MJ, Altshuler D: Integrated genotype calling and association analysis of SNPs, common copy number polymorphisms and rare CNVs. Nat Genet. 2008, 40: 1253-1260. 10.1038/ng.237. Zhao X, Weir BA, LaFramboise T, Lin M, Beroukhim R, Garraway L, Beheshti J, Lee JC, Naoki K, Richards WG, Sugarbaker D, Chen F, Rubin MA, Jänne PA, Girard L, Minna J, Christiani D, Li C, Sellers WR, Meyerson M: Homozygous deletions and chromosome amplifications in human lung carcinomas revealed by single nucleotide polymorphism array analysis. Cancer Res. 2005, 65: 5561-5570. 10.1158/0008-5472.CAN-04-4603. Olshen AB, Venkatraman ES, Lucito R, Wigler M: Circular binary segmentation for the analysis of array-based DNA copy number data. Biostatistics. 2004, 5: 557-572. 10.1093/biostatistics/kxh008. Venkatraman ES, Olshen AB: A faster circular binary segmentation algorithm for the analysis of array CGH data. Bioinformatics. 2007, 23: 657-663. 10.1093/bioinformatics/btl646. Polzehl J, Spokoiny S: Adaptive weights smoothing with applications to image restoration. J R Stat Soc, Ser B. 2000, 62: 335-354. 10.1111/1467-9868.00235. Hupé P, Stransky N, Thiery JP, Radvanyi F, Barillot E: Analysis of array CGH data: from signal ratio to gain and loss of DNA regions. Bioinformatics. 2004, 20: 3413-3422. 10.1093/bioinformatics/bth418. McCarroll SA, Kuruvilla FG, Korn JM, Cawley S, Nemesh J, Wysoker A, Shapero MH, de Bakker PI, Maller JB, Kirby A, Elliott AL, Parkin M, Hubbell E, Webster T, Mei R, Veitch J, Collins PJ, Handsaker R, Lincoln S, Nizzari M, Blume J, Jones KW, Rava R, Daly MJ, Gabriel SB, Altshuler D: Integrated detection and population-genetic analysis of SNPs and copy number variation. Nat Genet. 2008, 40: 1166-1174. 10.1038/ng.238. dChip Software Website. [http://www.dchip.org] Li C, Wong WH: Model-based analysis of oligonucleotide arrays: expression index computation and outlier detection. Proc Natl Acad Sci USA. 2001, 98: 31-36. 10.1073/pnas.011404098. Colella S, Yau C, Taylor JM, Mirza G, Butler H, Clouston P, Bassett AS, Seller A, Holmes CC, Ragoussis J: QuantiSNP: an objective Bayes hidden-Markov model to detect and accurately map copy number variation using SNP genotyping data. Nucleic Acids Res. 2007, 35: 2013-2025. 10.1093/nar/gkm076. Wang K, Li M, Hadley D, Liu R, Glessner J, Grant SF, Hakonarson H, Bucan M: PennCNV: an integrated hidden Markov model designed for high-resolution copy number variation detection in whole-genome SNP genotyping data. Genome Res. 2007, 17: 1665-1674. 10.1101/gr.6861907. Pinto D, Marshall C, Feuk L, Scherer SW: Copy-number variation in control population cohorts. Hum Mol Genet. 2007, 16: R168-173. 10.1093/hmg/ddm241. Active Perl. [http://www.activestate.com/activeperl/] Affymetrix Power Tools Software Package. [http://www.affymetrix.com/partners_programs/programs/developer/tools/powertools.affx] Database of Genomic Variants. [http://projects.tcag.ca/variation/] Schouten JP, McElgunn CJ, Waaijer R, Zwijnenburg D, Diepvens F, Pals G: Relative quantification of 40 nucleic acid sequences by multiplex ligation-dependent probe amplification. Nucleic Acids Res. 2002, 30: e57-10.1093/nar/gnf056. Dempster AP, Laird NM, Rubin DB: Maximum likelihood from incomplete data via the EM algorithm. J Roy Stat Soc. 1977, 39: 1-38. Beroukhim R, Lin M, Park Y, Hao K, Zhao X, Garraway LA, Fox EA, Hochberg EP, Mellinghoff IK, Hofer MD, Descazeaud A, Rubin MA, Meyerson M, Wong WH, Sellers WR, Li C: Inferring loss-of-heterozygosity from unpaired tumors using high-density oligonucleotide SNP arrays. PLoS Comput Biol. 2006, 2: e41-10.1371/journal.pcbi.0020041. Pique-Regi R, Tsau E-S, Ortega A, Seeger R, Asgharzadeh S: Wavelet footprints and sparse bayesian learning for DNA copy number change analysis. Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP): 15-20 April 2007; Honolulu, HI. 2007, IEEE Signal Processing Society Editors: IEEE Publications, Piscataway, NJ, USA, 1: 353-356. Alqallaf A, Tewfik A, Selleck S, Johnson R: Framework for the analysis of genetic variations across multiple DNA copy number samples. Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP): 30 March-4 April 2008; Las Vegas, NV. 2008, IEEE Signal Processing Society Editors: IEEE Publications, Piscataway, NJ, USA, 1: 553-556. Shen F, Huang J, Fitch KR, Truong VB, Kirby A, Chen W, Zhang J, Liu G, McCarroll SA, Jones KW, Shapero MH: Improved detection of global copy number variation using high density, non-polymorphic oligonucleotide probes. BMC Genet. 2008, 9: 27-10.1186/1471-2156-9-27. Wang J, Wang W, Li R, Li Y, Tian G, Goodman L, Fan W, Zhang J, Li J, Zhang J, Guo Y, Feng B, Li H, Lu Y, Fang X, Liang H, Du Z, Li D, Zhao Y, Hu Y, Yang Z, Zheng H, Hellmann I, Inouye M, Pool J, Yi X, Zhao J, Duan J, Zhou Y, Qin J, et al: The diploid genome sequence of an Asian individual. Nature. 2008, 456: 60-65. 10.1038/nature07484. Affymetrix File Parsing SDK. [http://www.bioconductor.org/packages/2.2/bioc/html/affxparser.html] Kirkpatrick S, Gelatt CD, Vecchi MP: Optimization by simulated annealing. Science. 1983, 220: 671-680. 10.1126/science.220.4598.671. Macconaill LE, Aldred MA, Lu X, Laframboise T: Toward accurate high-throughput SNP genotyping in the presence of inherited copy number variation. BMC Genomics. 2007, 8: 211-10.1186/1471-2164-8-211. Birdsuite: Downloads. [http://www.broad.mit.edu/science/programs/medical-and-population-genetics/birdsuite/birdsuite-downloads-0] PennCNV Download. [http://www.openbioinformatics.org/penncnv/penncnv_download.html] PennCNV-Affy Tutorials. [http://www.openbioinformatics.org/penncnv/penncnv_tutorial_affy_gw6.html] QuantiSNP Download. [http://www.well.ox.ac.uk/~ioannisr/quantisnp/] QuantiSNP for Affymetrix Tutorials. [http://groups.google.co.uk/group/quantisnp/files]