Keyphrase Extraction Using Knowledge Graphs

Data Science and Engineering - Tập 2 - Trang 275-288 - 2017
Wei Shi1, Weiguo Zheng1, Jeffrey Xu Yu1, Hong Cheng1, Lei Zou2
1The Chinese University of Hong Kong, Shatin, Hong Kong, China
2Peking University, Beijing, China

Tóm tắt

Extracting keyphrases from documents automatically is an important and interesting task since keyphrases provide a quick summarization for documents. Although lots of efforts have been made on keyphrase extraction, most of the existing methods (the co-occurrence-based methods and the statistic-based methods) do not take semantics into full consideration. The co-occurrence-based methods heavily depend on the co-occurrence relations between two words in the input document, which may ignore many semantic relations. The statistic-based methods exploit the external text corpus to enrich the document, which introduce more unrelated relations inevitably. In this paper, we propose a novel approach to extract keyphrases using knowledge graphs, based on which we could detect the latent relations of two keyterms (i.e., noun words and named entities) without introducing many noises. Extensive experiments over real data show that our method outperforms the state-of-the-art methods including the graph-based co-occurrence methods and statistic-based clustering methods.

Tài liệu tham khảo

Bizer C, Lehmann J, Kobilarov G, Auer S, Becker C, Cyganiak R, Hellmann S (2009) Dbpedia—a crystallization point for the web of data. J Web Sem 7(3):154–165 Blei DM, Ng AY, Jordan MI (2003) Latent Dirichlet allocation. J Mach Learn Res 3:993–1022 Boudin F (2013) A comparison of centrality measures for graph-based keyphrase extraction. In: IJCNLP 2013, pp 834–838 Chen Z, Cafarella MJ, Jagadish HV (2016) Long-tail vocabulary dictionary extraction from the web. In: Proceedings of the ninth ACM international conference on Web search and data mining, San Francisco, CA, USA, February 22–25, 2016, pp 625–634 Cilibrasi R, Vitányi PMB (2004) The google similarity distance. CoRR, abs/cs/0412098 Freeman LC (1977) A set of measures of centrality based on betweenness. Sociometry 40:35–41 Grineva MP, Grinev MN, Lizorkin D (2009) Extracting key terms from noisy and multitheme documents. In: WWW 2009, pp 661–670 Hammouda KM, Matute DN, Kamel MS (2005) Corephrase: keyphrase extraction for document clustering. In: MLDM 2005, pp 265–274 Haveliwala TH (2002) Topic-sensitive pagerank. In: WWW 2002, pp 517–526 Huang C, Tian Y, Zhou Z, Ling CX, Huang T (2006) Keyphrase extraction using semantic networks structure analysis. In: ICDM 2006, pp 275–284 Hulth A (2003) Improved automatic keyword extraction given more linguistic knowledge. In: EMNLP 2003, pp 216–223 Hulth A (2003) Reducing false positives by expert combination in automatic keyword indexing. In: RANLP 2003, pp 367–376 Jeh G, Widom J (2002) Simrank: a measure of structural-context similarity. In: Proceedings of the eighth ACM SIGKDD international conference on knowledge discovery and data mining, July 23–26, 2002, Edmonton, Alberta, Canada, pp 538–543 Jiang X, Hu Y, Li H (2009) A ranking approach to keyphrase extraction. In: SIGIR 2009, pp 756–757 Kusner MJ, Sun Y, Kolkin NI, Weinberger KQ (2015) From word embeddings to document distances. In: ICML 2015, pp 957–966 Liu Z, Huang W, Zheng Y, Sun M (2010) Automatic keyphrase extraction via topic decomposition. In: EMNLP, pp 366–376 Liu Z, Li P, Zheng Y, Sun M (2009) Clustering to find exemplar terms for keyphrase extraction. In: EMNLP 2009, pp 257–266 Mann GS, McCallum A (2010) Generalized expectation criteria for semi-supervised learning with weakly labeled data. J Mach Learn Res 11:955–984 Manning CD, Surdeanu M, Bauer J, Finkel JR, Bethard S, McClosky D (2014) The Stanford CoreNLP natural language processing toolkit. In: ACL 2014, pp 55–60 Mendes PN, Jakob M, García-Silva A, Bizer C (2011) Dbpedia spotlight: shedding light on the web of documents. In: I-SEMANTICS, pp 1–8 Mihalcea R, Csomai A (2007) Wikify!: linking documents to encyclopedic knowledge. In: CIKM 2007, pp 233–242 Mihalcea R, Tarau P (2004) Textrank: bringing order into text. In: EMNLP 2004, pp 404–411 Mikolov T, Chen K, Corrado G, Dean J (2013) Efficient estimation of word representations in vector space. CoRR Page L, Brin S, Motwani R, Winograd T (1999) The pagerank citation ranking: bringing order to the web Rong X, Chen Z, Mei Q, Adar E (2016) Egoset: exploiting word ego-networks and user-generated ontology for multifaceted set expansion. In: Proceedings of the ninth ACM international conference on Web search and data mining, San Francisco, CA, USA, February 22–25, 2016, pp 645–654 Russell SJ, Norvig P (2003) Artificial intelligence—a modern approach: the intelligent agent book. Prentice Hall Shi W, Zheng W, Yu JX, Cheng H, Zou L (2017) Keyphrase extraction using knowledge graphs. In: Web and Big Data—first international joint conference, APWeb-WAIM 2017, Beijing, China, July 7–9, 2017, Proceedings, Part I, pp 132–148 Tsatsaronis G, Varlamis I, Nørvåg K (2010) Semanticrank: ranking keywords and sentences using semantic graphs. In: COLING 2010, pp 1074–1082 Turney PD (2002) Learning algorithms for keyphrase extraction. CoRR, cs.LG/0212020 Turney PD (2002) Learning to extract keyphrases from text. CoRR, cs.LG/0212013 Wan X, Xiao J (2010) Exploiting neighborhood knowledge for single document summarization and keyphrase extraction. ACM Trans Inf Syst 28(2):8. https://doi.org/10.1145/1740592.1740596 Wan X, Yang J, Xiao J (2007) Towards an iterative reinforcement approach for simultaneous document summarization and keyword extraction. In: Proceedings of the 45th Annual Meeting of the Association of Computational Linguistics, Prague, Czech Republic, pp 552–559 Wang RC, Cohen WW (2007) Language-independent set expansion of named entities using the web. In: Proceedings of the 7th IEEE international conference on data mining (ICDM 2007), October 28–31, 2007, Omaha, Nebraska, USA, pp 342–350 Witten IH, Paynter GW, Frank E, Gutwin C, Nevill-Manning CG (1999) KEA: practical automatic keyphrase extraction. In: Proceedings of the fourth ACM conference on digital libraries, pp 254–255 Youn E, Jeong MK (2009) Class dependent feature scaling method using naive Bayes classifier for text datamining. Pattern Recognit Lett 30(5):477–485