Extracting indices from Japanese legal documents

Tho Thi Ngoc Le1, Kiyoaki Shirai1, Minh Le Nguyen1, Akira Shimazu1
1School of Information Science, Japan Advanced Institute of Science and Technology, Nomi, Japan

Tóm tắt

This article addresses the problem of automatically extracting legal indices which express the important contents of legal documents. Legal indices are not limited to single-word keywords and compound-word (or phrase) keywords, they are also clause keywords. We approach index extraction using structural information of Japanese sentences, i.e. chunks and clauses. Based on the assumption that legal indices are composed of important tokens from the documents, extracting legal indices is treated as a problem of collecting chunks and clauses that contain as many important tokens as possible. Each token is assigned a weight which is a statistical score, e.g. TF–IDF and Okapi BM25, to indicate its importance. The importance of a chunk or clause is determined based on the average weight of tokens included in that chunk or clause. Then, highly weighted chunks and clauses are recognized as the indices for legal documents. The experimental results on Japanese National Pension Act data show that our proposed method achieves better performance (8.6 % higher on F1-score) than TextRank, the most popular unsupervised method in extracting single-word and compound-word keywords. In addition, this approach is also applicable to extract clause keywords with high performance.

Từ khóa


Tài liệu tham khảo

Ashley KD, Brüninghaus S (2003) A predictive role for intermediate legal concepts. In: Proceedings of conference on JURIX’03, pp 1–10 Biber D, Johansson S, Leech G, Conrad S, Finegan E (1999) Longman grammar of spoken and written English. Pearson Education, England, pp 120 Blair DC, Maron ME (1985) An evaluation of retrieval effectiveness for a full-text document-retrieval system. Commun ACM 28(3):289–299 Brüninghaus S, Ashley KD (1999) Toward adding knowledge to learning algorithms for indexing legal cases. In: Proceedings of conference on ICAIL’99, pp 9–17 Frank E, Paynter GW, Witten IH, Gutwin C, Nevill-Manning CG (1999) Domain-specific keyphrase extraction. In: Proceedings of 16th international joint conference on artificial intelligence, pp 668–673 Grabmair M, Ashley KD (2011) Facilitating case comparison using value judgments and intermediate legal concepts. In: Proceedings of conference on ICAIL’11, pp 161–170 Hulth A (2003) Improved automatic keyword extraction given more linguistic knowledge. In: Proceedings of conference on EMNLP-ACL’03, pp 216–223 Katayama T (2007) Legal engineering—an engineering approach to laws in e-society age. In: Proceedings of the 1st international workshop on JURISIN, pp 1–5 Kudo T, Matsumoto Y (2002) Japanese dependency analysis using cascaded chunking. In: Proceedings of conference on ACL-CoNLL’02, pp 63–69 Kurohashi S, Nagao M (1994a) A syntactic analysis method of long Japanese sentences based on coordinate structures’ detection. Nat Lang Process 1(1):35–57 (In Japanese) Kurohashi S, Nagao M (1994b) A syntactic analysis method of long Japanese sentences based on the detection of conjunctive structures. Comput Linguist 20(4):507–534 Le TTN, Nguyen ML, Shimazu A (2013) Unsupervised keyword extraction for Japanese legal documents. In: Proceedings of conference on JURIX’13, pp 97–106 Litvak M, Last M (2008) Graph-based keyword extraction for single-document summarization. In: Proceedings of conference on COLING’08, pp 17–24 Liu Z, Li P, Zheng Y, Sun M (2009) Clustering to find exemplar terms for keyphrase extraction. In: Proceedings of conference on EMNLP-ACL’09, pp 257–266 Liu Z, Huang W, Zheng Y, Sun M (2010) Automatic keyphrase extraction via topic decomposition. In: Proceedings of conference on EMNLP-ACL’10, pp 366–376 Manning CD, Schütze H (1999) Foundations of statistical natural language processing. MIT Press, Cambridge Maruyama T, Kashioka H, Kumano T, Tanaka H (2004) Development and evaluation of Japanese clause boundaries annotation program. Nat Lang Process 11(3):39–68 (in Japanese) Mathieu J (1999) Adaptation of a keyphrase extractor for Japanese text. In: Proceedings of conference on CAIS’99, pp 182–189 Matsuo Y, Ishizuka M (2004) Keyword extraction from a single document using word co-occurrence statistical information. Int J Artif Intell Tools 13:157–169 Maxwell KT, Schafer B (2008) Concept and context in legal information retrieval. In: Proceedings of conference on JURIX’08, pp 63–72 Mihalcea R, Tarau P (2004) TextRank: bringing order into texts. In: Proceedings of conference on EMNLP-ACL’04, pp 404–411 Moens MF, Angheluta R (2003) Concept extraction from legal cases: the use of a statistic of coincidence. In: Proceedings of conference on ICAIL’03, pp 142–146 Nakagawa H, Mori T (2002) A simple but powerful automatic term extraction method. In: COLING-02 on COMPUTERM 2002: 2nd international workshop on computational terminology, vol 14, pp 1–7 Nakagawa H, Mori T (2003) Automatic term recognition based on statistics of compound nouns and their components. Terminology 9(2):201–219 Ogawa Y, Matsuda T (1997) Overlapping statistical word indexing: a new indexing method for Japanese text. In: Proceedings of conference on ACM-SIGIR’97, pp 226–234 Robertson SE, Walker S, Jones S, Hancock-Beaulieu MM, Gatford M (1994) Okapi at TREC-3. In: Proceedings of TREC-3, pp 109–126 Saravanan M, Ravindran B, Raman S (2009) Improving legal information retrieval using an ontological framework. Artif Intell Law 17(2):101–124 Suzuki Y, Fukumoto F, Sekiguchi Y (1997) Keyword extraction of radio news using term weighting for speech recognition. In: Proceedings of conference on ACM-NLPRS’97, pp 301–306 Suzuki Y, Fukumoto F, Sekiguchi Y (1998) Keyword extraction using term-domain interdependence for dictation of radio news. In: Proceedings of conference on COLING’98, pp 1272–1276 Turney PD (2000) Learning algorithms for keyphrase extraction. Inf Retr 2:303–336 Turney PD (1999) Learning to extract keyphrases from text. National Research Council of Canada, Institute for Information Technology, technical report ERB-1057 Wan X, Xiao J (2008) Single document keyphrase extraction using neighborhood knowledge. In: Proceedings of conference on AAAI’08, pp 855–860 Wu W, Zhang B, Ostendorf M (2010) Automatic generation of personalized annotation tags for twitter users. In: Proceedings of conference on HLT-ACL’10, pp 689–692 Yoshida M, Nakagawa H (2005) Automatic term extraction based on perplexity of compound words. In: Proceedings of conference on IJCNLP’05, pp 269–279 Zhao WX, Jiang J, He J, Song Y, Achananuparp P, Lim EP, Li X (2011) Topical keyphrase extraction from twitter. In: Proceedings of conference on HLT-ACL’11, pp 379–388