A sequence labeling model for catchphrase identification from legal case documents

Artificial Intelligence and Law - Tập 30 - Trang 325-358 - 2021
Arpan Mandal1, Kripabandhu Ghosh2, Saptarshi Ghosh3, Sekhar Mandal1
1Department of Computer Science and Technology, Indian Institute of Engineering Science and Technology, Shibpur, Howrah, India
2Department of Computational and Data Sciences (CDS), Indian Institute of Science Education and Research (IISER) Kolkata, Kolkata, India
3Department of Computer Science and Engineering, Indian Institute of Technology Kharagpur, Kharagpur, India

Tóm tắt

In a Common Law system, legal practitioners need frequent access to prior case documents that discuss relevant legal issues. Case documents are generally very lengthy, containing complex sentence structures, and reading them fully is a strenuous task even for legal practitioners. Having a concise overview of these documents can relieve legal practitioners from the task of reading the complete case statements. Legal catchphrases are (multi-word) phrases that provide a concise overview of the contents of a case document, and automated generation of catchphrases is a challenging problem in legal analytics. In this paper, we propose a novel supervised neural sequence tagging model for the extraction of catchphrases from legal case documents. Specifically, we show that incorporating document-specific information along with a sequence tagging model can enhance the performance of catchphrase extraction. We perform experiments over a set of Indian Supreme Court case documents, for which the gold-standard catchphrases (annotated by legal practitioners) are obtained from a popular legal information system. The performance of our proposed method is compared with that of several existing supervised and unsupervised methods, and our proposed method is empirically shown to be superior to all baselines.

Tài liệu tham khảo

Breiman L (2001) Random forests. Mach learn 45(1):5–32

Truong S, Le Minh N, Satoh K, Satoshi T, Shimazu A (2017) Single and multiple layer bi-lstmcrf for recognizing requisite and effectuation parts in legal texts. In: Proceedings of Automated Semantic Analysis of Information in Legal Texts