Hate speech recognition in multilingual text: hinglish documents

International Journal of Information Technology - Tập 15 - Trang 1319-1331 - 2023
Arun Kumar Yadav1, Mohit Kumar1, Abhishek Kumar1, Shivani1, Kusum1, Divakar Yadav2
1Department of Computer Science & Engineering, NIT Hamirpur (HP), Hamirpur, India
2School of Computer and Information Sciences, Indira Gandhi National Open University, Delhi, India

Tóm tắt

The Internet is a boon for mankind but its misuse has been increasing drastically. Social networking platforms such as Facebook, Twitter and Instagram play a predominant role in expressing views by the users. Sometimes users wield abusive or inflammatory language, that may provoke readers. This paper aims to evaluate various machine learning and deep learning techniques to detect hate speech on various social media platforms in the Hinglish (English-Hindi code-mix) language. In this paper, we apply and evaluate several machine learning and deep learning methods, along with various feature extraction and word-embedding techniques, on a consolidated dataset of 20600 instances, for hate speech detection from tweets and comments in Hinglish. The experimental results reveal that deep learning models perform better than machine learning models in general. Among the deep learning models, the CNN-BiLSTM model with word2vec word embedding provides the best results. The model yields 0.876 accuracy, 0.830 precision, 0.840 recall and 0.835 F1-score. These results surpass the recent state-of-art approaches.

Tài liệu tham khảo

India Social Media Statistics (2021) The Global Statistics. 2021 Dec. Available from: https://www.theglobalstatistics.com/india-social-media-statistics/ Hate speech (2021) Wikimedia Foundation; Available from: https://en.wikipedia.org/w/index.php?title=Hate_speech &oldid=1059042962 Council of Europe;. Available from: https://www.coe.int/en/web/portal/home Available from: https://www.coe.int/en/web/portal/home 2020.Available from: https://backlinko.com/instagram-users Khanday AMUD, Khan QR, Rabani ST (2021) Identifying propaganda from online social networks during COVID-19 using machine learning techniques. Int J Inform Technol 13:115–22 Hamid Y, Elyassami S, Gulzar Y, Balasaraswathi VR, Habuza T, Wani S (2022) An improvised CNN model for fake image detection. Int J Informat Technol 15:5–15 Yadav AK, Singh A, Dhiman M, Kaundal R, Verma A, Yadav D (2022) Extractive text summarization using deep learning approach. Int J Informat Technol 14(5):2407–15 Bharti S, Yadav AK, Kumar M, Yadav D (2021) Cyberbullying detection from tweets using deep learning. Kybernetes 51(9):2695–2711 Poletto F, Basile V, Sanguinetti M, Bosco C, Patti V (2020) Resources and benchmark corpora for hate speech detection: a systematic review. Lang Res Eval 55:477–523 Shah SR, Kaushik A (2019) Sentiment analysis on indian indigenous languages: a review on multilingual opinion mining. arXiv preprint arXiv:1911.12848 Kaur S, Singh S, Kaushal S (2021) Abusive content detection in online user-generated data: a survey. Procedia Comp Sci 189:274–81 Yadav A, Vishwakarma DK (2020) Sentiment analysis using deep learning architectures: a review. Artifi Intel Rev 53(6):4335–85 Drias HH, Drias Y (2020) Mining twitter data on COVID-19 for sentiment analysis and frequent patterns discovery. medRxiv 18:2020 Thakur V, Sahu R, Omer S (2020) Current State of Hinglish Text Sentiment Analysis. In: Proceedings of the International Conference on Innovative Computing & Communications (ICICC) Srivastava V, Singh M (2021) Hinge: A dataset for generation and evaluation of code-mixed hinglish text. arXiv preprint arXiv:2107.03760 Akuma S, Lubem T, Adom IT (2022) Comparing Bag of Words and TF-IDF with different models for hate speech detection from live tweets. International Journal of Information Technology 14, 3629–3635. Kumar P, Vardhan M (2022) PWEBSA: Twitter sentiment analysis by combining Plutchik wheel of emotion and word embedding. International Journal of Information Technology 14, 69–77. Kumar R, Reganti AN, Bhatia A, Maheshwari T (2018) Aggression-annotated corpus of hindi-english code-mixed data. arXiv preprint arXiv:1803.09402 Li T, Lin L, Choi M, Fu K, Gong S, Wang J (2018) Youtube av 50k: an annotated corpus for comments in autonomous vehicles. In:2018 international joint symposium on artificial intelligence and natural language processing (iSAI-NLP). IEEE 2018:1–5 Ravi K, Ravi V (2016) Sentiment classification of Hinglish text. In: 2016 3rd International Conference on Recent Advances in Information Technology (RAIT). IEEE; p. 641-5 Bohra A, Vijay D, Singh V, Akhtar SS, Shrivastava M (2018) A dataset of hindi–english code-mixed social media text for hate speech detection. Proceeding of second workshop on computational modeling of people’s opinions personality and emotions in social media. IEEE Pamungkas EW, Basile V, Patti V (2020) Misogyny detection in twitter: a multilingual and cross-domain study. Info Process Manag 57(6):102360 Sreelakshmi K, Premjith B, Soman K (2020) Detection of hate speech text in Hindi-English code-mixed data. Procedia Comp Sci 171:737–44 Kamble S, Joshi A (2018) Hate speech detection from code-mixed hindi-english tweets using deep learning models. arXiv preprint arXiv:1811.05145 Mathur P, Sawhney R, Ayyar M, Shah R (2018) Did you offend me? classification of offensive tweets in hinglish language. In: Proceedings of the 2nd Workshop on Abusive Language Online (ALW2); p. 138-48 Mathur P, Shah R, Sawhney R, Mahata D (2018) Detecting offensive tweets in hindi-english code-switched language. In: Proceedings of the Sixth International Workshop on Natural Language Processing for Social Media; p. 18-26 Singh P, Lefever E (2020) Sentiment analysis for hinglish code-mixed tweets by means of cross-lingual word embeddings. In: Proceedings of the The 4th Workshop on Computational Approaches to Code Switching; p. 45-51 Kovács G, Alonso P, Saini R (2021) Challenges of hate speech detection in social media. SN Comp Sci 2(2):1–15 Chopra S, Sawhney R, Mathur P, Shah RR (2020) Hindi–english hate speech detection: Author profiling, debiasing, and practical perspectives. Proc AAAI Conf Artif Intell 34:386–93 Gupta V, Sehra V, Vardhan YR, et al (2021) Hindi-English Code Mixed Hate Speech Detection using Character Level Embeddings. In: 2021 5th International Conference on Computing Methodologies and Communication (ICCMC). IEEE; p. 1112-8 Santosh T, Aravind K (2019) Hate speech detection in hindi-english code-mixed social media text. In: Proceedings of the ACM India joint international conference on data science and management of data; p. 310-3 Sasidhar TT, Premjith B, Soman K (2020) Emotion detection in hinglish (hindi+ english) code-mixed social media text. Procedia Comp Sci 171:1346–52 Kapoor R, Kumar Y, Rajput K, Shah RR, Kumaraguru P, Zimmermann R (2019) Mind your language: abuse and offense detection for code-switched languages. Proc AAAI conf Artif Intell 33:9951–2 Sengupta A, Bhattacharjee SK, Akhtar MS, Chakraborty T (2021) Does aggression lead to hate? Detecting and reasoning offensive traits in hinglish code-mixed texts. Neurocomputing 488:598–617 Sharma A, Kabra A, Jain M (2022) Ceasing hate with MoH: hate speech detection in hindi-english code-switched language. Inf Proc Manag 59(1):102760 Zhu AZ, Thakur D, Özaslan T, Pfrommer B, Kumar V, Daniilidis K (2018) The multivehicle stereo event camera dataset: an event camera dataset for 3D perception. IEEE Robot Automat Lett 3(3):2032–9 Mandl T, Modha S, Shahi GK, Madhu H, Satapara S, Majumder P, et al (2021) Overview of the HASOC subtrack at FIRE 2021: Hate speech and offensive content identification in English and Indo-Aryan languages. arXiv preprint arXiv:2112.09301 Grimm LG, Yarnold PR (1995) Reading and understanding multivariate statistics. American Psychological Association; Fürnkranz J (2010) In: Sammut C, Webb GI, editors. Decision Tree. Boston, MA: Springer US; p. 263-7. Available from: https://doi.org/10.1007/978-0-387-30164-8_204 Chen T, Guestrin C (2016) XGBoost: A Scalable Tree Boosting System. In: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. KDD ’16. New York, NY, USA: ACM; p. 785-94. Available from: http://doi.acm.org/10.1145/2939672.2939785 Liaw A, Wiener M et al (2002) Classification and regression by randomForest. R news. 2(3):18–22 Bahlmann C, Haasdonk B, Burkhardt H (2002) Online handwriting recognition with support vector machines-a kernel approach. In: Proceedings Eighth International Workshop on Frontiers in Handwriting Recognition. IEEE; p. 49-54 Samet H (2007) K-nearest neighbor finding using MaxNearestDist. IEEE Transactions on Pattern Analysis and Machine Intelligence. 30(2):243–52 Webb GI (2010) In: Sammut C, Webb GI, editors. Naïve Bayes. Boston, MA: Springer US; p. 713-4. Available from: https://doi.org/10.1007/978-0-387-30164-8_576 Hochreiter S, Schmidhuber J (1997) Long short-term memory. Neural computation. 9(8):1735–80 Xu G, Meng Y, Qiu X, Yu Z, Wu X (2019) Sentiment analysis of comment texts based on BiLSTM. Ieee Access. 7:51522–32 Albawi S, Mohammed TA, Al-Zawi S, Understanding of a convolutional neural network. In, (2017) international conference on engineering and technology (ICET). Ieee 2017:1–6 Ghosh S, Priyankar A, Ekbal A, Bhattacharyya P (2023) Multitasking of sentiment detection and emotion recognition in code-mixed Hinglish data. Knowledge-Based Systems. 260:110182