Nghiên cứu Định kiến Giới tính trong BERT

Cognitive Computation - Tập 13 - Trang 1008-1018 - 2021
Rishabh Bhardwaj1, Navonil Majumder1, Soujanya Poria1
1DeCLaRe Lab, Singapore University of Technology and Design, Singapore, Singapore

Tóm tắt

Trong công trình này, chúng tôi phân tích định kiến giới tính do BERT gây ra trong các tác vụ hạ nguồn. Chúng tôi cũng đề xuất các giải pháp để giảm thiểu định kiến giới tính. Các mô hình ngôn ngữ theo ngữ cảnh (CLMs) đã đưa các tiêu chuẩn xử lý ngôn ngữ tự nhiên (NLP) lên một tầm cao mới. Việc sử dụng các nhúng từ do CLMs cung cấp trong các tác vụ hạ nguồn như phân loại văn bản đã trở thành một chuẩn mực mới. Tuy nhiên, nếu không được giải quyết, CLMs có khả năng học các định kiến giới tính vốn có trong tập dữ liệu. Kết quả là, dự đoán của các mô hình NLP hạ nguồn có thể thay đổi đáng kể khi thay đổi các từ giới tính, chẳng hạn như thay “he” thành “she”, hoặc thậm chí các từ trung lập về giới. Trong bài báo này, chúng tôi tập trung phân tích vào một CLM phổ biến, tức là BERT. Chúng tôi phân tích định kiến giới tính mà nó gây ra trong năm tác vụ hạ nguồn liên quan đến dự đoán cảm xúc và cường độ cảm xúc. Đối với mỗi tác vụ, chúng tôi huấn luyện một hồi quy đơn giản sử dụng nhúng từ của BERT. Sau đó, chúng tôi đánh giá định kiến giới tính trong các hồi quy bằng cách sử dụng một tập dữ liệu đánh giá công bằng. Lý tưởng và theo thiết kế cụ thể, các mô hình nên loại bỏ các đặc điểm thông tin về giới từ đầu vào. Tuy nhiên, kết quả cho thấy sự phụ thuộc đáng kể của các dự đoán của hệ thống vào các từ và cụm từ cụ thể về giới. Chúng tôi cho rằng những định kiến như vậy có thể được giảm thiểu bằng cách loại bỏ các đặc điểm giới tính từ nhúng từ. Do đó, đối với mỗi lớp trong BERT, chúng tôi xác định các hướng chủ yếu mã hóa thông tin giới tính. Không gian được hình thành bởi các hướng này được gọi là không gian con giới tính trong không gian ngữ nghĩa của nhúng từ. Chúng tôi đề xuất một thuật toán tìm kiếm các hướng giới tính chi tiết, tức là một hướng chính cho mỗi lớp BERT. Điều này giúp loại bỏ nhu cầu về việc thực hiện không gian con giới tính trong nhiều chiều và ngăn chặn các thông tin quan trọng khác bị bỏ sót. Các thí nghiệm cho thấy việc loại bỏ các thành phần nhúng trong các hướng giới tính đạt được thành công lớn trong việc giảm thiểu định kiến do BERT gây ra trong các tác vụ hạ nguồn. Điều tra tiết lộ một định kiến giới tính đáng kể mà một mô hình ngôn ngữ theo ngữ cảnh (tức là BERT) gây ra trong các tác vụ hạ nguồn. Giải pháp được đề xuất dường như đầy hứa hẹn trong việc giảm thiểu những định kiến như vậy.

Từ khóa

#định kiến giới tính #BERT #xử lý ngôn ngữ tự nhiên #mô hình ngôn ngữ theo ngữ cảnh #nhúng từ

Tài liệu tham khảo

Hussain A, Tahir A, Hussain Z, Sheikh Z, Gogate M, Dashtipour K, Ali A, Sheikh A. Artificial intelligence-enabled analysis of public attitudes on facebook and twitter toward covid-19 vaccines in the united kingdom and the united states: Observational study. J Med Intern Res. 2021;23(4):e26627. Basta C, Costa-jussà MR, Casas N. Evaluating the underlying gender bias in contextualized word embeddings. In Proceedings of the First Workshop on Gender Bias in Natural Language Processing. Association for Computational Linguistics. 2019:33–39. Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser Ł, Polosukhin I. Attention is all you need. In Advances in neural information processing systems 2017:5998–6008. Caliskan A, Bryson JJ, Narayanan A. Semantics derived automatically from language corpora contain human-like biases. Science. 2017;356(6334):83–186. Dashtipour K, Gogate M, Cambria E, Hussain A. A novel context-aware multimodal framework for persian sentiment analysis. 2021. https://arxiv.org/abs/2103.02636. Devlin J, Chang MW, Lee K, Toutanova K. Bert: Pre-training of deep bidirectional transformers for language understanding. 2018. https://arxiv.org/abs/1810.04805. Bolukbasi T, Chang KW, Zou JY, Saligrama V, Kalai AT. Man is to computer programmer as woman is to homemaker? debiasing word embeddings. In Advances in neural information processing systems. 2016:4349–4357. Gonen H, Goldberg Y. Lipstick on a pig: Debiasing methods cover up systematic gender biases in word embeddings but do not remove them. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). Association for Computational Linguistics. 2019:609–614. Caliskan A, Bryson JJ, Narayanan A. Semantics derived automatically from language corpora contain human-like biases. Science. 2017;356(6334):183–6. Hussain A, Tahir A, Hussain Z, Sheikh Z, Gogate M, Dashtipour K, Ali A, Sheikh A. Artificial intelligence–enabled analysis of public attitudes on facebook and twitter toward covid-19 vaccines in the united kingdom and the united states: Observational study. Journal of medical Internet research. 2021;23(4):e26627. Kim Y. Convolutional neural networks for sentence classification. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP). Association for Computational Linguistics. 2014:1746–1751. Kiritchenko S, Mohammad S. Examining gender and race bias in two hundred sentiment analysis systems. In Proceedings of the Seventh Joint Conference on Lexical and Computational Semantics. Association for Computational Linguistics. 2018:43–53. Kurita K, Vyas N, Pareek A, Black AW, Tsvetkov Y. Measuring bias in contextualized word representations. In Proceedings of the First Workshop on Gender Bias in Natural Language Processing. Association for Computational Linguistics. 2019:166–172. Liu Y, Ott M, Goyal N, Du J, Joshi M, Chen D, Levy O, Lewis M, Zettlemoyer L, Stoyanov V. Roberta: A robustly optimized bert pretraining approach. 2019. https://arxiv.org/abs/1907.11692. Lu J, Batra D, Parikh D, Lee S. Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. In Advances in Neural Information Processing Systems. 2019:13–23. Mohammad S, Bravo-Marquez F, Salameh M, Kiritchenko S. SemEval-2018 task 1: Affect in tweets. In Proceedings of The 12th International Workshop on Semantic Evaluation. Association for Computational Linguistics. 2018:pp. 1–17. Morcos A, Raghu M, Bengio S. Insights on representational similarity in neural networks with canonical correlation. In Advances in Neural Information Processing Systems. 2018:5727–5736. Peters M, Neumann M, Iyyer M, Gardner M, Clark C, Lee K, Zettlemoyer L. Deep contextualized word representations. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). Association for Computational Linguistics. 2018:2227–2237. Elazar Y, Ravfogel S, Jacovi A, Goldberg Y. Amnesic probing: Behavioral explanation with amnesic counterfactuals. Transactions of the Association for Computational Linguistics. 2021;9:160–75. Morcos A, Raghu M, Bengio S. Insights on representational similarity in neural networks with canonical correlation. In Advances in Neural Information Processing Systems. 2018:5727–5736. Tenney I, Xia P, Chen B, Wang A, Poliak A, McCoy RT, Kim N, Van Durme B, Bowman SR, Das D, et al. What do you learn from context? probing for sentence structure in contextualized word representations. 2019. https://arxiv.org/abs/1905.06316. Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser Ł, Polosukhin I. Attention is all you need. In Advances in neural information processing systems. 2017:5998–6008. Voita E, Titov I. Information-theoretic probing with minimum description length. 2020. https://arxiv.org/abs/2003.12298. Wolf T, Debut L, Sanh V, Chaumond J, Delangue C, Moi A, Cistac P, Rault T, Louf R, Funtowicz M, Brew J. Huggingface’s transformers: State-of-the-art natural language processing. 2019. https://arxiv.org/abs/1910.03771. Yang Z, Dai Z, Yang Y, Carbonell J, Salakhutdinov R, Le Q. V. Xlnet: Generalized autoregressive pretraining for language understanding. 2019. https://arxiv.org/abs/1906.08237. Zhao J, Wang T, Yatskar M, Cotterell R, Ordonez V, Chang KW. Gender bias in contextualized word embeddings. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). Association for Computational Linguistics. 2019:629–634. Zhao J, Wang T, Yatskar M, Ordonez V, Chang KW. Gender bias in coreference resolution: Evaluation and debiasing methods. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers). Association for Computational Linguistics. 2018:15–20. Zhao J, Zhou Y, Li Z, Wang W, Chang KW. Learning gender-neutral word embeddings. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics. 2018:4847–4853. Zhao J, Zhou Y, Li Z, Wang W, Chang KW. Learning gender-neutral word embeddings. 2018. https://arxiv.org/abs/1809.01496.