中文
相关论文

相关论文: The Arabic Parallel Gender Corpus 2.0: Extensions …

200 篇论文

In academia, plagiarism is certainly not an emerging concern, but it became of a greater magnitude with the popularisation of the Internet and the ease of access to a worldwide source of content, rendering human-only intervention…

计算与语言 · 计算机科学 2022-01-11 Mehdi Abdelhamid , Faical Azouaou , Sofiane Batata

There is a growing body of work in recent years to develop pre-trained language models (PLMs) for the Arabic language. This work concerns addressing two major problems in existing Arabic PLMs which constraint progress of the Arabic NLU and…

Arabic dialects form a diverse continuum, yet NLP models often treat them as discrete categories. Recent work addresses this issue by modeling dialectness as a continuous variable, notably through the Arabic Level of Dialectness (ALDi).…

计算与语言 · 计算机科学 2025-08-26 Sanad Shaban , Nizar Habash

In this paper we propose a novel method of augmenting parallel text corpora which promises good quality and is also capable of producing many fold larger corpora than the seed corpus we start with. We do not need any additional monolingual…

计算与语言 · 计算机科学 2024-10-07 Vibhuti Kumari , Narayana Murthy Kavi

Recent studies in the field of Machine Translation (MT) and Natural Language Processing (NLP) have shown that existing models amplify biases observed in the training data. The amplification of biases in language technology has mainly been…

计算与语言 · 计算机科学 2021-02-02 Eva Vanmassenhove , Dimitar Shterionov , Matthew Gwilliam

Gender bias represents a form of systematic negative treatment that targets individuals based on their gender. This discrimination can range from subtle sexist remarks and gendered stereotypes to outright hate speech. Prior research has…

计算与语言 · 计算机科学 2024-03-19 Karolina Stańczak

The use of Project Gutenberg (PG) as a text corpus has been extremely popular in statistical analysis of language for more than 25 years. However, in contrast to other major linguistic datasets of similar importance, no consensual full…

计算与语言 · 计算机科学 2018-12-20 Martin Gerlach , Francesc Font-Clos

This study is an attempt to build a contemporary linguistic corpus for Arabic language. The corpus produced, is a text corpus includes more than five million newspaper articles. It contains over a billion and a half words in total, out of…

计算与语言 · 计算机科学 2016-11-15 Ibrahim Abu El-khair

Most current machine translation models are mainly trained with parallel corpora, and their translation accuracy largely depends on the quality and quantity of the corpora. Although there are billions of parallel sentences for a few…

计算与语言 · 计算机科学 2022-03-01 Makoto Morishita , Katsuki Chousa , Jun Suzuki , Masaaki Nagata

This paper is devoted to the development of a localized Large Language Model (LLM) specifically for Arabic, a language imbued with unique cultural characteristics inadequately addressed by current mainstream models. Significant concerns…

Pretrained multilingual models exhibit the same social bias as models processing English texts. This systematic review analyzes emerging research that extends bias evaluation and mitigation approaches into multilingual and non-English…

计算与语言 · 计算机科学 2025-09-08 Lance Calvin Lim Gamboa , Yue Feng , Mark Lee

Large language models (LLMs) are being increasingly used in urban planning, but since gendered space theory highlights how gender hierarchies are embedded in spatial organization, there is concern that LLMs may reproduce or amplify such…

计算与语言 · 计算机科学 2026-04-17 Binxian Su , Haoye Lou , Shucheng Zhu , Weikang Wang , Ying Liu , Dong Yu , Pengyuan Liu

Language in the Arab world presents a complex diglossic and multilingual setting, involving the use of Modern Standard Arabic, various dialects and sub-dialects, as well as multiple European languages. This diverse linguistic landscape has…

计算与语言 · 计算机科学 2025-01-24 Injy Hamed , Caroline Sabty , Slim Abdennadher , Ngoc Thang Vu , Thamar Solorio , Nizar Habash

The successful application of neural methods to machine translation has realized huge quality advances for the community. With these improvements, many have noted outstanding challenges, including the modeling and treatment of gendered…

计算与语言 · 计算机科学 2020-10-16 Hila Gonen , Kellie Webster

The quality and accessibility of multilingual datasets are crucial for advancing machine translation. However, previous corpora built from United Nations documents have suffered from issues such as opaque process, difficulty of…

计算与语言 · 计算机科学 2025-09-22 Qiuyang Lu , Fangjian Shen , Zhengkai Tang , Qiang Liu , Hexuan Cheng , Hui Liu , Wushao Wen

The processing of the Arabic language is a complex field of research. This is due to many factors, including the complex and rich morphology of Arabic, its high degree of ambiguity, and the presence of several regional varieties that need…

计算与语言 · 计算机科学 2022-05-20 Karim El Haff , Mustafa Jarrar , Tymaa Hammouda , Fadi Zaraket

While multilingual large language models (LLMs) perform well on high-level tasks like translation and question answering, their ability to handle grammatical gender and morphological agreement remains underexplored. In morphologically rich…

计算与语言 · 计算机科学 2026-04-22 Mehul Agarwal , Aditya Aggarwal , Arnav Goel , Medha Hira , Anubha Gupta

As natural language processing systems become more widespread, it is necessary to address fairness issues in their implementation and deployment to ensure that their negative impacts on society are understood and minimized. However, there…

计算与语言 · 计算机科学 2022-04-08 António Câmara , Nina Taneja , Tamjeed Azad , Emily Allaway , Richard Zemel

Gender bias is largely recognized as a problematic phenomenon affecting language technologies, with recent studies underscoring that it might surface differently across languages. However, most of current evaluation practices adopt a…

计算与语言 · 计算机科学 2022-03-21 Beatrice Savoldi , Marco Gaido , Luisa Bentivogli , Matteo Negri , Marco Turchi

This paper presents the Arabic Women and Society Corpus, a ten year collection of 252,487 public Arabic Facebook posts related to women's empowerment and social wellbeing. The corpus was collected from 51,660 pages across 77 countries…

计算与语言 · 计算机科学 2026-05-22 Wajdi Zaghouani , Mabrouka Bessghaier , MD. Rafiul Biswas , Shimaa Amer Ibrahim