中文
相关论文

相关论文: Gender Prediction Based on Vietnamese Names with M…

200 篇论文

We introduce MTet, the largest publicly available parallel corpus for English-Vietnamese translation. MTet consists of 4.2M high-quality training sentence pairs and a multi-domain test set refined by the Vietnamese research community.…

计算与语言 · 计算机科学 2022-10-20 Chinh Ngo , Trieu H. Trinh , Long Phan , Hieu Tran , Tai Dang , Hieu Nguyen , Minh Nguyen , Minh-Thang Luong

Gender bias in artificial intelligence (AI) and natural language processing has garnered significant attention due to its potential impact on societal perceptions and biases. This research paper aims to analyze gender bias in Large Language…

计算与语言 · 计算机科学 2023-09-04 Vishesh Thakur

We employ an audit design to investigate biases in state-of-the-art large language models, including GPT-4. In our study, we prompt the models for advice involving a named individual across a variety of scenarios, such as during car…

计算与语言 · 计算机科学 2025-01-27 Alejandro Salinas , Amit Haim , Julian Nyarko

Large Language Models (LLMs) are increasingly being used to generate text across various languages, for tasks such as translation, customer support, and education. Despite these advancements, LLMs show notable gender biases in English,…

计算与语言 · 计算机科学 2024-09-23 Ishika Joshi , Ishita Gupta , Adrita Dey , Tapan Parikh

Machine translation (MT) systems often translate terms with ambiguous gender (e.g., English term "the nurse") into the gendered form that is most prevalent in the systems' training data (e.g., "enfermera", the Spanish term for a female…

计算与语言 · 计算机科学 2024-07-31 Sarthak Garg , Mozhdeh Gheini , Clara Emmanuel , Tatiana Likhomanenko , Qin Gao , Matthias Paulik

Can we leverage LLMs to model the process of discovering novel language model (LM) architectures? Inspired by real research, we propose a multi-agent LLM approach that simulates the conventional stages of research, from ideation and…

人工智能 · 计算机科学 2025-06-26 Junyan Cheng , Peter Clark , Kyle Richardson

In the text classification problem, the imbalance of labels in datasets affect the performance of the text-classification models. Practically, the data about user comments on social networking sites not altogether appeared - the…

计算与语言 · 计算机科学 2020-10-12 Son T. Luu , Kiet Van Nguyen , Ngan Luu-Thuy Nguyen

We introduce PhoWhisper in five versions for Vietnamese automatic speech recognition. PhoWhisper's robustness is achieved through fine-tuning the Whisper model on an 844-hour dataset that encompasses diverse Vietnamese accents. Our…

音频与语音处理 · 电气工程与系统科学 2024-06-06 Thanh-Thien Le , Linh The Nguyen , Dat Quoc Nguyen

This paper studies gender bias in machine translation through the lens of Large Language Models (LLMs). Four widely-used test sets are employed to benchmark various base LLMs, comparing their translation quality and gender bias against…

计算与语言 · 计算机科学 2024-07-29 Aleix Sant , Carlos Escolano , Audrey Mash , Francesca De Luca Fornaciari , Maite Melero

Deep learning techniques have gained a lot of traction in the field of NLP research. The aim of this paper is to predict the age and gender of an individual by inspecting their written text. We propose a supervised BERT-based classification…

计算与语言 · 计算机科学 2023-05-16 Vishesh Thakur , Aneesh Tickoo

Deep learning (DL) models are widely used to provide a more convenient and smarter life. However, biased algorithms will negatively influence us. For instance, groups targeted by biased algorithms will feel unfairly treated and even fearful…

计算机视觉与模式识别 · 计算机科学 2021-08-24 Xuyang Shen , Jo Plested , Sabrina Caldwell , Tom Gedeon

In this paper, we pose the question: do people talk about women and men in different ways? We introduce two datasets and a novel integration of approaches for automatically inferring gender associations from language, discovering coherent…

计算与语言 · 计算机科学 2019-09-04 Serina Chang , Kathleen McKeown

Short Message Service (SMS) spam is a serious problem in Vietnam because of the availability of very cheap pre-paid SMS packages. There are some systems to detect and filter spam messages for English, most of which use machine learning…

计算与语言 · 计算机科学 2017-05-12 Thai-Hoang Pham , Phuong Le-Hong

Nowadays, Large Language Models (LLMs) have attracted widespread attention due to their powerful performance. However, due to the unavoidable exposure to socially biased data during training, LLMs tend to exhibit social biases, particularly…

计算与语言 · 计算机科学 2025-05-22 Zhanyue Qin , Yue Ding , Deyuan Liu , Qingbin Liu , Junxian Cai , Xi Chen , Zhiying Tu , Dianhui Chu , Cuiyun Gao , Dianbo Sui

Vietnamese document analysis and recognition (DAR) is a crucial field with applications in digitization, information retrieval, and automation. Despite advancements in OCR and NLP, Vietnamese text recognition faces unique challenges due to…

计算机视觉与模式识别 · 计算机科学 2025-06-06 Anh Le , Thanh Lam , Dung Nguyen

Large Language Models (LLMs) have shown remarkable capabilities in a multitude of Natural Language Processing (NLP) tasks. However, these models are still not immune to limitations such as social biases, especially gender bias. This work…

计算与语言 · 计算机科学 2024-10-15 Divij Bajaj , Yuanyuan Lei , Jonathan Tong , Ruihong Huang

Image Captioning, the task of automatic generation of image captions, has attracted attentions from researchers in many fields of computer science, being computer vision, natural language processing and machine learning in recent years.…

计算与语言 · 计算机科学 2020-02-04 Quan Hoang Lam , Quang Duy Le , Kiet Van Nguyen , Ngan Luu-Thuy Nguyen

We investigate whether it is feasible to remove gendered information from resumes to mitigate potential bias in algorithmic resume screening. Using a corpus of 709k resumes from IT firms, we first train a series of models to classify the…

计算与语言 · 计算机科学 2022-07-14 Prasanna Parasurama , João Sedoc

Vietnamese labor market has been under an imbalanced development. The number of university graduates is growing, but so is the unemployment rate. This situation is often caused by the lack of accurate and timely labor market information,…

计算与语言 · 计算机科学 2022-10-27 Viet-Trung Tran , Hai-Nam Cao , Tuan-Dung Cao

Recent advances in contextualized word embeddings have greatly improved semantic tasks such as Word Sense Disambiguation (WSD) and contextual similarity, but most progress has been limited to high-resource languages like English.…

计算与语言 · 计算机科学 2025-11-18 Khang T. Huynh , Dung H. Nguyen , Binh T. Nguyen