中文
相关论文

相关论文: Gender Prediction Based on Vietnamese Names with M…

200 篇论文

This study investigates gender bias in large language models (LLMs) by comparing their gender perception to that of human respondents, U.S. Bureau of Labor Statistics data, and a 50% no-bias benchmark. We created a new evaluation set using…

计算与语言 · 计算机科学 2024-11-22 Tetiana Bas

Universities face surging applications and heightened expectations for fairness, making accurate admission prediction increasingly vital. This work presents a comprehensive framework that fuses machine learning, deep learning, and large…

计算机与社会 · 计算机科学 2025-09-29 Mohammad Abbadi , Yassine Himeur , Shadi Atalla , Dahlia Mansoor , Wathiq Mansoor

Large language models increasingly support multiple languages, yet most benchmarks for gender bias remain English-centric. We introduce EuroGEST, a dataset designed to measure gender-stereotypical reasoning in LLMs across English and 29…

计算与语言 · 计算机科学 2026-02-24 Jacqueline Rowe , Mateusz Klimaszewski , Liane Guillou , Shannon Vallor , Alexandra Birch

Large Language Models (LLMs) have demonstrated remarkable proficiency in general medical domains. However, their performance significantly degrades in specialized, culturally specific domains such as Vietnamese Traditional Medicine (VTM),…

计算与语言 · 计算机科学 2026-01-08 Huynh Trung Kiet , Dao Sy Duy Minh , Nguyen Dinh Ha Duong , Le Hoang Minh Huy , Long Nguyen , Dien Dinh

The current COVID-19 pandemic has lead to the creation of many corpora that facilitate NLP research and downstream applications to help fight the pandemic. However, most of these corpora are exclusively for English. As the pandemic is a…

计算与语言 · 计算机科学 2021-04-09 Thinh Hung Truong , Mai Hoang Dao , Dat Quoc Nguyen

Visual Question Answering (VQA) is an intricate and demanding task that integrates natural language processing (NLP) and computer vision (CV), capturing the interest of researchers. The English language, renowned for its wealth of…

计算与语言 · 计算机科学 2023-07-31 Khiem Vinh Tran , Kiet Van Nguyen , Ngan Luu Thuy Nguyen

This work explores the biases in learning processes based on deep neural network architectures. We analyze how bias affects deep learning processes through a toy example using the MNIST database and a case study in gender detection from…

计算机视觉与模式识别 · 计算机科学 2020-07-23 Ignacio Serna , Alejandro Peña , Aythami Morales , Julian Fierrez

This paper presents a new method for automatically detecting words with lexical gender in large-scale language datasets. Currently, the evaluation of gender bias in natural language processing relies on manually compiled lexicons of…

计算与语言 · 计算机科学 2022-06-29 Marion Bartl , Susan Leavy

In this paper, we propose a span labeling approach to model n-gram information for Vietnamese word segmentation, namely SPAN SEG. We compare the span labeling approach with the conditional random field by using encoders with the same…

计算与语言 · 计算机科学 2021-10-04 Duc-Vu Nguyen , Linh-Bao Vo , Dang Van Thin , Ngan Luu-Thuy Nguyen

Large Language Models (LLMs) excel in Natural Language Processing (NLP) tasks, but they often propagate biases embedded in their training data, which is potentially impactful in sensitive domains like healthcare. While existing benchmarks…

计算与语言 · 计算机科学 2026-03-11 Trung Hieu Ngo , Adrien Bazoge , Solen Quiniou , Pierre-Antoine Gourraud , Emmanuel Morin

We study the presence of heteronormative biases and prejudice against interracial romantic relationships in large language models by performing controlled name-replacement experiments for the task of relationship prediction. We show that…

计算与语言 · 计算机科学 2024-10-08 Abhilasha Sancheti , Haozhe An , Rachel Rudinger

In this article, we introduce ViLegalNLI, the first large-scale Vietnamese Natural Language Inference (NLI) dataset specifically constructed for the legal domain. The dataset consists of 42,012 premise-hypothesis pairs derived from official…

计算与语言 · 计算机科学 2026-05-04 Nhung Thi-Hong Duong , Mai Ngoc Ho , Tin Van Huynh , Kiet Van Nguyen

In recent years, generative artificial intelligence (GenAI) systems have assumed increasingly crucial roles in selection processes, personnel recruitment and analysis of candidates' profiles. However, the employment of large language models…

人工智能 · 计算机科学 2026-03-13 Martina Ullasci , Marco Rondina , Riccardo Coppola , Antonio Vetrò

In recent years, Vietnam witnesses the mass development of social network users on different social platforms such as Facebook, Youtube, Instagram, and Tiktok. On social medias, hate speech has become a critical problem for social network…

计算与语言 · 计算机科学 2021-07-21 Son T. Luu , Kiet Van Nguyen , Ngan Luu-Thuy Nguyen

Existing Vietnamese Natural Language Inference (NLI) datasets lack adversarial complexity, limiting their ability to evaluate model robustness against challenging linguistic phenomena. In this article, we address the gap in robust…

计算与语言 · 计算机科学 2025-10-24 Tin Van Huynh , Kiet Van Nguyen , Ngan Luu-Thuy Nguyen

The successful application of neural methods to machine translation has realized huge quality advances for the community. With these improvements, many have noted outstanding challenges, including the modeling and treatment of gendered…

计算与语言 · 计算机科学 2020-10-16 Hila Gonen , Kellie Webster

Visual Question Answering (VQA) is a challenging task that requires the joint understanding of natural language and visual content. While early research primarily focused on recognizing objects and scene context, it often overlooked scene…

Visual Question Answering (VQA) has recently emerged as a potential research domain, captivating the interest of many in the field of artificial intelligence and computer vision. Despite the prevalence of approaches in English, there is a…

计算机视觉与模式识别 · 计算机科学 2024-08-01 Ngoc Son Nguyen , Van Son Nguyen , Tung Le

Image captioning is a crucial task with applications in a wide range of domains, including healthcare and education. Despite extensive research on English image captioning datasets, the availability of such datasets for Vietnamese remains…

计算机视觉与模式识别 · 计算机科学 2024-01-17 Anh-Cuong Pham , Van-Quang Nguyen , Thi-Hong Vuong , Quang-Thuy Ha

In this study, we present a novel and challenging multilabel Vietnamese dataset (RMDM) designed to assess the performance of large language models (LLMs), in verifying electronic information related to legal contexts, focusing on fake news…

计算与语言 · 计算机科学 2023-09-19 Hai-Long Nguyen , Thi-Kieu-Trang Pham , Thai-Son Le , Tan-Minh Nguyen , Thi-Hai-Yen Vuong , Ha-Thanh Nguyen