English
Related papers

Related papers: Gender Prediction Based on Vietnamese Names with M…

200 papers

The rapid development research of Large Language Models (LLMs) based on transformer architectures raises key challenges, one of them being the task of distinguishing between human-written text and LLM-generated text. As LLM-generated…

Computation and Language · Computer Science 2025-10-01 Trieu Hai Nguyen , Sivaswamy Akilesh

Masked Language Models (MLMs) pre-trained by predicting masked tokens on large corpora have been used successfully in natural language processing tasks for a variety of languages. Unfortunately, it was reported that MLMs also learn…

Computation and Language · Computer Science 2022-05-05 Masahiro Kaneko , Aizhan Imankulova , Danushka Bollegala , Naoaki Okazaki

In this paper, we present a feature-based named-entity recognition (NER) model that achieves the start-of-the-art accuracy for Vietnamese language. We combine word, word-shape features, PoS, chunk, Brown-cluster-based features, and…

Computation and Language · Computer Science 2018-03-13 Pham Quang Nhat Minh

The rapid advancement of information and communication technology has facilitated easier access to information. However, this progress has also necessitated more stringent verification measures to ensure the accuracy of information,…

Computation and Language · Computer Science 2025-03-04 Bao Tran , T. N. Khanh , Khang Nguyen Tuong , Thien Dang , Quang Nguyen , Nguyen T. Thinh , Vo T. Hung

Large Language Models (LLMs) are increasingly leveraged for translation tasks but often fall short when translating inclusive language -- such as texts containing the singular 'they' pronoun or otherwise reflecting fair linguistic…

Computation and Language · Computer Science 2025-05-06 Fanny Jourdan , Yannick Chevalier , Cécile Favre

Sentiment analysis is one of the most crucial tasks in Natural Language Processing (NLP), involving the training of machine learning models to classify text based on the polarity of opinions. Pre-trained Language Models (PLMs) can be…

Computation and Language · Computer Science 2025-01-16 Hong-Viet Tran , Van-Tan Bui , Lam-Quan Tran

VietJobs is the first large-scale, publicly available corpus of Vietnamese job advertisements, comprising 48,092 postings and over 15 million words collected from all 34 provinces and municipalities across Vietnam. The dataset provides…

Computation and Language · Computer Science 2026-03-06 Hieu Pham Dinh , Hung Nguyen Huy , Mo El-Haj

The development of natural language processing (NLP) in general and machine reading comprehension in particular has attracted the great attention of the research community. In recent years, there are a few datasets for machine reading…

Computation and Language · Computer Science 2021-06-14 Phong Nguyen-Thuan Do , Nhat Duy Nguyen , Tin Van Huynh , Kiet Van Nguyen , Anh Gia-Tuan Nguyen , Ngan Luu-Thuy Nguyen

We present GEST -- a new manually created dataset designed to measure gender-stereotypical reasoning in language models and machine translation systems. GEST contains samples for 16 gender stereotypes about men and women (e.g., Women are…

Computation and Language · Computer Science 2024-10-02 Matúš Pikuliak , Andrea Hrckova , Stefan Oresko , Marián Šimko

The rapid spread of information in the digital age highlights the critical need for effective fact-checking tools, particularly for languages with limited resources, such as Vietnamese. In response to this challenge, we introduce…

Computation and Language · Computer Science 2024-12-23 Tran Thai Hoa , Tran Quang Duy , Khanh Quoc Tran , Kiet Van Nguyen

Predictive algorithms have a powerful potential to offer benefits in areas as varied as medicine or education. However, these algorithms and the data they use are built by humans, consequently, they can inherit the bias and prejudices…

Human-Computer Interaction · Computer Science 2022-03-22 Cristina Manresa-Yee , Silvia Ramis

Large language models (LLMs) are the foundation of the current successes of artificial intelligence (AI), however, they are unavoidably biased. To effectively communicate the risks and encourage mitigation efforts these models need adequate…

Computation and Language · Computer Science 2025-01-14 Carolin M. Schuster , Maria-Alexandra Dinisor , Shashwat Ghatiwala , Georg Groh

Predicting nationality from personal names has practical value in marketing, demographic research, and genealogical studies. Conventional neural models learn statistical correspondences between names and nationalities from task-specific…

Computation and Language · Computer Science 2026-01-21 Keito Inoshita

Large language models (LLMs) often inherit and amplify social biases embedded in their training data. A prominent social bias is gender bias. In this regard, prior work has mainly focused on gender stereotyping bias - the association of…

Computation and Language · Computer Science 2025-06-18 Erik Derner , Sara Sansalvador de la Fuente , Yoan Gutiérrez , Paloma Moreda , Nuria Oliver

Student's feedback is an important source of collecting students' opinions to improve the quality of training activities. Implementing sentiment analysis into student feedback data, we can determine sentiments polarities which express all…

Computation and Language · Computer Science 2019-11-19 Phu X. V. Nguyen , Tham T. T. Hong , Kiet Van Nguyen , Ngan Luu-Thuy Nguyen

With the rapid development of large language models (LLMs), they have significantly improved efficiency across a wide range of domains. However, recent studies have revealed that LLMs often exhibit gender bias, leading to serious social…

Computation and Language · Computer Science 2025-06-17 Xiaoqing Cheng , Hongying Zan , Lulu Kong , Jinwang Song , Min Peng

Recent studies of gender bias in computing use large datasets involving automatic predictions of gender to analyze computing publications, conferences, and other key populations. Gender bias is partly defined by software-driven algorithmic…

Computers and Society · Computer Science 2022-10-18 Thomas J. Misa

Due to privacy restrictions, there's a shortage of publicly available speech recognition datasets in the medical domain. In this work, we present VietMed - a Vietnamese speech recognition dataset in the medical domain comprising 16h of…

Computation and Language · Computer Science 2025-04-07 Khai Le-Duc

Artificial Intelligence has the capacity to amplify and perpetuate societal biases and presents profound ethical implications for society. Gender bias has been identified in the context of employment advertising and recruitment tools, due…

Computation and Language · Computer Science 2020-05-19 Susan Leavy , Gerardine Meaney , Karen Wade , Derek Greene

Human gender classification based on biometric features is a major concern for computer vision due to its vast variety of applications. The human ear is popular among researchers as a soft biometric trait, because it is less affected by age…

Computer Vision and Pattern Recognition · Computer Science 2023-08-21 Ritwiz Singh , Keshav Kashyap , Rajesh Mukherjee , Asish Bera , Mamata Dalui Chakraborty