中文
相关论文

相关论文: Dialect and Gender Bias in YouTube's Spanish Capti…

200 篇论文

We examine the possibility that recent promising results in automatic caption generation are due primarily to language models. By varying image representation quality produced by a convolutional neural network, we find that a…

计算与语言 · 计算机科学 2015-08-11 Jack Hessel , Nicolas Savva , Michael J. Wilber

Contrary to Google Search's mission of delivering information from "many angles so you can form your own understanding of the world," we find that Google and its most prominent returned results - Wikipedia and YouTube - simply reflect a…

计算机与社会 · 计算机科学 2024-03-11 Queenie Luo , Michael J. Puett , Michael D. Smith

This paper describes a web-based corpus of global language use with a focus on how this corpus can be used for data-driven language mapping. First, the corpus provides a representation of where national varieties of major languages are used…

计算与语言 · 计算机科学 2020-04-03 Jonathan Dunn

The number of user reviews of tourist attractions, restaurants, mobile apps, etc. is increasing for all languages; yet, research is lacking on how reviews in multiple languages should be aggregated and displayed. Speakers of different…

人机交互 · 计算机科学 2016-05-09 Scott A. Hale

The appearance of complex attention-based language models such as BERT, Roberta or GPT-3 has allowed to address highly complex tasks in a plethora of scenarios. However, when applied to specific domains, these models encounter considerable…

计算与语言 · 计算机科学 2022-06-14 Javier Huertas-Tato , Alejandro Martin , David Camacho

Understanding the impact of digital platforms on user behavior presents foundational challenges, including issues related to polarization, misinformation dynamics, and variation in news consumption. Comparative analyses across platforms and…

The Multi-language Video Subtitle Dataset is a comprehensive collection designed to support research in text recognition across multiple languages. This dataset includes 4,224 subtitle images extracted from 24 videos sourced from online…

计算机视觉与模式识别 · 计算机科学 2024-11-11 Thanadol Singkhornart , Olarik Surinta

When captioning an image, people describe objects in diverse ways, such as by using different terms and/or including details that are perceptually noteworthy to them. Descriptions can be especially unique across languages and cultures.…

计算机视觉与模式识别 · 计算机科学 2025-11-12 Kyle Buettner , Jacob T. Emmerson , Adriana Kovashka

Sign language translation (SLT) is an active field of study that encompasses human-computer interaction, computer vision, natural language processing and machine learning. Progress on this field could lead to higher levels of integration of…

计算机视觉与模式识别 · 计算机科学 2022-11-29 Pedro Dal Bianco , Gastón Ríos , Franco Ronchetti , Facundo Quiroga , Oscar Stanchi , Waldo Hasperué , Alejandro Rosete

The task of image captioning implicitly involves gender identification. However, due to the gender bias in data, gender identification by an image captioning model suffers. Also, the gender-activity bias, owing to the word-by-word…

计算机视觉与模式识别 · 计算机科学 2019-12-03 Shruti Bhargava , David Forsyth

Millions of people use platforms such as YouTube, Facebook, Twitter, and other mass media. Due to the accessibility of these platforms, they are often used to establish a narrative, conduct propaganda, and disseminate misinformation. This…

机器学习 · 计算机科学 2021-07-05 Raj Jagtap , Abhinav Kumar , Rahul Goel , Shakshi Sharma , Rajesh Sharma , Clint P. George

Image captioning, a fundamental task in vision-language understanding, seeks to generate accurate natural language descriptions for provided images. Current image captioning approaches heavily rely on high-quality image-caption pairs, which…

计算机视觉与模式识别 · 计算机科学 2023-11-03 Chuanyang Jin

In the dynamic realm of social media, diverse topics are discussed daily, transcending linguistic boundaries. However, the complexities of understanding and categorising this content across various languages remain an important challenge…

计算与语言 · 计算机科学 2024-10-07 Dimosthenis Antypas , Asahi Ushio , Francesco Barbieri , Jose Camacho-Collados

Image captioning has made substantial progress with huge supporting image collections sourced from the web. However, recent studies have pointed out that captioning datasets, such as COCO, contain gender bias found in web corpora. As a…

计算机视觉与模式识别 · 计算机科学 2021-04-22 Ruixiang Tang , Mengnan Du , Yuening Li , Zirui Liu , Na Zou , Xia Hu

Gender stereotypes are manifest in most of the world's languages and are consequently propagated or amplified by NLP systems. Although research has focused on mitigating gender stereotypes in English, the approaches that are commonly…

计算与语言 · 计算机科学 2020-05-28 Ran Zmigrod , Sabrina J. Mielke , Hanna Wallach , Ryan Cotterell

In recent years, large language models (LLMs) have demonstrated a high capacity for understanding and generating text in Spanish. However, with five hundred million native speakers, Spanish is not a homogeneous language but rather one rich…

计算与语言 · 计算机科学 2025-05-22 Marina Mayor-Rocher , Cristina Pozo , Nina Melero , Gonzalo Martínez , María Grandury , Pedro Reviriego

Energy production and management face significant political, economic, and environmental challenges, yet the rise in information consumption through social media undermines the availability of reliable knowledge to the general public. This…

物理与社会 · 物理学 2025-11-21 Aleix Bassolas , Piero Birello , Julian Vicens

There are many Language Models for the English language according to its worldwide relevance. However, for the Spanish language, even if it is a widely spoken language, there are very few Spanish Language Models which result to be small and…

计算与语言 · 计算机科学 2021-10-26 Asier Gutiérrez-Fandiño , Jordi Armengol-Estapé , Aitor Gonzalez-Agirre , Marta Villegas

This paper presents the participation of the MiniTrue team in the EXIST 2021 Challenge on the sexism detection in social media task for English and Spanish. Our approach combines the language models with a simple voting mechanism for the…

计算与语言 · 计算机科学 2021-06-01 Chao Feng

Neural Machine Translation models tend to perpetuate gender bias present in their training data distribution. Context-aware models have been previously suggested as a means to mitigate this type of bias. In this work, we examine this claim…

计算与语言 · 计算机科学 2024-06-19 Harritxu Gete , Thierry Etchegoyhen