中文
相关论文

相关论文: Uncurated Image-Text Datasets: Shedding Light on D…

200 篇论文

We introduce a dataset of concept learning tasks that helps uncover implicit biases in large language models. Using in-context concept learning experiments, we found that language models may have a bias toward upward monotonicity in…

计算与语言 · 计算机科学 2025-11-27 Leroy Z. Wang

Build accurate DNN models requires training on large labeled, context specific datasets, especially those matching the target scenario. We believe advances in wireless localization, working in unison with cameras, can produce automated…

计算机视觉与模式识别 · 计算机科学 2018-10-11 Zhujun Xiao , Yanzi Zhu , Yuxin Chen , Ben Y. Zhao , Junchen Jiang , Haitao Zheng

With the recent proliferation of the use of text classifications, researchers have found that there are certain unintended biases in text classification datasets. For example, texts containing some demographic identity-terms (e.g., "gay",…

计算与语言 · 计算机科学 2020-08-21 Guanhua Zhang , Bing Bai , Junqi Zhang , Kun Bai , Conghui Zhu , Tiejun Zhao

In recent years, the rapid development of artificial intelligence (AI) systems has raised concerns about our ability to ensure their fairness, that is, how to avoid discrimination based on protected characteristics such as gender, race, or…

计算机视觉与模式识别 · 计算机科学 2025-03-06 Iris Dominguez-Catena , Daniel Paternain , Mikel Galar , MaryBeth Defrance , Maarten Buyl , Tijl De Bie

The race to develop image generation models is intensifying, with a rapid increase in the number of text-to-image models available. This is coupled with growing public awareness of these technologies. Though other generative AI…

计算机与社会 · 计算机科学 2024-07-24 Amelia Katirai , Noa Garcia , Kazuki Ide , Yuta Nakashima , Atsuo Kishimoto

This paper aims to shed light on the ethical problems of creating and deploying computer vision tech, particularly in using publicly available datasets. Due to the rapid growth of machine learning and artificial intelligence, computer…

计算机视觉与模式识别 · 计算机科学 2025-08-12 Ghalib Ahmed Tahir

Most machine learning methods are known to capture and exploit biases of the training data. While some biases are beneficial for learning, others are harmful. Specifically, image captioning models tend to exaggerate biases present in…

计算机视觉与模式识别 · 计算机科学 2019-03-15 Kaylee Burns , Lisa Anne Hendricks , Kate Saenko , Trevor Darrell , Anna Rohrbach

We have now entered the era of trillion parameter machine learning models trained on billion-sized datasets scraped from the internet. The rise of these gargantuan datasets has given rise to formidable bodies of critical work that has…

计算机与社会 · 计算机科学 2021-10-06 Abeba Birhane , Vinay Uday Prabhu , Emmanuel Kahembwe

This paper collects a set of open research questions on how to visualize sociodemographic data. Sociodemographic data is a common part of datasets related to people, including institutional censuses, health data systems, and human-resources…

人机交互 · 计算机科学 2023-08-24 Florent Cabric , Margrét Vilborg Bjarnadóttir , Anne-Flore Cabouat , Petra Isenberg

Image-classification datasets have been used to pretrain image recognition models. Recently, web-scale image-caption datasets have emerged as a source of powerful pretraining alternative. Image-caption datasets are more ``open-domain'',…

计算机视觉与模式识别 · 计算机科学 2023-05-17 Kuniaki Saito , Kihyuk Sohn , Xiang Zhang , Chun-Liang Li , Chen-Yu Lee , Kate Saenko , Tomas Pfister

Automatically describing images using natural sentences is an important task to support visually impaired people's inclusion onto the Internet. It is still a big challenge that requires understanding the relation of the objects present in…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Gabriel Oliveira dos Santos , Esther Luna Colombini , Sandra Avila

Generative AI models have recently achieved astonishing results in quality and are consequently employed in a fast-growing number of applications. However, since they are highly data-driven, relying on billion-sized datasets randomly…

Recent works in image captioning have shown very promising raw performance. However, we realize that most of these encoder-decoder style networks with attention do not scale naturally to large vocabulary size, making them difficult to be…

计算机视觉与模式识别 · 计算机科学 2019-06-13 Jia Huei Tan , Chee Seng Chan , Joon Huang Chuah

Text-to-image generative models are becoming increasingly popular and accessible to the general public. As these models see large-scale deployments, it is necessary to deeply investigate their safety and fairness to not disseminate and…

计算机视觉与模式识别 · 计算机科学 2024-08-06 Moreno D'Incà , Elia Peruzzo , Massimiliano Mancini , Dejia Xu , Vidit Goel , Xingqian Xu , Zhangyang Wang , Humphrey Shi , Nicu Sebe

Despite their prevalence in society, social biases are difficult to identify, primarily because human judgements in this domain can be unreliable. We take an unsupervised approach to identifying gender bias against women at a comment level…

计算与语言 · 计算机科学 2020-10-07 Anjalie Field , Yulia Tsvetkov

Vision and vision-language applications of neural networks, such as image classification and captioning, rely on large-scale annotated datasets that require non-trivial data-collecting processes. This time-consuming endeavor hinders the…

Social bias in machine learning has drawn significant attention, with work ranging from demonstrations of bias in a multitude of applications, curating definitions of fairness for different contexts, to developing algorithms to mitigate…

计算与语言 · 计算机科学 2019-11-06 Yi Chern Tan , L. Elisa Celis

Attention mechanisms have attracted considerable interest in image captioning because of its powerful performance. Existing attention-based models use feedback information from the caption generator as guidance to determine which of the…

计算机视觉与模式识别 · 计算机科学 2018-07-11 Zhihao Zhu , Zhan Xue , Zejian Yuan

Automated computer vision systems have been applied in many domains including security, law enforcement, and personal devices, but recent reports suggest that these systems may produce biased results, discriminating against people in…

计算机视觉与模式识别 · 计算机科学 2020-05-22 Jungseock Joo , Kimmo Kärkkäinen

Language models have been shown to reproduce underlying biases existing in their training data, which is the majority perspective by default. Proposed solutions aim to capture minority perspectives by either modelling annotator…

计算与语言 · 计算机科学 2024-07-22 Nikolas Vitsakis , Amit Parekh , Ioannis Konstas