中文
相关论文

相关论文: Uncurated Image-Text Datasets: Shedding Light on D…

200 篇论文

The presence of a bias in each image data collection has recently attracted a lot of attention in the computer vision community showing the limits in generalization of any learning method trained on a specific dataset. At the same time,…

计算机视觉与模式识别 · 计算机科学 2015-05-07 Tatiana Tommasi , Novi Patricia , Barbara Caputo , Tinne Tuytelaars

Computer vision applications like automated face detection are used for a variety of purposes ranging from unlocking smart devices to tracking potential persons of interest for surveillance. Audits of these applications have revealed that…

计算机视觉与模式识别 · 计算机科学 2021-11-18 Siddharth D Jaiswal , Karthikeya Duggirala , Abhisek Dash , Animesh Mukherjee

We propose the ambiguity problem for the foreground object segmentation task and motivate the importance of estimating and accounting for this ambiguity when designing vision systems. Specifically, we distinguish between images which lead…

计算机视觉与模式识别 · 计算机科学 2017-05-02 Danna Gurari , Kun He , Bo Xiong , Jianming Zhang , Mehrnoosh Sameki , Suyog Dutt Jain , Stan Sclaroff , Margrit Betke , Kristen Grauman

Societal bias towards certain communities is a big problem that affects a lot of machine learning systems. This work aims at addressing the racial bias present in many modern gender recognition systems. We learn race invariant…

机器学习 · 计算机科学 2019-11-21 Komal K. Teru , Aishik Chakraborty

Constructing an organized dataset comprised of a large number of images and several captions for each image is a laborious task, which requires vast human effort. On the other hand, collecting a large number of images and sentences…

计算机视觉与模式识别 · 计算机科学 2019-11-22 Dong-Jin Kim , Jinsoo Choi , Tae-Hyun Oh , In So Kweon

Gender classification systems often inherit and amplify demographic imbalances in their training data. We first audit five widely used gender classification datasets, revealing that all suffer from significant intersectional…

计算机视觉与模式识别 · 计算机科学 2026-01-23 Tadesse K Bahiru , Natnael Tilahun Sinshaw , Teshager Hailemariam Moges , Dheeraj Kumar Singh

Information availability affects people's behavior and perception of the world. Notably, people rely on search engines to satisfy their need for information. Search engines deliver results relevant to user requests usually without being or…

信息检索 · 计算机科学 2021-10-19 Aldo Lipani , Florina Piroi , Emine Yilmaz

It has been shown that accurate representation in media improves the well-being of the people who consume it. By contrast, inaccurate representations can negatively affect viewers and lead to harmful perceptions of other cultures. To…

计算机视觉与模式识别 · 计算机科学 2023-04-27 Zhixuan Liu , Youeun Shin , Beverley-Claire Okogwu , Youngsik Yun , Lia Coleman , Peter Schaldenbrand , Jihie Kim , Jean Oh

Recent studies have shown that generative language models often reflect and amplify societal biases in their outputs. However, these studies frequently conflate observed biases with other task-specific shortcomings, such as comprehension…

计算与语言 · 计算机科学 2024-12-17 Akshita Jha , Sanchit Kabra , Chandan K. Reddy

Applications based on Machine Learning models have now become an indispensable part of the everyday life and the professional world. A critical question then recently arised among the population: Do algorithmic decisions convey any type of…

By supporting multi-modal retrieval training and evaluation, image captioning datasets have spurred remarkable progress on representation learning. Unfortunately, datasets have limited cross-modal associations: images are not paired with…

计算与语言 · 计算机科学 2021-03-25 Zarana Parekh , Jason Baldridge , Daniel Cer , Austin Waters , Yinfei Yang

Machine learning (ML) datasets, often perceived as neutral, inherently encapsulate abstract and disputed social constructs. Dataset curators frequently employ value-laden terms such as diversity, bias, and quality to characterize datasets.…

机器学习 · 计算机科学 2024-07-12 Dora Zhao , Jerone T. A. Andrews , Orestis Papakyriakopoulos , Alice Xiang

In this paper we investigate problematic practices and consequences of large scale vision datasets. We examine broad issues such as the question of consent and justice as well as specific concerns such as the inclusion of verifiably…

计算机与社会 · 计算机科学 2020-07-27 Vinay Uday Prabhu , Abeba Birhane

Identifying and mitigating bias in deep learning algorithms has gained significant popularity in the past few years due to its impact on the society. Researchers argue that models trained on balanced datasets with good representation…

计算机视觉与模式识别 · 计算机科学 2021-08-17 Puspita Majumdar , Surbhi Mittal , Richa Singh , Mayank Vatsa

State-of-the-art approaches for image captioning require supervised training data consisting of captions with paired image data. These methods are typically unable to use unsupervised data such as textual data with no corresponding images,…

计算机视觉与模式识别 · 计算机科学 2017-06-27 Wenhu Chen , Aurelien Lucchi , Thomas Hofmann

Large datasets of paired images and text have become increasingly popular for learning generic representations for vision and vision-and-language tasks. Such datasets have been built by querying search engines or collecting HTML alt-text --…

计算机视觉与模式识别 · 计算机科学 2021-11-23 Karan Desai , Gaurav Kaul , Zubin Aysola , Justin Johnson

We introduce a new large-scale dataset that links the assessment of image quality issues to two practical vision tasks: image captioning and visual question answering. First, we identify for 39,181 images taken by people who are blind…

计算机视觉与模式识别 · 计算机科学 2020-03-31 Tai-Yin Chiu , Yinan Zhao , Danna Gurari

Reference texts such as encyclopedias and news articles can manifest biased language when objective reporting is substituted by subjective writing. Existing methods to detect bias mostly rely on annotated data to train machine learning…

计算与语言 · 计算机科学 2021-12-20 Timo Spinde , David Krieger , Manuel Plank , Bela Gipp

Automatically generating descriptive captions for images is a well-researched area in computer vision. However, existing evaluation approaches focus on measuring the similarity between two sentences disregarding fine-grained semantics of…

计算机视觉与模式识别 · 计算机科学 2019-08-07 Philipp Harzig , Dan Zecha , Rainer Lienhart , Carolin Kaiser , René Schallner

We investigate the potential for nationality biases in natural language processing (NLP) models using human evaluation methods. Biased NLP models can perpetuate stereotypes and lead to algorithmic discrimination, posing a significant…