中文
相关论文

相关论文: Mitigating Dataset Harms Requires Stewardship: Les…

200 篇论文

We investigate the contents of web-scraped data for training AI systems, at sizes where human dataset curators and compilers no longer manually annotate every sample. Building off of prior privacy concerns in machine learning models, we…

密码学与安全 · 计算机科学 2026-04-08 Rachel Hong , Jevan Hutson , William Agnew , Imaad Huda , Tadayoshi Kohno , Jamie Morgenstern

Various face image datasets intended for facial biometrics research were created via web-scraping, i.e. the collection of images publicly available on the internet. This work presents an approach to detect both exactly and nearly identical…

计算机视觉与模式识别 · 计算机科学 2024-03-07 Torsten Schlett , Christian Rathgeb , Juan Tapia , Christoph Busch

Human-centric computer vision (HCCV) data curation practices often neglect privacy and bias concerns, leading to dataset retractions and unfair models. HCCV datasets constructed through nonconsensual web scraping lack crucial metadata for…

计算机视觉与模式识别 · 计算机科学 2023-12-12 Jerone T. A. Andrews , Dora Zhao , William Thong , Apostolos Modas , Orestis Papakyriakopoulos , Alice Xiang

Facial Expression Recognition faces two core challenges. The first is class imbalance in public datasets, which skews the learning process and weakens generalization. The second is related to privacy and data collection constraints, which…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Ali Azmoudeh , Erdi Sarıtaş , Ömer Yıldırım , Hazım Kemal Ekenel

Due to the data-driven nature of current face identity (FaceID) customization methods, all state-of-the-art models rely on large-scale datasets containing millions of high-quality text-image pairs for training. However, none of these…

计算机视觉与模式识别 · 计算机科学 2025-03-28 Shuhe Wang , Xiaoya Li , Jiwei Li , Guoyin Wang , Xiaofei Sun , Bob Zhu , Han Qiu , Mo Yu , Shengjie Shen , Tianwei Zhang , Eduard Hovy

Datasets are central to training machine learning (ML) models. The ML community has recently made significant improvements to data stewardship and documentation practices across the model development life cycle. However, the act of…

计算机与社会 · 计算机科学 2022-05-11 Alexandra Sasha Luccioni , Frances Corry , Hamsini Sridharan , Mike Ananny , Jason Schultz , Kate Crawford

Recent deep face hallucination methods show stunning performance in super-resolving severely degraded facial images, even surpassing human ability. However, these algorithms are mainly evaluated on non-public synthetic datasets. It is thus…

计算机视觉与模式识别 · 计算机科学 2022-06-09 Kaihao Zhang , Dongxu Li , Wenhan Luo , Jingyu Liu , Jiankang Deng , Wei Liu , Stefanos Zafeiriou

The deployment of facial recognition systems has created an ethical dilemma: achieving high accuracy requires massive datasets of real faces collected without consent, leading to dataset retractions and potential legal liabilities under…

In an ideal world, deployed machine learning models will enhance our society. We hope that those models will provide unbiased and ethical decisions that will benefit everyone. However, this is not always the case; issues arise during the…

计算机与社会 · 计算机科学 2021-11-25 Jasmine DeHart , Chenguang Xu , Lisa Egede , Christan Grant

To achieve good performance in face recognition, a large scale training dataset is usually required. A simple yet effective way to improve recognition performance is to use a dataset as large as possible by combining multiple datasets in…

计算机视觉与模式识别 · 计算机科学 2021-01-15 Gaoang Wang , Lin Chen , Tianqiang Liu , Mingwei He , Jiebo Luo

In difficult decision-making scenarios, it is common to have conflicting opinions among expert human decision-makers as there may not be a single right answer. Such decisions may be guided by different attributes that can be used to…

计算与语言 · 计算机科学 2024-06-11 Brian Hu , Bill Ray , Alice Leung , Amy Summerville , David Joy , Christopher Funk , Arslan Basharat

As datasets become critical assets in modern machine learning systems, ensuring robust copyright protection has emerged as an urgent challenge. Traditional legal mechanisms often fail to address the technical complexities of digital data…

密码学与安全 · 计算机科学 2025-09-09 Kun Li , Cheng Wang , Minghui Xu , Yue Zhang , Xiuzhen Cheng

Large-scale datasets have played a crucial role in the advancement of computer vision. However, they often suffer from problems such as class imbalance, noisy labels, dataset bias, or high resource costs, which can inhibit model performance…

计算机视觉与模式识别 · 计算机科学 2023-10-09 Zhijing Wan , Zhixiang Wang , CheukTing Chung , Zheng Wang

Bias in data can have unintended consequences that propagate to the design, development, and deployment of machine learning models. In the financial services sector, this can result in discrimination from certain financial instruments and…

密码学与安全 · 计算机科学 2019-11-12 Reginald Bryant , Celia Cintas , Isaac Wambugu , Andrew Kinai , Komminist Weldemariam

[Context] Generative AI technologies, particularly Large Language Models (LLMs), have transformed numerous domains by enhancing convenience and efficiency in information retrieval, content generation, and decision-making processes. However,…

计算机与社会 · 计算机科学 2025-08-25 Yutan Huang , Chetan Arora , Wen Cheng Houng , Tanjila Kanij , Anuradha Madulgalla , John Grundy

Data is a crucial component of machine learning. The field is reliant on data to train, validate, and test models. With increased technical capabilities, machine learning research has boomed in both academic and industry settings, and one…

计算机视觉与模式识别 · 计算机科学 2021-09-20 Morgan Klaus Scheuerman , Emily Denton , Alex Hanna

As new machine learning methods demand larger training datasets, researchers and developers face significant challenges in dataset management. Although ethics reviews, documentation, and checklists have been established, it remains…

机器学习 · 计算机科学 2024-11-04 Yiwei Wu , Leah Ajmani , Shayne Longpre , Hanlin Li

Face recognition applications have grown in parallel with the size of datasets, complexity of deep learning models and computational power. However, while deep learning models evolve to become more capable and computational power keeps…

计算机视觉与模式识别 · 计算机科学 2024-04-24 Pedro C. Neto , Rafael M. Mamede , Carolina Albuquerque , Tiago Gonçalves , Ana F. Sequeira

It is well known that deep learning approaches to face recognition and facial landmark detection suffer from biases in modern training datasets. In this work, we propose to use synthetic face images to reduce the negative effects of dataset…

计算机视觉与模式识别 · 计算机科学 2019-06-25 Adam Kortylewski , Bernhard Egger , Andreas Morel-Forster , Andreas Schneider , Thomas Gerig , Clemens Blumer , Corius Reyneke , Thomas Vetter

Web-scraped, in-the-wild datasets have become the norm in face recognition research. The numbers of subjects and images acquired in web-scraped datasets are usually very large, with number of images on the millions scale. A variety of…

计算机视觉与模式识别 · 计算机科学 2020-04-08 Kai Zhang , Vítor Albiero , Kevin W. Bowyer