中文
相关论文

相关论文: Auditing for Diversity using Representative Exampl…

200 篇论文

We consider the problem of active domain adaptation (ADA) to unlabeled target data, of which subset is actively selected and labeled given a budget constraint. Inspired by recent analysis on a critical issue from label distribution mismatch…

机器学习 · 计算机科学 2022-08-16 Sehyun Hwang , Sohyun Lee , Sungyeon Kim , Jungseul Ok , Suha Kwak

Computational social science studies often contextualize content analysis within standard demographics. Since demographics are unavailable on many social media platforms (e.g. Twitter) numerous studies have inferred demographics…

计算与语言 · 计算机科学 2021-07-13 Zach Wood-Doughty , Paiheng Xu , Xiao Liu , Mark Dredze

Large language models (LLMs) can be used to generate text data for training and evaluating other models. However, creating high-quality datasets with LLMs can be challenging. In this work, we explore human-AI partnerships to facilitate high…

计算与语言 · 计算机科学 2023-08-11 John Joon Young Chung , Ece Kamar , Saleema Amershi

Finding a small set of representatives from an unlabeled dataset is a core problem in a broad range of applications such as dataset summarization and information extraction. Classical exemplar selection methods such as $k$-medoids work…

机器学习 · 计算机科学 2020-06-09 Chong You , Chi Li , Daniel P. Robinson , Rene Vidal

Unsupervised feature selection aims to identify a compact subset of features that captures the intrinsic structure of data without supervised label. Most existing studies evaluate the performance of methods using the single-label dataset…

机器学习 · 计算机科学 2026-02-10 Gyu-Il Kim , Dae-Won Kim , Jaesung Lee

As acquiring reliable ground-truth labels is usually costly, or infeasible, crowdsourcing and aggregation of noisy human annotations is the typical resort. Aggregating subjective labels, though, may amplify individual biases, particularly…

机器学习 · 计算机科学 2026-02-02 Gabriel Singer , Samuel Gruffaz , Olivier Vo Van , Nicolas Vayatis , Argyris Kalogeratos

Machine learning (ML) models often exhibit bias that can exacerbate inequities in biomedical applications. Fairness auditing, the process of evaluating a model's performance across subpopulations, is critical for identifying and mitigating…

统计方法学 · 统计学 2026-05-19 Jianhui Gao , Jessica Gronsbell

With social media datasets being increasingly shared by researchers, it also presents the caveat that those datasets are not always completely replicable. Having to adhere to requirements of platforms like Twitter, researchers cannot…

数字图书馆 · 计算机科学 2018-03-08 Arkaitz Zubiaga

Annotating data via crowdsourcing is time-consuming and expensive. Due to these costs, dataset creators often have each annotator label only a small subset of the data. This leads to sparse datasets with examples that are marked by few…

计算与语言 · 计算机科学 2023-10-06 London Lowmanstone , Ruyuan Wan , Risako Owan , Jaehyung Kim , Dongyeop Kang

The sharp increase in data-related expenses has motivated research into condensing datasets while retaining the most informative features. Dataset distillation has thus recently come to the fore. This paradigm generates synthetic datasets…

机器学习 · 计算机科学 2024-11-20 Jiawei Du , Xin Zhang , Juncheng Hu , Wenxin Huang , Joey Tianyi Zhou

Many operational AI systems depend on large-scale human annotation to detect rare but consequential events (e.g., fraud, defects, and medical abnormalities). When positives are rare, the prevalence effect induces systematic cognitive biases…

Diversity is an important consideration in the construction of robust neural network ensembles. A collection of well trained models will generalize better if they are diverse in the patterns they respond to and the predictions they make.…

机器学习 · 计算机科学 2023-02-14 Tim Whitaker , Darrell Whitley

In the last few years, Artificial Intelligence systems have become increasingly widespread. Unfortunately, these systems can share many biases with human decision-making, including demographic biases. Often, these biases can be traced back…

计算机视觉与模式识别 · 计算机科学 2024-11-27 Iris Dominguez-Catena , Daniel Paternain , Mikel Galar

Measuring similarity between two images often requires performing complex reasoning along different axes (e.g., color, texture, or shape). Insights into what might be important for measuring similarity can can be provided by annotated…

计算机视觉与模式识别 · 计算机科学 2021-08-23 Samarth Mishra , Zhongping Zhang , Yuan Shen , Ranjitha Kumar , Venkatesh Saligrama , Bryan Plummer

The learning from imbalanced data is a deeply studied problem in standard classification and, in recent times, also in multilabel classification. A handful of multilabel resampling methods have been proposed in late years, aiming to balance…

机器学习 · 计算机科学 2018-02-15 Francisco Charte , Antonio J. Rivera , María J. del Jesus , Francisco Herrera

Large scale image classification models trained on top of popular datasets such as Imagenet have shown to have a distributional skew which leads to disparities in prediction accuracies across different subsections of population…

计算机视觉与模式识别 · 计算机科学 2021-07-21 Rohan Mahadev , Anindya Chakravarti

Human-annotated data plays a critical role in the fairness of AI systems, including those that deal with life-altering decisions or moderating human-created web/social media content. Conventionally, annotator disagreements are resolved…

Most of the literature around text classification treats it as a supervised learning problem: given a corpus of labeled documents, train a classifier such that it can accurately predict the classes of unseen documents. In industry, however,…

计算与语言 · 计算机科学 2018-04-09 Katherine Bailey , Sunny Chopra

We present a method to improve the calibration of deep ensembles in the small training data regime in the presence of unlabeled data. Our approach is extremely simple to implement: given an unlabeled set, for each unlabeled data point, we…

机器学习 · 计算机科学 2023-10-05 Konstantinos Pitas , Julyan Arbel

Safe artificial intelligence for perception tasks remains a major challenge, partly due to the lack of data with high-quality labels. Annotations themselves are subject to aleatoric and epistemic uncertainty, which is typically ignored…