English
Related papers

Related papers: Fundamentals of Task-Agnostic Data Valuation

200 papers

Assessing whether a sample survey credibly represents the population is a critical question for ensuring the validity of downstream research. Generally, this problem reduces to estimating the distance between two high-dimensional…

Machine Learning · Computer Science 2025-08-29 Debabrota Basu , Sourav Chakraborty , Debarshi Chanda , Buddha Dev Das , Arijit Ghosh , Arnab Ray

How can we assess the reliability of a dataset without access to ground truth? We introduce the problem of reliability scoring for datasets collected from potentially strategic sources. The true data are unobserved, but we see outcomes of…

Machine Learning · Computer Science 2025-10-21 Yiling Chen , Shi Feng , Paul Kattuman , Fang-Yi Yu

Personal data is essential in showing users targeted ads - the economic backbone of the web. Still, there are major inefficiencies in how data is transacted online: (1) users don't decide what information is released nor get paid for this…

Social and Information Networks · Computer Science 2019-06-17 Anish Agarwal , Munther Dahleh , Devavrat Shah , Dylan Sleeper , Andrew Tsai , Madeline Wong

We study a setting in which a data buyer seeks to estimate an unknown parameter by purchasing samples from one of K data sellers. Each seller has privately known data quality (e.g., high vs. low variance) and a private per-sample cost. We…

Computer Science and Game Theory · Computer Science 2026-02-20 Nivasini Ananthakrishnan , Alireza Fallah , Michael I. Jordan

We formalise the essential data of objective functions as equality constraints on composites of learners. We call these constraints "tasks", and we investigate the idealised view that such tasks determine model behaviours. We develop a…

Machine Learning · Computer Science 2025-05-06 Benjamin Rodatz , Ian Fan , Tuomas Laakkonen , Neil John Ortega , Thomas Hoffmann , Vincent Wang-Mascianica

One of the most effective approaches to improving the performance of a machine learning model is to procure additional training data. A model owner seeking relevant training data from a data owner needs to appraise the data before acquiring…

Machine Learning · Computer Science 2022-03-15 Mimee Xu , Laurens van der Maaten , Awni Hannun

We consider the problem of purchasing data for machine learning or statistical estimation. The data analyst has a budget to purchase datasets from multiple data providers. She does not have any test data that can be used to evaluate the…

Computer Science and Game Theory · Computer Science 2020-10-30 Yiling Chen , Yiheng Shen , Shuran Zheng

Data values in a dataset can be missing or anomalous due to mishandling or human error. Analysing data with missing values can create bias and affect the inferences. Several analysis methods, such as principle components analysis or…

Artificial Intelligence · Computer Science 2022-05-11 Sandeep Hans , Diptikalyan Saha , Aniya Aggarwal

Data valuation plays a crucial role in machine learning. Existing data valuation methods, mainly focused on discriminative models, overlook generative models that have gained attention recently. In generative models, data valuation measures…

Computer Vision and Pattern Recognition · Computer Science 2025-04-15 Jiaxi Yang , Wenglong Deng , Benlin Liu , Yangsibo Huang , James Zou , Xiaoxiao Li

We study statistical parameter estimation in the setting of data markets. A buyer seeks to estimate a parameter based on samples that can be purchased from competing providers that differ in their data quality and provision costs. When…

Computer Science and Game Theory · Computer Science 2026-04-13 Yuchen Hu , Martin J. Wainwright , Stephen Bates

Data subsampling is widely used to speed up the training of large-scale recommendation systems. Most subsampling methods are model-based and often require a pre-trained pilot model to measure data importance via e.g. sample hardness.…

Information Retrieval · Computer Science 2023-06-19 Xiaohui Chen , Jiankai Sun , Taiqing Wang , Ruocheng Guo , Li-Ping Liu , Aonan Zhang

Traditional learning approaches for classification implicitly assume that each mistake has the same cost. In many real-world problems though, the utility of a decision depends on the underlying context $x$ and decision $y$. However,…

Machine Learning · Computer Science 2021-04-20 Kush Bhatia , Peter L. Bartlett , Anca D. Dragan , Jacob Steinhardt

A data marketplace is an online venue that brings data owners, data brokers, and data consumers together and facilitates commoditisation of data amongst them. Data pricing, as a key function of a data marketplace, demands quantifying the…

Computer Science and Game Theory · Computer Science 2023-03-10 Mengxiao Zhang , Fernando Beltran , Jiamou Liu

It has been reported that deep learning models are extremely vulnerable to small but intentionally chosen perturbations of its input. In particular, a deep network, despite its near-optimal accuracy on the clean images, often mis-classifies…

Machine Learning · Computer Science 2022-03-16 A. Tuan Nguyen , Ser Nam Lim , Philip Torr

Many recent works on understanding deep learning try to quantify how much individual data instances influence the optimization and generalization of a model. Such attempts reveal characteristics and importance of individual instances, which…

Machine Learning · Computer Science 2023-03-08 Nohyun Ki , Hoyong Choi , Hye Won Chung

A challenge that data analysts face is building a data analysis that is useful for a given consumer. Previously, we defined a set of principles for describing data analyses that can be used to create a data analysis and to characterize the…

Methodology · Statistics 2023-12-14 Lucy D'Agostino McGowan , Roger D. Peng , Stephanie C. Hicks

Missing data, the data value that is not recorded for a variable, occurs in almost all statistical analyses and may be caused by many reasons, such as lack of collection or a lack of documentation. Researchers need to adequately deal with…

Human-Computer Interaction · Computer Science 2024-10-08 Sarah Alsufyani , Matthew Forshaw , Sara Johansson Fernstad

We present a system to support generalized SQL workload analysis and management for multi-tenant and multi-database platforms. Workload analysis applications are becoming more sophisticated to support database administration, model user…

Databases · Computer Science 2018-08-28 Shrainik Jain , Jiaqi Yan , Thierry Cruane , Bill Howe

Quantifying the value of data is a fundamental problem in machine learning. Data valuation has multiple important use cases: (1) building insights about the learning task, (2) domain adaptation, (3) corrupted sample discovery, and (4)…

Machine Learning · Computer Science 2019-09-27 Jinsung Yoon , Sercan O. Arik , Tomas Pfister

Authentication is the task of confirming the matching relationship between a data instance and a given identity. Typical examples of authentication problems include face recognition and person re-identification. Data-driven authentication…

Machine Learning · Statistics 2020-11-24 Jian Liang , Yuren Cao , Shuang Li , Bing Bai , Hao Li , Fei Wang , Kun Bai