English
Related papers

Related papers: Fundamentals of Task-Agnostic Data Valuation

200 papers

Estimating causal effects from observational data requires identifying valid adjustment sets. This task is especially challenging in realistic settings where latent confounding and feedback loops are present. Existing approaches typically…

Machine Learning · Computer Science 2026-05-08 Ana Leticia Garcez Vicente , Gijs van Seeventer , Saber Salehkaleybar

Since there is, in principle, no reason why third parties should not pay individuals for the use of their data, we introduce a realistic market that would allow these payments to be made while taking into account the privacy attitude of the…

Computers and Society · Computer Science 2012-05-02 Christina Aperjis , Bernardo A. Huberman

Econometricians have usefully separated study of estimation into identification and statistical components. Identification analysis, which assumes knowledge of the probability distribution generating observable data, places an upper bound…

Econometrics · Economics 2025-09-03 Charles F. Manski

Ranking evaluation metrics are a fundamental element of design and improvement efforts in information retrieval. We observe that most popular metrics disregard information portrayed in the scores used to derive rankings, when available.…

Information Retrieval · Computer Science 2016-12-20 Nuno Moniz , Luís Torgo , João Vinagre

Continuously learning to solve unseen tasks with limited experience has been extensively pursued in meta-learning and continual learning, but with restricted assumptions such as accessible task distributions, independently and identically…

Machine Learning · Computer Science 2020-12-01 Mengdi Xu , Wenhao Ding , Jiacheng Zhu , Zuxin Liu , Baiming Chen , Ding Zhao

Given well-shuffled data, can we determine whether the data items are statistically (in)dependent? Formally, we consider the problem of testing whether a set of exchangeable random variables are independent. We will show that this is…

Statistics Theory · Mathematics 2022-10-25 Marcus Hutter

Extraction of missing attribute values is to find values describing an attribute of interest from a free text input. Most past related work on extraction of missing attribute values work with a closed world assumption with the possible set…

Computation and Language · Computer Science 2018-10-09 Guineng Zheng , Subhabrata Mukherjee , Xin Luna Dong , Feifei Li

Missing data arises when certain values are not recorded or observed for variables of interest. However, most of the statistical theory assume complete data availability. To address incomplete databases, one approach is to fill the gaps…

Health economic evaluations face the issues of non-compliance and missing data. Here, non-compliance is defined as non-adherence to a specific treatment, and occurs within randomised controlled trials (RCTs) when participants depart from…

Applications · Statistics 2019-02-26 Karla DiazOrdaz , Richard Grieve

As large language models increasingly rely on external data sources, compensating data contributors has become a central concern. But how should these payments be devised? We revisit data valuations from a $\textit{market-design…

Computer Science and Game Theory · Computer Science 2025-09-29 Dongyang Fan , Tyler J. Rotello , Sai Praneeth Karimireddy

We consider a model of unreliable or crowdsourced data where there is an underlying set of $n$ binary variables, each evaluator contributes a (possibly unreliable or adversarial) estimate of the values of some subset of $r$ of the…

Machine Learning · Computer Science 2017-08-10 Michela Meister , Gregory Valiant

Data valuation is concerned with determining a fair valuation of data from data sources to compensate them or to identify training examples that are the most or least useful for predictions. With the rising interest in personal data…

Machine Learning · Computer Science 2024-01-23 Xiao Tian , Rachael Hwee Ling Sim , Jue Fan , Bryan Kian Hsiang Low

Managers often believe that collecting more data will continually improve the accuracy of their machine learning models. However, we argue in this paper that when data lose relevance over time, it may be optimal to collect a limited amount…

Machine Learning · Computer Science 2022-03-18 Ehsan Valavi , Joel Hestness , Newsha Ardalani , Marco Iansiti

To increase autonomy in reinforcement learning, agents need to learn useful behaviours without reliance on manually designed reward functions. To that end, skill discovery methods have been used to learn the intrinsic options available to…

Artificial Intelligence · Computer Science 2021-08-05 Even Klemsdal , Sverre Herland , Abdulmajid Murad

We consider a data analyst's problem of purchasing data from strategic agents to compute an unbiased estimate of a statistic of interest. Agents incur private costs to reveal their data and the costs can be arbitrarily correlated with their…

Computer Science and Game Theory · Computer Science 2018-09-06 Yiling Chen , Nicole Immorlica , Brendan Lucier , Vasilis Syrgkanis , Juba Ziani

As data becomes the fuel driving technological and economic growth, a fundamental challenge is how to quantify the value of data in algorithmic predictions and decisions. For example, in healthcare and consumer markets, it has been…

Machine Learning · Statistics 2019-06-11 Amirata Ghorbani , James Zou

Eliciting relevance judgments for ranking evaluation is labor-intensive and costly, motivating careful selection of which documents to judge. Unlike traditional approaches that make this selection deterministically, probabilistic sampling…

Information Retrieval · Computer Science 2016-04-26 Tobias Schnabel , Adith Swaminathan , Peter Frazier , Thorsten Joachims

We propose data thinning, an approach for splitting an observation into two or more independent parts that sum to the original observation, and that follow the same distribution as the original observation, up to a (known) scaling of a…

Methodology · Statistics 2023-11-22 Anna Neufeld , Ameer Dharamshi , Lucy L. Gao , Daniela Witten

We aim to understand the value of additional labeled or unlabeled target data in transfer learning, for any given amount of source data; this is motivated by practical questions around minimizing sampling costs, whereby, target data is…

Machine Learning · Computer Science 2020-02-13 Steve Hanneke , Samory Kpotufe

Missing data theory deals with the statistical methods in the occurrence of missing data. Missing data occurs when some values are not stored or observed for variables of interest. However, most of the statistical theory assumes that data…

‹ Prev 1 4 5 6 7 8 10 Next ›