English
Related papers

Related papers: Data Distribution Valuation

200 papers

Machine learning (ML) datasets, often perceived as neutral, inherently encapsulate abstract and disputed social constructs. Dataset curators frequently employ value-laden terms such as diversity, bias, and quality to characterize datasets.…

Machine Learning · Computer Science 2024-07-12 Dora Zhao , Jerone T. A. Andrews , Orestis Papakyriakopoulos , Alice Xiang

We study the problem of imputing missing values in a dataset, which has important applications in many domains. The key to missing value imputation is to capture the data distribution with incomplete samples and impute the missing values…

Machine Learning · Computer Science 2023-06-26 He Zhao , Ke Sun , Amir Dezfouli , Edwin Bonilla

For data pricing, data quality is a factor that must be considered. To keep the fairness of data market from the aspect of data quality, we proposed a fair data market that considers data quality while pricing. To ensure fairness, we first…

Databases · Computer Science 2018-08-07 Dan Zhang , Hongzhi Wang , Xiaoou Ding , Yice Zhang , Jianzhong Li , Hong Gao

Semivalue-based data valuation uses cooperative-game theory intuitions to assign each data point a value reflecting its contribution to a downstream task. Still, those values depend on the practitioner's choice of utility, raising the…

Artificial Intelligence · Computer Science 2026-03-11 Mélissa Tamine , Benjamin Heymann , Maxime Vono , Patrick Loiseau

Data is the new oil of the 21st century. The growing trend of trading data for greater welfare has led to the emergence of data markets. A data market is any mechanism whereby the exchange of data products including datasets and data…

The Shapley value provides a principled framework for fairly distributing rewards among participants according to their individual contributions. While prior work has applied this concept to data valuation in machine learning, existing…

Computer Science and Game Theory · Computer Science 2026-01-22 Zhuofan Jia , Jian Pei

Methods for quantifying the similarity of datasets are relevant in applications where two or more datasets, or their underlying distributions, need to be compared, ranging from two- and k-sample testing to applications in machine learning…

Methodology · Statistics 2026-04-15 Marieke Stolte , Jörg Rahnenführer , Andrea Bommert

This paper presents a selective review of statistical computation methods for massive data analysis. A huge amount of statistical methods for massive data computation have been rapidly developed in the past decades. In this work, we focus…

Data collection is a fundamental problem in the scenario of big data, where the size of sampling sets plays a very important role, especially in the characterization of data structure. This paper considers the information collection process…

Information Theory · Computer Science 2018-01-23 Shanyun Liu , Rui She , Pingyi Fan

We identify the task of measuring data to quantitatively characterize the composition of machine learning data and datasets. Similar to an object's height, width, and volume, data measurements quantify different attributes of data along…

Distinguishing the importance of views has proven to be quite helpful for semi-supervised multi-view learning models. However, existing strategies cannot take advantage of semi-supervised information, only distinguishing the importance of…

Computer Vision and Pattern Recognition · Computer Science 2022-01-04 Yuyuan Yu , Guoxu Zhou , Haonan Huang , Shengli Xie , Qibin Zhao

In data mining, estimating the number of distinct values (NDV) is a fundamental problem with various applications. Existing methods for estimating NDV can be broadly classified into two categories: i) scanning-based methods, which scan the…

Databases · Computer Science 2022-06-14 Jiajun Li , Zhewei Wei , Bolin Ding , Xiening Dai , Lu Lu , Jingren Zhou

Information collection is a fundamental problem in big data, where the size of sampling sets plays a very important role. This work considers the information collection process by taking message importance into account. Similar to…

Information Theory · Computer Science 2018-01-15 Shanyun Liu , Rui She , Pingyi Fan

This paper describes valuation-based systems for representing and solving discrete optimization problems. In valuation-based systems, we represent information in an optimization problem using variables, sample spaces of variables, a set of…

Artificial Intelligence · Computer Science 2013-04-05 Prakash P. Shenoy , Glenn Shafer

Mean absolute deviation function is used to explore the pattern and the distribution of the data graphically to enable analysts gaining greater understanding of raw data and to foster quick and a deep understanding of the data as an…

Methodology · Statistics 2022-06-22 Elsayed A. H. Elamir

We introduce methods to bound the mean of a discrete distribution (or finite population) based on sample data, for random variables with a known set of possible values. In particular, the methods can be applied to categorical data with…

Statistics Theory · Mathematics 2021-11-16 Eric Bax , Frédéric Ouimet

In recent years, dataset distillation has provided a reliable solution for data compression, where models trained on the resulting smaller synthetic datasets achieve performance comparable to those trained on the original datasets. To…

Estimating statistical models within sensor networks requires distributed algorithms, in which both data and computation are distributed across the nodes of the network. We propose a general approach for distributed learning based on…

Machine Learning · Computer Science 2012-07-03 Qiang Liu , Alexander Ihler

We study statistical parameter estimation in the setting of data markets. A buyer seeks to estimate a parameter based on samples that can be purchased from competing providers that differ in their data quality and provision costs. When…

Computer Science and Game Theory · Computer Science 2026-04-13 Yuchen Hu , Martin J. Wainwright , Stephen Bates

Data valuation plays a crucial role in machine learning. Existing data valuation methods, mainly focused on discriminative models, overlook generative models that have gained attention recently. In generative models, data valuation measures…

Computer Vision and Pattern Recognition · Computer Science 2025-04-15 Jiaxi Yang , Wenglong Deng , Benlin Liu , Yangsibo Huang , James Zou , Xiaoxiao Li