English
Related papers

Related papers: Data Distribution Valuation

200 papers

User behavior records serve as the foundation for recommender systems. While the behavior data exhibits ease of acquisition, it often suffers from varying quality. Current methods employ data valuation to discern high-quality data from…

Machine Learning · Computer Science 2025-02-14 Renqi Jia , Xiaokun Zhang , Bowei He , Qiannan Zhu , Weitao Xu , Jiehao Chen , Chen Ma

In this paper, we propose a novel data valuation method for a Dataset Retrieval (DR) use case in Ireland's National mapping agency. To the best of our knowledge, data valuation has not yet been applied to Dataset Retrieval. By leveraging…

Information Retrieval · Computer Science 2024-07-23 Malick Ebiele , Malika Bendechache , Eamonn Clinton , Rob Brennan

The quality of the data in a dataset can have a substantial impact on the performance of a machine learning model that is trained and/or evaluated using the dataset. Effective dataset management, including tasks such as data cleanup,…

Databases · Computer Science 2023-03-16 Ze Mao , Yang Xu , Erick Suarez

One of the biggest challenges of value-based decision-making is dealing with the subjective nature of values. The relative importance of a value for a particular decision varies between individuals, and people may also have different…

Multiagent Systems · Computer Science 2026-03-30 Arturo Hernandez-Sanchez , Natalia Criado , Stella Heras , Miguel Rebollo , Jose Such

In machine learning, knowing the impact of a given datum on model training is a fundamental task referred to as Data Valuation. Building on previous works from the literature, we have designed a novel canonical decomposition allowing…

Machine Learning · Computer Science 2025-06-05 Clément Bénesse , Patrick Mesana , Athénaïs Gautier , Sébastien Gambs

A key element in transfer learning is representation learning; if representations can be developed that expose the relevant factors underlying the data, then new tasks and domains can be learned readily based on mappings of these salient…

Machine Learning · Computer Science 2014-12-18 Yujia Li , Kevin Swersky , Richard Zemel

Machine learning (ML) methods are widely used in industrial applications, which usually require a large amount of training data. However, data collection needs extensive time costs and investments in the manufacturing system, and data…

Machine Learning · Computer Science 2024-04-02 Yue Zhao , Yuxuan Li , Chenang Liu , Yinan Wang

Problem definition: Traditional monopoly pricing assumes sellers have full information about consumer valuations. We consider monopoly pricing under limited information, where a seller only knows the mean, variance and support of the…

Optimization and Control · Mathematics 2026-03-30 Tim S. G. van Eck , Pieter Kleer , Johan S. H. van Leeuwaarden

The $\textit{data market design}$ problem is a problem in economic theory to find a set of signaling schemes (statistical experiments) to maximize expected revenue to the information seller, where each experiment reveals some of the…

Computer Science and Game Theory · Computer Science 2023-11-01 Sai Srivatsa Ravindranath , Yanchen Jiang , David C. Parkes

Machine learning has been proven to be effective in various application areas, such as object and speech recognition on mobile systems. Since a critical key to machine learning success is the availability of large training data, many…

Machine Learning · Computer Science 2021-01-06 Hyeongmin Cho , Sangkyun Lee

Data valuation quantifies data importance, but existing methods cannot ensure validity in a single training process. The neural dynamic data valuation (NDDV) method [3] addresses this limitation. Based on NDDV, we are the first to explore…

Machine Learning · Computer Science 2025-12-19 Zhangyong Liang , Huanhuan Gao , Ji Zhang

A big data service is any data-originated resource that is offered over the Internet. The performance of a big data service depends on the data bought from the data collectors. However, the problem of optimal pricing and data allocation in…

Computer Science and Game Theory · Computer Science 2017-04-13 Yutao Jiao , Ping Wang , Dusit Niyato , Mohammad Abu Alsheikh , Shaohan Feng

Data-driven machine learning (ML) has witnessed great successes across a variety of application domains. Since ML model training are crucially relied on a large amount of data, there is a growing demand for high quality data to be collected…

Databases · Computer Science 2020-03-31 Jinfei Liu

Data Mining is the process of extracting useful patterns from the huge amount of database and many data mining techniques are used for mining these patterns. Recently, one of the remarkable facts in higher educational institute is the rapid…

Artificial Intelligence · Computer Science 2014-05-16 Priyanka Saini

Quantifying the similarity between datasets has widespread applications in statistics and machine learning. The performance of a predictive model on novel datasets, referred to as generalizability, depends on how similar the training and…

Methodology · Statistics 2025-06-18 Marieke Stolte , Franziska Kappenberg , Jörg Rahnenführer , Andrea Bommert

In distributed machine learning, data is dispatched to multiple machines for processing. Motivated by the fact that similar data points often belong to the same or similar classes, and more generally, classification rules of high accuracy…

Machine Learning · Computer Science 2016-12-16 Travis Dick , Mu Li , Venkata Krishna Pillutla , Colin White , Maria Florina Balcan , Alex Smola

We present techniques to characterize which data is important to a recommender system and which is not. Important data is data that contributes most to the accuracy of the recommendation algorithm, while less important data contributes less…

Information Retrieval · Computer Science 2013-10-04 Richard Chow , Hongxia Jin , Bart Knijnenburg , Gokay Saldamli

Data assets are data commodities that have been processed, produced, priced, and traded based on actual demand. Reasonable pricing mechanism for data assets is essential for developing the data market and realizing their value. Most…

Mathematical Finance · Quantitative Finance 2025-05-23 Xiaoshan Chen , Chen Yang , Zhou Yang

Statistical divergence is widely applied in multimedia processing, basically due to regularity and interpretable features displayed in data. However, in a broader range of data realm, these advantages may no longer be feasible, and…

Databases · Computer Science 2020-11-20 Ruoyu Wang , Xiaobo Hu , Daniel Sun , Guoqiang Li , Raymond Wong , Shiping Chen , Jianquan Liu

Data mining has been widely used to identify potential customers for a new product or service. In this article is done a study of previous work relating to the application of data mining methodologies for software projects, specifically for…

Computers and Society · Computer Science 2016-09-06 Jorge Luis Rivero Pérez , Yaimara Peñate Santana , Pedro Harenton Martínez López