English
Related papers

Related papers: Private, Augmentation-Robust and Task-Agnostic Dat…

200 papers

Modeling human personality is important for several AI challenges, from the engineering of artificial psychotherapists to the design of persona bots. However, the field of computational personality analysis heavily relies on labeled data,…

Computation and Language · Computer Science 2023-01-23 Yair Neuman , Vladyslav Kozhukhov , Dan Vilenchik

We consider the problem of purchasing data for machine learning or statistical estimation. The data analyst has a budget to purchase datasets from multiple data providers. She does not have any test data that can be used to evaluate the…

Computer Science and Game Theory · Computer Science 2020-10-30 Yiling Chen , Yiheng Shen , Shuran Zheng

Data augmentation is known to contribute significantly to the robustness of machine learning models. In most instances, data augmentation is utilized during the training phase. Test-Time Augmentation (TTA) is a technique that instead…

Machine Learning · Statistics 2024-09-20 Masanari Kimura , Howard Bondell

Privacy is crucial in many applications of machine learning. Legal, ethical and societal issues restrict the sharing of sensitive data making it difficult to learn from datasets that are partitioned between many parties. One important…

Machine Learning · Statistics 2018-09-21 Christina Heinze-Deml , Brian McWilliams , Nicolai Meinshausen

We consider a data analyst's problem of purchasing data from strategic agents to compute an unbiased estimate of a statistic of interest. Agents incur private costs to reveal their data and the costs can be arbitrarily correlated with their…

Computer Science and Game Theory · Computer Science 2018-09-06 Yiling Chen , Nicole Immorlica , Brendan Lucier , Vasilis Syrgkanis , Juba Ziani

The emerging public awareness and government regulations of data privacy motivate new paradigms of collecting and analyzing data that are transparent and acceptable to data owners. We present a new concept of privacy and corresponding data…

Cryptography and Security · Computer Science 2022-06-08 Jie Ding , Bangjun Ding

The 'old world' instrument, survey, remains a tool of choice for firms to obtain ratings of satisfaction and experience that customers realize while interacting online with firms. While avenues for survey have evolved from emails and links…

Artificial Intelligence · Computer Science 2020-06-14 Atanu R Sinha , Deepali Jain , Nikhil Sheoran , Sopan Khosla , Reshmi Sasidharan

This paper studies two design tasks faced by a geo-distributed cloud data market: which data to purchase (data purchasing) and where to place/replicate the data for delivery (data placement). We show that the joint problem of data…

Distributed, Parallel, and Cluster Computing · Computer Science 2016-04-12 Xiaoqi Ren , Palma London , Juba Ziani , Adam Wierman

Given the vital role that smart meter data could play in handling uncertainty in energy markets, data markets have been proposed as a means to enable increased data access. However, most extant literature considers energy markets and data…

Systems and Control · Electrical Eng. & Systems 2024-12-11 Saurab Chhachhi , Fei Teng

A key challenge with machine learning approaches for ranking is the gap between the performance metrics of interest and the surrogate loss functions that can be optimized with gradient-based methods. This gap arises because ranking metrics…

Machine Learning · Computer Science 2021-11-30 Robin Swezey , Aditya Grover , Bruno Charron , Stefano Ermon

The widespread availability of large public datasets is a key factor behind the recent successes of statistical inference and machine learning methods. However, these datasets often contain some low-quality or contaminated data, to which…

Machine Learning · Statistics 2025-07-11 Kristian Minchev , Dimitar Iliev Dimitrov , Nikola Konstantinov

Synthetic data is increasingly used to support research without exposing sensitive user content. Social media data is one of the types of datasets that would hugely benefit from representative synthetic equivalents that can be used to…

Cryptography and Security · Computer Science 2026-03-06 Henry Tari , Adriana Iamnitchi

In matching markets such as job posting and online dating platforms, the recommender system plays a critical role in the success of the platform. Unlike standard recommender systems that suggest items to users, reciprocal recommender…

Information Retrieval · Computer Science 2023-07-28 Yoji Tomita , Riku Togashi , Yuriko Hashizume , Naoto Ohsaka

Data augmentation has been widely applied as an effective methodology to improve generalization in particular when training deep neural networks. Recently, researchers proposed a few intensive data augmentation techniques, which indeed…

Machine Learning · Computer Science 2019-11-22 Zhuoxun He , Lingxi Xie , Xin Chen , Ya Zhang , Yanfeng Wang , Qi Tian

Learning-augmented data structures use predicted frequency estimates to retrieve frequently occurring database elements faster than standard data structures. Recent work has developed data structures that optimally exploit these frequency…

Information Retrieval · Computer Science 2025-10-02 Prabhav Goyal , Vinesh Sridhar , Wilson Zheng

The financial market is a particularly challenging playground for deep reinforcement learning due to its unique feature of dynamic datasets. Building high-quality market environments for training financial reinforcement learning (FinRL)…

Machine Learning · Computer Science 2023-04-27 Xiao-Yang Liu , Ziyi Xia , Hongyang Yang , Jiechao Gao , Daochen Zha , Ming Zhu , Christina Dan Wang , Zhaoran Wang , Jian Guo

Assessing whether a sample survey credibly represents the population is a critical question for ensuring the validity of downstream research. Generally, this problem reduces to estimating the distance between two high-dimensional…

Machine Learning · Computer Science 2025-08-29 Debabrota Basu , Sourav Chakraborty , Debarshi Chanda , Buddha Dev Das , Arijit Ghosh , Arnab Ray

Data preparation, also called data wrangling, is considered one of the most expensive and time-consuming steps when performing analytics or building machine learning models. Preparing data typically involves collecting and merging data from…

Computation and Language · Computer Science 2023-06-22 Michael Glass , Xueqing Wu , Ankita Rajaram Naik , Gaetano Rossiello , Alfio Gliozzo

Estimating causal effects from observational data is essential in fields such as medicine, economics and social sciences, where privacy concerns are paramount. We propose a general, model-agnostic framework for differentially private…

Machine Learning · Computer Science 2026-02-02 Christian Janos Lebeda , Mathieu Even , Aurélien Bellet , Julie Josse

Labeling a large set of data is expensive. Active learning aims to tackle this problem by asking to annotate only the most informative data from the unlabeled set. We propose a novel active learning approach that utilizes self-supervised…

Computer Vision and Pattern Recognition · Computer Science 2022-07-27 John Seon Keun Yi , Minseok Seo , Jongchan Park , Dong-Geol Choi