English
Related papers

Related papers: How much does your data exploration overfit? Contr…

200 papers

Observational data is often readily available in large quantities, but can lead to biased causal effect estimates due to the presence of unobserved confounding. Recent works attempt to remove this bias by supplementing observational data…

In many predictive contexts (e.g., credit lending), true outcomes are only observed for samples that were positively classified in the past. These past observations, in turn, form training datasets for classifiers that make future…

Machine Learning · Computer Science 2024-06-04 Vijay Keswani , Anay Mehrotra , L. Elisa Celis

Adaptive data analysis is frequently criticized for its pessimistic generalization guarantees. The source of these pessimistic bounds is a model that permits arbitrary, possibly adversarial analysts that optimally use information to bias…

Machine Learning · Computer Science 2019-05-14 Tijana Zrnic , Moritz Hardt

Information flow analysis has largely ignored the setting where the analyst has neither control over nor a complete model of the analyzed system. We formalize such limited information flow analyses and study an instance of it: detecting the…

Cryptography and Security · Computer Science 2014-05-13 Michael Carl Tschantz , Amit Datta , Anupam Datta , Jeannette M. Wing

As large language models (LLMs) are adopted into frameworks that grant them the capacity to make real decisions, it is increasingly important to ensure that they are unbiased. In this paper, we argue that the predominant approach of simply…

Computers and Society · Computer Science 2026-01-13 Addison J. Wu , Ryan Liu , Xuechunzi Bai , Thomas L. Griffiths

Fairness-aware recommender systems often mitigate bias by increasing exposure to under-represented or long-tail content, commonly through mechanisms that promote novelty and diversity. In practice, the strength of such interventions is…

Information Retrieval · Computer Science 2026-04-21 Enock O. Ayiku , Evelyn Osei , Emebo Onyeka

Tasks that require information about the world imply a trade-off between the time spent on observation and the variance of the response. In particular, fast decisions need to rely on uncertain information. However, standard estimates of…

Neurons and Cognition · Quantitative Biology 2023-07-18 Sahel Azizpour , Viola Priesemann , Johannes Zierenberg , Anna Levina

We consider information filtering, in which we face a stream of items too voluminous to process by hand (e.g., scientific articles, blog posts, emails), and must rely on a computer system to automatically filter out irrelevant items. Such…

Optimization and Control · Mathematics 2015-02-10 Xiaoting Zhao , Peter I. Frazier

Information foraging connects optimal foraging theory in ecology with how humans search for information. The theory suggests that, following an information scent, the information seeker must optimize the tradeoff between exploration by…

Information Retrieval · Computer Science 2016-11-18 Peter Wittek , Ying-Hsang Liu , Sándor Darányi , Tom Gedeon , Ik Soo Lim

Living in the 'Information Age' means that not only access to information has become easier but also that the distribution of information is more dynamic than ever. Through a large-scale online field experiment, we provide new empirical…

Computers and Society · Computer Science 2023-01-05 Taha Yasseri , Jannie Reher

Recommender systems rely on user behavior data like ratings and clicks to build personalization model. However, the collected data is observational rather than experimental, causing various biases in the data which significantly affect the…

Machine Learning · Computer Science 2021-10-29 Jiawei Chen , Hande Dong , Yang Qiu , Xiangnan He , Xin Xin , Liang Chen , Guli Lin , Keping Yang

The increasing availability of passively observed data has yielded a growing methodological interest in "data fusion." These methods involve merging data from observational and experimental sources to draw causal conclusions -- and they…

Methodology · Statistics 2021-12-15 Evan Rosenman , Art B. Owen

While data-driven decision-making is transforming modern operations, most large-scale data is of an observational nature, such as transactional records. These data pose unique challenges in a variety of operational problems posed as…

Optimization and Control · Mathematics 2017-05-23 Dimitris Bertsimas , Nathan Kallus

Recommender systems often struggle with over-specialization, which severely limits users' exposure to diverse content and creates filter bubbles that reduce serendipitous discovery. To address this fundamental limitation, this paper…

Information Retrieval · Computer Science 2026-05-27 Edoardo Bianchi

Exploration has been a crucial part of reinforcement learning, yet several important questions concerning exploration efficiency are still not answered satisfactorily by existing analytical frameworks. These questions include exploration…

Machine Learning · Computer Science 2016-12-06 Liangpeng Zhang , Ke Tang , Xin Yao

The closed feedback loop in recommender systems is a common setting that can lead to different types of biases. Several studies have dealt with these biases by designing methods to mitigate their effect on the recommendations. However, most…

Information Retrieval · Computer Science 2020-09-01 Sami Khenissi , Mariem Boujelbene , Olfa Nasraoui

Randomized experiments have long been the gold standard for scientists seeking to learn about cause and effect. When randomized experiments are infeasible, scientists often resort to observational studies, which are widely available and…

Methodology · Statistics 2026-04-13 Bohan Wu , Sebastian Salazar , Donald P. Green , David M. Blei

Experimental design is crucial for inference where limitations in the data collection procedure are present due to cost or other restrictions. Optimal experimental designs determine parameters that in some appropriate sense make the data…

Machine Learning · Statistics 2016-03-11 Panagiotis Tsilifis , Roger G. Ghanem , Paris Hajali

Binary classification is a task that involves the classification of data into one of two distinct classes. It is widely utilized in various fields. However, conventional classifiers tend to make overconfident predictions for data that…

Machine Learning · Computer Science 2025-03-13 Shoma Yokura , Akihisa Ichiki

There is an increasing concern that most current published research findings are false. The main cause seems to lie in the fundamental disconnection between theory and practice in data analysis. While the former typically relies on…

Machine Learning · Statistics 2019-03-06 Amedeo Roberto Esposito , Michael Gastpar , Ibrahim Issa