English
Related papers

Related papers: Do We Really Even Need Data?

200 papers

Big data and algorithmic risk prediction tools promise to improve criminal justice systems by reducing human biases and inconsistencies in decision making. Yet different, equally-justifiable choices when developing, testing, and deploying…

Computers and Society · Computer Science 2022-09-23 Travis Greene , Galit Shmueli , Jan Fell , Ching-Fu Lin , Han-Wei Liu

Much artificial intelligence research focuses on the problem of deducing the validity of unobservable propositions or hypotheses from observable evidence.! Many of the knowledge representation techniques designed for this problem encode the…

Artificial Intelligence · Computer Science 2013-04-12 Ross D. Shachter , David Heckerman

To analyze unstructured data (text, images, audio, video), economists typically first extract low-dimensional structured features with a neural network. Neural networks do not make generically unbiased predictions, and biases will propagate…

Econometrics · Economics 2026-02-20 Jacob Carlson , Melissa Dell

Due to the widespread use of data-powered systems in our everyday lives, concepts like bias and fairness gained significant attention among researchers and practitioners, in both industry and academia. Such issues typically emerge from the…

Machine Learning · Computer Science 2023-05-18 Gianluca Demartini , Kevin Roitero , Stefano Mizzaro

Identifying causal relationships from observation data is difficult, in large part, due to the presence of hidden common causes. In some cases, where just the right patterns of conditional independence and dependence lie in the data---for…

Artificial Intelligence · Computer Science 2018-01-08 David Heckerman

Semi-supervised learning is a setting in which one has labeled and unlabeled data available. In this survey we explore different types of theoretical results when one uses unlabeled data in classification and regression tasks. Most methods…

Machine Learning · Computer Science 2020-07-31 Alexander Mey , Marco Loog

In some multivariate problems with missing data, pairs of variables exist that are never observed together. For example, some modern biological tools can produce data of this form. As a result of this structure, the covariance matrix is…

Methodology · Statistics 2013-08-13 Max Grazier G'Sell , Shai S. Shen-Orr , Robert Tibshirani

Model-free reinforcement learning algorithms can compute policy gradients given sampled environment transitions, but require large amounts of data. In contrast, model-based methods can use the learned model to generate new data, but model…

Machine Learning · Computer Science 2022-03-04 Lukas P. Fröhlich , Maksym Lefarov , Melanie N. Zeilinger , Felix Berkenkamp

Many scholars have called for raising statistical hurdles to guard against false discoveries in academic publications. I show these calls may be difficult to justify empirically. Published data exhibit bias: results that fail to meet…

General Finance · Quantitative Finance 2024-04-09 Andrew Y. Chen

Causal inference is a key research area in machine learning, yet confusion reigns over the tools needed to tackle it. There are prevalent claims in the machine learning literature that you need a bespoke causal framework or notation to…

Machine Learning · Statistics 2025-12-30 Bruno Mlodozeniec , David Krueger , Richard E. Turner

One of the most crucial issues in data mining is to model human behaviour in order to provide personalisation, adaptation and recommendation. This usually involves implicit or explicit knowledge, either by observing user interactions, or by…

Human-Computer Interaction · Computer Science 2017-08-21 Kevin Jasberg , Sergej Sizov

We consider the problem of assessing whether, in an individual case, there is a causal relationship between an observed exposure and a response variable. When data are available on similar individuals we may be able to estimate prospective…

Statistics Theory · Mathematics 2023-11-15 Monica Musio , Philip Dawid

Causal effect estimation from observational data is a crucial but challenging task. Currently, only a limited number of data-driven causal effect estimation methods are available. These methods either provide only a bound estimation of the…

Methodology · Statistics 2020-11-10 Debo Cheng , Jiuyong Li , Lin Liu , Kui Yu , Thuc Duy Lee , Jixue Liu

Data-driven algorithms play a large role in decision making across a variety of industries. Increasingly, these algorithms are being used to make decisions that have significant ramifications for people's social and economic well-being,…

Machine Learning · Computer Science 2018-09-26 J. Henry Hinnefeld , Peter Cooman , Nat Mammo , Rupert Deese

Learning with limited data is one of the biggest problems of machine learning. Current approaches to this issue consist in learning general representations from huge amounts of data before fine-tuning the model on a small dataset of…

Machine Learning · Computer Science 2023-02-22 Grégoire Mialon

When machine learning systems meet real world applications, accuracy is only one of several requirements. In this paper, we assay a complementary perspective originating from the increasing availability of pre-trained and regularly…

A growing body of literature attempts to learn about contagion using observational (i.e. non-experimental) data collected from a single social network. While the conclusions of these studies may be correct, the methods rely on assumptions…

Applications · Statistics 2017-06-30 Elizabeth L. Ogburn

Quantifying and managing uncertainties that occur when data-driven models such as those provided by AI and machine learning methods are applied is crucial. This whitepaper provides a brief motivation and first overview of the state of the…

Machine Learning · Computer Science 2018-11-29 Michael Kläs

Advances in machine learning and the increasing availability of high-dimensional data have led to the proliferation of social science research that uses the predictions of machine learning models as proxies for measures of human activity or…

Machine Learning · Computer Science 2025-02-19 Luke C Sanford , Megan Ayers , Matthew Gordon , Eliana Stone

This paper discusses the problem of causal query in observational data with hidden variables, with the aim of seeking the change of an outcome when "manipulating" a variable while given a set of plausible confounding variables which affect…

Artificial Intelligence · Computer Science 2020-11-25 Debo Cheng , Jiuyong Li , Lin Liu , Jixue Liu , Kui Yu , Thuc Duy Le