English
Related papers

Related papers: Statistical Distortion: Consequences of Data Clean…

200 papers

Frequentist statistical methods, such as hypothesis testing, are standard practice in papers that provide benchmark comparisons. Unfortunately, these methods have often been misused, e.g., without testing for their statistical test…

Methodology · Statistics 2021-05-18 David Issa Mattos , Jan Bosch , Helena Holmström Olsson

We establish sharp upper and lower bounds for distortion risk metrics under distributional uncertainty. The uncertainty sets are characterized by four key features of the underlying distribution: mean, variance, unimodality, and Wasserstein…

Risk Management · Quantitative Finance 2025-11-13 Peng Liu , Steven Vanduffel , Yi Xia

In this paper, we consider the problem of estimating the distance between any two large data streams in small- space constraint. This problem is of utmost importance in data intensive monitoring applications where input streams are…

Data Structures and Algorithms · Computer Science 2012-08-01 Emmanuelle Anceaume , Yann Busnel

As datasets grow it becomes infeasible to process them completely with a desired model. For giant datasets, we frame the order in which computation is performed as a decision problem. The order is designed so that partial computations are…

Computation · Statistics 2014-03-18 Daniel John Lawson , Niall M Adams

Data cleaning is often framed as a technical preprocessing step, yet in practice it relies heavily on human judgment. We report results from a controlled survey study in which participants performed error detection, data repair and…

Databases · Computer Science 2026-03-26 Hazim AbdElazim , Shadman Islam , Mostafa Milani

We address the problem of merging graph and feature-space information while learning a metric from structured data. Existing algorithms tackle the problem in an asymmetric way, by either extracting vectorized summaries of the graph…

Machine Learning · Computer Science 2020-02-17 Nicolo Colombo

Many modern data-intensive computational problems either require, or benefit from distance or similarity data that adhere to a metric. The algorithms run faster or have better performance guarantees. Unfortunately, in real applications, the…

Machine Learning · Statistics 2017-10-31 Anna C. Gilbert , Lalit Jain

Contemporary statistical publications rely on simulation to evaluate performance of new methods and compare them with established methods. In the context of meta-analysis of log-odds-ratios, we investigate how the ways in which simulations…

Methodology · Statistics 2020-07-06 Elena Kulinskaya , David C. Hoaglin , Ilyas Bakbergenuly

We establish a profound connection between coherent risk measures, a prominent object in quantitative finance, and uniform integrability, a fundamental concept in probability theory. Instead of working with absolute values of random…

Risk Management · Quantitative Finance 2025-04-08 Muqiao Huang , Ruodu Wang

Mapping political party systems to metric policy spaces is one of the major methodological problems in political science. At present, in most political science project this task is performed by domain experts relying on purely qualitative…

Computation and Language · Computer Science 2023-06-06 Daria Boratyn , Damian Brzyski , Beata Kosowska-Gąstoł , Jan Rybicki , Wojciech Słomczyński , Dariusz Stolicki

Synthetic data generation, a cornerstone of Generative Artificial Intelligence, promotes a paradigm shift in data science by addressing data scarcity and privacy while enabling unprecedented performance. As synthetic data becomes more…

Machine Learning · Statistics 2024-03-12 Xiaotong Shen , Yifei Liu , Rex Shen

We consider the problem of quantitatively evaluating missing value imputation algorithms. Given a dataset with missing values and a choice of several imputation algorithms to fill them in, there is currently no principled way to rank the…

Machine Learning · Computer Science 2013-11-12 Vinod Nair , Rahul Kidambi , Sundararajan Sellamanickam , S. Sathiya Keerthi , Johannes Gehrke , Vijay Narayanan

High-dimensional big data appears in many research fields such as image recognition, biology and collaborative filtering. Often, the exploration of such data by classic algorithms is encountered with difficulties due to `curse of…

Machine Learning · Computer Science 2016-07-13 Amit Bermanis , Aviv Rotbart , Moshe Salhov , Amir Averbuch

As synthetic data becomes widely used in language model development, understanding its impact on model behavior is crucial. This paper investigates the impact of the diversity of sources of synthetic data on fine-tuned large language…

Computation and Language · Computer Science 2026-04-29 Max Schaffelder , Albert Gatt

In microarray technology, a number of critical steps are required to convert the raw measurements into the data relied upon by biologists and clinicians. These data manipulations, referred to as preprocessing, influence the quality of the…

Applications · Statistics 2009-09-29 Zhijin Wu , Rafael A. Irizarry

Statistical agencies face a dual mandate to publish accurate statistics while protecting respondent privacy. Increasing privacy protection requires decreased accuracy. Recognizing this as a resource allocation problem, we propose an…

Cryptography and Security · Computer Science 2019-03-12 John M. Abowd , Ian M. Schmutte

Ratio statistics--such as relative risk and odds ratios--play a central role in hypothesis testing, model evaluation, and decision-making across many areas of machine learning, including causal inference and fairness analysis. However,…

Machine Learning · Statistics 2025-05-28 Tomer Shoham , Katrina Ligettt

The robustness of risk measures to changes in underlying loss distributions (distributional uncertainty) is of crucial importance in making well-informed decisions. In this paper, we quantify, for the class of distortion risk measures with…

Risk Management · Quantitative Finance 2023-03-14 Carole Bernard , Silvana M. Pesenti , Steven Vanduffel

We study the performance of voting mechanisms from a utilitarian standpoint, under the recently introduced framework of metric-distortion, offering new insights along three main lines. First, if $d$ represents the doubling dimension of the…

Computer Science and Game Theory · Computer Science 2022-03-25 Ioannis Anagnostides , Dimitris Fotakis , Panagiotis Patsilinakos

Statistical analysis is the tool of choice to turn data into information, and then information into empirical knowledge. To be valid, the process that goes from data to knowledge should be supported by detailed, rigorous guidelines, which…

Software Engineering · Computer Science 2024-10-03 Carlo A. Furia , Richard Torkar , Robert Feldt