中文
相关论文

相关论文: Statistical Distortion: Consequences of Data Clean…

200 篇论文

Frequentist statistical methods, such as hypothesis testing, are standard practice in papers that provide benchmark comparisons. Unfortunately, these methods have often been misused, e.g., without testing for their statistical test…

统计方法学 · 统计学 2021-05-18 David Issa Mattos , Jan Bosch , Helena Holmström Olsson

We establish sharp upper and lower bounds for distortion risk metrics under distributional uncertainty. The uncertainty sets are characterized by four key features of the underlying distribution: mean, variance, unimodality, and Wasserstein…

风险管理 · 定量金融 2025-11-13 Peng Liu , Steven Vanduffel , Yi Xia

In this paper, we consider the problem of estimating the distance between any two large data streams in small- space constraint. This problem is of utmost importance in data intensive monitoring applications where input streams are…

数据结构与算法 · 计算机科学 2012-08-01 Emmanuelle Anceaume , Yann Busnel

As datasets grow it becomes infeasible to process them completely with a desired model. For giant datasets, we frame the order in which computation is performed as a decision problem. The order is designed so that partial computations are…

统计计算 · 统计学 2014-03-18 Daniel John Lawson , Niall M Adams

Data cleaning is often framed as a technical preprocessing step, yet in practice it relies heavily on human judgment. We report results from a controlled survey study in which participants performed error detection, data repair and…

数据库 · 计算机科学 2026-03-26 Hazim AbdElazim , Shadman Islam , Mostafa Milani

We address the problem of merging graph and feature-space information while learning a metric from structured data. Existing algorithms tackle the problem in an asymmetric way, by either extracting vectorized summaries of the graph…

机器学习 · 计算机科学 2020-02-17 Nicolo Colombo

Many modern data-intensive computational problems either require, or benefit from distance or similarity data that adhere to a metric. The algorithms run faster or have better performance guarantees. Unfortunately, in real applications, the…

机器学习 · 统计学 2017-10-31 Anna C. Gilbert , Lalit Jain

Contemporary statistical publications rely on simulation to evaluate performance of new methods and compare them with established methods. In the context of meta-analysis of log-odds-ratios, we investigate how the ways in which simulations…

统计方法学 · 统计学 2020-07-06 Elena Kulinskaya , David C. Hoaglin , Ilyas Bakbergenuly

We establish a profound connection between coherent risk measures, a prominent object in quantitative finance, and uniform integrability, a fundamental concept in probability theory. Instead of working with absolute values of random…

风险管理 · 定量金融 2025-04-08 Muqiao Huang , Ruodu Wang

Mapping political party systems to metric policy spaces is one of the major methodological problems in political science. At present, in most political science project this task is performed by domain experts relying on purely qualitative…

Synthetic data generation, a cornerstone of Generative Artificial Intelligence, promotes a paradigm shift in data science by addressing data scarcity and privacy while enabling unprecedented performance. As synthetic data becomes more…

机器学习 · 统计学 2024-03-12 Xiaotong Shen , Yifei Liu , Rex Shen

We consider the problem of quantitatively evaluating missing value imputation algorithms. Given a dataset with missing values and a choice of several imputation algorithms to fill them in, there is currently no principled way to rank the…

High-dimensional big data appears in many research fields such as image recognition, biology and collaborative filtering. Often, the exploration of such data by classic algorithms is encountered with difficulties due to `curse of…

机器学习 · 计算机科学 2016-07-13 Amit Bermanis , Aviv Rotbart , Moshe Salhov , Amir Averbuch

As synthetic data becomes widely used in language model development, understanding its impact on model behavior is crucial. This paper investigates the impact of the diversity of sources of synthetic data on fine-tuned large language…

计算与语言 · 计算机科学 2026-04-29 Max Schaffelder , Albert Gatt

In microarray technology, a number of critical steps are required to convert the raw measurements into the data relied upon by biologists and clinicians. These data manipulations, referred to as preprocessing, influence the quality of the…

应用统计 · 统计学 2009-09-29 Zhijin Wu , Rafael A. Irizarry

Statistical agencies face a dual mandate to publish accurate statistics while protecting respondent privacy. Increasing privacy protection requires decreased accuracy. Recognizing this as a resource allocation problem, we propose an…

密码学与安全 · 计算机科学 2019-03-12 John M. Abowd , Ian M. Schmutte

Ratio statistics--such as relative risk and odds ratios--play a central role in hypothesis testing, model evaluation, and decision-making across many areas of machine learning, including causal inference and fairness analysis. However,…

机器学习 · 统计学 2025-05-28 Tomer Shoham , Katrina Ligettt

The robustness of risk measures to changes in underlying loss distributions (distributional uncertainty) is of crucial importance in making well-informed decisions. In this paper, we quantify, for the class of distortion risk measures with…

风险管理 · 定量金融 2023-03-14 Carole Bernard , Silvana M. Pesenti , Steven Vanduffel

We study the performance of voting mechanisms from a utilitarian standpoint, under the recently introduced framework of metric-distortion, offering new insights along three main lines. First, if $d$ represents the doubling dimension of the…

计算机科学与博弈论 · 计算机科学 2022-03-25 Ioannis Anagnostides , Dimitris Fotakis , Panagiotis Patsilinakos

Statistical analysis is the tool of choice to turn data into information, and then information into empirical knowledge. To be valid, the process that goes from data to knowledge should be supported by detailed, rigorous guidelines, which…

软件工程 · 计算机科学 2024-10-03 Carlo A. Furia , Richard Torkar , Robert Feldt