中文
相关论文

相关论文: Statistical Distortion: Consequences of Data Clean…

200 篇论文

In many machine learning for healthcare tasks, standard datasets are constructed by amassing data across many, often fundamentally dissimilar, sources. But when does adding more data help, and when does it hinder progress on desired model…

机器学习 · 计算机科学 2024-08-09 Judy Hanwen Shen , Inioluwa Deborah Raji , Irene Y. Chen

When working with real-world insurance data, practitioners often encounter challenges during the data preparation stage that can undermine the statistical validity and reliability of downstream modeling. This study illustrates that…

机器学习 · 统计学 2026-03-20 Jiayi Guo , Panyi Dong , Zhiyu Quan

Rigorous assessment of uncertainty is crucial to the utility of DNS results. Uncertainties in the computed statistics arise from two sources: finite statistical sampling and the discretization of the Navier-Stokes equations. Due to the…

流体动力学 · 物理学 2015-06-17 Todd A. Oliver , Nicholas Malaya , Rhys Ulerich , Robert D. Moser

Existing effect measures for compositional features are inadequate for many modern applications, for example, in microbiome research, since they display traits such as high-dimensionality and sparsity that can be poorly modelled with…

统计方法学 · 统计学 2025-06-02 Anton Rask Lundborg , Niklas Pfister

We study the design of voting rules in the metric distortion framework. It is known that any deterministic rule suffers distortion of at least $3$, and that randomized rules can achieve distortion strictly less than $3$, often at the cost…

计算机科学与博弈论 · 计算机科学 2026-02-10 Ziyi Cai , D. D. Gao , Prasanna Ramakrishnan , Kangning Wang

Data is inherently dirty and there has been a sustained effort to come up with different approaches to clean it. A large class of data repair algorithms rely on data-quality rules and integrity constraints to detect and repair the data. A…

数据库 · 计算机科学 2017-12-29 El Kindi Rezig , Mourad Ouzzani , Walid G. Aref , Ahmed K. Elmagarmid , Ahmed R. Mahmood

We determine the quality of randomized social choice mechanisms in a setting in which the agents have metric preferences: every agent has a cost for each alternative, and these costs form a metric. We assume that these costs are unknown to…

人工智能 · 计算机科学 2016-09-27 Elliot Anshelevich , John Postl

Assessing whether a sample survey credibly represents the population is a critical question for ensuring the validity of downstream research. Generally, this problem reduces to estimating the distance between two high-dimensional…

机器学习 · 计算机科学 2025-08-29 Debabrota Basu , Sourav Chakraborty , Debarshi Chanda , Buddha Dev Das , Arijit Ghosh , Arnab Ray

We present the first diffusion-based framework that can learn an unknown distribution using only highly-corrupted samples. This problem arises in scientific applications where access to uncorrupted samples is impossible or expensive to…

机器学习 · 计算机科学 2023-05-31 Giannis Daras , Kulin Shah , Yuval Dagan , Aravind Gollakota , Alexandros G. Dimakis , Adam Klivans

Modern technologies are generating ever-increasing amounts of data. Making use of these data requires methods that are both statistically sound and computationally efficient. Typically, the statistical and computational aspects are treated…

统计方法学 · 统计学 2022-09-15 Mahsa Taheri , Néhémy Lim , Johannes Lederer

Data visualization is the process by which data of any size or dimensionality is processed to produce an understandable set of data in a lower dimensionality, allowing it to be manipulated and understood more easily by people. The goal of…

图形学 · 计算机科学 2021-07-06 Alexander Kiefer , Md. Khaledur Rahman

This paper presents a selective survey of recent developments in statistical inference and multiple testing for high-dimensional regression models, including linear and logistic regression. We examine the construction of confidence…

统计方法学 · 统计学 2023-01-26 T. Tony Cai , Zijian Guo , Yin Xia

This article explores the critical role of statistical analysis in precision medicine. It discusses how personalized healthcare is enhanced by statistical methods that interpret complex, multidimensional datasets, focusing on predictive…

机器学习 · 计算机科学 2024-01-17 Xiaofei Chen

Many major works in social science employ matching to make causal conclusions, but different matches on the same data may produce different treatment effect estimates, even when they achieve similar balance or minimize the same loss…

应用统计 · 统计学 2023-03-23 Marco Morucci , Cynthia Rudin

Data sharing in the medical image analysis field has potential yet remains underappreciated. The aim is often to share datasets efficiently with other sites to train models effectively. One possible solution is to avoid transferring the…

图像与视频处理 · 电气工程与系统科学 2025-02-25 Muyang Li , Can Cui , Quan Liu , Ruining Deng , Tianyuan Yao , Marilyn Lionts , Yuankai Huo

We study a setting where a data holder wishes to share data with a receiver, without revealing certain summary statistics of the data distribution (e.g., mean, standard deviation). It achieves this by passing the data through a…

密码学与安全 · 计算机科学 2023-10-31 Zinan Lin , Shuaiqi Wang , Vyas Sekar , Giulia Fanti

Linear regression is a fundamental building block of statistical data analysis. It amounts to estimating the parameters of a linear model that maps input features to corresponding outputs. In the classical setting where the precision of…

计算机科学与博弈论 · 计算机科学 2019-12-16 Nicolas Gast , Stratis Ioannidis , Patrick Loiseau , Benjamin Roussillon

We study the utilitarian distortion of social choice mechanisms under the recently proposed learning-augmented framework where some (possibly unreliable) predicted information about the preferences of the agents is given as input. In…

计算机科学与博弈论 · 计算机科学 2025-02-11 Aris Filos-Ratsikas , Georgios Kalantzis , Alexandros A. Voudouris

A powerful approach to detecting erroneous data is to check which potentially dirty data records are incompatible with a user's domain knowledge. Previous approaches allow the user to specify domain knowledge in the form of logical…

数据库 · 计算机科学 2019-02-27 Jing Nathan Yan , Oliver Schulte , Jiannan Wang , Reynold Cheng

Deep ensembles excel in large-scale image classification tasks both in terms of prediction accuracy and calibration. Despite being simple to train, the computation and memory cost of deep ensembles limits their practicability. While some…

机器学习 · 计算机科学 2021-10-28 Giung Nam , Jongmin Yoon , Yoonho Lee , Juho Lee
‹ 上一页 1 8 9 10 下一页 ›