中文
相关论文

相关论文: Statistical Distortion: Consequences of Data Clean…

200 篇论文

A fundamental problem in the practice and teaching of data science is how to evaluate the quality of a given data analysis, which is different than the evaluation of the science or question underlying the data analysis. Previously, we…

其他统计学 · 统计学 2019-04-29 Stephanie C. Hicks , Roger D. Peng

Informatics and technological advancements have triggered generation of huge volume of data with varied complexity in its management and analysis. Big Data analytics is the practice of revealing hidden aspects of such data and making…

数据库 · 计算机科学 2018-03-30 Bikram Karmakar , Indranil Mukhopadhyay

For high dimensional data, some of the standard statistical techniques do not work well. So modification or further development of statistical methods are necessary. In this paper, we explore these modifications. We start with the important…

统计金融 · 定量金融 2024-05-29 Arnab Chakrabarti , Rituparna Sen

Robust statistics aims to compute quantities to represent data where a fraction of it may be arbitrarily corrupted. The most essential statistic is the mean, and in recent years, there has been a flurry of theoretical advancement for…

机器学习 · 统计学 2025-02-18 Cullen Anderson , Jeff M. Phillips

Data quality issues have attracted widespread attention due to the negative impacts of dirty data on data mining and machine learning results. The relationship between data quality and the accuracy of results could be applied on the…

数据库 · 计算机科学 2021-04-27 Zhixin Qi , Hongzhi Wang , Jianzhong Li , Hong Gao

The notion of distortion in social choice problems has been defined to measure the loss in efficiency -- typically measured by the utilitarian social welfare, the sum of utilities of the participating agents -- due to having access only to…

计算机科学与博弈论 · 计算机科学 2021-03-02 Elliot Anshelevich , Aris Filos-Ratsikas , Nisarg Shah , Alexandros A. Voudouris

Data sharing enables critical advances in many research areas and business applications, but it may lead to inadvertent disclosure of sensitive summary statistics (e.g., means or quantiles). Existing literature only focuses on protecting a…

密码学与安全 · 计算机科学 2024-06-14 Shuaiqi Wang , Rongzhe Wei , Mohsen Ghassemi , Eleonora Kreacic , Vamsi K. Potluru

Datasets serve as crucial training resources and model performance trackers. However, existing datasets have exposed a plethora of problems, inducing biased models and unreliable evaluation results. In this paper, we propose a…

计算与语言 · 计算机科学 2022-12-20 Chengwen Wang , Qingxiu Dong , Xiaochen Wang , Haitao Wang , Zhifang Sui

Social choice theory offers a wealth of approaches for selecting a candidate on behalf of voters based on their reported preference rankings over options. When voters have underlying utilities for these options, however, using preference…

计算机科学与博弈论 · 计算机科学 2025-10-24 Luise Ge , Gregory Kehne , Yevgeniy Vorobeychik

Statistical matching is an effective method for estimating causal effects in which treated units are paired with control units with ``similar'' values of confounding covariates prior to performing estimation. In this way, matching helps…

统计方法学 · 统计学 2023-09-13 Sanjeewani Weerasingha , Michael J. Higgins

Data cleaning is the initial stage of any machine learning project and is one of the most critical processes in data analysis. It is a critical step in ensuring that the dataset is devoid of incorrect or erroneous data. It can be done…

数据库 · 计算机科学 2021-09-16 Ga Young Lee , Lubna Alzamil , Bakhtiyar Doskenov , Arash Termehchy

Data Cleaning refers to the process of detecting and fixing errors in the data. Human involvement is instrumental at several stages of this process, e.g., to identify and repair errors, to validate computed repairs, etc. There is currently…

数据库 · 计算机科学 2018-01-03 El Kindi Rezig , Mourad Ouzzani , Ahmed K. Elmagarmid , Walid G. Aref

The US Census Bureau will deliberately corrupt data sets derived from the 2020 US Census, enhancing the privacy of respondents while potentially reducing the precision of economic analysis. To investigate whether this trade-off is…

计量经济学 · 经济学 2024-02-13 Anish Agarwal , Rahul Singh

Data cleaning, whether manual or algorithmic, is rarely perfect leaving a dataset with an unknown number of false positives and false negatives after cleaning. In many scenarios, quantifying the number of remaining errors is challenging…

数据库 · 计算机科学 2017-05-30 Yeounoh Chung , Sanjay Krishnan , Tim Kraska

Dimension reduction of data sets is a standard problem in the realm of machine learning and knowledge reasoning. They affect patterns in and dependencies on data dimensions and ultimately influence any decision-making processes. Therefore,…

机器学习 · 计算机科学 2022-04-26 Tom Hanika , Johannes Hirth

The collection, transfer and integration of research information into different research Information systems can result in different data errors that can have a variety of negative effects on data quality. In order to detect errors at an…

数据库 · 计算机科学 2019-01-21 Otmane Azeroual , Gunter Saake , Mohammad Abuosba

We propose a general statistical inference framework to capture the privacy threat incurred by a user that releases data to a passive but curious adversary, given utility constraints. We show that applying this general framework to the…

信息论 · 计算机科学 2012-10-09 Flavio du Pin Calmon , Nadia Fawaz

The concept of depth has proved very important for multivariate and functional data analysis, as it essentially acts as a surrogate for the notion a ranking of observations which is absent in more than one dimension. Motivated by the rapid…

统计方法学 · 统计学 2021-07-30 Gery Geenens , Alicia Nieto-Reyes , Giacomo Francisci

Statistical divergence is widely applied in multimedia processing, basically due to regularity and interpretable features displayed in data. However, in a broader range of data realm, these advantages may no longer be feasible, and…

数据库 · 计算机科学 2020-11-20 Ruoyu Wang , Xiaobo Hu , Daniel Sun , Guoqiang Li , Raymond Wong , Shiping Chen , Jianquan Liu

Much of statistics relies upon four key elements: a law of large numbers, a calculus to operationalize stochastic convergence, a central limit theorem, and a framework for constructing local approximations. These elements are…

最优化与控制 · 数学 2018-01-09 Anil Aswani
‹ 上一页 1 2 3 10 下一页 ›