English
Related papers

Related papers: Statistical Distortion: Consequences of Data Clean…

200 papers

This paper studies the problem of multivariate linear regression where a portion of the observations is grossly corrupted or is missing, and the magnitudes and locations of such occurrences are unknown in priori. To deal with this problem,…

Machine Learning · Statistics 2017-01-12 Xiaowei Zhang , Chi Xu , Yu Zhang , Tingshao Zhu , Li Cheng

One fundamental statistical question for research areas such as precision medicine and health disparity is about discovering effect modification of treatment or exposure by observed covariates. We propose a semiparametric framework for…

Methodology · Statistics 2020-08-04 Muxuan Liang , Menggang Yu

Big data are data on a massive scale in terms of volume, intensity, and complexity that exceed the capacity of standard software tools. They present opportunities as well as challenges to statisticians. The role of computational…

Computation · Statistics 2018-06-13 Chun Wang , Ming-Hui Chen , Elizabeth Schifano , Jing Wu , Jun Yan

'Big' high-dimensional data are commonly analyzed in low-dimensions, after performing a dimensionality-reduction step that inherently distorts the data structure. For the same purpose, clustering methods are also often used. These methods…

Machine Learning · Statistics 2019-02-20 Tom Lorimer , Karlis Kanders , Ruedi Stoop

Unsupervised machine learning lacks ground truth by definition. This poses a major difficulty when designing metrics to evaluate the performance of such algorithms. In sharp contrast with supervised learning, for which plenty of quality…

Machine Learning · Computer Science 2023-03-20 Raúl Lara-Cabrera , Ángel González-Prieto , Diego Pérez-López , Diego Trujillo , Fernando Ortega

Software systems are expansive, exhibiting behaviors characteristic of complex systems, such as self-organization and emergence. These systems, highlighted by advancements in Large Language Models (LLMs) and other AI applications developed…

Software Engineering · Computer Science 2025-04-01 Jan Žižka

This paper provides a detailed survey of synthetic data techniques. We first discuss the expected goals of using synthetic data in data augmentation, which can be divided into four parts: 1) Improving Diversity, 2) Data Balancing, 3)…

Machine Learning · Computer Science 2024-07-08 Hsin-Yu Chang , Pei-Yu Chen , Tun-Hsiang Chou , Chang-Sheng Kao , Hsuan-Yun Yu , Yen-Ting Lin , Yun-Nung Chen

Randomized algorithms depend on accurate sampling from probability distributions, as their correctness and performance hinge on the quality of the generated samples. However, even for common distributions like Binomial, exact sampling is…

Computation · Statistics 2025-06-17 Uddalok Sarkar , Sourav Chakraborty , Kuldeep S. Meel

Statistical estimation in many contemporary settings involves the acquisition, analysis, and aggregation of datasets from multiple sources, which can have significant differences in character and in value. Due to these variations, the…

Applications · Statistics 2014-12-23 Quentin Berthet , Venkat Chandrasekaran

Complex, high-dimensional data is ubiquitous across many scientific disciplines, including machine learning, biology, and the social sciences. One of the primary methods of visualizing these datasets is with two-dimensional scatter plots…

Machine Learning · Computer Science 2025-10-13 Kiran Smelser , Kaviru Gunaratne , Jacob Miller , Stephen Kobourov

Context: Empirical Software Engineering (ESE) drives innovation in SE through qualitative and quantitative studies. However, concerns about the correct application of empirical methodologies have existed since the 2006 Dagstuhl seminar on…

We study the problem of robust mean estimation and introduce a novel Hamming distance-based measure of distribution shift for coordinate-level corruptions. We show that this measure yields adversary models that capture more realistic…

Machine Learning · Computer Science 2021-06-14 Zifan Liu , Jongho Park , Theodoros Rekatsinas , Christos Tzamos

We consider models for social choice where voters rank a set of choices (or alternatives) by deliberating in small groups of size at most $k$, and these outcomes are aggregated by a social choice rule to find the winning alternative. We…

Computer Science and Game Theory · Computer Science 2025-03-21 Ashish Goel , Mohak Goyal , Kamesh Munagala

While sensitivity analysis improves the transparency and reliability of mathematical models, its uptake by modelers is still scarce. This is partially explained by its technical requirements, which may be hard to understand and implement by…

Applications · Statistics 2023-03-20 Arnald Puy , Pamphile T. Roy , Andrea Saltelli

Data inherently possesses dual attributes: samples and targets. For targets, knowledge distillation has been widely employed to accelerate model convergence, primarily relying on teacher-generated soft target supervision. Conversely, recent…

Machine Learning · Computer Science 2025-06-04 Runkang Yang , Peng Sun , Xinyi Shang , Yi Tang , Tao Lin

How should we evaluate the effect of a policy on the likelihood of an undesirable event, such as conflict? The significance test has three limitations. First, relying on statistical significance misses the fact that uncertainty is a…

Methodology · Statistics 2022-05-03 Akisato Suzuki

We study social choice rules under the utilitarian distortion framework, with an additional metric assumption on the agents' costs over the alternatives. In this approach, these costs are given by an underlying metric on the set of all…

Computer Science and Game Theory · Computer Science 2021-01-15 Ashish Goel , Anilesh Kollagunta Krishnaswamy , Kamesh Munagala

In recent years, dataset distillation has provided a reliable solution for data compression, where models trained on the resulting smaller synthetic datasets achieve performance comparable to those trained on the original datasets. To…

We precisely quantify the impact of statistical error in the quality of a numerical approximation to a random matrix eigendecomposition, and under mild conditions, we use this to introduce an optimal numerical tolerance for residual error…

We present techniques to characterize which data is important to a recommender system and which is not. Important data is data that contributes most to the accuracy of the recommendation algorithm, while less important data contributes less…

Information Retrieval · Computer Science 2013-10-04 Richard Chow , Hongxia Jin , Bart Knijnenburg , Gokay Saldamli