English
Related papers

Related papers: Minor Issues Escalated to Critical Levels in Large…

200 papers

Study samples often differ from the target populations of inference and policy decisions in non-random ways. Researchers typically believe that such departures from random sampling -- due to changes in the population over time and space, or…

Methodology · Statistics 2023-07-20 Tamara Broderick , Ryan Giordano , Rachael Meager

Informatics and technological advancements have triggered generation of huge volume of data with varied complexity in its management and analysis. Big Data analytics is the practice of revealing hidden aspects of such data and making…

Databases · Computer Science 2018-03-30 Bikram Karmakar , Indranil Mukhopadhyay

How should researchers analyze randomized experiments in which the main outcome is latent and measured in multiple ways but each measure contains some degree of error? We first identify a critical study-specific noncomparability problem in…

Econometrics · Economics 2026-01-13 Jiawei Fu , Donald P. Green

This paper develops an interpretive framework for divergence P-values and S-values within a descriptive frequentist perspective. Statistical analysis is framed as operating within idealized worlds defined by a set of assumptions and a…

Other Statistics · Statistics 2026-03-31 Alessandro Rovetta

It is common for genomic data analysis to use $p$-values from a large number of permutation tests. The multiplicity of tests may require very tiny $p$-values in order to reject any null hypotheses and the common practice of using randomly…

Statistics Theory · Mathematics 2017-08-10 Hera Yu He , Kinjal Basu , Qingyuan Zhao , Art B. Owen

We consider one of the most basic multiple testing problems that compares expectations of multivariate data among several groups. As a test statistic, a conventional (approximate) $t$-statistic is considered, and we determine its rejection…

Methodology · Statistics 2016-12-20 Yoshiyuki Ninomiya , Satoshi Kuriki , Toshihiko Shiroishi , Toyoyuki Takada

A novel heuristic approach is proposed here for time series data analysis, dubbed Generalized weighted permutation entropy, which amalgamates and generalizes beyond their original scope two well established data analysis methods:…

Statistical Mechanics · Physics 2022-10-19 Darko Stosic , Dusan Stosic , Tatijana Stosic , Borko Stosic

We study a structured permutation scheme for two-sample testing that restricts permutations to single cross-swaps between block-selected representatives. Our analysis yields three main results. First, we provide an exact validity…

Machine Learning · Statistics 2025-12-02 Jungwoo Ho

In this paper, we study inference for high-dimensional data characterized by small sample sizes relative to the dimension of the data. In particular, we provide an infinite-dimensional framework to study statistical models that involve…

Statistics Theory · Mathematics 2010-02-25 Jim Kuelbs , Anand N. Vidyashankar

Identifying anomalies and contamination in datasets is important in a wide variety of settings. In this paper, we describe a new technique for estimating contamination in large, discrete valued datasets. Our approach considers the normal…

Information Theory · Computer Science 2015-06-16 Matthew L. Malloy , Scott Alfeld , Paul Barford

Clinical prediction models must be developed using sufficiently large datasets to minimise overfitting and ensure robust predictive performance. Existing sample size calculations assume complete predictor data for all included participants,…

In randomised trials, continuous endpoints are often measured with some degree of error. This study explores the impact of ignoring measurement error, and proposes methods to improve statistical inference in the presence of measurement…

Methodology · Statistics 2019-08-30 Linda Nab , Rolf H. H. Groenwold , Paco M. J. Welsing , Maarten van Smeden

Data augmentation is commonly used to encode invariances in learning methods. However, this process is often performed in an inefficient manner, as artificial examples are created by applying a number of transformations to all points in the…

Machine Learning · Computer Science 2019-03-04 Michael Kuchnik , Virginia Smith

We propose a permutation-based method for testing a large collection of hypotheses simultaneously. Our method provides lower bounds for the number of true discoveries in any selected subset of hypotheses. These bounds are simultaneously…

Applications · Statistics 2023-01-30 Angela Andreella , Jesse Hemerik , Wouter Weeda , Livio Finos , Jelle Goeman

Measurement error is a pervasive issue which renders the results of an analysis unreliable. The measurement error literature contains numerous correction techniques, which can be broadly divided into those which aim to produce exactly…

Methodology · Statistics 2021-11-08 Dylan Spicker , Michael P Wallace , Grace Y Yi

Increasing accessibility of data to researchers makes it possible to conduct massive amounts of statistical testing. Rather than follow a carefully crafted set of scientific hypotheses with statistical analysis, researchers can now test…

Genomics · Quantitative Biology 2016-09-08 Olga A. Vsevolozhskaya , Chia-Ling Kuo , Gabriel Ruiz , Luda Diatchenko , Dmitri V. Zaykin

Data can be collected in scientific studies via a controlled experiment or passive observation. Big data is often collected in a passive way, e.g. from social media. In studies of causation great efforts are made to guard against bias and…

Methodology · Statistics 2018-11-21 Elena Pesce , Eva Riccomagno , Henry P. Wynn

Data augmentation has been widely applied as an effective methodology to improve generalization in particular when training deep neural networks. Recently, researchers proposed a few intensive data augmentation techniques, which indeed…

Machine Learning · Computer Science 2019-11-22 Zhuoxun He , Lingxi Xie , Xin Chen , Ya Zhang , Yanfeng Wang , Qi Tian

When machine learning systems meet real world applications, accuracy is only one of several requirements. In this paper, we assay a complementary perspective originating from the increasing availability of pre-trained and regularly…

Databases in domains such as healthcare are routinely released to the public in aggregated form. Unfortunately, naive modeling with aggregated data may significantly diminish the accuracy of inferences at the individual level. This paper…

Machine Learning · Statistics 2016-05-17 Avradeep Bhowmik , Joydeep Ghosh , Oluwasanmi Koyejo