English
Related papers

Related papers: Validating Approximate Slope Homogeneity in Large …

200 papers

In genetic studies, haplotype data provide more refined information than data about separate genetic markers. However, large-scale studies that genotype hundreds to thousands of individuals may only provide results of pooled data, where…

Methodology · Statistics 2023-09-01 Yong See Foo , Jennifer A. Flegg

We engineer a new probabilistic Monte-Carlo algorithm for isomorphism testing. Most notably, as opposed to all other solvers, it implicitly exploits the presence of symmetries without explicitly computing them. We provide extensive…

Data Structures and Algorithms · Computer Science 2020-11-19 Markus Anders , Pascal Schweitzer

Particularly in genomics, but also in other fields, it has become commonplace to undertake highly multiple Student's $t$-tests based on relatively small sample sizes. The literature on this topic is continually expanding, but the main…

Statistics Theory · Mathematics 2010-10-11 Peter Hall , Qiying Wang

Micro-panel data are collected and analysed in many research and industry areas. Cluster analysis of micro-panel data is an unsupervised learning exploratory method identifying subgroup clusters in a data set which include homogeneous…

Machine Learning · Statistics 2018-07-17 Lukas Sobisek , Maria Stachova , Jan Fojtik

Feature generation is an open topic of investigation in graph machine learning. In this paper, we study the use of graph homomorphism density features as a scalable alternative to homomorphism numbers which retain similar theoretical…

Machine Learning · Computer Science 2021-04-12 Paul Beaujean , Florian Sikora , Florian Yger

In this paper, we review state-of-the-art methods for feature selection in statistics with an application-oriented eye. Indeed, sparsity is a valuable property and the profusion of research on the topic might have provided little guidance…

Methodology · Statistics 2021-11-08 Dimitris Bertsimas , Jean Pauphilet , Bart Van Parys

Entropy is a measure of heterogeneity widely used in applied sciences, often when data are collected over space. Recently, a number of approaches has been proposed to include spatial information in entropy. The aim of entropy is to…

Statistics Theory · Mathematics 2019-11-12 Linda Altieri , Daniela Cocchi , Giulia Roli

In many complex applications, data heterogeneity and homogeneity exist simultaneously. Ignoring either one will result in incorrect statistical inference. In addition, coping with complex data that are non-Euclidean becomes more common. To…

Methodology · Statistics 2021-05-28 Zixuan Han , Tao Li , Jinhong You

Economic choices are often stochastic: the same person may make a different choice when facing the same alternatives repeatedly. Standard models assume that the degree of randomness reflects the size of utility differences, but choice…

Theoretical Economics · Economics 2026-05-05 Shuhua Si

We generalize Levene's test for variance (scale) heterogeneity between $k$ groups for more complex data, which includes sample correlation and group membership uncertainty. Following a two-stage regression framework, we show that least…

Methodology · Statistics 2016-05-19 David Soave , Lei Sun

In computational mechanics, multiple models are often present to describe a physical system. While Bayesian model selection is a helpful tool to compare these models using measurement data, it requires the computationally expensive…

Computation · Statistics 2025-04-14 Subhayan De , Reza Farzad , Patrick T. Brewick , Erik A. Johnson , Steven F. Wojtkiewicz

We develop randomization-based tests for heterogeneous treatment effects in the presence of network interference. Leveraging the exposure mapping framework, we study a broad class of null hypotheses that represent various forms of constant…

Econometrics · Economics 2025-06-25 Julius Owusu

In functional linear regression, the slope ``parameter'' is a function. Therefore, in a nonparametric context, it is determined by an infinite number of unknowns. Its estimation involves solving an ill-posed problem and has points of…

Statistics Theory · Mathematics 2007-08-07 Peter Hall , Joel L. Horowitz

Fairness is steadily becoming a crucial requirement of Machine Learning (ML) systems. A particularly important notion is subgroup fairness, i.e., fairness in subgroups of individuals that are defined by more than one attributes. Identifying…

Machine Learning · Computer Science 2024-04-30 Giorgos Giannopoulos , Dimitris Sacharidis , Nikolas Theologitis , Loukas Kavouras , Ioannis Emiris

We propose a change-point detection method for large scale multiple testing problems with data having clustered signals. Unlike the classic change-point setup, the signals can vary in size within a cluster. The clustering structure on the…

Methodology · Statistics 2021-10-07 Hongyuan Cao , Wei Biao Wu

In panel data we observe a usually high number N of individuals over a time period T. Even if T is large one often assumes stability of the model over time. We propose a nonparametric and robust test for a change in location and derive its…

Statistics Theory · Mathematics 2017-03-22 Alexander Dürre , Roland Fried

One popular method for dealing with large-scale data sets is sampling. For example, by using the empirical statistical leverage scores as an importance sampling distribution, the method of algorithmic leveraging samples and rescales…

Methodology · Statistics 2013-06-25 Ping Ma , Michael W. Mahoney , Bin Yu

Persistent homology is an important methodology in topological data analysis which adapts theory from algebraic topology to data settings. Computing persistent homology produces persistence diagrams, which have been successfully used in…

Machine Learning · Statistics 2026-01-13 Yueqi Cao , Anthea Monod

The $\beta$-model has been extensively utilized to model degree heterogeneity in networks, wherein each node is assigned a unique parameter. In this article, we consider the hypothesis testing problem that two nodes $i$ and $j$ of a…

Statistics Theory · Mathematics 2024-03-12 Kang Fu , Jianwei Hu , Meng Sun

Network data is prevalent in many contemporary big data applications in which a common interest is to unveil important latent links between different pairs of nodes. Yet a simple fundamental question of how to precisely quantify the…

Methodology · Statistics 2021-08-31 Jianqing Fan , Yingying Fan , Xiao Han , Jinchi Lv