English
Related papers

Related papers: Tests for high dimensional data based on means, sp…

200 papers

When the outcome of interest is semicontinuous and collected longitudinally, efficient testing can be difficult. Daily rainfall data is an excellent example which we use to illustrate the various challenges. Even under the simplest…

Methodology · Statistics 2018-07-10 Harlan Campbell

We propose exploiting symmetries (exact or approximate) of the Standard Model (SM) to search for physics Beyond the Standard Model (BSM) using the data-directed paradigm (DDP). Symmetries are very powerful because they provide two samples…

High Energy Physics - Phenomenology · Physics 2022-06-09 Mattias Birman , Benjamin Nachman , Raphael Sebbah , Gal Sela , Ophir Turetz , Shikma Bressler

Spatial data is playing an emerging role in new technologies such as web and mobile mapping and Geographic Information Systems (GIS). Important decisions in political, social and many other aspects of modern human life are being made using…

Databases · Computer Science 2016-05-17 Bagher Saberi , Nasser Ghadiri

We consider multivariate two-sample tests of means, where the location shift between the two populations is expected to be related to a known graph structure. An important application of such tests is the detection of differentially…

Quantitative Methods · Quantitative Biology 2014-05-16 Laurent Jacob , Pierre Neuvial , Sandrine Dudoit

This article deals with the problem of testing conditional independence between two random vectors ${\bf X}$ and ${\bf Y}$ given a confounding random vector ${\bf Z}$. Several authors have considered this problem for multivariate data.…

Statistics Theory · Mathematics 2025-09-16 Bilol Banerjee

Evaluating spatial patterns in data is an integral task across various domains, including geostatistics, astronomy, and spatial tissue biology. The analysis of transcriptomics data in particular relies on methods for detecting…

Methodology · Statistics 2025-05-26 Katharina Limbeck , Bastian Rieck

Non-deterministic measurements are common in real-world scenarios: the performance of a stochastic optimization algorithm or the total reward of a reinforcement learning agent in a chaotic environment are just two examples in which…

Machine Learning · Statistics 2022-08-31 Etor Arza , Josu Ceberio , Ekhiñe Irurozki , Aritz Pérez

In this article, we investigate the asymptotic properties of Bayesian multiple testing procedures under general dependent setup, when the sample size and the number of hypotheses both tend to infinity. Specifically, we investigate strong…

Statistics Theory · Mathematics 2020-05-14 Noirrit Kiran Chandra , Sourabh Bhattacharya

The comparison of alternative rankings of a set of items is a general and prominent task in applied statistics. Predictor variables are ranked according to magnitude of association with an outcome, prediction models rank subjects according…

The advent of modern data collection and processing techniques has seen the size, scale, and complexity of data grow exponentially. A seminal step in leveraging these rich datasets for downstream inference is understanding the…

Applications · Statistics 2024-07-30 Zeyi Wang , Eric Bridgeford , Shangsi Wang , Joshua T. Vogelstein , Brian Caffo

Motivated by the increasing use of kernel-based metrics for high-dimensional and large-scale data, we study the asymptotic behavior of kernel two-sample tests when the dimension and sample sizes both diverge to infinity. We focus on the…

Statistics Theory · Mathematics 2024-10-31 Jian Yan , Xianyang Zhang

Robust statistics aims to compute quantities to represent data where a fraction of it may be arbitrarily corrupted. The most essential statistic is the mean, and in recent years, there has been a flurry of theoretical advancement for…

Machine Learning · Statistics 2025-02-18 Cullen Anderson , Jeff M. Phillips

Two semimetrics on probability distributions are proposed, given as the sum of differences of expectations of analytic functions evaluated at spatial or frequency locations (i.e, features). The features are chosen so as to maximize the…

Machine Learning · Statistics 2016-10-31 Wittawat Jitkrittum , Zoltan Szabo , Kacper Chwialkowski , Arthur Gretton

We propose novel methodology for testing equality of model parameters between two high-dimensional populations. The technique is very general and applicable to a wide range of models. The method is based on sample splitting: the data is…

Methodology · Statistics 2013-01-17 Nicolas Städler , Sach Mukherjee

In this paper we consider testing the equality of probability vectors of two independent multinomial distributions in high dimension. The classical chi-square test may have some drawbacks in this case since many of cell counts may be zero…

Statistics Theory · Mathematics 2017-11-16 Amanda Plunkett , Junyong Park

A number of applications require two-sample testing on ranked preference data. For instance, in crowdsourcing, there is a long-standing question of whether pairwise comparison data provided by people is distributed similar to…

Machine Learning · Statistics 2020-11-20 Charvi Rastogi , Sivaraman Balakrishnan , Nihar B. Shah , Aarti Singh

This article conducts a large dimensional study of a simple yet quite versatile classification model, encompassing at once multi-task and semi-supervised learning, and taking into account uncertain labeling. Using tools from random matrix…

Machine Learning · Statistics 2024-02-22 Victor Leger , Romain Couillet

In this paper, we consider testing the martingale difference hypothesis for high-dimensional time series. Our test is built on the sum of squares of the element-wise max-norm of the proposed matrix-valued nonlinear dependence measure at…

Econometrics · Economics 2023-11-15 Jinyuan Chang , Qing Jiang , Xiaofeng Shao

When applying multivariate extreme value statistics to analyze tail risk in compound events defined by a multivariate random vector, one often assumes that all dimensions share the same extreme value index. While such an assumption can be…

Methodology · Statistics 2026-02-16 Liujun Chen , Chen Zhou

It is a standard assumption that datasets in high dimension have an internal structure which means that they in fact lie on, or near, subsets of a lower dimension. In many instances it is important to understand the real dimension of the…

Machine Learning · Statistics 2025-07-21 James A. D. Binnie , Paweł Dłotko , John Harvey , Jakub Malinowski , Ka Man Yim
‹ Prev 1 8 9 10 Next ›