English
Related papers

Related papers: Multiple-Splitting Projection Test for High-Dimens…

200 papers

Data splitting divides data into two parts. One part is reserved for model selection. In some applications, the second part is used for model validation but we use this part for estimating the parameters of the chosen model. We focus on the…

Methodology · Statistics 2016-01-20 Julian J. Faraway

We study equivalence, in the context of a variable diffusion problem, between (conforming) mixed methods and (primal) nonconforming methods defined on potentially general polytopal partitions. In this first paper of a series of two, we…

Numerical Analysis · Mathematics 2026-02-18 Simon Lemaire

Jittered Sampling is a refinement of the classical Monte Carlo sampling method. Instead of picking $n$ points randomly from $[0,1]^2$, one partitions the unit square into $n$ regions of equal measure and then chooses a point randomly from…

Numerical Analysis · Mathematics 2017-04-20 Florian Pausinger , Manas Rachh , Stefan Steinerberger

Even though a train/test split of the dataset randomly performed is a common practice, could not always be the best approach for estimating performance generalization under some scenarios. The fact is that the usual machine learning…

Machine Learning · Computer Science 2022-09-09 Carlos Catania , Jorge Guerra , Juan Manuel Romero , Gabriel Caffaratti , Martin Marchetta

A new method based on the rejection sampling for finding statistical tests is proposed. This method is conceptually intuitive, easy to implement, and applicable for arbitrary dimension. To illustrate its potential applicability, three…

Methodology · Statistics 2026-03-11 Markku Kuismin

Prototypical-part methods, e.g., ProtoPNet, enhance interpretability in image recognition by linking predictions to training prototypes, thereby offering intuitive insights into their decision-making. Existing methods, which rely on a…

Computer Vision and Pattern Recognition · Computer Science 2025-04-14 Chong Wang , Yuanhong Chen , Fengbei Liu , Yuyuan Liu , Davis James McCarthy , Helen Frazer , Gustavo Carneiro

Power-enhanced tests with high-dimensional data have received growing attention in theoretical and applied statistics in recent years. Existing tests possess their respective high-power regions, and we may lack prior knowledge about the…

Methodology · Statistics 2021-10-01 Xiufan Yu , Danning Li , Lingzhou Xue , Runze Li

Predictive Mutation Testing (PMT) is a technique to predict whether a mutant will be killed by using machine learning approaches. Researchers have proposed various machine learning methods for PMT under the cross-project setting. However,…

Software Engineering · Computer Science 2020-05-26 Alireza Aghamohammadi , Seyed-Hassan Mirian-Hosseinabadi

The problem of characterizing a multivariate distribution of a random vector using examination of univariate combinations of vector components is an essential issue of multivariate analysis. The likelihood principle plays a prominent role…

Methodology · Statistics 2019-10-29 Albert Vexler

The notion of concept drift refers to the phenomenon that the distribution, which is underlying the observed data, changes over time; as a consequence machine learning models may become inaccurate and need adjustment. Many unsupervised…

Machine Learning · Computer Science 2022-02-22 Fabian Hinder , Valerie Vaquet , Barbara Hammer

In this paper we develop a methodology that we call split sampling methods to estimate high dimensional expectations and rare event probabilities. Split sampling uses an auxiliary variable MCMC simulation and expresses the expectation of…

Computation · Statistics 2013-11-04 John R. Birge , Changgee Chang , Nicholas G. Polson

Space-filling experimental designs are widely used in engineering computer experiments, where only a limited number of expensive model evaluations can be afforded. Distance-based designs such as Maximin or Minimax ensure global…

Computational Engineering, Finance, and Science · Computer Science 2026-03-30 Miroslav Vořechovský , Jan Mašek

Split-plot or repeated measures designs are frequently used for planning experiments in the life or social sciences. Typical examples include the comparison of different treatments over time, where both factors may possess an additional…

Statistics Theory · Mathematics 2017-10-13 Maria Umlauft , Marius Placzek , Frank Konietschke , Markus Pauly

We bring a control perspective to the problem of identifying paths of measures for sampling via dynamic measure transport (DMT). We highlight the fact that commonly used paths may be poor choices for DMT and connect existing methods for…

Machine Learning · Statistics 2025-11-07 Aimee Maurais , Bamdad Hosseini , Youssef Marzouk

High-dimensional changepoint inference that adapts to various change patterns has received much attention recently. We propose a simple, fast yet effective approach for adaptive changepoint testing. The key observation is that two…

Methodology · Statistics 2022-05-03 Guanghui Wang , Long Feng

In this work, we study distance metric learning (DML) for high dimensional data. A typical approach for DML with high dimensional data is to perform the dimensionality reduction first before learning the distance metric. The main…

Machine Learning · Computer Science 2015-09-16 Qi Qian , Rong Jin , Lijun Zhang , Shenghuo Zhu

Multi-parameter one-sided hypothesis test problems arise naturally in many applications. We are particularly interested in effective tests for monitoring multiple quality indices in forestry products. Our search reveals that there are many…

Statistics Theory · Mathematics 2017-03-16 Guangyu Zhu , Jiahua Chen

We propose a high dimensional mean test framework for shrinking random variables, where the underlying random variables shrink to zero as the sample size increases. By pooling observations across overlapping subsets of dimensions, we…

Methodology · Statistics 2026-02-11 Liujun Chen , Chen Zhou

In this paper we consider testing the equality of probability vectors of two independent multinomial distributions in high dimension. The classical chi-square test may have some drawbacks in this case since many of cell counts may be zero…

Statistics Theory · Mathematics 2017-11-16 Amanda Plunkett , Junyong Park

Change point testing for high-dimensional data has attracted a lot of attention in statistics and machine learning owing to the emergence of high-dimensional data with structural breaks from many fields. In practice, when the dimension is…

Methodology · Statistics 2023-12-05 Hanjia Gao , Runmin Wang , Xiaofeng Shao