English
Related papers

Related papers: Identification testing via sample splitting -- an …

200 papers

In modern data analysis, statistical efficiency improvement is expected via effective collaboration among multiple data holders with non-shared data. In this article, we propose a collaborative score-type test (CST) for testing linear…

Methodology · Statistics 2025-04-30 Yifan Gu , Hanfang Yang , Songshan Yang , Hui Zou

This paper reconsiders the problem of testing the equality of two unspecified continuous distributions. The framework, which we propose, allows for readable and insightful data visualisation and helps to understand and quantify how two…

Methodology · Statistics 2025-03-04 Bogdan Ćmiel , Teresa Ledwina

Data used for training structural health monitoring (SHM) systems are expensive and often impractical to obtain, particularly labelled data. Population-based SHM presents a potential solution to this issue by considering the available data…

Machine Learning · Computer Science 2025-07-29 J. Poole , P. Gardner , A. J. Hughes , N. Dervilis , R. S. Mills , T. A. Dardeno , K. Worden

Maximum Mean Discrepancy (MMD) is a widely used concept in machine learning research which has gained popularity in recent years as a highly effective tool for comparing (finite-dimensional) distributions. Since it is designed as a…

Machine Learning · Statistics 2025-06-03 Andrew Alden , Blanka Horvath , Zacharia Issa

We study the problem of testing whether the missing values of a potentially high-dimensional dataset are Missing Completely at Random (MCAR). We relax the problem of testing MCAR to the problem of testing the compatibility of a collection…

Statistics Theory · Mathematics 2024-12-13 Alberto Bordino , Thomas B. Berrett

Semi-Supervised Learning (SSL) is implemented when algorithms are trained on both labeled and unlabeled data. This is a very common application of ML as it is unrealistic to obtain a fully labeled dataset. Researchers have tackled three…

Machine Learning · Computer Science 2023-08-16 Jason Lu , Michael Ma , Huaze Xu , Zixi Xu

Given a set of incomplete observations, we study the nonparametric problem of testing whether data are Missing Completely At Random (MCAR). Our first contribution is to characterise precisely the set of alternatives that can be…

Statistics Theory · Mathematics 2022-05-19 Thomas B Berrett , Richard J Samworth

Markov parameters play a key role in system identification. There exists many algorithms where these parameters are estimated using least-squares in a first, pre-processing, step, including subspace identification and multi-step…

Systems and Control · Electrical Eng. & Systems 2024-05-08 Jiabao He , Cristian R. Rojas , Håkan Hjalmarsson

The vector autoregressive (VAR) model has been widely used for modeling temporal dependence in a multivariate time series. For large (and even moderate) dimensions, the number of AR coefficients can be prohibitively large, resulting in…

Applications · Statistics 2013-10-21 Richard A. Davis , Pengfei Zang , Tian Zheng

In this paper, we propose a model averaging approach for addressing model uncertainty in the context of partial linear functional additive models. These models are designed to describe the relation between a response and mixed-types of…

Methodology · Statistics 2023-06-12 Shishi Liu , Jingxiao Zhang

High-dimensional vector autoregression with measurement error is frequently encountered in a large variety of scientific and business applications. In this article, we study statistical inference of the transition matrix under this model.…

Methodology · Statistics 2020-09-18 Xiang Lyu , Jian Kang , Lexin Li

We formulate nonparametric and semiparametric hypothesis testing of multivariate stationary linear time series in a unified fashion and propose new test statistics based on estimators of the spectral density matrix. The limiting…

Statistics Theory · Mathematics 2009-09-03 Yoshihiro Yajima , Yasumasa Matsuda

In many statistical problems, the data distribution is specified through a generative process for which the likelihood function is analytically intractable, yet inference on the associated model parameters remains of primary interest. We…

Methodology · Statistics 2026-04-01 Haoyu Jiang , Yuexi Wang , Yun Yang

Identification-robust hypothesis tests are commonly based on the continuous updating GMM objective function. When the number of moment conditions grows proportionally with the sample size, the large-dimensional weighting matrix prohibits…

Econometrics · Economics 2025-10-10 Tom Boot , Johannes W. Ligtenberg

We propose a structural vector autoregressive model with a new and flexible specification of the volatility process which we call Sparse Heterogeneous Markov-Switching Heteroskedasticity. In this model, the conditional variance of each…

Econometrics · Economics 2026-03-18 Fei Shang , Tomasz Woźniak

A novel variational inference based resampling framework is proposed to evaluate the robustness and generalization capability of deep learning models with respect to distribution shift. We use Auto Encoding Variational Bayes to find a…

Machine Learning · Computer Science 2019-10-29 Xudong Sun , Alexej Gossmann , Yu Wang , Bernd Bischl

By studying the family of $p$-dimensional scale mixtures, this paper shows for the first time a non trivial example where the eigenvalue distribution of the corresponding sample covariance matrix {\em does not converge} to the celebrated…

Methodology · Statistics 2017-05-16 Weiming Li , Jianfeng Yao

Given a random sample of observations, mixtures of normal densities are often used to estimate the unknown continuous distribution from which the data come. Here we propose the use of this semiparametric framework for testing symmetry about…

Methodology · Statistics 2012-04-23 Silvia Bacci , Francesco Bartolucci

In regression models for spatial data, it is often assumed that the marginal effects of covariates on the response are constant over space. In practice, this assumption might often be questionable. In this article, we show how a Gaussian…

Methodology · Statistics 2020-11-13 Jakob A. Dambon , Fabio Sigrist , Reinhard Furrer

A hybrid censoring scheme is a mixture of Type-I and Type-II censoring schemes. We study the estimation of parameters of weighted exponential distribution based on Type-II hybrid censored data. By applying EM algorithm, maximum likelihood…

Statistics Theory · Mathematics 2012-03-02 Akram Kohansal , Saeid Rezakhah