English
Related papers

Related papers: Testing Parametric Distribution Family Assumptions…

200 papers

A new nonparametric approach for system identification has been recently proposed where the impulse response is modeled as the realization of a zero-mean Gaussian process whose covariance (kernel) has to be estimated from data. In this…

Optimization and Control · Mathematics 2016-11-17 Francesca Paola Carli , Tianshi Chen , Lennart Ljung

This paper introduces a robust estimation framework based solely on the copula function. We begin by introducing a family of divergence measures tailored for copulas, including the \(\alpha\)-, \(\beta\)-, and \(\gamma\)-copula divergences,…

Methodology · Statistics 2025-09-18 Shinto Eguchi , Shogo Kato

We consider the problem of hypothesis testing for discrete distributions. In the standard model, where we have sample access to an underlying distribution $p$, extensive research has established optimal bounds for uniformity testing,…

Machine Learning · Computer Science 2024-12-03 Maryam Aliakbarpour , Piotr Indyk , Ronitt Rubinfeld , Sandeep Silwal

We consider the problem of change point detection for high-dimensional distributions in a location family when the dimension can be much larger than the sample size. In change point analysis, the widely used cumulative sum (CUSUM)…

Statistics Theory · Mathematics 2021-10-14 Mengjia Yu , Xiaohui Chen

An experiment to study the entropy method for an anomaly detection system has been performed. The study has been conducted using real data generated from the distributed sensor networks at the Intel Berkeley Research Laboratory. The…

Cryptography and Security · Computer Science 2017-03-14 A. A. Waskita , H. Suhartanto , L. T. Handoko

The purpose of unconditional text generation is to train a model with real sentences, then generate novel sentences of the same quality and diversity as the training data. However, when different metrics are used for comparing the methods…

Computation and Language · Computer Science 2020-07-03 Ping Cai , Xingyuan Chen , Peng Jin , Hongjun Wang , Tianrui Li

In this paper, we propose a general method for testing inequality restrictions on nonparametric functions. Our framework includes many nonparametric testing problems in a unified framework, with a number of possible applications in auction…

Statistics Theory · Mathematics 2015-06-18 Sokbae Lee , Kyungchul Song , Yoon-Jae Whang

We propose a simple and intuitive test for arguably the most prevailing hypothesis in statistics that data are independent and identically distributed (IID), based on a newly introduced off-diagonal sequential U-process. This IID test is…

Methodology · Statistics 2025-06-30 Tongyu Li , Jonas Mueller , Fang Yao

We introduce kernel density machines (KDM), an agnostic kernel-based framework for learning the Radon-Nikodym derivative (density) between probability measures under minimal assumptions. KDM applies to general measurable spaces and avoids…

Machine Learning · Statistics 2026-03-27 Andrea Della Vecchia , Damir Filipovic , Paul Schneider

Learning probabilistic models that can estimate the density of a given set of samples, and generate samples from that density, is one of the fundamental challenges in unsupervised machine learning. We introduce a new generative model based…

Machine Learning · Computer Science 2020-06-11 Siavash A. Bigdeli , Geng Lin , Tiziano Portenier , L. Andrea Dunbar , Matthias Zwicker

Modal regression estimates the local modes of the distribution of $Y$ given $X=x$, instead of the mean, as in the usual regression sense, and can hence reveal important structure missed by usual regression methods. We study a simple…

Methodology · Statistics 2016-03-31 Yen-Chi Chen , Christopher R. Genovese , Ryan J. Tibshirani , Larry Wasserman

Many differentially private (DP) data release systems either output DP synthetic data and leave analysts to perform inference as usual, which can lead to severe miscalibration, or output a DP point estimate without a principled way to do…

Machine Learning · Computer Science 2026-03-03 Amir Asiaee , Samhita Pal

This work aims at making a comprehensive contribution in the general area of parametric inference for discretely observed diffusion processes. Established approaches for likelihood-based estimation invoke a time-discretisation scheme for…

Methodology · Statistics 2024-01-30 Yuga Iguchi , Alexandros Beskos , Matthew M. Graham

We propose a new class of semiparametric exponential family graphical models for the analysis of high dimensional mixed data. Different from the existing mixed graphical models, we allow the nodewise conditional distributions to be…

Machine Learning · Statistics 2015-10-16 Zhuoran Yang , Yang Ning , Han Liu

We introduce a powerful scan statistic and the corresponding test for detecting the presence and pinpointing the location of a change point within the distribution of a data sequence with the data elements residing in a separable metric…

Methodology · Statistics 2026-01-27 Paromita Dubey , Minxing Zheng

Symmetry plays a central role in the sciences, machine learning, and statistics. While statistical tests for the presence of distributional invariance with respect to groups have a long history, tests for conditional symmetry in the form of…

Methodology · Statistics 2025-12-12 Kenny Chiu , Alex Sharp , Benjamin Bloem-Reddy

Data collection is a fundamental problem in the scenario of big data, where the size of sampling sets plays a very important role, especially in the characterization of data structure. This paper considers the information collection process…

Information Theory · Computer Science 2018-01-23 Shanyun Liu , Rui She , Pingyi Fan

In this paper we propose a two-sample test based on copula entropy (CE). The proposed test statistic is defined as the difference between the CEs of the null hypothesis and the alternative. The estimator of the test statistic is proposed…

Methodology · Statistics 2023-07-24 Jian Ma

A massive dataset often consists of a growing number of (potentially) heterogeneous sub-populations. This paper is concerned about testing various forms of heterogeneity arising from massive data. In a general nonparametric framework, a set…

Statistics Theory · Mathematics 2016-01-26 Junwei Lu , Guang Cheng , Han Liu

We develop a nonparametric two-sample test for distributions supported on the cone of symmetric positive definite matrices. The procedure relies on the Wishart kernel density estimator (KDE) introduced by Belzile et al. (2025), whose…

Statistics Theory · Mathematics 2026-03-17 Frédéric Ouimet