English
Related papers

Related papers: Shifted Composition III: Local Error Framework for…

200 papers

Bayesian nonparametric statistics is an area of considerable research interest. While recently there has been an extensive concentration in developing Bayesian nonparametric procedures for model checking, the use of the Dirichlet process,…

Statistics Theory · Mathematics 2019-03-15 Luai Al-Labadi , Viskakh Patel , Kasra Vakiloroayaei , Clement Wan

Variational inference (VI) seeks to approximate a target distribution $\pi$ by an element of a tractable family of distributions. Of key interest in statistics and machine learning is Gaussian VI, which approximates $\pi$ by minimizing the…

Statistics Theory · Mathematics 2023-04-13 Michael Diao , Krishnakumar Balasubramanian , Sinho Chewi , Adil Salim

Two geometrical structures have been extensively studied for a manifold of probability distributions. One is based on the Fisher information metric, which is invariant under reversible transformations of random variables, while the other is…

Optimization and Control · Mathematics 2017-10-02 Shun-ichi Amari , Ryo Karakida , Masafumi Oizumi

We propose a new approach for metric learning by framing it as learning a sparse combination of locally discriminative metrics that are inexpensive to generate from the training data. This flexible framework allows us to naturally derive…

Machine Learning · Computer Science 2019-01-25 Yuan Shi , Aurélien Bellet , Fei Sha

In this paper, we study the strong consistency of a bias reduced kernel density estimator and derive a strongly con- sistent Kullback-Leibler divergence (KLD) estimator. As application, we formulate a goodness-of-fit test and an…

Methodology · Statistics 2018-05-21 Papa Ngom , Freedath Djibril Moussa , Jean de Dieu Nkurunziza

Langevin diffusion is a commonly used tool for sampling from a given distribution. In this work, we establish that when the target density $p^*$ is such that $\log p^*$ is $L$ smooth and $m$ strongly convex, discrete Langevin diffusion…

Machine Learning · Statistics 2017-11-02 Xiang Cheng , Peter Bartlett

Knowledge distillation is widely used to improve generalization in practice, yet its theoretical understanding remains elusive. In the standard distillation setting, a teacher model provides soft predictions to guide the training of a…

Information Theory · Computer Science 2026-05-18 Bingying Li , Haiyun He

The unadjusted Langevin algorithm is widely used for sampling from complex high-dimensional distributions. It is well known to be biased, with the bias typically scaling linearly with the dimension when measured in squared Wasserstein…

Machine Learning · Statistics 2025-09-11 Daniel Lacker , Fuzhong Zhou

Particle-based methods include a variety of techniques, such as Markov Chain Monte Carlo (MCMC) and Sequential Monte Carlo (SMC), for approximating a probabilistic target distribution with a set of weighted particles. In this paper, we…

Machine Learning · Statistics 2024-12-03 Hadi Mohasel Afshar , Gilad Francis , Sally Cripps

The stochastic gradient Langevin Dynamics is one of the most fundamental algorithms to solve sampling problems and non-convex optimization appearing in several machine learning applications. Especially, its variance reduced versions have…

Machine Learning · Computer Science 2022-11-22 Yuri Kinoshita , Taiji Suzuki

In many applications, data come with a natural ordering. This ordering can often induce local dependence among nearby variables. However, in complex data, the width of this dependence may vary, making simple assumptions such as a constant…

Statistics Theory · Mathematics 2017-12-11 Guo Yu , Jacob Bien

We address the problem of {\it adaptivity} in the framework of reproducing kernel Hilbert space (RKHS) regression. More precisely, we analyze estimators arising from a linear regularization scheme $g_\lam$. In practical applications, an…

Machine Learning · Statistics 2018-04-17 Nicole Mücke

We provide a new convergence analysis of stochastic gradient Langevin dynamics (SGLD) for sampling from a class of distributions that can be non-log-concave. At the core of our approach is a novel conductance analysis of SGLD using an…

Machine Learning · Computer Science 2021-02-24 Difan Zou , Pan Xu , Quanquan Gu

The variational framework for learning inducing variables (Titsias, 2009a) has had a large impact on the Gaussian process literature. The framework may be interpreted as minimizing a rigorously defined Kullback-Leibler divergence between…

Machine Learning · Statistics 2015-12-07 Alexander G. de G. Matthews , James Hensman , Richard E. Turner , Zoubin Ghahramani

A well-known technique in estimating probabilities of rare events in general and in information theory in particular (used, e.g., in the sphere-packing bound), is that of finding a reference probability measure under which the event of…

Information Theory · Computer Science 2014-12-23 Rami Atar , Neri Merhav

Diffusion models have emerged as a powerful paradigm for modern generative modeling, demonstrating strong potential for large language models (LLMs). Unlike conventional autoregressive (AR) models that generate tokens sequentially,…

Machine Learning · Computer Science 2026-01-09 Gen Li , Changxiao Cai

The maximum entropy principle is a powerful tool for solving underdetermined inverse problems. This paper considers the problem of discretizing a continuous distribution, which arises in various applied fields. We obtain the approximating…

Numerical Analysis · Mathematics 2020-08-05 Ken'ichiro Tanaka , Alexis Akira Toda

When we consider discretization of continuous probability distributions, it inevitably induces irreversible geometric distortion of local measure on the discretized support. While such discretziation-induced distortion is extrinsic to…

Statistical Mechanics · Physics 2026-05-28 Koretaka Yuge

Knowledge distillation (KD) has been widely adopted to compress large language models (LLMs). Existing KD methods investigate various divergence measures including the Kullback-Leibler (KL), reverse Kullback-Leibler (RKL), and…

Machine Learning · Computer Science 2024-03-03 Xiao Cui , Yulei Qin , Yuting Gao , Enwei Zhang , Zihan Xu , Tong Wu , Ke Li , Xing Sun , Wengang Zhou , Houqiang Li

Researchers from different areas have independently defined extensions of the usual weak convergence of laws of stochastic processes with the goal of adequately accounting for the flow of information. Natural approaches are convergence of…

Probability · Mathematics 2025-01-27 Daniel Bartl , Mathias Beiglböck , Gudmund Pammer , Stefan Schrott , Xin Zhang