English
Related papers

Related papers: Momentum SVGD-EM for Accelerated Maximum Marginal …

200 papers

Recently, {\it stochastic momentum} methods have been widely adopted in training deep neural networks. However, their convergence analysis is still underexplored at the moment, in particular for non-convex optimization. This paper fills the…

Optimization and Control · Mathematics 2016-05-06 Tianbao Yang , Qihang Lin , Zhe Li

Normal mean-variance mixture distributions are widely applied to simplify a model's implementation and improve their computational efficiency under the Maximum Likelihood (ML) approach. Especially for distributions with normal mean-variance…

Methodology · Statistics 2015-06-18 Thanakorn Nitithumbundit , Jennifer S. K. Chan

Stein variational gradient descent (SVGD) and its variants have shown promising successes in approximate inference for complex distributions. In practice, we notice that the kernel used in SVGD-based methods has a decisive effect on the…

Machine Learning · Computer Science 2022-11-29 Qingzhong Ai , Shiyu Liu , Lirong He , Zenglin Xu

Identifying important features linked to a response variable is a fundamental task in various scientific domains. This article explores statistical inference for simulated Markov random fields in high-dimensional settings. We introduce a…

Machine Learning · Statistics 2024-01-23 Haoyu Wei , Xiaoyu Lei , Yixin Han , Huiming Zhang

Due to its heavy-tailed and fully parametric form, the multivariate generalized Gaussian distribution (MGGD) has been receiving much attention for modeling extreme events in signal and image processing applications. Considering the…

Applications · Statistics 2017-02-27 F. Pascal , L. Bombrun , J. Y. Tourneret , Y. Berthoumieu

Zhu and Melnykov (2018) develop a model to fit mixture models when the components are derived from the Manly transformation. Their EM algorithm utilizes Nelder-Mead optimization in the M-step to update the skew parameter,…

Machine Learning · Statistics 2025-08-04 Katharine M. Clark , Paul D. McNicholas

Estimators derived from a divergence criterion such as $\varphi-$divergences are generally more robust than the maximum likelihood ones. We are interested in particular in the so-called MD$\varphi$DE, an estimator built using a dual…

Computation · Statistics 2016-06-14 Diaa Al Mohamad , Michel Broniatowski

Nonlinear Mixed Effects models (NLME) models are widely used in pharmacometrics and related fields to analyze hierarchical and longitudinal data. However, as the number of parameters and random effects increases, traditional methods for…

Methodology · Statistics 2026-04-30 Mohamed Tarek , Pedro Afonso

Accelerated gradient methods play a central role in optimization, achieving optimal rates in many settings. While many generalizations and extensions of Nesterov's original acceleration method have been proposed, it is not yet clear what is…

Optimization and Control · Mathematics 2022-06-08 Andre Wibisono , Ashia C. Wilson , Michael I. Jordan

The EM (Expectation-Maximization) algorithm is regarded as an MM (Majorization-Minimization) algorithm for maximum likelihood estimation of statistical models. Expanding this view, this paper demonstrates that by choosing an appropriate…

Optimization and Control · Mathematics 2026-02-12 Kensuke Asai , Jun-ya Gotoh

We develop interacting particle algorithms for learning latent variable models with energy-based priors. To do so, we leverage recent developments in particle-based methods for solving maximum marginal likelihood estimation (MMLE) problems.…

Machine Learning · Statistics 2025-10-15 Joanna Marks , Tim Y. J. Wang , O. Deniz Akyildiz

In various practical situations, we encounter data from stochastic processes which can be efficiently modelled by an appropriate parametric model for subsequent statistical analyses. Unfortunately, the most common estimation and inference…

Methodology · Statistics 2022-04-12 Rohan Hore , Abhik Ghosh

We study the problem of computing the maximum likelihood estimator (MLE) of multivariate log-concave densities. Our main result is the first computationally efficient algorithm for this problem. In more detail, we give an algorithm that, on…

Data Structures and Algorithms · Computer Science 2018-12-14 Ilias Diakonikolas , Anastasios Sidiropoulos , Alistair Stewart

This paper deals with nonparametric maximum likelihood estimation for Gaussian locally stationary processes. Our nonparametric MLE is constructed by minimizing a frequency domain likelihood over a class of functions. The asymptotic behavior…

Statistics Theory · Mathematics 2011-11-10 Rainer Dahlhaus , Wolfgang Polonik

Orthogonal group synchronization aims to recover orthogonal group elements from their noisy pairwise measurements. It has found numerous applications including computer vision, imaging science, and community detection. Due to the orthogonal…

Statistics Theory · Mathematics 2025-02-21 Ziliang Samuel Zhong , Shuyang Ling

Asynchronous stochastic gradient descent (ASGD) is a popular parallel optimization algorithm in machine learning. Most theoretical analysis on ASGD take a discrete view and prove upper bounds for their convergence rates. However, the…

Machine Learning · Statistics 2018-05-09 Li He , Qi Meng , Wei Chen , Zhi-Ming Ma , Tie-Yan Liu

The fluctuation effect of gradient expectation and variance caused by parameter update between consecutive iterations is neglected or confusing by current mainstream gradient optimization algorithms.Using this fluctuation effect, combined…

Machine Learning · Statistics 2022-02-23 Aixiang , Chen , Jinting Zhang , Zanbo Zhang , Zhihong Li

In coherent imaging, speckle is statistically modeled as multiplicative noise, posing a fundamental challenge for image reconstruction. While maximum likelihood estimation (MLE) provides a principled framework for speckle mitigation, its…

Computer Vision and Pattern Recognition · Computer Science 2026-02-12 Xi Chen , Arian Maleki , Shirin Jalali

We present a coupled system of ODEs which, when discretized with a constant time step/learning rate, recovers Nesterov's accelerated gradient descent algorithm. The same ODEs, when discretized with a decreasing learning rate, leads to novel…

Optimization and Control · Mathematics 2020-09-02 Maxime Laborde , Adam M. Oberman

Estimation of generalized linear mixed models (GLMMs) with non-nested random effects structures requires approximation of high-dimensional integrals. Many existing methods are tailored to the low-dimensional integrals produced by nested…

Computation · Statistics 2014-04-01 Andrew T. Karl , Yan Yang , Sharon L. Lohr