English
Related papers

Related papers: Deriving the Scaled-Dot-Function via Maximum Likel…

200 papers

We propose a Monte Carlo algorithm to sample from high dimensional probability distributions that combines Markov chain Monte Carlo and importance sampling. We provide a careful theoretical analysis, including guarantees on robustness to…

Computation · Statistics 2019-09-18 Giacomo Zanella , Gareth Roberts

Maximum entropy method is a constructive criterion for setting up a probability distribution maximally non-committal to missing information on the basis of partial knowledge, usually stated as constrains on expectation values of some…

Statistical Mechanics · Physics 2015-07-20 Jorge Fernandez-de-Cossio , Jorge Fernandez-de-Cossio Diaz

We consider forecasting functional time series of extreme values within a generalised extreme value distribution (GEV). The GEV distribution can be characterised using the three parameters (location, scale and shape). As a result, the…

Methodology · Statistics 2020-12-22 Han Lin Shang , Ruofan Xu

High-dimensional limit theorems have been shown useful to derive tuning rules for finding the optimal scaling in random-walk Metropolis algorithms. The assumptions under which weak convergence results are proved are however restrictive: the…

Methodology · Statistics 2022-02-16 Sebastian M Schmon , Philippe Gagnon

Semantic sentence embedding models encode natural language sentences into vectors, such that closeness in embedding space indicates closeness in the semantics between the sentences. Bilingual data offers a useful signal for learning such…

Computation and Language · Computer Science 2020-11-20 John Wieting , Graham Neubig , Taylor Berg-Kirkpatrick

The Cluster Variation Method known in statistical mechanics and condensed matter is revived for weighted bipartite networks. The decomposition of a Hamiltonian through a finite number of components, whence serving to define variable…

Physics and Society · Physics 2010-03-16 Marcel Ausloos , Mircea Gligor

In most data-scientific approaches, the principle of Maximum Entropy (MaxEnt) is used to a posteriori justify some parametric model which has been already chosen based on experience, prior knowledge or computational simplicity. In a…

Methodology · Statistics 2022-06-29 Orestis Loukas , Ho Ryun Chung

Max-stable processes are a popular tool for the study of environmental extremes, and the extremal skew-$t$ process is a general model that allows for a flexible extremal dependence structure. For inference on max-stable processes with…

Methodology · Statistics 2020-04-21 B. Beranger , A. G. Stephenson , S. A. Sisson

The $k$ principal points of a random vector $\mathbf{X}$ are defined as a set of points which minimize the expected squared distance between $\mathbf{X}$ and the nearest point in the set. They are thoroughly studied in Flury (1990, 1993),…

Probability · Mathematics 2020-06-09 Juan Lucas Bali , Graciela Boente

Theoretical guarantees are established for a standard estimator in a semi-parametric finite mixture model, where each component density is modeled as a product of univariate densities under a conditional independence assumption. The focus…

Statistics Theory · Mathematics 2025-11-07 Marie Du Roy de Chaumaray , Michael Levine , Matthieu Marbac

We propose the tensorizing flow method for estimating high-dimensional probability density functions from the observed data. The method is based on tensor-train and flow-based generative modeling. Our method first efficiently constructs an…

Machine Learning · Computer Science 2022-12-02 Yinuo Ren , Hongli Zhao , Yuehaw Khoo , Lexing Ying

We review here {\it Maximum Caliber} (Max Cal), a general variational principle for inferring distributions of paths in dynamical processes and networks. Max Cal is to dynamical trajectories what the principle of {\it Maximum Entropy} (Max…

Statistical Mechanics · Physics 2018-01-17 Purushottam D. Dixit , Jason Wagoner , Corey Weistuch , Steve Pressé , Kingshuk Ghosh , Ken A. Dill

We present an algorithmic approach to estimate the value distributions of random variables of probabilistic loops whose statistical moments are (partially) known. Based on these moments, we apply two statistical methods, Maximum Entropy and…

Estimating statistical models within sensor networks requires distributed algorithms, in which both data and computation are distributed across the nodes of the network. We propose a general approach for distributed learning based on…

Machine Learning · Computer Science 2012-07-03 Qiang Liu , Alexander Ihler

Attention mechanisms have been extensively employed in various applications, including time series modeling, owing to their capacity to capture intricate dependencies; however, their utility is often constrained by quadratic computational…

Machine Learning · Computer Science 2025-11-06 Mingtao Zhang , Guoli Yang , Zhanxing Zhu , Mengzhu Wang , Xiaoying Bai

In recent years, feature selection has become a challenging problem in several machine learning fields, such as classification problems. Support Vector Machine (SVM) is a well-known technique applied in classification tasks. Various…

Machine Learning · Computer Science 2021-01-18 Asunción Jiménez-Cordero , Juan Miguel Morales , Salvador Pineda

Supplement 1 to GUM (GUM-S1) recommends the use of maximum entropy principle (MaxEnt) in determining the probability distribution of a quantity having specified properties, e.g., specified central moments. When we only know the mean value…

Data Analysis, Statistics and Probability · Physics 2012-07-20 Stefano Olivares , Matteo G. A. Paris

Monte Carlo methods are widely used importance sampling techniques for studying complex physical systems. Integrating these methods with deep learning has significantly improved efficiency and accuracy in high-dimensional problems and…

Disordered Systems and Neural Networks · Physics 2024-12-24 Yixiong Ren , Jianhui Zhou

Processing high-volume, streaming data is increasingly common in modern statistics and machine learning, where batch-mode algorithms are often impractical because they require repeated passes over the full dataset. This has motivated…

Multivariate circular observations, i.e. points on a torus are nowadays very common. Multivariate wrapped models are often appropriate to describe data points scattered on p-dimensional torus. However, statistical inference based on this…

Computation · Statistics 2018-11-16 Anahita Nodehi , Mousa Golalizadeh , Mehdi Maadooliat , Claudio Agostinelli