English
Related papers

Related papers: Deriving the Scaled-Dot-Function via Maximum Likel…

200 papers

This paper studies the problem of estimation from relative measurements in a graph, in which a vector indexed over the nodes has to be reconstructed from pairwise measurements of differences between its components associated to nodes…

Systems and Control · Computer Science 2018-07-27 Chiara Ravazzi , Nelson P. K. Chan , Paolo Frasca

Over the past decades, there has been a surge of interest in studying low-dimensional structures within high-dimensional data. Statistical factor models $-$ i.e., low-rank plus diagonal covariance structures $-$ offer a powerful framework…

Machine Learning · Statistics 2025-05-20 Daniel Cederberg

We consider the problem of finding an input to a stochastic black box function such that the scalar output of the black box function is as close as possible to a target value in the sense of the expected squared error. While the…

Machine Learning · Computer Science 2023-05-16 Johannes G. Hoffer , Sascha Ranftl , Bernhard C. Geiger

We prove a formula for the evaluation of expectations containing a scalar function of a Gaussian random vector multiplied by a product of the random vector components, each one raised to a non-negative integer power. Some of the powers…

Probability · Mathematics 2025-08-21 Konstantinos Mamis

Mobility entropy is proposed to measure predictability of human movements, based on which, the upper and lower bound of prediction accuracy is deduced, but corresponding mathematical expressions of prediction accuracy keeps yet open. In…

Social and Information Networks · Computer Science 2019-01-29 Lu Liu , Wuyang Zhou , Sihai Zhang , Wei Cai

Dot product kernels, such as polynomial and exponential (softmax) kernels, are among the most widely used kernels in machine learning, as they enable modeling the interactions between input features, which is crucial in applications like…

Machine Learning · Statistics 2024-08-14 Jonas Wacker , Motonobu Kanagawa , Maurizio Filippone

In this paper, we address the classical problem of maximum-likelihood (ML) detection of data in the presence of random phase noise. We consider a system, where the random phase noise affecting the received signal is first compensated by a…

Information Theory · Computer Science 2016-11-15 Rajet Krishnan , M. Reza Khanzadi , Thomas Eriksson , Tommy Svensson

We estimate the global minimum variance (GMV) portfolio in the high-dimensional case using results from random matrix theory. This approach leads to a shrinkage-type estimator which is distribution-free and it is optimal in the sense of…

Statistical Finance · Quantitative Finance 2023-04-19 Taras Bodnar , Nestor Parolya , Wolfgang Schmid

We find large deviations rates for consensus-based distributed inference for directed networks. When the topology is deterministic, we establish the large deviations principle and find exactly the corresponding rate function, equal at all…

Information Theory · Computer Science 2016-06-29 Dragana Bajović , José M. F. Moura , João Xavier , Bruno Sinopoli

Probability density function estimation with weighted samples is the main foundation of all adaptive importance sampling algorithms. Classically, a target distribution is approximated either by a non-parametric model or within a parametric…

Machine Learning · Computer Science 2023-10-16 Julien Demange-Chryst , François Bachoc , Jérôme Morio , Timothé Krauth

In this paper we study the distribution of the scaled largest eigenvalue of complexWishart matrices, which has diverse applications both in statistics and wireless communications. Exact expressions, valid for any matrix dimensions, have…

Information Theory · Computer Science 2012-02-06 Lu Wei , Olav Tirkkonen , Prathapasinghe Dharmawansa , Matthew McKay

We study the entanglement entropy of a random tensor network (RTN) using tools from free probability theory. Random tensor networks are simple toy models that help the understanding of the entanglement behavior of a boundary region in the…

Quantum Physics · Physics 2024-07-04 Khurshed Fitter , Faedi Loulidi , Ion Nechita

The principle of maximum entropy is a broadly applicable technique for computing a distribution with the least amount of information possible constrained to match empirical data, for instance, feature expectations. We seek to generalize…

Information Theory · Computer Science 2022-05-30 Kenneth Bogert

The principle of maximum entropy (Maxent) is often used to obtain prior probability distributions as a method to obtain a Gibbs measure under some restriction giving the probability that a system will be in a certain state compared to the…

Information Theory · Computer Science 2019-06-26 Hector Zenil , Narsis A. Kiani , Jesper Tegnér

Following [1], the aim of this paper is to analyze the relative weighted entropy involving the central moments weight functions. We compare the standard relative entropy with the weighted case in two particular forms of Gaussian…

Information Theory · Computer Science 2015-06-23 Salimeh Yasaei Sekeh , Adriano Polpo

We introduce MESSY estimation, a Maximum-Entropy based Stochastic and Symbolic densitY estimation method. The proposed approach recovers probability density functions symbolically from samples using moments of a Gradient flow in which the…

Machine Learning · Computer Science 2024-02-13 Tony Tohme , Mohsen Sadr , Kamal Youcef-Toumi , Nicolas G. Hadjiconstantinou

Our work introduces an approach for estimating the contribution of attachment mechanisms to the formation of growing networks. We present a generic model in which growth is driven by the continuous attachment of new nodes according to…

Probability · Mathematics 2019-02-20 Jan Medina , Jorge Finke , Camilo Rocha

This paper derives analytic expressions for the expected value of sample information (EVSI), the expected value of distribution information (EVDI), and the optimal sample size when data consists of independent draws from a bounded sequence…

Methodology · Statistics 2024-03-13 Adam Fleischhacker , Pak-Wing Fok , Mokshay Madiman , Nan Wu

The direct sampling method proposed by Walker et al. (JCGS 2011) can generate draws from weighted distributions possibly having intractable normalizing constants. The method may be of interest as a useful tool in situations which require…

Computation · Statistics 2024-01-19 Andrew M. Raim

While deep generative models have succeeded in image processing, natural language processing, and reinforcement learning, training that involves discrete random variables remains challenging due to the high variance of its gradient…

Machine Learning · Computer Science 2022-06-16 Ting-Han Fan , Ta-Chung Chi , Alexander I. Rudnicky , Peter J. Ramadge