English
Related papers

Related papers: Deriving the Scaled-Dot-Function via Maximum Likel…

200 papers

The ability of many powerful machine learning algorithms to deal with large data sets without compromise is often hampered by computationally expensive linear algebra tasks, of which calculating the log determinant is a canonical example.…

Machine Learning · Statistics 2017-09-11 Diego Granziol , Stephen Roberts

We describe a Bayesian approach to estimating luminosity functions. We derive the likelihood function and posterior probability distribution for the luminosity function, given the observed data, and we compare the Bayesian approach with…

Astrophysics · Physics 2009-11-13 Brandon C. Kelly , Xiaohui Fan , Marianne Vestergaard

The literature on multivariate time series is, largely, limited to either models based on the multivariate Gaussian distribution or models specifically developed for a given application. In this paper we develop a general approach which is…

Methodology · Statistics 2025-12-02 Jonas Andersson , Dimitris Karlis

We explore a method of statistical estimation called Maximum Entropy on the Mean (MEM) which is based on an information-driven criterion that quantifies the compliance of a given point with a reference prior probability measure. At the core…

Statistics Theory · Mathematics 2022-12-20 Yakov Vaisbourd , Rustum Choksi , Ariel Goodwin , Tim Hoheisel , Carola-Bibiane Schönlieb

Given a sample of independent and identically distributed random variables, a novel nonparametric maximum entropy method is presented to estimate the underlying continuous univariate probability density function (pdf). Estimates are found…

Probability · Mathematics 2016-06-30 Jenny Farmer , Donald J. Jacobs

The maximum element of the vector output by the Softmax function approaches zero as the input vector size increases. Transformer-based language models rely on Softmax to compute attention scores, causing the attention distribution to…

Computation and Language · Computer Science 2025-02-03 Ken M. Nakanishi

Transformers are widely used for their ability to capture data relations in sequence processing, with great success for a wide range of static tasks. However, the computational and memory footprint of their main component, i.e., the Scaled…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Ginés Carreto Picón , Illia Oleksiienko , Lukas Hedegaard , Arian Bakhtiarnia , Alexandros Iosifidis

The softmax function combined with a cross-entropy loss is a principled approach to modeling probability distributions that has become ubiquitous in deep learning. The softmax function is defined by a lone hyperparameter, the temperature,…

Machine Learning · Computer Science 2020-10-16 Atish Agarwala , Jeffrey Pennington , Yann Dauphin , Sam Schoenholz

We propose and study properties of maximum likelihood estimators in the class of conditional transformation models. Based on a suitable explicit parameterisation of the unconditional or conditional transformation function, we establish a…

Methodology · Statistics 2019-10-22 Torsten Hothorn , Lisa Möst , Peter Bühlmann

We study a novel large dimensional approximate factor model with regime changes in the loadings driven by a latent first order Markov process. By exploiting the equivalent linear representation of the model, we first recover the latent…

Econometrics · Economics 2024-12-04 Matteo Barigozzi , Daniele Massacci

We study the problem of identifying change points in high-dimensional generalized linear models, and propose an approach based on sample-weighted empirical risk minimization. Our method, Weighted ERM, encodes priors on the change points via…

Methodology · Statistics 2026-04-14 Gabriel Arpino , Ramji Venkataramanan

To efficiently evaluate system reliability based on Monte Carlo simulation, importance sampling is used widely. The optimal importance sampling density was derived in 1950s for the deterministic simulation model, which maps an input to an…

Methodology · Statistics 2019-06-04 Quoc Dung Cao , Youngjun Choe

In this paper, we propose and study random maxout features, which are constructed by first projecting the input data onto sets of randomly generated vectors with Gaussian elements, and then outputing the maximum projection value for each…

Machine Learning · Computer Science 2015-06-15 Youssef Mroueh , Steven Rennie , Vaibhava Goel

Maximum entropy distributions with discrete support in $m$ dimensions arise in machine learning, statistics, information theory, and theoretical computer science. While structural and computational properties of max-entropy distributions…

Data Structures and Algorithms · Computer Science 2019-06-04 Damian Straszak , Nisheeth K. Vishnoi

The statistical analysis of data stemming from dynamical systems, including, but not limited to, time series, routinely relies on the estimation of information theoretical quantities, most notably Shannon entropy. To this purpose, possibly…

Information Theory · Computer Science 2021-09-01 Leonardo Ricci , Alessio Perinelli , Michele Castelluzzo

Many problems involve the use of models which learn probability distributions or incorporate randomness in some way. In such problems, because computing the true expected gradient may be intractable, a gradient estimator is used to update…

Machine Learning · Computer Science 2022-12-29 Ronan Keane , H. Oliver Gao

The block maxima method in extreme value theory consists of fitting an extreme value distribution to a sample of block maxima extracted from a time series. Traditionally, the maxima are taken over disjoint blocks of observations.…

Statistics Theory · Mathematics 2018-02-28 Axel Bücher , Johan Segers

Inferring the input parameters of simulators from observations is a crucial challenge with applications from epidemiology to molecular dynamics. Here we show a simple approach in the regime of sparse data and approximately correct models,…

Methodology · Statistics 2022-04-06 Rainier Barrett , Mehrad Ansari , Gourab Ghoshal , Andrew D White

This paper focuses on the problem of determining as large a region as possible where a function exceeds a given threshold with high probability. We assume that we only have access to a noise-corrupted version of the function and that…

Machine Learning · Statistics 2018-11-27 Andrea Zanette , Junzi Zhang , Mykel J. Kochenderfer

The optimal value function is one of the basic objects in the field of mathematical optimization, as it allows the evaluation of the variations in the cost/revenue generated while minimizing/maximizing a given function under some…

Optimization and Control · Mathematics 2021-11-29 Alain B. Zemkoho