English
Related papers

Related papers: Deriving the Scaled-Dot-Function via Maximum Likel…

200 papers

We consider the estimation of the value of a linear functional of the slope parameter in functional linear regression, where scalar responses are modeled in dependence of random functions. The theory in this paper covers in particular…

Statistics Theory · Mathematics 2011-12-19 J. Johannes , R. Schenk

A common approach for modeling extremes, such as peak flow or high temperatures, is the three-parameter Generalized Extreme-Value distribution. This is typically fit to extreme observations, here defined as maxima over disjoint blocks. This…

Applications · Statistics 2025-10-07 Nathan Huet , Ilaria Prosdocimi

According to the concept of typicality, an ensemble average can be accurately approximated by an expectation value with respect to a single pure state drawn at random from a high-dimensional Hilbert space. This random-vector approximation,…

Statistical Mechanics · Physics 2020-05-22 J. Schnack , J. Richter , T. Heitmann , J. Richter , R. Steinigeweg

For complex latent variable models, the likelihood function is not available in closed form. In this context, a popular method to perform parameter estimation is Importance Weighted Variational Inference. It essentially maximizes the…

Statistics Theory · Mathematics 2025-01-16 Badr-Eddine Cherief-Abdellatif , Randal Douc , Arnaud Doucet , Hugo Marival

Semi-continuous data comes from a distribution that is a mixture of the point mass at zero and a continuous distribution with support on the positive real line. A clear example is the daily rainfall data. In this paper, we present a novel…

Methodology · Statistics 2021-06-17 Sai K. Popuri , Nagaraj K. Neerchal , Amita Mehta , Ahmad Mousavi

Entropy estimation is of practical importance in information theory and statistical science. Many existing entropy estimators suffer from fast growing estimation bias with respect to dimensionality, rendering them unsuitable for…

Information Theory · Computer Science 2023-08-22 Ziqiao Ao , Jinglai Li

A common statistical situation concerns inferring an unknown distribution Q(x) from a known distribution P(y), where X (dimension n), and Y (dimension m) have a known functional relationship. Most commonly, n<m, and the task is relatively…

Quantitative Methods · Quantitative Biology 2016-02-01 Jayajit Das , Sayak Mukherjee , Susan E. Hodge

We exploit the idea to use the maximal-entropy method, successfully tested in information theory and statistical thermodynamics, to determine approximating function's coefficients and squared errors' weights simultaneously as output of one…

Numerical Analysis · Mathematics 2021-03-04 Domenico Giordano , Felice Iavernaro

Scaled dot-product attention applies a softmax function on the scaled dot-product of queries and keys to calculate weights and then multiplies the weights and values. In this work, we study how to improve the learning of scaled dot-product…

Computer Vision and Pattern Recognition · Computer Science 2023-03-17 Kaikai Zhao , Norimichi Ukita

This paper studies fundamental aspects of modelling data using multivariate Watson distributions. Although these distributions are natural for modelling axially symmetric data (i.e., unit vectors where $\pm \x$ are equivalent), for…

Computation · Statistics 2012-05-28 Suvrit Sra , Dmitrii Karp

We propose the Value Gradient Sampler (VGS), a diffusion sampler parameterized by value functions. VGS generates samples from an unnormalized target density (i.e., energy) by evolving randomly initialized particles along the gradient of the…

Machine Learning · Computer Science 2026-04-01 Himchan Hwang , Hyeokju Jeong , Dong Kyu Shin , Che-Sang Park , Sehee Kweon , Sangwoong Yoon , Frank Chongwoo Park

Increasing the size of a Transformer does not always lead to enhanced performance. This phenomenon cannot be explained by the empirical scaling laws. Furthermore, the model's enhanced performance is closely associated with its memorization…

Machine Learning · Computer Science 2024-12-02 Xueyan Niu , Bo Bai , Lei Deng , Wei Han

With an eye towards human-centered automation, we contribute to the development of a systematic means to infer features of human decision-making from behavioral data. Motivated by the common use of softmax selection in models of human…

Optimization and Control · Mathematics 2015-09-01 Paul Reverdy , Naomi E. Leonard

Following the success of dot-product attention in Transformers, numerous approximations have been recently proposed to address its quadratic complexity with respect to the input length. However, all approximations thus far have ignored the…

Machine Learning · Computer Science 2021-03-19 Ankit Gupta , Jonathan Berant

Effective utilization of flexible loads for grid services, while satisfying end-user preferences and constraints, requires an accurate estimation of the aggregated predictive flexibility offered by the electrical loads. Virtual battery (VB)…

Systems and Control · Electrical Eng. & Systems 2020-03-20 Indrasis Chakraborty , Sai Pushpak Nandanoori , Soumya Kundu , Karanjit Kalsi

The majority of machine learning methods can be regarded as the minimization of an unavailable risk function. To optimize the latter, given samples provided in a streaming fashion, we define a general stochastic Newton algorithm and its…

Statistics Theory · Mathematics 2023-06-30 Claire Boyer , Antoine Godichon-Baggioni

In this paper, we propose an optimization-based mechanism to explain power law distributions, where the function that the optimization process is seeking to optimize is derived mathematically, then the behavior and interpretation of this…

Physics and Society · Physics 2018-12-27 A. M. Khalili

This paper estimates the break point for large-dimensional factor models with a single structural break in factor loadings at a common unknown date. First, we propose a quasi-maximum likelihood (QML) estimator of the change point based on…

Econometrics · Economics 2021-04-01 Jiangtao Duan , Jushan Bai , Xu Han

In dealing with high-dimensional data, factor models are often used for reducing dimensions and extracting relevant information. The spectrum of covariance matrices from power data exhibits two aspects: 1) bulk, which arises from random…

Applications · Statistics 2019-10-22 Xin Shi , Robert Qiu

The hidden variable formalism (based on the assumption of some intrinsic node parameters) turned out to be a remarkably efficient and powerful approach in describing and analyzing the topology of complex networks. Owing to one of its most…

Physics and Society · Physics 2019-08-13 Sámuel G. Balogh , Péter Pollner , Gergely Palla
‹ Prev 1 3 4 5 6 7 10 Next ›