Related papers: Probabilistic Transformers
Several variations of the Watterson estimator of variability for Next Generation Sequencing (NGS) data have been proposed in the literature. We present a unified framework for generalized Watterson estimators based on Maximum Composite…
In the 1940s, Wiener introduced a linear predictor, where the future prediction is computed by linearly combining the past data. A transformer generalizes this idea: it is a nonlinear predictor where the next-token prediction is computed by…
Mixture models are widely used in Bayesian statistics and machine learning, in particular in computational biology, natural language processing and many other fields. Variational inference, a technique for approximating intractable…
Transformers have become an important workhorse of machine learning, with numerous applications. This necessitates the development of reliable methods for increasing their transparency. Multiple interpretability methods, often based on…
We establish the characterizations of commutators of several versions of maximal functions on spaces of homogeneous type. In addition, with the aid of interpolation theory, we provide weighted version of the commutator theorems by…
Probabilistic circuits (PCs) have gained prominence in recent years as a versatile framework for discussing probabilistic models that support tractable queries and are yet expressive enough to model complex probability distributions.…
Gaussian Process (GP) models are a powerful tool in probabilistic machine learning with a solid theoretical foundation. Thanks to current advances, modeling complex data with GPs is becoming increasingly feasible, which makes them an…
Finite mixture models are statistical models which appear in many problems in statistics and machine learning. In such models it is assumed that data are drawn from random probability measures, called mixture components, which are…
Variational inference is a popular technique to approximate a possibly intractable Bayesian posterior with a more tractable one. Recently, boosting variational inference has been proposed as a new paradigm to approximate the posterior by a…
Bayesian probabilistic numerical methods are a set of tools providing posterior distributions on the output of numerical methods. The use of these methods is usually motivated by the fact that they can represent our uncertainty due to…
Transformers have had a significant impact on natural language processing and have recently demonstrated their potential in computer vision. They have shown promising results over convolution neural networks in fundamental computer vision…
In the paper is discussed complete probabilistic description of quantum systems with application to multiqubit quantum computations. In simplest case it is a set of probabilities of transitions to some fixed set of states. The probabilities…
Due to their conjugate posteriors, Gaussian process priors are attractive for estimating the drift of stochastic differential equations with continuous time observations. However, their performance strongly depends on the choice of the…
We propose a new approach to Bayesian prediction that caters for models with a large number of parameters and is robust to model misspecification. Given a class of high-dimensional (but parametric) predictive models, this new approach…
We obtain approximation results for general positive linear operators satisfying mild conditions, when acting on discontinuous functions and absolutely continuous functions having discontinuous derivatives. The upper bounds, given in terms…
The experimentally measured multiplicity distributions exhibit, after closer inspection, peculiarly enhanced void probability and oscillatory behavior of the modified combinants. We show that both these features can be used as additional…
We present a new nonparametric mixture-of-experts model for multivariate regression problems, inspired by the probabilistic k-nearest neighbors algorithm. Using a conditionally specified model, predictions for out-of-sample inputs are based…
The q-Gaussians are discussed from the point of view of variance mixtures of normals and exchangeability. For each q< 3, there is a q-Gaussian distribution that maximizes the Tsallis entropy under suitable constraints. This paper shows that…
Transformers have become the architecture of choice for learning long-range dependencies, yet their adoption in hyperspectral imaging (HSI) is still emerging. We reviewed more than 300 papers published up to 2025 and present the first…
Observables in random tensor theory are polynomials in the entries of a tensor of rank $d$ which are invariant under $U(N)^d$. It is notoriously difficult to evaluate the expectations of such polynomials, even in the Gaussian distribution.…