English
Related papers

Related papers: Strong Asymptotic Assertions for Discrete MDL in R…

200 papers

We show a Talagrand-type concentration inequality for Multi-Task Learning (MTL), using which we establish sharp excess risk bounds for MTL in terms of distribution- and data-dependent versions of the Local Rademacher Complexity (LRC). We…

Machine Learning · Computer Science 2017-02-13 Niloofar Yousefi , Yunwen Lei , Marius Kloft , Mansooreh Mollaghasemi , Georgios Anagnostopoulos

We consider the problem of learning to optimize an unknown Markov decision process (MDP). We show that, if the MDP can be parameterized within some known function class, we can obtain regret bounds that scale with the dimensionality, rather…

Machine Learning · Statistics 2014-11-04 Ian Osband , Benjamin Van Roy

We prove weak convergence in a separable Hilbert space for estimators of high-dimensional regression coefficients, which yields asymptotic normality and enables direct use of standard asymptotic tools such as the continuous mapping theorem.…

Statistics Theory · Mathematics 2026-05-05 Kou Fujimori , Koji Tsukuda

Many functionals of interest in statistics and machine learning can be written as minimizers of expected loss functions. Such functionals are called $M$-estimands, and can be estimated by $M$-estimators -- minimizers of empirical average…

Statistics Theory · Mathematics 2024-11-27 Arunav Bhowmick , Arun Kumar Kuchibhotla

Solomonoff's central result on induction is that the posterior of a universal semimeasure M converges rapidly and with probability 1 to the true sequence generating posterior mu, if the latter is computable. Hence, M is eligible as a…

Information Theory · Computer Science 2007-08-20 Marcus Hutter , Andrej Muchnik

In this paper, under mild assumptions, we derive a law of large numbers, a central limit theorem with an error estimate, an almost sure invariance principle and a variant of Chernoff bound in finite-state hidden Markov models. These limit…

Information Theory · Computer Science 2012-04-13 Guangyue Han

The Adaptive Multilevel Splitting algorithm is a very powerful and versatile iterative method to estimate the probability of rare events, based on an interacting particle systems. In an other article, in a so-called idealized setting, the…

Probability · Mathematics 2019-10-21 Charles-Edouard Bréhier , Ludovic Goudenège , Loic Tudela

We bound the future loss when predicting any (computably) stochastic sequence online. Solomonoff finitely bounded the total deviation of his universal predictor M from the true distribution m by the algorithmic complexity of m. Here we…

Machine Learning · Computer Science 2007-07-16 Alexey Chernov , Marcus Hutter

Highly robust and efficient estimators for the generalized linear model with a dispersion parameter are proposed. The estimators are based on three steps. In the first step the maximum rank correlation estimator is used to consistently…

Methodology · Statistics 2017-03-29 Michael Amiguet , Alfio Marazzi , Marina Valdora , Victor Yohai

This paper considers a finite sample perspective on the problem of identifying an LTI system from a finite set of possible systems using trajectory data. To this end, we use the maximum likelihood estimator to identify the true system and…

Systems and Control · Electrical Eng. & Systems 2024-12-03 Nicolas Chatzikiriakos , Andrea Iannelli

The maximum likelihood estimator (MLE) is pivotal in statistical inference, yet its application is often hindered by the absence of closed-form solutions for many models. This poses challenges in real-time computation scenarios,…

Methodology · Statistics 2025-04-16 Pedro L. Ramos , Eduardo Ramos , Francisco A. Rodrigues , Francisco Louzada

We consider supervised learning (regression/classification) problems with tensor-valued input. We derive multi-linear sufficient reductions for the regression or classification problem by modeling the conditional distribution of the…

Methodology · Statistics 2025-02-28 Daniel Kapla , Efstathia Bura

In this paper we provide an asymptotic theory for the symmetric version of the Kullback--Leibler (KL) divergence. We define a estimator for this divergence and study its asymptotic properties. In particular, we prove Law of Large Numbers…

Probability · Mathematics 2024-01-31 Helder Rojas , Artem Logachov

Mixed linear regression (MLR) is a powerful model for characterizing nonlinear relationships by utilizing a mixture of linear regression sub-models. The identification of MLR is a fundamental problem, where most of the existing results…

Machine Learning · Statistics 2023-12-01 Yujing Liu , Zhixin Liu , Lei Guo

We bound the future loss when predicting any (computably) stochastic sequence online. Solomonoff finitely bounded the total deviation of his universal predictor $M$ from the true distribution $mu$ by the algorithmic complexity of $mu$. Here…

Machine Learning · Computer Science 2007-07-16 A. Chernov , M. Hutter , J. Schmidhuber

We establish decidability for the infinitely many axiomatic extensions of the commutative Full Lambek logic with weakening FLew (i.e. IMALLW) that have a cut-free hypersequent proof calculus (specifically: every analytic structural rule…

Logic in Computer Science · Computer Science 2021-04-21 A. R. Balasubramanian , Timo Lang , Revantha Ramanayake

We revisit the problem of the existence of the maximum likelihood estimate for multi-class logistic regression. We show that one method of ensuring its existence is by assigning positive probability to every class in the sample dataset. The…

Machine Learning · Computer Science 2024-05-09 Dwight Nwaigwe , Marek Rychlik

Recent works have proposed various explanations for the ability of modern large language models (LLMs) to perform in-context prediction. We propose an alternative conceptual viewpoint from an information-geometric and statistical…

Information Theory · Computer Science 2026-02-23 Sreejith Sreekumar , Nir Weinberger

State-of-the-art neural networks can be trained to become remarkable solutions to many problems. But while these architectures can express symbolic, perfect solutions, trained models often arrive at approximations instead. We show that the…

Machine Learning · Computer Science 2025-09-09 Matan Abudy , Orr Well , Emmanuel Chemla , Roni Katzir , Nur Lan

The Minimum Description Length (MDL) principle offers a formal framework for applying Occam's razor in machine learning. However, its application to neural networks such as Transformers is challenging due to the lack of a principled,…

Machine Learning · Computer Science 2026-03-04 Peter Shaw , James Cohan , Jacob Eisenstein , Kristina Toutanova
‹ Prev 1 4 5 6 7 8 10 Next ›