English
Related papers

Related papers: Codage arithmetique pour la description d'une dist…

200 papers

For linear time-invariant systems with uncertain parameters belonging to a finite set, we present a purely deterministic approach to multiple-model estimation and propose an algorithm based on the minimax criterion using constrained…

Optimization and Control · Mathematics 2022-07-18 Olle Kjellqvist , Anders Rantzer

In the era of big data, it is necessary to split extremely large data sets across multiple computing nodes and construct estimators using the distributed data. When designing distributed estimators, it is desirable to minimize the amount of…

Statistics Theory · Mathematics 2022-04-25 Azeem Zaman , Botond Szabó

This work introduces interpretable regional descriptors, or IRDs, for local, model-agnostic interpretations. IRDs are hyperboxes that describe how an observation's feature values can be changed without affecting its prediction. They justify…

Machine Learning · Statistics 2023-11-09 Susanne Dandl , Giuseppe Casalicchio , Bernd Bischl , Ludwig Bothmann

This paper proposes an information theory approach to estimate the number of changepoints and their locations in a climatic time series. A model is introduced that has an unknown number of changepoints and allows for series…

Applications · Statistics 2010-10-08 QiQi Lu , Robert Lund , Thomas C. M. Lee

Statistical inference is considered for variables of interest, called primary variables, when auxiliary variables are observed along with the primary variables. We consider the setting of incomplete data analysis, where some primary…

Methodology · Statistics 2019-03-27 Shinpei Imori , Hidetoshi Shimodaira

Linear mixed effects models are highly flexible in handling a broad range of data types and are therefore widely used in applications. A key part in the analysis of data is model selection, which often aims to choose a parsimonious model…

Methodology · Statistics 2013-06-12 Samuel Müller , J. L. Scealy , A. H. Welsh

Although multi-scales representation learning enables elastic-dimension embeddings, nested subspaces often suffer from dimensional redundancy and spectral collapse. To address this, we introduce MIC, a framework that optimizes the geometric…

Machine Learning · Computer Science 2026-05-29 Dang Hong Nguyen , Nhi Ngoc-Yen Nguyen , Huy-Hieu Pham

The question of selecting the "best" amongst different choices is a common problem in statistics. In drug development, our motivating setting, the question becomes, for example: what is the dose that gives me a pre-specified risk of…

Statistics Theory · Mathematics 2018-03-15 Pavel Mozgunov , Thomas Jaki

Symbolic Data Analysis works with variables for which each unit or class of units takes a finite set of values/categories, an interval or a distribution (an histogram, for instance). When to each observation corresponds an empirical…

Methodology · Statistics 2013-05-01 Sónia Dias , Paula Brito

Information set decoding (ISD) algorithms are the best known procedures to solve the decoding problem for general linear codes. These algorithms are hence used for codes without a visible structure, or for which efficient decoders…

Cryptography and Security · Computer Science 2021-02-19 Violetta Weger , Massimo Battaglioni , Paolo Santini , Franco Chiaraluce , Marco Baldi , Edoardo Persichetti

We introduce and define the novel problem of multi-distribution information retrieval (IR) where given a query, systems need to retrieve passages from within multiple collections, each drawn from a different distribution. Some of these…

Information Retrieval · Computer Science 2023-06-23 Soumya Chatterjee , Omar Khattab , Simran Arora

Distributed Image Compression (DIC) is crucial for multi-view transmission, especially when operating at extremely low bitrates (< 0.1 bpp). Its core challenge is effectively utilizing side information to achieve high-quality reconstruction…

Computer Vision and Pattern Recognition · Computer Science 2026-05-22 Guojun Xu , Mingyang Zhang , Jianwen Xiang , Cheng Tan , Yanchao Yang , Junwei Zhou

We introduce the Randomized Dependence Coefficient (RDC), a measure of non-linear dependence between random variables of arbitrary dimension based on the Hirschfeld-Gebelein-R\'enyi Maximum Correlation Coefficient. RDC is defined in terms…

Machine Learning · Statistics 2013-06-04 David Lopez-Paz , Philipp Hennig , Bernhard Schölkopf

We emphasize that it is possible to improve the principle of unbiased risk estimation for model selection by addressing excess risk deviations in the design of penalization procedures. Indeed, we propose a modification of Akaike's…

Statistics Theory · Mathematics 2018-07-23 Adrien Saumard , Fabien Navarro

Recent sequential pattern mining methods have used the minimum description length (MDL) principle to define an encoding scheme which describes an algorithm for mining the most compressing patterns in a database. We present a novel…

Machine Learning · Statistics 2016-11-14 Jaroslav Fowkes , Charles Sutton

Several applications in communication, control, and learning require approximating target distributions to within small informational divergence (I-divergence). The additional requirement of invertibility usually leads to using encoders…

Information Theory · Computer Science 2020-10-22 Patrick Schulte , Rana Ali Amjad , Thomas Wiegart , Gerhard Kramer

Mixture model-based clustering has become an increasingly popular data analysis technique since its introduction over fifty years ago, and is now commonly utilized within a family setting. Families of mixture models arise when the component…

Methodology · Statistics 2019-11-11 Sanjeena Subedi , Paul D. McNicholas

Several recent publications report advances in training optimal decision trees (ODT) using mixed-integer programs (MIP), due to algorithmic advances in integer programming and a growing interest in addressing the inherent suboptimality of…

Machine Learning · Computer Science 2020-11-09 Haoran Zhu , Pavankumar Murali , Dzung T. Phan , Lam M. Nguyen , Jayant R. Kalagnanam

Reliable density estimation is fundamental for numerous applications in statistics and machine learning. In many practical scenarios, data are best modeled as mixtures of component densities that capture complex and multimodal patterns.…

Machine Learning · Computer Science 2025-09-30 Mustafa Musab , Joseph K. Chege , Arie Yeredor , Martin Haardt

For many scientific questions, understanding the underlying mechanism is the goal. To help investigators better understand the underlying mechanism, variable selection is a crucial step that permits the identification of the most associated…

Methodology · Statistics 2025-10-06 Shuangshuang Xu , Marco A. R. Ferreira , Allison N. Tegge