English
Related papers

Related papers: A Closed-Form EVSI Expression for a Multinomial Da…

200 papers

This article introduces tools to analyze set-valued data statistically. The tools were initially developed to analyze results from an interlaboratory comparison made by the Electromagnetic Compatibility Working Group of Eurolab France,…

Methodology · Statistics 2026-01-27 Sébastien Petit , Sébastien Marmin , Nicolas Fischer

The ability of many powerful machine learning algorithms to deal with large data sets without compromise is often hampered by computationally expensive linear algebra tasks, of which calculating the log determinant is a canonical example.…

Machine Learning · Statistics 2017-09-11 Diego Granziol , Stephen Roberts

In the present work we have selected a collection of statistical and mathematical tools useful for the exploration of multivariate data and we present them in a form that is meant to be particularly accessible to a classically trained…

Statistics Theory · Mathematics 2010-09-01 Magnus Fontes

Attributing model behavior to training data is an evolving research field. A common benchmark is data removal, which involves eliminating data instances with either low or high values, then assessing a model's performance trained on the…

Artificial Intelligence · Computer Science 2026-05-13 Danilo Brajovic , David A. Kreplin , Marco F. Huber

The estimation of the Extreme Value Index (EVI) is fundamental in extreme value analysis but suffers from high variance due to reliance on only a few extreme observations. We propose a control variates based transfer learning approach in a…

Methodology · Statistics 2025-11-20 Louison Bocquet-Nouaille , Jérôme Morio , Benjamin Bobbia

In emotion recognition, it is difficult to recognize human's emotional states using just a single modality. Besides, the annotation of physiological emotional data is particularly expensive. These two aspects make the building of effective…

Artificial Intelligence · Computer Science 2017-04-26 Changde Du , Changying Du , Jinpeng Li , Wei-long Zheng , Bao-liang Lu , Huiguang He

Recent studies have highlighted the benefits of generating multiple synthetic datasets for supervised learning, from increased accuracy to more effective model selection and uncertainty estimation. These benefits have clear empirical…

Machine Learning · Computer Science 2025-04-28 Ossi Räisä , Antti Honkela

We describe a method to computationally estimate the probability density function of a univariate random variable by applying the maximum entropy principle with some local conditions given by Gaussian functions. The estimation errors and…

Statistics Theory · Mathematics 2012-06-21 Mihail-Ioan Pop

Quantifying the value of data within a machine learning workflow can play a pivotal role in making more strategic decisions in machine learning initiatives. The existing Shapley value based frameworks for data valuation in machine learning…

Machine Learning · Computer Science 2024-07-10 Ayush K Tarun , Vikram S Chundawat , Murari Mandal , Hong Ming Tan , Bowei Chen , Mohan Kankanhalli

We investigate the data distribution valuation problem, which aims to quantify the values of data distributions from their samples. This is a recently proposed problem that is related to but different from classical data valuation and can…

Machine Learning · Computer Science 2026-04-08 Cuong N. Nguyen , Cuong V. Nguyen

Extracting low-dimensional summary statistics from large datasets is essential for efficient (likelihood-free) inference. We characterize three different classes of summaries and demonstrate their importance for correctly analyzing…

Methodology · Statistics 2025-11-25 Till Hoffmann , Jukka-Pekka Onnela

Optimization of expensive computer models with the help of Gaussian process emulators in now commonplace. However, when several (competing) objectives are considered, choosing an appropriate sampling strategy remains an open question. We…

Optimization and Control · Mathematics 2013-10-03 Victor Picheny

Compression is at the heart of effective representation learning. However, lossy compression is typically achieved through simple parametric models like Gaussian noise to preserve analytic tractability, and the limitations this imposes on…

Machine Learning · Computer Science 2019-11-15 Rob Brekelmans , Daniel Moyer , Aram Galstyan , Greg Ver Steeg

In the Bayesian approach to structure learning of graphical models, the equivalent sample size (ESS) in the Dirichlet prior over the model parameters was recently shown to have an important effect on the maximum-a-posteriori estimate of the…

Machine Learning · Computer Science 2012-06-18 Harald Steck

In this article we derive the best possible upper bound for $E[\max{X_i}-\min_i{X_i}]$ under given means and variances on $n$ random variables $X_i$. The random vector $(X_1,...,X_n)$ is allowed to have any dependence structure, provided $E…

Methodology · Statistics 2016-11-18 Nickos Papadatos

The general relationship between an arbitrary frequency distribution and the expectation value of the frequency distributions of its samples is discussed. A wide set of measurable quantities ("invariant moments") whose expectation value…

Data Analysis, Statistics and Probability · Physics 2015-06-15 Paolo Rossi

In this paper will be presented methodology of encoding information in valuations of discrete lattice with some translational invariant constrains in asymptotically optimal way. The method is based on finding statistical description of such…

Information Theory · Computer Science 2008-11-02 Jarek Duda

Data is a precious resource in today's society, and is generated at an unprecedented and constantly growing pace. The need to store, analyze, and make data promptly available to a multitude of users introduces formidable challenges in…

Distributed, Parallel, and Cluster Computing · Computer Science 2023-06-08 Alessandro Margara , Gianpaolo Cugola , Nicolò Felicioni , Stefano Cilloni

Probit models are useful for modeling correlated discrete responses in many disciplines, including consumer choice data in economics and marketing. However, the Gaussian latent variable feature of probit models coupled with identification…

Methodology · Statistics 2024-09-30 Patrick Ding , Guido Imbens , Zhaonan Qu , Yinyu Ye

A new shrinkage-based construction is developed for a compressible vector $\boldsymbol{x}\in\mathbb{R}^n$, for cases in which the components of $\xv$ are naturally associated with a tree structure. Important examples are when $\xv$…

Machine Learning · Statistics 2014-01-14 Xin Yuan , Vinayak Rao , Shaobo Han , Lawrence Carin