中文
相关论文

相关论文: A Closed-Form EVSI Expression for a Multinomial Da…

200 篇论文

We propose data thinning, an approach for splitting an observation into two or more independent parts that sum to the original observation, and that follow the same distribution as the original observation, up to a (known) scaling of a…

统计方法学 · 统计学 2023-11-22 Anna Neufeld , Ameer Dharamshi , Lucy L. Gao , Daniela Witten

Determining the sample size of an experiment can be challenging, even more so when incorporating external information via a prior distribution. Such information is increasingly used to reduce the size of the control group in randomized…

应用统计 · 统计学 2019-07-10 Beat Neuenschwander , Sebastian Weber , Heinz Schmidli , Anthony O'Hagan

How much value does a dataset or a data production process have to an agent who wishes to use the data to assist decision-making? This is a fundamental question towards understanding the value of data as well as further pricing of data.…

计算机科学与博弈论 · 计算机科学 2024-12-25 Rui Ai , Boxiang Lyu , Zhaoran Wang , Zhuoran Yang , Haifeng Xu

We propose a new approach for estimating the parameters of a probability distribution. It consists on combining two new methods of estimation. The first is based on the definition of a new distance measuring the difference between…

统计方法学 · 统计学 2008-12-30 Ahmed Guellil , Tewfik Kernane

Several strategies have been developed recently to ensure valid inference after model selection; some of these are easy to compute, while others fare better in terms of inferential power. In this paper, we consider a selective inference…

统计方法学 · 统计学 2022-07-13 Snigdha Panigrahi , Jonathan Taylor

A widely used method to create a continuous representation of a discrete data-set is regression analysis. When the regression model is not based on a mathematical description of the physics underlying the data, heuristic techniques play a…

统计理论 · 数学 2013-07-18 Giovanni Mana , Paolo Alberto Giuliano Albo , Simona Lago

The estimation of a density profile from experimental data points is a challenging problem, usually tackled by plotting a histogram. Prior assumptions on the nature of the density, from its smoothness to the specification of its form, allow…

统计方法学 · 统计学 2015-03-13 Alberto Bernacchia , Simone Pigolotti

This work presents a distributed estimation algorithm that efficiently uses the available communication resources. The approach is based on Bayesian filtering that is distributed across a network by using the logarithmic opinion pool…

机器人学 · 计算机科学 2022-04-04 Miguel Calvo-Fullana , Jonathan P. How

Suppose we have a Bayesian model which combines evidence from several different sources. We want to know which model parameters most affect the estimate or decision from the model, or which of the parameter uncertainties drive the decision…

应用统计 · 统计学 2021-11-25 Christopher Jackson , Anne Presanis , Stefano Conti , Daniela De Angelis

We deal with the efficient parallelization of Bayesian global optimization algorithms, and more specifically of those based on the expected improvement criterion and its variants. A closed form formula relying on multivariate Gaussian…

机器学习 · 统计学 2016-09-12 Sébastien Marmin , Clément Chevalier , David Ginsbourger

Deep directed generative models have attracted much attention recently due to their expressive representation power and the ability of ancestral sampling. One major difficulty of learning directed models with many latent variables is the…

机器学习 · 计算机科学 2015-06-16 Siqi Nie , Qiang Ji

This paper deals with the problem of finding suboptimal values of an unknown function on the basis of measured data corrupted by bounded noise. As a prior, we assume that the unknown function is parameterized in terms of a number of basis…

最优化与控制 · 数学 2025-06-10 Jaap Eising , Jorge Cortes

This work is about recovering an analysis-sparse vector, i.e. sparse vector in some transform domain, from under-sampled measurements. In real-world applications, there often exist random analysis-sparse vectors whose distribution in the…

信息论 · 计算机科学 2022-12-29 Raziyeh Takbiri , Sajad Daei

Multi-class classification problem is among the most popular and well-studied statistical frameworks. Modern multi-class datasets can be extremely ambiguous and single-output predictions fail to deliver satisfactory performance. By allowing…

机器学习 · 统计学 2021-02-25 Evgenii Chzhen , Christophe Denis , Mohamed Hebiri , Titouan Lorieul

In ecological and environmental contexts, management actions must sometimes be chosen urgently. Value of information (VoI) analysis provides a quantitative toolkit for projecting the improved management outcomes expected after making…

Stochastic variational inference (SVI), the state-of-the-art algorithm for scaling variational inference to large-datasets, is inherently serial. Moreover, it requires the parameters to fit in the memory of a single processor; this is…

The multinomial and related distributions have long been used to model categorical, count-based data in fields ranging from bioinformatics to natural language processing. Commonly utilized variants include the standard multinomial and the…

机器学习 · 统计学 2020-06-16 Steven Michael Lakin , Zaid Abdo

We propose a method for the release of differentially private synthetic datasets. In many contexts, data contain sensitive values which cannot be released in their original form in order to protect individuals' privacy. Synthetic data is a…

统计方法学 · 统计学 2018-05-25 Joshua Snoke , Aleksandra Slavković

Modelling block maxima using the generalised extreme value (GEV) distribution is a classical and widely used method for studying univariate extremes. It allows for theoretically motivated estimation of return levels, including extrapolation…

统计方法学 · 统计学 2026-02-02 Emma S. Simpson , Paul J. Northrop

The Dirichlet distribution, also known as multivariate beta, is the most used to analyse frequencies or proportions data. Maximum likelihood is widespread for estimation of Dirichlet's parameters. However, for small sample sizes, the…

统计方法学 · 统计学 2021-03-04 Vincenzo Gioia , Euloge Clovis Kenne Pagui