中文
相关论文

相关论文: A Closed-Form EVSI Expression for a Multinomial Da…

200 篇论文

In the era of big data, it is necessary to split extremely large data sets across multiple computing nodes and construct estimators using the distributed data. When designing distributed estimators, it is desirable to minimize the amount of…

统计理论 · 数学 2022-04-25 Azeem Zaman , Botond Szabó

Bayesian optimal experimental design provides a principled framework for selecting experimental settings that maximize obtained information. In this work, we focus on estimating the expected information gain in the setting where the…

机器学习 · 统计学 2025-10-02 Chuntao Chen , Tapio Helin , Nuutti Hyvönen , Yuya Suzuki

We consider the distribution of the sum and the maximum of a collection of independent exponentially distributed random variables. The focus is laid on the explicit form of the density functions (pdf) of non-i.i.d. sequences. Those are…

概率论 · 数学 2013-07-16 Markus Bibinger

We consider the problem of estimating the expected value of information (the knowledge gradient) for Bayesian learning problems where the belief model is nonlinear in the parameters. Our goal is to maximize some metric, while simultaneously…

机器学习 · 统计学 2016-11-23 Xinyu He , Warren B. Powell

Simulation is increasingly being used for generating large labelled datasets in many machine learning problems. Recent methods have focused on adjusting simulator parameters with the goal of maximising accuracy on a validation task, usually…

计算机视觉与模式识别 · 计算机科学 2020-08-20 Harkirat Singh Behl , Atılım Güneş Baydin , Ran Gal , Philip H. S. Torr , Vibhav Vineet

We endeavour to estimate numerous multi-dimensional means of various probability distributions on a common space based on independent samples. Our approach involves forming estimators through convex combinations of empirical means derived…

机器学习 · 统计学 2025-03-11 Gilles Blanchard , Jean-Baptiste Fermanian , Hannah Marienwald

The aim of survey statistics is to produce estimates with a minimal bias and a corresponding acceptable variance given a specific budget, preferable with a minor response burden for the participants. In recent years, considerable efforts…

统计方法学 · 统计学 2026-04-02 Martin Hyllienmark , Gustaf Strandell

Computing value of information (VOI) is a crucial task in various aspects of decision-making under uncertainty, such as in meta-reasoning for search; in selecting measurements to make, prior to choosing a course of action; and in managing…

人工智能 · 计算机科学 2015-03-13 David Tolpin , Solomon Eyal Shimony

The integrative analysis of multiple datasets is an important strategy in data analysis. It is increasingly popular in genomics, which enjoys a wealth of publicly available datasets that can be compared, contrasted, and combined in order to…

统计方法学 · 统计学 2019-11-20 Sihai Dave Zhao

As data becomes the fuel driving technological and economic growth, a fundamental challenge is how to quantify the value of data in algorithmic predictions and decisions. For example, in healthcare and consumer markets, it has been…

机器学习 · 统计学 2019-06-11 Amirata Ghorbani , James Zou

When optimizing against the mean loss over a distribution of predictions in the context of a regression task, then even if there is a distribution of targets the optimal prediction distribution is always a delta function at a single value.…

机器学习 · 计算机科学 2019-02-11 Nicholas Guttenberg

This paper focuses on the privacy paradigm of providing access to researchers to remotely carry out analyses on sensitive data stored behind firewalls. We address the situation where the analysis demands data from multiple physically…

统计方法学 · 统计学 2017-10-20 Joshua Snoke , Timothy R. Brick , Aleksandra Slavkovic , Michael D. Hunter

For multi-source data, blocks of variable information from certain sources are likely missing. Existing methods for handling missing data do not take structures of block-wise missing data into consideration. In this paper, we propose a…

统计方法学 · 统计学 2020-04-07 Fei Xue , Annie Qu

Contemporary sample size calculations for external validation of risk prediction models require users to specify fixed values of assumed model performance metrics alongside target precision levels (e.g., 95% CI widths). However, due to the…

We consider the estimation of Dirichlet Process Mixture Models (DPMMs) in distributed environments, where data are distributed across multiple computing nodes. A key advantage of Bayesian nonparametric models such as DPMMs is that they…

机器学习 · 统计学 2017-09-20 Ruohui Wang , Dahua Lin

One important issue commonly encountered in the analysis of microarray data is to decide which and how many genes should be selected for further studies. For discriminant microarray data analyses based on statistical models, such as the…

定量方法 · 定量生物学 2009-11-09 Wentian Li , Fengzhu Sun , Ivo Grosse

We study a data analyst's problem of acquiring data from self-interested individuals to obtain an accurate estimation of some statistic of a population, subject to an expected budget constraint. Each data holder incurs a cost, which is…

计算机科学与博弈论 · 计算机科学 2019-05-15 Yiling Chen , Shuran Zheng

We propose a self-improving algorithm for computing Voronoi diagrams under a given convex distance function with constant description complexity. The $n$ input points are drawn from a hidden mixture of product distributions; we are only…

计算几何 · 计算机科学 2021-10-26 Siu-Wing Cheng , Man Ting Wong

Measuring the value of individual samples is critical for many data-driven tasks, e.g., the training of a deep learning model. Recent literature witnesses the substantial efforts in developing data valuation methods. The primary data…

机器学习 · 计算机科学 2024-06-06 Ou Wu , Weiyao Zhu , Mengyang Li

Variational Autoencoder is a scalable method for learning latent variable models of complex data. It employs a clear objective that can be easily optimized. However, it does not explicitly measure the quality of learned representations. We…

机器学习 · 计算机科学 2020-05-29 Andriy Serdega , Dae-Shik Kim