中文
相关论文

相关论文: How many data clusters are in the Galaxy data set?…

200 篇论文

We marshall the arguments for preferring Bayesian hypothesis testing and confidence sets to frequentist ones. We define admissible solutions to inference problems, noting that Bayesian solutions are admissible. We give seven weaker…

统计理论 · 数学 2024-05-22 Roger Sewell

Archetypal analysis represents each individual member of a set of data vectors as a mixture (a constrained linear combination) of the pure types or archetypes of the data set. The archetypes are themselves required to be mixtures of the…

天体物理学 · 物理学 2009-11-07 B. H. P. Chan , D. A. Mitchell , L. E. Cram

Clustering is a widely used technique with a long and rich history in a variety of areas. However, most existing algorithms do not scale well to large datasets, or are missing theoretical guarantees of convergence. This paper introduces a…

机器学习 · 统计学 2024-10-16 Yijia Zhou , Kyle A. Gallivan , Adrian Barbu

We propose a two-component mixture of a noninformative (diffuse) and an informative prior distribution, weighted through the data in such a way to prefer the first component if a prior-data conflict arises. The data-driven approach for…

统计方法学 · 统计学 2017-08-02 Leonardo Egidi , Francesco Pauli , Nicola Torelli

We consider the problem of Gaussian mixture clustering in the high-dimensional limit where the data consists of $m$ points in $n$ dimensions, $n,m \rightarrow \infty$ and $\alpha = m/n$ stays finite. Using exact but non-rigorous methods…

In order to improve forecasts, a decisionmaker often combines probabilities given by various sources, such as human experts and machine learning classifiers. When few training data are available, aggregation can be improved by incorporating…

机器学习 · 计算机科学 2012-07-19 Joseph Kahn

A key sticking point of Bayesian analysis is the choice of prior distribution, and there is a vast literature on potential defaults including uniform priors, Jeffreys' priors, reference priors, maximum entropy priors, and weakly informative…

统计方法学 · 统计学 2017-11-22 Andrew Gelman , Daniel Simpson , Michael Betancourt

Bayesian models are a powerful tool for studying complex data, allowing the analyst to encode rich hierarchical dependencies and leverage prior information. Most importantly, they facilitate a complete characterization of uncertainty…

机器学习 · 统计学 2023-04-25 Steven Winter , Trevor Campbell , Lizhen Lin , Sanvesh Srivastava , David B. Dunson

Mixture models provide a flexible representation of heterogeneity in a finite number of latent classes. From the Bayesian point of view, Markov Chain Monte Carlo methods provide a way to draw inferences from these models. In particular,…

统计方法学 · 统计学 2020-05-06 Carolina Valani Cavalcante , Kelly Cristina Mota Gonçalves

Genetic association analyses often involve data from multiple potentially-heterogeneous subgroups. The expected amount of heterogeneity can vary from modest (e.g., a typical meta-analysis) to large (e.g., a strong gene--environment…

统计方法学 · 统计学 2014-04-15 Xiaoquan Wen , Matthew Stephens

The design of an experiment can be always be considered at least implicitly Bayesian, with prior knowledge used informally to aid decisions such as the variables to be studied and the choice of a plausible relationship between the…

统计方法学 · 统计学 2017-01-03 David C. Woods , Antony M. Overstall , Maria Adamou , Timothy W. Waite

Databases often contain corrupted, degraded, and noisy data with duplicate entries across and within each database. Such problems arise in citations, medical databases, genetics, human rights databases, and a variety of other applied…

统计方法学 · 统计学 2015-04-29 Rebecca C. Steorts

Bayesian model comparison requires the specification of a prior distribution on the parameter space of each candidate model. In this connection two concerns arise: on the one hand the elicitation task rapidly becomes prohibitive as the…

统计方法学 · 统计学 2011-02-16 Guido Consonni , Piero Veronese

Numerous papers ask how difficult it is to cluster data. We suggest that the more relevant and interesting question is how difficult it is to cluster data sets {\em that can be clustered well}. More generally, despite the ubiquity and the…

机器学习 · 计算机科学 2012-05-23 Amit Daniely , Nati Linial , Michael Saks

Discrete data such as counts of microbiome taxa resulting from next-generation sequencing are routinely encountered in bioinformatics. Taxa count data in microbiome studies are typically high-dimensional, over-dispersed, and can only reveal…

统计方法学 · 统计学 2022-06-23 Yuan Fang , Sanjeena Subedi

Discovering causal genetic variants from large genetic association studies poses many difficult challenges. Assessing which genetic markers are involved in determining trait status is a computationally demanding task, especially in the…

基因组学 · 定量生物学 2015-04-09 Andrew L. Beam , Alison Motsinger-Reif , Jon Doyle

Motivated by examples from genetic association studies, this paper considers the model selection problem in a general complex linear model system and in a Bayesian framework. We discuss formulating model selection problems and incorporating…

统计方法学 · 统计学 2014-03-14 Xiaoquan Wen

Incorporating graph side information into recommender systems has been widely used to better predict ratings, but relatively few works have focused on theoretical guarantees. Ahn et al. (2018) firstly characterized the optimal sample…

信息论 · 计算机科学 2021-09-09 Changhun Jo , Kangwook Lee

A common approach in computational science is to use a set of of highly precise but expensive calculations to parameterize a model that allows less precise, but more rapid calculations on larger scale systems. Least-squares fitting on a…

材料科学 · 物理学 2015-05-13 Eric Cockayne , Axel van de Walle

Replication studies are essential for assessing the credibility of claims from original studies. A critical aspect of designing replication studies is determining their sample size; a too small sample size may lead to inconclusive studies…

统计方法学 · 统计学 2023-08-14 Samuel Pawel , Guido Consonni , Leonhard Held
‹ 上一页 1 8 9 10 下一页 ›