English
Related papers

Related papers: How many data clusters are in the Galaxy data set?…

200 papers

We examine the potential improvements in constraints on the dark energy equation of state parameter $w$ and matter density $\Omega_M$ from using clustering information along with number counts for future samples of thermal…

Cosmology and Nongalactic Astrophysics · Physics 2026-01-13 Junhao Zhan , Christian L. Reichardt

Frequentist statistical methods, such as hypothesis testing, are standard practice in papers that provide benchmark comparisons. Unfortunately, these methods have often been misused, e.g., without testing for their statistical test…

Methodology · Statistics 2021-05-18 David Issa Mattos , Jan Bosch , Helena Holmström Olsson

The paper presents the algorithm for clustering a dataset by grouping the optimal, from the point of view of the BIC criterion, number of Gaussian clusters into the optimal, from the point of view of their statistical separability,…

Machine Learning · Computer Science 2023-10-31 Oleg I. Berngardt

In this paper, we study the accuracy of values aggregated over classes predicted by a classification algorithm. The problem is that the resulting aggregates (e.g., sums of a variable) are known to be biased. The bias can be large even for…

Machine Learning · Statistics 2019-12-02 Q. A. Meertens , C. G. H. Diks , H. J. van den Herik , F W Takes

We consider a weighted biasing scheme for galaxy clustering. This differs from previous treatments in the fact that the biased density field coincides with the background mass--density whenever the latter exceeds a given threshold value.…

Astrophysics · Physics 2015-06-24 P. Catelan , P. Coles , S. Matarrese , L. Moscardini

A Gaussian process has been one of the important approaches for emulating computer simulations. However, the stationarity assumption for a Gaussian process and the intractability for large-scale dataset limit its availability in practice.…

Methodology · Statistics 2020-11-06 Chih-Li Sung , Benjamin Haaland , Youngdeok Hwang , Siyuan Lu

All lens modeling methods, simply-parametrized, hybrid, and free-form, use assumptions to reconstruct galaxy clusters with multiply imaged sources, though the nature of these assumptions (priors) can differ considerably between methods.…

Cosmology and Nongalactic Astrophysics · Physics 2023-09-15 Kekoa Lasko , Liliya L. R. Williams , Agniva Ghosh

Analyzes of next-generation galaxy data require accurate treatment of systematic effects such as the bias between observed galaxies and the underlying matter density field. However, proposed models of the phenomenon are either numerically…

Cosmology and Nongalactic Astrophysics · Physics 2021-04-28 Guilhem Lavaux , Jens Jasche

A model-based approach is developed for clustering categorical data with no natural ordering. The proposed method exploits the Hamming distance to define a family of probability mass functions to model the data. The elements of this family…

Methodology · Statistics 2024-07-02 Raffaele Argiento , Edoardo Filippi-Mazzola , Lucia Paci

We propose a new method for conducting Bayesian prediction that delivers accurate predictions without correctly specifying the unknown true data generating process. A prior is defined over a class of plausible predictive models. After…

Methodology · Statistics 2020-08-24 Ruben Loaiza-Maya , Gael M. Martin , David T. Frazier

We consider the problem of inferring an unknown number of clusters in replicated multinomial data. Under a model based clustering point of view, this task can be treated by estimating finite mixtures of multinomial distributions with or…

Methodology · Statistics 2023-07-07 Panagiotis Papastamoulis

In recent years, there has been a growing demand to discern clusters of subjects in datasets characterized by a large set of features. Often, these clusters may be highly variable in size and present partial hierarchical structures. In this…

Methodology · Statistics 2024-07-01 Lorenzo Schiavon , Mattia Stival

The assumption of group heterogeneity has become popular in panel data models. We develop a constrained Bayesian grouped estimator that exploits researchers' prior beliefs on groups in a form of pairwise constraints, indicating whether a…

Econometrics · Economics 2023-10-31 Boyuan Zhang

Probabilistic clustering models (or equivalently, mixture models) are basic building blocks in countless statistical models and involve latent random variables over discrete spaces. For these models, posterior inference methods can be…

Machine Learning · Statistics 2020-06-24 Ari Pakman , Yueqi Wang , Catalin Mitelut , JinHyung Lee , Liam Paninski

The behavior of many Bayesian models used in machine learning critically depends on the choice of prior distributions, controlled by some hyperparameters that are typically selected by Bayesian optimization or cross-validation. This…

Machine Learning · Statistics 2023-10-09 Eliezer de Souza da Silva , Tomasz Kuśmierczyk , Marcelo Hartmann , Arto Klami

In this paper, we consider the task of clustering a set of individual time series while modeling each cluster, that is, model-based time series clustering. The task requires a parametric model with sufficient flexibility to describe the…

Machine Learning · Computer Science 2023-02-23 Ryohei Umatani , Takashi Imai , Kaoru Kawamoto , Shutaro Kunimasa

Galaxy groups and clusters are formidable cosmological probes. They permit the studying of the environmental effects on galaxy formation. A reliable detection of galaxy groups is an open problem and is important for ongoing and future…

Cosmology and Nongalactic Astrophysics · Physics 2018-10-17 Elmo Tempel , Maarja Kruuse , Rain Kipper , Taavi Tuvikene , Jenny G. Sorce , Radu S. Stoica

Gibbs-type priors are widely used as key components in several Bayesian nonparametric models. By virtue of their flexibility and mathematical tractability, they turn out to be predominant priors in species sampling problems, clustering and…

Methodology · Statistics 2021-08-30 Federico Camerlenghi , Riccardo Corradin , Andrea Ongaro

To cluster data is to separate samples into distinctive groups that should ideally have some cohesive properties. Today, numerous clustering algorithms exist, and their differences lie essentially in what can be perceived as ``cohesive…

Machine Learning · Statistics 2025-05-08 Louis Ohl , Pierre-Alexandre Mattei , Frédéric Precioso

Performing optimal Bayesian design for discriminating between competing models is computationally intensive as it involves estimating posterior model probabilities for thousands of simulated datasets. This issue is compounded further when…

Methodology · Statistics 2022-04-07 Markus Hainy , David J. Price , Olivier Restif , Christopher Drovandi