English
Related papers

Related papers: The N-ary in the Coal Mine: Avoiding Mixture Model…

200 papers

During the ion bombardment of targets containing multiple component species, highly-ordered arrays of nanostructures are sometimes observed. Models incorporating coupled partial differential equations, describing both morphological and…

Mathematical Physics · Physics 2015-06-15 Scott A. Norris , Juha Samela , Matias Vestberg , Kai Nordlund , Michael J. Aziz

Confirmation bias, the tendency to interpret information in a way that aligns with one's preconceptions, can profoundly impact scientific research, leading to conclusions that reflect the researcher's hypotheses even when the observational…

Machine Learning · Statistics 2025-09-09 Amnon Balanov , Tamir Bendory , Wasim Huleihel

Lipid membranes have complex compositions and modeling the thermodynamic properties of multi-component lipid systems remains a remote goal. In this work we attempt to describe the thermodynamics of binary lipid mixtures by mapping…

Soft Condensed Matter · Physics 2024-09-04 L. Berezovska , R. Kociurzynski , F. Thalmann

This paper proposes a new nonparametric Bayesian bootstrap for a mixture model, by developing the traditional Bayesian bootstrap. We first reinterpret the Bayesian bootstrap, which uses the P\'olya-urn scheme, as a gradient ascent algorithm…

Methodology · Statistics 2025-01-28 Fuheng Cui , Stephen G. Walker

A mixture of multivariate contaminated normal distributions is developed for model-based clustering. In addition to the parameters of the classical normal mixture, our contaminated mixture has, for each cluster, a parameter controlling the…

Methodology · Statistics 2016-05-20 Antonio Punzo , Paul D. McNicholas

Motivated by problems in data clustering, we establish general conditions under which families of nonparametric mixture models are identifiable, by introducing a novel framework involving clustering overfitted \emph{parametric} (i.e.…

Statistics Theory · Mathematics 2020-02-19 Bryon Aragam , Chen Dan , Eric P. Xing , Pradeep Ravikumar

The problem of missing values in multivariable time series is a key challenge in many applications such as clinical data mining. Although many imputation methods show their effectiveness in many applications, few of them are designed to…

Machine Learning · Computer Science 2020-03-04 Ye Xue , Diego Klabjan , Yuan Luo

Despite the flexibility and popularity of mixture models, their associated parameter spaces are often difficult to represent due to fundamental identification problems. This paper looks at a novel way of representing such a space for…

Methodology · Statistics 2015-10-16 Vahed Maroufy , Paul Marriott

In the field of modeling, the word validation refers to simple comparisons between model outputs and experimental data. Usually, this comparison constitutes plotting the model results against data on the same axes to provide a visual…

Applications · Statistics 2021-06-11 Farid Mohammadi

Mixtures of Linear Regressions (MLR) is an important mixture model with many applications. In this model, each observation is generated from one of the several unknown linear regression components, where the identity of the generated…

Machine Learning · Computer Science 2020-03-31 Yuanzhi Li , Yingyu Liang

This paper considers the problem of mismeasured categorical covariates in the context of regression modeling; if unaccounted for, such misclassification is known to result in misestimation of model parameters. Here, we exploit the fact that…

Statistics Theory · Mathematics 2017-04-28 P. Richard Hahn , Michelle Xia

In computational materials science, mechanical properties are typically extracted from simulations by means of analysis routines that seek to mimic their experimental counterparts. However, simulated data often exhibit uncertainties that…

Data Analysis, Statistics and Probability · Physics 2017-12-07 Paul N. Patrone , Anthony J. Kearsley , Andrew M. Dienstfrey

The analysis of experimental data with mixed-effects models requires decisions about the specification of the appropriate random-effects structure. Recently, Barr, Levy, Scheepers, and Tily, 2013 recommended fitting `maximal' models with…

Methodology · Statistics 2018-05-29 Douglas Bates , Reinhold Kliegl , Shravan Vasishth , Harald Baayen

In many contexts it is extremely costly to perform enough high quality experimental measurements to accurately parameterize a predictive quantitative model. However, it is often much easier to carry out large numbers of experiments that…

Data Analysis, Statistics and Probability · Physics 2017-11-22 Alpha A. Lee , Michael P. Brenner , Lucy J. Colwell

When a missing-data mechanism is NMAR or non-ignorable, missingness is itself vital information and it must be taken into the likelihood, which, however, needs to introduce additional parameters to be estimated. The incompleteness of the…

Methodology · Statistics 2014-05-15 Kosuke Morikawa , Yutaka Kano

Mixing datasets for fine-tuning large models (LMs) has become critical for maximizing performance on downstream tasks. However, composing effective dataset mixtures typically relies on heuristics and trial-and-error, often requiring…

Machine Learning · Computer Science 2025-05-23 Zhixu Silvia Tao , Kasper Vinken , Hao-Wei Yeh , Avi Cooper , Xavier Boix

The general principles of Bayesian data analysis imply that models for survey responses should be constructed conditional on all variables that affect the probability of inclusion and nonresponse, which are also the variables used in survey…

Methodology · Statistics 2007-11-06 Andrew Gelman

Mathematical models are increasingly being used to understand complex biochemical systems, to analyze experimental data and make predictions about unobserved quantities. However, we rarely know how robust our conclusions are with respect to…

Molecular Networks · Quantitative Biology 2015-11-06 Elisenda Feliu , Carsten Wiuf

Data sets obtained from linking multiple files are frequently affected by mismatch error, as a result of non-unique or noisy identifiers used during record linkage. Accounting for such mismatch error in downstream analysis performed on the…

Scientific fields such as insider-threat detection and highway-safety planning often lack sufficient amounts of time-series data to estimate statistical models for the purpose of scientific discovery. Moreover, the available limited data…

Machine Learning · Statistics 2018-03-16 Daniel Emaasit , Matthew Johnson