English
Related papers

Related papers: Correcting misclassification errors in crowdsource…

200 papers

Two key challenges in modern statistical applications are the large amount of information recorded per individual, and that such data are often not collected all at once but in batches. These batch effects can be complex, causing…

Applications · Statistics 2019-05-21 Alejandra Avalos-Pacheco , David Rossell , Richard S. Savage

Within a supervised classification framework, labeled data are used to learn classifier parameters. Prior to that, it is generally required to perform dimensionality reduction via feature extraction. These preprocessing steps have motivated…

Computer Vision and Pattern Recognition · Computer Science 2017-12-04 Adrien Lagrange , Mathieu Fauvel , Stéphane May , Nicolas Dobigeon

The problem of selecting a model given a set of candidates remains a challenging one that pervades many scientific fields. We employ techniques from the theory of Lie groups to analyse the symmetries in differential equation models of…

Quantitative Methods · Quantitative Biology 2023-10-10 Reemon Spector

Learning unbiased models on imbalanced datasets is a significant challenge. Rare classes tend to get a concentrated representation in the classification space which hampers the generalization of learned boundaries to new test examples. In…

Computer Vision and Pattern Recognition · Computer Science 2019-04-11 Salman Khan , Munawar Hayat , Waqas Zamir , Jianbing Shen , Ling Shao

In ecology, the description of species composition and biodiversity calls for statistical methods that involve estimating features of interest in unobserved samples based on an observed one. In the last decade, the Bayesian nonparametrics…

Methodology · Statistics 2026-04-28 Alessandro Colombi , Raffaele Argiento , Federico Camerlenghi , Lucia Paci

In this article, we present a Bayesian hierarchical model for predicting a latent health state from longitudinal clinical measurements. Model development is motivated by the need to integrate multiple sources of data to improve clinical…

Bayesian analysis is increasingly popular for use in social science and other application areas where the data are observations from an informative sample. An informative sampling design leads to inclusion probabilities that are correlated…

Statistics Theory · Mathematics 2016-06-07 Terrance D. Savitsky , Daniell Toth

Hierarchical parametric models consisting of observable and latent variables are widely used for unsupervised learning tasks. For example, a mixture model is a representative hierarchical model for clustering. From the statistical point of…

Machine Learning · Statistics 2014-01-24 Keisuke Yamazaki

Citizen science projects in which volunteers collect data are increasingly popular due to their ability to engage the public with scientific questions. The scientific value of these data are however hampered by several biases. In this…

Applications · Statistics 2024-02-01 Peter Lugtig , Erik-Jan van Kesteren , Annemarie Timmers

Webcam-based eye tracking is a cost-effective, scalable method for remote research that effectively reaches broader populations. However, uncontrolled environments and hardware diversity lead to inconsistent data quality in crowdsourcing.…

Human-Computer Interaction · Computer Science 2026-05-06 Ka Hei Carrie Lau , Enkelejda Kasneci

Bayesian inference for inverse problems hinges critically on the choice of priors. In the absence of specific prior information, population-level distributions can serve as effective priors for parameters of interest. With the advent of…

Instrumentation and Methods for Astrophysics · Physics 2025-02-11 Gabriel Missael Barco , Alexandre Adam , Connor Stone , Yashar Hezaveh , Laurence Perreault-Levasseur

Big Data often presents as massive non-probability samples. Not only is the selection mechanism often unknown, but larger data volume amplifies the relative contribution of selection bias to total error. Existing bias adjustment approaches…

Methodology · Statistics 2022-03-29 Ali Rafei , Carol A. C. Flannagan , Brady T. West , Michael R. Elliott

In many statistical problems, a more coarse-grained model may be suitable for population-level behaviour, whereas a more detailed model is appropriate for accurate modelling of individual behaviour. This raises the question of how to…

Machine Learning · Statistics 2015-11-02 Mingjun Zhong , Nigel Goddard , Charles Sutton

The recent success of machine learning models, especially large-scale classifiers and language models, relies heavily on training with massive data. These data are often collected from online sources. This raises serious concerns about the…

Artificial Intelligence · Computer Science 2025-11-12 Ruihan Zhang , Jun Sun , Ee-Peng Lim , Peixin Zhang

Datasets are rarely a realistic approximation of the target population. Say, prevalence is misrepresented, image quality is above clinical standards, etc. This mismatch is known as sampling bias. Sampling biases are a major hindrance for…

Statistical estimates from survey samples have traditionally been obtained via design-based estimators. In many cases, these estimators tend to work well for quantities such as population totals or means, but can fall short as sample sizes…

Methodology · Statistics 2020-09-15 Paul A. Parker , Scott H. Holan , Ryan Janicki

Growing anthropogenic pressures have increased the need for robust predictive models. Meeting this demand requires approaches that can handle bigger data to yield forecasts that capture the variability and underlying uncertainty of…

Quantitative Methods · Quantitative Biology 2024-08-06 EM Wolkovich , T Jonathan Davies , William D Pearse , Michael Betancourt

We propose to learn latent graphical models when data have mixed variables and missing values. This model could be used for further data analysis, including regression, classification, ranking etc. It also could be used for imputing missing…

Methodology · Statistics 2015-11-17 Xiao Li , Jinzhu Jia , Yuan Yao

This study presents a novel approach to quantifying uncertainties in Bayesian model updating, which is effective in sparse or single observations. Conventional uncertainty quantification metrics such as the Euclidean and Bhattacharyya…

Applications · Statistics 2024-10-14 Sangwon Lee , Taro Yaoyama , Yuma Matsumoto , Takenori Hida , Tatsuya Itoi

We propose a cautious Bayesian variable selection routine by investigating the sensitivity of a hierarchical model, where the regression coefficients are specified by spike and slab priors. We exploit the use of latent variables to…

Methodology · Statistics 2022-06-20 Tathagata Basu , Matthias C. M. Troffaes , Jochen Einbeck