中文
相关论文

相关论文: Correcting misclassification errors in crowdsource…

200 篇论文

Two key challenges in modern statistical applications are the large amount of information recorded per individual, and that such data are often not collected all at once but in batches. These batch effects can be complex, causing…

应用统计 · 统计学 2019-05-21 Alejandra Avalos-Pacheco , David Rossell , Richard S. Savage

Within a supervised classification framework, labeled data are used to learn classifier parameters. Prior to that, it is generally required to perform dimensionality reduction via feature extraction. These preprocessing steps have motivated…

计算机视觉与模式识别 · 计算机科学 2017-12-04 Adrien Lagrange , Mathieu Fauvel , Stéphane May , Nicolas Dobigeon

The problem of selecting a model given a set of candidates remains a challenging one that pervades many scientific fields. We employ techniques from the theory of Lie groups to analyse the symmetries in differential equation models of…

定量方法 · 定量生物学 2023-10-10 Reemon Spector

Learning unbiased models on imbalanced datasets is a significant challenge. Rare classes tend to get a concentrated representation in the classification space which hampers the generalization of learned boundaries to new test examples. In…

计算机视觉与模式识别 · 计算机科学 2019-04-11 Salman Khan , Munawar Hayat , Waqas Zamir , Jianbing Shen , Ling Shao

In ecology, the description of species composition and biodiversity calls for statistical methods that involve estimating features of interest in unobserved samples based on an observed one. In the last decade, the Bayesian nonparametrics…

统计方法学 · 统计学 2026-04-28 Alessandro Colombi , Raffaele Argiento , Federico Camerlenghi , Lucia Paci

In this article, we present a Bayesian hierarchical model for predicting a latent health state from longitudinal clinical measurements. Model development is motivated by the need to integrate multiple sources of data to improve clinical…

Bayesian analysis is increasingly popular for use in social science and other application areas where the data are observations from an informative sample. An informative sampling design leads to inclusion probabilities that are correlated…

统计理论 · 数学 2016-06-07 Terrance D. Savitsky , Daniell Toth

Hierarchical parametric models consisting of observable and latent variables are widely used for unsupervised learning tasks. For example, a mixture model is a representative hierarchical model for clustering. From the statistical point of…

机器学习 · 统计学 2014-01-24 Keisuke Yamazaki

Citizen science projects in which volunteers collect data are increasingly popular due to their ability to engage the public with scientific questions. The scientific value of these data are however hampered by several biases. In this…

应用统计 · 统计学 2024-02-01 Peter Lugtig , Erik-Jan van Kesteren , Annemarie Timmers

Webcam-based eye tracking is a cost-effective, scalable method for remote research that effectively reaches broader populations. However, uncontrolled environments and hardware diversity lead to inconsistent data quality in crowdsourcing.…

人机交互 · 计算机科学 2026-05-06 Ka Hei Carrie Lau , Enkelejda Kasneci

Bayesian inference for inverse problems hinges critically on the choice of priors. In the absence of specific prior information, population-level distributions can serve as effective priors for parameters of interest. With the advent of…

天体物理仪器与方法 · 物理学 2025-02-11 Gabriel Missael Barco , Alexandre Adam , Connor Stone , Yashar Hezaveh , Laurence Perreault-Levasseur

Big Data often presents as massive non-probability samples. Not only is the selection mechanism often unknown, but larger data volume amplifies the relative contribution of selection bias to total error. Existing bias adjustment approaches…

统计方法学 · 统计学 2022-03-29 Ali Rafei , Carol A. C. Flannagan , Brady T. West , Michael R. Elliott

In many statistical problems, a more coarse-grained model may be suitable for population-level behaviour, whereas a more detailed model is appropriate for accurate modelling of individual behaviour. This raises the question of how to…

机器学习 · 统计学 2015-11-02 Mingjun Zhong , Nigel Goddard , Charles Sutton

The recent success of machine learning models, especially large-scale classifiers and language models, relies heavily on training with massive data. These data are often collected from online sources. This raises serious concerns about the…

人工智能 · 计算机科学 2025-11-12 Ruihan Zhang , Jun Sun , Ee-Peng Lim , Peixin Zhang

Datasets are rarely a realistic approximation of the target population. Say, prevalence is misrepresented, image quality is above clinical standards, etc. This mismatch is known as sampling bias. Sampling biases are a major hindrance for…

Statistical estimates from survey samples have traditionally been obtained via design-based estimators. In many cases, these estimators tend to work well for quantities such as population totals or means, but can fall short as sample sizes…

统计方法学 · 统计学 2020-09-15 Paul A. Parker , Scott H. Holan , Ryan Janicki

Growing anthropogenic pressures have increased the need for robust predictive models. Meeting this demand requires approaches that can handle bigger data to yield forecasts that capture the variability and underlying uncertainty of…

定量方法 · 定量生物学 2024-08-06 EM Wolkovich , T Jonathan Davies , William D Pearse , Michael Betancourt

We propose to learn latent graphical models when data have mixed variables and missing values. This model could be used for further data analysis, including regression, classification, ranking etc. It also could be used for imputing missing…

统计方法学 · 统计学 2015-11-17 Xiao Li , Jinzhu Jia , Yuan Yao

This study presents a novel approach to quantifying uncertainties in Bayesian model updating, which is effective in sparse or single observations. Conventional uncertainty quantification metrics such as the Euclidean and Bhattacharyya…

应用统计 · 统计学 2024-10-14 Sangwon Lee , Taro Yaoyama , Yuma Matsumoto , Takenori Hida , Tatsuya Itoi

We propose a cautious Bayesian variable selection routine by investigating the sensitivity of a hierarchical model, where the regression coefficients are specified by spike and slab priors. We exploit the use of latent variables to…

统计方法学 · 统计学 2022-06-20 Tathagata Basu , Matthias C. M. Troffaes , Jochen Einbeck