English
Related papers

Related papers: On the Interplay Between Exposure Misclassificatio…

200 papers

The health effects of environmental exposures have been studied for decades, typically using standard regression models to assess exposure-outcome associations found in observational non-experimental data. We propose and illustrate a…

Applications · Statistics 2017-09-20 Marie-Abele C. Bind , Donald B. Rubin

Meta-analyses frequently include trials that report multiple effect sizes based on a common set of study participants. These effect sizes will generally be correlated. Cluster-robust variance-covariance estimators are a fruitful approach…

Methodology · Statistics 2022-03-07 Thilo Welz , Wolfgang Viechtbauer , Markus Pauly

Adaptive sample size re-estimation, early stopping, and trial re-design at interim analyses can reduce expected sample sizes in randomised trials. Cluster randomised trials, in which groups of participants are randomly allocated to…

Methodology · Statistics 2026-03-09 Samuel I. Watson , James Martin

An optical cluster finder inevitably suffers from projection effects, where it misidentifies a superposition of galaxies in multiple halos along the line-of-sight as a single cluster. Using mock cluster catalogs built from cosmological…

Cosmology and Nongalactic Astrophysics · Physics 2020-06-17 Tomomi Sunayama , Youngsoo Park , Masahiro Takada , Yosuke Kobayashi , Takahiro Nishimichi , Toshiki Kurita , Surhud More , Masamune Oguri , Ken Osato

The misclassification error distance and the adjusted Rand index are two of the most commonly used criteria to evaluate the performance of clustering algorithms. This paper provides an in-depth comparison of the two criteria, aimed to…

Machine Learning · Statistics 2019-07-29 José E. Chacón

The goal of contrasting learning is to learn a representation that preserves underlying clusters by keeping samples with similar content, e.g. the ``dogness'' of a dog, close to each other in the space generated by the representation. A…

Machine Learning · Computer Science 2023-02-17 Advait Parulekar , Liam Collins , Karthikeyan Shanmugam , Aryan Mokhtari , Sanjay Shakkottai

The primary difficulty in measuring dynamical masses of galaxy clusters from galaxy data lies in the separation between true cluster members from interloping galaxies along the line of sight. We study the impact of membership contamination…

Stress testing poses a causal question: how would portfolio credit losses change if the macroeconomy followed an adverse counterfactual path? Yet standard practice remains predictive and might be therefore vulnerable to omitted-variable…

Artificial Intelligence · Computer Science 2026-05-19 Yu Wang , Xiangchen Liu , Siguang Li

The digital spread of misinformation is one of the leading threats to democracy, public health, and the global economy. Popular strategies for mitigating misinformation include crowdsourcing, machine learning, and media literacy programs…

Social and Information Networks · Computer Science 2021-06-09 Douglas Guilbeault , Samuel Woolley , Joshua Becker

An extension of the latent class model is presented for clustering categorical data by relaxing the classical "class conditional independence assumption" of variables. This model consists in grouping the variables into inter-independent and…

Computation · Statistics 2015-10-01 Matthieu Marbac , Christophe Biernacki , Vincent Vandewalle

Widely distributed misinformation shared across social media channels is a pressing issue that poses a significant threat to many aspects of society's well-being. Inaccurate shared information causes confusion, can adversely affect mental…

Social and Information Networks · Computer Science 2024-09-27 Juanita Zainudin , Nazlena Mohamad Ali , Alan F. Smeaton , Mohamad Taha Ijab

Multivariable Mendelian randomization estimates the causal effect of multiple exposures on an outcome, typically using summary statistics of genetic variant associations. However, exposures of interest in Mendelian randomization…

Methodology · Statistics 2022-03-17 Jiazheng Zhu , Stephen Burgess , Andrew J. Grant

Propensity score weighting is a tool for causal inference to adjust for measured confounders. Survey data are often collected under complex sampling designs such as multistage cluster sampling, which presents challenges for propensity score…

Methodology · Statistics 2016-07-27 Shu Yang

Not all instances in a data set are equally beneficial for inferring a model of the data. Some instances (such as outliers) are detrimental to inferring a model of the data. Several machine learning techniques treat instances in a data set…

Machine Learning · Computer Science 2013-12-19 Michael R. Smith , Tony Martinez

Identification of disease subtypes and corresponding biomarkers can substantially improve clinical diagnosis and treatment selection. Discovering these subtypes in noisy, high dimensional biomedical data is often impossible for humans and…

Quantitative Methods · Quantitative Biology 2020-05-18 Marc-Andre Schulz , Matt Chapman-Rounds , Manisha Verma , Danilo Bzdok , Konstantinos Georgatzis

We study the problem of explainability-first clustering where explainability becomes a first-class citizen for clustering. Previous clustering approaches use decision trees for explanation, but only after the clustering is completed. In…

Machine Learning · Computer Science 2022-12-13 Hyunseung Hwang , Steven Euijong Whang

Distributed processing over networks relies on in-network processing and cooperation among neighboring agents. Cooperation is beneficial when agents share a common objective. However, in many applications agents may belong to different…

Optimization and Control · Mathematics 2023-07-19 Xiaochuan Zhao , Ali H. Sayed

The behaviour of sharing information on social media should be fulfilled only when a user is exhibiting attentive behaviour. So that the useful information can be consumed constructively, and misinformation can be identified and ignored.…

Social and Information Networks · Computer Science 2020-12-29 Zaid Amin , Nazlena Mohamad Ali , Alan F. Smeaton

This paper provides a simple theoretical framework to evaluate the effect of key parameters of ranking algorithms, namely popularity and personalization parameters, on measures of platform engagement, misinformation and polarization. The…

Social and Information Networks · Computer Science 2022-10-06 Fabrizio Germano , Vicenç Gómez , Francesco Sobbrio

The challenge of clustering short text data lies in balancing informativeness with interpretability. Traditional evaluation metrics often overlook this trade-off. Inspired by linguistic principles of communicative efficiency, this paper…

Computation and Language · Computer Science 2025-04-08 Justin Miller , Tristram Alexander