English
Related papers

Related papers: Estimating intracluster correlation for ordinal da…

200 papers

Clustered data arise naturally in many scientific and applied research settings where units are grouped within clusters. They are commonly analyzed using linear mixed models to account for within-cluster correlations. This article focuses…

Methodology · Statistics 2025-10-10 Zhi Yang Tho , Raymond Chambers , A. H. Welsh

Finite mixture models are widely used in econometric analyses to capture unobserved heterogeneity. This paper shows that maximum likelihood estimation of finite mixtures of parametric densities can suffer from substantial finite-sample bias…

Methodology · Statistics 2026-02-04 Raphaël Langevin

Embedding models, which learn latent representations of users and items based on user-item interaction patterns, are a key component of recommendation systems. In many applications, contextual constraints need to be applied to refine…

Information Retrieval · Computer Science 2019-07-04 Syrine Krichene , Mike Gartrell , Clement Calauzenes

Compositional data sets are ubiquitous in science, including geology, ecology, and microbiology. In microbiome research, compositional data primarily arise from high-throughput sequence-based profiling experiments. These data comprise…

Statistics Theory · Mathematics 2019-03-05 Patrick L. Combettes , Christian L. Müller

This article develops the theoretical framework needed to study the multinomial logistic regression model for complex sample design with pseudo minimum phi-divergence estimators. Through a numerical example and simulation study new…

Methodology · Statistics 2016-06-06 Elena Castilla , Nirian Martin , Leandro Pardo

Clustering is a fundamental approach to understanding data patterns, wherein the intuitive Euclidean distance space is commonly adopted. However, this is not the case for implicit cluster distributions reflected by qualitative attribute…

Machine Learning · Statistics 2026-03-05 Mingjie Zhao , Sen Feng , Yiqun Zhang , Mengke Li , Yang Lu , Yiu-ming Cheung

Multimodal large language models (LLMs) are increasingly explored as automated evaluators in clinical settings, yet their scoring behavior on ordinal clinical scales remains poorly understood. We benchmark three frontier LLM families…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Jiaqing Zhang , Sandeep Elluri , Bhanu Cherukuvada , Yonah Joffe , Jessica Sena , Miguel Contreras , Scott Siegel , Subhash Nerella , Catherine Price , Parisa Rashidi

Estimating causal effects from observational data is not always possible due to confounding. Identifying a set of appropriate covariates (adjustment set) and adjusting for their influence can remove confounding bias; however, such a set is…

Methodology · Statistics 2020-11-19 Sofia Triantafillou , Gregory Cooper

In longitudinal studies where units are embedded in space or a social network, interference may arise, meaning that a unit's outcome can depend on treatment histories of others. The presence of interference poses significant challenges for…

Methodology · Statistics 2025-08-26 Ye Wang , Michael Jetsupphasuk

Manual annotation of audio datasets is labour intensive, and it is challenging to balance label granularity with acoustic separability. We introduce AuditoryHuM, a novel framework for the unsupervised discovery and clustering of auditory…

Sound · Computer Science 2026-02-24 Henry Zhong , Jörg M. Buchholz , Julian Maclaren , Simon Carlile , Richard F. Lyon

We study a general clustering setting in which we have $n$ elements to be clustered, and we aim to perform as few queries as possible to an oracle that returns a noisy sample of the weighted similarity between two elements. Our setting…

Machine Learning · Statistics 2024-11-05 Yuko Kuroki , Atsushi Miyauchi , Francesco Bonchi , Wei Chen

Local density-based score normalization is an effective component of distance-based embedding methods for anomalous sound detection, particularly when data densities vary across conditions or domains. In practice, however, performance…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-24 Kevin Wilkinghoff , Gordon Wichern , Jonathan Le Roux , Zheng-Hua Tan

The research is about a systematic investigation on the following issues. First, we construct different outcome regression-based estimators for conditional average treatment effect under, respectively, true (oracle), parametric,…

Statistics Theory · Mathematics 2020-09-23 Lu Li , Niwen Zhou , Lixing Zhu

Classical wisdom suggests that estimators should avoid fitting noise to achieve good generalization. In contrast, modern overparameterized models can yield small test error despite interpolating noise -- a phenomenon often called "benign…

Machine Learning · Statistics 2023-03-02 Michael Aerni , Marco Milanta , Konstantin Donhauser , Fanny Yang

Large language models (LLMs) exhibit impressive in-context learning (ICL) capabilities, yet the quality of their predictions is fundamentally limited by the few costly labeled demonstrations that can fit into a prompt. Meanwhile, there…

Machine Learning · Computer Science 2026-01-16 Renpu Liu , Jing Yang

In many studies multivariate event time data are generated from clusters having a possibly complex association pattern. Flexible models are needed to capture this dependence. Vine copulas serve this purpose. Inference methods for vine…

Applications · Statistics 2017-07-25 Nicole Barthel , Candida Geerdens , Matthias Killiches , Paul Janssen , Claudia Czado

Deep clustering aims to learn a clustering representation through deep architectures. Most of the existing methods usually conduct clustering with the unique goal of maximizing clustering performance, that ignores the personalized demand of…

Computer Vision and Pattern Recognition · Computer Science 2022-11-02 Mengdie Wang , Liyuan Shang , Suyun Zhao , Yiming Wang , Hong Chen , Cuiping Li , Xizhao Wang

The consumption of music has its specificities in comparison with other media, especially in relation to listening durations and replays. Music recommendation can take these properties into account in order to predict the behaviours of the…

Information Retrieval · Computer Science 2017-11-15 Pierre Hanna

In-context learning (ICL) suffers from oversensitivity to the prompt, making it unreliable in real-world scenarios. We study the sensitivity of ICL with respect to multiple perturbation types. First, we find that label bias obscures the…

Computation and Language · Computer Science 2024-01-30 Yanda Chen , Chen Zhao , Zhou Yu , Kathleen McKeown , He He

We investigate the problem of machine learning-based (ML) predictive inference on individual treatment effects (ITEs). Previous work has focused primarily on developing ML-based meta-learners that can provide point estimates of the…

Machine Learning · Computer Science 2023-08-30 Ahmed Alaa , Zaid Ahmad , Mark van der Laan