English
Related papers

Related papers: Measuring Feature-Label Dependence Using Projectio…

200 papers

Datasets may contain observations with multiple labels. If the labels are not mutually exclusive, and if the labels vary greatly in frequency, obtaining a sample that includes sufficient observations with scarcer labels to make inferences…

Machine Learning · Computer Science 2026-05-27 Simon Chung , Colby J. Vorland , Donna L. Maney , Andrew W. Brown

This paper presents and analyzes an approach to cluster-based inference for dependent data. The primary setting considered here is with spatially indexed data in which the dependence structure of observed random variables is characterized…

Statistics Theory · Mathematics 2022-11-16 Jianfei Cao , Christian Hansen , Damian Kozbur , Lucciano Villacorta

Several real-life applications require crafting concise, quantitative scoring functions (also called rating systems) from measured observations. For example, an effectiveness score needs to be created for advertising campaigns using a…

Machine Learning · Computer Science 2025-04-01 Ragja Palakkadavath , Sarath Sivaprasad , Shirish Karande , Niranjan Pedanekar

Classification is a major tool of statistics and machine learning. A classification method first processes a training set of objects with given classes (labels), with the goal of afterward assigning new objects to one of these classes. When…

Machine Learning · Statistics 2024-07-08 Jakob Raymaekers , Peter J. Rousseeuw , Mia Hubert

We propose a learning setting in which unlabeled data is free, and the cost of a label depends on its value, which is not known in advance. We study binary classification in an extreme case, where the algorithm only pays for negative…

Machine Learning · Computer Science 2015-07-14 Sivan Sabato , Anand D. Sarwate , Nathan Srebro

Log-linear models are a family of probability distributions which capture relationships between variables. They have been proven useful in a wide variety of fields such as epidemiology, economics and sociology. The interest in using these…

Machine Learning · Computer Science 2022-12-29 Jan Strappa , Facundo Bromberg

We study the data-driven selection of causal graphical models using constraint-based algorithms, which determine the existence or non-existence of edges (causal connections) in a graph based on testing a series of conditional independence…

Methodology · Statistics 2026-04-29 Daniel Malinsky

Following our previous work on copula-based nonsymmetric bivariate dependence measures, we propose a new set of conditions on nonsymmetric multivariate dependence measures which characterize both independence and complete dependence of one…

Methodology · Statistics 2015-12-04 Hui Li

We introduce a framework for filtering features that employs the Hilbert-Schmidt Independence Criterion (HSIC) as a measure of dependence between the features and the labels. The key idea is that good features should maximise such…

Machine Learning · Computer Science 2007-05-23 Le Song , Alex Smola , Arthur Gretton , Karsten Borgwardt , Justin Bedo

This paper proposes a new mutual independence test for a large number of high dimensional random vectors. The test statistic is based on the characteristic function of the empirical spectral distribution of the sample covariance matrix. The…

Statistics Theory · Mathematics 2012-05-31 G. M. Pan , J. Gao , Y. Yang , M. Guo

Label-free model evaluation, or AutoEval, estimates model accuracy on unlabeled test sets, and is critical for understanding model behaviors in various unseen environments. In the absence of image labels, based on dataset representations,…

Computer Vision and Pattern Recognition · Computer Science 2021-12-02 Xiaoxiao Sun , Yunzhong Hou , Hongdong Li , Liang Zheng

In this work we study the identification of spatial correlation in distributions of 2D scalar fields, presented across different forms of visual displays. We study simple visual displays that directly show color-mapped scalar fields, namely…

Human-Computer Interaction · Computer Science 2025-07-25 Yayan Zhao , Matthew Berger

In this letter, a novel method for change detection is proposed using neighborhood structure correlation. Because structure features are insensitive to the intensity differences between bi-temporal images, we perform the correlation…

Computer Vision and Pattern Recognition · Computer Science 2023-02-13 Mengmeng Wang , Zhiqiang Han , Peizhen Yang , Bai Zhu , Ming Hao , Jianwei Fan , Yuanxin Ye

It is of critical importance to be aware of the historical discrimination embedded in the data and to consider a fairness measure to reduce bias throughout the predictive modeling pipeline. Given various notions of fairness defined in the…

Machine Learning · Computer Science 2023-01-02 Hadis Anahideh , Nazanin Nezami , Abolfazl Asudeh

The expense of acquiring labels in large-scale statistical machine learning makes partially and weakly-labeled data attractive, though it is not always apparent how to leverage such data for model fitting or validation. We present a…

Machine Learning · Statistics 2022-02-10 Maxime Cauchois , Suyash Gupta , Alnur Ali , John Duchi

Graphical models describe associations between variables through the notion of conditional independence. Gaussian graphical models are a widely used class of such models where the relationships are formalized by non-null entries of the…

Methodology · Statistics 2023-08-08 Sagnik Bhadury , Riten Mitra , Jeremy T. Gaskins

Multi-label learning has attracted significant interests in computer vision recently, finding applications in many vision tasks such as multiple object recognition and automatic image annotation. Associating multiple labels to a complex…

Computer Vision and Pattern Recognition · Computer Science 2016-08-05 Hao Yang , Joey Tianyi Zhou , Jianfei Cai

We propose a high-dimensional white noise test that captures serial correlations within and across component series without specifying an alternative model. The test statistic is a U-statistic based on sample autocovariances. Under the…

Methodology · Statistics 2026-05-07 Yuanya Xu

This study introduces a data-driven, machine learning-based method to detect suitable control variables and instruments for assessing the causal effect of a treatment on an outcome in observational data. Our approach tests the joint…

Econometrics · Economics 2026-05-20 Nicolas Apfel , Julia Hatamyar , Martin Huber , Jannis Kueck

The rapid growth in feature dimension may introduce implicit associations between features and labels in multi-label datasets, making the relationships between features and labels increasingly complex. Moreover, existing methods often adopt…

Machine Learning · Computer Science 2025-05-30 Wanfu Gao , Jun Gao , Qingqi Han , Hanlin Pan , Kunpeng Liu