English
Related papers

Related papers: Bias Detection via Maximum Subgroup Discrepancy

200 papers

Measuring geometric similarity between high-dimensional network representations is a topic of longstanding interest to neuroscience and deep learning. Although many methods have been proposed, only a few works have rigorously analyzed their…

Machine Learning · Statistics 2023-12-12 Dean A. Pospisil , Brett W. Larsen , Sarah E. Harvey , Alex H. Williams

Distance-based methods involve the computation of distance values between features and are a well-established paradigm in machine learning. In anomaly detection, anomalies are identified by their large distance from normal data points.…

Instrumentation and Methods for Astrophysics · Physics 2025-10-29 Siddharth Chaini , Federica B. Bianco , Ashish Mahabal

Community detection becomes an important problem with the booming of social networks. The Medoid-Shift algorithm preserves the benefits of Mean-Shift and can be applied to problems based on distance matrix, such as community detection. One…

Social and Information Networks · Computer Science 2023-09-20 Jie Hou , Jiakang Li , Xiaokang Peng , Wei Ke , Yonggang Lu

Continuous learning from an immense volume of data streams becomes exceptionally critical in the internet era. However, data streams often do not conform to the same distribution over time, leading to a phenomenon called concept drift.…

Machine Learning · Computer Science 2024-07-09 Ke Wan , Yi Liang , Susik Yoon

The distance standard deviation, which arises in distance correlation analysis of multivariate data, is studied as a measure of spread. The asymptotic distribution of the empirical distance standard deviation is derived under the assumption…

Statistics Theory · Mathematics 2019-12-12 Dominic Edelmann , Donald Richards , Daniel Vogel

Bayesian Neural Networks (BNNs) are trained to optimize an entire distribution over their weights instead of a single set, having significant advantages in terms of, e.g., interpretability, multi-task learning, and calibration. Because of…

Machine Learning · Computer Science 2022-10-07 Jary Pomponi , Simone Scardapane , Aurelio Uncini

We introduce a new one-parameter family of divergence measures, called bounded Bhattacharyya distance (BBD) measures, for quantifying the dissimilarity between probability distributions. These measures are bounded, symmetric and positive…

Statistics Theory · Mathematics 2016-04-12 Shivakumar Jolad , Ahmed Roman , Mahesh C. Shastry , Mihir Gadgil , Ayanendranath Basu

We develop a novel method for carrying out model selection for Bayesian autoencoders (BAEs) by means of prior hyper-parameter optimization. Inspired by the common practice of type-II maximum likelihood optimization and its equivalence to…

Semi-supervised classification, where unlabeled data are massive but labeled data are limited, often arises in machine learning applications. We address this challenge under high-dimensional data by leveraging the manifold and cluster…

Machine Learning · Statistics 2026-04-28 Ruoxu Tan , Yiming Zang

Quantifying degrees of fusion and separability between data groups in representation space is a fundamental problem in representation learning, particularly under domain shift. A meaningful metric should capture fusion-altering factors like…

Machine Learning · Computer Science 2026-01-30 Xiaolong Zhang , Jianwei Zhang , Xubo Song

The capability of reliably detecting out-of-distribution samples is one of the key factors in deploying a good classifier, as the test distribution always does not match with the training distribution in most real-world applications. In…

Machine Learning · Computer Science 2021-04-05 Dongha Lee , Sehun Yu , Hwanjo Yu

In generative modeling, the Wasserstein distance (WD) has emerged as a useful metric to measure the discrepancy between generated and real data distributions. Unfortunately, it is challenging to approximate the WD of high-dimensional…

Computer Vision and Pattern Recognition · Computer Science 2019-04-16 Jiqing Wu , Zhiwu Huang , Dinesh Acharya , Wen Li , Janine Thoma , Danda Pani Paudel , Luc Van Gool

In generative modeling, the Wasserstein distance (WD) has emerged as a useful metric to measure the discrepancy between generated and real data distributions. Unfortunately, it is challenging to approximate the WD of high-dimensional…

Computer Vision and Pattern Recognition · Computer Science 2019-04-17 Jiqing Wu , Zhiwu Huang , Dinesh Acharya , Wen Li , Janine Thoma , Danda Pani Paudel , Luc Van Gool

The Wasserstein distance received a lot of attention recently in the community of machine learning, especially for its principled way of comparing distributions. It has found numerous applications in several hard problems, such as domain…

Machine Learning · Statistics 2017-10-23 Nicolas Courty , Rémi Flamary , Mélanie Ducoffe

In many classification models, data is discretized to better estimate its distribution. Existing discretization methods often target at maximizing the discriminant power of discretized data, while overlooking the fact that the primary…

Machine Learning · Computer Science 2023-04-06 Shihe Wang , Jianfeng Ren , Ruibin Bai , Yuan Yao , Xudong Jiang

Generative models have been successfully used for generating realistic signals. Because the likelihood function is typically intractable in most of these models, the common practice is to use "implicit" models that avoid likelihood…

Machine Learning · Computer Science 2024-05-07 Itai Alon , Amir Globerson , Ami Wiesel

We propose a novel supervised learning method to optimize the kernel in the maximum mean discrepancy generative adversarial networks (MMD GANs), and the kernel support vector machines (SVMs). Specifically, we characterize a distributionally…

Machine Learning · Computer Science 2020-02-25 Masoud Badiei Khuzani , Liyue Shen , Shahin Shahrampour , Lei Xing

Large language models (LLMs) such as ChatGPT have exhibited remarkable performance in generating human-like texts. However, machine-generated texts (MGTs) may carry critical risks, such as plagiarism issues, misleading information, or…

Computation and Language · Computer Science 2024-03-01 Shuhai Zhang , Yiliao Song , Jiahao Yang , Yuanqing Li , Bo Han , Mingkui Tan

A frequent problem in statistical science is how to properly handle missing data in matched paired observations. There is a large body of literature coping with the univariate case. Yet, the ongoing technological progress in measuring…

Methodology · Statistics 2022-06-06 Marcos Matabuena , Paulo Félix , Marc Ditzhaus , Juan Vidal , Francisco Gude

We propose a methodology for intercomparing climate models and evaluating their performance against benchmarks based on the use of the Wasserstein distance (WD). This distance provides a rigorous way to measure quantitatively the difference…

Atmospheric and Oceanic Physics · Physics 2020-11-16 Gabriele Vissio , Valerio Lembo , Valerio Lucarini , Michael Ghil
‹ Prev 1 4 5 6 7 8 10 Next ›