中文
相关论文

相关论文: Class Density and Dataset Quality in High-Dimensio…

200 篇论文

In high-dimension, low-sample size (HDLSS) data, it is not always true that closeness of two objects reflects a hidden cluster structure. We point out the important fact that it is not the closeness, but the "values" of distance that…

机器学习 · 统计学 2013-12-30 Yoshikazu Terada

This work proposes and evaluates a novel approach to determine interesting categorical attributes for lists of entities. Once identified, such categories are of immense value to allow constraining (filtering) a current view of a user to…

数据库 · 计算机科学 2017-11-30 Koninika Pal , Sebastian Michel

With the rapid increase of large-scale, real-world datasets, it becomes critical to address the problem of long-tailed data distribution (i.e., a few classes account for most of the data, while most classes are under-represented). Existing…

计算机视觉与模式识别 · 计算机科学 2019-01-18 Yin Cui , Menglin Jia , Tsung-Yi Lin , Yang Song , Serge Belongie

With the widespread application of Large Language Models (LLMs) to various domains, concerns regarding the trustworthiness of LLMs in safety-critical scenarios have been raised, due to their unpredictable tendency to hallucinate and…

计算与语言 · 计算机科学 2024-11-04 Xin Qiu , Risto Miikkulainen

We define a general method for finding a quasi-best approximant in sup-norm to a target density belonging to a given model, based on independent samples drawn from distributions which average to the target (which does not necessarily belong…

统计理论 · 数学 2025-06-26 Guillaume Maillard

We propose a way of transforming the problem of conditional density estimation into a single nonparametric regression task via the introduction of auxiliary samples. This allows leveraging regression methods that work well in high…

机器学习 · 统计学 2025-11-25 Alexander G. Reisach , Olivier Collier , Alex Luedtke , Antoine Chambaz

Some examples are easier for humans to classify than others. The same should be true for deep neural networks (DNNs). We use the term example perplexity to refer to the level of difficulty of classifying an example. In this paper, we…

机器学习 · 计算机科学 2022-03-18 Nevin L. Zhang , Weiyan Xie , Zhi Lin , Guanfang Dong , Xiao-Hui Li , Caleb Chen Cao , Yunpeng Wang

Data complexity is an important concept in the natural sciences and related areas, but lacks a rigorous and computable definition. In this paper, we focus on a particular sense of complexity that is high if the data is structured in a way…

计算机视觉与模式识别 · 计算机科学 2025-03-21 Louis Mahon

In this paper, we address an issue of finding explainable clusters of class-uniform data in labelled datasets. The issue falls into the domain of interpretable supervised clustering. Unlike traditional clustering, supervised clustering aims…

机器学习 · 计算机科学 2023-07-18 Natallia Kokash , Leonid Makhnist

Recent advances in measuring hardness-wise properties of data guide language models in sample selection within low-resource scenarios. However, class-specific properties are overlooked for task setup and learning. How will these properties…

计算与语言 · 计算机科学 2024-07-18 Fengyu Cai , Xinran Zhao , Hongming Zhang , Iryna Gurevych , Heinz Koeppl

Although generative models have made remarkable progress in recent years, their use in critical applications has been hindered by an inability to reliably evaluate the quality of their generated samples. Quality refers to at least two…

机器学习 · 计算机科学 2026-02-18 Nicolas Salvy , Hugues Talbot , Bertrand Thirion

This paper presents new methodology for computationally efficient kernel density estimation. It is shown that a large class of kernels allows for exact evaluation of the density estimates using simple recursions. The same methodology can be…

统计计算 · 统计学 2019-11-12 David P. Hofmeyr

Density estimation plays a crucial role in many data analysis tasks, as it infers a continuous probability density function (PDF) from discrete samples. Thus, it is used in tasks as diverse as analyzing population data, spatial locations in…

机器学习 · 计算机科学 2021-07-26 Patrik Puchert , Pedro Hermosilla , Tobias Ritschel , Timo Ropinski

The question of how best to estimate a continuous probability density from finite data is an intriguing open problem at the interface of statistics and physics. Previous work has argued that this problem can be addressed in a natural way…

数据分析、统计与概率 · 物理学 2014-07-16 Justin B. Kinney

The objective of goodness-of-fit testing is to assess whether a dataset of observations is likely to have been drawn from a candidate probability distribution. This paper presents a rank-based family of goodness-of-fit tests that is…

Hyperdimensional (HD) computing is built upon its unique data type referred to as hypervectors. The dimension of these hypervectors is typically in the range of tens of thousands. Proposed to solve cognitive tasks, HD computing aims at…

机器学习 · 计算机科学 2020-06-08 Lulu Ge , Keshab K. Parhi

We show that, for each of five datasets of increasing complexity, certain training samples are more informative of class membership than others. These samples can be identified a priori to training by analyzing their position in reduced…

机器学习 · 计算机科学 2022-02-08 Adam Byerly , Tatiana Kalganova

Data classification is a major machine learning paradigm, which has been widely applied to solve a large number of real-world problems. Traditional data classification techniques consider only physical features (e.g., distance, similarity,…

机器学习 · 计算机科学 2020-11-12 Esteban Vilca , Liang Zhao

High-quality data is key to interpretable and trustworthy data analytics and the basis for meaningful data-driven decisions. In practical scenarios, data quality is typically associated with data preprocessing, profiling, and cleansing for…

数据库 · 计算机科学 2019-07-19 Lisa Ehrlinger , Elisa Rusz , Wolfram Wöß

New proposed models are often compared to state-of-the-art using statistical significance testing. Literature is scarce for classifier comparison using metrics other than accuracy. We present a survey of statistical methods that can be used…

机器学习 · 计算机科学 2016-11-17 Lovedeep Gondara