中文
相关论文

相关论文: Class Density and Dataset Quality in High-Dimensio…

200 篇论文

Datasets serve as crucial training resources and model performance trackers. However, existing datasets have exposed a plethora of problems, inducing biased models and unreliable evaluation results. In this paper, we propose a…

计算与语言 · 计算机科学 2022-12-20 Chengwen Wang , Qingxiu Dong , Xiaochen Wang , Haitao Wang , Zhifang Sui

In the era of rapidly increasing amounts of time series data, classification of variable objects has become the main objective of time-domain astronomy. Classification of irregularly sampled time series is particularly difficult because the…

天体物理仪器与方法 · 物理学 2015-05-21 Sven Dennis Kügler , Nikos Gianniotis , Kai Lars Polsterer

Data Mining is the process of extracting useful patterns from the huge amount of database and many data mining techniques are used for mining these patterns. Recently, one of the remarkable facts in higher educational institute is the rapid…

人工智能 · 计算机科学 2014-05-16 Priyanka Saini

Interpreting experimental data in high school experiments can be a difficult task for students, especially when there is large variation in the data. At the same time, calculating the standard deviation poses a challenge for students. In…

物理教育 · 物理学 2022-10-18 Karel Kok , Burkhard Priemer

The idea underlying the modal formulation of density-based clustering is to associate groups with the regions around the modes of the probability density function underlying the data. This correspondence between clusters and dense regions…

社会与信息网络 · 计算机科学 2021-01-22 Giovanna Menardi , Domenico De Stefano

In the universal quest to optimize machine-learning classifiers, three factors -- model architecture, dataset size, and class balance -- have been shown to influence test-time performance but do not fully account for it. Previously,…

机器学习 · 计算机科学 2025-06-05 Josiah Couch , Miao Li , Rima Arnaout , Ramy Arnaout

Synthetic data generation is a promising technique to facilitate the use of sensitive data while mitigating the risk of privacy breaches. However, for synthetic data to be useful in downstream analysis tasks, it needs to be of sufficient…

机器学习 · 统计学 2024-08-26 Thom Benjamin Volker , Peter-Paul de Wolf , Erik-Jan van Kesteren

We introduce the idea of Data Readiness Level (DRL) to measure the relative richness of data to answer specific questions often encountered by data scientists. We first approach the problem in its full generality explaining its desired…

信息检索 · 计算机科学 2017-02-08 Hui Guan , Thanos Gentimis , Hamid Krim , James Keiser

The ratio between two probability density functions is an important component of various tasks, including selection bias correction, novelty detection and classification. Recently, several estimators of this ratio have been proposed. Most…

统计方法学 · 统计学 2014-04-30 Rafael Izbicki , Ann B. Lee , Chad M. Schafer

The notion of probability density for a random function is not as straightforward as in finite-dimensional cases. While a probability density function generally does not exist for functional data, we show that it is possible to develop the…

统计理论 · 数学 2010-03-01 Aurore Delaigle , Peter Hall

Data analysis in high-dimensional spaces aims at obtaining a synthetic description of a data set, revealing its main structure and its salient features. We here introduce an approach providing this description in the form of a topography of…

机器学习 · 统计学 2021-03-02 Maria d'Errico , Elena Facco , Alessandro Laio , Alex Rodriguez

Large scale image datasets are a growing trend in the field of machine learning. However, it is hard to quantitatively understand or specify how various datasets compare to each other - i.e., if one dataset is more complex or harder to…

计算机视觉与模式识别 · 计算机科学 2020-08-12 Ameet Annasaheb Rahane , Anbumani Subramanian

In deep learning, achieving high performance on image classification tasks requires diverse training sets. However, the current best practice$\unicode{x2013}$maximizing dataset size and class balance$\unicode{x2013}$does not guarantee…

计算机视觉与模式识别 · 计算机科学 2024-08-02 Josiah Couch , Rima Arnaout , Ramy Arnaout

Data acquisition processes for machine learning are often costly. To construct a high-performance prediction model with fewer data, a degree of difficulty in prediction is often deployed as the acquisition function in adding a new data…

机器学习 · 计算机科学 2022-04-27 Bongjoon Park , Eunkyung Koh

We first exhibit a multimodal image registration task, for which a neural network trained on a dataset with noisy labels reaches almost perfect accuracy, far beyond noise variance. This surprising auto-denoising phenomenon can be explained…

机器学习 · 计算机科学 2021-02-11 Guillaume Charpiat , Nicolas Girard , Loris Felardos , Yuliya Tarabalka

Supervised deep learning models require significant amount of labeled data to achieve an acceptable performance on a specific task. However, when tested on unseen data, the models may not perform well. Therefore, the models need to be…

计算机视觉与模式识别 · 计算机科学 2024-01-01 Akshit Achara , Ram Krishna Pandey

Multicellular systems play a key role in bioprocess and biomedical engineering. Cell ensembles encountered in these setups show phenotypic variability like size and biochemical composition. As this variability may result in undesired…

系统与控制 · 计算机科学 2018-07-16 Armin Küper , Robert Dürr , Steffen Waldherr

Deep Learning performs well when training data densely covers the experience space. For complex problems this makes data collection prohibitively expensive. We propose to intelligently select samples when constructing data sets in order to…

计算机视觉与模式识别 · 计算机科学 2020-04-01 Mark Philip Philipsen , Thomas Baltzer Moeslund

Given i.i.d samples from some unknown continuous density on hyper-rectangle $[0, 1]^d$, we attempt to learn a piecewise constant function that approximates this underlying density non-parametrically. Our density estimate is defined on a…

机器学习 · 统计学 2015-09-24 Kun Yang , Hao Su , Wing Hung Wang

Current algorithms and architecture can create excellent DNN classifier models from example data. In general, larger training datasets result in better model estimations, which improve test performance. Existing methods for predicting…

机器学习 · 计算机科学 2023-05-26 Nathaniel Dean , Dilip Sarkar