中文
相关论文

相关论文: Distance-based Data Cleaning: A Survey (Technical …

200 篇论文

This study introduces the Garbage Dataset (GD), a publicly available image dataset designed to advance automated waste segregation through machine learning and computer vision. It is a diverse dataset that covers 10 categories of common…

计算机视觉与模式识别 · 计算机科学 2026-03-04 Suman Kunwar

Classifying large-scale image data into object categories is an important problem that has received increasing research attention. Given the huge amount of data, non-parametric approaches such as nearest neighbor classifiers have shown…

计算机视觉与模式识别 · 计算机科学 2014-04-28 Zhaowen Wang , Jianchao Yang , Zhe Lin , Jonathan Brandt , Shiyu Chang , Thomas Huang

Classification is an important supervised machine learning method, which is necessary and challenging issue for ecological research. It offers a way to classify a dataset into subsets that share common patterns. Notably, there are many…

机器学习 · 统计学 2018-12-24 Md. Siraj-Ud-Doula , Md. Ashad Alam

Small, low-cost, wireless cameras are becoming increasingly commonplace making surreptitious observation of people more difficult to detect. Previous work in detecting hidden cameras has only addressed limited environments in small spaces…

密码学与安全 · 计算机科学 2020-03-31 Kevin Wu , Brent Lagesse

We survey permutation-based methods for approximate k-nearest neighbor search. In these methods, every data point is represented by a ranked list of pivots sorted by the distance to this point. Such ranked lists are called permutations. The…

机器学习 · 计算机科学 2016-11-01 Bilegsaikhan Naidan , Leonid Boytsov , Eric Nyberg

Missing data often exists in real-world datasets, requiring significant time and effort for data repair to learn accurate models. In this paper, we show that imputing all missing values is not always necessary to achieve an accurate ML…

机器学习 · 计算机科学 2026-03-19 Cheng Zhen , Prayoga , Nischal Aryal , Arash Termehchy , Garrett Biwer , Lubna Alzamil

Data corruption, including missing and noisy data, poses significant challenges in real-world machine learning. This study investigates the effects of data corruption on model performance and explores strategies to mitigate these effects…

机器学习 · 计算机科学 2025-05-22 Qi Liu , Wanjing Ma

Big data mining is well known to be an important task for data science, because it can provide useful observations and new knowledge hidden in given large datasets. Proximity-based data analysis is particularly utilized in many real-life…

数据库 · 计算机科学 2022-11-29 Daichi Amagata , Yusuke Arai , Sumio Fujita , Takahiro Hara

Data fusion describes the method of combining data from (at least) two initially independent data sources to allow for joint analysis of variables which are not jointly observed. The fundamental idea is to base inference on identifying…

统计方法学 · 统计学 2020-12-02 Florian Meinfelder , Jannik Schaller

Event data are prevalent in diverse domains such as financial trading, business workflows and industrial IoT nowadays. An event is often characterized by several attributes denoting the meaning associated with the corresponding occurrence…

数据库 · 计算机科学 2020-12-15 Ruihong Huang , Jianmin Wang

Multiplicity, the existence of equally good yet competing models, has received growing attention in recent years. While prior work has emphasized modelling choices, the critical role of data in shaping multiplicity has been largely…

机器学习 · 计算机科学 2026-02-03 Prakhar Ganesh , Hsiang Hsu , Golnoosh Farnadi

Data fusion has played an important role in data mining because high-quality data is required in a lot of applications. As on-line data may be out-of-date and errors in the data may propagate with copying and referring between sources, it…

数据库 · 计算机科学 2017-02-03 Yunfan Chen , Lei Chen , Chen Jason Zhang

Record linkage is aimed at the accurate and efficient identification of records that represent the same entity within or across disparate databases. It is a fundamental task in data integration and increasingly required for accurate…

数据库 · 计算机科学 2021-04-21 Thilina Ranbaduge , Peter Christen , Rainer Schnell

Link prediction is one of the fundamental problems in computational social science. A particularly common means to predict existence of unobserved links is via structural similarity metrics, such as the number of common neighbors; node…

社会与信息网络 · 计算机科学 2019-01-01 Kai Zhou , Tomasz P. Michalak , Talal Rahwan , Marcin Waniek , Yevgeniy Vorobeychik

Identifying similar materials, i.e., those sharing a certain property or feature, requires interoperable data of high quality. It also requires means to measure similarity. We demonstrate how a spectral fingerprint as a descriptor, combined…

材料科学 · 物理学 2022-09-21 Martin Kuban , Šimon Gabaj , Wahib Aggoune , Cecilia Vona , Santiago Rigamonti , Claudia Draxl

Nowadays, digital content is widespread and simply redistributable, either lawfully or unlawfully. For example, after images are posted on the internet, other web users can modify them and then repost their versions, thereby generating…

计算机视觉与模式识别 · 计算机科学 2020-09-08 K. K. Thyagharajan , G. Kalaiarasi

With the proliferation of multi-core hardware, parallel programs have become ubiquitous. These programs have their own type of bugs known as concurrency bugs and among them, data race bugs have been mostly in the focus of researchers over…

分布式、并行与集群计算 · 计算机科学 2019-07-17 Ali Tehrani , Mohammed Khaleel , Reza Akbari , Ali Jannesari

Spatial data is playing an emerging role in new technologies such as web and mobile mapping and Geographic Information Systems (GIS). Important decisions in political, social and many other aspects of modern human life are being made using…

数据库 · 计算机科学 2016-05-17 Bagher Saberi , Nasser Ghadiri

The Jaccard similarity index has often been employed in science and technology as a means to quantify the similarity between two sets. When modified to operate on real-valued values, the Jaccard similarity index can be applied to compare…

数据分析、统计与概率 · 物理学 2024-10-23 Gonzalo Travieso , Alexandre Benatti , Luciano da F. Costa

Data mining has various real-time applications in fields such as finance telecommunications, biology, and government. Classification is a primary task in data mining. With the rise of cloud computing, users can outsource and access their…

密码学与安全 · 计算机科学 2024-07-09 Gunjan Mishra , Kalyani Pathak , Yash Mishra , Pragati Jadhav , Vaishali Keshervani