中文
相关论文

相关论文: Identifying Outliers using Influence Function of M…

200 篇论文

The Projection Congruent Subset (PCS) Outlyingness is a new index of multivariate outlyingness obtained by considering univariate projections of the data. Like many other outlier detection procedures, PCS searches for a subset which…

统计方法学 · 统计学 2013-08-01 Kaveh Vakili , Eric Schmitt

The interactive exploration of large and evolving datasets is challenging as relationships between underlying variables may not be fully understood. There may be hidden trends and patterns in the data that are worthy of further exploration…

机器学习 · 计算机科学 2023-04-06 A. Ravishankar Rao , Daniel Clarke , Subrata Garai , Soumyabrata Dey

Identifying outlier documents, whose content is different from the majority of the documents in a corpus, has played an important role to manage a large text collection. However, due to the absence of explicit information about the inlier…

信息检索 · 计算机科学 2021-11-29 Dongha Lee , Dongmin Hyun , Jiawei Han , Hwanjo Yu

Anomaly detection based on one-class classification algorithms is broadly used in many applied domains like image processing (e.g. detection of whether a patient is "cancerous" or "healthy" from mammography image), network intrusion…

机器学习 · 统计学 2017-07-14 Evgeny Burnaev , Pavel Erofeev , Dmitry Smolyakov

Many traditional methods for identifying changepoints can struggle in the presence of outliers, or when the noise is heavy-tailed. Often they will infer additional changepoints in order to fit the outliers. To overcome this problem, data…

统计方法学 · 统计学 2017-07-12 Paul Fearnhead , Guillem Rigaill

The accuracy of machine learning interatomic potentials suffers from reference data that contains numerical noise. Often originating from unconverged or inconsistent electronic-structure calculations, this noise is challenging to identify.…

机器学习 · 统计学 2026-02-10 Terry C. W. Lam , Niamh O'Neill , Christoph Schran , Lars L. Schaaf

Canonical correlation analysis is a technique to extract common features from a pair of multivariate data. In complex situations, however, it does not extract useful features because of its linearity. On the other hand, kernel method used…

机器学习 · 计算机科学 2007-05-23 Shotaro Akaho

Recent work has sought to understand the behavior of neural networks by comparing representations between layers and between different trained models. We examine methods for comparing neural network representations based on canonical…

机器学习 · 计算机科学 2019-07-22 Simon Kornblith , Mohammad Norouzi , Honglak Lee , Geoffrey Hinton

Emerging integrative analysis of genomic and anatomical imaging data which has not been well developed, provides invaluable information for the holistic discovery of the genomic structure of disease and has the potential to open a new…

基因组学 · 定量生物学 2014-09-16 Junhai Jiang , Nan Lin , Shicheng Guo , Jinyun Chen , Momiao Xiong

Many machine learning classification systems lack competency awareness. Specifically, many systems lack the ability to identify when outliers (e.g., samples that are distinct from and not represented in the training data distribution) are…

机器学习 · 计算机科学 2020-07-03 Matthew Cook , Alina Zare , Paul Gader

Identifying complex phenotypes from high-dimensional biological data is challenging due to the intricate interdependencies among different physiological indicators. Traditional approaches often focus on detecting outliers in single…

机器学习 · 统计学 2024-10-24 Yafei Shen , Tao Zhang , Zhiwei Liu , Kalliopi Kostelidou , Ying Xu , Ling Yang

Using methods of statistical mechanics, we analyse the effect of outliers on the supervised learning of a classification problem. The learning strategy aims at selecting informative examples and discarding outliers. We compare two…

无序系统与神经网络 · 物理学 2009-10-31 Rainer Dietrich , Manfred Opper

Many estimation problems in robotics, computer vision, and learning require estimating unknown quantities in the face of outliers. Outliers are typically the result of incorrect data association or feature matching, and it is common to have…

计算机视觉与模式识别 · 计算机科学 2021-03-25 Jingnan Shi , Heng Yang , Luca Carlone

Identification of input data points relevant for the classifier (i.e. serve as the support vector) has recently spurred the interest of researchers for both interpretability as well as dataset debugging. This paper presents an in-depth…

机器学习 · 计算机科学 2020-09-30 Dominique Mercier , Shoaib Ahmed Siddiqui , Andreas Dengel , Sheraz Ahmed

Many computer vision tasks involve processing large amounts of data contaminated by outliers, which need to be detected and rejected. While outlier detection methods based on robust statistics have existed for decades, only recently have…

计算机视觉与模式识别 · 计算机科学 2017-04-14 Chong You , Daniel P. Robinson , René Vidal

Causal learning is a beneficial approach to analyze the cause and effect relationships among variables in a dataset. A causal graph can be generated from a dataset using a particular causal algorithm, for instance, the PC algorithm or Fast…

机器学习 · 计算机科学 2019-10-09 Teny Handhayani , James Cussens

Principal component analysis (PCA) is one of the most fundamental procedures in exploratory data analysis and is the basic step in applications ranging from quantitative finance and bioinformatics to image analysis and neuroscience.…

数据结构与算法 · 计算机科学 2019-05-13 Fedor V. Fomin , Petr A. Golovach , Fahad Panolan , Kirill Simonov

We introduce OpportunityFinder, a code-less framework for performing a variety of causal inference studies with panel data for non-expert users. In its current state, OpportunityFinder only requires users to provide raw observational data…

机器学习 · 计算机科学 2023-09-26 Huy Nguyen , Prince Grover , Devashish Khatwani

Outlier is the term that indicates in statistics an anomalous observation, aberrant, clearly distant from others collected observations. The outliers are the subject to animated discussions in various contexts with regard to be or not to be…

应用统计 · 统计学 2014-03-24 Gianluca Rosso

In a corpus of data, outliers are either errors: mistakes in the data that are counterproductive, or are unique: informative samples that improve model robustness. Identifying outliers can lead to better datasets by (1) removing noise in…