中文
相关论文

相关论文: Identifying Outliers using Influence Function of M…

200 篇论文

Many unsupervised kernel methods rely on the estimation of the kernel covariance operator (kernel CO) or kernel cross-covariance operator (kernel CCO). Both kernel CO and kernel CCO are sensitive to contaminated data, even when bounded…

机器学习 · 统计学 2017-05-12 Md. Ashad Alam , Kenji Fukumizu , Yu-Ping Wang

Outlying observations are frequently encountered across a wide spectrum of scientific domains, posing notable challenges to the generalizability of statistical models and the reproducibility of downstream analysis. They are identified…

统计方法学 · 统计学 2026-03-17 Dongliang Zhang , Masoud Asgharian , Martin A. Lindquist

Astronomers are increasingly faced with a deluge of information, and finding worthwhile targets of study in the sea of data can be difficult. Outlier identification studies are a method that can be used to focus investigations by presenting…

高能天体物理现象 · 物理学 2022-09-26 Dustin K. Swarm , Casey T. DeRoo , Yanan Liu , Samantha Watkins

Identifying significant subsets of the genes, gene shaving is an essential and challenging issue for biomedical research for a huge number of genes and the complex nature of biological networks,. Since positive definite kernel based methods…

机器学习 · 统计学 2018-09-06 Md. Ashad Alam , Mohammad Shahjama , Md. Ferdush Rahman

Multilinear Principal Component Analysis (MPCA) is an important tool for analyzing tensor data. It performs dimension reduction similar to PCA for multivariate data. However, standard MPCA is sensitive to outliers. It is highly influenced…

统计方法学 · 统计学 2026-03-18 Mehdi Hirari , Fabio Centofanti , Mia Hubert , Stefan Van Aelst

Principal component analysis (PCA) is a fundamental tool for analyzing multivariate data. Here the focus is on dimension reduction to the principal subspace, characterized by its projection matrix. The classical principal subspace can be…

统计方法学 · 统计学 2026-05-29 Fabio Centofanti , Mia Hubert , Peter J. Rousseeuw

Mendelian Randomisation (MR) uses genetic variants as instrumental variables to infer causal effects of exposures on an outcome. One key assumption of MR is that the genetic variants used as instrumental variables are independent of the…

统计方法学 · 统计学 2025-02-21 Maximilian M Mandl , Anne-Laure Boulesteix , Stephen Burgess , Verena Zuber

We propose a new method to visualize and detect shape outliers in samples of curves. In functional data analysis we observe curves defined over a given real interval and shape outliers are those curves that exhibit a different shape from…

统计计算 · 统计学 2013-10-01 Ana Arribas-Gil , Juan Romo

Correspondence analysis (CA) is a popular technique to visualize the relationship between two categorical variables. CA uses the data from a two-way contingency table and is affected by the presence of outliers. The supplementary points…

统计方法学 · 统计学 2026-01-05 Qianqian Qi , David J. Hessen , Aike N. Vonk , Peter G. M. van der Heijden

In medical imaging, outliers can contain hypo/hyper-intensities, minor deformations, or completely altered anatomy. To detect these irregularities it is helpful to learn the features present in both normal and abnormal images. However this…

计算机视觉与模式识别 · 计算机科学 2024-03-15 Jeremy Tan , Benjamin Hou , James Batten , Huaqi Qiu , Bernhard Kainz

We propose a novel procedure for outlier detection in functional data, in a semi-supervised framework. As the data is functional, we consider the coefficients obtained after projecting the observations onto orthonormal bases (wavelet, PCA).…

A core data-centric learning challenge is the identification of training samples that are detrimental to model performance. Influence functions serve as a prominent tool for this task and offer a robust framework for assessing training data…

机器学习 · 计算机科学 2025-11-04 Anshuman Chhabra , Bo Li , Jian Chen , Prasant Mohapatra , Hongfu Liu

Various technologies, including computer vision models, are employed for the automatic monitoring of manual assembly processes in production. These models detect and classify events such as the presence of components in an assembly area or…

计算机视觉与模式识别 · 计算机科学 2026-01-21 Anton Sergeev , Victor Minchenkov , Aleksei Soldatov , Vasiliy Kakurin , Yaroslav Mazikov

In statistics and machine learning, the traditional meaning of the terms `outlier' and `anomaly' is a case in the dataset that behaves differently from the bulk of the data. This raises suspicion that it may belong to a different…

统计方法学 · 统计学 2026-04-17 Mia Hubert , Jakob Raymaekers , Peter J. Rousseeuw

Smart metering infrastructures collect data almost continuously in the form of fine-grained long time series. These massive data series often have common daily patterns that are repeated between similar days or seasons and shared among…

统计方法学 · 统计学 2022-10-10 A. Elías , J. M. Morales , S. Pineda

The detection of outliers is of critical importance in the assurance of data quality. Outliers may exist in observed data or in data derived from these observed data, such as estimates and forecasts. An outlier may indicate a problem with…

统计方法学 · 统计学 2025-10-23 Charles D. Coleman , Thomas Bryan

Outlier detection for high-dimensional (HD) data is a popular topic in modern statistical research. However, one source of HD data that has received relatively little attention is functional magnetic resonance images (fMRI), which consists…

统计方法学 · 统计学 2016-10-25 Amanda F. Mejia , Mary Beth Nebel , Ani Eloyan , Brian Caffo , Martin A. Lindquist

Two central objects in constructive approximation, the Christoffel-Darboux kernel and the Christoffel function, are encoding ample information about the associated moment data and ultimately about the possible generating measures. We…

复变函数 · 数学 2019-04-30 Bernhard Beckermann , Mihai Putinar , Edward B. Saff , Nikos Stylianopoulos

It has become routine in neuroscience studies to measure brain networks for different individuals using neuroimaging. These networks are typically expressed as adjacency matrices, with each cell containing a summary of connectivity between…

统计方法学 · 统计学 2022-06-30 Pritam Dey , Zhengwu Zhang , David B. Dunson

A popular approach for comparing gene expression levels between (replicated) conditions of RNA sequencing data relies on counting reads that map to features of interest. Within such count-based methods, many flexible and advanced…

定量方法 · 定量生物学 2014-03-17 Xiaobei Zhou , Helen Lindsay , Mark D. Robinson
‹ 上一页 1 2 3 10 下一页 ›