中文
相关论文

相关论文: A robust multivariate, non-parametric outlier iden…

200 篇论文

Functional Magnetic Resonance Imaging (fMRI) is a neuroimaging technique with pivotal importance due to its scientific and clinical applications. As with any widely used imaging modality, there is a need to ensure the quality of the same,…

图像与视频处理 · 电气工程与系统科学 2020-09-29 David Calhas , Rui Henriques

Nonnegative Matrix Factorization (NMF) is a widely used technique in many applications such as face recognition, motion segmentation, etc. It approximates the nonnegative data in an original high dimensional space with a linear…

机器学习 · 计算机科学 2012-04-12 Bin Shen , Luo Si , Rongrong Ji , Baodi Liu

Large outliers break down linear and nonlinear regression models. Robust regression methods allow one to filter out the outliers when building a model. By replacing the traditional least squares criterion with the least trimmed squares…

最优化与控制 · 数学 2012-06-07 Gleb Beliakov , Andrei Kelarev , John Yearwood

We study the algorithmic problem of robust mean estimation of an identity covariance Gaussian in the presence of mean-shift contamination. In this contamination model, we are given a set of points in $\mathbb{R}^d$ generated i.i.d. via the…

数据结构与算法 · 计算机科学 2025-02-21 Ilias Diakonikolas , Giannis Iakovidis , Daniel M. Kane , Thanasis Pittas

Indoor localization is critical for IoT applications, yet challenges such as non-Gaussian noise, environmental interference, and measurement outliers hinder the robustness of traditional methods. Existing approaches, including Kalman…

系统与控制 · 电气工程与系统科学 2025-05-14 Zhiyi Zhou , Dongzhuo Liu , Songtao Guo , Yuanyuan Yang

The leading workhorse of anomaly (and attack) detection in the literature has been residual-based detectors, where the residual is the discrepancy between the observed output provided by the sensors (inclusive of any tampering along the…

系统与控制 · 电气工程与系统科学 2020-04-17 Navid Hashemi , Eduardo Verdugo German , Jonatan Pena Ramirez , Justin Ruths

We propose an inlier-based outlier detection method capable of both identifying the outliers and explaining why they are outliers, by identifying the outlier-specific features. Specifically, we employ an inlier-based outlier detection…

机器学习 · 统计学 2017-02-22 Makoto Yamada , Song Liu , Samuel Kaski

Improving the retrieval relevance on noisy datasets is an emerging need for the curation of a large-scale clean dataset in the medical domain. While existing methods can be applied for class-wise retrieval (aka. inter-class), they cannot…

计算机视觉与模式识别 · 计算机科学 2022-04-08 Xiaoyuan Guo , Jiali Duan , Saptarshi Purkayastha , Hari Trivedi , Judy Wawira Gichoya , Imon Banerjee

Robust estimators of location and dispersion are often used in the elliptical model to obtain an uncontaminated and highly representative subsample by trimming the data outside an ellipsoid based in the associated Mahalanobis distance. Here…

统计理论 · 数学 2016-08-14 Juan A. Cuesta-Albertos , Carlos Matrán , Agustín Mayo-Iscar

Distributional data analysis, concerned with statistical analysis and modeling for data objects consisting of random probability density functions (PDFs) in the framework of functional data analysis (FDA), has received considerable interest…

统计方法学 · 统计学 2021-10-05 Xinyi Lei , Zhicheng Chen , Hui Li

Advances in sensor technology have enabled the collection of large-scale datasets. Such datasets can be extremely noisy and often contain a significant amount of outliers that result from sensor malfunction or human operation faults. In…

机器学习 · 计算机科学 2018-08-28 Yu-Hsuan Kuo , Zhenhui Li , Daniel Kifer

Sources of bias in empirical studies can be separated in those coming from the modelling domain (e.g. multicollinearity) and those coming from outliers. We propose a two-step approach to counter both issues. First, by decontaminating data…

综合经济学 · 经济学 2019-02-14 Mathias Kloss , Thomas Kirschstein , Steffen Liebscher , Martin Petrick

This paper examines the problem of locating outlier columns in a large, otherwise low-rank matrix, in settings where {}{the data} are noisy, or where the overall matrix has missing elements. We propose a randomized two-step inference…

信息论 · 计算机科学 2016-12-12 Xingguo Li , Jarvis Haupt

Linear mixed models (LMMs) are a popular class of methods for analyzing longitudinal and clustered data. However, such models can be sensitive to outliers, and this can lead to biased inference on model parameters and inaccurate prediction…

统计方法学 · 统计学 2025-03-28 Shonosuke Sugasawa , Francis K. C. Hui , Alan H. Welsh

Learning expressive low-dimensional representations of ultrahigh-dimensional data, e.g., data with thousands/millions of features, has been a major way to enable learning methods to address the curse of dimensionality. However, existing…

机器学习 · 计算机科学 2018-06-14 Guansong Pang , Longbing Cao , Ling Chen , Huan Liu

Model averaging is an alternative to model selection for dealing with model uncertainty, which is widely used and very valuable. However, most of the existing model averaging methods are proposed based on the least squares loss function,…

统计方法学 · 统计学 2019-10-29 Miaomiao Wang , Guohua Zou

Data depth is an efficient tool for robustly summarizing the distribution of functional data and detecting potential magnitude and shape outliers. Commonly used functional data depth notions, such as the modified band depth and extremal…

统计方法学 · 统计学 2023-11-07 Cristian F. Jimenez-Varon , Fouzi Harrou , Ying Sun

Many computer vision tasks involve processing large amounts of data contaminated by outliers, which need to be detected and rejected. While outlier detection methods based on robust statistics have existed for decades, only recently have…

计算机视觉与模式识别 · 计算机科学 2017-04-14 Chong You , Daniel P. Robinson , René Vidal

Diagnosing and cleaning data is a crucial step for building robust machine learning systems. However, identifying problems within large-scale datasets with real-world distributions is challenging due to the presence of complex issues such…

机器学习 · 计算机科学 2023-10-31 Jang-Hyun Kim , Sangdoo Yun , Hyun Oh Song

Dialect classification is used in a variety of applications, such as machine translation and speech recognition, to improve the overall performance of the system. In a real-world scenario, a deployed dialect classification model can…

计算与语言 · 计算机科学 2024-03-26 Sourya Dipta Das , Yash Vadi , Abhishek Unnam , Kuldeep Yadav