English
Related papers

Related papers: Identifying Outliers using Influence Function of M…

200 papers

Kernel methods have been proven to be a powerful tool for the integration and analysis of highthroughput technologies generated data. Kernels offer a nonlinear version of any linear algorithm solely based on dot products. The kernelized…

Applications · Statistics 2024-11-27 Mitja Briscik , Marie-Agnès Dillies , Sébastien Déjean

We describe our methods that achieved the 3rd and 4th places in tasks 1 and 2, respectively, at ISIC challenge 2019. The goal of this challenge is to provide the diagnostic for skin cancer using images and meta-data. There are nine classes…

Machine Learning · Computer Science 2020-01-07 Andre G. C. Pacheco , Abder-Rahman Ali , Thomas Trappenberg

Robust PCA, the problem of PCA in the presence of outliers has been extensively investigated in the last few years. Here we focus on Robust PCA in the outlier model where each column of the data matrix is either an inlier or an outlier.…

Machine Learning · Statistics 2019-05-01 Vishnu Menon , Sheetal Kalyani

We aim to analyze the relation between two random vectors that may potentially have both different number of attributes as well as realizations, and which may even not have a joint distribution. This problem arises in many practical…

Machine Learning · Statistics 2015-11-12 Hoang-Vu Nguyen , Jilles Vreeken

This paper introduces a kernel discrepancy-based framework for rerandomization to enhance the precision of causal inference in controlled experiments. We demonstrate that the kernel discrepancy is the key part of the variance upper bound…

Methodology · Statistics 2025-11-05 Yiou Li , Lulu Kang

Due to the complexity of cancer, clustering algorithms have been used to disentangle the observed heterogeneity and identify cancer subtypes that can be treated specifically. While kernel based clustering approaches allow the use of more…

Machine Learning · Statistics 2018-11-21 Nora K. Speicher , Nico Pfeifer

In high reliability standards fields such as automotive, avionics or aerospace, the detection of anomalies is crucial. An efficient methodology for automatically detecting multivariate outliers is introduced. It takes advantage of the…

Methodology · Statistics 2018-08-01 Aurore Archimbaud , Klaus Nordhausen , Anne Ruiz-Gazen

We propose two new outlier detection methods, for identifying and classifying different types of outliers in (big) functional data sets. The proposed methods are based on an existing method called Massive Unsupervised Outlier Detection…

Methodology · Statistics 2021-10-15 Oluwasegun Taiwo Ojo , Antonio Fernández Anta , Rosa E. Lillo , Carlo Sguera

Inverse problems and, in particular, inferring unknown or latent parameters from data are ubiquitous in engineering simulations. A predominant viewpoint in identifying unknown parameters is Bayesian inference where both prior information…

Computation · Statistics 2022-08-31 Vahid Keshavarzzadeh , Robert M. Kirby , Akil Narayan

We address modelling and computational issues for multiple treatment effect inference under many potential confounders. Our main contribution is providing a trade-off between preventing the omission of relevant confounders, while not…

Methodology · Statistics 2025-02-04 Omiros Papaspiliopoulos , David Rossell , Miquel Torrens-i-Dinarès

Much of the natural variation for a complex trait can be explained by variation in DNA sequence levels. As part of sequence variation, gene-gene interaction has been ubiquitously observed in nature, where its role in shaping the development…

Applications · Statistics 2012-10-01 Shaoyu Li , Yuehua Cui

Outlier or anomaly detection is an important task in data analysis. We discuss the problem from a geometrical perspective and provide a framework that exploits the metric structure of a data set. Our approach rests on the manifold…

Machine Learning · Statistics 2022-08-01 Moritz Herrmann , Florian Pfisterer , Fabian Scheipl

This note investigates the problem of detecting outliers in longitudinal data. It compares well-known methods used in official statistics with proposals from the fields of data mining and machine learning that are based on the distance…

Methodology · Statistics 2025-07-30 Marcello D'Orazio

Integration of multi-omics data provides opportunities for revealing biological mechanisms related to certain phenotypes. We propose a novel method of multi-omics integration called supervised deep generalized canonical correlation analysis…

Quantitative Methods · Quantitative Biology 2022-04-21 Jeongyoung Hwang , Sehwan Moon , Hyunju Lee

Networks pervade many disciplines of science for analyzing complex systems with interacting components. In particular, this concept is commonly used to model interactions between genes and identify closely associated genes forming…

Methodology · Statistics 2015-06-02 Y. X. Rachel Wang , Keni Jiang , Lewis J. Feldman , Peter J. Bickel , Haiyan Huang

The additive genetic effect is arguably the most important quantity inferred in animal and plant breeding analyses. The term effect indicates that it represents causal information, which is different from standard statistical concepts as…

Quantitative Methods · Quantitative Biology 2015-04-27 Bruno Dourado Valente , Gota Morota , Guilherme Jordao Magalhaes Rosa , Daniel Gianola , Kent Weigel

Mendelian randomization (MR) has become an essential tool for causal inference in biomedical and public health research. By using genetic variants as instrumental variables, MR helps address unmeasured confounding and reverse causation,…

Methodology · Statistics 2025-11-04 Minhao Yao , Anqi Wang , Xihao Li , Zhonghua Liu

We develop a robust Bayesian functional principal component analysis (RB-FPCA) method that utilizes the skew elliptical class of distributions to model functional data, which are observed over a continuous domain. This approach effectively…

Methodology · Statistics 2025-04-15 Jiarui Zhang , Jiguo Cao , Liangliang Wang

We investigate the performance of robust estimates of multivariate location under nonstandard data contamination models such as componentwise outliers (i.e., contamination in each variable is independent from the other variables). This…

Statistics Theory · Mathematics 2009-03-04 Fatemah Alqallaf , Stefan Van Aelst , Victor J. Yohai , Ruben H. Zamar

Principal component analysis (PCA) is widely used for dimensionality reduction, with well-documented merits in various applications involving high-dimensional data, including computer vision, preference measurement, and bioinformatics. In…

Machine Learning · Statistics 2013-10-01 Gonzalo Mateos , Georgios B. Giannakis