中文
相关论文

相关论文: Detection of Two-Way Outliers in Multivariate Data…

200 篇论文

Anomaly detection aims to detect data that do not conform to regular patterns, and such data is also called outliers. The anomalies to be detected are often tiny in proportion, containing crucial information, and are suitable for…

机器学习 · 计算机科学 2023-06-06 Fan Xu , Nan Wang , Xibin Zhao

In this paper, we propose a general framework for combining evidence of varying quality to estimate underlying binary latent variables in the presence of restrictions imposed to respect the scientific context. The resulting algorithms…

统计方法学 · 统计学 2018-08-28 Zhenke Wu , Livia Casciola-Rosen , Antony Rosen , Scott L. Zeger

Anomalies in economic and financial data -- often linked to rare yet impactful events -- are of theoretical interest, but can also severely distort inference. Although outlier-robust methodologies can be used, many researchers prefer…

统计方法学 · 统计学 2025-09-01 Monica Billio , Roberto Casarin , Fausto Corradin , Antonio Peruzzi

Sources of bias in empirical studies can be separated in those coming from the modelling domain (e.g. multicollinearity) and those coming from outliers. We propose a two-step approach to counter both issues. First, by decontaminating data…

综合经济学 · 经济学 2019-02-14 Mathias Kloss , Thomas Kirschstein , Steffen Liebscher , Martin Petrick

There exist multiple methods to detect outliers in multivariate data in the literature, but most of them require to estimate the covariance matrix. The higher the dimension, the more complex the estimation of the matrix becoming impossible…

统计方法学 · 统计学 2020-12-01 P. Navarro-Esteban , J. A. Cuesta-Albertos

This paper proposes an approach for anomalous sound detection that incorporates outlier exposure and inlier modeling within a unified framework by multitask learning. While outlier exposure-based methods can extract features efficiently, it…

声音 · 计算机科学 2023-09-15 Yucong Zhang , Hongbin Suo , Yulong Wan , Ming Li

We consider the problem of parameter estimation using weakly supervised datasets, where a training sample consists of the input and a partially specified annotation, which we refer to as the output. The missing information in the annotation…

机器学习 · 计算机科学 2012-06-22 M. Pawan Kumar , Ben Packer , Daphne Koller

We propose a general approach to handle data contaminations that might disrupt the performance of feature selection and estimation procedures for high-dimensional linear models. Specifically, we consider the co-occurrence of mean-shift and…

统计方法学 · 统计学 2021-06-23 Luca Insolia , Francesca Chiaromonte , Runze Li , Marco Riani

The spread of the Coronavirus disease-2019 epidemic has caused many courses and exams to be conducted online. The cheating behavior detection model in examination invigilation systems plays a pivotal role in guaranteeing the equality of…

计算机视觉与模式识别 · 计算机科学 2024-02-12 Yemeng Liu , Jing Ren , Jianshuo Xu , Xiaomei Bai , Roopdeep Kaur , Feng Xia

Measurement non-invariance arises when the psychometric properties of a scale differ across subgroups, undermining the validity of group comparisons. At the item level, such non-invariance manifests as differential item functioning (DIF),…

统计方法学 · 统计学 2026-01-27 Gabriel Wallin , Qi Huang

In an industrial context, the activity of sensors is recorded at a high frequency. A challenge is to automatically detect abnormal measurement behavior. Considering the sensor measures as functional data, the problem can be formulated as…

统计理论 · 数学 2022-03-09 Martial Amovin-Assagba , Irène Gannaz , Julien Jacques

This paper proposes an adaptive penalized weighted mean regression for outlier detection of high-dimensional data. In comparison to existing approaches based on the mean shift model, the proposed estimators demonstrate robustness against…

统计理论 · 数学 2023-06-27 Jiaqi Li , Linglong Kong , Bei Jiang , Wei Tu

Latent variable models are popularly used to measure latent factors (e.g., abilities and personalities) from large-scale assessment data. Beyond understanding these latent factors, the covariate effect on responses controlling for latent…

统计方法学 · 统计学 2026-01-12 Jing Ouyang , Chengyu Cui , Kean Ming Tan , Gongjun Xu

Hierarchical learning models, such as mixture models and Bayesian networks, are widely employed for unsupervised learning tasks, such as clustering analysis. They consist of observable and hidden variables, which represent the given data…

机器学习 · 统计学 2018-01-08 Keisuke Yamazaki

Dynamic factor models have a wide range of applications in econometrics and applied economics. The basic motivation resides in their capability of reducing a large set of time series to only few indicators (factors). If the number of time…

统计理论 · 数学 2009-09-29 Roberto Baragona , Francesco Battaglia

In large scale multiple testing problems, a two-class empirical Bayes approach can be used to control the false discovery rate (Fdr) for the entire array of hypotheses under study. A sample splitting step is incorporated to modify that…

统计计算 · 统计学 2019-12-13 Paramita Chakraborty , Chong Ma , John Grego , James Lynch

This work proposes an unsupervised learning framework for trajectory (sequence) outlier detection that combines ranking tests with user sequence models. The overall framework identifies sequence outliers at a desired false positive rate…

机器学习 · 计算机科学 2021-11-09 Mohamed A. Zahran , Leonardo Teixeira , Vinayak Rao , Bruno Ribeiro

This paper studies the construction of p-values for nonparametric outlier detection, taking a multiple-testing perspective. The goal is to test whether new independent samples belong to the same distribution as a reference data set or are…

统计方法学 · 统计学 2024-03-12 Stephen Bates , Emmanuel Candès , Lihua Lei , Yaniv Romano , Matteo Sesia

A general framework of latent trait item response models for continuous responses is given. In contrast to classical test theory models, which traditionally distinguish between true scores and error scores, the responses are clearly linked…

统计方法学 · 统计学 2022-04-11 Gerhard Tutz , Pascal Jordan

Subgroup analysis evaluates treatment effects across multiple sub-populations. When subgroups are defined by latent memberships inferred from imperfect measurements, the analysis typically involves two inter-connected models, a latent class…

统计方法学 · 统计学 2026-01-05 Yuanhui Luo , Xinzhou Guo , Yuqi Gu