中文
相关论文

相关论文: MacroPARAFAC for handling rowwise and cellwise out…

200 篇论文

Existing methods for anomaly detection often fall short due to their inability to handle the complexity, heterogeneity, and high dimensionality inherent in real-world mobility data. In this paper, we propose DeepBayesic, a novel framework…

机器学习 · 计算机科学 2024-10-07 Minxuan Duan , Yinlong Qian , Lingyi Zhao , Zihao Zhou , Zeeshan Rasheed , Rose Yu , Khurram Shafique

Heterogeneous multi-typed, multimodal relational data is increasingly available in many domains and their exploratory analysis poses several challenges. We advance the state-of-the-art in neural unsupervised learning to analyze such data.…

机器学习 · 统计学 2021-09-28 Ragunathan Mariappan , Vaibhav Rajan

There is a great need for robust techniques in data mining and machine learning contexts where many standard techniques such as principal component analysis and linear discriminant analysis are inherently susceptible to outliers.…

统计方法学 · 统计学 2015-09-28 Garth Tarr , Samuel Müller , Neville C. Weber

Principal component analysis (PCA) is a classical and widely used method for dimensionality reduction, with applications in data compression, computer vision, pattern recognition, and signal processing. However, PCA is designed for…

统计方法学 · 统计学 2025-10-01 Wenhui Wu , Changchun Shang , Jianhua Zhao , Xuan Ma , Yue Wang

We consider functional outlier detection from a geometric perspective, specifically: for functional data sets drawn from a functional manifold which is defined by the data's modes of variation in amplitude and phase. Based on this manifold,…

机器学习 · 统计学 2021-09-15 Moritz Herrmann , Fabian Scheipl

This paper develops an inferential theory for high-dimensional matrix-variate factor models with missing observations. We propose an easy-to-use all-purpose method that involves two straightforward steps. First, we perform principal…

统计方法学 · 统计学 2025-03-26 Yongxia Zhang , Jinwen Liang , Liwen Xu , Keming Yu , Maozai Tian

This paper presents a new modeling strategy for joint unsupervised analysis of multiple high-throughput biological studies. As in Multi-study Factor Analysis, our goals are to identify both common factors shared across studies and…

应用统计 · 统计学 2018-06-27 Roberta De Vito , Ruggero Bellio , Lorenzo Trippa , Giovanni Parmigiani

In one-class-learning tasks, only the normal case (foreground) can be modeled with data, whereas the variation of all possible anomalies is too erratic to be described by samples. Thus, due to the lack of representative data, the…

计算机视觉与模式识别 · 计算机科学 2019-06-03 Duc Tam Nguyen , Zhongyu Lou , Michael Klar , Thomas Brox

Computer-based assessments routinely generate detailed interaction logs -- commonly referred to as process data -- that record every action a respondent performs during task completion, yet systematic preprocessing guidance, integrated…

应用统计 · 统计学 2026-04-21 Daeun Hwangbo , Junyeong Park , Minjeong Jeon , Ick Hoon Jin

We propose VarFA, a variational inference factor analysis framework that extends existing factor analysis models for educational data mining to efficiently output uncertainty estimation in the model's estimated factors. Such uncertainty…

机器学习 · 统计学 2020-08-18 Zichao Wang , Yi Gu , Andrew Lan , Richard Baraniuk

Existing studies on identifying outliers in wind speed-power datasets are often challenged by the complicated and irregular distributions of outliers, especially those being densely stacked yet staying close to normal data. This could…

信号处理 · 电气工程与系统科学 2025-05-01 Limengqian Zheng , Lipeng Zhu , Weijia Wen , Jiayong Li , Cong Zhang

We propose a new class of semiparametric exponential family graphical models for the analysis of high dimensional mixed data. Different from the existing mixed graphical models, we allow the nodewise conditional distributions to be…

机器学习 · 统计学 2015-10-16 Zhuoran Yang , Yang Ning , Han Liu

A well-known method for completing low-rank matrices based on convex optimization has been established by Cand{\`e}s and Recht. Although theoretically complete, the method may not entirely solve the low-rank matrix completion problem. This…

统计方法学 · 统计学 2014-07-17 Guangcan Liu , Ping Li

Factor models are widely applied to the analysis of multivariate data across disparate fields of research. However, modern scientific data are often incomplete, and estimating a factor model from partially observed data can be very…

统计方法学 · 统计学 2026-02-24 Giuseppe Vinci

Multiway datasets are commonly analyzed using unsupervised matrix and tensor factorization methods to reveal underlying patterns. Frequently, such datasets include timestamps and could correspond to, for example, health-related measurements…

机器学习 · 计算机科学 2025-02-27 Christos Chatzis , Carla Schenker , Jérémy E. Cohen , Evrim Acar

Many tasks in data mining and related fields can be formalized as matching between objects in two heterogeneous domains, including collaborative filtering, link prediction, image tagging, and web search. Machine learning techniques,…

机器学习 · 计算机科学 2014-10-24 Jingbo Shang , Tianqi Chen , Hang Li , Zhengdong Lu , Yong Yu

We study the underlying structure of data (approximately) generated from a union of independent subspaces. Traditional methods learn only one subspace, failing to discover the multi-subspace structure, while state-of-the-art methods analyze…

计算机视觉与模式识别 · 计算机科学 2018-01-30 Binghui Wang , Chuang Lin

In many inverse problems, model parameters cannot be precisely determined from observational data. Bayesian inference provides a mechanism for capturing the resulting parameter uncertainty, but typically at a high computational cost. This…

统计计算 · 统计学 2019-03-28 Matthew Parno , Tarek Moselhy , Youssef Marzouk

In an industrial context, the activity of sensors is recorded at a high frequency. A challenge is to automatically detect abnormal measurement behavior. Considering the sensor measures as functional data, the problem can be formulated as…

统计理论 · 数学 2022-03-09 Martial Amovin-Assagba , Irène Gannaz , Julien Jacques

The problem of outlier detection is extremely challenging in many domains such as text, in which the attribute values are typically non-negative, and most values are zero. In such cases, it often becomes difficult to separate the outliers…

信息检索 · 计算机科学 2017-01-06 Ramakrishnan Kannan , Hyenkyun Woo , Charu C. Aggarwal , Haesun Park