English
Related papers

Related papers: Archetypal Analysis for Binary Data

200 papers

The AI revolution is data driven. AI "data wrangling" is the process by which unusable data is transformed to support AI algorithm development (training) and deployment (inference). Significant time is devoted to translating diverse data…

Databases · Computer Science 2020-01-22 Jeremy Kepner , Vijay Gadepally , Hayden Jananthan , Lauren Milechin , Siddharth Samsi

Topological Data Analysis (TDA) can broadly be described as a collection of data analysis methods that find structure in data. This includes: clustering, manifold estimation, nonlinear dimension reduction, mode estimation, ridge estimation…

Methodology · Statistics 2016-09-28 Larry Wasserman

We revisit a pioneer unsupervised learning technique called archetypal analysis, which is related to successful data analysis methods such as sparse coding and non-negative matrix factorization. Since it was proposed, archetypal analysis…

Computer Vision and Pattern Recognition · Computer Science 2014-05-27 Yuansi Chen , Julien Mairal , Zaid Harchaoui

Correspondence analysis (CA) is a popular technique to visualize the relationship between two categorical variables. CA uses the data from a two-way contingency table and is affected by the presence of outliers. The supplementary points…

Methodology · Statistics 2026-01-05 Qianqian Qi , David J. Hessen , Aike N. Vonk , Peter G. M. van der Heijden

We present a novel method for predicting binary phase diagrams through the automatic construction of a minimal basis set of representative templates. The core assumption is that any materials space can be divided into a small number of…

Materials Science · Physics 2024-10-03 Caja Annweiler , Simone Di Cataldo , Maurits W. Haverkort , Lilia Boeri

Given a collection of data points, non-negative matrix factorization (NMF) suggests to express them as convex combinations of a small set of `archetypes' with non-negative entries. This decomposition is unique only if the true archetypes…

Machine Learning · Statistics 2017-05-09 Hamid Javadi , Andrea Montanari

At the crossway of machine learning and data analysis, anomaly detection aims at identifying observations that exhibit abnormal behaviour. Be it measurement errors, disease development, severe weather, production quality default(s) (items)…

Methodology · Statistics 2025-06-06 Romain Valla , Pavlo Mozharovskyi , Florence d'Alché-Buc

We consider the problem of identifying patterns in a data set that exhibit anomalous behavior, often referred to as anomaly detection. In most anomaly detection algorithms, the dissimilarity between data samples is calculated by a single…

Machine Learning · Computer Science 2013-01-08 Ko-Jen Hsiao , Kevin S. Xu , Jeff Calder , Alfred O. Hero

Principal component analysis (PCA) for binary data, known as logistic PCA, has become a popular alternative to dimensionality reduction of binary data. It is motivated as an extension of ordinary PCA by means of a matrix factorization, akin…

Machine Learning · Statistics 2020-09-08 Andrew J. Landgraf , Yoonkyung Lee

Rational approximation is a powerful tool to obtain accurate surrogates for nonlinear functions that are easy to evaluate and linearize. The interpolatory adaptive Antoulas--Anderson (AAA) method is one approach to construct such…

Numerical Analysis · Mathematics 2024-06-27 Stefan Güttel , Daniel Kressner , Bart Vandereycken

The adaptive Antoulas-Anderson (AAA) algorithm for rational approximation is a widely used method for the efficient construction of highly accurate rational approximations to given data. While AAA can often produce rational approximations…

Numerical Analysis · Mathematics 2026-01-28 Michael S. Ackermann , Linus Balicki , Serkan Gugercin , Steffen W. R. Werner

We develop a new principal components analysis (PCA) type dimension reduction method for binary data. Different from the standard PCA which is defined on the observed data, the proposed PCA is defined on the logit transform of the success…

Applications · Statistics 2010-11-17 Seokho Lee , Jianhua Z. Huang , Jianhua Hu

Handling contaminated data poses a critical challenge in anomaly detection, as traditional models assume training on purely normal data. Conventional methods mitigate contamination by relying on fixed contamination ratios, but discrepancies…

Machine Learning · Computer Science 2025-11-27 Jungi Lee , Jungkwon Kim , Chi Zhang , Kwangsun Yoo , Seok-Joo Byun

Data augmentation methods are commonly integrated into the training of anomaly detection models. Previous approaches have primarily focused on replicating real-world anomalies or enhancing diversity, without considering that the standard of…

Artificial Intelligence · Computer Science 2024-12-30 Jiang Lin , Yaping Yan

Artificial neural networks that learn to perform Principal Component Analysis (PCA) and related tasks using strictly local learning rules have been previously derived based on the principle of similarity matching: similar pairs of inputs…

Computation · Statistics 2018-11-06 Victor Minden , Cengiz Pehlevan , Dmitri B. Chklovskii

Binary concepts are empirically used by humans to generalize efficiently. And they are based on Bernoulli distribution which is the building block of information. These concepts span both low-level and high-level features such as "large vs…

Machine Learning · Computer Science 2023-03-23 Zizhao Hu , Mohammad Rostami

Complex prediction models such as deep learning are the output from fitting machine learning, neural networks, or AI models to a set of training data. These are now standard tools in science. A key challenge with the current generation of…

Machine Learning · Computer Science 2022-10-21 Meng Liu , Tamal K. Dey , David F. Gleich

We propose the Autoencoding Binary Classifiers (ABC), a novel supervised anomaly detector based on the Autoencoder (AE). There are two main approaches in anomaly detection: supervised and unsupervised. The supervised approach accurately…

Machine Learning · Statistics 2019-03-27 Yuki Yamanaka , Tomoharu Iwata , Hiroshi Takahashi , Masanori Yamada , Sekitoshi Kanai

Cross-modal Retrieval methods build similarity relations between vision and language modalities by jointly learning a common representation space. However, the predictions are often unreliable due to the Aleatoric uncertainty, which is…

Computer Vision and Pattern Recognition · Computer Science 2024-01-17 Hao Li , Jingkuan Song , Lianli Gao , Xiaosu Zhu , Heng Tao Shen

Canonical correlation analysis (CCA) is a statistical learning method that seeks to build view-independent latent representations from multi-view data. This method has been successfully applied to several pattern analysis tasks such as…

Computer Vision and Pattern Recognition · Computer Science 2018-12-24 Hichem Sahbi