English
Related papers

Related papers: RLE Plots: Visualising Unwanted Variation in High …

200 papers

Disentangled representation learning (DRL) aims to identify and decompose underlying factors behind observations, thus facilitating data perception and generation. However, current DRL approaches often rely on the unrealistic assumption…

Computer Vision and Pattern Recognition · Computer Science 2024-07-30 Baao Xie , Qiuyu Chen , Yunnan Wang , Zequn Zhang , Xin Jin , Wenjun Zeng

A powerful data transformation method named guided projections is proposed creating new possibilities to reveal the group structure of high-dimensional data in the presence of noise variables. Utilising projections onto a space spanned by a…

In this article, we introduce a two-way factor model for a high-dimensional data matrix and study the properties of the maximum likelihood estimation (MLE). The proposed model assumes separable effects of row and column attributes and…

Methodology · Statistics 2021-03-17 Gao Zhigen , Yuan Chaofeng , Jing Bingyi , Huang Wei , Guo Jianhua

Investigating relationships between variables in multi-dimensional data sets is a common task for data analysts and engineers. More specifically, it is often valuable to understand which ranges of which input variables lead to particular…

Machine Learning · Computer Science 2020-09-14 Johannes Knittel , Andres Lalama , Steffen Koch , Thomas Ertl

Hidden regular variation defines a subfamily of distributions satisfying multivariate regular variation on $\mathbb{E} = [0, \infty]^d \backslash \{(0,0, ..., 0) \} $ and models another regular variation on the sub-cone $\mathbb{E}^{(2)} =…

Probability · Mathematics 2010-09-07 Abhimanyu Mitra , Sidney I. Resnick

We propose Blue Noise Plots, two-dimensional dot plots that depict data points of univariate data sets. While often one-dimensional strip plots are used to depict such data, one of their main problems is visual clutter which results from…

Graphics · Computer Science 2021-02-25 Christian van Onzenoodt , Gurprit Singh , Timo Ropinski , Tobias Ritschel

One important issue commonly encountered in the analysis of microarray data is to decide which and how many genes should be selected for further studies. For discriminant microarray data analyses based on statistical models, such as the…

Quantitative Methods · Quantitative Biology 2009-11-09 Wentian Li , Fengzhu Sun , Ivo Grosse

Parallel coordinates plotting is one of the most popular methods for multivariate visualization. However, when applied to larger data sets, there tends to be a "black screen problem," with the screen becoming so cluttered and full that…

Human-Computer Interaction · Computer Science 2017-09-05 Vincent Yang , Harrison Nguyen , Norman Matloff , Yingkang Xie

Nonlinear dimension reduction (NLDR) techniques such as tSNE, and UMAP provide a low-dimensional representation of high-dimensional data ($p\text{-}D$) by applying a nonlinear transformation. NLDR often exaggerates random patterns. But NLDR…

Visualizing high dimensional data by projecting them into two or three dimensional space is one of the most effective ways to intuitively understand the data's underlying characteristics, for example their class neighborhood structure.…

Machine Learning · Computer Science 2020-04-06 Pitoyo Hartono

The study of multivariate extremes is dominated by multivariate regular variation, although it is well known that this approach does not provide adequate distinction between random vectors whose components are not always simultaneously…

Statistics Theory · Mathematics 2021-08-17 Natalia Nolde , Jennifer L. Wadsworth

Gene expression-based heterogeneity analysis has been extensively conducted. In recent studies, it has been shown that network-based analysis, which takes a system perspective and accommodates the interconnections among genes, can be more…

Methodology · Statistics 2023-08-09 Rong Li , Qingzhao Zhang , Shuangge Ma

A widely used tool in the study of risk, insurance and extreme values is the mean excess plot. One use is for validating a generalized Pareto model for the excess distribution. This paper investigates some theoretical and practical aspects…

Probability · Mathematics 2010-06-03 Souvik Ghosh , Sidney I Resnick

When fitting black box supervised learning models (e.g., complex trees, neural networks, boosted trees, random forests, nearest neighbors, local kernel-weighted methods, etc.), visualizing the main effects of the individual predictor…

Methodology · Statistics 2019-08-21 Daniel W. Apley , Jingyu Zhu

Feature selection of high-dimensional labeled data with limited observations is critical for making powerful predictive modeling accessible, scalable, and interpretable for domain experts. Spectroscopy data, which records the interaction…

Machine Learning · Computer Science 2022-02-10 Frantishek Akulich , Hadis Anahideh , Manaf Sheyyab , Dhananjay Ambre

Extremes play a special role in Anomaly Detection. Beyond inference and simulation purposes, probabilistic tools borrowed from Extreme Value Theory (EVT), such as the angular measure, can also be used to design novel statistical learning…

Machine Learning · Statistics 2016-04-01 Nicolas Goix , Anne Sabourin , Stéphan Clémençon

Machine learning algorithms using deep architectures have been able to implement increasingly powerful and successful models. However, they also become increasingly more complex, more difficult to comprehend and easier to fool. So far, most…

Machine Learning · Computer Science 2020-08-20 Alexander Schulz , Fabian Hinder , Barbara Hammer

In machine learning, a bias occurs whenever training sets are not representative for the test data, which results in unreliable models. The most common biases in data are arguably class imbalance and covariate shift. In this work, we aim to…

Machine Learning · Computer Science 2018-04-04 Patrick Glauner , Radu State , Petko Valtchev , Diogo Duarte

The trace plot is seldom used in meta-analysis, yet it is a very informative plot. In this article we define and illustrate what the trace plot is, and discuss why it is important. The Bayesian version of the plot combines the posterior…

Methodology · Statistics 2024-09-04 Christian Röver , David Rindskopf , Tim Friede

Observational studies of treatment effects require adjustment for confounding variables. However, causal inference methods typically cannot deliver perfect adjustment on all measured baseline variables, and there is often ambiguity about…

Methodology · Statistics 2024-02-16 Lauren D. Liao , Yeyi Zhu , Amanda L. Ngo , Rana F. Chehab , Samuel D. Pimentel