English
Related papers

Related papers: RLE Plots: Visualising Unwanted Variation in High …

200 papers

Simulated high-dimensional data is useful for testing, validating, and improving algorithms used in dimension reduction, supervised and unsupervised learning. High-dimensional data is characterized by multiple variables that are dependent…

The central aim in this paper is to address variable selection questions in nonlinear and nonparametric regression. Motivated by statistical genetics, where nonlinear interactions are of particular interest, we introduce a novel and…

Methodology · Statistics 2018-08-28 Lorin Crawford , Seth R. Flaxman , Daniel E. Runcie , Mike West

In this work we present a visualization tool specifically tailored to deal with skewed data. The technique is based upon the use of two types of notched boxplots (the usual one, and one which is tuned for the skewness of the data), the…

Computation · Statistics 2014-03-04 R. Ospina , A. M. Larangeiras , A. C. Frery

Automated data extraction from research texts has been steadily improving, with the emergence of large language models (LLMs) accelerating progress even further. Extracting data from plots in research papers, however, has been such a…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Maciej P. Polak , Dane Morgan

Logistic regression is commonly used for modeling dichotomous outcomes. In the classical setting, where the number of observations is much larger than the number of parameters, properties of the maximum likelihood estimator in logistic…

Machine Learning · Statistics 2019-11-14 Fariborz Salehi , Ehsan Abbasi , Babak Hassibi

Gene expression analysis aims at identifying the genes able to accurately predict biological parameters like, for example, disease subtyping or progression. While accurate prediction can be achieved by means of many different techniques,…

Methodology · Statistics 2008-09-11 Christine De Mol , Sofia Mosci , Magali Traskine , Alessandro Verri

This study introduces a new method of visualizing complex tree structured objects. The usefulness of this method is illustrated in the context of detecting unexpected features in a data set of very large trees. The major contribution is a…

Applications · Statistics 2012-02-14 Burcu Aydin , Gabor Pataki , Haonan Wang , Alim Ladha , Elizabeth Bullitt , J. S. Marron

General log-linear models specified by non-negative integer design matrices have a potentially wide range of applications, although using models without the genuine overall effect, that is, ones which cannot be reparameterized to include a…

Methodology · Statistics 2023-01-02 Anna Klimova , Matthias Kuhn

Principal component analysis is a useful dimension reduction and data visualization method. However, in high dimension, low sample size asymptotic contexts, where the sample size is fixed and the dimension goes to infinity,a paradox has…

Applications · Statistics 2012-11-21 Dan Shen , Haipeng Shen , Hongtu Zhu , J. S. Marron

In this time when biased information, deep fakes, and propaganda proliferate, the accessibility of reliable data sources is more important than ever. National statistical institutes provide curated data that contain quantitative information…

Information Retrieval · Computer Science 2025-10-03 Gadir Suleymanli , Alexander Rogiers , Lucas Lageweg , Jefrey Lijffijt

One of the challenges in analyzing high-dimensional expression data is the detection of important biological signals. A common approach is to apply a dimension reduction method, such as principal component analysis. Typically, after…

Quantitative Methods · Quantitative Biology 2012-06-05 Andreas Lehrmann , Michael Huber , Aydin C. Polatkan , Albert Pritzkau , Kay Nieselt

As regression is a widely studied problem, many methods have been proposed to solve it, each of them often requiring setting different hyper-parameters. Therefore, selecting the proper method for a given application may be very difficult…

Machine Learning · Computer Science 2026-03-23 Nassime Mountasir , Baptiste Lafabregue , Bruno Albert , Nicolas Lachiche

As the complexity and volume of datasets have increased along with the capabilities of modular, open-source, easy-to-implement, visualization tools, scientists' need for, and appreciation of, data visualization has risen too. Until…

Instrumentation and Methods for Astrophysics · Physics 2018-05-30 Alyssa A. Goodman , Michelle A. Borkin , Thomas P. Robitaille

We present Clusterplot, a multi-class high-dimensional data visualization tool designed to visualize cluster-level information offering an intuitive understanding of the cluster inter-relations. Our unique plots leverage 2D blobs devised to…

Graphics · Computer Science 2021-03-05 Or Malkai , Min Lu , Daniel Cohen-Or

Developing Machine Learning (ML) algorithms for heterogeneous/mixed data is a longstanding problem. Many ML algorithms are not applicable to mixed data, which include numeric and non-numeric data, text, graphs and so on to generate…

Machine Learning · Computer Science 2022-06-15 Boris Kovalerchuk , Elijah McCoy

The ability to detect log anomalies from system logs is a vital activity needed to ensure cyber resiliency of systems. It is applied for fault identification or facilitate cyber investigation and digital forensics. However, as logs…

Cryptography and Security · Computer Science 2023-11-10 Jonathan Pan , Swee Liang Wong , Yidi Yuan

Visualization of high-dimensional data is counter-intuitive using conventional graphs. Parallel coordinates are proposed as an alternative to explore multivariate data more effectively. However, it is difficult to extract relevant…

Computation · Statistics 2019-05-27 Shaima Tilouche , Vahid Partovi Nia , Samuel Bassetto

This paper investigates the effectiveness of using the Random Projection Ensemble (RPE) approach in Quadratic Discriminant Analysis (QDA) for ultrahigh-dimensional classification problems. Classical methods such as Linear Discriminant…

Methodology · Statistics 2025-07-10 Annesha Deb , Minerva Mukhopadhyay , Subhajit Dutta

Undirected graphs are often used to describe high dimensional distributions. Under sparsity conditions, the graph can be estimated using $\ell_1$ penalization methods. However, current methods assume that the data are independent and…

Machine Learning · Statistics 2008-04-29 Shuheng Zhou , John Lafferty , Larry Wasserman

We study the problem of visualizing large-scale and high-dimensional data in a low-dimensional (typically 2D or 3D) space. Much success has been reported recently by techniques that first compute a similarity structure of the data points…

Machine Learning · Computer Science 2016-04-06 Jian Tang , Jingzhou Liu , Ming Zhang , Qiaozhu Mei