English
Related papers

Related papers: ggskewboxplots: Enhanced Boxplots for Skewed Data …

200 papers

In this work we present a visualization tool specifically tailored to deal with skewed data. The technique is based upon the use of two types of notched boxplots (the usual one, and one which is tuned for the skewness of the data), the…

Computation · Statistics 2014-03-04 R. Ospina , A. M. Larangeiras , A. C. Frery

Tukey's boxplot is widely used for outlier detection; however, its classic fixed-fence rule tends to flag an excessive number of outliers as the sample size grows. To address this, we introduce two new R packages, ChauBoxplot and…

Methodology · Statistics 2026-03-04 Tiejun Tong , Hongmei Lin , Bowen Gang , Riquan Zhang

Skewness plays a relevant role in several multivariate statistical techniques. Sometimes it is used to recover data features, as in cluster analysis. In other circumstances, skewness impairs the performances of statistical methods, as in…

Computation · Statistics 2019-03-26 Cinzia Franceschini , Nicola Loperfido

The g-and-k and (generalised) g-and-h distributions are flexible univariate distributions which can model highly skewed or heavy tailed data through only four parameters: location and scale, and two shape parameters influencing the skewness…

Computation · Statistics 2017-06-22 Dennis Prangle

Because of its mathematical tractability, the Gaussian mixture model holds a special place in the literature for clustering and classification. For all its benefits, however, the Gaussian mixture model poses problems when the data is skewed…

Applications · Statistics 2020-11-19 Michael P. B. Gallaugher , Paul D. McNicholas , Volodymyr Melnykov , Xuwen Zhu

Tukey's boxplot is a foundational tool for exploratory data analysis, but its classic outlier-flagging rule does not account for the sample size, and subsequent modifications have often been presented as separate, heuristic adjustments. In…

Methodology · Statistics 2025-10-24 Bowen Gang , Hongmei Lin , Tiejun Tong

To solve the problems in measuring coefficient of skewness related to extreme value, irregular distance from the middle point and distance between two consecutive numbers, "Rank skewness" a new measure of the coefficient of skewness has…

Methodology · Statistics 2019-08-20 Ummay Salma Shorna , Md. Forhad Hossain

Kernel smoothers are essential tools for data analysis due to their ability to convey complex statistical information with concise graphical visualisations. Their inclusion in the base distribution and in the many user-contributed add-on…

Computation · Statistics 2024-09-17 Tarn Duong

Boxplots and related visualization methods are widely used exploratory tools for taking a first look at collections of univariate variables. In this note an extension is provided that is specifically designed to detect and display…

Methodology · Statistics 2026-05-05 Camille M. Montalcini , Peter J. Rousseeuw

The box-and-whisker plot, introduced by Tukey (1977), is one of the most popular graphical methods in descriptive statistics. On the other hand, however, Tukey's boxplot is free of sample size, yielding the so-called "one-size-fits-all"…

Methodology · Statistics 2025-06-10 Hongmei Lin , Riquan Zhang , Tiejun Tong

Scatterplots are a common tool for exploring multidimensional datasets, especially in the form of scatterplot matrices (SPLOMs). However, scatterplots suffer from overplotting when categorical variables are mapped to one or two axes, or the…

Human-Computer Interaction · Computer Science 2025-11-18 Deokgun Park , Sung-Hee Kim , Niklas Elmqvist

In recent work, robust mixture modelling approaches using skewed distributions have been explored to accommodate asymmetric data. We introduce parsimony by developing skew-t and skew-normal analogues of the popular GPCM family that employ…

Methodology · Statistics 2013-11-12 Irene Vrbik , Paul D. McNicholas

We introduce a new family of multivariate distributions by taking the component-wise Tukey-h transformation of a random vector following a skew-normal distribution. The proposed distribution is named the skew-normal-Tukey-h distribution and…

Methodology · Statistics 2023-10-19 Sagnik Mondal , Marc G. Genton

The bagplot, also known as the "bag-and-bolster plot", is a notable extension of the boxplot from univariate to bivariate data. Although widely used, its practical application is hindered by two key limitations: the fixed inflation factor…

Methodology · Statistics 2025-12-09 Shenghao Qin , Bowen Gang , Tiejun Tong , Hengjian Cui

With the progress of information technology, large amounts of asymmetric, leptokurtic and heavy-tailed data are arising in various fields, such as finance, engineering, genetics and medicine. It is very challenging to model those kinds of…

Methodology · Statistics 2024-01-26 Chengdi Lian , Yaohua Rong , Weihu Cheng

This article proposes a new class of Real Elliptically Skewed (RESK) distributions and associated clustering algorithms that allow for integrating robustness and skewness into a single unified cluster analysis framework. Non-symmetrically…

Signal Processing · Electrical Eng. & Systems 2021-07-05 Christian A. Schroth , Michael Muma

In 2017-2020 Jordanova and co-authors investigate probabilities for p-outside values and determine them in many particular cases. They show that these probabilities are closely related to the concept for heavy tails. Tukey's boxplots are…

Methodology · Statistics 2024-10-22 Pavlina K. Jordanova

In a world with data that change rapidly and abruptly, it is important to detect those changes accurately. In this paper we describe an R package implementing a generalized version of an algorithm recently proposed by Hocking et al. [2020]…

One aim of data mining is the identification of interesting structures in data. For better analytical results, the basic properties of an empirical distribution, such as skewness and eventual clipping, i.e. hard limits in value ranges, need…

Applications · Statistics 2020-09-08 Michael C. Thrun , Tino Gehlert , Alfred Ultsch

We present results from a preregistered and crowdsourced user study where we asked members of the general population to determine whether two samples represented using different forms of data visualizations are drawn from the same or…

Human-Computer Interaction · Computer Science 2023-09-25 Eric Newburger , Niklas Elmqvist
‹ Prev 1 2 3 10 Next ›