English
Related papers

Related papers: Boxplots and quartile plots for grouped and period…

200 papers

Whether an extreme observation is an outlier or not, depends strongly on the corresponding tail behaviour of the underlying distribution. We develop an automatic, data-driven method to identify extreme tail behaviour that deviates from the…

Methodology · Statistics 2019-12-06 Shrijita Bhattacharya , Jan Beirlant

Ensemble control offers rich and diverse opportunities in mathematical systems theory. In this paper, we present a new paradigm of ensemble control, referred to as distributional control, for ensemble systems. We shift the focus from…

Optimization and Control · Mathematics 2025-04-08 Jr-Shin Li , Wei Zhang

Consider a panel data setting where repeated observations on individuals are available. Often it is reasonable to assume that there exist groups of individuals that share similar effects of observed characteristics, but the grouping is…

Methodology · Statistics 2024-02-09 Lu Yu , Jiaying Gu , Stanislav Volgushev

Scatterplots are frequently shared across different displays in collaborative and communicative visual analytics. However, variations in displays diversify scatterplot sizes. Such variations can influence the perception of clustering…

Human-Computer Interaction · Computer Science 2024-07-24 Taehyun Yang , Hyeon Jeon , Jinwook Seo

We develop a clustering framework for observations from a population with a smooth probability distribution function and derive its asymptotic properties. A clustering criterion based on a linear combination of order statistics is proposed.…

Statistics Theory · Mathematics 2013-04-16 Karthik Bharath , Vladimir Pozdnyakov , Dipak K Dey

We present a comprehensive insight into counting distributions from the perspective of the combinants extracted from them. In particular, we focus on cases where these combinants exhibit oscillatory behavior that can provide an invaluable…

High Energy Physics - Phenomenology · Physics 2021-05-26 Grzegorz Wilk , Zbigniew Włodarczyk

Clustering large datasets is a fundamental problem with a number of applications in machine learning. Data is often collected on different sites and clustering needs to be performed in a distributed manner with low communication. We would…

Data Structures and Algorithms · Computer Science 2017-02-02 Jiecao Chen , He Sun , David P. Woodruff , Qin Zhang

In multiple correspondence analysis, both individuals (observations) and categories can be represented in a biplot that jointly depicts the relationships across categories or individuals, as well as the associations between them. Additional…

Methodology · Statistics 2019-01-10 Mariko Takagishi , Michel van de Velden

Overplotting of data points is a common problem when visualizing large datasets in a scatterplot, particularly when mapping nominal dimensions to one of the scatterplot axes. Transparency, aggregation, and jittering have previously been…

Human-Computer Interaction · Computer Science 2017-08-29 Deokgun Park , Sung-Hee Kim , Niklas Elmqvist

A diagrammatic method is presented for averaging over the circular ensemble of random-matrix theory. The method is applied to phase-coherent conduction through a chaotic cavity (a ``quantum dot'') and through the interface between a normal…

Condensed Matter · Physics 2007-05-23 P. W. Brouwer , C. W. J. Beenakker

We present a data-driven method to infer the redshift distribution of an arbitrary dataset based on spatial cross-correlation with a reference population and we apply it to various datasets across the electromagnetic spectrum to show its…

Cosmology and Nongalactic Astrophysics · Physics 2014-07-31 Brice Ménard , Ryan Scranton , Samuel Schmidt , Chris Morrison , Donghui Jeong , Tamas Budavari , Mubdi Rahman

Understanding the global organization of complicated and high dimensional data is of primary interest for many branches of applied sciences. It is typically achieved by applying dimensionality reduction techniques mapping the considered…

Computational Geometry · Computer Science 2024-11-11 Paweł Dłotko , Davide Gurnari , Mathis Hallier , Anna Jurek-Loughrey

Cluster interpretation after dimensionality reduction (DR) is a ubiquitous part of exploring multidimensional datasets. DR results are frequently represented by scatterplots, where spatial proximity encodes similarity among data samples. In…

Human-Computer Interaction · Computer Science 2021-08-18 Wilson E. Marcílio-Jr , Danilo M. Eler , Rogério E. Garcia

Tiny fluctuations of the Cosmic Microwave Background as well as various observable quantities obtained by spin raising and spin lowering of the effective gravitational lensing potential of distant galaxies and galaxy clusters, are described…

Probability · Mathematics 2018-05-04 Anatoliy Malyarenko

The angular measure on the unit sphere characterizes the first-order dependence structure of the components of a random vector in extreme regions and is defined in terms of standardized margins. Its statistical recovery is an important step…

Statistics Theory · Mathematics 2022-10-18 Stéphan Clémençon , Hamid Jalalzai , Stéphane Lhaut , Anne Sabourin , Johan Segers

For an ensemble of data points in a multi-parameter space, we present a visual analytics technique to select a representative distribution of parameter values, and analyse how representative this distribution is in all ensemble members. A…

Human-Computer Interaction · Computer Science 2020-07-31 Alexander Kumpf , Josef Stumpfegger , Patrick Fabian Härtl , Rüdiger Westermann

The focus of this paper is on the quantification of sampling variation in frequentist probabilistic forecasts. We propose a method of constructing confidence sets that respects the functional nature of the forecast distribution, and use…

Methodology · Statistics 2017-08-09 David Harris , Gael M. Martin , Indeewara Perera , D. S. Poskitt

One key use of k-means clustering is to identify cluster prototypes which can serve as representative points for a dataset. However, a drawback of using k-means cluster centers as representative points is that such points distort the…

Machine Learning · Statistics 2019-11-15 Arvind Krishna , Simon Mak , Roshan Joseph

Density-based clustering methodology has been widely considered in the statistical literature for classifying Euclidean observations. However, this approach has not been contemplated for directional data yet. In this work, directional…

Methodology · Statistics 2023-03-07 Paula Saavedra-Nieves , Martín Fernández-Pérez

We propose a model-based clustering algorithm for a general class of functional data for which the components could be curves or images. The random functional data realizations could be measured with error at discrete, and possibly random,…

Machine Learning · Statistics 2022-03-14 Steven Golovkine , Nicolas Klutchnikoff , Valentin Patilea