English
Related papers

Related papers: A method for visual identification of small sample…

200 papers

Parallel coordinate plots (PCPs) are among the most useful techniques for the visualization and exploration of high-dimensional data spaces. They are especially useful for the representation of correlations among the dimensions, which…

Human-Computer Interaction · Computer Science 2016-09-20 Takayuki Itoh , Ashnil Kumar , Karsten Klein , Jinman Kim

In Data Science, entities are typically represented by single valued measurements. Symbolic Data Analysis extends this framework to more complex structures, such as intervals and histograms, that express internal variability. We propose an…

Machine Learning · Statistics 2025-12-16 Diogo Pinheiro , M. Rosário Oliveira , Igor Kravchenko , Lina Oliveira

Visualizations support rapid analysis of scientific datasets, allowing viewers to glean aggregate information (e.g., the mean) within split-seconds. While prior research has explored this ability in conventional charts, it is unclear if…

Human-Computer Interaction · Computer Science 2024-06-21 Victor A. Mateevitsi , Michael E. Papka , Khairi Reda

Subgroup discovery is a descriptive and exploratory data mining technique to identify subgroups in a population that exhibit interesting behavior with respect to a variable of interest. Subgroup discovery has numerous applications in…

Machine Learning · Computer Science 2022-07-19 Ali Arab , Dev Arora , Jialin Lu , Martin Ester

This paper reconsiders the problem of testing the equality of two unspecified continuous distributions. The framework, which we propose, allows for readable and insightful data visualisation and helps to understand and quantify how two…

Methodology · Statistics 2025-03-04 Bogdan Ćmiel , Teresa Ledwina

Similarity-driven multi-view linear reconstruction (SiMLR) is an algorithm that exploits inter-modality relationships to transform large scientific datasets into smaller, more well-powered and interpretable low-dimensional spaces. SiMLR…

Machine Learning · Statistics 2021-01-22 Brian B. Avants , Nicholas J. Tustison , James R. Stone

Classification in the dissimilarity space has become a very active research area since it provides a possibility to learn from data given in the form of pairwise non-metric dissimilarities, which otherwise would be difficult to cope with.…

Despite advances in representation learning, high-dimensional classification remains challenging in low-sample-size regimes, where the dominant signal may vary across applications and labeled data are often limited. We propose a…

Methodology · Statistics 2026-05-18 Xiangbo Mo , Hao Chen

One of the challenges in analyzing high-dimensional expression data is the detection of important biological signals. A common approach is to apply a dimension reduction method, such as principal component analysis. Typically, after…

Quantitative Methods · Quantitative Biology 2012-06-05 Andreas Lehrmann , Michael Huber , Aydin C. Polatkan , Albert Pritzkau , Kay Nieselt

In repeated Measure Designs with multiple groups, the primary purpose is to compare different groups in various aspects. For several reasons, the number of measurements and therefore the dimension of the observation vectors can depend on…

Statistics Theory · Mathematics 2022-07-20 Paavo Sattler , Markus Pauly

Statisticians increasingly face the problem to reconsider the adaptability of classical inference techniques. In particular, divers types of high-dimensional data structures are observed in various research areas; disclosing the boundaries…

Statistics Theory · Mathematics 2017-06-09 Paavo Sattler , Markus Pauly

In biomedical Subgroup Discovery, practitioners are interested in discovering interpretable and homogeneous subgroups within a group of patients. In this paper, assuming that healthy subjects (i.e., controls) share common but irrelevant…

Machine Learning · Computer Science 2026-05-21 Robin Louiset , Edouard Duchesnay , Benoit Dufumier , Antoine Grigis , Pietro Gori

Graphic designers explore large stock image collections during open-ended or early-stage design tasks, yet common tools emphasize relevance and similarity, limiting designers' ability to overview the design space or discover visual…

Human-Computer Interaction · Computer Science 2026-03-10 Antonio Tejero-de-Pablos , Sichao Song , Naoto Ohsaka , Mayu Otani , Shin'ichi Satoh

Distance-based methods involve the computation of distance values between features and are a well-established paradigm in machine learning. In anomaly detection, anomalies are identified by their large distance from normal data points.…

Instrumentation and Methods for Astrophysics · Physics 2025-10-29 Siddharth Chaini , Federica B. Bianco , Ashish Mahabal

Analyzing data subgroups is a common data science task to build intuition about a dataset and identify areas to improve model performance. However, subgroup analysis is prohibitively difficult in datasets with many features, and existing…

Human-Computer Interaction · Computer Science 2025-02-18 Venkatesh Sivaraman , Zexuan Li , Adam Perer

Determining the strength of non-linear statistical dependencies between two variables is a crucial matter in many research fields. The established measure for quantifying such relations is the mutual information. However, estimating mutual…

Data Analysis, Statistics and Probability · Physics 2019-07-24 Damián G. Hernández , Inés Samengo

In this paper we provide new methodology for inference of the geometric features of a multivariate density in deconvolution. Our approach is based on multiscale tests to detect significant directional derivatives of the unknown density at…

Methodology · Statistics 2016-11-21 Konstantin Eckle , Nicolai Bissantz , Holger Dette

Recently, researches related to unsupervised disentanglement learning with deep generative models have gained substantial popularity. However, without introducing supervision, there is no guarantee that the factors of interest can be…

Machine Learning · Computer Science 2020-03-13 Junxiang Chen , Kayhan Batmanghelich

Classification and clustering are both important topics in statistical learning. A natural question herein is whether predefined classes are really different from one another, or whether clusters are really there. Specifically, we may be…

Machine Learning · Statistics 2015-09-22 Qiyi Lu , Xingye Qiao

When data is unlabelled and the target task is not known a priori, divergent search offers a strategy for learning a wide range of skills. Having such a repertoire allows a system to adapt to new, unforeseen tasks. Unlabelled image data is…

Neural and Evolutionary Computing · Computer Science 2020-04-20 Jeremy Tan , Bernhard Kainz