English
Related papers

Related papers: Multivariate Analysis of Mixed Data: The R Package…

200 papers

Urban analytics utilizes extensive datasets with diverse urban information to simulate, predict trends, and uncover complex patterns within cities. While these data enables advanced analysis, it also presents challenges due to its…

Machine Learning · Computer Science 2025-09-09 Ximena Pocco , Waqar Hassan , Karelia Salinas , Vladimir Molchanov , Luis G. Nonato

Process data refer to data recorded in the log files of computer-based items. These data, represented as timestamped action sequences, keep track of respondents' response processes of solving the items. Process data analysis aims at…

Computation · Statistics 2020-06-11 Xueying Tang , Susu Zhang , Zhi Wang , Jingchen Liu , Zhiliang Ying

Researchers in the behavioral and social sciences use linear discriminant analysis (LDA) for predictions of group membership (classification) and for identifying the variables most relevant to group separation among a set of continuous…

Methodology · Statistics 2025-05-28 Ricarda Graf , Marina Zeldovich , Sarah Friedrich

Visual Analytics (VA) tools and techniques have been instrumental in supporting users to build better classification models, interpret models' overall logic, and audit results. In a different direction, VA has recently been applied to…

Machine Learning · Computer Science 2022-11-21 Mário Popolin Neto , Fernando V. Paulovich

Image data are increasingly encountered and are of growing importance in many areas of science. Much of these data are quantitative image data, which are characterized by intensities that represent some measurement of interest in the…

ProfileGLMM is an R package integrating Generalised Linear Mixed Models (GLMMs) as the outcome model for Bayesian profile regression. This statistical framework simultaneously i) explains the variation in the outcome and ii) clusters the…

Methodology · Statistics 2026-04-23 Matteo Amestoy , Mark A. van de Wiel , Wessel N. van Wieringen

The growing volume of data usually creates an interesting challenge for the need of data analysis tools that discover regularities in these data. Data mining has emerged as disciplines that contribute tools for data analysis, discovery of…

Databases · Computer Science 2011-08-30 Abhishek Taneja , R. K. Chauhan

Histograms provide a powerful means of summarizing large data sets by representing their distribution in a compact, binned form. The HistogramTools R package enhances R built-in histogram functionality, offering advanced methods for…

Databases · Computer Science 2025-04-02 Shubham Malhotra

Multivariate data are typically represented by a rectangular matrix (table) in which the rows are the objects (cases) and the columns are the variables (measurements). When there are many variables one often reduces the dimension by…

Methodology · Statistics 2021-01-13 Mia Hubert , Peter J. Rousseeuw , Wannes Van den Bossche

In many contexts, missing data and disclosure control are ubiquitous and challenging issues. In particular at statistical agencies, the respondent-level data they collect from surveys and censuses can suffer from high rates of missingness.…

Computation · Statistics 2021-09-06 Jingchen Hu , Olanrewaju Akande , Quanli Wang

In the future, competitive advantages will be given to organisations that can extract valuable information from massive data and make better decisions. In most cases, this data comes from multiple sources. Therefore, the challenge is to…

Applications · Statistics 2016-05-11 Igor Barahona , Judith Cavazos , Jian-Bo Yang

High dimensional data can contain multiple scales of variance. Analysis tools that preferentially operate at one scale can be ineffective at capturing all the information present in this cross-scale complexity. We propose a multiscale joint…

Machine Learning · Statistics 2022-03-15 Daniel Sousa , Christopher Small

Multivariate statistical methods are widely used throughout the sciences, including microscopy, however, their utilisation for analysis of electron backscatter diffraction (EBSD) data has not been adequately explored. The basic aim of most…

With the continuous increase in the computational power and resources of modern high-performance computing (HPC) systems, large-scale ensemble simulations have become widely used in various fields of science and engineering, and especially…

Human-Computer Interaction · Computer Science 2022-09-23 Keita Watanabe , Naohisa Sakamoto , Jorji Nonaka , Yasumitsu Maejima

MSmix is a recently developed R package implementing maximum likelihood estimation of finite mixtures of Mallows models with Spearman distance for full and partial rankings. The package is designed to implement computationally tractable…

Methodology · Statistics 2024-06-24 Marta Crispino , Cristina Mollica , Lucia Modugno

In this work we present a visualization tool specifically tailored to deal with skewed data. The technique is based upon the use of two types of notched boxplots (the usual one, and one which is tuned for the skewness of the data), the…

Computation · Statistics 2014-03-04 R. Ospina , A. M. Larangeiras , A. C. Frery

In fields such as ecology, microbiology, and genomics, non-Euclidean distances are widely applied to describe pairwise dissimilarity between samples. Given these pairwise distances, principal coordinates analysis (PCoA) is commonly used to…

Quantitative Methods · Quantitative Biology 2020-03-24 Yushu Shi , Liangliang Zhang , Kim-Anh Do , Christine Peterson , Robert Jenq

Machine learning (ML) methods have proved to be a very successful tool in physical sciences, especially when applied to experimental data analysis. Artificial intelligence is particularly good at recognizing patterns in high dimensional…

Materials Science · Physics 2022-08-19 T. Tula , G. Möller , J. Quintanilla , S. R. Giblin , A. D. Hillier , E. E. McCabe , S. Ramos , D. S. Barker , S. Gibson

We illustrate the use of the R-package "rstiefel" for matrix-variate data analysis in the context of two examples. The first example considers estimation of a reduced-rank mean matrix in the presence of normally distributed noise. The…

Computation · Statistics 2013-04-15 Peter D. Hoff

Incomplete data is a persistent challenge in real-world datasets, often governed by complex and unobservable missing mechanisms. Simulating missingness has become a standard approach for understanding its impact on learning and analysis.…

Machine Learning · Computer Science 2025-08-08 Youran Zhou , Mohamed Reda Bouadjenek , Sunil Aryal