中文
相关论文

相关论文: RLE Plots: Visualising Unwanted Variation in High …

200 篇论文

Simulated high-dimensional data is useful for testing, validating, and improving algorithms used in dimension reduction, supervised and unsupervised learning. High-dimensional data is characterized by multiple variables that are dependent…

统计方法学 · 统计学 2025-12-23 Jayani P. Gamage , Dianne Cook , Paul Harrison , Michael Lydeamore , Thiyanga S. Talagala

The central aim in this paper is to address variable selection questions in nonlinear and nonparametric regression. Motivated by statistical genetics, where nonlinear interactions are of particular interest, we introduce a novel and…

统计方法学 · 统计学 2018-08-28 Lorin Crawford , Seth R. Flaxman , Daniel E. Runcie , Mike West

In this work we present a visualization tool specifically tailored to deal with skewed data. The technique is based upon the use of two types of notched boxplots (the usual one, and one which is tuned for the skewness of the data), the…

统计计算 · 统计学 2014-03-04 R. Ospina , A. M. Larangeiras , A. C. Frery

Automated data extraction from research texts has been steadily improving, with the emergence of large language models (LLMs) accelerating progress even further. Extracting data from plots in research papers, however, has been such a…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Maciej P. Polak , Dane Morgan

Logistic regression is commonly used for modeling dichotomous outcomes. In the classical setting, where the number of observations is much larger than the number of parameters, properties of the maximum likelihood estimator in logistic…

机器学习 · 统计学 2019-11-14 Fariborz Salehi , Ehsan Abbasi , Babak Hassibi

Gene expression analysis aims at identifying the genes able to accurately predict biological parameters like, for example, disease subtyping or progression. While accurate prediction can be achieved by means of many different techniques,…

统计方法学 · 统计学 2008-09-11 Christine De Mol , Sofia Mosci , Magali Traskine , Alessandro Verri

This study introduces a new method of visualizing complex tree structured objects. The usefulness of this method is illustrated in the context of detecting unexpected features in a data set of very large trees. The major contribution is a…

应用统计 · 统计学 2012-02-14 Burcu Aydin , Gabor Pataki , Haonan Wang , Alim Ladha , Elizabeth Bullitt , J. S. Marron

General log-linear models specified by non-negative integer design matrices have a potentially wide range of applications, although using models without the genuine overall effect, that is, ones which cannot be reparameterized to include a…

统计方法学 · 统计学 2023-01-02 Anna Klimova , Matthias Kuhn

Principal component analysis is a useful dimension reduction and data visualization method. However, in high dimension, low sample size asymptotic contexts, where the sample size is fixed and the dimension goes to infinity,a paradox has…

应用统计 · 统计学 2012-11-21 Dan Shen , Haipeng Shen , Hongtu Zhu , J. S. Marron

In this time when biased information, deep fakes, and propaganda proliferate, the accessibility of reliable data sources is more important than ever. National statistical institutes provide curated data that contain quantitative information…

信息检索 · 计算机科学 2025-10-03 Gadir Suleymanli , Alexander Rogiers , Lucas Lageweg , Jefrey Lijffijt

One of the challenges in analyzing high-dimensional expression data is the detection of important biological signals. A common approach is to apply a dimension reduction method, such as principal component analysis. Typically, after…

定量方法 · 定量生物学 2012-06-05 Andreas Lehrmann , Michael Huber , Aydin C. Polatkan , Albert Pritzkau , Kay Nieselt

As regression is a widely studied problem, many methods have been proposed to solve it, each of them often requiring setting different hyper-parameters. Therefore, selecting the proper method for a given application may be very difficult…

机器学习 · 计算机科学 2026-03-23 Nassime Mountasir , Baptiste Lafabregue , Bruno Albert , Nicolas Lachiche

As the complexity and volume of datasets have increased along with the capabilities of modular, open-source, easy-to-implement, visualization tools, scientists' need for, and appreciation of, data visualization has risen too. Until…

天体物理仪器与方法 · 物理学 2018-05-30 Alyssa A. Goodman , Michelle A. Borkin , Thomas P. Robitaille

We present Clusterplot, a multi-class high-dimensional data visualization tool designed to visualize cluster-level information offering an intuitive understanding of the cluster inter-relations. Our unique plots leverage 2D blobs devised to…

图形学 · 计算机科学 2021-03-05 Or Malkai , Min Lu , Daniel Cohen-Or

Developing Machine Learning (ML) algorithms for heterogeneous/mixed data is a longstanding problem. Many ML algorithms are not applicable to mixed data, which include numeric and non-numeric data, text, graphs and so on to generate…

机器学习 · 计算机科学 2022-06-15 Boris Kovalerchuk , Elijah McCoy

The ability to detect log anomalies from system logs is a vital activity needed to ensure cyber resiliency of systems. It is applied for fault identification or facilitate cyber investigation and digital forensics. However, as logs…

密码学与安全 · 计算机科学 2023-11-10 Jonathan Pan , Swee Liang Wong , Yidi Yuan

Visualization of high-dimensional data is counter-intuitive using conventional graphs. Parallel coordinates are proposed as an alternative to explore multivariate data more effectively. However, it is difficult to extract relevant…

统计计算 · 统计学 2019-05-27 Shaima Tilouche , Vahid Partovi Nia , Samuel Bassetto

This paper investigates the effectiveness of using the Random Projection Ensemble (RPE) approach in Quadratic Discriminant Analysis (QDA) for ultrahigh-dimensional classification problems. Classical methods such as Linear Discriminant…

统计方法学 · 统计学 2025-07-10 Annesha Deb , Minerva Mukhopadhyay , Subhajit Dutta

Undirected graphs are often used to describe high dimensional distributions. Under sparsity conditions, the graph can be estimated using $\ell_1$ penalization methods. However, current methods assume that the data are independent and…

机器学习 · 统计学 2008-04-29 Shuheng Zhou , John Lafferty , Larry Wasserman

We study the problem of visualizing large-scale and high-dimensional data in a low-dimensional (typically 2D or 3D) space. Much success has been reported recently by techniques that first compute a similarity structure of the data points…

机器学习 · 计算机科学 2016-04-06 Jian Tang , Jingzhou Liu , Ming Zhang , Qiaozhu Mei