中文
相关论文

相关论文: Dimensionality Reduction Using pseudo-Boolean poly…

200 篇论文

AI-enabled precision medicine promises a transformational improvement in healthcare outcomes by enabling data-driven personalized diagnosis, prognosis, and treatment. However, the well-known "curse of dimensionality" and the clustered…

机器学习 · 计算机科学 2023-05-19 Amanda M. Buch , Conor Liston , Logan Grosenick

Dimension reduction (DR) algorithms have proven to be extremely useful for gaining insight into large-scale high-dimensional datasets, particularly finding clusters in transcriptomic data. The initial phase of these DR methods often…

机器学习 · 计算机科学 2025-10-15 Yingfan Wang , Yiyang Sun , Haiyang Huang , Cynthia Rudin

Principal component analysis continues to be a powerful tool in dimension reduction of high dimensional data. We assume a variance-diverging model and use the high-dimension, low-sample-size asymptotics to show that even though the…

统计理论 · 数学 2020-09-28 Sungkyu Jung

We introduce a novel approach for image edge detection based on pseudo-Boolean polynomials for image patches. We show that patches covering edge regions in the image result in pseudo-Boolean polynomials with higher degrees compared to…

计算机视觉与模式识别 · 计算机科学 2023-08-31 Tendai Mapungwana Chikake , Boris Goldengorin

Identification of disease subtypes and corresponding biomarkers can substantially improve clinical diagnosis and treatment selection. Discovering these subtypes in noisy, high dimensional biomedical data is often impossible for humans and…

定量方法 · 定量生物学 2020-05-18 Marc-Andre Schulz , Matt Chapman-Rounds , Manisha Verma , Danilo Bzdok , Konstantinos Georgatzis

Axis-aligned subspace clustering generally entails searching through enormous numbers of subspaces (feature combinations) and evaluation of cluster quality within each subspace. In this paper, we tackle the problem of identifying subsets of…

机器学习 · 计算机科学 2019-07-17 Ruben Becker , Imane Hafnaoui , Michael E. Houle , Pan Li , Arthur Zimek

OPTICS is a density-based clustering algorithm that performs well in a wide variety of applications. For a set of input objects, the algorithm creates a so-called reachability plot that can be either used to produce cluster membership…

定量方法 · 定量生物学 2013-09-10 Gabor Ivan , Vince Grolmusz

Datasets in high-dimension do not typically form clusters in their original space; the issue is worse when the number of points in the dataset is small. We propose a low-computation method to find statistically significant clustering…

机器学习 · 统计学 2020-08-24 Alden Bradford , Tarun Yellamraju , Mireille Boutin

This paper introduces a novel framework for dynamic classification in high dimensional spaces, addressing the evolving nature of class distributions over time or other index variables. Traditional discriminant analysis techniques are…

统计方法学 · 统计学 2025-02-18 Wenbo Ouyang , Ruiyang Wu , Ning Hao , Hao Helen Zhang

The Johnson-Lindenstrauss transform is a fundamental method for dimension reduction in Euclidean spaces, that can map any dataset of $n$ points into dimension $O(\log n)$ with low distortion of their distances. This dimension bound is tight…

数据结构与算法 · 计算机科学 2026-02-20 Shaofeng H. -C. Jiang , Robert Krauthgamer , Shay Sapir , Sandeep Silwal , Di Yue

One potential solution to combat the scarcity of tail observations in extreme value analysis is to integrate information from multiple datasets sharing similar tail properties, for instance, a common extreme value index. In other words, for…

统计方法学 · 统计学 2025-06-25 Liujun Chen , Marco Oesting , Chen Zhou

Clustering, as an unsupervised technique, plays a pivotal role in various data analysis applications. Among clustering algorithms, Spectral Clustering on Euclidean Spaces has been extensively studied. However, with the rapid evolution of…

机器学习 · 计算机科学 2024-12-09 Sagar Ghosh , Swagatam Das

We consider the problem of clustering grouped data for which the observations may include group-specific variables in addition to the variables that are shared across groups. This type of data is common in cancer genomics where the…

统计方法学 · 统计学 2025-09-30 Arhit Chakrabarti , Yang Ni , Debdeep Pati , Bani K. Mallick

To cluster, classify and represent are three fundamental objectives of learning from high-dimensional data with intrinsic structure. To this end, this paper introduces three interpretable approaches, i.e., segmentation (clustering) via the…

计算机视觉与模式识别 · 计算机科学 2023-06-21 Kai-Liang Lu , Avraham Chapman

While nowadays visual anomaly detection algorithms use deep neural networks to extract salient features from images, the high dimensionality of extracted features makes it difficult to apply those algorithms to large data with 1000s of…

计算机视觉与模式识别 · 计算机科学 2026-03-02 Teng-Yok Lee

In statistical dimensionality reduction, it is common to rely on the assumption that high dimensional data tend to concentrate near a lower dimensional manifold. There is a rich literature on approximating the unknown manifold, and on…

机器学习 · 统计学 2022-02-22 Didong Li , Minerva Mukhopadhyay , David B. Dunson

Deep learning models have become widely adopted in various domains, but their performance heavily relies on a vast amount of data. Datasets often contain a large number of irrelevant or redundant samples, which can lead to computational…

音频与语音处理 · 电气工程与系统科学 2023-09-22 Boris Bergsma , Marta Brzezinska , Oleg V. Yazyev , Milos Cernak

We present Clusterplot, a multi-class high-dimensional data visualization tool designed to visualize cluster-level information offering an intuitive understanding of the cluster inter-relations. Our unique plots leverage 2D blobs devised to…

图形学 · 计算机科学 2021-03-05 Or Malkai , Min Lu , Daniel Cohen-Or

Machine learning-based analysis of medical images often faces several hurdles, such as the lack of training data, the curse of dimensionality problem, and the generalization issues. One of the main difficulties is that there exists…

机器学习 · 计算机科学 2019-07-29 Chang Min Hyun , Kang Cheol Kim , Hyun Cheol Cho , Jae Kyu Choi , Jin Keun Seo

Visualizing high dimensional data by projecting them into two or three dimensional space is one of the most effective ways to intuitively understand the data's underlying characteristics, for example their class neighborhood structure.…

机器学习 · 计算机科学 2020-04-06 Pitoyo Hartono