中文
相关论文

相关论文: PCA and K-Means decipher genome

200 篇论文

Principal Component Analysis (PCA) is a fundamental data preprocessing tool in the world of machine learning. While PCA is often thought of as a dimensionality reduction method, the purpose of PCA is actually two-fold: dimension reduction…

机器学习 · 计算机科学 2023-01-25 Arpita Gang , Waheed U. Bajwa

Bi-clustering is a useful approach in analyzing biological data when observations come from heterogeneous groups and have a large number of features. We outline a general Bayesian approach in tackling bi-clustering problems in moderate to…

应用统计 · 统计学 2021-02-11 Han Yan , Jiexing Wu , Yang Li , Jun S. Liu

In this paper, we present a new approach for decomposing scan paths and its utility for generating new scan paths. For this purpose, we use the K-Means clustering procedure to the raw gaze data and subsequently iteratively to find more…

图形学 · 计算机科学 2022-01-21 Wolfgang Fuhl

Clustering analysis plays an important role in scientific research and commercial application. K-means algorithm is a widely used partition method in clustering. However, it is known that the K-means algorithm may get stuck at suboptimal…

神经与进化计算 · 计算机科学 2014-05-26 M. H. Marghny , Rasha M. Abd El-Aziz , Ahmed I. Taloba

We describe a robust, fast, and memory-efficient procedure that can cluster millions of structures derived from molecular dynamics simulations. The essence of the method is based on a peak-picking algorithm applied to three- and…

生物大分子 · 定量生物学 2015-12-15 Athanasios S. Baltzis , Panagiotis I. Koukos , Nicholas M. Glykos

Rearrangements of bacterial chromosomes can be studied mathematically at several levels, most prominently at a local, or sequence level, as well as at a topological level. The biological changes involved locally are inversions, deletions,…

群论 · 数学 2013-12-10 Andrew R. Francis

Microarrays are made it possible to simultaneously monitor the expression profiles of thousands of genes under various experimental conditions. Identification of co-expressed genes and coherent patterns is the central goal in microarray or…

计算工程、金融与科学 · 计算机科学 2013-07-15 T. Chandrasekhar , K. Thangavel , E. Elayaraja , E. N. Sathishkumar

Unsupervised clustering algorithms for vectors has been widely used in the area of machine learning. Many applications, including the biological data we studied in this paper, contain some boundary datapoints which show combination…

机器学习 · 计算机科学 2022-05-23 Yingcong Li , Chandra Sekhar Mukherjee , Jiapeng Zhang

For open vocabulary recognition of ingredients in food images, segmenting the ingredients is a crucial step. This paper proposes a novel approach that explores PCA-based feature representations of image pixels using a convolutional neural…

计算机视觉与模式识别 · 计算机科学 2025-10-10 Ying Dai

With increasing urbanization, transportation plays an increasingly critical role in city development. The number of studies on modeling, optimization, simulation, and data analysis of transportation systems is on the rise. Many of these…

机器学习 · 计算机科学 2023-11-27 Sina Sabzekar , Mohammad Reza Valipour Malakshah , Zahra Amini

Principal components analysis (PCA) is a widely used dimension reduction technique with an extensive range of applications. In this paper, an online distributed algorithm is proposed for recovering the principal eigenspaces. We further…

Principal Component Analysis (PCA) is a well-known multivariate technique used to decorrelate a set of vectors. PCA has been extensively applied in the past to the classification of stellar and galaxy spectra. Here we apply PCA to the…

天体物理学 · 物理学 2007-05-23 I. Ferreras , B. Rogers , O. Lahav , .

We propose kernel PCA as a method for analyzing the dependence structure of multivariate extremes and demonstrate that it can be a powerful tool for clustering and dimension reduction. Our work provides some theoretical insight into the…

机器学习 · 统计学 2022-11-28 Marco Avella-Medina , Richard A. Davis , Gennady Samorodnitsky

Detecting variation in the evolutionary process along chromosomes is increasingly important as whole-genome data becomes more widely available. For example, factors such as incomplete lineage sorting, horizontal gene transfer, and…

种群与进化 · 定量生物学 2017-01-03 Elizabeth S. Allman , Laura S. Kubatko , John A. Rhodes

Bioinformatics, which is now a well known field of study, originated in the context of biological sequence analysis. Recently graphical representation takes place for the research on DNA sequence. Research in biological sequence is mainly…

数据结构与算法 · 计算机科学 2022-11-28 Probir Mondal

In several recent papers new gene-detection algorithms were proposed for detecting protein-coding regions without requiring learning dataset of already known genes. The fact that unsupervised gene-detection is possible closely connected to…

无序系统与神经网络 · 物理学 2007-05-23 A. N. Gorban , A. Yu. Zinovyev , T. G. Popova

Computational methods for discovering patterns of local correlations in sequences are important in computational biology. Here we show how to determine the optimal partitioning of aligned sequences into non-overlapping segments such that…

计算工程、金融与科学 · 计算机科学 2012-06-26 Joseph Bockhorst , Nebojsa Jojic

Cognitive diagnosis models (CDMs) are a popular tool for assessing students' mastery of sets of skills. Given a set of $K$ skills tested on an assessment, students are classified into one of $2^K$ latent skill set profiles that represent…

应用统计 · 统计学 2021-04-07 Alan Mishler , Rebecca Nugent

The $k$-means is one of the most important unsupervised learning techniques in statistics and computer science. The goal is to partition a data set into many clusters, such that observations within clusters are the most homogeneous and…

机器学习 · 统计学 2022-11-21 Tonglin Zhang

In this paper, we propose and study a Nystr\"om based approach to efficient large scale kernel principal component analysis (PCA). The latter is a natural nonlinear extension of classical PCA based on considering a nonlinear feature map or…

机器学习 · 统计学 2019-07-12 Nicholas Sterge , Bharath Sriperumbudur , Lorenzo Rosasco , Alessandro Rudi