中文
相关论文

相关论文: PCA and K-Means decipher genome

200 篇论文

When working with large biological data sets, exploratory analysis is an important first step for understanding the latent structure and for generating hypotheses to be tested in subsequent analyses. However, when the number of variables is…

统计方法学 · 统计学 2017-02-03 Julia Fukuyama

Principal component analysis (PCA) is commonly used in genetics to infer and visualize population structure and admixture between populations. PCA is often interpreted in a way similar to inferred admixture proportions, where it is assumed…

统计方法学 · 统计学 2023-02-10 Jan van Waaij , Song Li , Genís Garcia-Erill , Anders Albrechtsen , Carsten Wiuf

Relation of genome sizes to organisms complexity is still described rather equivocally. Neither the number of genes (G-value), nor the total amount of DNA (C-value) correlates consistently with phenotype complexity. Using information theory…

基因组学 · 定量生物学 2007-05-23 Dmitri V. Parkhomchuk

Network models provide a powerful framework for analysing single-cell count data, facilitating the characterisation of cellular identities, disease mechanisms, and developmental trajectories. However, uncertainty modeling in unsupervised…

基因组学 · 定量生物学 2026-04-27 Shanshan Ren , Thomas E. Bartlett , Lina Gerontogianni , Swati Chandna

This paper presents a probabilistic approach for DNA sequence analysis. A DNA sequence consists of an arrangement of the four nucleotides A, C, T and G and different representation schemes are presented according to a probability measure…

定量方法 · 定量生物学 2010-02-12 Amrita Priyam , B. M. Karan , G. Sahoo

Probabilistic graphical models (PGMs) have become a popular tool for computational analysis of biological data in a variety of domains. But, what exactly are they and how do they work? How can we use PGMs to discover patterns that are…

定量方法 · 定量生物学 2010-02-22 Edoardo M Airoldi

We describe how cross-kernel matrices, that is, kernel matrices between the data and a custom chosen set of `feature spanning points' can be used for learning. The main potential of cross-kernels lies in the fact that (a) only one side of…

机器学习 · 计算机科学 2014-06-11 Franz J Király , Martin Kreuzer , Louis Theran

While many good textbooks are available on Protein Structure, Molecular Simulations, Thermodynamics and Bioinformatics methods in general, there is no good introductory level book for the field of Structural Bioinformatics. This book aims…

In special coordinates (codon position--specific nucleotide frequencies) bacterial genomes form two straight lines in 9-dimensional space: one line for eubacterial genomes, another for archaeal genomes. All the 348 distinct bacterial…

基因组学 · 定量生物学 2007-11-13 A. N. Gorban , A. Yu. Zinovyev

The advent of digital pathology presents opportunities for computer vision for fast, accurate, and objective solutions for histopathological images and aid in knowledge discovery. This work uses deep learning to predict genomic biomarkers -…

计算机视觉与模式识别 · 计算机科学 2021-03-16 Ruchi Chauhan , PK Vinod , CV Jawahar

Data integration, or the strategic analysis of multiple sources of data simultaneously, can often lead to discoveries that may be hidden in individualistic analyses of a single data source. We develop a new unsupervised data integration…

统计方法学 · 统计学 2021-04-06 Tiffany M. Tang , Genevera I. Allen

The objectives of this paper are to explore ways to analyze breast cancer dataset in the context of unsupervised learning without prior training model. The paper investigates different ways of clustering techniques as well as preprocessing.…

机器学习 · 计算机科学 2021-09-06 Somenath Chakraborty , Beddhu Murali

Recent developed deep unsupervised methods allow us to jointly learn representation and cluster unlabelled data. These deep clustering methods mainly focus on the correlation among samples, e.g., selecting high precision pairs to gradually…

计算机视觉与模式识别 · 计算机科学 2019-08-13 Jianlong Wu , Keyu Long , Fei Wang , Chen Qian , Cheng Li , Zhouchen Lin , Hongbin Zha

The properties of certain networks are determined by hidden variables that are not explicitly measured. The conditional probability (propagator) that a vertex with a given value of the hidden variable is connected to k of other vertices…

定量方法 · 定量生物学 2009-11-13 Gerald A. Miller , Yi Y. Shi , Hong Qian , Karol Bomsztyk

Clustering large amount of data is becoming increasingly important in the current times. Due to the large sizes of data, clustering algorithm often take too much time. Sampling this data before clustering is commonly used to reduce this…

机器学习 · 计算机科学 2021-08-24 Seemandhar Jain , Aditya A. Shastri , Kapil Ahuja , Yann Busnel , Navneet Pratap Singh

Principal component analysis (PCA) is by far the most widespread tool for unsupervised learning with high-dimensional data sets. Its application is popularly studied for the purpose of exploratory data analysis and online process…

应用统计 · 统计学 2019-02-12 Stefania Russo , Guangyu Li , Kris Villez

Principal component analysis (PCA) aims at estimating the direction of maximal variability of a high-dimensional dataset. A natural question is: does this task become easier, and estimation more accurate, when we exploit additional…

信息论 · 计算机科学 2014-06-19 Andrea Montanari , Emile Richard

Inference in clustering is paramount to uncovering inherent group structure in data. Clustering methods which assess statistical significance have recently drawn attention owing to their importance for the identification of patterns in high…

统计方法学 · 统计学 2021-06-18 Debora Zava Bello , Marcio Valk , Gabriela Bettella Cybis

Recent studies on computer vision mainly focus on natural images that express real-world scenes. They achieve outstanding performance on diverse tasks such as visual question answering. Diagram is a special form of visual expression that…

计算机视觉与模式识别 · 计算机科学 2021-03-11 Shaowei Wang , LingLing Zhang , Xuan Luo , Yi Yang , Xin Hu , Jun Liu

We present *K-means clustering algorithm and source code by expanding statistical clustering methods applied in https://ssrn.com/abstract=2802753 to quantitative finance. *K-means is statistically deterministic without specifying initial…

基因组学 · 定量生物学 2017-10-05 Zura Kakushadze , Willie Yu