English
Related papers

Related papers: Towards a Taxonomical Consensus: Diversity and Ric…

200 papers

Molecular and morphological characters, as important parts of biological taxonomy, are contradictory but need to be integrated. Organism's image recognition and bioinformatics are emerging and hot problems nowadays but with a gap between…

Computer Vision and Pattern Recognition · Computer Science 2022-06-29 Jiewen Xiao , Wenbin Liao , Ming Zhang , Jing Wang , Jianxin Wang , Yihua Yang

Statistical methods for analyzing large-scale biomolecular data are commonplace in computational biology. A notable example is phenotype prediction from gene expression data, for instance, detecting human cancers, differentiating subtypes…

Genomics · Quantitative Biology 2014-11-24 Bahman Afsari , Ulisses M. Braga-Neto , Donald Geman

Like clustering analysis, community detection aims at assigning nodes in a network into different communities. Fdp is a recently proposed density-based clustering algorithm which does not need the number of clusters as prior input and the…

Social and Information Networks · Computer Science 2016-09-21 Tao You , Ben-Chang Shia , Zhong-Yuan Zhang

Networks often exhibit structure at disparate scales. We propose a method for identifying community structure at different scales based on multiresolution modularity and consensus clustering. Our contribution consists of two parts. First,…

Social and Information Networks · Computer Science 2018-02-01 Lucas G. S. Jeub , Olaf Sporns , Santo Fortunato

We investigate how to select the number of communities for weighted networks without a full likelihood modeling. First, we propose a novel weighted degree-corrected stochastic block model (DCSBM), where the mean adjacency matrix is modeled…

Methodology · Statistics 2025-03-12 Yucheng Liu , Xiaodong Li

A probabilistic reconstruction of genealogies in a polyploid population (from 2x to 4x) is investigated, by considering genetic data analyzed as the probability of allele presence in a given genotype. Based on the likelihood of all possible…

Populations and Evolution · Quantitative Biology 2018-11-29 Frédéric Proïa , Fabien Panloup , Chiraz Trabelsi , Jérémy Clotault

We present a catalog of galaxy cluster masses derived by exploiting the tight correlation between mass and richness, i.e., a properly computed number of bright cluster galaxies. The richness definition adopted in this work is properly…

Cosmology and Nongalactic Astrophysics · Physics 2016-03-23 S. Andreon

We present a new algorithm, CAMIRA, to identify clusters of galaxies in wide-field imaging survey data. We base our algorithm on the stellar population synthesis model to predict colours of red-sequence galaxies at a given redshift for an…

Cosmology and Nongalactic Astrophysics · Physics 2015-06-22 Masamune Oguri

This paper proposes a hierarchical Bayesian multitask learning model that is applicable to the general multi-task binary classification learning problem where the model assumes a shared sparsity structure across different tasks. We derive a…

The mixed membership stochastic blockmodel (MMSB) is a popular Bayesian network model for community detection. Fitting such large Bayesian network models quickly becomes computationally infeasible when the number of nodes grows into…

Social and Information Networks · Computer Science 2024-05-24 Timothy Jones , Owen G. Ward , Yiran Jiang , John Paisley , Tian Zheng

The quality of the inferences we make from pathogen sequence data is determined by the number and composition of pathogen sequences that make up the sample used to drive that inference. However, there remains limited guidance on how to best…

Populations and Evolution · Quantitative Biology 2023-06-13 Lucy D'Agostino McGowan , Shirlee Wohl , Justin Lessler

We introduce an optimization algorithm for resource allocation in the LIPI Public Cluster to optimize its usage according to incoming requests from users. The tool is an extended and modified genetic algorithm developed to match specific…

Distributed, Parallel, and Cluster Computing · Computer Science 2007-08-09 Z. Akbar , L. T. Handoko

Modern biological techniques enable very dense genetic sampling of unfolding evolutionary histories, and thus frequently sample some genotypes multiple times. This motivates strategies to incorporate genotype abundance information in…

Populations and Evolution · Quantitative Biology 2018-04-09 William S. DeWitt , Luka Mesin , Gabriel D. Victora , Vladimir N. Minin , Frederick A. Matsen

Finetuning large language models on instruction data is crucial for enhancing pre-trained knowledge and improving instruction-following capabilities. As instruction datasets proliferate, selecting optimal data for effective training becomes…

Computation and Language · Computer Science 2024-09-18 Simon Yu , Liangyu Chen , Sara Ahmadian , Marzieh Fadaee

High throughput sequencing (HTS)-based technology enables identifying and quantifying non-culturable microbial organisms in all environments. Microbial sequences have enhanced our understanding of the human microbiome, the soil and plant…

Applications · Statistics 2021-03-09 Pratheepa Jeganathan , Susan P. Holmes

Using a sample from a population to estimate the proportion of the population with a certain category label is a broadly important problem. In the context of microbiome studies, this problem arises when researchers wish to use a sample from…

Methodology · Statistics 2019-02-08 Bryan D. Martin , Daniela Witten , Amy D. Willis

Richness estimation of an interesting area is always a challenge statistical work due to small sample size or species identity error. In the literatures, most richness estimators were only proposed to tackle the underestimation of the…

Applications · Statistics 2020-12-18 Jai-Hua Yen , Chun-Huo Chiu

We propose a robust, scalable, integrated methodology for community detection and community comparison in graphs. In our procedure, we first embed a graph into an appropriate Euclidean space to obtain a low-dimensional representation, and…

Machine Learning · Statistics 2016-08-29 Vince Lyzinski , Minh Tang , Avanti Athreya , Youngser Park , Carey E. Priebe

DNA databases are widely used in forensic science to identify unknown offenders. When no exact match is found, familial DNA searches can help by identifying first-degree relatives using likelihood ratios. If multiple subpopulations are…

Applications · Statistics 2025-12-08 Monchai Kooakachai , Tiwakorn Chapalee , Chairat Thitiyan , Patsaya Jumnongwut

The ongoing explosion of genome sequence data is transforming how we reconstruct and understand the histories of biological systems. Across biological scales, from individual cells to populations and species, trees-based models provide a…

Populations and Evolution · Quantitative Biology 2025-12-08 Yun Deng , Shing H. Zhan , Yulin Zhang , Chao Zhang , Bingjie Chen