English
Related papers

Related papers: Using Genetic Data to Build Intuition about Popula…

200 papers

Phylogenetics uses alignments of molecular sequence data to learn about evolutionary trees relating species. Along branches, sequence evolution is modelled using a continuous-time Markov process characterised by an instantaneous rate…

In the era of big data, analysts usually explore various statistical models or machine learning methods for observed data in order to facilitate scientific discoveries or gain predictive power. Whatever data and fitting procedures are…

Machine Learning · Statistics 2018-10-24 Jie Ding , Vahid Tarokh , Yuhong Yang

Genetic data collection has become ubiquitous, producing genetic information about health, ancestry, and social traits. However, unregulated use, especially amid evolving scientific understanding, poses serious privacy and discrimination…

Computers and Society · Computer Science 2025-06-03 Vivek Ramanan , Ria Vinod , Cole Williams , Sohini Ramachandran , Suresh Venkatasubramanian

Genealogical networks, also known as family trees or population pedigrees, are commonly studied by genealogists wanting to know about their ancestry, but they also provide a valuable resource for disciplines such as digital demography,…

Social and Information Networks · Computer Science 2018-02-19 Eric Malmi , Aristides Gionis , Arno Solin

The evolutionary dynamics of molecular populations are strongly dependent on the structure of genotype spaces. The map between genotype and phenotype determines how easily genotype spaces can be navigated and the accessibility of…

Populations and Evolution · Quantitative Biology 2019-07-03 Juan Antonio García-Martín , Pablo Catalán , Susanna Manrubia , José A. Cuesta

Modern population genetics studies typically involve genome-wide genotyping of individuals from a diverse network of ancestries. An important, unsolved problem is how to formulate and estimate probabilistic models of observed genotypes that…

Populations and Evolution · Quantitative Biology 2017-01-10 Wei Hao , Minsun Song , John D. Storey

Machine learning and deep learning have been celebrating many successes in the application to biological problems, especially in the domain of protein folding. Another equally complex and important question has received relatively little…

Machine Learning · Computer Science 2023-10-09 Lucie Bourguignon , Caroline Weis , Catherine R. Jutzeler , Michael Adamer

Statistical analysis of DNA mixtures is known to pose computational challenges due to the enormous state space of possible DNA profiles. We propose a Bayesian network representation for genotypes, allowing computations to be performed…

Methodology · Statistics 2014-02-21 Therese Graversen , Steffen Lauritzen

Gene set analysis, a popular approach for analysing high-throughput gene expression data, aims to identify sets of genes that show enriched expression patterns between two conditions. In addition to the multitude of methods available for…

Probabilistic graphical models (PGMs) have become a popular tool for computational analysis of biological data in a variety of domains. But, what exactly are they and how do they work? How can we use PGMs to discover patterns that are…

Quantitative Methods · Quantitative Biology 2010-02-22 Edoardo M Airoldi

Complex analyses involving multiple, dependent random quantities often lead to graphical models - a set of nodes denoting variables of interest, and corresponding edges denoting statistical interactions between nodes. To develop statistical…

Computer Vision and Pattern Recognition · Computer Science 2021-04-05 Xiaoyang Guo , Anuj Srivastava , Sudeep Sarkar

The language commonly used in human genetics can inadvertently pose problems for multiple reasons. Terms like "ancestry", "ethnicity", and other ways of grouping people can have complex, often poorly understood, or multiple meanings within…

Populations and Evolution · Quantitative Biology 2021-06-21 Ewan Birney , Michael Inouye , Jennifer Raff , Adam Rutherford , Aylwyn Scally

Exploratory data analysis is a fundamental aspect of knowledge discovery that aims to find the main characteristics of a dataset. Dimensionality reduction, such as manifold learning, is often used to reduce the number of features in a…

Neural and Evolutionary Computing · Computer Science 2019-10-24 Andrew Lensen , Bing Xue , Mengjie Zhang

Applying the concepts and formalisms from Evolutionary Game Theory to the data regime, the fundamental paradigms of Evolutionary Data Theory are introduced. Interpreting data in matrix form as evolutionary entities, input data is mapped to…

Neural and Evolutionary Computing · Computer Science 2026-05-27 Philipp Wissgott

Genotype networks are a method used in systems biology to study the "innovability" of a set of genotypes having the same phenotype. In the past they have been applied to determine the genetic heterogeneity, and stability to mutations, of…

Populations and Evolution · Quantitative Biology 2015-06-17 Giovanni Marco Dall'Olio , Jaume Bertranpetit , Andreas Wagner , Hafid Laayouni

Frequentist statistical methods, such as hypothesis testing, are standard practice in papers that provide benchmark comparisons. Unfortunately, these methods have often been misused, e.g., without testing for their statistical test…

Methodology · Statistics 2021-05-18 David Issa Mattos , Jan Bosch , Helena Holmström Olsson

The increasing availability of high throughput data arising from gene expression studies leads to the necessity of methods for summarizing the available information. As annotation quality improves it is becoming common to rely on the Gene…

Genomics · Quantitative Biology 2007-05-23 Alex Sanchez-Pla , Miquel Salicru , Jordi Ocanya

The amount of sequence data obtained from ancient samples has dramatically expanded in the last decade, and so have the types of questions that can now be addressed using ancient DNA. In the field of human history, while ancient DNA has…

Populations and Evolution · Quantitative Biology 2020-01-08 Fernando Racimo , Martin Sikora , Hannes Schroeder , Carles Lalueza-Fox

Genome-wide association studies (GWAS) have identified hundreds of loci at very stringent levels of statistical significance across many different human traits. However, it is now clear that very large samples (n~10^4-10^5) are needed to…

Genomics · Quantitative Biology 2013-08-20 Inti Pedroso

Increasingly used high throughput experimental techniques, like DNA or protein microarrays give as a result groups of interesting, e.g. differentially regulated genes which require further biological interpretation. With the systematic…

Genomics · Quantitative Biology 2007-05-23 Nils Blüthgen , Karsten Brand , Branka Čajavec , Maciej Swat , Hanspeter Herzel , Dieter Beule