English
Related papers

Related papers: YHap: software for probabilistic assignment of Y h…

200 papers

Low-coverage short-read resequencing experiments have the potential to expand our understanding of Y chromosome haplogroups. However, the uncertainty associated with these experiments mean that haplogroups must be assigned probabilistically…

Populations and Evolution · Quantitative Biology 2014-07-31 Luke Jostins , Yali Xu , Shane McCarthy , Qasim Ayub , Richard Durbin , Jeff Barrett , Chris Tyler-Smith

The computational problem of inferring the full haplotype of a cell starting from read sequencing data is known as haplotype assembly, and consists in assigning all heterozygous Single Nucleotide Polymorphisms (SNPs) to exactly one of the…

Computing haplotypes from sequencing data, i.e. haplotype assembly, is an important component of molecular and population genetics problems, including interpreting the effects of genetic variation on complex traits and reconstructing…

Genomics · Quantitative Biology 2026-03-12 Marjan Hosseini , Ella Veiner , Thomas Bergendahl , Tala Yasenpoor , Zane Smith , Margaret Staton , Derek Aguiar

Probabilistic Answer Set Programming under the credal semantics (PASP) extends Answer Set Programming with probabilistic facts that represent uncertain information. The probabilistic facts are discrete with Bernoulli distributions. However,…

Artificial Intelligence · Computer Science 2025-02-19 Damiano Azzolini , Fabrizio Riguzzi

Genetic variation in human populations is influenced by geographic ancestry due to spatial locality in historical mating and migration patterns. Spatial population structure in genetic datasets has been traditionally analyzed using either…

Populations and Evolution · Quantitative Biology 2016-10-26 Anand Bhaskar , Adel Javanmard , Thomas A. Courtade , David Tse

Analyzing a functional genomics experiment, such as ATAC-, ChIP- or RNA-sequencing, requires reference data including a genome assembly and gene annotation. These resources can generally be retrieved from different organizations and in…

Genomics · Quantitative Biology 2022-09-05 Siebren Frölich , Maarten van der Sande , Tilman Schäfers , Simon J. van Heeringen

Feature selection remains a major challenge in medical prediction, where existing approaches such as LASSO often lack robustness and interpretability. We introduce GRASP, a novel framework that couples Shapley value driven attribution with…

Machine Learning · Computer Science 2026-05-01 Yuheng Luo , Shuyan Li , Zhong Cao

There is little debate about the importance of the ancestral recombination graph in population genetics. An important theoretical tool, the main obstacle to its widespread usage is the computational cost required to match the…

Populations and Evolution · Quantitative Biology 2026-05-14 Patrick Fournier , Fabrice Larribe

Machine learning models often perform poorly under subpopulation shifts in the data distribution. Developing methods that allow machine learning models to better generalize to such shifts is crucial for safe deployment in real-world…

Machine Learning · Statistics 2024-03-18 Tim G. J. Rudner , Ya Shi Zhang , Andrew Gordon Wilson , Julia Kempe

We present an application, EasyScan_HEP, for connecting programs to scan the parameter space of High Energy Physics (HEP) models using various sampling algorithms. We develop EasyScan_HEP according to the principle of flexibility and…

High Energy Physics - Phenomenology · Physics 2023-12-04 Liangliang Shang , Yang Zhang

Genetic data obtained on population samples convey information about their evolutionary history. Inference methods can extract this information (at least partially) but they require sophisticated statistical techniques that have been made…

The detection of molecular signatures of selection is one of the major concerns of modern population genetics. A widely used strategy in this context is to compare samples from several populations, and to look for genomic regions with…

Populations and Evolution · Quantitative Biology 2013-01-24 Marìa Inès Fariello , Simon Boitard , Hugo Naya , Magali SanCristobal , Bertrand Servin

Clusters of genes that have evolved by repeated segmental duplication present difficult challenges throughout genomic analysis, from sequence assembly to functional analysis. Improved understanding of these clusters is of utmost importance,…

Machine Learning · Computer Science 2010-01-25 Tomáš Vinař , Broňa Brejová , Giltae Song , Adam Siepel

Processing high-throughput DNA sequencing data of individuals or populations requires stringing together independent software tools with many parameters, often leading to non-reproducible pipelines and datasets. We developed grenepipe to…

Genomics · Quantitative Biology 2025-01-09 Lucas Czech , Moises Exposito-Alonso

The mathematical software \texttt{GAP} (Groups, Algorithms, Programming) offers a powerful set of tools to investigate computationally group theory. Using this software package we investigate a variation of a well-known problem in…

Group Theory · Mathematics 2017-11-03 Ignacio P. Navarro

There are numerous approaches to building analysis applications across the high-energy physics community. Among them are Python-based, or at least Python-driven, analysis workflows. We aim to ease the adoption of a Python-based analysis…

Computational Physics · Physics 2018-04-25 David Lange

Background: Significance analysis plays a major role in identifying and ranking genes, transcription factor binding sites, DNA methylation regions, and other high-throughput features for association with disease. We propose a new approach,…

Methodology · Statistics 2017-01-10 Andrew E. Jaffe , John D. Storey , Hongkai Ji , Jeffrey T. Leek

Non-sharable sensitive data collection and analysis in large-scale consortia for genomic research is complicated. Time consuming issues in installing software arise due to different operating systems, software dependencies and running the…

In modern drug development, the broader availability of high-dimensional observational data provides opportunities for scientist to explore subgroup heterogeneity, especially when randomized clinical trials are unavailable due to cost and…

Methodology · Statistics 2021-02-24 Xinzhou Guo , Linqing Wei , Chong Wu , Jingshen Wang

The perennial problem of "how many clusters?" remains an issue of substantial interest in data mining and machine learning communities, and becomes particularly salient in large data sets such as populational genomic data where the number…

Machine Learning · Statistics 2009-08-20 Kyung-Ah Sohn , Eric P. Xing
‹ Prev 1 2 3 10 Next ›