English
Related papers

Related papers: Finding easy regions for short-read variant callin…

200 papers

Large-scale pretrained language models have led to dramatic improvements in text generation. Impressive performance can be achieved by finetuning only on a small number of instances (few-shot setting). Nonetheless, almost all previous work…

Computation and Language · Computer Science 2021-07-08 Ernie Chang , Xiaoyu Shen , Hui-Syuan Yeh , Vera Demberg

Random Indexing is a simple implementation of Random Projections with a wide range of applications. It can solve a variety of problems with good accuracy without introducing much complexity. Here we use it for identifying the language of…

Computation and Language · Computer Science 2015-03-02 Aditya Joshi , Johan Halseth , Pentti Kanerva

Machine learning models have prevalent applications in many real-world problems, which increases the importance of correctness in the behaviour of these trained models. Finding a good test case that can reveal the potential failure in these…

Machine Learning · Computer Science 2022-06-14 Harsh Vardhan , Janos Sztipanovits

The problem of detecting a binding site -- a substring of DNA where transcription factors attach -- on a long DNA sequence requires the recognition of a small pattern in a large background. For short binding sites, the matching probability…

Genomics · Quantitative Biology 2009-11-13 Daniela Bianchi , Brunello Tirozzi

Mapping gene expression as a quantitative trait using whole genome-sequencing and transcriptome analysis allows to discover the functional consequences of genetic variation. We developed a novel method and ultra-fast software Findr for…

Genomics · Quantitative Biology 2017-08-23 Lingfei Wang , Tom Michoel

To segment a sequence of independent random variables at an unknown number of change-points, we introduce new procedures that are based on thresholding the likelihood ratio statistic. We also study confidence regions based on the likelihood…

Statistics Theory · Mathematics 2018-10-16 Xiao Fang , Jian Li , David Siegmund

Random forest is a classification algorithm well suited for microarray data: it shows excellent performance even when most predictive variables are noise, can be used when the number of variables is much larger than the number of…

Quantitative Methods · Quantitative Biology 2007-05-23 Ramon Diaz-Uriarte , Sara Alvarez de Andres

Consider the observation of n iid realizations of an experiment with d>1 possible outcomes, which corresponds to a single observation of a multinomial distribution M(n,p) where p is an unknown discrete distribution on {1,...,d}. In many…

Computation · Statistics 2010-06-15 Djalil Chafai , Didier Concordet

DNA data storage systems encode digital data into DNA strands, enabling dense and durable storage. Efficient data retrieval depends on coverage depth, a key performance metric. We study the random access coverage depth problem and focus on…

Information Theory · Computer Science 2025-07-29 Şeyma Bodur , Stefano Lia , Hiram H. López , Rati Ludhani , Alberto Ravagnani , Lisa Seccia

A general approach to selective inference is considered for hypothesis testing of the null hypothesis represented as an arbitrary shaped region in the parameter space of multivariate normal model. This approach is useful for hierarchical…

Statistics Theory · Mathematics 2018-03-28 Yoshikazu Terada , Hidetoshi Shimodaira

Clinical adoption of human genome sequencing requires methods with known accuracy of genotype calls at millions or billions of positions across a genome. Previous work showing discordance amongst sequencing methods and algorithms has made…

Genomics · Quantitative Biology 2014-02-18 Justin M. Zook , Brad Chapman , Jason Wang , David Mittelman , Oliver Hofmann , Winston Hide , Marc Salit

In the last few years, deep learning classifiers have shown promising results in image-based medical diagnosis. However, interpreting the outputs of these models remains a challenge. In cancer diagnosis, interpretability can be achieved by…

Computer Vision and Pattern Recognition · Computer Science 2021-06-16 Kangning Liu , Yiqiu Shen , Nan Wu , Jakub Chłędowski , Carlos Fernandez-Granda , Krzysztof J. Geras

Motivation: Genome-wide association studies (GWASs), which assay more than a million single nucleotide polymorphisms (SNPs) in thousands of individuals, have been widely used to identify genetic risk variants for complex diseases. However,…

Computational Engineering, Finance, and Science · Computer Science 2015-01-27 Ben Teng , Can Yang , Jiming Liu , Zhipeng Cai , Xiang Wan

This work proposes a new method for computing acceptance regions of exact multinomial tests. From this an algorithm is derived, which finds exact p-values for tests of simple multinomial hypotheses. Using concepts from discrete convex…

Computation · Statistics 2023-05-31 Johannes Resin

Recent progress in multi-object filtering has led to algorithms that compute the first-order moment of multi-object distributions based on sensor measurements. The number of targets in arbitrarily selected regions can be estimated using the…

Applications · Statistics 2013-10-11 Emmanuel Delande , Murat Uney , Jeremie Houssineau , Daniel Clark

There are currently plenty of programs available for mapping short sequences (reads) to a genome. Most of them, however, including such popular and actively developed programs as Bowtie, BWA, TopHat and many others, are based on…

Genomics · Quantitative Biology 2019-08-06 Igor Seledtsov , Jaroslav Efremov , Vladimir Molodtsov , Victor Solovyev

Human populations have experienced dramatic growth since the Neolithic revolution. Recent studies that sequenced a very large number of individuals observed an extreme excess of rare variants, and provided clear evidence of recent rapid…

Populations and Evolution · Quantitative Biology 2015-06-17 Elodie Gazave , Li Ma , Diana Chang , Alex Coventry , Feng Gao , Donna Muzny , Eric Boerwinkle , Richard Gibbs , Charles F. Sing , Andrew G. Clark , Alon Keinan

The discovery of genetic risk factors has transformed human genetics, yet the pace of new gene identification has slowed despite the exponential expansion of sequencing and biobank resources. Current approaches are optimized for the…

Genomics · Quantitative Biology 2025-11-11 Madison Caballero , Behrang Mahjani

Efficient and consistent string processing is critical in the exponentially growing genomic data era. Locally Consistent Parsing (LCP) addresses this need by partitioning an input genome string into short, exactly matching substrings (e.g.,…

An open question in \emph{Imprecise Probabilistic Machine Learning} is how to empirically derive a credal region (i.e., a closed and convex family of probabilities on the output space) from the available data, without any prior knowledge or…

Machine Learning · Statistics 2025-01-29 Michele Caprio , David Stutz , Shuo Li , Arnaud Doucet
‹ Prev 1 3 4 5 6 7 10 Next ›