English
Related papers

Related papers: Fast ordered sampling of DNA sequence variants

200 papers

The recent super-exponential growth in the amount of sequencing data generated worldwide has put techniques for compressed storage into the focus. Most available solutions, however, are strictly tied to specific bioinformatics formats,…

Genomics · Quantitative Biology 2021-11-01 Łukasz Roguski , Paolo Ribeca

Dynamic mode decomposition (DMD) is a popular technique for modal decomposition, flow analysis, and reduced-order modeling. In situations where a system is time varying, one would like to update the system's description online as time…

Optimization and Control · Mathematics 2017-07-11 Hao Zhang , Clarence W. Rowley , Eric A. Deem , Louis N. Cattafesta

Disentangling factors of variation has become a very challenging problem on representation learning. Existing algorithms suffer from many limitations, such as unpredictable disentangling factors, poor quality of generated images from…

Computer Vision and Pattern Recognition · Computer Science 2018-03-29 Taihong Xiao , Jiapeng Hong , Jinwen Ma

This paper introduces a novel framework for DNA sequence generation, comprising two key components: DiscDiff, a Latent Diffusion Model (LDM) tailored for generating discrete DNA sequences, and Absorb-Escape, a post-training algorithm…

Genomics · Quantitative Biology 2024-04-18 Zehui Li , Yuhao Ni , William A V Beardall , Guoxuan Xia , Akashaditya Das , Guy-Bart Stan , Yiren Zhao

Cheap high-throughput DNA sequencing may soon become routine not only for human genomes but also for practically anything requiring the identification of living organisms from their DNA: tracking of infectious agents, control of food…

Genomics · Quantitative Biology 2014-03-05 Laurent Gautier , Ole Lund

We present a performant, general-purpose gradient-guided nested sampling algorithm, ${\tt GGNS}$, combining the state of the art in differentiable programming, Hamiltonian slice sampling, clustering, mode separation, dynamic nested…

Machine Learning · Computer Science 2023-12-08 Pablo Lemos , Nikolay Malkin , Will Handley , Yoshua Bengio , Yashar Hezaveh , Laurence Perreault-Levasseur

Subsampling from a large data set is useful in many supervised learning contexts to provide a global view of the data based on only a fraction of the observations. Diverse (or space-filling) subsampling is an appealing subsampling approach…

Methodology · Statistics 2023-11-27 Boyang Shang , Daniel W. Apley , Sanjay Mehrotra

Genomic datasets generated with massively parallel sequencing methods have the potential to propel systematics in new and exciting directions, but selecting appropriate markers and methods is not straightforward. We applied two approaches…

Genomics · Quantitative Biology 2017-03-28 Michael G. Harvey , Brian Tilston Smith , Travis C. Glenn , Brant C. Faircloth , Robb T. Brumfield

With the development of high throughput sequencing technology, it becomes possible to directly analyze mutation distribution in a genome-wide fashion, dissociating mutation rate measurements from the traditional underlying assumptions.…

Genomics · Quantitative Biology 2015-05-14 D. Parkhomchuk , V. S. Amstislavskiy , A. Soldatov , V. Ogryzko

Relative compression, where a set of similar strings are compressed with respect to a reference string, is a very effective method of compressing DNA datasets containing multiple similar sequences. Relative compression is fast to perform…

Quantitative Methods · Quantitative Biology 2011-06-21 Shanika Kuruppu , Simon Puglisi , Justin Zobel

Powerful deep learning tools, such as convolutional neural networks (CNN), are able to learn the input-output relationships of large complicated systems directly from data. Encoder-decoder deep CNNs are able to extract features directly…

Machine Learning · Statistics 2021-06-08 Alexander Scheinker , Frederick Cropp , Sergio Paiagua , Daniele Filippetto

The surge in availability of genomic data holds promise for enabling determination of genetic causes of observed individual traits, with applications to problems such as discovery of the genetic roots of phenotypes, be they molecular…

Distributed, Parallel, and Cluster Computing · Computer Science 2018-04-23 Wayne Joubert , James Nance , Deborah Weighill , Daniel Jacobson

For many machine learning problems, data is abundant and it may be prohibitive to make multiple passes through the full training set. In this context, we investigate strategies for dynamically increasing the effective sample size, when…

Machine Learning · Computer Science 2016-10-10 Hadi Daneshmand , Aurelien Lucchi , Thomas Hofmann

We propose discrete Langevin proposal (DLP), a simple and scalable gradient-based proposal for sampling complex high-dimensional discrete distributions. In contrast to Gibbs sampling-based methods, DLP is able to update all coordinates in…

Machine Learning · Computer Science 2022-06-22 Ruqi Zhang , Xingchao Liu , Qiang Liu

Since the release of human genome sequences, one of the most important research issues is about indexing the genome sequences, and the suffix tree is most widely adopted for that purpose. The traditional suffix tree construction algorithms…

Databases · Computer Science 2015-05-20 Woong-Kee Loh , Yang-Sae Moon , Wookey Lee

We propose a new alignment-free algorithm by constructing a compact vector representation on $\mathbb{R}^{24}$ of a DNA sequence of arbitrary length. Each component of this vector is obtained from a representative sequence, the elements of…

Data Structures and Algorithms · Computer Science 2024-09-27 Probir Mondal , Pratyay Banerjee , Debranjan Pal , Krishnendu Basuli

Quantum advantage, benchmarking the computational power of quantum machines outperforming all classical computers in a specific task, represents a crucial milestone in developing quantum computers and has been driving different physical…

Splice sites play a crucial role in gene expression, and accurate prediction of these sites in DNA sequences is essential for diagnosing and treating genetic disorders. We address the challenge of splice site prediction by introducing…

Genomics · Quantitative Biology 2023-11-23 Asmita Poddar , Vladimir Uzun , Elizabeth Tunbridge , Wilfried Haerty , Alejo Nevado-Holgado

High read depth can be used to assemble short sequence repeats. The existing genome assemblers fail in repetitive regions of longer than average read. I propose a new algorithm for a DNA assembly which uses the relative frequency of reads…

Genomics · Quantitative Biology 2015-01-08 Robert M. Nowak

Multi-organ segmentation of 3D medical images is fundamental with meaningful applications in various clinical automation pipelines. Although deep learning has achieved superior performance, the time and memory consumption of segmenting the…

Computer Vision and Pattern Recognition · Computer Science 2025-11-12 Xueqi Guo , Halid Ziya Yerebakan , Yoshihisa Shinagawa , Kritika Iyer , Gerardo Hermosillo Valadez