English
Related papers

Related papers: Informed and Automated k-Mer Size Selection for Ge…

200 papers

The $k$-core decomposition is a widely studied summary statistic that describes a graph's global connectivity structure. In this paper, we move beyond using $k$-core decomposition as a tool to summarize a graph and propose using $k$-core…

Statistics Theory · Mathematics 2016-11-29 Vishesh Karwa , Michael J. Pelsmajer , Sonja Petrović , Despina Stasi , Dane Wilburne

We present a framework for uncovering and exploiting dependencies among tools and documents to enhance exemplar artifact generation. Our method begins by constructing a tool knowledge graph from tool schemas,including descriptions,…

Artificial Intelligence · Computer Science 2025-10-29 Shengjie Liu , Li Dong , Zhenyu Zhang

Given $K$ uncertainty sets that are arbitrarily dependent -- for example, confidence intervals for an unknown parameter obtained with $K$ different estimators, or prediction sets obtained via conformal prediction based on $K$ different…

Methodology · Statistics 2024-11-15 Matteo Gasparin , Aaditya Ramdas

We consider the problem of estimating the number of clusters (k) in a dataset. We propose a non-parametric approach to the problem that utilizes similarity graphs to construct a robust statistic that effectively captures similarity…

Methodology · Statistics 2025-06-13 Yichuan Bai , Lynna Chu

This study proposes a data condensation method for multivariate kernel density estimation by genetic algorithm. First, our proposed algorithm generates multiple subsamples of a given size with replacement from the original sample. The…

Methodology · Statistics 2022-03-04 Kiheiji Nishida

Deep neural network-based architectures give promising results in various domains including pattern recognition. Finding the optimal combination of the hyper-parameters of such a large-sized architecture is tedious and requires a large…

Computer Vision and Pattern Recognition · Computer Science 2020-03-17 Animesh Singh , Sandip Saha , Ritesh Sarkhel , Mahantapas Kundu , Mita Nasipuri , Nibaran Das

Retrieval-Augmented Generation (RAG) enhances large language models by incorporating external knowledge. However, existing vector-based methods often fail on global sensemaking tasks that require reasoning across many documents. GraphRAG…

Information Retrieval · Computer Science 2026-03-06 Jakir Hossain , Ahmet Erdem Sarıyüce

Distances between sequences based on their $k$-mer frequency counts can be used to reconstruct phylogenies without first computing a sequence alignment. Past work has shown that effective use of k-mer methods depends on 1) model-based…

Populations and Evolution · Quantitative Biology 2017-05-22 Chris Durden , Seth Sullivant

Motivation: A Genomic Dictionary, i.e., the set of the k-mers appearing in a genome, is a fundamental source of genomic information: its collection is the first step in strategic computational methods ranging from assembly to sequence…

Data Structures and Algorithms · Computer Science 2022-12-07 Raffaele Giancarlo , Gennaro Grimaudo

An important way to make large training sets is to gather noisy labels from crowds of non experts. We propose a method to aggregate noisy labels collected from a crowd of workers or annotators. Eliciting labels is important in tasks such as…

Machine Learning · Computer Science 2016-11-18 Abhay Gupta

We study the problem of partitioning a small sample of $n$ individuals from a mixture of $k$ product distributions over a Boolean cube $\{0, 1\}^K$ according to their distributions. Each distribution is described by a vector of allele…

Machine Learning · Computer Science 2008-02-21 Shuheng Zhou

Autoregressive (AR) architectures have achieved significant successes in LLMs, inspiring explorations for video generation. In LLMs, top-p/top-k sampling strategies work exceptionally well: language tokens have high semantic density and low…

Computer Vision and Pattern Recognition · Computer Science 2026-02-02 Yizhao Han , Tianxing Shi , Zhao Wang , Zifan Xu , Zhiyuan Pu , Mingxiao Li , Qian Zhang , Wei Yin , Xiao-Xiao Long

A Knowledge Graph (KG) is the directed graphical representation of entities and relations in the real world. KG can be applied in diverse Natural Language Processing (NLP) tasks where knowledge is required. The need to scale up and complete…

Computation and Language · Computer Science 2024-04-19 Xincan Feng , Zhi Qu , Yuchang Cheng , Taro Watanabe , Nobuhiro Yugami

Machine learning models are increasingly used to predict material properties and accelerate atomistic simulations, but the reliability of their predictions depends on the representativeness of the training data. We present a scalable,…

Chemical Physics · Physics 2025-10-20 Daniel Willimetz , Lukáš Grajciar

We propose a new algorithm for merging succinct representations of de Bruijn graphs introduced in [Bowe et al. WABI 2012]. Our algorithm is based on the lightweight BWT merging approach by Holt and McMillan [Bionformatics 2014, ACM-BCB…

Data Structures and Algorithms · Computer Science 2020-09-10 Lavinia Egidi , Felipe A. Louza , Giovanni Manzini

Background: In the metagenome assembly of a microbiome community, we may think abundant species would be easier to assemble due to their deeper coverage. However, this conjucture is rarely tested. We often do not know how many abundant…

Genomics · Quantitative Biology 2022-11-23 Xiaowen Feng , Heng Li

Plant breeding programs use data obtained from multi-environment selection experiments to produce improved varieties with the ultimate aim of maintaining high levels of genetic gain. Selection accuracy can be improved with the use of…

Methodology · Statistics 2026-05-13 Brian R Cullis , Alison B Smith , David GD Hughes , David Butler

Ensuring a satisfactory statistical convergence of anharmonic thermodynamic properties requires sampling of many atomic configurations, however the methods to obtain those necessarily produce correlated samples, thereby reducing the…

Statistical Mechanics · Physics 2022-06-07 Erki Metsanurk

The estimation of probability densities based on available data is a central task in many statistical applications. Especially in the case of large ensembles with many samples or high-dimensional sample spaces, computationally efficient…

Methodology · Statistics 2017-05-04 Daniel W. Meyer

Metagenomic profiling is challenging in part because of the highly uneven sampling of the tree of life by genome sequencing projects and the limitations imposed by performing phylogenetic inference at fixed taxonomic ranks. We present the…

Genomics · Quantitative Biology 2016-02-18 David Koslicki , Daniel Falush
‹ Prev 1 3 4 5 6 7 10 Next ›