English
Related papers

Related papers: Impossibility of consistent distance estimation fr…

200 papers

In multiple classification, one aims to determine whether a testing sequence is generated from the same distribution as one of the M training sequences or not. Unlike most of existing studies that focus on discrete-valued sequences with…

Machine Learning · Statistics 2024-10-30 Lina Zhu , Lin Zhou

Phylogenetic inference-the derivation of a hypothesis for the common evolutionary history of a group of species- is an active area of research at the intersection of biology, computer science, mathematics, and statistics. One assumes the…

Populations and Evolution · Quantitative Biology 2016-06-21 Ruth Davidson , Joseph Rusinko , Zoe Vernon , Jing Xi

We study the evolution of the graph distance and weighted distance between two fixed vertices in dynamically growing random graph models. More precisely, we consider preferential attachment models with power-law exponent $\tau\in(2,3)$,…

Probability · Mathematics 2023-08-15 Joost Jorritsma , Júlia Komjáthy

We consider a binary sequence generated by thresholding a hidden continuous sequence. The hidden variables are assumed to have a compound symmetry covariance structure with a single parameter characterizing the common correlation. We study…

Statistics Theory · Mathematics 2019-09-04 Haolei Weng , Yang Feng

A fundamental challenge in probabilistic modeling is to balance expressivity and inference efficiency. Tractable probabilistic models (TPMs) aim to directly address this tradeoff by imposing constraints that guarantee efficient inference of…

Artificial Intelligence · Computer Science 2025-10-28 John Leland , YooJung Choi

The problem of change-point estimation is considered under a general framework where the data are generated by unknown stationary ergodic process distributions. In this context, the consistent estimation of the number of change-points is…

Machine Learning · Statistics 2013-02-15 Azaden Khaleghi , Daniil Ryabko

The problem of reconstructing a sequence of independent and identically distributed symbols from a set of equal size, consecutive, fragments, as well as a dependent reference sequence, is considered. First, in the regime in which the…

Information Theory · Computer Science 2023-07-20 Nir Weinberger , Ilan Shomorony

We propose a tree ensemble method, referred to as time series forest (TSF), for time series classification. TSF employs a combination of the entropy gain and a distance measure, referred to as the Entrance (entropy and distance) gain, for…

Machine Learning · Computer Science 2013-06-04 Houtao Deng , George Runger , Eugene Tuv , Martyanov Vladimir

Phylogenetic inference, the task of reconstructing how related sequences evolved from common ancestors, is a central objective in evolutionary genomics. The current state-of-the-art methods exploit probabilistic models of sequence evolution…

Populations and Evolution · Quantitative Biology 2026-02-19 Luc Blassel , Noémie Sauvage , Pierre Barrat-Charlaix , Bastien Boussau , Nicolas Lartillot , Laurent Jacob

We consider deconvolution from repeated observations with unknown error distribution. So far, this model has mostly been studied under the additional assumption that the errors are symmetric. We construct an estimator for the non-symmetric…

Statistics Theory · Mathematics 2014-07-15 Johanna Kappus , Fabienne Comte

We study the complexity of approximations to the normalized information distance. We introduce a hierarchy of computable approximations by considering the number of oscillations. This is a function version of the difference hierarchy for…

Logic · Mathematics 2019-11-15 Klaus Ambos-Spies , Wolfgang Merkle , Sebastiaan A. Terwijn

A sequence $x_1,\dots,x_n,\dots$ of discrete-valued observations is generated according to some unknown probabilistic law (measure) $\mu$. After observing each outcome, one is required to give conditional probabilities of the next…

Machine Learning · Computer Science 2014-12-30 Daniil Ryabko

Typing methods are widely used in the surveillance of infectious diseases, outbreaks investigation and studies of the natural history of an infection. And their use is becoming standard, in particular with the introduction of High…

Data Structures and Algorithms · Computer Science 2020-06-16 Cátia Vaz , Marta Nascimento , João A. Carriço , Tatiana Rocher , Alexandre P. Francisco

Frequencies of $k$-mers in sequences are sometimes used as a basis for inferring phylogenetic trees without first obtaining a multiple sequence alignment. We show that a standard approach of using the squared-Euclidean distance between…

Populations and Evolution · Quantitative Biology 2016-01-15 Elizabeth S. Allman , John A. Rhodes , Seth Sullivant

The presence of reticulate evolutionary events in phylogenies turn phylogenetic trees into phylogenetic networks. These events imply in particular that there may exist multiple evolutionary paths from a non-extant species to an extant one,…

Populations and Evolution · Quantitative Biology 2008-03-21 Gabriel Cardona , Merce Llabres , Francesc Rossello , Gabriel Valiente

We consider the problem of sequential change detection, where the goal is to design a scheme for detecting any changes in a parameter or functional $\theta$ of the data stream distribution that has small detection delay, but guarantees…

Statistics Theory · Mathematics 2023-11-28 Shubhanshu Shekhar , Aaditya Ramdas

The Longest Common Subsequence (LCS) problem is a very important problem in math- ematics, which has a broad application in scheduling problems, physics and bioinformatics. It is known that the given two random sequences of infinite…

Discrete Mathematics · Computer Science 2013-06-19 Kang Ning , Kwok Pui Choi

This paper investigates the soft covering lemma under both the relative entropy and the total variation distance as the measures of deviation. The exact order of the expected deviation of the random i.i.d. code for the soft covering problem…

Information Theory · Computer Science 2019-02-22 Mohammad Hossein Yassaee

This paper focuses on a setting with observations having a cluster dependence structure and presents two main impossibility results. First, we show that when there is only one large cluster, i.e., the researcher does not have any knowledge…

Econometrics · Economics 2023-06-07 Denis Kojevnikov , Kyungchul Song

Canonical distances such as Euclidean distance often fail to capture the appropriate relationships between items, subsequently leading to subpar inference and prediction. Many algorithms have been proposed for automated learning of suitable…

Machine Learning · Statistics 2020-08-24 Tyler M. Tomita , Joshua T. Vogelstein