English
Related papers

Related papers: Beyond cognacy

200 papers

Phylogenetic inference is an intractable statistical problem on a complex space. Markov chain Monte Carlo methods are the primary tool for Bayesian phylogenetic inference but it is challenging to construct efficient schemes to explore the…

Methodology · Statistics 2022-10-11 Luke J. Kelly , Robin J. Ryder , Grégoire Clarté

In this paper we identify several serious problems that arise in the use of syntactic data from the SSWL database for the purpose of computational phylogenetic reconstruction. We show that the most naive approach fails to produce reliable…

Computation and Language · Computer Science 2016-07-12 Kevin Shu , Sharjeel Aziz , Vy-Luan Huynh , David Warrick , Matilde Marcolli

We construct a family of genealogy-valued Markov processes that are induced by a continuous-time Markov population process. We derive exact expressions for the likelihood of a given genealogy conditional on the history of the underlying…

Probability · Mathematics 2022-01-26 Aaron A. King , Qianying Lin , Edward L. Ionides

Phylogenetics, the inference of evolutionary trees from molecular sequence data such as DNA, is an enterprise that yields valuable evolutionary understanding of many biological systems. Bayesian phylogenetic algorithms, which approximate a…

Populations and Evolution · Quantitative Biology 2016-10-27 Vu Dinh , Aaron E. Darling , Frederick A. Matsen

We introduce a novel evolutionary algorithm (EA) with a semantic network-based representation. For enabling this, we establish new formulations of EA variation operators, crossover and mutation, that we adapt to work on semantic networks.…

Neural and Evolutionary Computing · Computer Science 2015-03-02 Atilim Gunes Baydin , Ramon Lopez de Mantaras , Santiago Ontanon

Generative models, from diffusion models to large language models, achieve remarkable performance but at a cost in training data orders of magnitude larger than what biological learners require. An alternative paradigm has emerged in which…

Machine Learning · Computer Science 2026-05-28 Daniel J. Korchinski , Alessandro Favero , Matthieu Wyart

Automated interlinear gloss prediction with neural networks is a promising approach to accelerate language documentation efforts. However, while state-of-the-art models like GlossLM achieve high scores on glossing benchmarks, user studies…

Computation and Language · Computer Science 2026-01-26 Michael Ginn , Lindia Tjuatja , Enora Rice , Ali Marashian , Maria Valentini , Jasmine Xu , Graham Neubig , Alexis Palmer

Languages evolve over time in a process in which reproduction, mutation and extinction are all possible, similar to what happens to living organisms. Using this similarity it is possible, in principle, to build family trees which show the…

Computation and Language · Computer Science 2012-07-03 Maurizio Serva

In this paper we explore the use of unsupervised methods for detecting cognates in multilingual word lists. We use online EM to train sound segment similarity weights for computing similarity between two words. We tested our online systems…

Computation and Language · Computer Science 2017-02-17 Taraka Rama , Johannes Wahle , Pavel Sofroniev , Gerhard Jäger

Multi-label classification is an important yet challenging task in natural language processing. It is more complex than single-label classification in that the labels tend to be correlated. Existing methods tend to ignore the correlations…

Computation and Language · Computer Science 2018-06-18 Pengcheng Yang , Xu Sun , Wei Li , Shuming Ma , Wei Wu , Houfeng Wang

We introduce a filtration-based framework for studying when and why adding taxa improves phylodynamic inference, by constructing a natural ordering of observed tips and applying sequential Bayesian analysis to the resulting filtration. We…

Quantitative Methods · Quantitative Biology 2026-04-07 David J Pascall

Typing methods are widely used in the surveillance of infectious diseases, outbreaks investigation and studies of the natural history of an infection. And their use is becoming standard, in particular with the introduction of High…

Data Structures and Algorithms · Computer Science 2020-06-16 Cátia Vaz , Marta Nascimento , João A. Carriço , Tatiana Rocher , Alexandre P. Francisco

Causal discovery seeks to uncover causal relations from data, typically represented as causal graphs, and is essential for predicting the effects of interventions. While expert knowledge is required to construct principled causal graphs,…

Artificial Intelligence · Computer Science 2026-02-19 Zihao Li , Fabrizio Russo

The evolutionary relationships between species are typically represented in the biological literature by rooted phylogenetic trees. However, a tree fails to capture ancestral reticulate processes, such as the formation of hybrid species or…

Populations and Evolution · Quantitative Biology 2024-12-09 Johanna Heiss , Daniel H. Huson , Mike Steel

Automatic detection of cognates helps downstream NLP tasks of Machine Translation, Cross-lingual Information Retrieval, Computational Phylogenetics and Cross-lingual Named Entity Recognition. Previous approaches for the task of cognate…

Computation and Language · Computer Science 2021-12-16 Diptesh Kanojia , Prashant Sharma , Sayali Ghodekar , Pushpak Bhattacharyya , Gholamreza Haffari , Malhar Kulkarni

We study phylogenetic signal present in syntactic information by considering the syntactic structures data from Longobardi (2017b), Collins (2010), Ceolin et al. (2020) and Koopman (2011). Focusing first on the general Markov models, we…

Computation and Language · Computer Science 2022-10-19 Sitanshu Gakkhar , Matilde Marcolli

We introduce layered automata, a subclass of alternating parity automata that generalises deterministic automata. Assuming a consistency property, these automata are history deterministic and 0-1 probabilistic. We show that every…

Formal Languages and Automata Theory · Computer Science 2026-01-23 Antonio Casares , Christof Löding , Igor Walukiewicz

Multi-model inference covers a wide range of modern statistical applications such as variable selection, model confidence set, model averaging and variable importance. The performance of multi-model inference depends on the availability of…

Statistics Theory · Mathematics 2019-06-07 Ching-Wei Cheng , Guang Cheng

The identification of syllables within phonetic sequences is known as syllabification. This task is thought to play an important role in natural language understanding, speech production, and the development of speech recognition systems.…

Computation and Language · Computer Science 2019-10-01 Jacob Krantz , Maxwell Dulin , Paul De Palma

Traditionally, phylogeny and sequence alignment are estimated separately: first estimate a multiple sequence alignment and then infer a phylogeny based on the sequence alignment estimated in the previous step. However, uncertainty in the…

Applications · Statistics 2014-11-25 Heejung Shim , Bret Larget