English
Related papers

Related papers: Chinese Restaurant Process for cognate clustering:…

200 papers

One of the most used priors in Bayesian clustering is the Dirichlet prior. It can be expressed as a Chinese Restaurant Process. This process allows nonparametric estimation of the number of clusters when partitioning datasets. Its key…

Machine Learning · Computer Science 2021-04-27 Gaël Poux-Médard , Julien Velcin , Sabine Loudcher

Dirichlet process mixture (DPM) models tend to produce many small clusters regardless of whether they are needed to accurately characterize the data - this is particularly true for large data sets. However, interpretability, parsimony, data…

Machine Learning · Computer Science 2018-02-16 Jun Lu , Meng Li , David Dunson

We develop the distance dependent Chinese restaurant process (CRP), a flexible class of distributions over partitions that allows for non-exchangeability. This class can be used to model many kinds of dependencies between data in infinite…

Machine Learning · Statistics 2011-08-11 David M. Blei , Peter I. Frazier

We present the nested Chinese restaurant process (nCRP), a stochastic process which assigns probability distributions to infinitely-deep, infinitely-branching trees. We show how this stochastic process can be used as a prior distribution in…

Machine Learning · Statistics 2009-08-27 David M. Blei , Thomas L. Griffiths , Michael I. Jordan

Learning from a continuous stream of non-stationary data in an unsupervised manner is arguably one of the most common and most challenging settings facing intelligent agents. Here, we attack learning under all three conditions…

Machine Learning · Computer Science 2023-05-23 Rylan Schaeffer , Gabrielle Kaili-May Liu , Yilun Du , Scott Linderman , Ila Rani Fiete

In this article I proposed a new model to achieve Chinese word segmentation(CWS),which may have the potentiality to apply in other domains in the future.It is a new thinking in CWS compared to previous works,to consider it as a clustering…

Computation and Language · Computer Science 2020-02-19 Yuze Zhao

The Generalized Chinese Restaurant Process (GCRP) describes a sequence of exchangeable random partitions of the numbers $\{1,\dots,n\}$. This process is related to the Ewens sampling model in Genetics and to Bayesian nonparametric methods…

Probability · Mathematics 2018-06-27 Alan Pereira , Roberto I. Oliveira , Rodrigo Ribeiro

Chinese word segmentation (CWS) is a fundamental step of Chinese natural language processing. In this paper, we build a new toolkit, named PKUSEG, for multi-domain word segmentation. Unlike existing single-model toolkits, PKUSEG targets…

Computation and Language · Computer Science 2022-05-31 Ruixuan Luo , Jingjing Xu , Yi Zhang , Zhiyuan Zhang , Xuancheng Ren , Xu Sun

This paper investigates the potential benefits of language-specific fact-checking models, focusing on the case of Chinese. We first demonstrate the limitations of translation-based methods and multilingual large language models (e.g.,…

Computation and Language · Computer Science 2024-10-07 Caiqi Zhang , Zhijiang Guo , Andreas Vlachos

In this paper we explore the use of unsupervised methods for detecting cognates in multilingual word lists. We use online EM to train sound segment similarity weights for computing similarity between two words. We tested our online systems…

Computation and Language · Computer Science 2017-02-17 Taraka Rama , Johannes Wahle , Pavel Sofroniev , Gerhard Jäger

Due to the absence of labeled data, discourse parsing still remains challenging in some languages. In this paper, we present a simple and efficient method to conduct zero-shot Chinese text-level dependency parsing by leveraging English…

Computation and Language · Computer Science 2019-11-28 Yi Cheng , Sujian Li

The Chinese restaurant process is a basic sequential construction of consistent random partitions. We consider random point measures describing the composition of small blocks in such partitions and show that their scaling limit is given by…

Probability · Mathematics 2025-10-09 Oleksii Galganov , Andrii Ilienko

In this paper, we present a dynamic semantic clustering approach inspired by the Chinese Restaurant Process, aimed at addressing uncertainty in the inference of Large Language Models (LLMs). We quantify uncertainty of an LLM on a given…

We have developed a high-performance Chinese Chess AI that operates without reliance on search algorithms. This AI has demonstrated the capability to compete at a level commensurate with the top 0.1\% of human players. By eliminating the…

Machine Learning · Computer Science 2024-10-08 Yu Chen , Juntong Lin , Zhichao Shu

Most previous approaches to Chinese word segmentation formalize this problem as a character-based sequence labeling task where only contextual information within fixed sized local windows and simple interactions between adjacent tags can be…

Computation and Language · Computer Science 2016-12-05 Deng Cai , Hai Zhao

We establish scaling limit theorems for the up-down ordered Chinese restaurant processes (oCRPs) of Rogers and Winkel as processes in a space of interval partitions. As previously conjectured, the limits are self-similar diffusions…

Probability · Mathematics 2025-12-09 Quan Shi , Matthias Winkel

This article proposes a Bayesian nonparametric method for forecasting, imputation, and clustering in sparsely observed, multivariate time series data. The method is appropriate for jointly modeling hundreds of time series with widely…

Methodology · Statistics 2019-02-27 Feras A. Saad , Vikash K. Mansinghka

The Recurrent Chinese Restaurant Process (RCRP) is a powerful statistical method for modeling evolving clusters in large scale social media data. With the RCRP, one can allow both the number of clusters and the cluster parameters in a model…

Artificial Intelligence · Computer Science 2017-08-22 Wei Wei , Kennth Joseph , Kathleen Carley

This paper focuses on the problem of hierarchical non-overlapping clustering of a dataset. In such a clustering, each data item is associated with exactly one leaf node and each internal node is associated with all the data items stored in…

Machine Learning · Statistics 2021-05-26 Weipeng Huang , Nishma Laitonjam , Guangyuan Piao , Neil Hurley

This paper proposes a fully unsupervised approach to the construction of verb collostruction database for Chinese language, aimed at complementing LLMs by providing explicit and interpretable rules for application scenarios where…

Computation and Language · Computer Science 2026-01-09 Xuri Tang , Daohuan Liu
‹ Prev 1 2 3 10 Next ›