中文
相关论文

相关论文: Measures of lexical distance between languages

200 篇论文

Human history leaves fingerprints in human languages. Little is known over language evolution and its study is of great importance. Here, we construct a simple stochastic model and compare its results to statistical data of real languages.…

物理与社会 · 物理学 2015-05-13 V. Schwämmle , P. M. C. de Oliveira

A novel type of permutation tests for dendrogram data is studied with respect to two types of metrics for measuring the difference between dendrograms. First, the Frobenius norm is used, and we prove the consistency and efficiency of the…

统计理论 · 数学 2014-03-13 Kei Kobayashi , Mitsuru Orita

A new word usage measure is proposed. It is based on psychophysical relations and allows to reveal words by its degree of "importance" for making basic dictionaries of sublanguages.

计算与语言 · 计算机科学 2007-05-23 V. Kromer

We survey a new area of parameter-free similarity distance measures useful in data-mining, pattern recognition, learning and automatic semantics extraction. Given a family of distances on a set of objects, a distance is universal up to a…

信息检索 · 计算机科学 2007-05-23 Paul Vitanyi

The measure of distance between two fuzzy sets is a fundamental tool within fuzzy set theory. However, current distance measures within the literature do not account for the direction of change between fuzzy sets; a useful concept in a…

人工智能 · 计算机科学 2013-08-26 Josie McCulloch , Christian Wagner , Uwe Aickelin

Does the grammatical gender of a language interfere when measuring the semantic gender information captured by its word embeddings? A number of anomalous gender bias measurements in the embeddings of gendered languages suggest this…

计算机与社会 · 计算机科学 2022-06-06 Shiva Omrani Sabbaghi , Aylin Caliskan

This work presents a novel methodology for calculating the phonetic similarity between words taking motivation from the human perception of sounds. This metric is employed to learn a continuous vector embedding space that groups similar…

计算与语言 · 计算机科学 2021-10-01 Rahul Sharma , Kunal Dhawan , Balakrishna Pailla

A new class of distances appropriate for measuring similarity relations between sequences, say one type of similarity per distance, is studied. We propose a new ``normalized information distance'', based on the noncomputable notion of…

计算复杂性 · 计算机科学 2011-11-09 Ming Li , Xin Chen , Xin Li , Bin Ma , Paul Vitanyi

Dempster-Shafer theory of evidence (D-S theory) is widely used in uncertain information process. The basic probability assignment(BPA) is a key element in D-S theory. How to measure the distance between two BPAs is an open issue. In this…

人工智能 · 计算机科学 2013-11-19 Hongming Mo , Xiaoyan Su , Yong Hu , Yong Deng

The paper proposes a computationally feasible method for measuring context-sensitive semantic distance between words. The distance is computed by adaptive scaling of a semantic space. In the semantic space, each word in the vocabulary V is…

cmp-lg · 计算机科学 2008-02-03 Hideki Kozima , Akira Ito

Words and phrases acquire meaning from the way they are used in society, from their relative semantics to other words and phrases. For computers the equivalent of `society' is `database,' and the equivalent of `use' is `way to search the…

计算与语言 · 计算机科学 2007-06-13 Rudi Cilibrasi , Paul M. B. Vitanyi

Given a set of sequences, the distance between pairs of them helps us to find their similarity and derive structural relationship amongst them. For genomic sequences such measures make it possible to construct the evolution tree of…

信息论 · 计算机科学 2012-08-29 Sandeep Hosangadi

Liu et al. (2017) provide a comprehensive account of research on dependency distance in human languages. While the article is a very rich and useful report on this complex subject, here I will expand on a few specific issues where research…

计算与语言 · 计算机科学 2017-12-14 Carlos Gómez-Rodríguez

We show that short-range phoneme dependencies encode large-scale patterns of linguistic relatedness, with direct implications for quantitative typology and evolutionary linguistics. Specifically, using an information-theoretic framework, we…

计算与语言 · 计算机科学 2026-04-14 Marius Mavridis , Juan De Gregorio , Raul Toral , David Sanchez

Dialect groupings can be discovered objectively and automatically by cluster analysis of phonetic transcriptions such as those found in a linguistic atlas. The first step in the analysis, the computation of linguistic distance between each…

cmp-lg · 计算机科学 2016-08-31 Brett Kessler

Handwritten numerals of different languages have various characteristics. Similarities and dissimilarities of the languages can be measured by analyzing the extracted features of the numerals. Handwritten numeral datasets are available and…

计算机视觉与模式识别 · 计算机科学 2020-12-15 Md. Rahat-uz-Zaman , Shadmaan Hye

Lexical resemblances among a group of languages indicate that the languages could be genetically related, i.e., they could have descended from a common ancestral language. However, such resemblances can arise by chance and, hence, need not…

计算与语言 · 计算机科学 2024-04-02 V. S. D. S. Mahesh Akavarapu , Arnab Bhattacharya

Dialog is a core building block of human natural language interactions. It contains multi-party utterances used to convey information from one party to another in a dynamic and evolving manner. The ability to compare dialogs is beneficial…

计算与语言 · 计算机科学 2021-10-13 Ofer Lavi , Ella Rabinovich , Segev Shlomov , David Boaz , Inbal Ronen , Ateret Anaby-Tavor

We probe the layers in multilingual BERT (mBERT) for phylogenetic and geographic language signals across 100 languages and compute language distances based on the mBERT representations. We 1) employ the language distances to infer and…

计算与语言 · 计算机科学 2020-11-05 Taraka Rama , Lisa Beinborn , Steffen Eger

This paper applies a new model and analytical tool to measure and study contemporary globalization processes in collaborative science - a world in which scientists, scholars, technicians and engineers interact within a 'grid' of…

数字图书馆 · 计算机科学 2012-03-20 Robert J. W. Tijssen , Ludo Waltman , Nees Jan van Eck