中文
相关论文

相关论文: Measures of lexical distance between languages

200 篇论文

A set of ontology matching algorithms (for finding correspondences between concepts) is based on a thesaurus that provides the source data for the semantic distance calculations. In this wiki era, new resources may spring up and improve…

信息检索 · 计算机科学 2009-10-12 A. A. Krizhanovsky , Feiyu Lin

Knowing the degree of semantic contrast between words has widespread application in natural language processing, including machine translation, information retrieval, and dialogue systems. Manually-created lexicons focus on opposites, such…

计算与语言 · 计算机科学 2013-08-30 Saif M. Mohammad , Bonnie J. Dorr , Graeme Hirst , Peter D. Turney

Today, with the emergence of semantic web technologies and increasing of information quantity, searching for information based on the semantic web has become a fertile area of research. For this reason, a large number of studies are…

计算机视觉与模式识别 · 计算机科学 2021-10-05 Noreddine Gherabi , Abdelhadi Daoui , Abderrahim Marzouk

The measurement of distance between two objects is generalized to the case where the objects are no longer points but are one-dimensional. Additional concepts such as non-extensibility, curvature constraints, and non-crossing become central…

软凝聚态物质 · 物理学 2008-03-04 Steven S. Plotkin

Artificial Neural networks are mathematical models at their core. This truismpresents some fundamental difficulty when networks are tasked with Natural Language Processing. A key problem lies in measuring the similarity or distance among…

计算与语言 · 计算机科学 2021-06-07 Thomas Conley , Jugal Kalita

The average uncertainty associated with words is an information-theoretic concept at the heart of quantitative and computational linguistics. The entropy has been established as a measure of this average uncertainty - also called average…

计算与语言 · 计算机科学 2016-06-23 Christian Bentz , Dimitrios Alikaniotis

In recent geospatial research, the importance of modeling large-scale human mobility data and predicting trajectories is rising, in parallel with progress in text generation using large-scale corpora in natural language processing. Whereas…

机器学习 · 计算机科学 2022-11-02 Toru Shimizu , Kota Tsubouchi , Takahiro Yabe

As subjects perceive the sensory world, different stimuli elicit a number of neural representations. Here, a subjective distance between stimuli is defined, measuring the degree of similarity between the underlying representations. As an…

神经元与认知 · 定量生物学 2007-05-23 D. Oliva , I. Samengo , S. Leutgeb , S. Mizumori

This paper addresses the problem of determining the distance between two regular languages. It will show how to expand Jaccard distance, which works on finite sets, to potentially-infinite regular languages. The entropy of a regular…

形式语言与自动机理论 · 计算机科学 2016-02-26 Austin J. Parker , Kelly B. Yancey , Matthew P. Yancey

The metric properties of the set in which random variables take their values lead to relevant probabilistic concepts. For example, the mean of a random variable is a best predictor in that it minimizes the standard Euclidean distance or…

概率论 · 数学 2018-09-21 Henryk Gzyl

Transformer-based language models have recently achieved remarkable results in many natural language tasks. However, performance on leaderboards is generally achieved by leveraging massive amounts of training data, and rarely by encoding…

计算与语言 · 计算机科学 2022-07-21 Bai Li

In this paper, we present a novel interpretation of the so-called Weisfeiler-Lehman (WL) distance, introduced by Chen et al. (2022), using concepts from stochastic processes. The WL distance aims at comparing graphs with node features, has…

机器学习 · 计算机科学 2023-10-03 Samantha Chen , Sunhyuk Lim , Facundo Mémoli , Zhengchao Wan , Yusu Wang

Modern large language models (LLMs) such as GPT, Claude, and Gemini have transformed the way we learn, work, and communicate. Yet, their ability to produce highly human-like text raises serious concerns about misinformation and academic…

计算与语言 · 计算机科学 2026-03-03 Hongyi Zhou , Jin Zhu , Kai Ye , Ying Yang , Erhan Xu , Chengchun Shi

A geometric graph is a combinatorial graph, endowed with a geometry that is inherited from its embedding in a Euclidean space. Formulation of a meaningful measure of (dis-)similarity in both the combinatorial and geometric structures of two…

计算几何 · 计算机科学 2022-09-27 Sushovan Majhi , Carola Wenk

Many real world tasks where Large Language Models (LLMs) can be used require spatial reasoning, like Point of Interest (POI) recommendation and itinerary planning. However, on their own LLMs lack reliable spatial reasoning capabilities,…

计算与语言 · 计算机科学 2025-06-05 Nicole R Schneider , Nandini Ramachandran , Kent O'Sullivan , Hanan Samet

Ten years ago a single metric, BLEU, governed progress in machine translation research. For better or worse, there is no such consensus today, and consequently it is difficult for researchers to develop and retain the kinds of heuristic…

计算与语言 · 计算机科学 2024-06-11 Tom Kocmi , Vilém Zouhar , Christian Federmann , Matt Post

A longstanding debate in semiotics centers on the relationship between linguistic signs and their corresponding semantics: is there an arbitrary relationship between a word form and its meaning, or does some systematic phenomenon pervade?…

计算与语言 · 计算机科学 2019-07-30 Tiago Pimentel , Arya D. McCarthy , Damián E. Blasi , Brian Roark , Ryan Cotterell

We propose an unsupervised solution to the Authorship Verification task that utilizes pre-trained deep language models to compute a new metric called DV-Distance. The proposed metric is a measure of the difference between the two authors…

计算与语言 · 计算机科学 2021-03-15 Yifan Zhang , Dainis Boumber , Marjan Hosseinia , Fan Yang , Arjun Mukherjee

A classical problem in grammatical inference is to identify a language from a set of examples. In this paper, we address the problem of identifying a union of languages from examples that belong to several different unknown languages.…

形式语言与自动机理论 · 计算机科学 2018-12-21 Alexis Linard , Colin de la Higuera , Frits Vaandrager

Cilibrasi and Vitanyi have demonstrated that it is possible to extract the meaning of words from the world-wide web. To achieve this, they rely on the number of webpages that are found through a Google search containing a given word and…

计算与语言 · 计算机科学 2015-01-29 Bjørn Kjos-Hanssen , Alberto J. Evangelista