中文
相关论文

相关论文: Information Retrieval of Jumbled Words

200 篇论文

Learning word representations has garnered greater attention in the recent past due to its diverse text applications. Word embeddings encapsulate the syntactic and semantic regularities of sentences. Modelling word embedding as multi-sense…

计算与语言 · 计算机科学 2019-11-15 P. Jayashree , Ballijepalli Shreya , P. K. Srijith

We address the problem of predicting similarity between a pair of handwritten document images written by different individuals. This has applications related to matching and mining in image collections containing handwritten content. A…

计算机视觉与模式识别 · 计算机科学 2016-05-20 Praveen Krishnan , C. V. Jawahar

We introduce a model for the retrieval of information hidden in legal texts. These are typically organised in a hierarchical (tree) structure, which a reader interested in a given provision needs to explore down to the "deepest" level…

物理与社会 · 物理学 2024-09-17 Yanik-Pascal Förster , Alessia Annibale , Luca Gamberi , Evan Tzanis , Pierpaolo Vivo

Assessing the degree of semantic relatedness between words is an important task with a variety of semantic applications, such as ontology learning for the Semantic Web, semantic search or query expansion. To accomplish this in an automated…

计算与语言 · 计算机科学 2017-05-25 Thomas Niebler , Martin Becker , Christian Pölitz , Andreas Hotho

Hidden structural patterns in written texts have been subject of considerable research in the last decades. In particular, mapping a text into a time series of sentence lengths is a natural way to investigate text structure. Typically,…

计算与语言 · 计算机科学 2018-05-07 Denner S. Vieira , Sergio Picoli , Renio S. Mendes

Increasingly, critical decisions in public policy, governance, and business strategy rely on a deeper understanding of the needs and opinions of constituent members (e.g. citizens, shareholders). While it has become easier to collect a…

计算与语言 · 计算机科学 2020-01-28 Saket Gurukar , Deepak Ajwani , Sourav Dutta , Juho Lauri , Srinivasan Parthasarathy , Alessandra Sala

There have been several efforts to extend distributional semantics beyond individual words, to measure the similarity of word pairs, phrases, and sentences (briefly, tuples; ordered sets of words, contiguous or noncontiguous). One way to…

机器学习 · 计算机科学 2013-10-21 Peter D. Turney

We consider the problem of similarity search within a set of top-k lists under the Kendall's Tau distance function. This distance describes how related two rankings are in terms of concordantly and discordantly ordered items. As top-k lists…

数据库 · 计算机科学 2014-09-03 Koninika Pal , Sebastian Michel

We define a measure of redundant information based on projections in the space of probability distributions. Redundant information between random variables is information that is shared between those variables. But in contrast to mutual…

信息论 · 计算机科学 2013-05-30 Malte Harder , Christoph Salge , Daniel Polani

Most works related to unithood were conducted as part of a larger effort for the determination of termhood. Consequently, the number of independent research that study the notion of unithood and produce dedicated techniques for measuring…

人工智能 · 计算机科学 2008-10-02 Wilson Wong , Wei Liu , Mohammed Bennamoun

The problem of how people find information is studied extensively; however, the problem of how people organize, re-use, and re-find information that they have found is not as well understood. Recently, several projects have conducted…

人机交互 · 计算机科学 2007-05-23 Robert G. Capra , Manuel A. Perez-Quinones

An edit distance is a metric between words that quantifies how two words differ by counting the number of edit operations needed to transform one word into the other one. A word f is said isometric with respect to an edit distance if, for…

形式语言与自动机理论 · 计算机科学 2023-03-07 Marcella Anselmo , Giuseppa Castiglione , Manuela Flores , Dora Giammarresi , Maria Madonia , Sabrina Mantaci

In this paper we exploit concepts of information theory to address the fundamental problem of identifying and defining the most suitable tools to extract, in a automatic and agnostic way, information from a generic string of characters. We…

统计力学 · 物理学 2009-11-10 Andrea Baronchelli , Emanuele Caglioti , Vittorio Loreto

In language recognition, the task of rejecting/differentiating closely spaced versus acoustically far spaced languages remains a major challenge. For confusable closely spaced languages, the system needs longer input test duration material…

声音 · 计算机科学 2016-09-22 Suwon Shon , Seongkyu Mun , John H. L. Hansen , Hanseok Ko

As major progress is made in open-ended text generation, measuring how close machine-generated text is to human language remains a critical open problem. We introduce MAUVE, a comparison measure for open-ended text generation, which…

计算与语言 · 计算机科学 2021-11-24 Krishna Pillutla , Swabha Swayamdipta , Rowan Zellers , John Thickstun , Sean Welleck , Yejin Choi , Zaid Harchaoui

Measuring the congruence between two texts has several useful applications, such as detecting the prevalent deceptive and misleading news headlines on the web. Many works have proposed machine learning based solutions such as text…

计算与语言 · 计算机科学 2020-10-09 Rahul Mishra , Piyush Yadav , Remi Calizzano , Markus Leippold

The problem of guessing a random string is revisited. A close relation between guessing and compression is first established. Then it is shown that if the sequence of distributions of the information spectrum satisfies the large deviation…

信息论 · 计算机科学 2010-08-12 Manjesh Kumar Hanawal , Rajesh Sundaresan

Sentence similarity is considered the basis of many natural language tasks such as information retrieval, question answering and text summarization. The semantic meaning between compared text fragments is based on the words semantic…

信息检索 · 计算机科学 2016-10-17 Issa Atoum , Ahmed Otoom , Narayanan Kulathuramaiyer

Learning high-quality embeddings for rare words is a hard problem because of sparse context information. Mimicking (Pinter et al., 2017) has been proposed as a solution: given embeddings learned by a standard algorithm, a model is first…

计算与语言 · 计算机科学 2019-04-08 Timo Schick , Hinrich Schütze

In this paper, we present an algorithm for evaluating lexical similarity between a given language and several reference language clusters. As an input, we have a list of concepts and the corresponding translations in all considered…

计算与语言 · 计算机科学 2025-04-10 Karol Mikula , Mariana Sarkociová Remešíková