中文
相关论文

相关论文: Improving Pointwise Mutual Information (PMI) by In…

200 篇论文

In distributional semantics, the pointwise mutual information ($\mathit{PMI}$) weighting of the cooccurrence matrix performs far better than raw counts. There is, however, an issue with unobserved pair cooccurrences as $\mathit{PMI}$ goes…

计算与语言 · 计算机科学 2019-08-20 Alexandre Salle , Aline Villavicencio

Unsupervised approaches to extractive summarization usually rely on a notion of sentence importance defined by the semantic similarity between a sentence and the document. We propose new metrics of relevance and redundancy using pointwise…

计算与语言 · 计算机科学 2021-03-24 Vishakh Padmakumar , He He

Lexical co-occurrence is an important cue for detecting word associations. We present a theoretical framework for discovering statistically significant lexical co-occurrences from a given corpus. In contrast with the prevalent practice of…

计算与语言 · 计算机科学 2010-09-01 Dipak Chaudhari , Om P. Damani , Srivatsan Laxman

Recent work suggests that large language models enhanced with retrieval-augmented generation are easily influenced by the order, in which the retrieved documents are presented to the model when solving tasks such as question answering (QA).…

计算与语言 · 计算机科学 2025-12-12 Tianyu Liu , Jirui Qi , Paul He , Arianna Bisazza , Mrinmaya Sachan , Ryan Cotterell

CLIP and large multimodal models (LMMs) have better accuracy on examples involving concepts that are highly represented in the training data. However, the role of concept combinations in the training data on compositional generalization is…

计算机视觉与模式识别 · 计算机科学 2025-07-11 Helen Qu , Sang Michael Xie

A major concern in using deep learning based generative models for document-grounded dialogs is the potential generation of responses that are not \textit{faithful} to the underlying document. Existing automated metrics used for evaluating…

计算与语言 · 计算机科学 2023-12-04 Yatin Nandwani , Vineet Kumar , Dinesh Raghu , Sachindra Joshi , Luis A. Lastras

The degree of success in document summarization processes depends on the performance of the method used in identifying significant sentences in the documents. The collection of unique words characterizes the major signature of the document,…

信息检索 · 计算机科学 2012-05-09 Aji S , Ramachandra Kaimal

In this paper, we propose a new kernel-based co-occurrence measure that can be applied to sparse linguistic expressions (e.g., sentences) with a very short learning time, as an alternative to pointwise mutual information (PMI). As well as…

计算与语言 · 计算机科学 2020-10-13 Sho Yokoi , Sosuke Kobayashi , Kenji Fukumizu , Jun Suzuki , Kentaro Inui

The Mutual Reinforcement Effect (MRE) investigates the synergistic relationship between word-level and text-level classifications in text classification tasks. It posits that the performance of both classification levels can be mutually…

计算与语言 · 计算机科学 2024-06-06 Chengguang Gan , Xuzheng He , Qinghao Zhang , Tatsunori Mori

To measure the similarity of two documents in the bag-of-words (BoW) vector representation, different term weighting schemes are used to improve the performance of cosine similarity---the most widely used inter-document similarity measure…

信息检索 · 计算机科学 2019-02-12 Sunil Aryal , Kai Ming Ting , Takashi Washio , Gholamreza Haffari

Barlow (1985) hypothesized that the co-occurrence of two events $A$ and $B$ is "suspicious" if $P(A,B) \gg P(A) P(B)$. We first review classical measures of association for $2 \times 2$ contingency tables, including Yule's $Y$ (Yule, 1912),…

机器学习 · 计算机科学 2023-03-03 Christopher K. I. Williams

Query expansion is a well known method to improve the performance of information retrieval systems. In this work we have tested different approaches to extract the candidate query terms from the top ranked documents returned by the…

信息检索 · 计算机科学 2008-12-18 José R. Pérez-Agüera , Lourdes Araujo

The Mutual Reinforcement Effect (MRE) investigates the synergistic relationship between word-level and text-level classifications in text classification tasks. It posits that the performance of both classification levels can be mutually…

计算与语言 · 计算机科学 2024-10-15 Chengguang Gan , Tatsunori Mori

Search techniques make use of elementary information such as term frequencies and document lengths in computation of similarity weighting. They can also exploit richer statistics, in particular the number of documents in which any two terms…

信息检索 · 计算机科学 2020-07-20 Bodo Billerbeck , Justin Zobel , Nicholas Lester , Nick Craswell

Mutual Information (MI) is a powerful statistical measure that quantifies shared information between random variables, particularly valuable in high-dimensional data analysis across fields like genomics, natural language processing, and…

机器学习 · 计算机科学 2024-12-02 Andre O. Falcao

Since its inception, the neural estimation of mutual information (MI) has demonstrated the empirical success of modeling expected dependency between high-dimensional random variables. However, MI is an aggregate statistic and cannot be used…

机器学习 · 计算机科学 2020-10-16 Yao-Hung Hubert Tsai , Han Zhao , Makoto Yamada , Louis-Philippe Morency , Ruslan Salakhutdinov

The persistent mutual information (PMI) is a complexity measure for stochastic processes. It is related to well-known complexity measures like excess entropy or statistical complexity. Essentially it is a variation of the excess entropy so…

数学物理 · 物理学 2012-10-19 Peter Gmeiner

Nowadays, classical count-based word embeddings using positive pointwise mutual information (PPMI) weighted co-occurrence matrices have been widely superseded by machine-learning-based methods like word2vec and GloVe. But these methods are…

计算与语言 · 计算机科学 2020-06-23 Jakob Jungmaier , Nora Kassner , Benjamin Roth

One of the main problems that emerges in the classic approach to semantics is the difficulty in acquisition and maintenance of ontologies and semantic annotations. On the other hand, the Internet explosion and the massive diffusion of…

人工智能 · 计算机科学 2017-01-12 Valentina Franzoni

Metrics based on percentile ranks (PRs) for measuring scholarly impact involves complex treatment because of various defects such as overvaluing or devaluing an object caused by percentile ranking schemes, ignoring precise citation…

数字图书馆 · 计算机科学 2012-05-14 Ping Zhou , Yongfeng Zhong
‹ 上一页 1 2 3 10 下一页 ›