English
Related papers

Related papers: Improving Pointwise Mutual Information (PMI) by In…

200 papers

In distributional semantics, the pointwise mutual information ($\mathit{PMI}$) weighting of the cooccurrence matrix performs far better than raw counts. There is, however, an issue with unobserved pair cooccurrences as $\mathit{PMI}$ goes…

Computation and Language · Computer Science 2019-08-20 Alexandre Salle , Aline Villavicencio

Unsupervised approaches to extractive summarization usually rely on a notion of sentence importance defined by the semantic similarity between a sentence and the document. We propose new metrics of relevance and redundancy using pointwise…

Computation and Language · Computer Science 2021-03-24 Vishakh Padmakumar , He He

Lexical co-occurrence is an important cue for detecting word associations. We present a theoretical framework for discovering statistically significant lexical co-occurrences from a given corpus. In contrast with the prevalent practice of…

Computation and Language · Computer Science 2010-09-01 Dipak Chaudhari , Om P. Damani , Srivatsan Laxman

Recent work suggests that large language models enhanced with retrieval-augmented generation are easily influenced by the order, in which the retrieved documents are presented to the model when solving tasks such as question answering (QA).…

Computation and Language · Computer Science 2025-12-12 Tianyu Liu , Jirui Qi , Paul He , Arianna Bisazza , Mrinmaya Sachan , Ryan Cotterell

CLIP and large multimodal models (LMMs) have better accuracy on examples involving concepts that are highly represented in the training data. However, the role of concept combinations in the training data on compositional generalization is…

Computer Vision and Pattern Recognition · Computer Science 2025-07-11 Helen Qu , Sang Michael Xie

A major concern in using deep learning based generative models for document-grounded dialogs is the potential generation of responses that are not \textit{faithful} to the underlying document. Existing automated metrics used for evaluating…

Computation and Language · Computer Science 2023-12-04 Yatin Nandwani , Vineet Kumar , Dinesh Raghu , Sachindra Joshi , Luis A. Lastras

The degree of success in document summarization processes depends on the performance of the method used in identifying significant sentences in the documents. The collection of unique words characterizes the major signature of the document,…

Information Retrieval · Computer Science 2012-05-09 Aji S , Ramachandra Kaimal

In this paper, we propose a new kernel-based co-occurrence measure that can be applied to sparse linguistic expressions (e.g., sentences) with a very short learning time, as an alternative to pointwise mutual information (PMI). As well as…

Computation and Language · Computer Science 2020-10-13 Sho Yokoi , Sosuke Kobayashi , Kenji Fukumizu , Jun Suzuki , Kentaro Inui

The Mutual Reinforcement Effect (MRE) investigates the synergistic relationship between word-level and text-level classifications in text classification tasks. It posits that the performance of both classification levels can be mutually…

Computation and Language · Computer Science 2024-06-06 Chengguang Gan , Xuzheng He , Qinghao Zhang , Tatsunori Mori

To measure the similarity of two documents in the bag-of-words (BoW) vector representation, different term weighting schemes are used to improve the performance of cosine similarity---the most widely used inter-document similarity measure…

Information Retrieval · Computer Science 2019-02-12 Sunil Aryal , Kai Ming Ting , Takashi Washio , Gholamreza Haffari

Barlow (1985) hypothesized that the co-occurrence of two events $A$ and $B$ is "suspicious" if $P(A,B) \gg P(A) P(B)$. We first review classical measures of association for $2 \times 2$ contingency tables, including Yule's $Y$ (Yule, 1912),…

Machine Learning · Computer Science 2023-03-03 Christopher K. I. Williams

Query expansion is a well known method to improve the performance of information retrieval systems. In this work we have tested different approaches to extract the candidate query terms from the top ranked documents returned by the…

Information Retrieval · Computer Science 2008-12-18 José R. Pérez-Agüera , Lourdes Araujo

The Mutual Reinforcement Effect (MRE) investigates the synergistic relationship between word-level and text-level classifications in text classification tasks. It posits that the performance of both classification levels can be mutually…

Computation and Language · Computer Science 2024-10-15 Chengguang Gan , Tatsunori Mori

Search techniques make use of elementary information such as term frequencies and document lengths in computation of similarity weighting. They can also exploit richer statistics, in particular the number of documents in which any two terms…

Information Retrieval · Computer Science 2020-07-20 Bodo Billerbeck , Justin Zobel , Nicholas Lester , Nick Craswell

Mutual Information (MI) is a powerful statistical measure that quantifies shared information between random variables, particularly valuable in high-dimensional data analysis across fields like genomics, natural language processing, and…

Machine Learning · Computer Science 2024-12-02 Andre O. Falcao

Since its inception, the neural estimation of mutual information (MI) has demonstrated the empirical success of modeling expected dependency between high-dimensional random variables. However, MI is an aggregate statistic and cannot be used…

Machine Learning · Computer Science 2020-10-16 Yao-Hung Hubert Tsai , Han Zhao , Makoto Yamada , Louis-Philippe Morency , Ruslan Salakhutdinov

The persistent mutual information (PMI) is a complexity measure for stochastic processes. It is related to well-known complexity measures like excess entropy or statistical complexity. Essentially it is a variation of the excess entropy so…

Mathematical Physics · Physics 2012-10-19 Peter Gmeiner

Nowadays, classical count-based word embeddings using positive pointwise mutual information (PPMI) weighted co-occurrence matrices have been widely superseded by machine-learning-based methods like word2vec and GloVe. But these methods are…

Computation and Language · Computer Science 2020-06-23 Jakob Jungmaier , Nora Kassner , Benjamin Roth

One of the main problems that emerges in the classic approach to semantics is the difficulty in acquisition and maintenance of ontologies and semantic annotations. On the other hand, the Internet explosion and the massive diffusion of…

Artificial Intelligence · Computer Science 2017-01-12 Valentina Franzoni

Metrics based on percentile ranks (PRs) for measuring scholarly impact involves complex treatment because of various defects such as overvaluing or devaluing an object caused by percentile ranking schemes, ignoring precise citation…

Digital Libraries · Computer Science 2012-05-14 Ping Zhou , Yongfeng Zhong
‹ Prev 1 2 3 10 Next ›