English
Related papers

Related papers: Improving Pointwise Mutual Information (PMI) by In…

200 papers

Estimating mutual information (MI) from samples is a fundamental problem in statistics, machine learning, and data analysis. Recently it was shown that a popular class of non-parametric MI estimators perform very poorly for strongly…

Information Theory · Computer Science 2016-02-18 Shuyang Gao , Greg Ver Steeg , Aram Galstyan

Many of the classical and recent relations between information and estimation in the presence of Gaussian noise can be viewed as identities between expectations of random quantities. These include the I-MMSE relationship of Guo et al.; the…

Information Theory · Computer Science 2012-05-02 Kartik Venkat , Tsachy Weissman

Providing natural language-based explanations to justify recommendations helps to improve users' satisfaction and gain users' trust. However, as current explanation generation methods are commonly trained with an objective to mimic existing…

Information Retrieval · Computer Science 2024-08-22 Yurou Zhao , Yiding Sun , Ruidong Han , Fei Jiang , Lu Guan , Xiang Li , Wei Lin , Weizhi Ma , Jiaxin Mao

We propose the conditional predictive impact (CPI), a consistent and unbiased estimator of the association between one or several features and a given outcome, conditional on a reduced feature set. Building on the knockoff framework of…

Methodology · Statistics 2021-05-14 David S. Watson , Marvin N. Wright

We address in this paper the co-clustering and co-classification of bilingual data laying in two linguistic similarity spaces when a comparability measure defining a mapping between these two spaces is available. A new approach that we can…

Information Retrieval · Computer Science 2015-02-27 Pierre-François Marteau , Guiyao Ke

The abundance of training data is not guaranteed in various supervised learning applications. One of these situations is the post-earthquake regional damage assessment of buildings. Querying the damage label of each building requires a…

Machine Learning · Computer Science 2021-08-17 Mohamadreza Sheibani , Ge Ou

We investigate the similarities of pairs of articles which are co-cited at the different co-citation levels of the journal, article, section, paragraph, sentence and bracket. Our results indicate that textual similarity, intellectual…

Digital Libraries · Computer Science 2017-10-30 Giovanni Colavizza , Kevin W. Boyack , Nees Jan van Eck , Ludo Waltman

The use of Mutual Information (MI) as a measure to evaluate the efficiency of cryptosystems has an extensive history. However, estimating MI between unknown random variables in a high-dimensional space is challenging. Recent advances in…

We address the problem of recommending relevant items to a user in order to "complete" a partial set of items already known. We consider the two scenarios of citation and subject label recommendation, which resemble different semantics of…

Information Retrieval · Computer Science 2021-05-11 Iacopo Vagliano , Lukas Galke , Ansgar Scherp

A central challenge in the study of complex systems is the quantification of emergence -- understood as the ability of the system to exhibit collective behaviours that cannot be traced down to the individual components. While recent work…

Objective: Semantic indexing of biomedical literature is usually done at the level of MeSH descriptors with several related but distinct biomedical concepts often grouped together and treated as a single topic. This study proposes a new…

Computation and Language · Computer Science 2023-10-06 Anastasios Nentidis , Thomas Chatzopoulos , Anastasia Krithara , Grigorios Tsoumakas , Georgios Paliouras

This work focuses on learning useful and robust deep world models using multiple, possibly unreliable, sensors. We find that current methods do not sufficiently encourage a shared representation between modalities; this can cause poor…

Machine Learning · Computer Science 2021-07-07 Kaiqi Chen , Yong Lee , Harold Soh

In Science, Reshef et al. (2011) proposed the concept of equitability for measures of dependence between two random variables. To this end, they proposed a novel measure, the maximal information coefficient (MIC). Recently a PNAS paper…

Methodology · Statistics 2023-04-17 A. Adam Ding , Yi Li

We present the results of our system for the CoMeDi Shared Task, which predicts majority votes (Subtask 1) and annotator disagreements (Subtask 2). Our approach combines model ensemble strategies with MLP-based and threshold-based methods…

Computation and Language · Computer Science 2024-12-31 Zhu Liu , Zhen Hu , Ying Liu

Mutual information is widely used, in a descriptive way, to measure the stochastic dependence of categorical random variables. In order to address questions such as the reliability of the descriptive value, one must consider…

Machine Learning · Computer Science 2007-07-13 Marcus Hutter , Marco Zaffalon

How to leverage cross-document interactions to improve ranking performance is an important topic in information retrieval (IR) research. However, this topic has not been well-studied in the learning-to-rank setting and most of the existing…

Information Retrieval · Computer Science 2019-10-24 Rama Kumar Pasumarthi , Xuanhui Wang , Michael Bendersky , Marc Najork

State-of-the-art coreference resolutions systems depend on multiple LLM calls per document and are thus prohibitively expensive for many use cases (e.g., information extraction with large corpora). The leading word-level coreference system…

Computation and Language · Computer Science 2023-10-20 Karel D'Oosterlinck , Semere Kiros Bitew , Brandon Papineau , Christopher Potts , Thomas Demeester , Chris Develder

A comparison was made of vectors derived by using ordinary co-occurrence statistics from large text corpora and of vectors derived by measuring the inter-word distances in dictionary definitions. The precision of word sense disambiguation…

cmp-lg · Computer Science 2016-08-31 Yoshiki Niwa , Yoshihiko Nitta

We argue that citation is a composed indicator: short-term citations can be considered as currency at the research front, whereas long-term citations can contribute to the codification of knowledge claims into concept symbols. Knowledge…

Digital Libraries · Computer Science 2016-07-22 Loet Leydesdorff , Lutz Bornmann , Jordan Comins , Staša Milojević

In this paper, I propose a novel word sense disambiguation method based on the global co-occurrence information using NMF. When I calculate the dependency relation matrix, the existing method tends to produce very sparse co-occurrence…

Computation and Language · Computer Science 2014-03-06 Minoru Sasaki