English
Related papers

Related papers: An Update to the Minho Quotation Resource

200 papers

Current approaches to automatic summarization of scientific papers generate informative summaries in the form of abstracts. However, abstracts are not intended to show the relationship between a paper and the references cited in it. We…

Computation and Language · Computer Science 2023-11-14 Shahbaz Syed , Ahmad Dawar Hakimi , Khalid Al-Khatib , Martin Potthast

To maintain the company's talent pool, recruiters need to continuously search for resumes from third-party websites (e.g., LinkedIn, Indeed). However, fetched resumes are often incomplete and inaccurate. To improve the quality of…

Artificial Intelligence · Computer Science 2025-09-08 Yu Li , Zulong Chen , Wenjian Xu , Hong Wen , Yipeng Yu , Man Lung Yiu , Yuyu Yin

Is software obsolescence a significant risk? To explore this issue, we analysed a corpus of over 2.5 billion resources corresponding to the UK Web domain, as crawled between 1996 and 2010. Using the DROID and Apache Tika identification…

Digital Libraries · Computer Science 2012-10-08 Andrew N. Jackson

This study investigates the potential of language models to improve the classification of labor market information by linking job vacancy texts to two major European frameworks: the European Skills, Competences, Qualifications and…

In bibliometrics studies, a common challenge is how to deal with incorrect or incomplete data. However, given a large volume of data, there often exists certain relationships between the data items that can allow us to recover missing data…

Digital Libraries · Computer Science 2014-07-23 Tom Z. J. Fu , Qiufang Ying , Dah Ming Chiu

Benchmarking modern large language models (LLMs) on complex and realistic tasks is critical to advancing their development. In this work, we evaluate the factual accuracy and citation performance of state-of-the-art LLMs on the task of…

Computation and Language · Computer Science 2024-12-25 Maya Patel , Aditi Anand

The proliferation of misinformation necessitates scalable, automated fact-checking solutions. Yet, current benchmarks often overlook multilingual and topical diversity. This paper introduces a novel, dynamically extensible data set that…

Computers and Society · Computer Science 2025-10-22 Lorraine Saju , Arnim Bleier , Jana Lasser , Claudia Wagner

Similes are natural language expressions used to compare unlikely things, where the comparison is not taken literally. They are often used in everyday communication and are an important part of cultural heritage. Having an up-to-date corpus…

Computation and Language · Computer Science 2016-05-23 Nikola Milosevic , Goran Nenadic

This paper proposes some modest improvements to Extractor, a state-of-the-art keyphrase extraction system, by using a terabyte-sized corpus to estimate the informativeness and semantic similarity of keyphrases. We present two techniques to…

Computation and Language · Computer Science 2012-04-03 Mario Jarmasz , Caroline Barrière

The Mutual Reinforcement Effect (MRE) describes a phenomenon in information extraction where word-level and sentence-level tasks can mutually improve each other when jointly modeled. While prior work has reported MRE in Japanese, its…

Large language models (LLMs) have emerged as a widely-used tool for information seeking, but their generated outputs are prone to hallucination. In this work, our aim is to allow LLMs to generate text with citations, improving their factual…

Computation and Language · Computer Science 2023-11-01 Tianyu Gao , Howard Yen , Jiatong Yu , Danqi Chen

To seek reliable information sources for news events, we introduce a novel task of expert recommendation, which aims to identify trustworthy sources based on their previously quoted statements. To achieve this, we built a novel dataset,…

Information Retrieval · Computer Science 2024-06-18 Wenjia Zhang , Lin Gui , Rob Procter , Yulan He

Efforts to make research results open and reproducible are increasingly reflected by journal policies encouraging or mandating authors to provide data availability statements. As a consequence of this, there has been a strong uptake of data…

Digital Libraries · Computer Science 2020-07-01 Giovanni Colavizza , Iain Hrynaszkiewicz , Isla Staden , Kirstie Whitaker , Barbara McGillivray

Misinformation is becoming increasingly prevalent on social media and in news articles. It has become so widespread that we require algorithmic assistance utilising machine learning to detect such content. Training these machine learning…

Machine Learning · Computer Science 2022-03-09 Dan Saattrup Nielsen , Ryan McConville

Asking clarification questions is an active area of research; however, resources for training and evaluating search clarification methods are not sufficient. To address this issue, we describe MIMICS-Duo, a new freely available dataset of…

Information Retrieval · Computer Science 2022-06-10 Leila Tavakoli , Johanne R. Trippas , Hamed Zamani , Falk Scholer , Mark Sanderson

We present two novel datasets for the low-resource language Vietnamese to assess models of semantic similarity: ViCon comprises pairs of synonyms and antonyms across word classes, thus offering data to distinguish between similarity and…

Computation and Language · Computer Science 2018-04-20 Kim Anh Nguyen , Sabine Schulte im Walde , Ngoc Thang Vu

The timeline generation task summarises an entity's biography by selecting stories representing key events from a large pool of relevant documents. This paper addresses the lack of a standard dataset and evaluative methodology for the…

Computation and Language · Computer Science 2016-11-08 Xavier Holt , Will Radford , Ben Hachey

This paper introduces a new task in Natural Language Processing (NLP) and Digital Humanities (DH): Mining Asymmetric Intertextuality. Asymmetric intertextuality refers to one-sided relationships between texts, where one text cites, quotes,…

Information Retrieval · Computer Science 2024-10-22 Pak Kin Lau , Stuart Michael McManus

There are different citation habits in the research fields that influence the obsolescence of the research literature. We analyze the distinctive obsolescence of research literature in disciplinary journals in eight scientific subfields…

Digital Libraries · Computer Science 2022-03-17 Pablo Dorta-González , Emilio Gómez-Déniz

In this article, we discuss the outcomes of an experiment where we analysed whether and to what extent the introduction, in 2012, of the new research assessment exercise in Italy (a.k.a. Italian Scientific Habilitation) affected…

Digital Libraries · Computer Science 2020-02-20 Silvio Peroni , Paolo Ciancarini , Aldo Gangemi , Andrea Giovanni Nuzzolese , Francesco Poggi , Valentina Presutti