English
Related papers

Related papers: An Update to the Minho Quotation Resource

200 papers

Quotation extraction aims to extract quotations from written text. There are three components in a quotation: source refers to the holder of the quotation, cue is the trigger word(s), and content is the main body. Existing solutions for…

Computation and Language · Computer Science 2022-09-21 Yequan Wang , Xiang Li , Aixin Sun , Xuying Meng , Huaming Liao , Jiafeng Guo

Here a collection of 1169 abstracts, which corresponds to articles that the Journal of Marketing Research has published from 2005 to 2014, are analysed under a novel approach. We apply several statistical methods, such as Principal…

Digital Libraries · Computer Science 2016-06-14 Igor Barahona , Daria Micaela Hernandez , Hector Hugo Perez-Villarreal

MMEAD, or MS MARCO Entity Annotations and Disambiguations, is a resource for entity links for the MS MARCO datasets. We specify a format to store and share links for both document and passage collections of MS MARCO. Following this…

Information Retrieval · Computer Science 2023-09-15 Chris Kamphuis , Aileen Lin , Siwen Yang , Jimmy Lin , Arjen P. de Vries , Faegheh Hasibi

We present NEWSROOM, a summarization dataset of 1.3 million articles and summaries written by authors and editors in newsrooms of 38 major news publications. Extracted from search and social media metadata between 1998 and 2017, these…

Computation and Language · Computer Science 2020-05-19 Max Grusky , Mor Naaman , Yoav Artzi

The One Billion Word Benchmark is a dataset derived from the WMT 2011 News Crawl, commonly used to measure language modeling ability in natural language processing. We train models solely on Common Crawl web scrapes partitioned by year, and…

Computation and Language · Computer Science 2021-10-26 Helen Ngo , João G. M. Araújo , Jeffrey Hui , Nicholas Frosst

Current citation practices observed in articles are very noisy, confusing, and not standardised, making identifying the cited works problematic for hu-mans and any reference extraction software. In this work, we want to investigate such…

Digital Libraries · Computer Science 2022-07-22 Erika Alves dos Santos , Silvio Peroni , Marcos Luiz Mucheroni

As Large Language Models (LLMs) continue to evolve, more are being designed to handle long-context inputs. Despite this advancement, most of them still face challenges in accurately handling long-context tasks, often showing the "lost in…

Computation and Language · Computer Science 2024-12-13 Yijiong Yu , Yongfeng Huang , Zhixiao Qi , Zhe Zhou

Large language models (LLMs) are increasingly applied to complex reasoning tasks that require executing several complex steps before receiving any reward. Properly assigning credit to these steps is essential for enhancing model…

Large language models (LLMs) have demonstrated remarkable capabilities in text analysis tasks, yet their evaluation on complex, real-world applications remains challenging. We define a set of tasks, Multi-Insight Multi-Document Extraction…

Computation and Language · Computer Science 2024-12-02 John Francis , Saba Esnaashari , Anton Poletaev , Sukankana Chakraborty , Youmna Hashem , Jonathan Bright

The goal of this working paper is to summarize the main empirical evidences provided by the scientific community as regards the comparison between the two main citation based academic search engines: Google Scholar and Microsoft Academic…

Digital Libraries · Computer Science 2015-09-10 Enrique Orduna-Malea , Juan Manuel Ayllon , Alberto Martin-Martin , Emilio Delgado Lopez-Cozar

Reinforcement learning in large language models (LLMs) often relies on scalar rewards, a practice that discards valuable textual rationale buried in the rollouts, forcing the model to explore \textit{de novo} with each attempt and hindering…

Machine Learning · Computer Science 2025-10-21 Ang Li , Yifei Wang , Zhihang Yuan , Stefanie Jegelka , Yisen Wang

Accurate parsing of citations is necessary for machine-readable scholarly infrastructure. But, despite sustained interest in this problem, existing evaluation techniques are often not generalizable, based on synthetic data, or not publicly…

Digital Libraries · Computer Science 2026-03-27 Parth Sarin , Juan Pablo Alperin , Adam Buttrick , Dione Mentis

This paper presents the formal release of MedMentions, a new manually annotated resource for the recognition of biomedical concepts. What distinguishes MedMentions from other annotated biomedical corpora is its size (over 4,000 abstracts…

Computation and Language · Computer Science 2019-02-26 Sunil Mohan , Donghui Li

This report describes the feature introduced by Google to provide standardized access to institutional affiliations within Google Scholar Citations. First, this new tool is described, pointing out its main characteristics and functioning.…

Digital Libraries · Computer Science 2015-09-16 Enrique Orduna-Malea , Juan Manuel Ayllón , Alberto Martín-Martín , Emilio Delgado López-Cózar

This paper describes a method for multi-document update summarization that relies on a double maximization criterion. A Maximal Marginal Relevance like criterion, modified and so called Smmr, is used to select sentences that are close to…

Information Retrieval · Computer Science 2010-04-21 Florian Boudin , Juan-Manuel Torres-Moreno , Marc El-Bèze

An important editing policy in Wikipedia is to provide citations for added statements in Wikipedia pages, where statements can be arbitrary pieces of text, ranging from a sentence to a paragraph. In many cases citations are either outdated…

Information Retrieval · Computer Science 2017-04-26 Besnik Fetahu , Katja Markert , Wolfgang Nejdl , Avishek Anand

Large language models often solve tasks from a fully specified prompt but degrade when the same requirements unfold over multiple turns, known as the lost-in-conversation (LiC) gap. We trace part of this degradation to self-contamination:…

Computation and Language · Computer Science 2026-05-27 Haoyu Zheng , Yun Zhu , Shu Yuan , Shangming Chen , Qing Wang , Wenqiao Zhang , Jun Xiao , Yueting Zhuang

Finding previously debunked narratives involves identifying claims that have already undergone fact-checking. The issue intensifies when similar false claims persist in multiple languages, despite the availability of debunks for several…

Computation and Language · Computer Science 2024-08-21 Iknoor Singh , Carolina Scarton , Xingyi Song , Kalina Bontcheva

Group Relative Policy Optimization (GRPO) is widely used for training reasoning models, but updating all sampled completions in each group incurs substantial cost and can reinforce verbose reasoning trajectories. In this paper, we study…

Machine Learning · Computer Science 2026-05-28 Qingfei Zhao , Huan Song , Shuyu Tian , Jiawei Shao , Xuelong Li

Using three years of the Journal Citation Reports (2011, 2012, and 2013), indicators of transitions in 2012 (between 2011 and 2013) are studied using methodologies based on entropy statistics. Changes can be indicated at the level of…

Digital Libraries · Computer Science 2015-08-18 Loet Leydesdorff , Wouter de Nooy
‹ Prev 1 3 4 5 6 7 10 Next ›