English
Related papers

Related papers: An Update to the Minho Quotation Resource

200 papers

We introduce Multi-SimLex, a large-scale lexical resource and evaluation benchmark covering datasets for 12 typologically diverse languages, including major languages (e.g., Mandarin Chinese, Spanish, Russian) as well as less-resourced ones…

Data from online job postings are difficult to access and are not built in a standard or transparent manner. Data included in the standard taxonomy and occupational information database (O*NET) are updated infrequently and based on small…

Computers and Society · Computer Science 2025-10-03 Stephen Meisenbacher , Svetlozar Nestorov , Peter Norlander

Automatic and accurate classification of items enables numerous downstream applications in many domains. These applications can range from faceted browsing of items to product recommendations and big data analytics. In the online…

Artificial Intelligence · Computer Science 2016-09-21 Yun Zhu , Faizan Javed , Ozgur Ozturk

We introduce entity post-modifier generation as an instance of a collaborative writing task. Given a sentence about a target entity, the task is to automatically generate a post-modifier phrase that provides contextually relevant…

Computation and Language · Computer Science 2019-04-10 Jun Seok Kang , Robert L. Logan , Zewei Chu , Yang Chen , Dheeru Dua , Kevin Gimpel , Sameer Singh , Niranjan Balasubramanian

Quotes of public figures can mark turning points in history. A quote can explain its originator's actions, foreshadowing political or personal decisions and revealing character traits. Impactful quotes cross language barriers and influence…

Computation and Language · Computer Science 2022-07-21 Tin Kuculo , Simon Gottschalk , Elena Demidova

In recent years, the problem of misinformation on the web has become widespread across languages, countries, and various social media platforms. Although there has been much work on automated fake news detection, the role of images and…

Computation and Language · Computer Science 2022-05-05 Gullal S. Cheema , Sherzod Hakimov , Abdul Sittar , Eric Müller-Budack , Christian Otto , Ralph Ewerth

To enable building and testing models on long-document comprehension, we introduce QuALITY, a multiple-choice QA dataset with context passages in English that have an average length of about 5,000 tokens, much longer than typical current…

Legal citation in common-law systems depends not only on factual similarity, but also on the legal principle for which a precedent is invoked. However, existing benchmarks for legal citation retrieval use case facts, citation context, or…

Information Retrieval · Computer Science 2026-05-21 Shannon Lee Yueh Ern , Kaidong Feng , Yingpeng Du , Chloe Lee En Jia , Zhu Sun

There has been growing interest in utilizing occupational data mining and analysis. In today's job market, occupational data mining and analysis is growing in importance as it enables companies to predict employee turnover, model career…

Social and Information Networks · Computer Science 2020-05-07 Junhua Liu , Yung Chuen Ng , Kwan Hui Lim

The performance and usability of Large-Language Models (LLMs) are driving their use in explanation generation tasks. However, despite their widespread adoption, LLM explanations have been found to be unreliable, making it difficult for…

Computation and Language · Computer Science 2025-06-05 Diana Galvan-Sosa , Gabrielle Gaudeau , Pride Kavumba , Yunmeng Li , Hongyi gu , Zheng Yuan , Keisuke Sakaguchi , Paula Buttery

Job search through online matching engines nowadays are very prominent and beneficial to both job seekers and employers. But the solutions of traditional engines without understanding the semantic meanings of different resumes have not kept…

Computation and Language · Computer Science 2016-07-27 Yiou Lin , Hang Lei , Prince Clement Addo , Xiaoyu Li

Using a corpus of 17,000+ financial news reports (involving over 10M words), we perform an analysis of the argument-distributions of the UP and DOWN verbs used to describe movements of indices, stocks and shares. In Study 1 participants…

Computation and Language · Computer Science 2013-02-05 Aaron Gerow , Mark Keane

While a source sentence can be translated in many ways, most machine translation (MT) models are trained with only a single reference. Previous work has shown that using synthetic paraphrases can improve MT. This paper investigates best…

Computation and Language · Computer Science 2025-02-27 Si Wu , John Wieting , David A. Smith

In this paper, we introduce the MLM (Multiple Languages and Modalities) dataset - a new resource to train and evaluate multitask systems on samples in multiple modalities and three languages. The generation process and inclusion of semantic…

Machine Learning · Computer Science 2020-10-27 Jason Armitage , Endri Kacupaj , Golsa Tahmasebzadeh , Swati , Maria Maleshkova , Ralph Ewerth , Jens Lehmann

In this paper, we present an updated version of the NELA-GT-2019 dataset, entitled NELA-GT-2020. NELA-GT-2020 contains nearly 1.8M news articles from 519 sources collected between January 1st, 2020 and December 31st, 2020. Just as with…

Computers and Society · Computer Science 2021-02-10 Maurício Gruppi , Benjamin D. Horne , Sibel Adalı

In this paper, we revisit cognitive extent, originally defined as the number of unique phrases in a quota. We introduce Freshness and Informative Weighted Cognitive Extent (FICE), calculated based on two novel weighting factors, the…

Digital Libraries · Computer Science 2024-12-31 Zihe Wang , Jian Wu

One of the main problems that emerges in the classic approach to semantics is the difficulty in acquisition and maintenance of ontologies and semantic annotations. On the other hand, the Internet explosion and the massive diffusion of…

Artificial Intelligence · Computer Science 2017-01-12 Valentina Franzoni

The ability to automatically identify whether an entity is referenced in a future context can have multiple applications including decision making, planning and trend forecasting. This paper focuses on detecting implicit future references…

Computation and Language · Computer Science 2025-02-24 Puneet Prashar , Krishna Mohan Shukla , Adam Jatowt

We consider a resource-constrained updater, such as Google Scholar, which wishes to update the citation records of a group of researchers, who have different mean citation rates (and optionally, different importance coefficients), in such a…

Information Theory · Computer Science 2020-04-28 Melih Bastopcu , Sennur Ulukus

The Classification Literature Automated Search Service, an annual bibliography based on citation of one or more of a set of around 80 book or journal publications, ran from 1972 to 2012. We analyze here the years 1994 to 2011. The…

Digital Libraries · Computer Science 2013-08-20 Fionn Murtagh , Michael J. Kurtz
‹ Prev 1 4 5 6 7 8 10 Next ›