English
Related papers

Related papers: RuCoCo: a new Russian corpus with coreference anno…

200 papers

Popularized by the Differentiable Search Index, the emerging paradigm of generative retrieval re-frames the classic information retrieval problem into a sequence-to-sequence modeling task, forgoing external indices and encoding an entire…

Information Retrieval · Computer Science 2023-05-22 Ronak Pradeep , Kai Hui , Jai Gupta , Adam D. Lelkes , Honglei Zhuang , Jimmy Lin , Donald Metzler , Vinh Q. Tran

Distributed representations of words as real-valued vectors in a relatively low-dimensional space aim at extracting syntactic and semantic features from large text corpora. A recently introduced neural network, named word2vec (Mikolov et…

Computation and Language · Computer Science 2015-08-11 Adriaan M. J. Schakel , Benjamin J. Wilson

This paper presents a new selection-based question answering dataset, SelQA. The dataset consists of questions generated through crowdsourcing and sentence length answers that are drawn from the ten most prevalent topics in the English…

Computation and Language · Computer Science 2016-10-31 Tomasz Jurczyk , Michael Zhai , Jinho D. Choi

This paper presents a scheme for annotating coreference across news articles, extending beyond traditional identity relations by also considering near-identity and bridging relations. It includes a precise description of how to set up…

Computation and Language · Computer Science 2023-10-19 Jakob Vogel

We introduce Balalaika, an open-source, data-centric pipeline for processing audio and producing prosody-aware annotations. It combines semantic VAD for context-preserving segmentation, multi-ASR ensembling with ROVER consensus decoding,…

Computation and Language · Computer Science 2026-05-05 Kirill Borodin , Nikita Vasiliev , Vasiliy Kudryavtsev , Maxim Maslov , Mikhail Gorodnichev , Grach Mkrtchian

The Novelties corpus is a collection of novels (and parts of novels) annotated for Alias Resolution, among other tasks. This document describes the guidelines applied during the annotation process. It contains the instructions used by the…

Computation and Language · Computer Science 2024-10-02 Arthur Amalvy , Vincent Labatut

Taxonomies play a crucial role in various applications by providing a structural representation of knowledge. The task of taxonomy expansion involves integrating emerging concepts into existing taxonomies by identifying appropriate parent…

Computation and Language · Computer Science 2025-05-27 Qingkai Zeng , Yuyang Bai , Zhaoxuan Tan , Zhenyu Wu , Shangbin Feng , Meng Jiang

Current approaches to automatic summarization of scientific papers generate informative summaries in the form of abstracts. However, abstracts are not intended to show the relationship between a paper and the references cited in it. We…

Computation and Language · Computer Science 2023-11-14 Shahbaz Syed , Ahmad Dawar Hakimi , Khalid Al-Khatib , Martin Potthast

We present ShortScience.org, a platform for post-publication discussion of research papers. On ShortScience.org, the research community can read and write summaries of papers in order to increase accessible and reproducibility. Summaries…

Digital Libraries · Computer Science 2017-07-24 Joseph Paul Cohen , Henry Z. Lo

We present the Prepositions Annotated with Supersense Tags in Reddit International English ("PASTRIE") corpus, a new dataset containing manually annotated preposition supersenses of English data from presumed speakers of four L1s: English,…

Computation and Language · Computer Science 2021-10-26 Michael Kranzlein , Emma Manning , Siyao Peng , Shira Wein , Aryaman Arora , Bradford Salen , Nathan Schneider

The abundance of predicted and mined but uncertain biological data show huge needs for massive, efficient and scalable curation efforts. The human expertise warranted by any successful curation enterprize is often economically prohibitive…

Databases · Computer Science 2016-06-08 Hasan M. Jamil , Fereidoon Sadri

Reference resolution on extended texts (several thousand references) cannot be evaluated manually. An evaluation algorithm has been proposed for the MUC tests, using equivalence classes for the coreference relation. However, we show here…

Computation and Language · Computer Science 2007-05-23 Andrei Popescu-Belis , Isabelle Robba

This paper describes NEREL-BIO -- an annotation scheme and corpus of PubMed abstracts in Russian and smaller number of abstracts in English. NEREL-BIO extends the general domain dataset NEREL by introducing domain-specific entity types.…

When speaking or writing, people omit information that seems clear and evident, such that only part of the message is expressed in words. Especially in argumentative texts it is very common that (important) parts of the argument are implied…

Computation and Language · Computer Science 2019-12-24 Maria Becker , Katharina Korfhage , Anette Frank

The development of large and super-large language models, such as GPT-3, T5, Switch Transformer, ERNIE, etc., has significantly improved the performance of text generation. One of the important research directions in this area is the…

Computation and Language · Computer Science 2022-06-22 Irina Fishcheva , Dmitriy Osadchiy , Klavdiya Bochenina , Evgeny Kotelnikov

We introduce a data set called DCH-2, which contains 4,390 real customer-helpdesk dialogues in Chinese and their English translations. DCH-2 also contains dialogue-level annotations and turn-level annotations obtained independently from…

Computation and Language · Computer Science 2021-06-01 Zhaohao Zeng , Tetsuya Sakai

In spite of the development of content-based data management, text-based searching remains the primary means of multimedia retrieval in many areas. Automatic creation of text metadata is thus a crucial tool for increasing the findability of…

Information Retrieval · Computer Science 2014-12-19 Jan Botorek , Petra Budikova , Pavel Zezula

We introduce the Speak & Improve Corpus 2025, a dataset of L2 learner English data with holistic scores and language error annotation, collected from open (spontaneous) speaking tests on the Speak & Improve learning platform. The aim of the…

Computation and Language · Computer Science 2024-12-18 Kate Knill , Diane Nicholls , Mark J. F. Gales , Mengjie Qian , Pawel Stroinski

Deep learning based natural language processing model is proven powerful, but need large-scale dataset. Due to the significant gap between the real-world tasks and existing Chinese corpus, in this paper, we introduce a large-scale corpus of…

Computation and Language · Computer Science 2018-11-27 Jianyu Zhao , Zhuoran Ji

Despite biographies are widely spread within the Semantic Web, resources and approaches to automatically extract biographical events are limited. Such limitation reduces the amount of structured, machine-readable biographical information,…

Computation and Language · Computer Science 2022-06-09 Marco Antonio Stranisci , Enrico Mensa , Ousmane Diakite , Daniele Radicioni , Rossana Damiano