English
Related papers

Related papers: S2ORC: The Semantic Scholar Open Research Corpus

200 papers

The main objective of this paper is to identify the set of highly-cited documents in Google Scholar and to define their core characteristics (document types, language, free availability, source providers, and number of versions), under the…

Digital Libraries · Computer Science 2016-12-28 Alberto Martin-Martin , Enrique Orduna-Malea , Juan M. Ayllon , Emilio Delgado Lopez-Cozar

We present SciDMT, an enhanced and expanded corpus for scientific mention detection, offering a significant advancement over existing related resources. SciDMT contains annotated scientific documents for datasets (D), methods (M), and tasks…

Artificial Intelligence · Computer Science 2024-06-24 Huitong Pan , Qi Zhang , Cornelia Caragea , Eduard Dragut , Longin Jan Latecki

The Open Access movement in scientific publishing and search engines like Google Scholar have made scientific articles more broadly accessible. During the last decade, the availability of scientific papers in full text has become more and…

Digital Libraries · Computer Science 2015-05-07 Iana Atanassova , Marc Bertin , Philipp Mayr

Standing at the forefront of knowledge dissemination, digital libraries curate vast collections of scientific literature. However, these scholarly writings are often laden with jargon and tailored for domain experts rather than the general…

Computation and Language · Computer Science 2024-08-08 Haining Wang , Jason Clark

This paper proposes OCR++, an open-source framework designed for a variety of information extraction tasks from scholarly articles including metadata (title, author names, affiliation and e-mail), structure (section headings and body text,…

We present a new release of the Czech-English parallel corpus CzEng 2.0 consisting of over 2 billion words (2 "gigawords") in each language. The corpus contains document-level information and is filtered with several techniques to lower the…

Computation and Language · Computer Science 2020-07-08 Tom Kocmi , Martin Popel , Ondrej Bojar

This paper introduces OARelatedWork, the first large-scale multi-document summarization dataset for related work generation containing whole related work sections and full-texts of cited papers. The dataset includes 94 450 papers and 5 824…

Computation and Language · Computer Science 2024-05-06 Martin Docekal , Martin Fajcik , Pavel Smrz

We introduce the Self-Annotated Reddit Corpus (SARC), a large corpus for sarcasm research and for training and evaluating systems for sarcasm detection. The corpus has 1.3 million sarcastic statements -- 10 times more than any previous…

Computation and Language · Computer Science 2018-03-26 Mikhail Khodak , Nikunj Saunshi , Kiran Vodrahalli

Who has not read letters of recommendations that comment on a student's `broadness' and wondered what to make of it? We here propose a way to quantify scientific broadness by a semantic analysis of researchers' publications. We apply our…

Physics and Society · Physics 2018-08-06 Tom Price , Sabine Hossenfelder

We present a freely available, genre-balanced English web corpus totaling 4M tokens and featuring a large number of high-quality automatic annotation layers, including dependency trees, non-named entity annotations, coreference resolution,…

Computation and Language · Computer Science 2020-06-19 Luke Gessler , Siyao Peng , Yang Liu , Yilun Zhu , Shabnam Behzad , Amir Zeldes

The amount of scholarly data has been increasing dramatically over the last years. For newcomers to a particular science domain (e.g., IR, physics, NLP) it is often difficult to spot larger trends and to position the latest research in the…

Digital Libraries · Computer Science 2021-12-08 Naman Paharia , Muhammad Syafiq Mohd Pozi , Adam Jatowt

We present an overview of the SCIDOCA 2025 Shared Task, which focuses on citation discovery and prediction in scientific documents. The task is divided into three subtasks: (1) Citation Discovery, where systems must identify relevant…

Digital Libraries · Computer Science 2025-09-30 An Dao , Vu Tran , Le-Minh Nguyen , Yuji Matsumoto

Mentorship is a critical component of academia, but is not as visible as publications, citations, grants, and awards. Despite the importance of studying the quality and impact of mentorship, there are few large representative mentorship…

Digital Libraries · Computer Science 2022-05-03 Shaurya Rohatgi , Doug Downey , Daniel King , Sergey Feldman

We introduce the Elsevier OA CC-BY corpus. This is the first open corpus of Scientific Research papers which has a representative sample from across scientific disciplines. This corpus not only includes the full text of the article, but…

Computation and Language · Computer Science 2020-09-16 Daniel Kershaw , Rob Koeling

We present MMOCR-an open-source toolbox which provides a comprehensive pipeline for text detection and recognition, as well as their downstream tasks such as named entity recognition and key information extraction. MMOCR implements 14…

Computer Vision and Pattern Recognition · Computer Science 2021-08-17 Zhanghui Kuang , Hongbin Sun , Zhizhong Li , Xiaoyu Yue , Tsui Hin Lin , Jianyong Chen , Huaqiang Wei , Yiqin Zhu , Tong Gao , Wenwei Zhang , Kai Chen , Wayne Zhang , Dahua Lin

In Natural Language Processing, automatic extracting the meaning of texts constitutes an important problem. Our focus is the computational analysis of meaning of short scientific texts (abstracts or brief reports). In this paper, a vector…

Computation and Language · Computer Science 2020-04-30 Neslihan Suzen , Evgeny M. Mirkes , Alexander N. Gorban

Cross-Lingual SynthDocs is a large-scale synthetic corpus designed to address the scarcity of Arabic resources for Optical Character Recognition (OCR) and Document Understanding (DU). The dataset comprises over 2.5 million of samples,…

Computation and Language · Computer Science 2025-11-10 Haneen Al-Homoud , Asma Ibrahim , Murtadha Al-Jubran , Fahad Al-Otaibi , Yazeed Al-Harbi , Daulet Toibazar , Kesen Wang , Pedro J. Moreno

We introduce the Speak & Improve Corpus 2025, a dataset of L2 learner English data with holistic scores and language error annotation, collected from open (spontaneous) speaking tests on the Speak & Improve learning platform. The aim of the…

Computation and Language · Computer Science 2024-12-18 Kate Knill , Diane Nicholls , Mark J. F. Gales , Mengjie Qian , Pawel Stroinski

We present a novel corpus of 445 human- and computer-generated documents, comprising about 27,000 clauses, annotated for semantic clause types and coherence relations that allow for nuanced comparison of artificial and natural discourse…

We present a new, unique and freely available parallel corpus containing European Union (EU) documents of mostly legal nature. It is available in all 20 official EUanguages, with additional documents being available in the languages of the…

Computation and Language · Computer Science 2007-05-23 Ralf Steinberger , Bruno Pouliquen , Anna Widiger , Camelia Ignat , Tomaz Erjavec , Dan Tufis , Daniel Varga