English
Related papers

Related papers: S2ORC: The Semantic Scholar Open Research Corpus

200 papers

Understanding the role of citations is essential for research assessment and citation-aware digital libraries. However, existing citation classification frameworks often conflate citation intent (why a work is cited) with cited content type…

Digital Libraries · Computer Science 2026-01-09 Changxu Duan , Zhiyin Tan

In the scientific digital libraries, some papers from different research communities can be described by community-dependent keywords even if they share a semantically similar topic. Articles that are not tagged with enough keyword…

Digital Libraries · Computer Science 2018-06-22 Hussein T. Al-Natsheh , Lucie Martinet , Fabrice Muhlenbach , Fabien Rico , Djamel A. Zighed

Previous work for text summarization in scientific domain mainly focused on the content of the input document, but seldom considering its citation network. However, scientific papers are full of uncommon domain-specific terms, making it…

Computation and Language · Computer Science 2021-04-08 Chenxin An , Ming Zhong , Yiran Chen , Danqing Wang , Xipeng Qiu , Xuanjing Huang

The study of highly cited documents on Google Scholar (GS) has never been addressed to date in a comprehensive manner. The objective of this work is to identify the set of highly cited documents in Google Scholar and define their core…

The present comparative study examines the three main multidisciplinary bibliographic databases, Web of Science Core Collection, Scopus, and OpenAlex, with the aim of providing up-to-date evidence on coverage, metadata quality, and…

Digital Libraries · Computer Science 2026-01-30 Daniel Torres-Salinas , Wenceslao Arroyo-Machado

The number of scientific papers grows exponentially in many disciplines. The share of online available papers grows as well. At the same time, the period of time for a paper to loose at chance to be cited anymore shortens. The decay of the…

Computation and Language · Computer Science 2016-01-11 Rustam Tagiew

An effective structure helps an article to convey its core message. The optimal structure depends on the information to be conveyed and the expectations of the audience. In the current increasingly interdisciplinary era, structural norms…

Digital Libraries · Computer Science 2020-08-10 Mike Thelwall

The Teacher-Student Chatroom Corpus (TSCC) is a collection of written conversations captured during one-to-one lessons between teachers and learners of English. The lessons took place in an online chatroom and therefore involve more…

Computation and Language · Computer Science 2020-11-17 Andrew Caines , Helen Yannakoudakis , Helena Edmondson , Helen Allen , Pascual Pérez-Paredes , Bill Byrne , Paula Buttery

Classic Topic Models are built under the Bag Of Words assumption, in which word position is ignored for simplicity. Besides, symmetric priors are typically used in most applications. In order to easily learn topics with different properties…

Computation and Language · Computer Science 2018-06-27 Simón Roca-Sotelo , Jerónimo Arenas-García

Generating texts in scientific papers requires not only capturing the content contained within the given input but also frequently acquiring the external information called \textit{context}. We push forward the scientific text generation by…

Computation and Language · Computer Science 2021-10-22 Hong Chen , Hiroya Takamura , Hideki Nakayama

We introduce a corpus of 7,032 sentences rated by human annotators for formality, informativeness, and implicature on a 1-7 scale. The corpus was annotated using Amazon Mechanical Turk. Reliability in the obtained judgments was examined by…

Computation and Language · Computer Science 2016-09-29 Shibamouli Lahiri

This paper introduces a new open-source speech corpus named "speechocean762" designed for pronunciation assessment use, consisting of 5000 English utterances from 250 non-native speakers, where half of the speakers are children. Five…

Computation and Language · Computer Science 2021-06-03 Junbo Zhang , Zhiwen Zhang , Yongqing Wang , Zhiyong Yan , Qiong Song , Yukai Huang , Ke Li , Daniel Povey , Yujun Wang

We are presenting a set of multilingual text analysis tools that can help analysts in any field to explore large document collections quickly in order to determine whether the documents contain information of interest, and to find the…

Computation and Language · Computer Science 2007-05-23 Camelia Ignat , Bruno Pouliquen , Ralf Steinberger , Tomaz Erjavec

We use an information-theoretic measure of linguistic similarity to investigate the organization and evolution of scientific fields. An analysis of almost 20M papers from the past three decades reveals that the linguistic similarity is…

Digital Libraries · Computer Science 2018-01-30 Laercio Dias , Martin Gerlach , Joachim Scharloth , Eduardo G. Altmann

How to efficiently generate an accurate, well-structured overview report (ORPT) over thousands of related documents is challenging. A well-structured ORPT consists of sections of multiple levels (e.g., sections and subsections). None of the…

Computation and Language · Computer Science 2019-08-20 Jingwen Wang , Hao Zhang , Cheng Zhang , Wenjing Yang , Liqun Shao , Jie Wang

The rapid growth of open-access (OA) publications has intensified the challenge of identifying relevant scientific papers. Due to privacy constraints and limited access to user interaction data, recent efforts have shifted toward…

Information Retrieval · Computer Science 2025-11-06 Shenghua Wang , Zhen Yin

The rapid acceleration of scientific publishing has created substantial challenges for researchers attempting to discover, contextualize, and interpret relevant literature. Traditional keyword-based search systems provide limited semantic…

Information Retrieval · Computer Science 2025-12-16 Sina Jani , Arman Heidari , Amirmohammad Anvari , Zahra Rahimi

Many useful tasks on scientific documents, such as topic classification and citation prediction, involve corpora that span multiple scientific domains. Typically, such tasks are accomplished by representing the text with a vector embedding…

Computation and Language · Computer Science 2023-09-11 Ronald Seoh , Haw-Shiuan Chang , Andrew McCallum

In this paper, we present a corpus for use in automatic readability assessment and automatic text simplification of German. The corpus is compiled from web sources and consists of approximately 211,000 sentences. As a novel contribution, it…

Computation and Language · Computer Science 2019-09-20 Alessia Battisti , Sarah Ebling

Open set classification (OSC) tackles the problem of determining whether the data are in-class or out-of-class during inference, when only provided with a set of in-class examples at training time. Traditional OSC methods usually train…

Machine Learning · Computer Science 2020-08-12 Yang Yang , Zhen-Qiang Sun , Hui Xiong , Jian Yang