English
Related papers

Related papers: S2ORC: The Semantic Scholar Open Research Corpus

200 papers

Bibliometrics, whether used for research or research evaluation, relies on large multidisciplinary databases of research outputs and citation indices. The Web of Science (WoS) was the main supporting infrastructure of the field for more…

Digital Libraries · Computer Science 2025-08-27 Philippe Mongeon , Madelaine Hare , Poppy Riddle , Summer Wilson , Geoff Krause , Rebecca Marjoram , Rémi Toupin

Logic Mill is a scalable and openly accessible software system that identifies semantically similar documents within either one domain-specific corpus or multi-domain corpora. It uses advanced Natural Language Processing (NLP) techniques to…

Computation and Language · Computer Science 2024-10-14 Sebastian Erhardt , Mainak Ghosh , Erik Buunk , Michael E. Rose , Dietmar Harhoff

Clarivate's Web of Science (WoS) and Elsevier's Scopus have been for decades the main sources of bibliometric information. Although highly curated, these closed, proprietary databases are largely biased towards English-language…

This study explores the extent to which bibliometric indicators based on counts of highly-cited documents could be affected by the choice of data source. The initial hypothesis is that databases that rely on journal selection criteria for…

Digital Libraries · Computer Science 2018-11-20 Alberto Martín-Martín , Enrique Orduna-Malea , Emilio Delgado López-Cózar

A scientific paper can be divided into two major constructs which are Metadata and Full-body text. Metadata provides a brief overview of the paper while the Full-body text contains key-insights that can be valuable to fellow researchers. To…

Digital Libraries · Computer Science 2023-08-28 Azanzi Jiomekong , Sanju Tiwari

Sanskrit is a classical language with about 30 million extant manuscripts fit for digitisation, available in written, printed or scannedimage forms. However, it is still considered to be a low-resource language when it comes to available…

Computation and Language · Computer Science 2022-11-16 Ayush Maheshwari , Nikhil Singh , Amrith Krishna , Ganesh Ramakrishnan

Learner corpus collects language data produced by L2 learners, that is second or foreign-language learners. This resource is of great relevance for second language acquisition research, foreign-language teaching, and automatic grammatical…

Computation and Language · Computer Science 2022-01-03 Yingying Wang , Cunliang Kong , Liner Yang , Yijun Wang , Xiaorong Lu , Renfen Hu , Shan He , Zhenghao Liu , Yun Chen , Erhong Yang , Maosong Sun

This article presents a hybrid methodology for building a multilingual corpus designed to support the study of emerging concepts in the humanities and social sciences (HSS), illustrated here through the case of ``non-technological…

Computation and Language · Computer Science 2025-12-09 Revekka Kyriakoglou , Anna Pappa

CODEC is a document and entity ranking benchmark that focuses on complex research topics. We target essay-style information needs of social science researchers, i.e. "How has the UK's Open Banking Regulation benefited Challenger Banks?".…

Information Retrieval · Computer Science 2022-05-18 Iain Mackie , Paul Owoicho , Carlos Gemmell , Sophie Fischer , Sean MacAvaney , Jeffrey Dalton

The need for raw large raw corpora has dramatically increased in recent years with the introduction of transfer learning and semi-supervised learning methods to Natural Language Processing. And while there have been some recent attempts to…

Computation and Language · Computer Science 2022-01-19 Julien Abadji , Pedro Ortiz Suarez , Laurent Romary , Benoît Sagot

Deep learning based natural language processing model is proven powerful, but need large-scale dataset. Due to the significant gap between the real-world tasks and existing Chinese corpus, in this paper, we introduce a large-scale corpus of…

Computation and Language · Computer Science 2018-11-27 Jianyu Zhao , Zhuoran Ji

Determining semantic similarity between academic documents is crucial to many tasks such as plagiarism detection, automatic technical survey and semantic search. Current studies mostly focus on semantic similarity between concepts,…

Computation and Language · Computer Science 2017-12-01 Ming Liu , Bo Lang , Zepeng Gu

This paper describes a corpus of about 3000 English literary texts with about 250 million words extracted from the Gutenberg project that span a range of genres from both fiction and non-fiction written by more than 130 authors (e.g.,…

Computation and Language · Computer Science 2018-01-09 Arthur M. Jacobs

The Scielo database is an important source of scientific information in Latin America, containing articles from several research domains. A striking characteristic of Scielo is that many of its full-text contents are presented in more than…

Computation and Language · Computer Science 2019-05-07 Felipe Soares , Viviane Pereira Moreira , Karin Becker

This article uses Google Scholar (GS) as a source of data to analyse Open Access (OA) levels across all countries and fields of research. All articles and reviews with a DOI and published in 2009 or 2014 and covered by the three main…

Digital Libraries · Computer Science 2018-07-26 Alberto Martín-Martín , Rodrigo Costas , Thed van Leeuwen , Emilio Delgado López-Cózar

The large-scale analysis of scholarly artifact usage is constrained primarily by current practices in usage data archiving, privacy issues concerned with the dissemination of usage data, and the lack of a practical ontology for modeling the…

Digital Libraries · Computer Science 2007-08-09 Marko A. Rodriguez , Johah Bollen , Herbert Van de Sompel

"Open access" has become a central theme of journal reform in academic publishing. In this article, I examine the relationship between open access publishing and an important infrastructural element of a modern research enterprise,…

Computers and Society · Computer Science 2018-04-25 Gopal P. Sarma

Comprehensively evaluating and comparing researchers' academic performance is complicated due to the intrinsic complexity of scholarly data. Different scholarly evaluation tasks often require the publication and citation data to be…

Graphics · Computer Science 2022-03-25 Zhichun Guo , Jun Tao , Siming Chen , Nitesh V. Chawla , Chaoli Wang

We present the SAMER Corpus, the first manually annotated Arabic parallel corpus for text simplification targeting school-aged learners. Our corpus comprises texts of 159K words selected from 15 publicly available Arabic fiction novels most…

Computation and Language · Computer Science 2024-04-30 Bashar Alhafni , Reem Hazim , Juan Piñeros Liberato , Muhamed Al Khalil , Nizar Habash

Scientific article summarization is challenging: large, annotated corpora are not available, and the summary should ideally include the article's impacts on research community. This paper provides novel solutions to these two challenges. We…

Computation and Language · Computer Science 2019-09-17 Michihiro Yasunaga , Jungo Kasai , Rui Zhang , Alexander R. Fabbri , Irene Li , Dan Friedman , Dragomir R. Radev
‹ Prev 1 3 4 5 6 7 10 Next ›