中文
相关论文

相关论文: Elsevier OA CC-By Corpus

200 篇论文

The rapidly growing volume of scientific publications offers an interesting challenge for research on methods for analyzing the authorship of documents with one or more authors. However, most existing datasets lack scientific documents or…

计算与语言 · 计算机科学 2023-05-11 Janek Bevendorff , Philipp Sauer , Lukas Gienapp , Wolfgang Kircheis , Erik Körner , Benno Stein , Martin Potthast

We introduce S2ORC, a large corpus of 81.1M English-language academic papers spanning many academic disciplines. The corpus consists of rich metadata, paper abstracts, resolved bibliographic references, as well as structured full text for…

计算与语言 · 计算机科学 2020-07-08 Kyle Lo , Lucy Lu Wang , Mark Neumann , Rodney Kinney , Dan S. Weld

PubMed-OCR is an OCR-centric corpus of scientific articles derived from PubMed Central Open Access PDFs. Each page image is annotated with Google Cloud Vision and released in a compact JSON schema with word-, line-, and paragraph-level…

计算机视觉与模式识别 · 计算机科学 2026-01-19 Hunter Heidenreich , Yosheb Getachew , Olivia Dinica , Ben Elliott

This paper proposes OCR++, an open-source framework designed for a variety of information extraction tasks from scholarly articles including metadata (title, author names, affiliation and e-mail), structure (section headings and body text,…

This article uses Google Scholar (GS) as a source of data to analyse Open Access (OA) levels across all countries and fields of research. All articles and reviews with a DOI and published in 2009 or 2014 and covered by the three main…

数字图书馆 · 计算机科学 2018-07-26 Alberto Martín-Martín , Rodrigo Costas , Thed van Leeuwen , Emilio Delgado López-Cózar

Scientific literature serves as a high-quality corpus, supporting a lot of Natural Language Processing (NLP) research. However, existing datasets are centered around the English language, which restricts the development of Chinese…

计算与语言 · 计算机科学 2022-09-13 Yudong Li , Yuqing Zhang , Zhe Zhao , Linlin Shen , Weijie Liu , Weiquan Mao , Hui Zhang

We introduce the Cambridge Law Corpus (CLC), a dataset for legal AI research. It consists of over 250 000 court cases from the UK. Most cases are from the 21st century, but the corpus includes cases as old as the 16th century. This paper…

The research content hosted by arXiv is not fully accessible to everyone due to disabilities and other barriers. This matters because a significant proportion of people have reading and visual disabilities, it is important to our community…

We present a dataset of 833k paragraphs extracted from CC-BY licensed scientific publications, classified into four categories: acknowledgments, data mentions, software/code mentions, and clinical trial mentions. The paragraphs are…

计算与语言 · 计算机科学 2025-10-28 Eric Jeangirard

A multitude of factors are responsible for the overall quality of scientific papers, including readability, linguistic quality, fluency,semantic complexity, and of course domain-specific technical factors. These factors vary from one field…

信息检索 · 计算机科学 2019-08-13 Roman Vainshtein , Gilad Katz , Bracha Shapira , Lior Rokach

Open information extraction (OIE) systems extract relations and their arguments from natural language text in an unsupervised manner. The resulting extractions are a valuable resource for downstream tasks such as knowledge base…

计算与语言 · 计算机科学 2019-04-30 Kiril Gashteovski , Sebastian Wanner , Sven Hertling , Samuel Broscheit , Rainer Gemulla

The EcoLexicon English Corpus (EEC) is a 23.1-million-word corpus of contemporary environmental texts. It was compiled by the LexiCon research group for the development of EcoLexicon (Faber, Leon-Arauz & Reimerink 2016; San Martin et al.…

计算与语言 · 计算机科学 2018-07-17 Pilar Leon-Arauz , Antonio San Martin , Arianne Reimerink

The ongoing paradigm change in the scholarly publication system ('science is turning to e-science') makes it necessary to construct alternative evaluation criteria/metrics which appropriately take into account the unique characteristics of…

数字图书馆 · 计算机科学 2019-01-15 Philipp Mayr

With the growth of open access (OA), the financial flows in scholarly journal publishing have become increasingly complex, but comprehensive data and transparency into these flows are still lacking. The opaqueness is especially concerning…

数字图书馆 · 计算机科学 2024-02-28 Najko Jahn , Lisa Matthias , Mikael Laakso

Open Information Extraction (OIE) is the task of the unsupervised creation of structured information from text. OIE is often used as a starting point for a number of downstream tasks including knowledge base construction, relation…

计算与语言 · 计算机科学 2018-08-23 Paul Groth , Michael Lauruhn , Antony Scerri , Ron Daniel

The OSCOSS project (Opening Scholarly Communication in Social Sciences), which will be outlined, aims at providing integrated support for all steps of the scholarly communication process. Incl. collaborative writing of a scientific paper,…

数字图书馆 · 计算机科学 2017-10-19 Philipp Mayr , Christoph Lange

A scientific paper can be divided into two major constructs which are Metadata and Full-body text. Metadata provides a brief overview of the paper while the Full-body text contains key-insights that can be valuable to fellow researchers. To…

数字图书馆 · 计算机科学 2023-08-28 Azanzi Jiomekong , Sanju Tiwari

Authorship of scientific articles has profoundly changed from early science until now. While once upon a time a paper was authored by a handful of authors, scientific collaborations are much more prominent on average nowadays. As authorship…

数字图书馆 · 计算机科学 2022-09-16 Andrea Mannocci , Ornella Irrera , Paolo Manghi

We present ACL OCL, a scholarly corpus derived from the ACL Anthology to assist Open scientific research in the Computational Linguistics domain. Integrating and enhancing the previous versions of the ACL Anthology, the ACL OCL contributes…

计算与语言 · 计算机科学 2023-10-25 Shaurya Rohatgi , Yanxia Qin , Benjamin Aw , Niranjana Unnithan , Min-Yen Kan

The main objective of this paper is to identify the set of highly-cited documents in Google Scholar and to define their core characteristics (document types, language, free availability, source providers, and number of versions), under the…

数字图书馆 · 计算机科学 2016-12-28 Alberto Martin-Martin , Enrique Orduna-Malea , Juan M. Ayllon , Emilio Delgado Lopez-Cozar
‹ 上一页 1 2 3 10 下一页 ›