English
Related papers

Related papers: S2ORC: The Semantic Scholar Open Research Corpus

200 papers

Systems that can automatically define unfamiliar terms hold the promise of improving the accessibility of scientific texts, especially for readers who may lack prerequisite background knowledge. However, current systems assume a single…

With over 200 million published academic documents and millions of new documents being written each year, academic researchers face the challenge of searching for information within this vast corpus. However, existing retrieval systems…

Information Retrieval · Computer Science 2024-05-21 Gengchen Wei , Xinle Pang , Tianning Zhang , Yu Sun , Xun Qian , Chen Lin , Han-Sen Zhong , Wanli Ouyang

Scholarly documents have a great degree of variation, both in terms of content (semantics) and structure (pragmatics). Prior work in scholarly document understanding emphasizes semantics through document summarization and corpus topic…

Computation and Language · Computer Science 2023-10-03 Lee Kezar , Jay Pujara

Scientometric predictors of research performance need to be validated by showing that they have a high correlation with the external criterion they are trying to predict. The UK Research Assessment Exercise (RAE), together with the growing…

Information Retrieval · Computer Science 2007-05-23 Stevan Harnad

The wide adoption of electronic health records (EHRs) has enabled a wide range of applications leveraging EHR data. However, the meaningful use of EHR data largely depends on our ability to efficiently extract and consolidate information…

Information Retrieval · Computer Science 2018-08-29 Yanshan Wang , Naveed Afzal , Sunyang Fu , Liwei Wang , Feichen Shen , Majid Rastegar-Mojarad , Hongfang Liu

Materials science literature contains millions of materials synthesis procedures described in unstructured natural language text. Large-scale analysis of these synthesis procedures would facilitate deeper scientific understanding of…

Computation and Language · Computer Science 2019-07-16 Sheshera Mysore , Zach Jensen , Edward Kim , Kevin Huang , Haw-Shiuan Chang , Emma Strubell , Jeffrey Flanigan , Andrew McCallum , Elsa Olivetti

We introduce MCScript2.0, a machine comprehension corpus for the end-to-end evaluation of script knowledge. MCScript2.0 contains approx. 20,000 questions on approx. 3,500 texts, crowdsourced based on a new collection process that results in…

Computation and Language · Computer Science 2019-05-31 Simon Ostermann , Michael Roth , Manfred Pinkal

We present the HPLT (High Performance Language Technologies) language resources, a new massive multilingual dataset including both monolingual and bilingual corpora extracted from CommonCrawl and previously unused web crawls from the…

Analysis of acknowledgments is particularly interesting as acknowledgments may give information not only about funding, but they are also able to reveal hidden contributions to authorship and the researcher's collaboration patterns, context…

Digital Libraries · Computer Science 2023-05-26 Nina Smirnova , Philipp Mayr

Scientific abstracts contain what is considered by the author(s) as information that best describe documents' content. They represent a compressed view of the informational content of a document and allow readers to evaluate the relevance…

Information Retrieval · Computer Science 2016-04-12 Iana Atanassova , Marc Bertin , Vincent Larivière

Large-scale data sets on scholarly publications are the basis for a variety of bibliometric analyses and natural language processing (NLP) applications. Especially data sets derived from publication's full-text have recently gained…

Digital Libraries · Computer Science 2023-11-06 Tarek Saier , Johan Krause , Michael Färber

In the biotechnology and biomedical domains, recent text mining efforts advocate for machine-interpretable, and preferably, semantified, documentation formats of laboratory processes. This includes wet-lab protocols, (in)organic materials…

Digital Libraries · Computer Science 2020-09-17 Marco Anteghini , Jennifer D'Souza , Vitor A. P. Martins dos Santos , Sören Auer

The data article presents the large bilingual parallel corpus of low-resourced language pair Sanskrit-Hindi, named SAHAAYAK 2023. The corpus contains total of 1.5M sentence pairs between Sanskrit and Hindi. To make the universal usability…

Computation and Language · Computer Science 2023-07-04 Vishvajitsinh Bakrola , Jitendra Nasariwala

We are faced with an unprecedented production in scholarly publications worldwide. Stakeholders in the digital libraries posit that the document-based publishing paradigm has reached the limits of adequacy. Instead, structured,…

Computation and Language · Computer Science 2022-05-25 Jennifer D'Souza

In the recent years, transformer-based models have lead to significant advances in language modelling for natural language processing. However, they require a vast amount of data to be (pre-)trained and there is a lack of corpora in…

Computation and Language · Computer Science 2022-07-04 Asier Gutiérrez-Fandiño , David Pérez-Fernández , Jordi Armengol-Estapé , David Griol , Zoraida Callejas

Texts and their translations are a rich linguistic resource that can be used to train and test statistics-based Machine Translation systems and many other applications. In this paper, we present a working system that can identify…

Computation and Language · Computer Science 2007-05-23 Bruno Pouliquen , Ralf Steinberger , Camelia Ignat

Taxonomies and ontologies of research topics (e.g., MeSH, UMLS, CSO, NLM) play a central role in providing the primary framework through which intelligent systems can explore and interpret the literature. However, these resources have…

Digital Libraries · Computer Science 2025-08-07 Alessia Pisu , Livio Pompianu , Francesco Osborne , Diego Reforgiato Recupero , Daniele Riboni , Angelo Salatino

Anonymous peer review is used by the great majority of computer science conferences. OpenReview is such a platform that aims to promote openness in peer review process. The paper, (meta) reviews, rebuttals, and final decisions are all…

Digital Libraries · Computer Science 2021-04-07 Gang Wang , Qi Peng , Yanfeng Zhang , Mingyang Zhang

The Universal Knowledge Core (UKC) is a large multilingual lexical database with a focus on language diversity and covering over a thousand languages. The aim of the database, as well as its tools and data catalogue, is to make the somewhat…

In text analysis, Spherical K-means (SKM) is a specialized k-means clustering algorithm widely utilized for grouping documents represented in high-dimensional, sparse term-document matrices, often normalized using techniques like TF-IDF.…

Methodology · Statistics 2025-02-25 Ilaria Bombelli , Domenica Fioredistella Iezzi , Emiliano Seri , Maurizio Vichi