English
Related papers

Related papers: The ACL OCL Corpus: Advancing Open Science in Comp…

200 papers

The recent surge of language models (LMs) has rapidly expanded NLP/AI research, driving an exponential rise in submissions and acceptances at major conferences. Yet this growth has been shadowed by escalating concerns over conference…

Digital Libraries · Computer Science 2025-12-25 Jianglin Ma , Ben Yao , Xiang Li , Yazhou Zhang

PubMed-OCR is an OCR-centric corpus of scientific articles derived from PubMed Central Open Access PDFs. Each page image is annotated with Google Cloud Vision and released in a compact JSON schema with word-, line-, and paragraph-level…

Computer Vision and Pattern Recognition · Computer Science 2026-01-19 Hunter Heidenreich , Yosheb Getachew , Olivia Dinica , Ben Elliott

Citation information in scholarly data is an important source of insight into the reception of publications and the scholarly discourse. Outcomes of citation analyses and the applicability of citation based machine learning approaches…

Digital Libraries · Computer Science 2022-01-12 Tarek Saier , Michael Färber , Tornike Tsereteli

We describe a rule-based approach for the automatic acquisition of salient scientific entities from Computational Linguistics (CL) scholarly article titles. Two observations motivated the approach: (i) noting salient aspects of an article's…

Information Retrieval · Computer Science 2021-09-20 Jennifer D'Souza , Soeren Auer

In pace with developments in the research field of artificial intelligence, knowledge graphs (KGs) have attracted a surge of interest from both academia and industry. As a representation of semantic relations between entities, KGs have…

Computation and Language · Computer Science 2022-10-04 Phillip Schneider , Tim Schopf , Juraj Vladika , Mikhail Galkin , Elena Simperl , Florian Matthes

The CL-SciSumm Shared Task is the first medium-scale shared task on scientific document summarization in the computational linguistics~(CL) domain. In 2019, it comprised three tasks: (1A) identifying relationships between citing documents…

Computation and Language · Computer Science 2019-07-24 Muthu Kumar Chandrasekaran , Michihiro Yasunaga , Dragomir Radev , Dayne Freitag , Min-Yen Kan

Since late 2022, Large Language Models (LLMs) have become very prominent with LLMs like ChatGPT and Bard receiving millions of users. Hundreds of new LLMs are announced each week, many of which are deposited to Hugging Face, a repository of…

Digital Libraries · Computer Science 2023-07-20 Sarah Gao , Andrew Kean Gao

Online Continual Learning (OCL) studies learning over a continuous data stream without observing any single example more than once, a setting that is closer to the experience of humans and systems that must learn "on-the-wild". Yet,…

Computation and Language · Computer Science 2021-02-02 Germán Kruszewski , Ionut-Teodor Sorodoc , Tomas Mikolov

Cross-lingual document alignment aims to identify pairs of documents in two distinct languages that are of comparable content or translations of each other. In this paper, we exploit the signals embedded in URLs to label web documents at…

Computation and Language · Computer Science 2020-10-13 Ahmed El-Kishky , Vishrav Chaudhary , Francisco Guzman , Philipp Koehn

Diagrams are often used in scholarly communication. We analyse a corpus of diagrams found in scholarly computational linguistics conference proceedings (ACL 2017), and find inclusion of a system diagram to be correlated with higher numbers…

Digital Libraries · Computer Science 2022-11-22 Guy Clarke Marshall , Caroline Jay , Andre Freitas

Our languages are in constant flux driven by external factors such as cultural, societal and technological changes, as well as by only partially understood internal motivations. Words acquire new meanings and lose old senses, new words are…

Computation and Language · Computer Science 2019-03-14 Nina Tahmasebi , Lars Borin , Adam Jatowt

Identifying cross-language plagiarism is challenging, especially for distant language pairs and sense-for-sense translations. We introduce the new multilingual retrieval model Cross-Language Ontology-Based Similarity Analysis (CL-OSA) for…

Computation and Language · Computer Science 2021-12-17 Johannes Stegmüller , Fabian Bauer-Marquart , Norman Meuschke , Terry Ruas , Moritz Schubotz , Bela Gipp

Single-cell omics technologies have transformed our understanding of cellular diversity by enabling high-resolution profiling of individual cells. However, the unprecedented scale and heterogeneity of these datasets demand robust frameworks…

This paper proposes OCR++, an open-source framework designed for a variety of information extraction tasks from scholarly articles including metadata (title, author names, affiliation and e-mail), structure (section headings and body text,…

This chapter demonstrates how computational social science (CSS) tools are extending and expanding research on aging. The depth and context from traditionally qualitative methods such as participant observation, in-depth interviews, and…

Computers and Society · Computer Science 2026-01-15 Corey M. Abramson

With the enrichment of literature resources, researchers are facing the growing problem of information explosion and knowledge overload. To help scholars retrieve literature and acquire knowledge successfully, clarifying the semantic…

Computation and Language · Computer Science 2021-12-03 Bowen Ma , Chengzhi Zhang , Yuzhuo Wang , Sanhong Deng

We present DepCC, the largest-to-date linguistically analyzed corpus in English including 365 million documents, composed of 252 billion tokens and 7.5 billion of named entity occurrences in 14.3 billion sentences from a web-scale crawl of…

Computation and Language · Computer Science 2018-03-01 Alexander Panchenko , Eugen Ruppert , Stefano Faralli , Simone Paolo Ponzetto , Chris Biemann

We critically assess mainstream accounting and finance research applying methods from computational linguistics (CL) to study financial discourse. We also review common themes and innovations in the literature and assess the incremental…

Computation and Language · Computer Science 2019-04-01 Mahmoud El-Haj , Paul Rayson , Martin Walker , Steven Young , Vasiliki Simaki

We introduce the Cambridge Law Corpus (CLC), a dataset for legal AI research. It consists of over 250 000 court cases from the UK. Most cases are from the 21st century, but the corpus includes cases as old as the 16th century. This paper…

Computation and Language · Computer Science 2024-01-03 Andreas Östling , Holli Sargeant , Huiyuan Xie , Ludwig Bull , Alexander Terenin , Leif Jonsson , Måns Magnusson , Felix Steffek

In this study we analyzed a corpus of 8 million words academic literature from Computational lingustics' academic literature. the lexical bundles from this corpus are categorized based on structures and functions.

Computation and Language · Computer Science 2016-03-11 Adel Rahimi