English
Related papers

Related papers: Explorations in an English Poetry Corpus: A Neuroc…

200 papers

The Gutenberg Literary English Corpus (GLEC) provides a rich source of textual data for research in digital humanities, computational linguistics or neurocognitive poetics. However, so far only a small subcorpus, the Gutenberg English…

Computation and Language · Computer Science 2020-10-22 Arthur M. Jacobs , Annette Kinder

The Gutenberg Literary English Corpus (GLEC, Jacobs, 2018a) provides a rich source of textual data for research in digital humanities, computational linguistics or neurocognitive poetics. In this study we address differences among the…

Computation and Language · Computer Science 2022-01-13 Arthur M. Jacobs , Annette Kinder

The use of Project Gutenberg (PG) as a text corpus has been extremely popular in statistical analysis of language for more than 25 years. However, in contrast to other major linguistic datasets of similar importance, no consensual full…

Computation and Language · Computer Science 2018-12-20 Martin Gerlach , Francesc Font-Clos

We present DepCC, the largest-to-date linguistically analyzed corpus in English including 365 million documents, composed of 252 billion tokens and 7.5 billion of named entity occurrences in 14.3 billion sentences from a web-scale crawl of…

Computation and Language · Computer Science 2018-03-01 Alexander Panchenko , Eugen Ruppert , Stefano Faralli , Simone Paolo Ponzetto , Chris Biemann

We present the Project Dialogism Novel Corpus, or PDNC, an annotated dataset of quotations for English literary texts. PDNC contains annotations for 35,978 quotations across 22 full-length novels, and is by an order of magnitude the largest…

Computation and Language · Computer Science 2022-04-13 Krishnapriya Vishnubhotla , Adam Hammond , Graeme Hirst

I apply Schmidhuber's compression progress theory of interestingness at corpus scale, analyzing semantic novelty trajectories in more than 80,000 books spanning two centuries of English-language publishing. Using sentence-transformer…

Computation and Language · Computer Science 2026-03-03 Fred Zimmerman

This discussion paper reflects on how quantitative approaches to historical linguistics interact with dataset properties. Drawing on two worked examples, we examine English data using quad-based concept modelling of Early Modern English…

Computation and Language · Computer Science 2026-05-05 Catherine Wong , Bach Phan-Tat , Susan Fitzmaurice

In this paper, we try to explore the evolution of language through case calculations. First, we chose the novels of eleven British writers from 1400 to 2005 and found the corresponding works; Then, we use the natural language processing…

Computation and Language · Computer Science 2018-10-09 Zhu Gao , Yanhui Jiang , Junhui Gao

We present a novel corpus of 445 human- and computer-generated documents, comprising about 27,000 clauses, annotated for semantic clause types and coherence relations that allow for nuanced comparison of artificial and natural discourse…

Due to its semantic succinctness and novelty of expression, poetry is a great test bed for semantic change analysis. However, so far there is a scarcity of large diachronic corpora. Here, we provide a large corpus of German poetry which…

Computation and Language · Computer Science 2019-09-27 Thomas Haider , Steffen Eger

A simple method for finding the entropy and redundancy of a reasonable long sample of English text by direct computer processing and from first principles according to Shannon theory is presented. As an example, results on the entropy of…

Computation and Language · Computer Science 2009-11-19 Fabio G. Guerrero

Embedding models group text by semantic content, what text is about. We show that temporal co-occurrence within texts discovers a different kind of structure: recurrent transition-structure concepts or what text does. We train a…

Artificial Intelligence · Computer Science 2026-03-20 Jason Dury

We present in this work a new dataset of coreference annotations for works of literature in English, covering 29,103 mentions in 210,532 tokens from 100 works of fiction. This dataset differs from previous coreference datasets in containing…

Computation and Language · Computer Science 2020-05-18 David Bamman , Olivia Lewke , Anya Mansoor

We present a mechanistic analysis of literary style in GPT-2, identifying individual neurons that discriminate between exemplary prose and rigid AI-generated text. Using Herman Melville's Bartleby, the Scrivener as a corpus, we extract…

Computation and Language · Computer Science 2025-10-22 Tsogt-Ochir Enkhbayar

We introduce S2ORC, a large corpus of 81.1M English-language academic papers spanning many academic disciplines. The corpus consists of rich metadata, paper abstracts, resolved bibliographic references, as well as structured full text for…

Computation and Language · Computer Science 2020-07-08 Kyle Lo , Lucy Lu Wang , Mark Neumann , Rodney Kinney , Dan S. Weld

Entities like person, location, organization are important for literary text analysis. The lack of annotated data hinders the progress of named entity recognition (NER) in literary domain. To promote the research of literary NER, we build…

Computation and Language · Computer Science 2024-10-16 Hanjie Zhao , Jinge Xie , Yuchen Yan , Yuxiang Jia , Yawen Ye , Hongying Zan

This paper introduces the Multi-Genre Natural Language Inference (MultiNLI) corpus, a dataset designed for use in the development and evaluation of machine learning models for sentence understanding. In addition to being one of the largest…

Computation and Language · Computer Science 2018-02-21 Adina Williams , Nikita Nangia , Samuel R. Bowman

Publishing articles in high-impact English journals is difficult for scholars around the world, especially for non-native English-speaking scholars (NNESs), most of whom struggle with proficiency in English. In order to uncover the…

Computation and Language · Computer Science 2018-09-14 Chao Lu , Yi Bu , Jie Wang , Ying Ding , Vetle Torvik , Matthew Schnaars , Chengzhi Zhang

In this study, we investigated the academic literature on quantum technologies (QT) using bibliometric tools. We used a set of 49,823 articles obtained from the Web of Science (WoS) database using a search query constructed through expert…

Digital Libraries · Computer Science 2021-05-12 Zeki C. Seskir , Arsev U. Aydinoglu

In this paper, we present a scientific corpus of abstracts of academic papers in English -- Leicester Scientific Corpus (LSC). The LSC contains 1,673,824 abstracts of research articles and proceeding papers indexed by Web of Science (WoS)…

Computation and Language · Computer Science 2019-12-17 Neslihan Suzen , Evgeny M. Mirkes , Alexander N. Gorban
‹ Prev 1 2 3 10 Next ›