中文
相关论文

相关论文: Explorations in an English Poetry Corpus: A Neuroc…

200 篇论文

The Gutenberg Literary English Corpus (GLEC) provides a rich source of textual data for research in digital humanities, computational linguistics or neurocognitive poetics. However, so far only a small subcorpus, the Gutenberg English…

计算与语言 · 计算机科学 2020-10-22 Arthur M. Jacobs , Annette Kinder

The Gutenberg Literary English Corpus (GLEC, Jacobs, 2018a) provides a rich source of textual data for research in digital humanities, computational linguistics or neurocognitive poetics. In this study we address differences among the…

计算与语言 · 计算机科学 2022-01-13 Arthur M. Jacobs , Annette Kinder

The use of Project Gutenberg (PG) as a text corpus has been extremely popular in statistical analysis of language for more than 25 years. However, in contrast to other major linguistic datasets of similar importance, no consensual full…

计算与语言 · 计算机科学 2018-12-20 Martin Gerlach , Francesc Font-Clos

We present DepCC, the largest-to-date linguistically analyzed corpus in English including 365 million documents, composed of 252 billion tokens and 7.5 billion of named entity occurrences in 14.3 billion sentences from a web-scale crawl of…

计算与语言 · 计算机科学 2018-03-01 Alexander Panchenko , Eugen Ruppert , Stefano Faralli , Simone Paolo Ponzetto , Chris Biemann

We present the Project Dialogism Novel Corpus, or PDNC, an annotated dataset of quotations for English literary texts. PDNC contains annotations for 35,978 quotations across 22 full-length novels, and is by an order of magnitude the largest…

计算与语言 · 计算机科学 2022-04-13 Krishnapriya Vishnubhotla , Adam Hammond , Graeme Hirst

I apply Schmidhuber's compression progress theory of interestingness at corpus scale, analyzing semantic novelty trajectories in more than 80,000 books spanning two centuries of English-language publishing. Using sentence-transformer…

计算与语言 · 计算机科学 2026-03-03 Fred Zimmerman

This discussion paper reflects on how quantitative approaches to historical linguistics interact with dataset properties. Drawing on two worked examples, we examine English data using quad-based concept modelling of Early Modern English…

计算与语言 · 计算机科学 2026-05-05 Catherine Wong , Bach Phan-Tat , Susan Fitzmaurice

In this paper, we try to explore the evolution of language through case calculations. First, we chose the novels of eleven British writers from 1400 to 2005 and found the corresponding works; Then, we use the natural language processing…

计算与语言 · 计算机科学 2018-10-09 Zhu Gao , Yanhui Jiang , Junhui Gao

We present a novel corpus of 445 human- and computer-generated documents, comprising about 27,000 clauses, annotated for semantic clause types and coherence relations that allow for nuanced comparison of artificial and natural discourse…

Due to its semantic succinctness and novelty of expression, poetry is a great test bed for semantic change analysis. However, so far there is a scarcity of large diachronic corpora. Here, we provide a large corpus of German poetry which…

计算与语言 · 计算机科学 2019-09-27 Thomas Haider , Steffen Eger

A simple method for finding the entropy and redundancy of a reasonable long sample of English text by direct computer processing and from first principles according to Shannon theory is presented. As an example, results on the entropy of…

计算与语言 · 计算机科学 2009-11-19 Fabio G. Guerrero

Embedding models group text by semantic content, what text is about. We show that temporal co-occurrence within texts discovers a different kind of structure: recurrent transition-structure concepts or what text does. We train a…

人工智能 · 计算机科学 2026-03-20 Jason Dury

We present in this work a new dataset of coreference annotations for works of literature in English, covering 29,103 mentions in 210,532 tokens from 100 works of fiction. This dataset differs from previous coreference datasets in containing…

计算与语言 · 计算机科学 2020-05-18 David Bamman , Olivia Lewke , Anya Mansoor

We present a mechanistic analysis of literary style in GPT-2, identifying individual neurons that discriminate between exemplary prose and rigid AI-generated text. Using Herman Melville's Bartleby, the Scrivener as a corpus, we extract…

计算与语言 · 计算机科学 2025-10-22 Tsogt-Ochir Enkhbayar

We introduce S2ORC, a large corpus of 81.1M English-language academic papers spanning many academic disciplines. The corpus consists of rich metadata, paper abstracts, resolved bibliographic references, as well as structured full text for…

计算与语言 · 计算机科学 2020-07-08 Kyle Lo , Lucy Lu Wang , Mark Neumann , Rodney Kinney , Dan S. Weld

Entities like person, location, organization are important for literary text analysis. The lack of annotated data hinders the progress of named entity recognition (NER) in literary domain. To promote the research of literary NER, we build…

计算与语言 · 计算机科学 2024-10-16 Hanjie Zhao , Jinge Xie , Yuchen Yan , Yuxiang Jia , Yawen Ye , Hongying Zan

This paper introduces the Multi-Genre Natural Language Inference (MultiNLI) corpus, a dataset designed for use in the development and evaluation of machine learning models for sentence understanding. In addition to being one of the largest…

计算与语言 · 计算机科学 2018-02-21 Adina Williams , Nikita Nangia , Samuel R. Bowman

Publishing articles in high-impact English journals is difficult for scholars around the world, especially for non-native English-speaking scholars (NNESs), most of whom struggle with proficiency in English. In order to uncover the…

计算与语言 · 计算机科学 2018-09-14 Chao Lu , Yi Bu , Jie Wang , Ying Ding , Vetle Torvik , Matthew Schnaars , Chengzhi Zhang

In this study, we investigated the academic literature on quantum technologies (QT) using bibliometric tools. We used a set of 49,823 articles obtained from the Web of Science (WoS) database using a search query constructed through expert…

数字图书馆 · 计算机科学 2021-05-12 Zeki C. Seskir , Arsev U. Aydinoglu

In this paper, we present a scientific corpus of abstracts of academic papers in English -- Leicester Scientific Corpus (LSC). The LSC contains 1,673,824 abstracts of research articles and proceeding papers indexed by Web of Science (WoS)…

计算与语言 · 计算机科学 2019-12-17 Neslihan Suzen , Evgeny M. Mirkes , Alexander N. Gorban
‹ 上一页 1 2 3 10 下一页 ›