English
Related papers

Related papers: Modeling Language Change in Historical Corpora: Th…

200 papers

We provide a method for automatically detecting change in language across time through a chronologically trained neural language model. We train the model on the Google Books Ngram corpus to obtain word vector representations specific to…

Computation and Language · Computer Science 2014-08-26 Yoon Kim , Yi-I Chiu , Kentaro Hanaki , Darshan Hegde , Slav Petrov

Language change is a cultural evolutionary process in which variants of linguistic variables change in frequency through processes analogous to mutation, selection and genetic drift. In this work, we apply a recently-introduced method to…

Computation and Language · Computer Science 2023-08-22 Juan Guerrero Montero , Andres Karjus , Kenny Smith , Richard A. Blythe

We describe our participation in the PAN 2017 shared task on Author Profiling, identifying authors' gender and language variety for English, Spanish, Arabic and Portuguese. We describe both the final, submitted system, and a series of…

Computation and Language · Computer Science 2017-07-13 Angelo Basile , Gareth Dwyer , Maria Medvedeva , Josine Rawee , Hessel Haagsma , Malvina Nissim

Deep learning methods have been increasingly applied to computational linguistics to uncover patterns in text data. This study investigates author-specific word class distributions using part-of-speech (POS) tagging and bigram analysis. By…

Computation and Language · Computer Science 2025-01-20 Patrick Krauss , Achim Schilling

Languages change over time. Computational models can be trained to recognize such changes enabling them to estimate the publication date of texts. Despite recent advancements in Large Language Models (LLMs), their performance on automatic…

Computation and Language · Computer Science 2026-03-13 Nishat Raihan , Marcos Zampieri

The unigram distribution is the non-contextual probability of finding a specific word form in a corpus. While of central importance to the study of language, it is commonly approximated by each word's sample frequency in the corpus. This…

Computation and Language · Computer Science 2021-06-07 Irene Nikkarinen , Tiago Pimentel , Damián E. Blasi , Ryan Cotterell

We propose a theoretical framework within which information on the vocabulary of a given corpus can be inferred on the basis of statistical information gathered on that corpus. Inferences can be made on the categories of the words in the…

Computation and Language · Computer Science 2008-10-08 Pascal Vaillant , Richard Nock , Claudia Henry

Recent work has found that contemporary language models such as transformers can become so good at next-word prediction that the probabilities they calculate become worse for predicting reading time. In this paper, we propose that this can…

Computation and Language · Computer Science 2026-03-11 James A. Michaelov , Roger P. Levy

This article is devoted to the verification of the empirical Heaps law in European languages using Google Books Ngram corpus data. The connection between word distribution frequency and expected dependence of individual word number on text…

Computation and Language · Computer Science 2020-03-30 Vladimir V. Bochkarev , Eduard Yu. Lerner , Anna V. Shevlyakova

In this paper, a method for measuring synchronic corpus (dis-)similarity put forward by Kilgarriff (2001) is adapted and extended to identify trends and correlated changes in diachronic text data, using the Corpus of Historical American…

Computation and Language · Computer Science 2015-08-28 Alexander Koplenig

This paper presents a computational approach to author profiling taking gender and language variety into account. We apply an ensemble system with the output of multiple linear SVM classifiers trained on character and word $n$-grams. We…

Computation and Language · Computer Science 2017-07-04 Alina Maria Ciobanu , Marcos Zampieri , Shervin Malmasi , Liviu P. Dinu

The recent dramatic increase in online data availability has allowed researchers to explore human culture with unprecedented detail, such as the growth and diversification of language. In particular, it provides statistical tools to explore…

Historic variations of spelling poses a challenge for full-text search or natural language processing on historical digitized texts. To minimize the gap between the historic orthography and contemporary spelling, usually an automatic…

Computation and Language · Computer Science 2025-02-26 Anton Ehrmanntraut

Morphological and syntactic changes in word usage (as captured, e.g., by grammatical profiles) have been shown to be good predictors of a word's meaning change. In this work, we explore whether large pre-trained contextualised language…

Computation and Language · Computer Science 2022-04-13 Mario Giulianelli , Andrey Kutuzov , Lidia Pivovarova

An automatic word classification system has been designed which processes word unigram and bigram frequency statistics extracted from a corpus of natural language utterances. The system implements a binary top-down form of word clustering…

cmp-lg · Computer Science 2016-08-31 John McMahon , F. J. Smith

The availability of large linguistic data sets enables data-driven approaches to study linguistic change. The Google Books corpus unigram frequency data set is used to investigate the word rank dynamics in eight languages. We observed the…

Computation and Language · Computer Science 2022-02-15 Alex John Quijano , Rick Dale , Suzanne Sindi

Text normalization techniques based on rules, lexicons or supervised training requiring large corpora are not scalable nor domain interchangeable, and this makes them unsuitable for normalizing user-generated content (UGC). Current tools…

Computation and Language · Computer Science 2017-04-11 Thales Felipe Costa Bertaglia , Maria das Graças Volpe Nunes

This paper proposes methods of predicting dynamic time series (including non-stationary ones) based on a linguistic approach, namely, the study of occurrences and repetition of so-called N-grams. This approach is used in computational…

Numerical Analysis · Mathematics 2026-02-26 Dmytro Lande , Volodymyr Yuzefovych , Yevheniia Tsybulska

This research offers a new interdisciplinary approach to the field of Linguistics by using Computational Linguistics, NLP, Bayesian Statistics and Sociolinguistics methods. This thesis investigates word order change in infinitival clauses…

Computation and Language · Computer Science 2020-11-18 Olga Scrivner

In this paper we describe the use of text classification methods to investigate genre and method variation in an English - German translation corpus. For this purpose we use linguistically motivated features representing texts using a…

Computation and Language · Computer Science 2017-09-14 Ekaterina Lapshninova-Koltunski , Marcos Zampieri
‹ Prev 1 2 3 10 Next ›