中文
相关论文

相关论文: English verb regularization in books and tweets

200 篇论文

Spelling variation (e.g. funnnn vs. fun) can influence the social perception of texts and their writers: we often have various associations with different forms of writing (is the text informal? does the writer seem young?). In this study,…

计算与语言 · 计算机科学 2025-12-01 Dong Nguyen , Laura Rosseel

The frequency with which the letters of the English alphabet appear in writings has been applied to the field of cryptography, the development of keyboard mechanics, and the study of linguistics. We expanded on the statistical analysis of…

信息论 · 计算机科学 2024-01-30 Neil Zhao , Diana Zheng

Living languages are shaped by a host of conflicting internal and external evolutionary pressures. While some of these pressures are universal across languages and cultures, others differ depending on the social and conversational context:…

This document describes a sizable grammar of English written in the TAG formalism and implemented for use with the XTAG system. This report and the grammar described herein supersedes the TAG grammar described in an earlier 1995 XTAG…

计算与语言 · 计算机科学 2012-08-27 XTAG Research Group

Consider a person trying to spread an important message on a social network. He/she can spend hours trying to craft the message. Does it actually matter? While there has been extensive prior work looking into predicting popularity of…

社会与信息网络 · 计算机科学 2014-05-08 Chenhao Tan , Lillian Lee , Bo Pang

Transformer language models have received widespread public attention, yet their generated text is often surprising even to NLP researchers. In this survey, we discuss over 250 recent studies of English language model behavior before…

计算与语言 · 计算机科学 2023-08-29 Tyler A. Chang , Benjamin K. Bergen

The current study yielded a number of important findings. We managed to build a neural network that achieved an accuracy score of 91 per cent in classifying troll and genuine tweets. By means of regression analysis, we identified a number…

社会与信息网络 · 计算机科学 2019-11-21 Sergei Monakhov

Languages emerge and change over time at the population level though interactions between individual speakers. It is, however, hard to directly observe how a single speaker's linguistic innovation precipitates a population-wide change in…

计算与语言 · 计算机科学 2021-06-04 Richard A Blythe , William Croft

Our usage of language is not solely reliant on cognition but is arguably determined by myriad external factors leading to a global variability of linguistic patterns. This issue, which lies at the core of sociolinguistics and is backed by…

计算与语言 · 计算机科学 2018-04-05 Jacob Levy Abitbol , Márton Karsai , Jean-Philippe Magué , Jean-Pierre Chevrot , Eric Fleury

One advantage of neural ranking models is that they are meant to generalise well in situations of synonymity i.e. where two words have similar or identical meanings. In this paper, we investigate and quantify how well various ranking models…

信息检索 · 计算机科学 2023-08-02 Andreas Chari , Sean MacAvaney , Iadh Ounis

Words are malleable objects, influenced by events that are reflected in written texts. Situated in the global outbreak of COVID-19, our research aims at detecting semantic shifts in social media language triggered by the health crisis. With…

计算与语言 · 计算机科学 2021-02-17 Yanzhu Guo , Christos Xypolopoulos , Michalis Vazirgiannis

We provide a method for automatically detecting change in language across time through a chronologically trained neural language model. We train the model on the Google Books Ngram corpus to obtain word vector representations specific to…

计算与语言 · 计算机科学 2014-08-26 Yoon Kim , Yi-I Chiu , Kentaro Hanaki , Darshan Hegde , Slav Petrov

We introduce a dataset for studying the evolution of words, constructed from WordNet and the Google Books Ngram Corpus. The dataset tracks the evolution of 4,000 synonym sets (synsets), containing 9,000 English words, from 1800 AD to 2000…

计算与语言 · 计算机科学 2019-08-21 Peter D. Turney , Saif M. Mohammad

This article is devoted to the verification of the empirical Heaps law in European languages using Google Books Ngram corpus data. The connection between word distribution frequency and expected dependence of individual word number on text…

计算与语言 · 计算机科学 2020-03-30 Vladimir V. Bochkarev , Eduard Yu. Lerner , Anna V. Shevlyakova

The study uses the British National Corpus 2014, a large sample of contemporary spoken British English, to investigate language patterns across different age groups. Our research attempts to explore how language patterns vary between…

计算与语言 · 计算机科学 2025-06-24 MingZe Tang

When the world changes, so does the text that humans write about it. How do we build language models that can be easily updated to reflect these changes? One popular approach is retrieval-augmented generation, in which new documents are…

计算与语言 · 计算机科学 2024-06-18 Belinda Z. Li , Emmy Liu , Alexis Ross , Abbas Zeitoun , Graham Neubig , Jacob Andreas

Well-established cognitive models coming from anthropology have shown that, due to the cognitive constraints that limit our "bandwidth" for social interactions, humans organize their social relations according to a regular structure. In…

社会与信息网络 · 计算机科学 2023-04-04 Kilian Ollivier , Chiara Boldrini , Andrea Passarella , Marco Conti

With a sharp rise in fluency and users of "Hinglish" in linguistically diverse country, India, it has increasingly become important to analyze social content written in this language in platforms such as Twitter, Reddit, Facebook. This…

计算与语言 · 计算机科学 2020-01-01 Vivek Kumar Gupta

Maintaining the integrity of long-term data collection is an essential scientific practice. As a field evolves, so too will that field's measurement instruments and data storage systems, as they are invented, improved upon, and made…

Twitter is a well-known microblogging social site where users express their views and opinions in real-time. As a result, tweets tend to contain valuable information. With the advancements of deep learning in the domain of natural language…

计算与语言 · 计算机科学 2020-10-22 Mohiuddin Md Abdul Qudar , Vijay Mago