English
Related papers

Related papers: English verb regularization in books and tweets

200 papers

Recent efforts to consolidate guidelines and treebanks in the Universal Dependencies project raise the expectation that joint training and dataset comparison is increasingly possible for high-resource languages such as English, which have…

Computation and Language · Computer Science 2023-02-02 Amir Zeldes , Nathan Schneider

The proliferation of NLP-powered language technologies, AI-based natural language generation models, and English as a mainstream means of communication among both native and non-native speakers make the output of AI-powered tools especially…

Computation and Language · Computer Science 2025-02-07 Karolina Rudnicka

Studies of the overall structure of vocabulary and its dynamics became possible due to creation of diachronic text corpora, especially Google Books Ngram. This article discusses the question of core change rate and the degree to which the…

Computation and Language · Computer Science 2020-03-25 Valery D. Solovyev , Vladimir V. Bochkarev , Anna V. Shevlyakova

This study focuses on how scientifically-correct information is disseminated through social media, and how misinformation can be corrected. We have identified examples on Twitter where scientific terms that have been misused have been…

Social and Information Networks · Computer Science 2022-04-15 Dongwoo Lim , Fujio Toriumi , Mitsuo Yoshida

The availability of large linguistic data sets enables data-driven approaches to study linguistic change. The Google Books corpus unigram frequency data set is used to investigate the word rank dynamics in eight languages. We observed the…

Computation and Language · Computer Science 2022-02-15 Alex John Quijano , Rick Dale , Suzanne Sindi

The election forecasting 'industry' is a growing one, both in the volume of scholars producing forecasts and methodological diversity. In recent years a new approach has emerged that relies on social media and particularly Twitter data to…

Computers and Society · Computer Science 2015-05-08 Pete Burnap , Rachel Gibson , Luke Sloan , Rosalynd Southern , Matthew Williams

Twitter provides an open and rich source of data for studying human behaviour at scale and is widely used in social and network sciences. However, a major criticism of Twitter data is that demographic information is largely absent.…

Social and Information Networks · Computer Science 2017-02-27 Benjamin Paul Chamberlain , Clive Humby , Marc Peter Deisenroth

Transliteration is a task of translating named entities from a language to another, based on phonetic similarity. The task has embraced deep learning approaches in recent years, yet, most ignore the phonetic features of the involved…

Computation and Language · Computer Science 2022-01-24 Shi Cheng , Zhuofei Ding , Songpeng Yan

Scientific English has undergone rapid and unprecedented changes in recent years, with words such as "delve," "intricate," and "crucial" showing significant spikes in frequency since around 2022. These changes are widely attributed to the…

Computation and Language · Computer Science 2025-06-30 Riley Galpin , Bryce Anderson , Tom S. Juzek

Bragging is the act of uttering statements that are likely to be positively viewed by others and it is extensively employed in human communication with the aim to build a positive self-image of oneself. Social media is a natural platform…

Computation and Language · Computer Science 2024-03-26 Mali Jin , Daniel Preoţiuc-Pietro , A. Seza Doğruöz , Nikolaos Aletras

Recently, numerous approaches have emerged in the social sciences to exploit the opportunities made possible by the vast amounts of data generated by online social networks (OSNs). Having access to information about users on such a scale…

Social media currently provide a window on our lives, making it possible to learn how people from different places, with different backgrounds, ages, and genders use language. In this work we exploit a newly-created Arabic dataset with…

Computation and Language · Computer Science 2019-11-05 Muhammad Abdul-Mageed , Chiyu Zhang , Arun Rajendran , AbdelRahim Elmadany , Michael Przystupa , Lyle Ungar

Statement autoformalization, the automated translation of statements from natural language into formal languages, has become a subject of extensive research, yet the development of robust automated evaluation metrics remains limited.…

Machine Learning · Computer Science 2025-08-25 Yuntian Liu , Tao Zhu , Xiaoyang Liu , Yu Chen , Zhaoxuan Liu , Qingfeng Guo , Jiashuo Zhang , Kangjie Bao , Tao Luo

Social media has become extremely influential when it comes to policy making in modern societies, especially in the western world, where platforms such as Twitter allow users to follow politicians, thus making citizens more involved in…

Computation and Language · Computer Science 2023-04-05 Dimosthenis Antypas , Alun Preece , Jose Camacho-Collados

The pervasiveness of mobile devices, which is increasing daily, is generating a vast amount of geo-located data allowing us to gain further insights into human behaviors. In particular, this new technology enables users to communicate…

Physics and Society · Physics 2014-08-26 Maxime Lenormand , Antònia Tugores , Pere Colet , José J. Ramasco

Over the last million years, human language has emerged and evolved as a fundamental instrument of social communication and semiotic representation. People use language in part to convey emotional information, leading to the central and…

Code-mixing or code-switching are the effortless phenomena of natural switching between two or more languages in a single conversation. Use of a foreign word in a language; however, does not necessarily mean that the speaker is…

Computation and Language · Computer Science 2017-03-16 Jasabanta Patro , Bidisha Samanta , Saurabh Singh , Prithwish Mukherjee , Monojit Choudhury , Animesh Mukherjee

Variations in writing styles are commonly used to adapt the content to a specific context, audience, or purpose. However, applying stylistic variations is still by and large a manual process, and there have been little efforts towards…

Computation and Language · Computer Science 2017-07-24 Harsh Jhamtani , Varun Gangal , Eduard Hovy , Eric Nyberg

Many text corpora exhibit socially problematic biases, which can be propagated or amplified in the models trained on such data. For example, doctor cooccurs more frequently with male pronouns than female pronouns. In this study we (i)…

Computation and Language · Computer Science 2019-04-08 Shikha Bordia , Samuel R. Bowman

Linguistic uncertainty is common in social media, but its relationship with engagement remains unclear across languages and topics. Using 2,258 English-language posts on Federal Reserve policy, inflation, and electoral politics collected…

Computers and Society · Computer Science 2026-05-19 Mohamed Soufan