English
Related papers

Related papers: Lexical Normalisation of Twitter Data

200 papers

Maintaining the integrity of long-term data collection is an essential scientific practice. As a field evolves, so too will that field's measurement instruments and data storage systems, as they are invented, improved upon, and made…

Physics and Society · Physics 2020-08-31 P. S. Dodds , J. R. Minot , M. V. Arnold , T. Alshaabi , J. L. Adams , D. R. Dewhurst , A. J. Reagan , C. M. Danforth

We describe TweeTIME, a temporal tagger for recognizing and normalizing time expressions in Twitter. Most previous work in social media analysis has to rely on temporal resolvers that are designed for well-edited text, and therefore suffer…

Information Retrieval · Computer Science 2020-11-17 Jeniya Tabassum , Alan Ritter , Wei Xu

In recent years, social media has been criticized for yielding polarization. Identifying emerging disagreements and growing polarization is important for journalists to create alerts and provide more balanced coverage. While recent studies…

Social and Information Networks · Computer Science 2022-11-30 Tomoki Fukuma , Koki Noda , Hiroki Kumagai , Hiroki Yamamoto , Yoshiharu Ichikawa , Kyosuke Kambe , Yu Maubuchi , Fujio Toriumi

In real-time, social media data strongly imprints world events, popular culture, and day-to-day conversations by millions of ordinary people at a scale that is scarcely conventionalized and recorded. Vitally, and absent from many standard…

Lexical normalisation (LN) is the process of correcting each word in a dataset to its canonical form so that it may be more easily and more accurately analysed. Most lexical normalisation systems operate at the character-level, while…

Computation and Language · Computer Science 2019-11-15 Michael Stewart , Wei Liu , Rachel Cardell-Oliver

Semantic sentence embeddings are usually supervisedly built minimizing distances between pairs of embeddings of sentences labelled as semantically similar by annotators. Since big labelled datasets are rare, in particular for non-English…

Computation and Language · Computer Science 2021-10-06 Marco Di Giovanni , Marco Brambilla

To analyse large numbers of texts, social science researchers are increasingly confronting the challenge of text classification. When manual labeling is not possible and researchers have to find automatized ways to classify texts, computer…

Computation and Language · Computer Science 2023-10-10 Karina Shyrokykh , Maksym Girnyk , Lisa Dellmuth

This study addresses the problem of generating an optimised keyboard layout for single-finger typing on a smartphone. It offers Twitter users a tweet-typing experience that requires less effort and time. Bodies of tweet text for 85 popular…

Human-Computer Interaction · Computer Science 2023-10-04 E. Elson

Establishing authorship of online texts is fundamental to combat cybercrimes. Unfortunately, text length is limited on some platforms, making the challenge harder. We aim at identifying the authorship of Twitter messages limited to 140…

Computation and Language · Computer Science 2021-11-29 Fernando Alonso-Fernandez , Nicole Mariah Sharon Belvisi , Kevin Hernandez-Diaz , Naveed Muhammad , Josef Bigun

During sudden onset crisis events, the presence of spam, rumors and fake content on Twitter reduces the value of information contained on its messages (or "tweets"). A possible solution to this problem is to use machine learning to…

Cryptography and Security · Computer Science 2015-02-02 Aditi Gupta , Ponnurangam Kumaraguru , Carlos Castillo , Patrick Meier

City Logistics is characterized by multiple stakeholders that often have different views of such a complex system. From a public policy perspective, identifying stakeholders, issues and trends is a daunting challenge, only partially…

Machine Learning · Computer Science 2019-06-19 Simon Tamayo , François Combes , Gaudron Arthur

Language change is influenced by many factors, but often starts from synchronic variation, where multiple linguistic patterns or forms coexist, or where different speech communities use language in increasingly different ways. Besides…

Social and Information Networks · Computer Science 2023-09-06 Andres Karjus , Christine Cuskley

In this paper we present a method to identify tweets that a user may find interesting enough to retweet. The method is based on a global, but personalized classifier, which is trained on data from several users, represented in terms of…

Social and Information Networks · Computer Science 2017-09-20 Michail Vougioukas , Ion Androutsopoulos , Georgios Paliouras

The movements of ideas and content between locations and languages are unquestionably crucial concerns to researchers of the information age, and Twitter has emerged as a central, global platform on which hundreds of millions of people…

Computers and Society · Computer Science 2014-08-19 Mark Graham , Scott A. Hale , Devin Gaffney

The collection and examination of social media has become a useful mechanism for studying the mental activity and behavior tendencies of users. Through the analysis of collected Twitter data, models were developed for classifying…

Social and Information Networks · Computer Science 2020-03-26 Joseph Tassone , Peizhi Yan , Mackenzie Simpson , Chetan Mendhe , Vijay Mago , Salimur Choudhury

On daily basis, millions of Twitter accounts post a vast number of tweets including numerous Twitter entities (mentions, replies, hashtags, photos, URLs). Many of these entities are used in common by many accounts. The more common entities…

Social and Information Networks · Computer Science 2015-06-02 Gerasimos Razis , Ioannis Anagnostopoulos

In Twitter, a name, phrase, or topic that is mentioned at a greater rate than others is called a "trending topic" or simply "trend". Twitter trends list has a powerful ability to promote public events such as natural events, political…

Social and Information Networks · Computer Science 2020-08-31 Issa Annamoradnejad , Jafar Habibi

With the advancement of web technology and its growth, there is a huge volume of data present in the web for internet users and a lot of data is generated too. Internet has become a platform for online learning, exchanging ideas and sharing…

Computation and Language · Computer Science 2016-11-04 Vishal. A. Kharde , Prof. Sheetal. Sonawane

Social media data is a valuable resource for research, yet it contains a wide range of non-standard words (NSW). These irregularities hinder the effective operation of NLP tools. Current state-of-the-art methods for the Vietnamese language…

Computation and Language · Computer Science 2024-07-26 Anh Thi-Hoang Nguyen , Dung Ha Nguyen , Nguyet Thi Nguyen , Khanh Thanh-Duy Ho , Kiet Van Nguyen

Online hate speech is associated with substantial social harms, yet it remains unclear how consistently platforms enforce hate speech policies or whether enforcement is feasible at scale. We address these questions through a global audit of…