中文
相关论文

相关论文: Lexical Normalisation of Twitter Data

200 篇论文

In November 2017, Twitter doubled the maximum allowed tweet length from 140 to 280 characters, a drastic switch on one of the world's most influential social media platforms. In the first long-term study of how the new length limit was…

社会与信息网络 · 计算机科学 2020-09-17 Kristina Gligorić , Ashton Anderson , Robert West

Even though the Internet and social media have increased the amount of news and information people can consume, most users are only exposed to content that reinforces their positions and isolates them from other ideological communities.…

社会与信息网络 · 计算机科学 2021-12-21 Federico Albanese , Leandro Lombardi , Esteban Feuerstein , Pablo Balenzuela

In this paper we describe a dynamic normalization process applied to social network multilingual documents (Facebook and Twitter) to improve the performance of the Author profiling task for short texts. After the normalization process,…

With the increase in popularity of deep learning models for natural language processing (NLP) tasks, in the field of Pharmacovigilance, more specifically for the identification of Adverse Drug Reactions (ADRs), there is an inherent need for…

信息检索 · 计算机科学 2020-04-01 Ramya Tekumalla , Juan M. Banda

We present TweeNLP, a one-stop portal that organizes Twitter's natural language processing (NLP) data and builds a visualization and exploration platform. It curates 19,395 tweets (as of April 2021) from various NLP conferences and general…

计算与语言 · 计算机科学 2021-06-22 Viraj Shah , Shruti Singh , Mayank Singh

The prevalence of social media presents a growing opportunity to collect and analyse examples of English varieties. Whilst usage of these varieties was - and, in many cases, still is - used only in spoken contexts or hard-to-access private…

计算与语言 · 计算机科学 2024-01-23 Nhi Pham , Lachlan Pham , Adam L. Meyers

The need for a comprehensive study to explore various aspects of online social media has been instigated by many researchers. This paper gives an insight into the social platform, Twitter. In this present work, we have illustrated stepwise…

社会与信息网络 · 计算机科学 2021-05-24 Shahab Saquib Sohail , Mohammad Muzammil Khan , M. Afshar Alam

Our paper studies the predictability of online speech -- that is, how well language models learn to model the distribution of user generated content on X (previously Twitter). We define predictability as a measure of the model's…

计算与语言 · 计算机科学 2026-01-07 Mina Remeli , Moritz Hardt , Robert C. Williamson

Online media, such as blogs and social networking sites, generate massive volumes of unstructured data of great interest to analyze the opinions and sentiments of individuals and organizations. Novel approaches beyond Natural Language…

Over 500 million tweets are posted in Twitter each day, out of which about 11% tweets are deleted by the users posting them. This phenomenon of widespread deletion of tweets leads to a number of questions: what kind of content posted by…

社会与信息网络 · 计算机科学 2022-12-27 Parantapa Bhattacharya , Saptarshi Ghosh , Niloy Ganguly

One of the major sources of trending news, events and opinion in the current age is micro blogging. Twitter, being one of them, is extensively used to mine data about public responses and event updates. This paper intends to propose methods…

社会与信息网络 · 计算机科学 2015-06-22 Rishabh Jain , Abhishek B. S. , Satvik Jagannath

In recent years, social bots have been using increasingly more sophisticated, challenging detection strategies. While many approaches and features have been proposed, social bots evade detection and interact much like humans making it…

社会与信息网络 · 计算机科学 2018-12-20 Isa Inuwa-Dutse , Bello Shehu Bello , Ioannis Korkontzelos

Predicting personality is essential for social applications supporting human-centered activities, yet prior modeling methods with users written text require too much input data to be realistically used in the context of social media. In…

社会与信息网络 · 计算机科学 2017-04-20 Pierre-Hadrien Arnoux , Anbang Xu , Neil Boyette , Jalal Mahmud , Rama Akkiraju , Vibha Sinha

A major challenge in paraphrase research is the lack of parallel corpora. In this paper, we present a new method to collect large-scale sentential paraphrases from Twitter by linking tweets through shared URLs. The main advantage of our…

计算与语言 · 计算机科学 2017-08-02 Wuwei Lan , Siyu Qiu , Hua He , Wei Xu

An important part of the information gathering and data analysis is to find out what people think about, either a product or an entity. Twitter is an opinion rich social networking site. The posts or tweets from this data can be used for…

信息检索 · 计算机科学 2024-09-05 Dwarampudi Mahidhar Reddy , N V Subba Reddy , N V Subba Reddy

In the last decade, social networks became most popular medium for communication and interaction. As an example, micro-blogging service Twitter has more than 200 million registered users who exchange more than 65 million posts per day.…

信息检索 · 计算机科学 2020-01-29 Qadri Mishael , Aladdin Ayesh

A word embedding is a low-dimensional, dense and real- valued vector representation of a word. Word embeddings have been used in many NLP tasks. They are usually gener- ated from a large text corpus. The embedding of a word cap- tures both…

计算与语言 · 计算机科学 2017-08-15 Quanzhi Li , Sameena Shah , Xiaomo Liu , Armineh Nourbakhsh

Analysing sentiment of tweets is important as it helps to determine the users' opinion. Knowing people's opinion is crucial for several purposes starting from gathering knowledge about customer base, e-governance, campaigning and many more.…

计算与语言 · 计算机科学 2016-11-30 Lahari Poddar , Kishaloy Halder , Xianyan Jia

Computer-mediated communication is driving fundamental changes in the nature of written language. We investigate these changes by statistical analysis of a dataset comprising 107 million Twitter messages (authored by 2.7 million unique user…

计算与语言 · 计算机科学 2014-11-25 Jacob Eisenstein , Brendan O'Connor , Noah A. Smith , Eric P. Xing

With the rise in popularity of public social media and micro-blogging services, most notably Twitter, the people have found a venue to hear and be heard by their peers without an intermediary. As a consequence, and aided by the public…

计算与语言 · 计算机科学 2016-06-21 Prashanth Vijayaraghavan , Soroush Vosoughi , Deb Roy