中文
相关论文

相关论文: EveTAR: Building a Large-Scale Multi-Task Test Col…

200 篇论文

The prominence of figurative language devices, such as sarcasm and irony, poses serious challenges for Arabic Sentiment Analysis (SA). While previous research works tackle SA and sarcasm detection separately, this paper introduces an…

计算与语言 · 计算机科学 2021-06-24 Abdelkader El Mahdaouy , Abdellah El Mekki , Kabil Essefar , Nabil El Mamoun , Ismail Berrada , Ahmed Khoumsi

In recent work, we identified and studied a small cohort of Twitter users whose pregnancies with birth defect outcomes could be observed via their publicly available tweets. Exploiting social media's large-scale potential to complement the…

计算与语言 · 计算机科学 2019-10-03 Ari Z. Klein , Abeed Sarker , Davy Weissenbacher , Graciela Gonzalez-Hernandez

The role of social media in opinion formation has far-reaching implications in all spheres of society. Though social media provide platforms for expressing news and views, it is hard to control the quality of posts due to the sheer volumes…

机器学习 · 计算机科学 2021-09-08 Rini Anggrainingsih , Ghulam Mubashar Hassan , Amitava Datta

How can we study social interactions on evolving topics at a mass scale? Over the past decade, researchers from diverse fields such as economics, political science, and public health have often done this by querying Twitter's public API…

社会与信息网络 · 计算机科学 2022-09-23 Sacha Lévy , Farimah Poursafaei , Kellin Pelrine , Reihaneh Rabbany

The local event detection is to use posting messages with geotags on social networks to reveal the related ongoing events and their locations. Recent studies have demonstrated that the geo-tagged tweet stream serves as an unprecedentedly…

信息检索 · 计算机科学 2020-04-07 Sibo Zhang , Yuan Cheng , Deyuan Ke

Gender analysis of Twitter can reveal important socio-cultural differences between male and female users. There has been a significant effort to analyze and automatically infer gender in the past for most widely spoken languages' content,…

计算与语言 · 计算机科学 2022-03-02 Hamdy Mubarak , Shammur Absar Chowdhury , Firoj Alam

During broadcast events such as the Superbowl, the U.S. Presidential and Primary debates, etc., Twitter has become the de facto platform for crowds to share perspectives and commentaries about them. Given an event and an associated…

社会与信息网络 · 计算机科学 2012-12-24 Yuheng Hu , Ajita John , Fei Wang , Subbarao Kambhampati

Arabic dialect identification is a complex problem for a number of inherent properties of the language itself. In this paper, we present the experiments conducted, and the models developed by our competing team, Mawdoo3 AI, along the way to…

We tackle the challenge of topic classification of tweets in the context of analyzing a large collection of curated streams by news outlets and other organizations to deliver relevant content to users. Our approach is novel in applying…

信息检索 · 计算机科学 2017-04-25 Salman Mohammed , Nimesh Ghelani , Jimmy Lin

Receiving timely and relevant security information is crucial for maintaining a high-security level on an IT infrastructure. This information can be extracted from Open Source Intelligence published daily by users, security organisations,…

密码学与安全 · 计算机科学 2019-04-04 Fernando Alves , Aurélien Bettini , Pedro M. Ferreira , Alysson Bessani

In this paper we present a method to identify tweets that a user may find interesting enough to retweet. The method is based on a global, but personalized classifier, which is trained on data from several users, represented in terms of…

社会与信息网络 · 计算机科学 2017-09-20 Michail Vougioukas , Ion Androutsopoulos , Georgios Paliouras

This paper presents the model submitted by the NIT_COVID-19 team for identified informative COVID-19 English tweets at WNUT-2020 Task2. This shared task addresses the problem of automatically identifying whether an English tweet related to…

计算与语言 · 计算机科学 2021-06-01 Jagadeesh M S , Alphonse P J A

This paper describes our approach (ur-iw-hnt) for the Shared Task of GermEval2021 to identify toxic, engaging, and fact-claiming comments. We submitted three runs using an ensembling strategy by majority (hard) voting with multiple…

计算与语言 · 计算机科学 2021-10-06 Hoai Nam Tran , Udo Kruschwitz

In this paper, we, as the DS@GT team for CLEF 2025 CheckThat! Task 4a Scientific Web Discourse Detection, present the methods we explored for this task. For this multiclass classification task, we determined if a tweet contained a…

计算与语言 · 计算机科学 2025-07-09 Ayush Parikh , Hoang Thanh Thanh Truong , Jeanette Schofield , Maximilian Heil

With the growth of social media platform influence, the effect of their misuse becomes more and more impactful. The importance of automatic detection of threatening and abusive language can not be overestimated. However, most of the…

We study the problem of analyzing tweets with Universal Dependencies. We extend the UD guidelines to cover special constructions in tweets that affect tokenization, part-of-speech tagging, and labeled dependencies. Using the extended…

计算与语言 · 计算机科学 2018-04-24 Yijia Liu , Yi Zhu , Wanxiang Che , Bing Qin , Nathan Schneider , Noah A. Smith

This paper addresses the quality issues in existing Twitter-based paraphrase datasets, and discusses the necessity of using two separate definitions of paraphrase for identification and generation tasks. We present a new Multi-Topic…

计算与语言 · 计算机科学 2022-11-09 Yao Dou , Chao Jiang , Wei Xu

Event-argument extraction is a challenging task, particularly in Arabic due to sparse linguistic resources. To fill this gap, we introduce the \hadath corpus ($550$k tokens) as an extension of Wojood, enriched with event-argument…

计算与语言 · 计算机科学 2024-08-01 Alaa Aljabari , Lina Duaibes , Mustafa Jarrar , Mohammed Khalilia

Social media such as tweets are emerging as platforms contributing to situational awareness during disasters. Information shared on Twitter by both affected population (e.g., requesting assistance, warning) and those outside the impact zone…

信息检索 · 计算机科学 2017-05-08 Hien To , Sumeet Agrawal , Seon Ho Kim , Cyrus Shahabi

The development of social media user stance detection and bot detection methods rely heavily on large-scale and high-quality benchmarks. However, in addition to low annotation quality, existing benchmarks generally have incomplete user…

计算机视觉与模式识别 · 计算机科学 2023-03-14 Shuhao Shi , Kai Qiao , Jian Chen , Shuai Yang , Jie Yang , Baojie Song , Linyuan Wang , Bin Yan