English
Related papers

Related papers: A Crowd-Annotated Spanish Corpus for Humor Analysi…

200 papers

We introduce the Self-Annotated Reddit Corpus (SARC), a large corpus for sarcasm research and for training and evaluating systems for sarcasm detection. The corpus has 1.3 million sarcastic statements -- 10 times more than any previous…

Computation and Language · Computer Science 2018-03-26 Mikhail Khodak , Nikunj Saunshi , Kiran Vodrahalli

Usage of online textual media is steadily increasing. Daily, more and more news stories, blog posts and scientific articles are added to the online volumes. These are all freely accessible and have been employed extensively in multiple…

Computation and Language · Computer Science 2017-08-16 Nattapong Sanchan , Ahmet Aker , Kalina Bontcheva

Twitter has emerged as a global hub for engaging in online conversations and as a research corpus for various disciplines that have recognized the significance of its user-generated content. Argument mining is an important analytical task…

Computation and Language · Computer Science 2024-04-02 Marc Feger , Stefan Dietze

We present the Twitter Job/Employment Corpus, a collection of tweets annotated by a humans-in-the-loop supervised learning framework that integrates crowdsourcing contributions and expertise on the local community and employment…

Computation and Language · Computer Science 2019-01-31 Tong Liu , Christopher M. Homan

The study of Twitter as a means for analyzing social phenomena has gained interest in recent years due to the availability of large amounts of data in a relatively spontaneous environment. Within opinion-mining tasks, emotion detection is…

Computation and Language · Computer Science 2024-07-11 Juan Jose Iguaran Fernandez , Juan Manuel Perez , German Rosati

Social Media platforms have offered invaluable opportunities for linguistic research. The availability of up-to-date data, coming from any part in the world, and coming from natural contexts, has allowed researchers to study language in…

Computation and Language · Computer Science 2024-07-23 Simon Gonzalez

Hate speech on social media is a growing concern, and automated methods have so far been sub-par at reliably detecting it. A major challenge lies in the potentially evasive nature of hate speech due to the ambiguity and fast evolution of…

Computation and Language · Computer Science 2021-03-17 Maximilian Kupi , Michael Bodnar , Nikolas Schmidt , Carlos Eduardo Posada

In this article we describe our participation in TASS 2019, a shared task aimed at the detection of sentiment polarity of Spanish tweets. We combined different representations such as bag-of-words, bag-of-characters, and tweet embeddings.…

Computation and Language · Computer Science 2019-09-26 Franco M. Luque

Objective criteria for universal semantic components that distinguish a humorous utterance from a non-humorous one are presently under debate. In this article, we give an in-depth observation of our system of self-paced reading for…

Computation and Language · Computer Science 2024-07-11 Elena Mikhalkova , Nadezhda Ganzherli , Julia Murzina

This paper presents a large-scale corpus for non-task-oriented dialogue response selection, which contains over 27K distinct prompts more than 82K responses collected from social media. To annotate this corpus, we define a 5-grade rating…

Computation and Language · Computer Science 2018-05-16 Jing Li , Yan Song , Haisong Zhang , Shuming Shi

Sarcasm is the use of words usually used to either mock or annoy someone, or for humorous purposes. Sarcasm is largely used in social networks and microblogging websites, where people mock or censure in a way that makes it difficult even…

Computation and Language · Computer Science 2023-02-07 Alif Tri Handoyo , Hidayaturrahman , Derwin Suhartono

Reference texts such as encyclopedias and news articles can manifest biased language when objective reporting is substituted by subjective writing. Existing methods to detect bias mostly rely on annotated data to train machine learning…

Computation and Language · Computer Science 2021-12-20 Timo Spinde , David Krieger , Manuel Plank , Bela Gipp

Well-defined jokes can be divided neatly into a setup and a punchline. While most works on humor today talk about a joke as a whole, the idea of generating punchlines to a setup has applications in conversational humor, where funny remarks…

Computation and Language · Computer Science 2021-03-02 Tanishq Chaudhary , Mayank Goel , Radhika Mamidi

Fake news are affecting a large proportion of the population even becoming a danger to the society. Mostly, this disinformation flow take place through Internet. Being aware of that problem, in this work we propose a synthetic indicator…

Social and Information Networks · Computer Science 2022-01-24 Aarón López-García , Rafael Benítez

Parody is a figurative device used to imitate an entity for comedic or critical purposes and represents a widespread phenomenon in social media through many popular parody accounts. In this paper, we present the first computational study of…

Computation and Language · Computer Science 2020-05-04 Antonis Maronikolakis , Danae Sanchez Villegas , Daniel Preotiuc-Pietro , Nikolaos Aletras

Well-annotated data is a prerequisite for good Natural Language Processing models. Too often, though, annotation decisions are governed by optimizing time or annotator agreement. We make a case for nuanced efforts in an interdisciplinary…

Computation and Language · Computer Science 2022-10-31 Federico Bianchi , Stefanie Anja Hills , Patricia Rossini , Dirk Hovy , Rebekah Tromble , Nava Tintarev

A major challenge in paraphrase research is the lack of parallel corpora. In this paper, we present a new method to collect large-scale sentential paraphrases from Twitter by linking tweets through shared URLs. The main advantage of our…

Computation and Language · Computer Science 2017-08-02 Wuwei Lan , Siyu Qiu , Hua He , Wei Xu

As the interaction over the web has increased, incidents of aggression and related events like trolling, cyberbullying, flaming, hate speech, etc. too have increased manifold across the globe. While most of these behaviour like bullying or…

Computation and Language · Computer Science 2018-03-28 Ritesh Kumar , Aishwarya N. Reganti , Akshit Bhatia , Tushar Maheshwari

Information Extraction is a well-researched area of Natural Language Processing with applications in web search and question answering concerned with identifying entities and relationships between them as expressed in a given context,…

Information Retrieval · Computer Science 2020-11-17 Erin Macdonald , Denilson Barbosa

This paper presents Hotter and Colder, a dataset designed to analyze various types of online behavior in Icelandic blog comments. Building on previous work, we used GPT-4o mini to annotate approximately 800,000 comments for 25 tasks,…

Computation and Language · Computer Science 2025-02-25 Steinunn Rut Friðriksdóttir , Dan Saattrup Nielsen , Hafsteinn Einarsson