English
Related papers

Related papers: SaRoCo: Detecting Satire in a Novel Romanian Corpu…

200 papers

Satire and fake news can both contribute to the spread of false information, even though both have different purposes (one if for amusement, the other is to misinform). However, it is not enough to rely purely on text to detect the…

Computation and Language · Computer Science 2025-04-11 Răzvan-Alexandru Smădu , Andreea Iuga , Dumitru-Clementin Cercel

Satire, irony, and sarcasm are techniques typically used to express humor and critique, rather than deceive; however, they can occasionally be mistaken for factual reporting, akin to fake news. These techniques can be applied at a more…

Computation and Language · Computer Science 2025-10-16 Răzvan-Alexandru Smădu , Andreea Iuga , Dumitru-Clementin Cercel , Florin Pop

The primary goal of a news headline is to summarize an event in as few words as possible. Depending on the media outlet, a headline can serve as a means to objectively deliver a summary or improve its visibility. For the latter, specific…

Computation and Language · Computer Science 2025-09-03 Mihnea-Alexandru Vîrlan , Răzvan-Alexandru Smădu , Dumitru-Clementin Cercel , Florin Pop , Mihaela-Claudia Cercel

Satire detection and sentiment analysis are intensively explored natural language processing (NLP) tasks that study the identification of the satirical tone from texts and extracting sentiments in relationship with their targets. In…

Computation and Language · Computer Science 2023-06-14 Sebastian-Vasile Echim , Răzvan-Alexandru Smădu , Andrei-Marius Avram , Dumitru-Clementin Cercel , Florin Pop

We built models with Logistic Regression and linear Support Vector Machines on a large dataset consisting of regular news articles and news from satirical websites, and showed that such linear classifiers on a corpus with about 60,000…

Computation and Language · Computer Science 2018-10-02 Andreas Stöckl

In this work, we introduce the MOldavian and ROmanian Dialectal COrpus (MOROCO), which is freely available for download at https://github.com/butnaruandrei/MOROCO. The corpus contains 33564 samples of text (with over 10 million tokens)…

Computation and Language · Computer Science 2019-06-04 Andrei M. Butnaru , Radu Tudor Ionescu

Using supervised automatic summarisation methods requires sufficient corpora that include pairs of documents and their summaries. Similarly to many tasks in natural language processing, most of the datasets available for summarization are…

To increase revenue, news websites often resort to using deceptive news titles, luring users into clicking on the title and reading the full news. Clickbait detection is the task that aims to automatically detect this form of false…

Computation and Language · Computer Science 2023-10-11 Daria-Mihaela Broscoteanu , Radu Tudor Ionescu

Psychological corpora in NLP are collections of texts used to analyze human psychology, emotions, and mental health. These texts allow researchers to study psychological constructs, identify patterns related to mental health problems and…

Computation and Language · Computer Science 2026-03-13 Alexandra Ciobotaru , Ana-Maria Bucur , Liviu P. Dinu

Satirical news detection is an important yet challenging task to prevent spread of misinformation. Many feature based and end-to-end neural nets based satirical news detection systems have been proposed and delivered promising results.…

Computation and Language · Computer Science 2024-04-16 Yue Zhou , Yan Zhang , JingTao Yao

Satirical news is considered to be entertainment, but it is potentially deceptive and harmful. Despite the embedded genre in the article, not everyone can recognize the satirical cues and therefore believe the news as true news. We observe…

Computation and Language · Computer Science 2017-09-06 Fan Yang , Arjun Mukherjee , Eduard Dragut

We develop novel annotation guidelines for sentence-level subjectivity detection, which are not limited to language-specific cues. We use our guidelines to collect NewsSD-ENG, a corpus of 638 objective and 411 subjective sentences extracted…

Satire detection is essential for accurately extracting opinions from textual data and combating misinformation online. However, the lack of diverse corpora for satire leads to the problem of stylistic bias which impacts the models'…

Computation and Language · Computer Science 2024-12-13 Asli Umay Ozturk , Recep Firat Cekinel , Pinar Karagoz

Resources for Grammatical Error Correction (GEC) in non-English languages are scarce, while available spellcheckers in these languages are mostly limited to simple corrections and rules. In this paper we introduce a first GEC corpus for…

Computation and Language · Computer Science 2026-04-28 Teodor-Mihai Cotet , Stefan Ruseti , Mihai Dascalu

We introduce the Self-Annotated Reddit Corpus (SARC), a large corpus for sarcasm research and for training and evaluating systems for sarcasm detection. The corpus has 1.3 million sarcastic statements -- 10 times more than any previous…

Computation and Language · Computer Science 2018-03-26 Mikhail Khodak , Nikunj Saunshi , Kiran Vodrahalli

This work introduces HistNERo, the first Romanian corpus for Named Entity Recognition (NER) in historical newspapers. The dataset contains 323k tokens of text, covering more than half of the 19th century (i.e., 1817) until the late part of…

Sarcasm Detection has enjoyed great interest from the research community, however the task of predicting sarcasm in a text remains an elusive problem for machines. Past studies mostly make use of twitter datasets collected using hashtag…

Machine Learning · Computer Science 2022-10-17 Rishabh Misra , Prahal Arora

Online media aim for reaching ever bigger audience and for attracting ever longer attention span. This competition creates an environment that rewards sensational, fake, and toxic news. To help limit their spread and impact, we propose and…

Computation and Language · Computer Science 2019-08-27 Yoan Dinkov , Ivan Koychev , Preslav Nakov

We introduce PoPreRo, the first dataset for Popularity Prediction of Romanian posts collected from Reddit. The PoPreRo dataset includes a varied compilation of post samples from five distinct subreddits of Romania, totaling 28,107 data…

Computation and Language · Computer Science 2024-11-26 Ana-Cristina Rogoz , Maria Ilinca Nechita , Radu Tudor Ionescu

Satire is a form of humorous critique, but it is sometimes misinterpreted by readers as legitimate news, which can lead to harmful consequences. We observe that the images used in satirical news articles often contain absurd or ridiculous…

Computation and Language · Computer Science 2020-10-15 Lily Li , Or Levi , Pedram Hosseini , David A. Broniatowski
‹ Prev 1 2 3 10 Next ›