English
Related papers

Related papers: PoPreRo: A New Dataset for Popularity Prediction o…

200 papers

We introduce RoDia, the first dataset for Romanian dialect identification from speech. The RoDia dataset includes a varied compilation of speech samples from five distinct regions of Romania, covering both urban and rural environments,…

Computation and Language · Computer Science 2024-03-22 Codrut Rotaru , Nicolae-Catalin Ristea , Radu Tudor Ionescu

Social media creates crucial mass changes, as popular posts and opinions cast a significant influence on users' decisions and thought processes. For example, the recent Reddit uprising inspired by r/wallstreetbets which had remarkable…

Machine Learning · Computer Science 2021-06-18 Juno Kim

Authorship profiling is the process of identifying an author's characteristics based on their writings. This centuries old problem has become more intriguing especially with recent developments in Natural Language Processing (NLP). In this…

Computation and Language · Computer Science 2025-03-20 Ecaterina Ştefănescu , Alexandru-Iulius Jerpelea

Using supervised automatic summarisation methods requires sufficient corpora that include pairs of documents and their summaries. Similarly to many tasks in natural language processing, most of the datasets available for summarization are…

In this work, we introduce a corpus for satire detection in Romanian news. We gathered 55,608 public news articles from multiple real and satirical news sources, composing one of the largest corpora for satire detection regardless of…

Computation and Language · Computer Science 2021-07-01 Ana-Cristina Rogoz , Mihaela Gaman , Radu Tudor Ionescu

Personality and demographics are important variables in social sciences, while in NLP they can aid in interpretability and removal of societal biases. However, datasets with both personality and demographic labels are scarce. To address…

Computation and Language · Computer Science 2021-06-09 Matej Gjurković , Mladen Karan , Iva Vukojević , Mihaela Bošnjak , Jan Šnajder

Being around for decades, the problem of Authorship Attribution is still very much in focus currently. Some of the more recent instruments used are the pre-trained language models, the most prevalent being BERT. Here we used such a model to…

Artificial Intelligence · Computer Science 2023-01-31 Sanda-Maria Avram

A major task for moderators of online spaces is norm-setting, essentially creating shared norms for user behavior in their communities. Platform design principles emphasize the importance of highlighting norm-adhering examples and…

Human-Computer Interaction · Computer Science 2026-04-07 Agam Goyal , Charlotte Lambert , Yoshee Jain , Eshwar Chandrasekharan

Sequential recommenders are crucial to the success of online applications, \eg e-commerce, video streaming, and social media. While model architectures continue to improve, for every new application domain, we still have to train a new…

Information Retrieval · Computer Science 2025-07-01 Junting Wang , Praneet Rathi , Hari Sundaram

Multiple studies have focused on predicting the prospective popularity of an online document as a whole, without paying attention to the contributions of its individual parts. We introduce the task of proactively forecasting popularities of…

Computation and Language · Computer Science 2023-01-03 Sayar Ghosh Roy , Anshul Padhi , Risubh Jain , Manish Gupta , Vasudeva Varma

Understanding the sociodemographic composition of online platforms is essential for accurately interpreting digital behavior and its societal implications. Yet, current methods often lack the transparency and reliability required, risking…

Social and Information Networks · Computer Science 2025-11-04 Federico Cinus , Corrado Monti , Paolo Bajardi , Gianmarco De Francisci Morales

Developing natural language processing (NLP) systems for social media analysis remains an important topic in artificial intelligence research. This article introduces RoBERTweet, the first Transformer architecture trained on Romanian…

Computation and Language · Computer Science 2023-06-13 Iulian-Marius Tăiatu , Andrei-Marius Avram , Dumitru-Clementin Cercel , Florin Pop

This work introduces HistNERo, the first Romanian corpus for Named Entity Recognition (NER) in historical newspapers. The dataset contains 323k tokens of text, covering more than half of the 19th century (i.e., 1817) until the late part of…

Romanian is one of the understudied languages in computational linguistics, with few resources available for the development of natural language processing tools. In this paper, we introduce LaRoSeDa, a Large Romanian Sentiment Data Set,…

Computation and Language · Computer Science 2021-01-13 Anca Maria Tache , Mihaela Gaman , Radu Tudor Ionescu

Satire, irony, and sarcasm are techniques typically used to express humor and critique, rather than deceive; however, they can occasionally be mistaken for factual reporting, akin to fake news. These techniques can be applied at a more…

Computation and Language · Computer Science 2025-10-16 Răzvan-Alexandru Smădu , Andreea Iuga , Dumitru-Clementin Cercel , Florin Pop

In this paper we introduce a new natural language processing dataset and benchmark for predicting prosodic prominence from written text. To our knowledge this will be the largest publicly available dataset with prosodic labels. We describe…

Computation and Language · Computer Science 2019-08-07 Aarne Talman , Antti Suni , Hande Celikkanat , Sofoklis Kakouros , Jörg Tiedemann , Martti Vainio

Satire and fake news can both contribute to the spread of false information, even though both have different purposes (one if for amusement, the other is to misinform). However, it is not enough to rely purely on text to detect the…

Computation and Language · Computer Science 2025-04-11 Răzvan-Alexandru Smădu , Andreea Iuga , Dumitru-Clementin Cercel

Computational inference of aesthetics is an ill-defined task due to its subjective nature. Many datasets have been proposed to tackle the problem by providing pairs of images and aesthetic scores based on human ratings. However, humans are…

Computer Vision and Pattern Recognition · Computer Science 2022-09-22 Daniel Vera Nieto , Luigi Celona , Clara Fernandez-Labrador

We present a Few-Shot Relation Classification Dataset (FewRel), consisting of 70, 000 sentences on 100 relations derived from Wikipedia and annotated by crowdworkers. The relation of each sentence is first recognized by distant supervision…

Machine Learning · Computer Science 2018-10-30 Xu Han , Hao Zhu , Pengfei Yu , Ziyun Wang , Yuan Yao , Zhiyuan Liu , Maosong Sun

Conventional algorithms for training language models (LMs) with human feedback rely on preferences that are assumed to account for an "average" user, disregarding subjectivity and finer-grained variations. Recent studies have raised…

Computation and Language · Computer Science 2024-10-22 Sachin Kumar , Chan Young Park , Yulia Tsvetkov , Noah A. Smith , Hannaneh Hajishirzi
‹ Prev 1 2 3 10 Next ›