中文
相关论文

相关论文: PoPreRo: A New Dataset for Popularity Prediction o…

200 篇论文

We introduce RoDia, the first dataset for Romanian dialect identification from speech. The RoDia dataset includes a varied compilation of speech samples from five distinct regions of Romania, covering both urban and rural environments,…

计算与语言 · 计算机科学 2024-03-22 Codrut Rotaru , Nicolae-Catalin Ristea , Radu Tudor Ionescu

Social media creates crucial mass changes, as popular posts and opinions cast a significant influence on users' decisions and thought processes. For example, the recent Reddit uprising inspired by r/wallstreetbets which had remarkable…

机器学习 · 计算机科学 2021-06-18 Juno Kim

Authorship profiling is the process of identifying an author's characteristics based on their writings. This centuries old problem has become more intriguing especially with recent developments in Natural Language Processing (NLP). In this…

计算与语言 · 计算机科学 2025-03-20 Ecaterina Ştefănescu , Alexandru-Iulius Jerpelea

Using supervised automatic summarisation methods requires sufficient corpora that include pairs of documents and their summaries. Similarly to many tasks in natural language processing, most of the datasets available for summarization are…

In this work, we introduce a corpus for satire detection in Romanian news. We gathered 55,608 public news articles from multiple real and satirical news sources, composing one of the largest corpora for satire detection regardless of…

计算与语言 · 计算机科学 2021-07-01 Ana-Cristina Rogoz , Mihaela Gaman , Radu Tudor Ionescu

Personality and demographics are important variables in social sciences, while in NLP they can aid in interpretability and removal of societal biases. However, datasets with both personality and demographic labels are scarce. To address…

计算与语言 · 计算机科学 2021-06-09 Matej Gjurković , Mladen Karan , Iva Vukojević , Mihaela Bošnjak , Jan Šnajder

Being around for decades, the problem of Authorship Attribution is still very much in focus currently. Some of the more recent instruments used are the pre-trained language models, the most prevalent being BERT. Here we used such a model to…

人工智能 · 计算机科学 2023-01-31 Sanda-Maria Avram

A major task for moderators of online spaces is norm-setting, essentially creating shared norms for user behavior in their communities. Platform design principles emphasize the importance of highlighting norm-adhering examples and…

人机交互 · 计算机科学 2026-04-07 Agam Goyal , Charlotte Lambert , Yoshee Jain , Eshwar Chandrasekharan

Sequential recommenders are crucial to the success of online applications, \eg e-commerce, video streaming, and social media. While model architectures continue to improve, for every new application domain, we still have to train a new…

信息检索 · 计算机科学 2025-07-01 Junting Wang , Praneet Rathi , Hari Sundaram

Multiple studies have focused on predicting the prospective popularity of an online document as a whole, without paying attention to the contributions of its individual parts. We introduce the task of proactively forecasting popularities of…

计算与语言 · 计算机科学 2023-01-03 Sayar Ghosh Roy , Anshul Padhi , Risubh Jain , Manish Gupta , Vasudeva Varma

Understanding the sociodemographic composition of online platforms is essential for accurately interpreting digital behavior and its societal implications. Yet, current methods often lack the transparency and reliability required, risking…

社会与信息网络 · 计算机科学 2025-11-04 Federico Cinus , Corrado Monti , Paolo Bajardi , Gianmarco De Francisci Morales

Developing natural language processing (NLP) systems for social media analysis remains an important topic in artificial intelligence research. This article introduces RoBERTweet, the first Transformer architecture trained on Romanian…

计算与语言 · 计算机科学 2023-06-13 Iulian-Marius Tăiatu , Andrei-Marius Avram , Dumitru-Clementin Cercel , Florin Pop

This work introduces HistNERo, the first Romanian corpus for Named Entity Recognition (NER) in historical newspapers. The dataset contains 323k tokens of text, covering more than half of the 19th century (i.e., 1817) until the late part of…

Romanian is one of the understudied languages in computational linguistics, with few resources available for the development of natural language processing tools. In this paper, we introduce LaRoSeDa, a Large Romanian Sentiment Data Set,…

计算与语言 · 计算机科学 2021-01-13 Anca Maria Tache , Mihaela Gaman , Radu Tudor Ionescu

Satire, irony, and sarcasm are techniques typically used to express humor and critique, rather than deceive; however, they can occasionally be mistaken for factual reporting, akin to fake news. These techniques can be applied at a more…

计算与语言 · 计算机科学 2025-10-16 Răzvan-Alexandru Smădu , Andreea Iuga , Dumitru-Clementin Cercel , Florin Pop

In this paper we introduce a new natural language processing dataset and benchmark for predicting prosodic prominence from written text. To our knowledge this will be the largest publicly available dataset with prosodic labels. We describe…

计算与语言 · 计算机科学 2019-08-07 Aarne Talman , Antti Suni , Hande Celikkanat , Sofoklis Kakouros , Jörg Tiedemann , Martti Vainio

Satire and fake news can both contribute to the spread of false information, even though both have different purposes (one if for amusement, the other is to misinform). However, it is not enough to rely purely on text to detect the…

计算与语言 · 计算机科学 2025-04-11 Răzvan-Alexandru Smădu , Andreea Iuga , Dumitru-Clementin Cercel

Computational inference of aesthetics is an ill-defined task due to its subjective nature. Many datasets have been proposed to tackle the problem by providing pairs of images and aesthetic scores based on human ratings. However, humans are…

计算机视觉与模式识别 · 计算机科学 2022-09-22 Daniel Vera Nieto , Luigi Celona , Clara Fernandez-Labrador

We present a Few-Shot Relation Classification Dataset (FewRel), consisting of 70, 000 sentences on 100 relations derived from Wikipedia and annotated by crowdworkers. The relation of each sentence is first recognized by distant supervision…

机器学习 · 计算机科学 2018-10-30 Xu Han , Hao Zhu , Pengfei Yu , Ziyun Wang , Yuan Yao , Zhiyuan Liu , Maosong Sun

Conventional algorithms for training language models (LMs) with human feedback rely on preferences that are assumed to account for an "average" user, disregarding subjectivity and finer-grained variations. Recent studies have raised…

计算与语言 · 计算机科学 2024-10-22 Sachin Kumar , Chan Young Park , Yulia Tsvetkov , Noah A. Smith , Hannaneh Hajishirzi
‹ 上一页 1 2 3 10 下一页 ›