English
Related papers

Related papers: Multilingual Topic Classification in X: Dataset an…

200 papers

Many data sets contain rich information about objects, as well as pairwise relations between them. For instance, in networks of websites, scientific papers, and other documents, each node has content consisting of a collection of words, as…

Machine Learning · Computer Science 2014-10-30 Yaojia Zhu , Xiaoran Yan , Lise Getoor , Cristopher Moore

Offensive language detection is one of the most challenging problem in the natural language processing field, being imposed by the rising presence of this phenomenon in online social media. This paper describes our Transformer-based…

Computation and Language · Computer Science 2020-10-28 Mircea-Adrian Tanase , Dumitru-Clementin Cercel , Costin-Gabriel Chiru

Wikipedia serves as a globally accessible knowledge source with content in over 300 languages. Despite covering the same topics, the different versions of Wikipedia are written and updated independently. This leads to factual…

Computation and Language · Computer Science 2026-05-19 Silvia Cappa , Lingxiao Kong , Pille-Riin Peet , Fanfu Wei , Yuchen Zhou , Jan-Christoph Kalo

This work introduces a machine translation task where the output is aimed at audiences of different levels of target language proficiency. We collect a high quality dataset of news articles available in English and Spanish, written for…

Computation and Language · Computer Science 2019-11-05 Sweta Agrawal , Marine Carpuat

In this paper we share findings from our effort to build practical machine translation (MT) systems capable of translating across over one thousand languages. We describe results in three research domains: (i) Building clean, web-mined…

Traditional multitask learning methods basically can only exploit common knowledge in task- or language-wise, which lose either cross-language or cross-task knowledge. This paper proposes a general multilingual multitask model, named…

Computation and Language · Computer Science 2023-06-29 Zhangyin Feng , Yong Dai , Fan Zhang , Duyu Tang , Xiaocheng Feng , Shuangzhi Wu , Bing Qin , Yunbo Cao , Shuming Shi

In today's world, we follow news which is distributed globally. Significant events are reported by different sources and in different languages. In this work, we address the problem of tracking of events in a large multilingual stream.…

Information Retrieval · Computer Science 2015-12-23 Jan Rupnik , Andrej Muhic , Gregor Leban , Primoz Skraba , Blaz Fortuna , Marko Grobelnik

Topic modeling is a widely used approach for analyzing and exploring large document collections. Recent research efforts have incorporated pre-trained contextualized language models, such as BERT embeddings, into topic modeling. However,…

Computation and Language · Computer Science 2025-02-18 Suman Adhya , Debarshi Kumar Sanyal

Dynamic topic modeling is widely used to analyze evolving trends in scientific literature, medical records, and social media. Traditional topic models represent each topic through a single probability vector on the multinomial simplex and…

Machine Learning · Computer Science 2026-05-28 Hanjia Gao , Hanwen Ye , Qing Nie , Annie Qu

Probabilistic topic models are a powerful tool for extracting latent themes from large text datasets. In many text datasets, we also observe per-document covariates (e.g., source, style, political affiliation) that act as environments that…

Computation and Language · Computer Science 2024-11-04 Dominic Sobhani , Amir Feder , David Blei

We introduce the Multi30K dataset to stimulate multilingual multimodal research. Recent advances in image description have been demonstrated on English-language datasets almost exclusively, but image description should not be limited to…

Computation and Language · Computer Science 2016-05-03 Desmond Elliott , Stella Frank , Khalil Sima'an , Lucia Specia

Multiple business scenarios require an automated generation of descriptive human-readable text from structured input data. Hence, fact-to-text generation systems have been developed for various downstream tasks like generating soccer…

Computation and Language · Computer Science 2022-09-26 Shivprasad Sagare , Tushar Abhishek , Bhavyajeet Singh , Anubhav Sharma , Manish Gupta , Vasudeva Varma

In this paper, we address the problem of detection, classification and quantification of emotions of text in any form. We consider English text collected from social media like Twitter, which can provide information having utility in a…

Social and Information Networks · Computer Science 2019-06-13 Bharat Gaind , Varun Syal , Sneha Padgalwar

We test the hypothesis that the extent to which one obtains information on a given topic through Wikipedia depends on the language in which it is consulted. Controlling the size factor, we investigate this hypothesis for a number of 25…

Computation and Language · Computer Science 2021-06-01 Alexander Mehler , Wahed Hemati , Pascal Welke , Maxim Konca , Tolga Uslu

Fact-checkers are often hampered by the sheer amount of online content that needs to be fact-checked. NLP can help them by retrieving already existing fact-checks relevant to the content being investigated. This paper introduces a new…

The widespread use of social media necessitates reliable and efficient detection of offensive content to mitigate harmful effects. Although sophisticated models perform well on individual datasets, they often fail to generalize due to…

Computation and Language · Computer Science 2024-10-08 Huy Nghiem , Hal Daumé

As large language models (LLMs) are employed worldwide, existing evaluation paradigms for their multilingual capabilities primarily focus on factual task performance, neglecting the ability to judge content's deep-level values across…

Computation and Language · Computer Science 2026-05-12 Yukun Chen , Xinyu Zhang , Boyi Deng , Jialong Tang , Yu Wan , Fei Huang , Yuxi Zhou , Baosong Yang , Yiming Li

The increasing prevalence of mental disorders globally highlights the urgent need for effective digital screening methods that can be used in multilingual contexts. Most existing studies, however, focus on English data, overlooking critical…

Computation and Language · Computer Science 2026-01-27 Ana-Maria Bucur , Marcos Zampieri , Tharindu Ranasinghe , Fabio Crestani

The rise in popularity and ubiquity of Twitter has made sentiment analysis of tweets an important and well-covered area of research. However, the 140 character limit imposed on tweets makes it hard to use standard linguistic methods for…

Social and Information Networks · Computer Science 2021-01-05 Soroush Vosoughi , Helen Zhou , Deb Roy

This paper presents StoryDB - a broad multi-language dataset of narratives. StoryDB is a corpus of texts that includes stories in 42 different languages. Every language includes 500+ stories. Some of the languages include more than 20 000…

Computation and Language · Computer Science 2022-11-15 Alexey Tikhonov , Igor Samenko , Ivan P. Yamshchikov
‹ Prev 1 8 9 10 Next ›