中文
相关论文

相关论文: Evaluating Cross-Lingual Classification Approaches…

200 篇论文

We propose a new computational approach for tracking and detecting statistically significant linguistic shifts in the meaning and usage of words. Such linguistic shifts are especially prevalent on the Internet, where the rapid exchange of…

计算与语言 · 计算机科学 2014-11-13 Vivek Kulkarni , Rami Al-Rfou , Bryan Perozzi , Steven Skiena

With the widespread use of social networks, detecting the topics discussed on these platforms has become a significant challenge. Current approaches primarily rely on frequent pattern mining or semantic relations, often neglecting the…

计算与语言 · 计算机科学 2024-08-22 Mehrdad Ranjbar Khadivi , Shahin Akbarpour , Mohammad-Reza Feizi-Derakhshi , Babak Anari

The context-dependent nature of online aggression makes annotating large collections of data extremely difficult. Previously studied datasets in abusive language detection have been insufficient in size to efficiently train deep learning…

计算与语言 · 计算机科学 2018-08-31 Younghun Lee , Seunghyun Yoon , Kyomin Jung

In the recent past, social media platforms have helped people in connecting and communicating to a wider audience. But this has also led to a drastic increase in cyberbullying. It is essential to detect and curb hate speech to keep the…

计算与语言 · 计算机科学 2021-11-02 Ravindra Nayak , Raviraj Joshi

This paper introduces a large collection of time series data derived from Twitter, postprocessed using word embedding techniques, as well as specialized fine-tuned language models. This data comprises the past five years and captures…

In this paper, we address the problem of detection, classification and quantification of emotions of text in any form. We consider English text collected from social media like Twitter, which can provide information having utility in a…

社会与信息网络 · 计算机科学 2019-06-13 Bharat Gaind , Varun Syal , Sneha Padgalwar

Social-media data provides increasing opportunities for automated analysis of large sets of textual documents. So far, automated tools have been developed to account for either the social networks between the participants of the debates, or…

数字图书馆 · 计算机科学 2017-11-23 Iina Hellsten , Loet Leydesdorff

COVID-19 has created a major public health problem worldwide and other problems such as economic crisis, unemployment, mental distress, etc. The pandemic is deadly in the world and involves many people not only with infection but also with…

机器学习 · 计算机科学 2025-11-25 Khandaker Tayef Shahriar , Iqbal H. Sarker

Statements on social media can be analysed to identify individuals who are experiencing red flag medical symptoms, allowing early detection of the spread of disease such as influenza. Since disease does not respect cultural borders and may…

计算与语言 · 计算机科学 2019-10-11 Mattias Appelgren , Patrick Schrempf , Matúš Falis , Satoshi Ikeda , Alison Q O'Neil

Clustering news across languages enables efficient media monitoring by aggregating articles from multilingual sources into coherent stories. Doing so in an online setting allows scalable processing of massive news streams. To this end, we…

计算与语言 · 计算机科学 2018-09-05 Sebastião Miranda , Artūrs Znotiņš , Shay B. Cohen , Guntis Barzdins

Hashtag segmentation, also known as hashtag decomposition, is a common step in preprocessing pipelines for social media datasets. It usually precedes tasks such as sentiment analysis and hate speech detection. For sentiment analysis in…

In applications involving conversational speech, data sparsity is a limiting factor in building a better language model. We propose a simple, language-independent method to quickly harvest large amounts of data from Twitter to supplement a…

计算与语言 · 计算机科学 2015-04-13 Aaron Jaech , Mari Ostendorf

The identification and classification of political claims is an important step in the analysis of political newspaper reports; however, resources for this task are few and far between. This paper explores different strategies for the…

计算与语言 · 计算机科学 2023-10-16 Urs Zaberer , Sebastian Padó , Gabriella Lapesa

Applying natural language processing for mining and intelligent information access to tweets (a form of microblog) is a challenging, emerging research area. Unlike carefully authored news text and other longer content, tweets pose a number…

Sentiment analysis possesses the potential of diverse applicability on digital platforms. Sentiment analysis extracts the polarity to understand the intensity and subjectivity in the text. This work uses a lexicon-based method to perform…

计算与语言 · 计算机科学 2024-09-20 Muhammad Raees , Samina Fazilat

State-of-the-art natural language processing systems rely on supervision in the form of annotated data to learn competent models. These models are generally trained on data in a single language (usually English), and cannot be directly used…

Multilingual semantic parsing is a cost-effective method that allows a single model to understand different languages. However, researchers face a great imbalance of availability of training data, with English being resource rich, and other…

计算与语言 · 计算机科学 2021-06-15 Menglin Xia , Emilio Monti

Automatic identification of hateful and abusive content is vital in combating the spread of harmful online content and its damaging effects. Most existing works evaluate models by examining the generalization error on train-test splits on…

计算与语言 · 计算机科学 2025-04-07 Lanqin Yuan , Marian-Andrei Rizoiu

Massive web-crawled image-text datasets lay the foundation for recent progress in multimodal learning. These datasets are designed with the goal of training a model to do well on standard computer vision benchmarks, many of which, however,…

计算机视觉与模式识别 · 计算机科学 2025-09-16 Thao Nguyen , Matthew Wallingford , Sebastin Santy , Wei-Chiu Ma , Sewoong Oh , Ludwig Schmidt , Pang Wei Koh , Ranjay Krishna

With the growing use of social media and its availability, many instances of the use of offensive language have been observed across multiple languages and domains. This phenomenon has given rise to the growing need to detect the offensive…

计算与语言 · 计算机科学 2020-07-09 Kartikey Pant , Tanvi Dadu