English
Related papers

Related papers: Text Augmentations with R-drop for Classification …

200 papers

Identifying informative tweets is an important step when building information extraction systems based on social media. WNUT-2020 Task 2 was organised to recognise informative tweets from noise tweets. In this paper, we present our approach…

Computation and Language · Computer Science 2020-10-13 Hansi Hettiarachchi , Tharindu Ranasinghe

Words are malleable objects, influenced by events that are reflected in written texts. Situated in the global outbreak of COVID-19, our research aims at detecting semantic shifts in social media language triggered by the health crisis. With…

Computation and Language · Computer Science 2021-02-17 Yanzhu Guo , Christos Xypolopoulos , Michalis Vazirgiannis

To analyse large numbers of texts, social science researchers are increasingly confronting the challenge of text classification. When manual labeling is not possible and researchers have to find automatized ways to classify texts, computer…

Computation and Language · Computer Science 2023-10-10 Karina Shyrokykh , Maksym Girnyk , Lisa Dellmuth

Unsupervised pre-training has led to much recent progress in natural language understanding. In this paper, we study self-training as another way to leverage unlabeled data through semi-supervised learning. To obtain additional data for a…

Computation and Language · Computer Science 2020-10-06 Jingfei Du , Edouard Grave , Beliz Gunel , Vishrav Chaudhary , Onur Celebi , Michael Auli , Ves Stoyanov , Alexis Conneau

This paper formulates the problem of dynamically identifying key topics with proper labels from COVID-19 Tweets to provide an overview of wider public opinion. Nowadays, social media is one of the best ways to connect people through…

Information Retrieval · Computer Science 2021-09-07 Khandaker Tayef Shahriar , Iqbal H. Sarker , Muhammad Nazrul Islam , Mohammad Ali Moni

We analyze the process of creating word embedding feature representations designed for a learning task when annotated data is scarce, for example, in depressive language detection from Tweets. We start with a rich word embedding pre-trained…

Computation and Language · Computer Science 2021-06-25 Nawshad Farruque , Randy Goebel , Osmar Zaiane

As the Covid-19 outbreaks rapidly all over the world day by day and also affects the lives of million, a number of countries declared complete lock-down to check its intensity. During this lockdown period, social media plat-forms have…

Computation and Language · Computer Science 2021-06-15 Arunava Kumar Chakraborty , Sourav Das , Anup Kumar Kolya

This paper introduces a simple and effective form of data augmentation for recommender systems. A paraphrase similarity model is applied to widely available textual data, such as reviews and product descriptions, yielding new semantic…

Computation and Language · Computer Science 2021-09-21 Federico López , Martin Scholz , Jessica Yung , Marie Pellat , Michael Strube , Lucas Dixon

Pre-trained language model word representation, such as BERT, have been extremely successful in several Natural Language Processing tasks significantly improving on the state-of-the-art. This can largely be attributed to their ability to…

Computation and Language · Computer Science 2020-08-20 Wah Meng Lim , Harish Tayyar Madabushi

Since the beginning of coronavirus, the disease has spread worldwide and drastically changed many aspects of the human's lifestyle. Twitter as a powerful tool can help researchers measure public health in response to COVID-19. According to…

Computation and Language · Computer Science 2021-10-15 Mohamad Zamini

We propose a novel data augmentation for labeled sentences called contextual augmentation. We assume an invariance that sentences are natural even if the words in the sentences are replaced with other words with paradigmatic relations. We…

Computation and Language · Computer Science 2018-05-17 Sosuke Kobayashi

The emergence of cross-modal foundation models has introduced numerous approaches grounded in text-image retrieval. However, on some domain-specific retrieval tasks, these models fail to focus on the key attributes required. To address this…

Computer Vision and Pattern Recognition · Computer Science 2023-06-13 Yuguang Yang , Yiming Wang , Shupeng Geng , Runqi Wang , Yimi Wang , Sheng Wu , Baochang Zhang

Large language models (LLMs) play a crucial role in natural language processing (NLP) tasks, improving the understanding, generation, and manipulation of human language across domains such as translating, summarizing, and classifying text.…

Computation and Language · Computer Science 2025-03-04 Anna Glazkova , Olga Zakharova

The BioCreative VII Track 3 challenge focused on the identification of medication names in Twitter user timelines. For our submission to this challenge, we expanded the available training data by using several data augmentation techniques.…

Computation and Language · Computer Science 2021-11-15 Igor Kulev , Berkay Köprü , Raul Rodriguez-Esteban , Diego Saldana , Yi Huang , Alessandro La Torraca , Elif Ozkirimli

Text classification, as the task consisting in assigning categories to textual instances, is a very common task in information science. Methods learning distributed representations of words, such as word embeddings, have become popular in…

Computation and Language · Computer Science 2020-12-15 Arkaitz Zubiaga

Social media is an useful platform to share health-related information due to its vast reach. This makes it a good candidate for public-health monitoring tasks, specifically for pharmacovigilance. We study the problem of extraction of…

Information Retrieval · Computer Science 2017-09-07 Shashank Gupta , Sachin Pawar , Nitin Ramrakhiyani , Girish Palshikar , Vasudeva Varma

Research on data generation and augmentation has been focused majorly on enhancing generation models, leaving a notable gap in the exploration and refinement of methods for evaluating synthetic data. There are several text similarity…

Computation and Language · Computer Science 2023-11-09 Tiasa Singha Roy , Priyam Basu

This paper presents our submission to Task 2 of the Workshop on Noisy User-generated Text. We explore improving the performance of a pre-trained transformer-based language model fine-tuned for text classification through an ensemble…

Computation and Language · Computer Science 2020-10-19 Calum Perrio , Harish Tayyar Madabushi

The outbreak COVID-19 virus caused a significant impact on the health of people all over the world. Therefore, it is essential to have a piece of constant and accurate information about the disease with everyone. This paper describes our…

Computation and Language · Computer Science 2021-04-02 Tin Van Huynh , Luan Thanh Nguyen , Son T. Luu

We address claim normalization for multilingual misinformation detection - transforming noisy social media posts into clear, verifiable statements across 20 languages. The key contribution demonstrates how systematic decomposition of posts…

‹ Prev 1 4 5 6 7 8 10 Next ›