English
Related papers

Related papers: Improving Large-scale Paraphrase Acquisition and G…

200 papers

Automatically associating social media posts with topics is an important prerequisite for effective search and recommendation on many social media platforms. However, topic classification of such posts is quite challenging because of (a) a…

Computation and Language · Computer Science 2022-05-04 Vivek Kulkarni , Kenny Leung , Aria Haghighi

In online domain-specific customer service applications, many companies struggle to deploy advanced NLP models successfully, due to the limited availability of and noise in their datasets. While prior research demonstrated the potential of…

Computation and Language · Computer Science 2021-04-19 Amir Hadifar , Sofie Labat , Véronique Hoste , Chris Develder , Thomas Demeester

Existing paraphrase identification datasets lack sentence pairs that have high lexical overlap without being paraphrases. Models trained on such data fail to distinguish pairs like flights from New York to Florida and flights from Florida…

Computation and Language · Computer Science 2019-04-03 Yuan Zhang , Jason Baldridge , Luheng He

On social media platforms like Twitter, users regularly share their opinions and comments with software vendors and service providers. Popular software products might get thousands of user comments per day. Research has shown that such…

Software Engineering · Computer Science 2021-08-20 Christoph Stanik , Tim Pietz , Walid Maalej

Neural text generation, including neural machine translation, image captioning, and summarization, has been quite successful recently. However, during training time, typically only one reference is considered for each example, even though…

Computation and Language · Computer Science 2018-08-30 Renjie Zheng , Mingbo Ma , Liang Huang

A major hurdle on the road to conversational interfaces is the difficulty in collecting data that maps language utterances to logical forms. One prominent approach for data collection has been to automatically generate pseudo-language…

Computation and Language · Computer Science 2019-08-30 Jonathan Herzig , Jonathan Berant

Building a benchmark dataset for hate speech detection presents various challenges. Firstly, because hate speech is relatively rare, random sampling of tweets to annotate is very inefficient in finding hate speech. To address this, prior…

Computation and Language · Computer Science 2021-11-11 Md Mustafizur Rahman , Dinesh Balakrishnan , Dhiraj Murthy , Mucahid Kutlu , Matthew Lease

Twitter has grown to become an important platform to access immediate information about major events and dynamic topics. As one example, recent work has shown that classifiers trained to detect topical content on Twitter can generalize well…

Information Retrieval · Computer Science 2020-01-28 Kasra Safari , Scott Sanner

Acquiring training data to improve the robustness of dialog systems can be a painstakingly long process. In this work, we propose a method to reduce the cost and effort of creating new conversational agents by artificially generating more…

Computation and Language · Computer Science 2022-05-05 Louis Marceau , Raouf Belbahar , Marc Queudot , Nada Naji , Eric Charton , Marie-Jean Meurs

Identifying factors that make ad text attractive is essential for advertising success. This study proposes AdParaphrase v2.0, a dataset for ad text paraphrasing, containing human preference data, to enable the analysis of the linguistic…

Computation and Language · Computer Science 2025-05-28 Soichiro Murakami , Peinan Zhang , Hidetaka Kamigaito , Hiroya Takamura , Manabu Okumura

An intuitive way for a human to write paraphrase sentences is to replace words or phrases in the original sentence with their corresponding synonyms and make necessary changes to ensure the new sentences are fluent and grammatically…

Computation and Language · Computer Science 2018-06-22 Shaohan Huang , Yu Wu , Furu Wei , Ming Zhou

Keyphrase generation is the task of summarizing the contents of any given article into a few salient phrases (or keyphrases). Existing works for the task mostly rely on large-scale annotated datasets, which are not easy to acquire. Very few…

Computation and Language · Computer Science 2023-05-30 Krishna Garg , Jishnu Ray Chowdhury , Cornelia Caragea

We present a new benchmark dataset called PARADE for paraphrase identification that requires specialized domain knowledge. PARADE contains paraphrases that overlap very little at the lexical and syntactic level but are semantically…

Computation and Language · Computer Science 2020-10-09 Yun He , Zhuoer Wang , Yin Zhang , Ruihong Huang , James Caverlee

Keyphrase generation aims at generating important phrases (keyphrases) that best describe a given document. In scholarly domains, current approaches have largely used only the title and abstract of the articles to generate keyphrases. In…

Computation and Language · Computer Science 2022-10-24 Krishna Garg , Jishnu Ray Chowdhury , Cornelia Caragea

This paper describes our approach to hierarchical multi-label detection of persuasion techniques in meme texts. Our model, developed as a part of the recent SemEval task, is based on fine-tuning individual language models (BERT,…

Computation and Language · Computer Science 2024-07-04 Kota Shamanth Ramanath Nayak , Leila Kosseim

With the rise in popularity of public social media and micro-blogging services, most notably Twitter, the people have found a venue to hear and be heard by their peers without an intermediary. As a consequence, and aided by the public…

Computation and Language · Computer Science 2016-06-21 Prashanth Vijayaraghavan , Soroush Vosoughi , Deb Roy

Ironies can not only express stronger emotions but also show a sense of humor. With the development of social media, ironies are widely used in public. Although many prior research studies have been conducted in irony detection, few studies…

Computation and Language · Computer Science 2019-09-17 Mengdi Zhu , Zhiwei Yu , Xiaojun Wan

Our paper studies the predictability of online speech -- that is, how well language models learn to model the distribution of user generated content on X (previously Twitter). We define predictability as a measure of the model's…

Computation and Language · Computer Science 2026-01-07 Mina Remeli , Moritz Hardt , Robert C. Williamson

Paraphrasing is a useful natural language processing task that can contribute to more diverse generated or translated texts. Natural language inference (NLI) and paraphrasing share some similarities and can benefit from a joint approach. We…

Computation and Language · Computer Science 2021-11-16 Matej Klemen , Marko Robnik-Šikonja

In this paper, we propose a novel neural approach for paraphrase generation. Conventional para- phrase generation methods either leverage hand-written rules and thesauri-based alignments, or use statistical machine learning principles. To…

Computation and Language · Computer Science 2016-10-14 Aaditya Prakash , Sadid A. Hasan , Kathy Lee , Vivek Datla , Ashequl Qadir , Joey Liu , Oladimeji Farri