中文
相关论文

相关论文: NoticIA: A Clickbait Article Summarization Dataset…

200 篇论文

To advance understanding on how to engage readers, we advocate the novel task of automatic pull quote selection. Pull quotes are a component of articles specifically designed to catch the attention of readers with spans of text selected…

计算与语言 · 计算机科学 2020-10-15 Tanner Bohn , Charles X. Ling

We introduce WikiLingua, a large-scale, multilingual dataset for the evaluation of crosslingual abstractive summarization systems. We extract article and summary pairs in 18 languages from WikiHow, a high quality, collaborative resource of…

计算与语言 · 计算机科学 2020-10-08 Faisal Ladhak , Esin Durmus , Claire Cardie , Kathleen McKeown

More than 7,000 known languages are spoken around the world. However, due to the lack of annotated resources, only a small fraction of them are currently covered by speech technologies. Albeit self-supervised speech representations, recent…

计算机视觉与模式识别 · 计算机科学 2024-02-21 José-M. Acosta-Triana , David Gimeno-Gómez , Carlos-D. Martínez-Hinarejos

We introduce a new reading comprehension dataset, dubbed MultiWikiQA, which covers 306 languages and has 1,220,757 samples in total. We start with Wikipedia articles, which also provide the context for the dataset samples, and use an LLM to…

计算与语言 · 计算机科学 2026-03-05 Dan Saattrup Smart

This study presents a large multi-modal Bangla YouTube clickbait dataset consisting of 253,070 data points collected through an automated process using the YouTube API and Python web automation frameworks. The dataset contains 18 diverse…

机器学习 · 计算机科学 2023-10-19 Abdullah Al Imran , Md Sakib Hossain Shovon , M. F. Mridha

We present a computationally-grounded word similarity dataset based on two well-known Natural Language Processing resources; text corpora and knowledge bases. This dataset aims to fulfil a gap in psycholinguistic research by providing a…

计算与语言 · 计算机科学 2023-04-21 J. Goikoetxea , M. Arantzeta , I. San Martin

The use of alluring headlines (clickbait) to tempt the readers has become a growing practice nowadays. For the sake of existence in the highly competitive media industry, most of the on-line media including the mainstream ones, have started…

社会与信息网络 · 计算机科学 2017-03-29 Md Main Uddin Rony , Naeemul Hassan , Mohammad Yousuf

While the NLP community has produced numerous summarization benchmarks, none provide the rich annotations required to simultaneously address many important problems related to control and reliability. We introduce a Wikipedia-derived…

计算与语言 · 计算机科学 2023-12-05 Kundan Krishna , Prakhar Gupta , Sanjana Ramprasad , Byron C. Wallace , Jeffrey P. Bigham , Zachary C. Lipton

Clickbait is characterized by disproportionately high emotional intensity relative to informational content, often reinforced by specific structural patterns. However, current research considers clickbait as a static textual phenomenon…

计算与语言 · 计算机科学 2026-05-01 Syed Mhamudul Hasan , Mohd. Farhan Israk Soumik , Abdur R. Shahid

Clickbait headlines are frequently used to attract readers to read articles. Although this headline type has turned out to be a technique to engage readers with misleading items, it is still unknown whether the technique can be used to…

计算与语言 · 计算机科学 2019-11-27 Sima Bhowmik , Md Main Uddin Rony , Md Mahfuzul Haque , Kristen Alley Swain , Naeemul Hassan

We construct the first ever multimodal sarcasm dataset for Spanish. The audiovisual dataset consists of sarcasm annotated text that is aligned with video and audio. The dataset represents two varieties of Spanish, a Latin American variety…

计算与语言 · 计算机科学 2021-05-13 Khalid Alnajjar , Mika Hämäläinen

The widespread use of clickbait headlines, crafted to mislead and maximize engagement, poses a significant challenge to online credibility. These headlines employ sensationalism, misleading claims, and vague language, underscoring the need…

计算与语言 · 计算机科学 2026-04-09 Chhavi Dhiman , Naman Chawla , Riya Dhami , Gaurav Kumar , Ganesh Naik

Text summarization plays a crucial role in natural language processing by condensing large volumes of text into concise and coherent summaries. As digital content continues to grow rapidly and the demand for effective information retrieval…

Data-driven approaches to sequence-to-sequence modelling have been successfully applied to short text summarization of news articles. Such models are typically trained on input-summary pairs consisting of only a single or a few sentences,…

计算与语言 · 计算机科学 2018-04-25 Nikola I. Nikolov , Michael Pfeiffer , Richard H. R. Hahnloser

In this paper, we propose an approach for the detection of clickbait posts in online social media (OSM). Clickbait posts are short catchy phrases that attract a user's attention to click to an article. The approach is based on a machine…

社会与信息网络 · 计算机科学 2017-10-19 Aviad Elyashar , Jorge Bendahan , Rami Puzis

Automated news credibility and fact-checking at scale require accurately predicting news factuality and media bias. This paper introduces a large sentence-level dataset, titled "FactNews", composed of 6,191 sentences expertly annotated…

计算与语言 · 计算机科学 2024-09-16 Francielle Vargas , Kokil Jaidka , Thiago A. S. Pardo , Fabrício Benevenuto

Past studies in Sarcasm Detection mostly make use of Twitter datasets collected using hashtag-based supervision but such datasets are noisy in terms of labels and language. Furthermore, many tweets are replies to other tweets, and detecting…

计算与语言 · 计算机科学 2022-12-13 Rishabh Misra

Often clickbait articles have a title that is phrased as a question or vague teaser that entices the user to click on the link and read the article to find the explanation. We developed a system that will automatically find the answer or…

计算与语言 · 计算机科学 2022-12-19 Oliver Johnson , Beicheng Lou , Janet Zhong , Andrey Kurenkov

Clickbait, which aims to induce users with some surprising and even thrilling headlines for increasing click-through rates, permeates almost all online content publishers, such as news portals and social media. Recently, Large Language…

计算与语言 · 计算机科学 2025-05-13 Han Wang , Yi Zhu , Ye Wang , Yun Li , Yunhao Yuan , Jipeng Qiang

Automatic text summarization aims to produce a brief but crucial summary for the input documents. Both extractive and abstractive methods have witnessed great success in English datasets in recent years. However, there has been a minimal…

计算与语言 · 计算机科学 2021-10-22 Danqing Wang , Jiaze Chen , Xianze Wu , Hao Zhou , Lei Li