中文
相关论文

相关论文: Statistical Analysis on Bangla Newspaper Data to E…

200 篇论文

This paper introduces a large collection of time series data derived from Twitter, postprocessed using word embedding techniques, as well as specialized fine-tuned language models. This data comprises the past five years and captures…

Scientific research trends and interests evolve over time. The ability to identify and forecast these trends is vital for educational institutions, practitioners, investors, and funding organizations. In this study, we predict future trends…

数字图书馆 · 计算机科学 2023-09-22 Dan Ofer , Michal Linial

Social media platforms are thriving nowadays, so a huge volume of data is produced. As it includes brief and clear statements, millions of people post their thoughts on microblogging sites every day. This paper represents and analyze the…

社会与信息网络 · 计算机科学 2021-08-05 Suchandra Dutta , Dhrubasish Sarkar , Sohom Roy , Dipak K. Kole , Premananda Jana

Certain type of documents such as tweets are collected by specifying a set of keywords. As topics of interest change with time it is beneficial to adjust keywords dynamically. The challenge is that these need to be specified ahead of…

机器学习 · 统计学 2020-01-23 Xingyu Wang , Lida Zhang , Diego Klabjan

With the rise of social media and online news sources, fake news has become a significant issue globally. However, the detection of fake news in low resource languages like Bengali has received limited attention in research. In this paper,…

With rapidly evolving media narratives, it has become increasingly critical to not just extract narratives from a given corpus but rather investigate, how they develop over time. While popular narrative extraction methods such as Large…

计算与语言 · 计算机科学 2025-06-26 Kai-Robin Lange , Tobias Schmidt , Matthias Reccius , Henrik Müller , Michael Roos , Carsten Jentsch

Topic modeling refers to the task of discovering the underlying thematic structure in a text corpus, where the output is commonly presented as a report of the top terms appearing in each topic. Despite the diversity of topic modeling…

机器学习 · 计算机科学 2014-06-20 Derek Greene , Derek O'Callaghan , Pádraig Cunningham

Much information available on the web is copied, reused or rephrased. The phenomenon that multiple web sources pick up certain information is often called trend. A central problem in the context of web data mining is to detect those web…

机器学习 · 计算机科学 2012-07-03 Felix Biessmann , Jens-Michalis Papaioannou , Mikio Braun , Andreas Harth

Newspapers are a popular form of written discourse, read by many people, thanks to the novelty of the information provided by the news content in it. A headline is the most widely read part of any newspaper due to its appearance in a bigger…

计算与语言 · 计算机科学 2019-10-21 Elizabeth Jasmi George , Radhika Mamidi

Bangla, the seventh most widely spoken language worldwide with 300 million native speakers, faces digital under-representation due to limited resources and lack of annotated datasets. Stemming, a critical preprocessing step in language…

计算与语言 · 计算机科学 2025-08-22 Abhijit Paul , Mashiat Amin Farin , Sharif Md. Abdullah , Ahmedul Kabir , Zarif Masud , Shebuti Rayana

We present a software tool that employs state-of-the-art natural language processing (NLP) and machine learning techniques to help newspaper editors compose effective headlines for online publication. The system identifies the most salient…

计算与语言 · 计算机科学 2019-05-21 Terrence Szymanski , Claudia Orellana-Rodriguez , Mark T. Keane

Social media platforms contain a great wealth of information which provides opportunities for us to explore hidden patterns or unknown correlations, and understand people's satisfaction with what they are discussing. As one showcase, in…

信息检索 · 计算机科学 2017-05-24 Zhengkui Wang , Guangdong Bai , Soumyadeb Chowdhury , Quanqing Xu , Zhi Lin Seow

Determining the readability of a text is the first step to its simplification. In this paper, we present a readability analysis tool capable of analyzing text written in the Bengali language to provide in-depth information on its…

计算与语言 · 计算机科学 2020-12-15 Susmoy Chakraborty , Mir Tafseer Nayeem , Wasi Uddin Ahmad

Trending topics are the online conversations that grab collective attention on social media. They are continually changing and often reflect exogenous events that happen in the real world. Trends are localized in space and time as they are…

社会与信息网络 · 计算机科学 2017-03-07 Emilio Ferrara , Onur Varol , Filippo Menczer , Alessandro Flammini

Text analysis is an interesting research area in data science and has various applications, such as in artificial intelligence, biomedical research, and engineering. We review popular methods for text analysis, ranging from topic modeling…

应用统计 · 统计学 2024-02-08 Zheng Tracy Ke , Pengsheng Ji , Jiashun Jin , Wanshan Li

The task of predicting the publication period of text documents, such as news articles, is an important but less studied problem in the field of natural language processing. Predicting the year of a news article can be useful in various…

计算与语言 · 计算机科学 2023-04-26 Karthick Prasad Gunasekaran , B Chase Babrich , Saurabh Shirodkar , Hee Hwang

Monitoring online customer reviews is important for business organisations to measure customer satisfaction and better manage their reputations. In this paper, we propose a novel dynamic Brand-Topic Model (dBTM) which is able to…

信息检索 · 计算机科学 2023-01-19 Runcong Zhao , Lin Gui , Hanqi Yan , Yulan He

The time at which a message is communicated is a vital piece of metadata in many real-world natural language processing tasks such as Topic Detection and Tracking (TDT). TDT systems aim to cluster a corpus of news articles by event, and in…

计算与语言 · 计算机科学 2024-03-27 Hang Jiang , Doug Beeferman , Weiquan Mao , Deb Roy

This paper investigates advertising practices in print newspapers across India using a novel data-driven approach. We develop a pipeline employing image processing and OCR techniques to extract articles and advertisements from digital…

计算机与社会 · 计算机科学 2025-05-19 N Harsha Vardhan , Ponnurangam Kumaraguru , Kiran Garimella

Text data is inherently temporal. The meaning of words and phrases changes over time, and the context in which they are used is constantly evolving. This is not just true for social media data, where the language used is rapidly influenced…

计算与语言 · 计算机科学 2025-03-05 Kai-Robin Lange , Niklas Benner , Lars Grönberg , Aymane Hachcham , Imene Kolli , Jonas Rieger , Carsten Jentsch