中文
相关论文

相关论文: Topic Diffusion Discovery based on Sparseness-cons…

200 篇论文

As social networks are constantly changing and evolving, methods to analyze dynamic social networks are becoming more important in understanding social trends. However, due to the restrictions imposed by the social network service…

社会与信息网络 · 计算机科学 2018-01-09 Kaan Bingöl , Bahaeddin Eravcı , Çağrı Özgenç Etemoğlu , Hakan Ferhatosmanoğlu , Buğra Gedik

We propose a Concentrated Document Topic Model(CDTM) for unsupervised text classification, which is able to produce a concentrated and sparse document topic distribution. In particular, an exponential entropy penalty is imposed on the…

机器学习 · 统计学 2021-02-10 Hao Lei , Ying Chen

Mapping data from and/or onto a known family of distributions has become an important topic in machine learning and data analysis. Deep generative models (e.g., generative adversarial networks ) have been used effectively to match known and…

机器学习 · 计算机科学 2020-10-30 Surojit Saha , Shireen Elhabian , Ross T. Whitaker

Inferring topics from the overwhelming amount of short texts becomes a critical but challenging task for many content analysis tasks, such as content charactering, user interest profiling, and emerging topic detecting. Existing methods such…

计算与语言 · 计算机科学 2016-09-28 Jipeng Qiang , Ping Chen , Tong Wang , Xindong Wu

Recent advances in engineering technologies have enabled the collection of a large number of longitudinal features. This wealth of information presents unique opportunities for researchers to investigate the complex nature of diseases and…

统计方法学 · 统计学 2023-11-27 Zihang Lu , Noirrit Kiran Chandra

The short text has been the prevalent format for information of Internet in recent decades, especially with the development of online social media, whose millions of users generate a vast number of short messages everyday. Although…

计算与语言 · 计算机科学 2014-12-18 Yuan Zuo , Jichang Zhao , Ke Xu

Named entity classification is the task of classifying text-based elements into various categories, including places, names, dates, times, and monetary values. A bottleneck in named entity classification, however, is the data problem of…

计算与语言 · 计算机科学 2017-03-16 Ai Hirata , Mamoru Komachi

Detecting and tracking emerging trends and weak signals in large, evolving text corpora is vital for applications such as monitoring scientific literature, managing brand reputation, surveilling critical infrastructure and more generally to…

计算与语言 · 计算机科学 2024-11-22 Allaa Boutaleb , Jerome Picault , Guillaume Grosjean

The rapid and unregulated dissemination of information in the digital era has amplified the global "infodemic," complicating the identification of high quality information. We present a lightweight, interpretable and non-invasive framework…

社会与信息网络 · 计算机科学 2025-09-09 Anthony Lopes Temporao , Mickael Temporão , Corentin Vande Kerckhove , Flavio Abreu Araujo

Detection of emerging topics are now receiving renewed interest motivated by the rapid growth of social networks. Conventional term-frequency-based approaches may not be appropriate in this context, because the information exchanged are not…

机器学习 · 统计学 2011-10-14 Toshimitsu Takahashi , Ryota Tomioka , Kenji Yamanishi

We develop a unified and systematic framework for performing online nonnegative matrix factorization under a wide variety of important divergences. The online nature of our algorithm makes it particularly amenable to large-scale data. We…

机器学习 · 统计学 2016-08-17 Renbo Zhao , Vincent Y. F. Tan , Huan Xu

In this paper we present a model for unsupervised topic discovery in texts corpora. The proposed model uses documents, words, and topics lookup table embedding as neural network model parameters to build probabilities of words given topics,…

计算与语言 · 计算机科学 2019-11-26 Sileye 0. Ba

Probabilistic topic models are widely used to discover latent topics in document collections, while latent feature vector representations of words have been used to obtain high performance in many NLP tasks. In this paper, we extend two…

计算与语言 · 计算机科学 2018-10-16 Dat Quoc Nguyen , Richard Billingsley , Lan Du , Mark Johnson

Diffusion based approaches to long form text generation suffer from prohibitive computational cost and memory overhead as sequence length increases. We introduce SA-DiffuSeq, a diffusion framework that integrates sparse attention to…

计算与语言 · 计算机科学 2025-12-25 Alexandros Christoforos , Chadbourne Davis

Large Language Models have undoubtedly revolutionized the Natural Language Processing field, the current trend being to promote one-model-for-all tasks (sentiment analysis, translation, etc.). However, the statistical mechanisms at work in…

计算与语言 · 计算机科学 2024-08-26 Célia D'Cruz , Jean-Marc Bereder , Frédéric Precioso , Michel Riveill

Uncovering hidden topics from short texts is challenging for traditional and neural models due to data sparsity, which limits word co-occurrence patterns, and label sparsity, stemming from incomplete reconstruction targets. Although data…

计算与语言 · 计算机科学 2025-01-24 Quang Duc Nguyen , Tung Nguyen , Duc Anh Nguyen , Linh Ngo Van , Sang Dinh , Thien Huu Nguyen

We present a hybrid method for latent information discovery on the data sets containing both text content and connection structure based on constrained low rank approximation. The new method jointly optimizes the Nonnegative Matrix…

机器学习 · 计算机科学 2017-03-29 Rundong Du , Barry Drake , Haesun Park

Studying information diffusion in SNS (Social Networks Service) has remarkable significance in both academia and industry. Theoretically, it boosts the development of other subjects such as statistics, sociology, and data mining.…

社会与信息网络 · 计算机科学 2021-10-28 Huacheng Li , Chunhe Xia , Tianbo Wang , Sheng Wen , Chao Chen , Yang Xiang

Text Categorization is traditionally done by using the term frequency and inverse document frequency.This type of method is not very good because, some words which are not so important may appear in the document .The term frequency of…

信息检索 · 计算机科学 2016-11-25 Srikanth Bethu , G Charless Babu , J Vinoda , E Priyadarshini , M Raghavendra rao

Semi-Non-negative Matrix Factorization is a technique that learns a low-dimensional representation of a dataset that lends itself to a clustering interpretation. It is possible that the mapping between this new representation and our…

计算机视觉与模式识别 · 计算机科学 2015-09-11 George Trigeorgis , Konstantinos Bousmalis , Stefanos Zafeiriou , Bjoern W. Schuller