中文
相关论文

相关论文: Twitter-Network Topic Model: A Full Bayesian Treat…

200 篇论文

Many real systems have been modelled in terms of network concepts, and written texts are a particular example of information networks. In recent years, the use of network methods to analyze language has allowed the discovery of several…

计算与语言 · 计算机科学 2016-06-28 Henrique F. de Arruda , Luciano da F. Costa , Diego R. Amancio

With the rise in popularity of public social media and micro-blogging services, most notably Twitter, the people have found a venue to hear and be heard by their peers without an intermediary. As a consequence, and aided by the public…

计算与语言 · 计算机科学 2016-06-21 Prashanth Vijayaraghavan , Soroush Vosoughi , Deb Roy

Twitter data have become essential to Natural Language Processing (NLP) and social science research, driving various scientific discoveries in recent years. However, the textual data alone are often not enough to conduct studies: especially…

计算与语言 · 计算机科学 2022-01-27 Federico Bianchi , Vincenzo Cutrona , Dirk Hovy

In this paper, we consider the problem of latent sentiment detection in Online Social Networks such as Twitter. We demonstrate the benefits of using the underlying social network as an Ising prior to perform network aided sentiment…

社会与信息网络 · 计算机科学 2014-01-10 Rohit Negi , Vinay Uday Prabhu , Miguel Rodrigues

Our paper studies the predictability of online speech -- that is, how well language models learn to model the distribution of user generated content on X (previously Twitter). We define predictability as a measure of the model's…

计算与语言 · 计算机科学 2026-01-07 Mina Remeli , Moritz Hardt , Robert C. Williamson

This paper describes a pilot NER system for Twitter, comprising the USFD system entry to the W-NUT 2015 NER shared task. The goal is to correctly label entities in a tweet dataset, using an inventory of ten types. We employ structured…

计算与语言 · 计算机科学 2015-11-11 Leon Derczynski , Isabelle Augenstein , Kalina Bontcheva

Topic lifecycle analysis on Twitter, a branch of study that investigates Twitter topics from their birth through lifecycle to death, has gained immense mainstream research popularity. In the literature, topics are often treated as one of…

社会与信息网络 · 计算机科学 2018-01-19 Kuntal Dey , Saroj Kaushik , Kritika Garg , Ritvik Shrivastava

Recently, online social media has become a primary source for new information and misinformation or rumours. In the absence of an automatic rumour detection system the propagation of rumours has increased manifold leading to serious…

计算与语言 · 计算机科学 2022-12-21 Shaswat Patel , Prince Bansal , Preeti Kaur

Statistical topic models provide a general data-driven framework for automated discovery of high-level knowledge from large collections of text documents. While topic models can potentially discover a broad range of themes in a data set,…

人工智能 · 计算机科学 2008-08-08 Chaitanya Chemudugunta , Padhraic Smyth , Mark Steyvers

We present a novel response generation system that can be trained end to end on large quantities of unstructured Twitter conversations. A neural network architecture is used to address sparsity issues that arise when integrating contextual…

Text from social media provides a set of challenges that can cause traditional NLP approaches to fail. Informal language, spelling errors, abbreviations, and special characters are all commonplace in these posts, leading to a prohibitively…

机器学习 · 计算机科学 2016-05-18 Bhuwan Dhingra , Zhong Zhou , Dylan Fitzpatrick , Michael Muehl , William W. Cohen

The Topics over Time (ToT) model captures thematic changes in timestamped datasets by explicitly modeling publication dates jointly with word co-occurrence patterns. However, ToT was not approached in a fully Bayesian fashion, a flaw that…

计算与语言 · 计算机科学 2025-04-22 Julián Cendrero , Julio Gonzalo , Ivar Zapata

Existing deep hierarchical topic models are able to extract semantically meaningful topics from a text corpus in an unsupervised manner and automatically organize them into a topic hierarchy. However, it is unclear how to incorporate prior…

机器学习 · 计算机科学 2021-10-28 Zhibin Duan , Yishi Xu , Bo Chen , Dongsheng Wang , Chaojie Wang , Mingyuan Zhou

One of the main computational and scientific challenges in the modern age is to extract useful information from unstructured texts. Topic models are one popular machine-learning approach which infers the latent topical structure of a…

机器学习 · 统计学 2018-07-20 Martin Gerlach , Tiago P. Peixoto , Eduardo G. Altmann

Twitter is currently a popular online social media platform which allows users to share their user-generated content. This publicly-generated user data is also crucial to healthcare technologies because the discovered patterns would hugely…

机器学习 · 计算机科学 2021-05-25 Hamad Zogan , Imran Razzak , Shoaib Jameel , Guandong Xu

In the Humanities and Social Sciences, there is increasing interest in approaches to information extraction, prediction, intelligent linkage, and dimension reduction applicable to large text corpora. With approaches in these fields being…

应用统计 · 统计学 2019-04-16 Vanessa Glenny , Jonathan Tuke , Nigel Bean , Lewis Mitchell

Topic modelling is a text mining technique for identifying salient themes from a number of documents. The output is commonly a set of topics consisting of isolated tokens that often co-occur in such documents. Manual effort is often…

计算与语言 · 计算机科学 2024-04-26 Lowri Williams , Eirini Anthi , Laura Arman , Pete Burnap

Performance of neural models for named entity recognition degrades over time, becoming stale. This degradation is due to temporal drift, the change in our target variables' statistical properties over time. This issue is especially…

计算与语言 · 计算机科学 2021-04-21 Shuguang Chen , Leonardo Neves , Thamar Solorio

Topic modeling is a well-established technique for exploring text corpora. Conventional topic models (e.g., LDA) represent topics as bags of words that often require "reading the tea leaves" to interpret; additionally, they offer users…

计算与语言 · 计算机科学 2024-04-03 Chau Minh Pham , Alexander Hoyle , Simeng Sun , Philip Resnik , Mohit Iyyer

Tokenization is a foundational step in most natural language processing (NLP) pipelines, yet it introduces challenges such as vocabulary mismatch and out-of-vocabulary issues. Recent work has shown that models operating directly on raw text…

计算与语言 · 计算机科学 2025-05-05 Sumit Mamtani , Maitreya Sonawane , Kanika Agarwal , Nishanth Sanjeev