中文
相关论文

相关论文: CommunityFish: A Poisson-based Document Scaling Wi…

200 篇论文

Text clustering is a fundamental task in natural language processing, yet traditional clustering algorithms with pre-trained embeddings often struggle in domain-specific contexts without costly fine-tuning. Large language models (LLMs)…

计算与语言 · 计算机科学 2025-12-05 Yiming Xu , Yuan Yuan , Vijay Viswanathan , Graham Neubig

Speaker diarization based on bottom-up clustering of speech segments by acoustic similarity is often highly sensitive to the choice of hyperparameters, such as the initial number of clusters and feature weighting. Optimizing these…

计算与语言 · 计算机科学 2022-02-22 Andreas Stolcke

We replicate recent experiments attempting to demonstrate an attractive hypothesis about the use of the Fisher kernel framework and mixture models for aggregating word embeddings towards document representations and the use of these…

计算与语言 · 计算机科学 2020-01-15 Luca Papariello , Alexandros Bampoulidis , Mihai Lupu

The multi-document summarization task requires the designed summarizer to generate a short text that covers the important information of original documents and satisfies content diversity. This paper proposes a multi-document summarization…

计算与语言 · 计算机科学 2023-03-07 Bing Ma

Topic models are widely used for discovering latent thematic structures in large text corpora, yet traditional unsupervised methods often struggle to align with pre-defined conceptual domains. This paper introduces seeded Poisson…

统计方法学 · 统计学 2025-10-07 Bernd Prostmaier , Jan Vávra , Bettina Grün , Paul Hofmarcher

Temporal editing patterns on Wikipedia provide a unique computational lens to explore cultural dynamics across linguistic communities. This study analyses over a decade of editorial activity (2001-2010) across eleven Wikipedia language…

物理与社会 · 物理学 2026-03-03 David André Villamil Carrillo , Yérali Gandica

Understanding a user's motivations provides valuable information beyond the ability to recommend items. Quite often this can be accomplished by perusing both ratings and review texts, since it is the latter where the reasoning for specific…

机器学习 · 计算机科学 2015-12-08 Chao-Yuan Wu , Alex Beutel , Amr Ahmed , Alexander J. Smola

The paper presents a linguistic and computational model aiming at making the morphological structure of the lexicon emerge from the formal and semantic regularities of the words it contains. The model is word-based. The proposed…

计算与语言 · 计算机科学 2009-05-12 Nabil Hathout

Contextual large language model embeddings are increasingly utilized for topic modeling and clustering. However, current methods often scale poorly, rely on opaque similarity metrics, and struggle in multilingual settings. In this work, we…

计算与语言 · 计算机科学 2025-06-03 Hans W. A. Hanley , Zakir Durumeric

In this paper we propose a framework inspired by interacting particle physics and devised to perform clustering on multidimensional datasets. To this end, any given dataset is modeled as an interacting particle system, under the assumption…

统计力学 · 物理学 2012-07-26 Giuliano Armano , Marco Alberto Javarone

Subspace clustering is a problem of exploring the low-dimensional subspaces of high-dimensional data. State-of-the-arts approaches are designed by following the model of spectral clustering based method. These methods pay much attention to…

计算机视觉与模式识别 · 计算机科学 2019-04-30 Xuelong Li , Quanmao Lu , Yongsheng Dong , Dacheng Tao

We compare the performance of different clustering algorithms applied to the task of unsupervised text categorization. We consider agglomerative clustering algorithms, principal direction divisive partitioning and (for the first time)…

无序系统与神经网络 · 物理学 2007-05-23 D. Volk , M. G. Stepanov

The task of clustering a set of objects based on multiple sources of data arises in several modern applications. We propose an integrative statistical model that permits a separate clustering of the objects for each data source. These…

机器学习 · 统计学 2015-12-01 Eric F. Lock , David B. Dunson

Cross-lingual annotations of legislative texts enable us to explore major themes covered in multilingual legal data and are a key facilitator of semantic similarity when searching for similar documents. Multilingual probabilistic topic…

信息检索 · 计算机科学 2019-12-02 Carlos Badenes-Olmedo , Jose-Luis Redondo-Garcia , Oscar Corcho

The proliferation of tobacco-related content on social media platforms poses significant challenges for public health monitoring and intervention. This paper introduces a novel multi-modal deep learning framework named Flow-Attention…

计算机视觉与模式识别 · 计算机科学 2024-12-10 Naga VS Raviteja Chappa , Page Daniel Dobbs , Bhiksha Raj , Khoa Luu

This study introduces a novel methodology for mapping scientific communities at scale, addressing challenges associated with network analysis in large bibliometric datasets. By leveraging enriched publication metadata from the French…

数字图书馆 · 计算机科学 2025-01-20 Victor Barbier , Eric Jeangirard

Text Clustering is a text mining technique which divides the given set of text documents into significant clusters. It is used for organizing a huge number of text documents into a well-organized form. In the majority of the clustering…

信息检索 · 计算机科学 2015-03-12 G. Hannah Grace , Kalyani Desikan

Earlier techniques of text mining included algorithms like k-means, Naive Bayes, SVM which classify and cluster the text document for mining relevant information about the documents. The need for improving the mining techniques has us…

信息检索 · 计算机科学 2016-05-10 Jinju Joby , Jyothi Korra

The task of organizing and clustering multilingual news articles for media monitoring is essential to follow news stories in real time. Most approaches to this task focus on high-resource languages (mostly English), with low-resource…

计算与语言 · 计算机科学 2022-04-29 João Santos , Afonso Mendes , Sebastião Miranda

Natural data is often organized as a hierarchical composition of features. How many samples do generative models need in order to learn the composition rules, so as to produce a combinatorially large number of novel data? What signal in the…