中文
相关论文

相关论文: Dealing with Sparse Document and Topic Representat…

200 篇论文

Latent topic models have been successfully applied as an unsupervised topic discovery technique in large document collections. With the proliferation of hypertext document collection such as the Internet, there has also been great interest…

信息检索 · 计算机科学 2012-06-18 Amit Gruber , Michal Rosen-Zvi , Yair Weiss

This paper presents the results of an experiment to decide the question of authenticity of the supposedly spurious Rhesus - a attic tragedy sometimes credited to Euripides. The experiment involves use of statistics in order to test whether…

cmp-lg · 计算机科学 2008-02-03 Bernd Ludwig

We study retrieving a set of documents that covers various perspectives on a complex and contentious question (e.g., will ChatGPT do more harm than good?). We curate a Benchmark for Retrieval Diversity for Subjective questions (BERDS),…

计算与语言 · 计算机科学 2025-04-23 Hung-Ting Chen , Eunsol Choi

Authors' keyphrases assigned to scientific articles are essential for recognizing content and topic aspects. Most of the proposed supervised and unsupervised methods for keyphrase generation are unable to produce terms that are valuable but…

计算与语言 · 计算机科学 2019-08-22 Erion Çano , Ondřej Bojar

Document collections of various domains, e.g., legal, medical, or financial, often share some underlying collection-wide structure, which captures information that can aid both human users and structure-aware models. We propose to identify…

计算与语言 · 计算机科学 2025-08-27 Gili Lior , Yoav Goldberg , Gabriel Stanovsky

In this paper we present a model for unsupervised topic discovery in texts corpora. The proposed model uses documents, words, and topics lookup table embedding as neural network model parameters to build probabilities of words given topics,…

计算与语言 · 计算机科学 2019-11-26 Sileye 0. Ba

We describe a large-scale application of methods for finding plagiarism in research document collections. The methods are applied to a collection of 284,834 documents collected by arXiv.org over a 14 year period, covering a few different…

数据库 · 计算机科学 2007-05-23 Daria Sorokina , Johannes Gehrke , Simeon Warner , Paul Ginsparg

Topic models analyze text from a set of documents. Documents are modeled as a mixture of topics, with topics defined as probability distributions on words. Inferences of interest include the most probable topics and characterization of a…

信息检索 · 计算机科学 2021-04-19 Jason Wang , Robert E. Weiss

This paper presents a systematic literature review of image datasets for document image analysis, focusing on historical documents, such as handwritten manuscripts and early prints. Finding appropriate datasets for historical document…

计算机视觉与模式识别 · 计算机科学 2022-11-01 Konstantina Nikolaidou , Mathias Seuret , Hamam Mokayed , Marcus Liwicki

We introduce NaSGEC, a new dataset to facilitate research on Chinese grammatical error correction (CGEC) for native speaker texts from multiple domains. Previous CGEC research primarily focuses on correcting texts from a single domain,…

计算与语言 · 计算机科学 2023-05-26 Yue Zhang , Bo Zhang , Haochen Jiang , Zhenghua Li , Chen Li , Fei Huang , Min Zhang

This study explores the extent to which bibliometric indicators based on counts of highly-cited documents could be affected by the choice of data source. The initial hypothesis is that databases that rely on journal selection criteria for…

数字图书馆 · 计算机科学 2018-11-20 Alberto Martín-Martín , Enrique Orduna-Malea , Emilio Delgado López-Cózar

We introduce an approach to topic modelling with document-level covariates that remains tractable in the face of large text corpora. This is achieved by de-emphasizing the role of parameter estimation in an underlying probabilistic model,…

统计方法学 · 统计学 2025-11-05 Gabriel Phelan , David A. Campbell

This paper presents a new task of predicting the coverage of a text document for relation extraction (RE): does the document contain many relational tuples for a given entity? Coverage predictions are useful in selecting the best documents…

计算与语言 · 计算机科学 2021-11-29 Sneha Singhania , Simon Razniewski , Gerhard Weikum

Writing a survey paper on one research topic usually needs to cover the salient content from numerous related papers, which can be modeled as a multi-document summarization (MDS) task. Existing MDS datasets usually focus on producing the…

计算与语言 · 计算机科学 2023-02-10 Shuaiqi Liu , Jiannong Cao , Ruosong Yang , Zhiyuan Wen

Background: The clinical documentation of cystoscopy includes visual and textual materials. However, the secondary use of visual cystoscopic data for educational and research purposes remains limited due to inefficient data management in…

Efficiently identifying keyphrases that represent a given document is a challenging task. In the last years, plethora of keyword detection approaches were proposed. These approaches can be based on statistical (frequency-based) properties…

信息检索 · 计算机科学 2023-12-25 Blaž Škrlj , Boshko Koloski , Senja Pollak

The rapid expansion of medical informatics literature presents significant challenges in synthesizing and analyzing research trends. This study introduces a novel dataset derived from the Medical Informatics Europe (MIE) Conference…

信息检索 · 计算机科学 2024-10-08 Ehsan Bitaraf , Maryam Jafarpour

We address two challenges in topic models: (1) Context information around words helps in determining their actual meaning, e.g., "networks" used in the contexts "artificial neural networks" vs. "biological neuron networks". Generative topic…

计算与语言 · 计算机科学 2019-01-16 Pankaj Gupta , Yatin Chaudhary , Florian Buettner , Hinrich Schütze

For extracting meaningful topics from texts, their structures should be considered properly. In this paper, we aim to analyze structured time-series documents such as a collection of news articles and a series of scientific papers, wherein…

计算与语言 · 计算机科学 2018-05-08 Rem Hida , Naoya Takeishi , Takehisa Yairi , Koichi Hori

When people explore and manage information, they think in terms of topics and themes. However, the software that supports information exploration sees text at only the surface level. In this paper we show how topic modeling -- a technique…

人机交互 · 计算机科学 2011-11-07 Jacob Eisenstein , Duen Horng "Polo" Chau , Aniket Kittur , Eric P. Xing