中文
相关论文

相关论文: Analyzing Folktales of Different Regions Using Top…

200 篇论文

Topic modeling is a powerful technique to discover hidden topics and patterns within a collection of documents without prior knowledge. Traditional topic modeling and clustering-based techniques encounter challenges in capturing contextual…

计算与语言 · 计算机科学 2024-10-04 Melkamu Abay Mersha , Mesay Gemeda yigezu , Jugal Kalita

Natural language processing techniques are being applied to increasingly diverse types of electronic health records, and can benefit from in-depth understanding of the distinguishing characteristics of medical document types. We present a…

计算与语言 · 计算机科学 2019-10-02 Denis Newman-Griffis , Eric Fosler-Lussier

This work investigates style and topic aspects of language in online communities: looking at both utility as an identifier of the community and correlation with community reception of content. Style is characterized using a hybrid word and…

计算与语言 · 计算机科学 2016-09-16 Trang Tran , Mari Ostendorf

One way of getting a better view of data is using frequent patterns. In this paper frequent patterns are subsets that occur a minimal number of times in a stream of itemsets. However, the discovery of frequent patterns in streams has always…

人工智能 · 计算机科学 2007-05-23 Edgar H. de Graaf , Joost N. Kok , Walter A. Kosters

When exposed to human-generated data, language models are known to learn and amplify societal biases. While previous works introduced benchmarks that can be used to assess the bias in these models, they rely on assumptions that may not be…

计算与语言 · 计算机科学 2025-10-16 Angana Borah , Aparna Garimella , Rada Mihalcea

Note: This paper includes examples of potentially offensive content related to religious bias, presented solely for academic purposes. The widespread adoption of language models highlights the need for critical examinations of their…

计算与语言 · 计算机科学 2025-11-06 Ajwad Abrar , Nafisa Tabassum Oeshy , Mohsinul Kabir , Sophia Ananiadou

We study the problem of frequent itemset mining in domains where data is not recorded in a conventional database but only exists in human knowledge. We provide examples of such scenarios, and present a crowdsourcing model for them. The…

数据库 · 计算机科学 2016-07-19 Antoine Amarilli , Yael Amsterdamer , Tova Milo

Social bookmarking systems allow users to organise collections of resources on the Web in a collaborative fashion. The increasing popularity of these systems as well as first insights into their emergent semantics have made them relevant to…

数字图书馆 · 计算机科学 2008-05-15 Ciro Cattuto , Dominik Benz , Andreas Hotho , Gerd Stumme

Fairness has become a trending topic in natural language processing (NLP), which addresses biases targeting certain social groups such as genders and religions. However, regional bias in language models (LMs), a long-standing global…

计算与语言 · 计算机科学 2022-11-08 Yizhi Li , Ge Zhang , Bohao Yang , Chenghua Lin , Shi Wang , Anton Ragni , Jie Fu

The rapid expansion of biomedical publications creates challenges for organizing knowledge and detecting emerging trends, underscoring the need for scalable and interpretable methods. Common clustering and topic modeling approaches such as…

机器学习 · 计算机科学 2026-02-25 Lana E. Yeganova , Won G. Kim , Shubo Tian , Natalie Xie , Donald C. Comeau , W. John Wilbur , Zhiyong Lu

Folksonomy is an emerging technology that works to classify the information over WWW through tagging the bookmarks, photos or other web-based contents. It is understood to be organized by every user while not limited to the authors of the…

信息检索 · 计算机科学 2007-05-23 Kaikai Shen , Lide Wu

Poetic traditions across languages evolved differently, but we find that certain semantic topics occur in several of them, albeit sometimes with temporal delay, or with diverging trajectories over time. We apply Latent Dirichlet Allocation…

计算与语言 · 计算机科学 2020-08-31 Petr Plechac , Thomas N. Haider

Large language models (LLMs) have become increasingly pivotal in various domains due the recent advancements in their performance capabilities. However, concerns persist regarding biases in LLMs, including gender, racial, and cultural…

人工智能 · 计算机科学 2024-12-03 Mijntje Meijer , Hadi Mohammadi , Ayoub Bagheri

Topic models, such as latent Dirichlet allocation (LDA), can be useful tools for the statistical analysis of document collections and other discrete data. The LDA model assumes that the words of each document arise from a mixture of topics,…

应用统计 · 统计学 2009-09-29 David M. Blei , John D. Lafferty

One of the key issues in both natural language understanding and generation is the appropriate processing of Multiword Expressions (MWEs). MWEs pose a huge problem to the precise language processing due to their idiosyncratic nature and…

计算与语言 · 计算机科学 2014-01-24 Tanmoy Chakraborty , Dipankar Das , Sivaji Bandyopadhyay

Topic models can be useful tools to discover latent topics in collections of documents. Recent studies have shown the feasibility of approach topic modeling as a clustering task. We present BERTopic, a topic model that extends this process…

计算与语言 · 计算机科学 2022-03-14 Maarten Grootendorst

Topic models are probabilistic models for discovering topical themes in collections of documents. In real world applications, these models provide us with the means of organizing what would otherwise be unstructured collections. They can…

信息检索 · 计算机科学 2015-03-06 Wesam Elshamy

Online forums are rich sources of information about user communication activity over time. Finding temporal patterns in online forum communication threads can advance our understanding of the dynamics of conversations. The main challenge of…

社会与信息网络 · 计算机科学 2012-01-12 Andrey Kan , Jeffrey Chan , Conor Hayes , Bernie Hogan , James Bailey , Christopher Leckie

We use an information-theoretic measure of linguistic similarity to investigate the organization and evolution of scientific fields. An analysis of almost 20M papers from the past three decades reveals that the linguistic similarity is…

数字图书馆 · 计算机科学 2018-01-30 Laercio Dias , Martin Gerlach , Joachim Scharloth , Eduardo G. Altmann

Following Henry Small in his approach to co-citation analysis, highly cited sources are seen as concept symbols of research fronts. But instead of co-cited sources I cluster citation links, which are the thematically least heterogenous…

数字图书馆 · 计算机科学 2022-01-26 Frank Havemann
‹ 上一页 1 8 9 10 下一页 ›