中文
相关论文

相关论文: Use of diverse data sources to control which topic…

200 篇论文

Public knowledge graphs such as DBpedia and Wikidata have been recognized as interesting sources of background knowledge to build content-based recommender systems. They can be used to add information about the items to be recommended and…

信息检索 · 计算机科学 2021-05-04 Michael Matthias Voit , Heiko Paulheim

The determination of cluster centers generally depends on the scale that we use to analyze the data to be clustered. Inappropriate scale usually leads to unreasonable cluster centers and thus unreasonable results. In this study, we first…

机器学习 · 统计学 2016-10-20 Xiurui Geng , Hairong Tang

The growing need to manage and exploit the proliferation of online data sources is opening up new opportunities for bringing people closer to the resources they need. For instance, consider a recommendation service through which researchers…

信息检索 · 计算机科学 2011-06-02 C. Basu , W. W. Cohen , H. Hirsh , C. Nevill-Manning

The rapid expansion of biomedical publications creates challenges for organizing knowledge and detecting emerging trends, underscoring the need for scalable and interpretable methods. Common clustering and topic modeling approaches such as…

机器学习 · 计算机科学 2026-02-25 Lana E. Yeganova , Won G. Kim , Shubo Tian , Natalie Xie , Donald C. Comeau , W. John Wilbur , Zhiyong Lu

Mixed data comprises both numeric and categorical features, and mixed datasets occur frequently in many domains, such as health, finance, and marketing. Clustering is often applied to mixed datasets to find structures and to group similar…

机器学习 · 计算机科学 2019-03-20 Amir Ahmad , Shehroz S. Khan

This paper argues that maps of the Web's structure based solely on technical infrastructure such as hyperlinks may bear little resemblance to maps based on Web usage, as cultural factors drive the latter to a larger extent. To test this…

计算机与社会 · 计算机科学 2016-03-11 Harsh Taneja

Recent advances in data science, machine learning, and artificial intelligence, such as the emergence of large language models, are leading to an increasing demand for data that can be processed by such models. While data sources are…

机器学习 · 计算机科学 2023-09-13 Paul Bilokon , Oleksandr Bilokon , Saeed Amen

We analyze the publication records of individual scientists, aiming to quantify the topic switching dynamics of scientists and its influence. For each scientist, the relations among her publications are characterized via shared references.…

物理与社会 · 物理学 2019-09-11 An Zeng , Zhesi Shen , Jianlin Zhou , Ying Fan , Zengru Di , Yougui Wang , H. Eugene Stanley , Shlomo Havlin

There is a growing need for unbiased clustering methods, ideally automated. We have developed a topology-based analysis tool called Two-Tier Mapper (TTMap) to detect subgroups in global gene expression datasets and identify their…

基因组学 · 定量生物学 2018-01-08 Rachel Jeitziner , Mathieu Carrière , Jacques Rougemont , Steve Oudot , Kathryn Hess , Cathrin Brisken

Finding potential research collaborators is a challenging task, especially in today's fast-growing and interdisciplinary research landscape. While traditional methods often rely on observable relationships such as co-authorships and…

信息检索 · 计算机科学 2025-07-29 Md Asaduzzaman Noor , John Sheppard , Jason Clark

Traditionally a document is visualized by a word cloud. Recently, distributed representation methods for documents have been developed, which map a document to a set of topic embeddings. Visualizing such a representation is useful to…

信息检索 · 计算机科学 2017-02-07 Shaohua Li , Tat-Seng Chua

Are users who comment on a variety of matters more likely to achieve high influence than those who delve into one focused field? Do general Twitter hashtags, such as #lol, tend to be more popular than novel ones, such as #instantlyinlove?…

社会与信息网络 · 计算机科学 2015-06-18 Lilian Weng , Filippo Menczer

We provide an overview of tools enabling users to utilize data from open sources for decision-making support in weakly-structured subject domains. Presently, it is impossible to replace expert data with data from open sources in the process…

数据库 · 计算机科学 2019-11-14 Vitaliy Tsyganok , Sergii Kadenko , Oleh Andriichuk

Using data from a large laboratory experiment on problem solving in which we varied the structure of 16-person networks we investigate how an organization's network structure may be constructed to optimize performance in complex…

社会与信息网络 · 计算机科学 2014-07-01 Jesse Shore , Ethan Bernstein , David Lazer

We propose a new clustering approach, called optimality-based clustering, that clusters data points based on their latent decision-making preferences. We assume that each data point is a decision generated by a decision-maker who…

最优化与控制 · 数学 2022-02-15 Zahed Shahmoradi , Taewoo Lee

From social networks to P2P systems, network sampling arises in many settings. We present a detailed study on the nature of biases in network sampling strategies to shed light on how best to sample from networks. We investigate connections…

社会与信息网络 · 计算机科学 2011-09-20 Arun S. Maiya , Tanya Y. Berger-Wolf

This paper reports on the challenges and lessons we learned while running controlled experiments in crowdsourcing platforms. Crowdsourcing is becoming an attractive technique to engage a diverse and large pool of subjects in experimental…

人机交互 · 计算机科学 2020-11-06 Jorge Ramírez , Marcos Baez , Fabio Casati , Luca Cernuzzi , Boualem Benatallah

As large language models (LLMs) are more frequently used in retrieval-augmented generation pipelines, it is increasingly relevant to study their behavior under knowledge conflicts. Thus far, the role of the source of the retrieved…

计算与语言 · 计算机科学 2026-04-20 Jakob Schuster , Vagrant Gautam , Katja Markert

Scientific research trends and interests evolve over time. The ability to identify and forecast these trends is vital for educational institutions, practitioners, investors, and funding organizations. In this study, we predict future trends…

数字图书馆 · 计算机科学 2023-09-22 Dan Ofer , Michal Linial

Scientific document embeddings contain a variety of rich features which can be harnessed for downstream tasks such as recommendation, ranking, and clustering. We explore which tangible insights can be drawn from scientific document…

数字图书馆 · 计算机科学 2025-06-11 Brian D. Zimmerman , Joshua Folkins , Olga Vechtomova