中文
相关论文

相关论文: Use of diverse data sources to control which topic…

200 篇论文

Document clustering and topic modeling are two closely related tasks which can mutually benefit each other. Topic modeling can project documents into a topic space which facilitates effective document clustering. Cluster labels discovered…

机器学习 · 计算机科学 2013-09-27 Pengtao Xie , Eric P. Xing

Science is a social process with far-reaching impact on our modern society. In the recent years, for the first time we are able to scientifically study the science itself. This is enabled by massive amounts of data on scientific…

数字图书馆 · 计算机科学 2015-05-21 Lovro Šubelj , Marko Bajec , Biljana Mileva Boshkoska , Andrej Kastrin , Zoran Levnajić

Document clustering is an unsupervised approach in which a large collection of documents (corpus) is subdivided into smaller, meaningful, identifiable, and verifiable sub-groups (clusters). Meaningful representation of documents and…

信息检索 · 计算机科学 2014-12-08 Muhammad Rafi , Farnaz Amin , Mohammad Shahid Shaikh

Citation metrics are becoming pervasive in the quantitative evaluation of scholars, journals and institutions. More then ever before, hiring, promotion, and funding decisions rely on a variety of impact metrics that cannot disentangle…

数字图书馆 · 计算机科学 2015-09-03 Jasleen Kaur , Emilio Ferrara , Filippo Menczer , Alessandro Flammini , Filippo Radicchi

Tables are common and important in scientific documents, yet most text-based document search systems do not capture structures and semantics specific to tables. How to bridge different types of mismatch between keywords queries and…

信息检索 · 计算机科学 2017-07-13 Kyle Yingkai Gao , Jamie Callan

This paper presents results of topic modeling and network models of topics using the International Conference on Computational Science corpus, which contains domain-specific (computational science) papers over sixteen years (a total of 5695…

News sources undergo the process of selecting newsworthy information when covering a certain topic. The process inevitably exhibits selection biases, i.e. news sources' typical patterns of choosing what information to include in news…

计算与语言 · 计算机科学 2023-04-10 Sihao Chen , William Bruno , Dan Roth

Earlier techniques of text mining included algorithms like k-means, Naive Bayes, SVM which classify and cluster the text document for mining relevant information about the documents. The need for improving the mining techniques has us…

信息检索 · 计算机科学 2016-05-10 Jinju Joby , Jyothi Korra

Citation analysis of the scientific literature has been used to study and define disciplinary boundaries, to trace the dissemination of knowledge, and to estimate impact. Co-citation, the frequency with which pairs of publications are…

The increasing adoption of econometric and machine-learning approaches by empirical researchers has led to a widespread use of one data collection method: web scraping. Web scraping refers to the use of automated computer programs to access…

综合经济学 · 经济学 2023-08-07 Jens Foerderer

We propose a new technique to infer the structure and extract the tokens of data from the semi-structured web sources which are generated using a consistent template or layout with some implicit regularities. The attributes are extracted…

信息检索 · 计算机科学 2009-08-06 Z. Akbar , L. T. Handoko

The citation impact of Environment and Planning B can be visualized using its citation relations with journals in its environment as the links of a network. The size of the nodes is varied in correspondence to the relative citation impact…

数字图书馆 · 计算机科学 2009-11-17 Loet Leydesdorff

Twitter accounts have already been used in many scientometric studies, but the meaningfulness of the data for societal impact measurements in research evaluation has been questioned. Earlier research focused on social media counts and…

数字图书馆 · 计算机科学 2019-04-16 Robin Haunschild , Loet Leydesdorff , Lutz Bornmann , Iina Hellsten , Werner Marx

We tackle the challenge of topic classification of tweets in the context of analyzing a large collection of curated streams by news outlets and other organizations to deliver relevant content to users. Our approach is novel in applying…

信息检索 · 计算机科学 2017-04-25 Salman Mohammed , Nimesh Ghelani , Jimmy Lin

Tables in scientific papers contain a wealth of valuable knowledge for the scientific enterprise. To help the many of us who frequently consult this type of knowledge, we present Tab2Know, a new end-to-end system to build a Knowledge Base…

人工智能 · 计算机科学 2021-07-29 Benno Kruit , Hongyu He , Jacopo Urbani

The amount of scientific papers published every day is daunting and constantly increasing. Keeping up with literature represents a challenge. If one wants to start exploring new topics it is hard to have a big picture without reading lots…

信息检索 · 计算机科学 2020-11-10 Alberto Calderone

Probabilistic topic models are a powerful tool for extracting latent themes from large text datasets. In many text datasets, we also observe per-document covariates (e.g., source, style, political affiliation) that act as environments that…

计算与语言 · 计算机科学 2024-11-04 Dominic Sobhani , Amir Feder , David Blei

For a variety of inter-related cultural, organizational, and political reasons, progress in climate science and the actual solution of scientific problems in this field have moved at a much slower rate than would normally be possible. Not…

物理与社会 · 物理学 2012-10-09 Richard S. Lindzen

We study clustering on graphs with multiple edge types. Our main motivation is that similarities between objects can be measured in many different metrics. For instance similarity between two papers can be based on common authors, where…

社会与信息网络 · 计算机科学 2011-09-09 Matthew Rocklin , Ali Pinar

Although there are millions of transgender people in the world, a lack of information exists about their health issues. This issue has consequences for the medical field, which only has a nascent understanding of how to identify and meet…

计算机与社会 · 计算机科学 2018-10-01 Amir Karami , Frank Webb , Vanessa L. Kitzie