中文
相关论文

相关论文: Use of diverse data sources to control which topic…

200 篇论文

Comparative text mining extends from genre analysis and political bias detection to the revelation of cultural and geographic differences, through to the search for prior art across patents and scientific papers. These applications use…

信息检索 · 计算机科学 2019-11-27 Julian Risch , Ralf Krestel

Mining Time Series data has a tremendous growth of interest in today's world. To provide an indication various implementations are studied and summarized to identify the different problems in existing applications. Clustering time series is…

信息检索 · 计算机科学 2010-05-25 V. Kavitha , M. Punithavalli

Statistical topic models efficiently facilitate the exploration of large-scale data sets. Many models have been developed and broadly used to summarize the semantic structure in news, science, social media, and digital humanities. However,…

机器学习 · 计算机科学 2016-12-02 Jian Tang , Cheng Li , Ming Zhang , Qiaozhu Mei

Using administrative patient-care data such as Electronic Health Records (EHR) and medical/ pharmaceutical claims for population-based scientific research has become increasingly common. With vast sample sizes leading to very small standard…

统计方法学 · 统计学 2023-08-21 Ritoban Kundu , Xu Shi , Jean Morrison , Jessica Barrett , Bhramar Mukherjee

Online and in the real world, communities are bonded together by emotional consensus around core issues. Emotional responses to scientific findings often play a pivotal role in these core issues. When there is too much diversity of opinion…

社会与信息网络 · 计算机科学 2020-01-07 Cole Freeman , Hamed Alhoori , Murtuza Shahzad

The proliferation of high-dimensional data from sources such as social media, sensor networks, and online platforms has created new challenges for clustering algorithms. Multi-view clustering, which integrates complementary information from…

机器学习 · 计算机科学 2026-01-23 Chakib Fettal , Lazhar Labiod , Mohamed Nadif

In an information-rich world, people's time and attention must be divided among rapidly changing information sources and the diverse tasks demanded of them. How people decide which of the many sources, such as scientific articles or…

数字图书馆 · 计算机科学 2017-10-03 Kristina Lerman , Nathan Hodas , Hao Wu

In centralized countries, not only population, media and economic power are concentrated, but people give more attention to central locations. While this is not inherently bad, this behavior extends to micro-blogging platforms: central…

社会与信息网络 · 计算机科学 2016-01-05 Eduardo Graells-Garrido , Mounia Lalmas , Ricardo Baeza-Yates

Data mining is the task of discovering interesting, unexpected or valuable structures in large datasets and transforming them into an understandable structure for further use . Different approaches in the domain of data mining have been…

数据库 · 计算机科学 2020-09-21 Julie Bu Daher , Armelle Brun , Anne Boyer

Social media studies often collect data retrospectively to analyze public opinion. Social media data may decay over time and such decay may prevent the collection of the complete dataset. As a result, the collected dataset may differ from…

社会与信息网络 · 计算机科学 2023-03-03 Tuğrulcan Elmas

Due to the widespread use of data-powered systems in our everyday lives, concepts like bias and fairness gained significant attention among researchers and practitioners, in both industry and academia. Such issues typically emerge from the…

机器学习 · 计算机科学 2023-05-18 Gianluca Demartini , Kevin Roitero , Stefano Mizzaro

Machine learning approaches to multi-label document classification have to date largely relied on discriminative modeling techniques such as support vector machines. A drawback of these approaches is that performance rapidly drops off as…

机器学习 · 统计学 2011-11-11 Timothy N. Rubin , America Chambers , Padhraic Smyth , Mark Steyvers

The rapid growth of scientific literature has made it difficult for the researchers to quickly learn about the developments in their respective fields. Scientific document summarization addresses this challenge by providing summaries of the…

计算与语言 · 计算机科学 2017-06-13 Arman Cohan , Nazli Goharian

There are many scenarios where we may want to find pairs of textually similar documents in a large corpus (e.g. a researcher doing literature review, or an R&D project manager analyzing project proposals). To programmatically discover those…

计算与语言 · 计算机科学 2020-12-16 Carlos Badenes-Olmedo , Jose-Luis Redondo García , Oscar Corcho

Science policy is increasingly shifting towards an emphasis in societal problems or grand challenges. As a result, new evaluative tools are needed to help assess not only the knowledge production side of research programmes or…

数字图书馆 · 计算机科学 2017-10-16 Lorenzo Cassi , Agénor Lahatte , Ismael Rafols , Pierre Sautier , Élisabeth de Turckheim

Document clustering is a text mining technique used to provide better document search and browsing in digital libraries or online corpora. A lot of research has been done on biomedical document clustering that is based on using existing…

计算与语言 · 计算机科学 2018-10-24 Setu Shah , Xiao Luo

Citations are the cornerstone of knowledge propagation and the primary means of assessing the quality of research, as well as directing investments in science. Science is increasingly becoming "data-intensive", where large volumes of data…

数字图书馆 · 计算机科学 2017-09-28 Gianmaria Silvello

The representation of science as a citation-density landscape and the study of scaling rules with the field-specific citation-density as a main topological property was previously analysed at the level of research groups. Here the focus is…

物理与社会 · 物理学 2008-03-10 Rodrigo Costas , Maria Bordons , Thed N. van Leeuwen , Anthony F. J. van Raan

Scientific topics, claims and resources are increasingly debated as part of online discourse, where prominent examples include discourse related to COVID-19 or climate change. This has led to both significant societal impact and increased…

计算与语言 · 计算机科学 2022-07-07 Salim Hafid , Sebastian Schellhammer , Sandra Bringay , Konstantin Todorov , Stefan Dietze

We propose a novel clustering pipeline to detect and characterize influence campaigns from documents. This approach clusters parts of document, detects clusters that likely reflect an influence campaign, and then identifies documents linked…

计算与语言 · 计算机科学 2024-04-30 Zhengxiang Wang , Owen Rambow