English
Related papers

Related papers: A Bayesian Learning, Greedy agglomerative clusteri…

200 papers

We study a fundamental problem in Bayesian learning, where the goal is to select a set of data sources with minimum cost while achieving a certain learning performance based on the data streams provided by the selected data sources. First,…

Machine Learning · Computer Science 2021-05-04 Lintao Ye , Aritra Mitra , Shreyas Sundaram

Crowdsourcing has emerged as a powerful paradigm for efficiently labeling large datasets and performing various learning tasks, by leveraging crowds of human annotators. When additional information is available about the data,…

Machine Learning · Computer Science 2021-07-19 Panagiotis A. Traganitis , Georgios B. Giannakis

Author Name Disambiguation (AND) is the task of resolving which author mentions in a bibliographic database refer to the same real-world person, and is a critical ingredient of digital library applications such as search and citation…

Digital Libraries · Computer Science 2022-02-22 Shivashankar Subramanian , Daniel King , Doug Downey , Sergey Feldman

In supervised machine learning for author name disambiguation, negative training data are often dominantly larger than positive training data. This paper examines how the ratios of negative to positive training data can affect the…

Information Retrieval · Computer Science 2018-08-06 Jinseok Kim , Jenna Kim

Robust analysis of coauthorship networks is based on high quality data. However, ground-truth data are usually unavailable. Empirical data suffer several types of errors, a typical one of which is called merging error, identifying different…

Physics and Society · Physics 2018-12-27 Zheng Xie

In cluster analysis interest lies in probabilistically capturing partitions of individuals, items or observations into groups, such that those belonging to the same group share similar attributes or relational profiles. Bayesian posterior…

Methodology · Statistics 2017-03-23 Riccardo Rastelli , Nial Friel

We present a manually-labeled Author Name Disambiguation(AND) Dataset called WhoisWho, which consists of 399,255 documents and 45,187 distinct authors with 421 ambiguous author names. To label such a great amount of AND data of high…

Social and Information Networks · Computer Science 2020-07-07 Zhuoyue Xiao , Yutao Zhang , Bo Chen , Xiaozhao Liu , Jie Tang

In this article we propose a novel method to perform unsupervised clustering of different forms of Institute names. We use only author and affiliation metadata to perform the clustering without any string or pattern matching. After…

Digital Libraries · Computer Science 2025-10-21 Achal Agrawal , Jeet Mukherjee

Racial disparity in academia is a widely acknowledged problem. The quantitative understanding of racial based systemic inequalities is an important step towards a more equitable research system. However, because of the lack of robust…

Computers and Society · Computer Science 2022-03-09 Diego Kozlowski , Dakota S. Murray , Alexis Bell , Will Hulsey , Vincent Larivière , Thema Monroe-White , Cassidy R. Sugimoto

Several recent deep neural networks experiments leverage the generalist-specialist paradigm for classification. However, no formal study compared the performance of different clustering algorithms for class assignment. In this paper we…

Machine Learning · Computer Science 2016-09-14 Sébastien Arnold

Human annotations are an important source of information in the development of natural language understanding approaches. As under the pressure of productivity annotators can assign different labels to a given text, the quality of produced…

Computation and Language · Computer Science 2020-10-29 Kristian Miok , Gregor Pirs , Marko Robnik-Sikonja

Scholarly data is growing continuously containing information about the articles from a plethora of venues including conferences, journals, etc. Many initiatives have been taken to make scholarly data available as Knowledge Graphs (KGs).…

Artificial Intelligence · Computer Science 2022-06-02 Cristian Santini , Genet Asefa Gesese , Silvio Peroni , Aldo Gangemi , Harald Sack , Mehwish Alam

In this paper, we present a method to automatically build large labeled datasets for the author ambiguity problem in the academic world by leveraging the authoritative academic resources, ORCID and DOI. Using the method, we built LAGOS-AND,…

Digital Libraries · Computer Science 2022-07-15 Li Zhang , Wei Lu , Jinqing Yang

A recurrent neural network that has been trained to separately model the language of several documents by unknown authors is used to measure similarity between the documents. It is able to find clues of common authorship even when the…

Computation and Language · Computer Science 2016-08-17 Douglas Bagnall

In this paper, we study the problem of author identification under double-blind review setting, which is to identify potential authors given information of an anonymized paper. Different from existing approaches that rely heavily on feature…

Machine Learning · Computer Science 2016-12-20 Ting Chen , Yizhou Sun

An author name disambiguation (AND) algorithm identifies a unique author entity record from all similar or same publication records in scholarly or similar databases. Typically, a clustering method is used that requires calculation of…

Information Retrieval · Computer Science 2017-09-28 Kunho Kim , Athar Sefid , C. Lee Giles

We investigate how author name homonymy distorts clustered large-scale co-author networks, and present a simple, effective, scalable and generalizable algorithm to ameliorate such distortions. We evaluate the performance of the algorithm to…

Digital Libraries · Computer Science 2011-06-14 Theresa Velden , Asif-ul Haque , Carl Lagoze

Literature search is arguably one of the most important phases of the academic and non-academic research. The increase in the number of published papers each year makes manual search inefficient and furthermore insufficient. Hence,…

Information Retrieval · Computer Science 2012-09-27 Onur Küçüktunç , Erik Saule , Kamer Kaya , Ümit V. Çatalyürek

An important issue in clustering concerns the avoidance of false positives while searching for clusters. This work addressed this problem considering agglomerative methods, namely single, average, median, complete, centroid and Ward's…

Machine Learning · Computer Science 2020-06-30 Eric K. Tokuda , Cesar H. Comin , Luciano da F. Costa

Realizing when a model is right for a wrong reason is not trivial and requires a significant effort by model developers. In some cases an input salience method, which highlights the most important parts of the input, may reveal problematic…

Computation and Language · Computer Science 2023-01-12 Sebastian Ebert , Alice Shoshana Jakobovits , Katja Filippova