English
Related papers

Related papers: Scaling Author Name Disambiguation with CNF Blocki…

200 papers

Author Name Disambiguation (AND) is a critical task for digital libraries aiming to link existing authors with their respective publications. Due to the lack of persistent identifiers used by researchers and the presence of intrinsic…

Digital Libraries · Computer Science 2025-03-19 Francesca Cappelli , Giovanni Colavizza , Silvio Peroni

We investigate how author name homonymy distorts clustered large-scale co-author networks, and present a simple, effective, scalable and generalizable algorithm to ameliorate such distortions. We evaluate the performance of the algorithm to…

Digital Libraries · Computer Science 2011-06-14 Theresa Velden , Asif-ul Haque , Carl Lagoze

Entity resolution is a challenging and hot research area in the field of Information Systems since last decade. Author Name Disambiguation (AND) in Bibliographic Databases (BD) like DBLP , Citeseer , and Scopus is a specialized field of…

Social and Information Networks · Computer Science 2020-04-15 Muhammad Shoaib , Ali Daud , Tehmina Amjad

De-duplication---identification of distinct records referring to the same real-world entity---is a well-known challenge in data integration. Since very large datasets prohibit the comparison of every pair of records, {\em blocking} has been…

Databases · Computer Science 2011-11-17 Anish Das Sarma , Ankur Jain , Ashwin Machanavajjhala , Philip Bohannon

Several `edge-discovery' applications over graph-based data models are known to have worst-case quadratic time complexity in the nodes, even if the discovered edges are sparse. One example is the generic link discovery problem between two…

Artificial Intelligence · Computer Science 2017-07-04 Mayank Kejriwal

In this paper, we present a method to automatically build large labeled datasets for the author ambiguity problem in the academic world by leveraging the authoritative academic resources, ORCID and DOI. Using the method, we built LAGOS-AND,…

Digital Libraries · Computer Science 2022-07-15 Li Zhang , Wei Lu , Jinqing Yang

In real-world, our DNA is unique but many people share names. This phenomenon often causes erroneous aggregation of documents of multiple persons who are namesake of one another. Such mistakes deteriorate the performance of document…

Social and Information Networks · Computer Science 2017-09-12 Baichuan Zhang , Mohammad Al Hasan

Scholars have often relied on name initials to resolve name ambiguities in large-scale coauthorship network research. This approach bears the risk of incorrectly merging or splitting author identities. The use of initial-based…

Digital Libraries · Computer Science 2015-04-03 Jinseok Kim , Jana Diesner

Entity Resolution concerns identifying co-referent entity pairs across datasets. A typical workflow comprises two steps. In the first step, a blocking method uses a one-many function called a blocking scheme to map entities to blocks. In…

Databases · Computer Science 2015-01-09 Mayank Kejriwal , Daniel P. Miranker

Author disambiguation arises when different authors share the same name, which is a critical task in digital libraries, such as DBLP, CiteULike, CiteSeerX, etc. While the state-of-the-art methods have developed various paper embedding-based…

Information Retrieval · Computer Science 2020-12-01 Na Li , Renyu Zhu , Xiaoxu Zhou , Xiangnan He , Wenyuan Cai , Ming Gao , Aoying Zhou

Author Name Disambiguation (AND) is a long-standing challenge in bibliometrics and scientometrics, as name ambiguity undermines the accuracy of bibliographic databases and the reliability of research evaluation. This study addresses the…

Adequately disambiguating author names in bibliometric databases is a precondition for conducting reliable analyses at the author level. In the case of bibliometric studies that include many researchers, it is not possible to disambiguate…

Digital Libraries · Computer Science 2019-04-30 Alexander Tekles , Lutz Bornmann

There are a number of solutions that perform unsupervised name disambiguation based on the similarity of bibliographic records or common co-authorship patterns. Whether the use of these advanced methods, which are often difficult to…

Digital Libraries · Computer Science 2013-08-06 Staša Milojević

Author name ambiguity decreases the quality and reliability of information retrieved from digital libraries. Existing methods have tried to solve this problem by predefining a feature set based on expert's knowledge for a specific dataset.…

Digital Libraries · Computer Science 2020-02-24 Hung Nghiep Tran , Tin Huynh , Tien Do

Entity Resolution, also called record linkage or deduplication, refers to the process of identifying and merging duplicate versions of the same entity into a unified representation. The standard practice is to use a Rule based or Machine…

Artificial Intelligence · Computer Science 2016-09-22 Janani Balaji , Faizan Javed , Mayank Kejriwal , Chris Min , Sam Sander , Ozgur Ozturk

National exercises for the evaluation of research activity by universities are becoming regular practice in ever more countries. These exercises have mainly been conducted through the application of peer-review methods. Bibliometrics has…

Digital Libraries · Computer Science 2018-12-21 Ciriaco Andrea D'Angelo , Cristiano Giuffrida , Giovanni Abramo

Most NLP approaches to entity linking and coreference resolution focus on retrieving similar mentions using sparse or dense text representations. The common "Wikification" task, for instance, retrieves candidate Wikipedia articles for each…

Computation and Language · Computer Science 2022-09-02 Ryan Muther , David Smith

Unsupervised learning of the Dawid-Skene (D&S) model from noisy, incomplete and crowdsourced annotations has been a long-standing challenge, and is a critical step towards reliably labeling massive data. A recent work takes a coupled…

Machine Learning · Computer Science 2021-06-15 Shahana Ibrahim , Xiao Fu

Name disambiguation -- a fundamental problem in online academic systems -- is now facing greater challenges with the increasing growth of research papers. For example, on AMiner, an online academic search platform, about 10% of names own…

Information Retrieval · Computer Science 2023-06-07 Bo Chen , Jing Zhang , Fanjin Zhang , Tianyi Han , Yuqing Cheng , Xiaoyan Li , Yuxiao Dong , Jie Tang

The study of science at the individual micro-level frequently requires the disambiguation of author names. The creation of author's publication oeuvres involves matching the list of unique author names to names used in publication…

Digital Libraries · Computer Science 2013-04-23 Linda Reijnhoudt , Rodrigo Costas , Ed Noyons , Katy Boerner , Andrea Scharnhorst