English
Related papers

Related papers: Random Forest DBSCAN for USPTO Inventor Name Disam…

200 papers

Patent data represent a significant source of information on innovation and the evolution of technology through networks of citations, co-invention and co-assignment of new patents. A major obstacle to extracting useful information from…

Digital Libraries · Computer Science 2016-01-11 Greg Morrison , Massimo Riccaboni , Fabio Pammolli

To train algorithms for supervised author name disambiguation, many studies have relied on hand-labeled truth data that are very laborious to generate. This paper shows that labeled training data can be automatically generated using…

Digital Libraries · Computer Science 2021-02-08 Jinseok Kim , Jinmo Kim , Jason Owen-Smith

Author name ambiguity in a digital library may affect the findings of research that mines authorship data of the library. This study evaluates author name disambiguation in DBLP, a widely used but insufficiently evaluated digital library…

Digital Libraries · Computer Science 2018-07-31 Jinseok Kim

Patent classification is an essential task in patent information management and patent knowledge mining. It is very important to classify patents related to artificial intelligence, which is the biggest topic these days. However, artificial…

Computation and Language · Computer Science 2023-03-07 Yongmin Yoo , Tak-Sung Heo , Dongjin Lim , Deaho Seo

Name ambiguity is common in academic digital libraries, such as multiple authors having the same name. This creates challenges for academic data management and analysis, thus name disambiguation becomes necessary. The procedure of name…

Machine Learning · Computer Science 2024-04-02 Wenjin Xie , Siyuan Liu , Xiaomeng Wang , Tao Jia

The ability to distinctly and properly collate an individual researcher's publications is crucial for ensuring appropriate recognition, guiding the allocation of research funding and informing hiring decisions. However, accurately grouping…

Instrumentation and Methods for Astrophysics · Physics 2025-11-17 Vicente Amado Olivo , Wolfgang Kerzendorf , Bangjing Lu , Joshua V. Shields , Andreas Flörs , Nutan Chen

An author name disambiguation (AND) algorithm identifies a unique author entity record from all similar or same publication records in scholarly or similar databases. Typically, a clustering method is used that requires calculation of…

Information Retrieval · Computer Science 2017-09-28 Kunho Kim , Athar Sefid , C. Lee Giles

The wealth of data being gathered about humans and their surroundings drives new machine learning applications in various fields. Consequently, more and more often, classifiers are trained using not only numerical data but also complex data…

Machine Learning · Computer Science 2022-04-13 Maciej Piernik , Dariusz Brzezinski , Pawel Zawadzki

This work addresses the problem of author name homonymy in the Web of Science. Aiming for an efficient, simple and straightforward solution, we introduce a novel probabilistic similarity measure for author name disambiguation based on…

Information Retrieval · Computer Science 2018-08-14 Tobias Backes

In real-world, our DNA is unique but many people share names. This phenomenon often causes erroneous aggregation of documents of multiple persons who are namesake of one another. Such mistakes deteriorate the performance of document…

Social and Information Networks · Computer Science 2017-09-12 Baichuan Zhang , Mohammad Al Hasan

This research proposes a data segmentation algorithm which combines t-SNE, DBSCAN, and Random Forest classifier to form an end-to-end pipeline that separates data into natural clusters and produces a characteristic profile of each cluster…

Machine Learning · Computer Science 2021-01-14 Timothy DeLise

Author name disambiguation in bibliographic databases is the problem of grouping together scientific publications written by the same person, accounting for potential homonyms and/or synonyms. Among solutions to this problem, digital…

Digital Libraries · Computer Science 2016-05-05 Gilles Louppe , Hussein Al-Natsheh , Mateusz Susik , Eamonn Maguire

The problem of disambiguation of company names poses a significant challenge in extracting useful information from patents. This issue biases research outcomes as it mostly underestimates the number of patents attributed to companies,…

Information Retrieval · Computer Science 2024-03-20 Grazia Sveva Ascione , Valerio Sterzi

We present PatentsView-Evaluation, a Python package that enables researchers to evaluate the performance of inventor name disambiguation systems such as PatentsView.org. The package includes benchmark datasets and evaluation tools, and aims…

Digital Libraries · Computer Science 2023-01-11 Olivier Binette , Sarvo Madhavan , Jack Butler , Beth Anne Card , Emily Melluso , Christina Jones

We present a novel algorithm and validation method for disambiguating author names in very large bibliographic data sets and apply it to the full Web of Science (WoS) citation index. Our algorithm relies only upon the author and citation…

Digital Libraries · Computer Science 2014-12-11 Christian Schulz , Amin Mazloumian , Alexander M Petersen , Orion Penner , Dirk Helbing

An algorithm to improve performance parameter for unsupervised decision forest clustering and density estimation is presented. Specifically, a dual assignment parameter is introduced as a density estimator by combining Random Forest and…

Computer Vision and Pattern Recognition · Computer Science 2015-07-19 Hayder Albehadili , Naz Islam

DBSCAN is a well-known density-based clustering algorithm to discover arbitrary shape clusters. While conceptually simple in serial, the algorithm is challenging to efficiently parallelize on manycore GPU architectures. Common pitfalls,…

Distributed, Parallel, and Cluster Computing · Computer Science 2023-06-30 Andrey Prokopenko , Damien Lebrun-Grandie , Daniel Arndt

In this paper, we extend some usual techniques of classification resulting from a large-scale data-mining and network approach. This new technology, which in particular is designed to be suitable to big data, is used to construct an open…

Physics and Society · Physics 2017-07-05 Antonin Bergeaud , Yoann Potiron , Juste Raimbault

We introduce a novel method for converting text data into abstract image representations, which allows image-based processing techniques (e.g. image classification networks) to be applied to text-based comparison problems. We apply the…

Computation and Language · Computer Science 2020-02-07 Stephen M. Petrie , T'Mir D. Julius

Multi-label classification is a challenging task, particularly in domains where the number of labels to be predicted is large. Deep neural networks are often effective at multi-label classification of images and textual data. When dealing…

Machine Learning · Computer Science 2023-03-30 Nikolaos Mylonas , Ioannis Mollas , Nick Bassiliades , Grigorios Tsoumakas
‹ Prev 1 2 3 10 Next ›