中文
相关论文

相关论文: Random Forest DBSCAN for USPTO Inventor Name Disam…

200 篇论文

Patent data represent a significant source of information on innovation and the evolution of technology through networks of citations, co-invention and co-assignment of new patents. A major obstacle to extracting useful information from…

数字图书馆 · 计算机科学 2016-01-11 Greg Morrison , Massimo Riccaboni , Fabio Pammolli

To train algorithms for supervised author name disambiguation, many studies have relied on hand-labeled truth data that are very laborious to generate. This paper shows that labeled training data can be automatically generated using…

数字图书馆 · 计算机科学 2021-02-08 Jinseok Kim , Jinmo Kim , Jason Owen-Smith

Author name ambiguity in a digital library may affect the findings of research that mines authorship data of the library. This study evaluates author name disambiguation in DBLP, a widely used but insufficiently evaluated digital library…

数字图书馆 · 计算机科学 2018-07-31 Jinseok Kim

Patent classification is an essential task in patent information management and patent knowledge mining. It is very important to classify patents related to artificial intelligence, which is the biggest topic these days. However, artificial…

计算与语言 · 计算机科学 2023-03-07 Yongmin Yoo , Tak-Sung Heo , Dongjin Lim , Deaho Seo

Name ambiguity is common in academic digital libraries, such as multiple authors having the same name. This creates challenges for academic data management and analysis, thus name disambiguation becomes necessary. The procedure of name…

机器学习 · 计算机科学 2024-04-02 Wenjin Xie , Siyuan Liu , Xiaomeng Wang , Tao Jia

The ability to distinctly and properly collate an individual researcher's publications is crucial for ensuring appropriate recognition, guiding the allocation of research funding and informing hiring decisions. However, accurately grouping…

天体物理仪器与方法 · 物理学 2025-11-17 Vicente Amado Olivo , Wolfgang Kerzendorf , Bangjing Lu , Joshua V. Shields , Andreas Flörs , Nutan Chen

An author name disambiguation (AND) algorithm identifies a unique author entity record from all similar or same publication records in scholarly or similar databases. Typically, a clustering method is used that requires calculation of…

信息检索 · 计算机科学 2017-09-28 Kunho Kim , Athar Sefid , C. Lee Giles

The wealth of data being gathered about humans and their surroundings drives new machine learning applications in various fields. Consequently, more and more often, classifiers are trained using not only numerical data but also complex data…

机器学习 · 计算机科学 2022-04-13 Maciej Piernik , Dariusz Brzezinski , Pawel Zawadzki

This work addresses the problem of author name homonymy in the Web of Science. Aiming for an efficient, simple and straightforward solution, we introduce a novel probabilistic similarity measure for author name disambiguation based on…

信息检索 · 计算机科学 2018-08-14 Tobias Backes

In real-world, our DNA is unique but many people share names. This phenomenon often causes erroneous aggregation of documents of multiple persons who are namesake of one another. Such mistakes deteriorate the performance of document…

社会与信息网络 · 计算机科学 2017-09-12 Baichuan Zhang , Mohammad Al Hasan

This research proposes a data segmentation algorithm which combines t-SNE, DBSCAN, and Random Forest classifier to form an end-to-end pipeline that separates data into natural clusters and produces a characteristic profile of each cluster…

机器学习 · 计算机科学 2021-01-14 Timothy DeLise

Author name disambiguation in bibliographic databases is the problem of grouping together scientific publications written by the same person, accounting for potential homonyms and/or synonyms. Among solutions to this problem, digital…

数字图书馆 · 计算机科学 2016-05-05 Gilles Louppe , Hussein Al-Natsheh , Mateusz Susik , Eamonn Maguire

The problem of disambiguation of company names poses a significant challenge in extracting useful information from patents. This issue biases research outcomes as it mostly underestimates the number of patents attributed to companies,…

信息检索 · 计算机科学 2024-03-20 Grazia Sveva Ascione , Valerio Sterzi

We present PatentsView-Evaluation, a Python package that enables researchers to evaluate the performance of inventor name disambiguation systems such as PatentsView.org. The package includes benchmark datasets and evaluation tools, and aims…

数字图书馆 · 计算机科学 2023-01-11 Olivier Binette , Sarvo Madhavan , Jack Butler , Beth Anne Card , Emily Melluso , Christina Jones

We present a novel algorithm and validation method for disambiguating author names in very large bibliographic data sets and apply it to the full Web of Science (WoS) citation index. Our algorithm relies only upon the author and citation…

数字图书馆 · 计算机科学 2014-12-11 Christian Schulz , Amin Mazloumian , Alexander M Petersen , Orion Penner , Dirk Helbing

An algorithm to improve performance parameter for unsupervised decision forest clustering and density estimation is presented. Specifically, a dual assignment parameter is introduced as a density estimator by combining Random Forest and…

计算机视觉与模式识别 · 计算机科学 2015-07-19 Hayder Albehadili , Naz Islam

DBSCAN is a well-known density-based clustering algorithm to discover arbitrary shape clusters. While conceptually simple in serial, the algorithm is challenging to efficiently parallelize on manycore GPU architectures. Common pitfalls,…

分布式、并行与集群计算 · 计算机科学 2023-06-30 Andrey Prokopenko , Damien Lebrun-Grandie , Daniel Arndt

In this paper, we extend some usual techniques of classification resulting from a large-scale data-mining and network approach. This new technology, which in particular is designed to be suitable to big data, is used to construct an open…

物理与社会 · 物理学 2017-07-05 Antonin Bergeaud , Yoann Potiron , Juste Raimbault

We introduce a novel method for converting text data into abstract image representations, which allows image-based processing techniques (e.g. image classification networks) to be applied to text-based comparison problems. We apply the…

计算与语言 · 计算机科学 2020-02-07 Stephen M. Petrie , T'Mir D. Julius

Multi-label classification is a challenging task, particularly in domains where the number of labels to be predicted is large. Deep neural networks are often effective at multi-label classification of images and textual data. When dealing…

机器学习 · 计算机科学 2023-03-30 Nikolaos Mylonas , Ioannis Mollas , Nick Bassiliades , Grigorios Tsoumakas
‹ 上一页 1 2 3 10 下一页 ›