English
Related papers

Related papers: Deep Author Name Disambiguation using DBLP Data

200 papers

Disparities in authorship and citations across gender can have substantial adverse consequences not just on the disadvantaged genders, but also on the field of study as a whole. Measuring gender gaps is a crucial step towards addressing…

Digital Libraries · Computer Science 2020-09-07 Saif M. Mohammad

This paper examines the difference and similarities between the two on-line computer science citation databases DBLP and CiteSeer. The database entries in DBLP are inserted manually while the CiteSeer entries are obtained autonomously via a…

Digital Libraries · Computer Science 2007-05-23 Vaclav Petricek , Ingemar J. Cox , Hui Han , Isaac G. Councill , C. Lee Giles

Tools to explore scientific literature are essential for scientists, especially in biomedicine, where about a million new papers are published every year. Many such tools provide users the ability to search for specific entities (e.g.…

Computation and Language · Computer Science 2021-07-05 Sunil Mohan , Rico Angell , Nick Monath , Andrew McCallum

Massive-scale historical document collections are crucial for social science research. Despite increasing digitization, these documents typically lack unique cross-document identifiers for individuals mentioned within the texts, as well as…

Computation and Language · Computer Science 2024-06-25 Abhishek Arora , Emily Silcock , Leander Heldring , Melissa Dell

In this study, I analyze the DBLP bibliographic database to study role of single author publications in the computer science literature between 1940 and 2019. I examine the demographics and reception by computing the population fraction,…

Digital Libraries · Computer Science 2020-01-03 Brian K. Ryu

Deep Research Agents are a prominent category of LLM-based agents. By autonomously orchestrating multistep web exploration, targeted retrieval, and higher-order synthesis, they transform vast amounts of online information into…

Computation and Language · Computer Science 2025-06-16 Mingxuan Du , Benfeng Xu , Chiwei Zhu , Xiaorui Wang , Zhendong Mao

Conference publications in computer science (CS) have attracted scholarly attention due to their unique status as a main research outlet unlike other science fields where journals are dominantly used for communicating research findings. One…

Digital Libraries · Computer Science 2018-06-29 Jinseok Kim

Authorship attribution (AA) is the task of identifying the most likely author of a query document from a predefined set of candidate authors. We introduce a two-stage retrieve-and-rerank framework that finetunes LLMs for cross-genre AA.…

Computation and Language · Computer Science 2025-10-21 Shantanu Agarwal , Joel Barry , Steven Fincke , Scott Miller

Authorship attribution aims to identify the origin or author of a document. Traditional approaches have heavily relied on manual features and fail to capture long-range correlations, limiting their effectiveness. Recent advancements…

Computation and Language · Computer Science 2024-10-30 Zhengmian Hu , Tong Zheng , Heng Huang

As part of the NLP Scholar project, we created a single unified dataset of NLP papers and their meta-information (including citation numbers), by extracting and aligning information from the ACL Anthology and Google Scholar. In this paper,…

Digital Libraries · Computer Science 2020-06-03 Saif M. Mohammad

This work investigates the effectiveness of different pseudonymization techniques, ranging from rule-based substitutions to using pre-trained Large Language Models (LLMs), on a variety of datasets and models used for two widely used NLP…

Computation and Language · Computer Science 2023-06-12 Oleksandr Yermilov , Vipul Raheja , Artem Chernodub

With the rapid development of NLP, large-scale language models (LLMs) excel in various tasks across multiple domains now. However, existing benchmarks may not adequately measure these models' capabilities, especially when faced with new…

Computation and Language · Computer Science 2023-10-24 Xunjian Yin , Baizhou Huang , Xiaojun Wan

The automatic verification of document authorships is important in various settings. Researchers are for example judged and compared by the amount and impact of their publications and public figures are confronted by their posts on social…

Machine Learning · Computer Science 2022-08-25 Maximilian Stubbemann , Gerd Stumme

Acronym Disambiguation (AD) is crucial for natural language understanding on various sources, including biomedical reports, scientific papers, and search engine queries. However, existing acronym disambiguation benchmarks and tools are…

Computation and Language · Computer Science 2023-03-15 Lihu Chen , Gaël Varoquaux , Fabian M. Suchanek

Understanding how co-authors distribute credit is critical for accurately assessing scholarly collaboration. In this study, we uncover the implicit structures within scientific teamwork by systematically analyzing author contributions…

Digital Libraries · Computer Science 2026-02-26 Itai Assraf , Michael Fire

Academic and scientific publishing practices have changed significantly in recent years. This paper presents an analysis of 17 million research papers published since 2000 to explore changes in authorship and content practices. It shows a…

Digital Libraries · Computer Science 2025-10-28 Padraig Cunningham , Padhraic Smyth , Barry Smyth

We propose a novel deep learning model for joint document-level entity disambiguation, which leverages learned neural representations. Key components are entity embeddings, a neural attention mechanism over local context windows, and a…

Computation and Language · Computer Science 2017-08-02 Octavian-Eugen Ganea , Thomas Hofmann

We describe a deployed scalable system for organizing published scientific literature into a heterogeneous graph to facilitate algorithmic manipulation and discovery. The resulting literature graph consists of more than 280M nodes,…

On December 4, 2025, Anthropic released Anthropic Interviewer, an AI tool for running qualitative interviews at scale, along with a public dataset of 1,250 interviews with professionals, including 125 scientists, about their use of AI for…

Cryptography and Security · Computer Science 2026-01-12 Tianshi Li

Gender biases in scholarly metrics remain a persistent concern, despite numerous bibliometric studies exploring their presence and absence across productivity, impact, acknowledgment, and self-citations. However, methodological…

Digital Libraries · Computer Science 2025-01-31 HaeJin Lee , Shubhanshu Mishra , Apratim Mishra , Zhiwen You , Jinseok Kim , Jana Diesner