English
Related papers

Related papers: CiteRadar: A Citation Intelligence Platform for Re…

200 papers

We present HySemRAG, a framework that combines Extract, Transform, Load (ETL) pipelines with Retrieval-Augmented Generation (RAG) to automate large-scale literature synthesis and identify methodological research gaps. The system addresses…

Information Retrieval · Computer Science 2025-08-11 Alejandro Godinez

We investigate platform-native citation farming on ResearchGate by analyzing almost 3000 papers uploaded by five suspected boosting-service provider accounts. From the uploaded papers and associated metadata, we construct both paper-level…

Social and Information Networks · Computer Science 2026-04-17 Cenk Erdogan , Bennett Daniel , Benedikt Wotka , Ashish Sai , Adriana Iamnitchi

Identifying suitable datasets for a research question remains challenging because existing dataset search engines rely heavily on metadata quality and keyword overlap, which often fail to capture the semantic intent of scientific…

Digital Libraries · Computer Science 2026-01-09 Zhiyin Tan , Changxu Duan

Trustworthy language models should provide both correct and verifiable answers. However, citations generated directly by standalone LLMs are often unreliable. As a result, current systems insert citations by querying an external retriever…

Artificial Intelligence · Computer Science 2026-04-07 Yukun Huang , Sanxing Chen , Jian Pei , Manzil Zaheer , Bhuwan Dhingra

In quantitative SciSci (science of science) studies, accurately identifying individual scholars is paramount for scientific data analysis. However, the variability in how names are represented-due to commonality, abbreviations, and…

Information Retrieval · Computer Science 2025-03-05 Renyu Zhao , Yunxin Chen

In today's data-driven research landscape, dataset visibility and accessibility play a crucial role in advancing scientific knowledge. At the same time, data citation is essential for maintaining academic integrity, acknowledging…

Digital Libraries · Computer Science 2025-06-23 Fakhri Momeni , Janete Saldanha Bach , Brigitte Mathiak , Peter Mutschke

Large language models (LLMs) have created new opportunities to enhance the efficiency of scholarly activities; however, challenges persist in the ethical deployment of AI assistance, including (1) the trustworthiness of AI-generated…

Computation and Language · Computer Science 2026-02-27 Mengze Hong , Di Jiang , Chen Jason Zhang , Zichang Guo , Yawen Li , Jun Chen , Shaobo Cui , Zhiyang Su

h-index retrieved by citation indexes (Scopus, Google scholar, and Web of Science) is used to measure the scientific performance and the research impact studies based on the number of publications and citations of a scientist. It also is…

This study describes the methodology and analyses the results of the process of mapping entities between two large open bibliographic metadata collections, OpenCitations Meta and OpenAlex. The primary objective of this mapping is to…

Digital Libraries · Computer Science 2026-04-27 Elia Rizzetto , Silvio Peroni

Scientific progress depends on researchers' ability to synthesize the growing body of literature. Can large language models (LMs) assist scientists in this task? We introduce OpenScholar, a specialized retrieval-augmented LM that answers…

This article evaluates the quality of data collection in individual-level desktop information tracking used in the social sciences and shows that the existing approaches face sampling issues, validity issues due to the lack of content-level…

The paper presents first results of the CitEcCyr project funded by RANEPA. The project aims to create a source of open citation data for research papers written in Russian. Compared to existing sources of citation data, CitEcCyr is working…

Digital Libraries · Computer Science 2017-10-03 Jose Manuel Barrueco , Thomas Krichel , Sergey Parinov , Victor Lyapunov , Oxana Medvedeva , Varvara Sergeeva

Battery research is a rapidly growing and highly interdisciplinary field, making it increasingly difficult to track relevant expertise and identify potential collaborators across institutional boundaries. In this work, we present a pipeline…

Computation and Language · Computer Science 2026-04-23 Luca Foppiano , Sae Dieb , Malik Zain , Kazuki Kasama , Keitaro Sodeyama , Mikiko Tanifuji

SciLaD is a novel, large-scale dataset of scientific language constructed entirely using open-source frameworks and publicly available data sources. It comprises a curated English split containing over 10 million scientific publications and…

One compelling use of citation networks is to characterize papers by their relationships to the surrounding literature. We propose a method to characterize papers by embedding them into two distinct "co-factor" spaces: one describing how…

Methodology · Statistics 2025-10-07 Alex Hayes , Karl Rohe

In this article we propose a novel method to perform unsupervised clustering of different forms of Institute names. We use only author and affiliation metadata to perform the clustering without any string or pattern matching. After…

Digital Libraries · Computer Science 2025-10-21 Achal Agrawal , Jeet Mukherjee

We consider a resource-constrained updater, such as Google Scholar, which wishes to update the citation records of a group of researchers, who have different mean citation rates (and optionally, different importance coefficients), in such a…

Information Theory · Computer Science 2020-04-28 Melih Bastopcu , Sennur Ulukus

The launching of Scopus and Google Scholar, and methodological developments in Social Network Analysis have made many more indicators for evaluating journals available than the traditional Impact Factor, Cited Half-life, and Immediacy Index…

Physics and Society · Physics 2009-09-25 Loet Leydesdorff

New sources of citation data have recently become available, such as Microsoft Academic, Dimensions, and the OpenCitations Index of CrossRef open DOI-to-DOI citations (COCI). Although these have been compared to the Web of Science (WoS),…

Digital Libraries · Computer Science 2021-02-02 Alberto Martín-Martín , Mike Thelwall , Enrique Orduna-Malea , Emilio Delgado López-Cózar

Scientific metadata often suffer from incompleteness, inconsistency, and formatting errors, which hinder effective discovery and reuse of the associated datasets. We present a method that combines GPT-4 with structured metadata templates…

Information Retrieval · Computer Science 2025-06-10 Sowmya S Sundaram , Rafael S. Gonçalves , Mark A Musen