English
Related papers

Related papers: CiMaTe: Citation Count Prediction Effectively Leve…

200 papers

In this paper we address the challenge of extracting scientific references from patents. We approach the problem as a sequence labelling task and investigate the merits of BERT models to the extraction of these long sequences. References in…

Information Retrieval · Computer Science 2021-03-11 Ken Voskuil , Suzan Verberne

Even for domain experts, it is a non-trivial task to verify a scientific claim by providing supporting or refuting evidence rationales. The situation worsens as misinformation is proliferated on social media or news websites, manually or…

Computation and Language · Computer Science 2025-05-19 Xiangci Li , Gully Burns , Nanyun Peng

Finding hot topics in scholarly fields can help researchers to keep up with the latest concepts, trends, and inventions in their field of interest. Due to the rarity of complete large-scale scholarly data, earlier studies target this…

Social and Information Networks · Computer Science 2017-10-19 Jinghao Zhao , Hao Wu , Fengyu Deng , Wentian Bao , Wencheng Tang , Luoyi Fu , Xinbing Wang

Obtaining large-scale annotated data for NLP tasks in the scientific domain is challenging and expensive. We release SciBERT, a pretrained language model based on BERT (Devlin et al., 2018) to address the lack of high-quality, large-scale…

Computation and Language · Computer Science 2019-09-12 Iz Beltagy , Kyle Lo , Arman Cohan

Citation context analysis (CCA) is an important task in natural language processing that studies how and why scholars discuss each others' work. Despite decades of study, traditional frameworks for CCA have largely relied on…

Computation and Language · Computer Science 2021-08-03 Anne Lauscher , Brandon Ko , Bailey Kuehl , Sophie Johnson , David Jurgens , Arman Cohan , Kyle Lo

We report characteristics of in-text citations in over five million full text articles from two large databases - the PubMed Central Open Access subset and Elsevier journals - as functions of time, textual progression, and scientific field.…

Digital Libraries · Computer Science 2017-10-10 Kevin W. Boyack , Nees Jan van Eck , Giovanni Colavizza , Ludo Waltman

The identification of the most significant concepts in unstructured data is of critical importance in various practical applications. Despite the large number of methods that have been put forth to extract the main topics of texts, a…

Digital Libraries · Computer Science 2023-01-18 Jorge A. V. Tohalino , Thiago C. Silva , Diego R. Amancio

AI-generated text detection plays an increasingly important role in various fields. In this study, we developed an efficient AI-generated text detection model based on the BERT algorithm, which provides new ideas and methods for solving…

Computation and Language · Computer Science 2024-10-15 Hao Wang , Jianwei Li , Zhengyu Li

The attribution technique enhances the credibility of LLMs by adding citations to the generated sentences, enabling users to trace back to the original sources and verify the reliability of the output. However, existing instruction-tuned…

Information Retrieval · Computer Science 2026-03-24 Yue Yu , Ting Bai , HengZhi Lan , Li Qian , Li Peng , Jie Wu , Wei Liu , Jian Luan , Chuan Shi

Despite the increasing use of citation-based metrics for research evaluation purposes, we do not know yet which metrics best deliver on their promise to gauge the significance of a scientific paper or a patent. We assess 17 network-based…

Social and Information Networks · Computer Science 2020-07-10 Shuqi Xu , Manuel Sebastian Mariani , Linyuan Lü , Matúš Medo

Literature recommendation is essential for researchers to find relevant articles in an ever-growing academic field. However, traditional methods often struggle due to data limitations and methodological challenges. In this work, we…

Applications · Statistics 2025-03-04 Kun Liu , Yan Zhang , Rui Pan , Tianchen Gao , Hansheng Wang

Pre-trained language models have attracted increasing attention in the biomedical domain, inspired by their great success in the general natural language domain. Among the two main branches of pre-trained language models in the general…

Computation and Language · Computer Science 2023-04-04 Renqian Luo , Liai Sun , Yingce Xia , Tao Qin , Sheng Zhang , Hoifung Poon , Tie-Yan Liu

We introduce SelfCite, a novel self-supervised approach that aligns LLMs to generate high-quality, fine-grained, sentence-level citations for the statements in their generated responses. Instead of only relying on costly and labor-intensive…

Computation and Language · Computer Science 2025-06-17 Yung-Sung Chuang , Benjamin Cohen-Wang , Shannon Zejiang Shen , Zhaofeng Wu , Hu Xu , Xi Victoria Lin , James Glass , Shang-Wen Li , Wen-tau Yih

We here describe and present results of a simple neural network that predicts individual researchers' future citation counts based on a variety of data from the researchers' past. For publications available on the open access-server…

Digital Libraries · Computer Science 2019-08-09 Tobias Mistele , Tom Price , Sabine Hossenfelder

Most services built on powerful large-scale language models (LLMs) add citations to their output to enhance credibility. Recent research has paid increasing attention to the question of what reference documents to link to outputs. However,…

Computation and Language · Computer Science 2026-02-06 Kenichiro Ando , Tatsuya Harada

The distribution of the number of academic publications as a function of citation count for a given year is remarkably similar from year to year. We measure this similarity as a width of the distribution and find it to be approximately…

Physics and Society · Physics 2015-11-20 S. R. Goldberg , H. Anthony , T. S. Evans

Identifying articles that relate to infectious diseases is a necessary step for any automatic bio-surveillance system that monitors news articles from the Internet. Unlike scientific articles which are available in a strongly structured…

Computation and Language · Computer Science 2019-11-22 Son Doan , Mike Conway , Nigel Collier

The formulation of good academic paper titles in English is challenging for intermediate English authors (particularly students). This is because such authors are not aware of the type of titles that are generally in use. We aim to realize…

Computation and Language · Computer Science 2021-10-11 Kento Kaku , Masato Kikuchi , Tadachika Ozono , Toramatsu Shintani

This paper proposes a medical literature summary generation method based on the BERT model to address the challenges brought by the current explosion of medical information. By fine-tuning and optimizing the BERT model, we develop an…

Computation and Language · Computer Science 2024-10-29 Jiacheng Hu , Yiru Cang , Guiran Liu , Meiqi Wang , Weijie He , Runyuan Bao

If we want to assess whether the paper in question has had a particularly high or low citation impact compared to other papers, the standard practice in bibliometrics is to normalize citations in respect of the subject category and…

Digital Libraries · Computer Science 2013-07-30 Lutz Bornmann , Werner Marx , Andreas Barth
‹ Prev 1 8 9 10 Next ›