中文
相关论文

相关论文: SciDMT: A Large-Scale Corpus for Detecting Scienti…

200 篇论文

We present a content-based method for recommending citations in an academic paper draft. We embed a given query document into a vector space, then use its nearest neighbors as candidates, and rerank the candidates using a discriminative…

计算与语言 · 计算机科学 2018-02-26 Chandra Bhagavatula , Sergey Feldman , Russell Power , Waleed Ammar

Misalignment between claims and their cited evidence is a common failure mode in reports generated by large language models, limiting their reliability in scientific and other high-stakes settings. We present DeepSciVerify, a two-stage…

人工智能 · 计算机科学 2026-05-28 Shaghayegh Sadeghi , Khashayar Khajavi , Rise Adhikari , Alexander Tessier

Biomedical literature is a rapidly expanding field of science and technology. Classification of biomedical texts is an essential part of biomedicine research, especially in the field of biology. This work proposes the fine-tuned DistilBERT,…

计算与语言 · 计算机科学 2024-04-23 Ziqing Guo

The purpose of this paper is to update the review of Bornmann and Daniel (2008) presenting a narrative review of studies on citations in scientific documents. The current review covers 41 studies published between 2006 and 2018. Bornmann…

数字图书馆 · 计算机科学 2019-07-30 Iman Tahamtan , Lutz Bornmann

Many journals post accepted articles online before they are formally published in an issue. Early citation impact evidence for these articles could be helpful for timely research evaluation and to identify potentially important articles…

数字图书馆 · 计算机科学 2018-02-22 Kayvan Kousha , Mike Thelwall , Mahshid Abdoli

The field of scientometrics has shown the power of citation-based clusters for literature analysis, yet this technique has barely been used for information retrieval tasks. This work evaluates the performance of citation based-clusters for…

数字图书馆 · 计算机科学 2023-10-06 Juan Pablo Bascur , Suzan Verberne , Nees Jan van Eck , Ludo Waltman

Learned representations of scientific documents can serve as valuable input features for downstream tasks without further fine-tuning. However, existing benchmarks for evaluating these representations fail to capture the diversity of…

计算与语言 · 计算机科学 2023-11-14 Amanpreet Singh , Mike D'Arcy , Arman Cohan , Doug Downey , Sergey Feldman

We describe a strategy for identifying the universe of research publications relevant to the application and development of artificial intelligence. The approach leverages the arXiv corpus of scientific preprints, in which authors choose…

数字图书馆 · 计算机科学 2020-05-29 James Dunham , Jennifer Melot , Dewey Murdick

Large language models (LLMs) present a promising yet challenging frontier for automated source citation in scientific communication. Previous approaches to citation generation have been limited by citation ambiguity and LLM…

计算与语言 · 计算机科学 2025-04-14 Yash Saxena , Deepa Tilwani , Ali Mohammadi , Edward Raff , Amit Sheth , Srinivasan Parthasarathy , Manas Gaur

As large language models (LLMs) advance, the ultimate vision for their role in science is emerging: we could build an AI collaborator to effectively assist human beings throughout the entire scientific research process. We refer to this…

We introduce the STEM (Science, Technology, Engineering, and Medicine) Dataset for Scientific Entity Extraction, Classification, and Resolution, version 1.0 (STEM-ECR v1.0). The STEM-ECR v1.0 dataset has been developed to provide a…

信息检索 · 计算机科学 2020-07-29 Jennifer D'Souza , Anett Hoppe , Arthur Brack , Mohamad Yaser Jaradeh , Sören Auer , Ralph Ewerth

Scientist learn early on how to cite scientific sources to support their claims. Sometimes, however, scientists have challenges determining where a citation should be situated -- or, even worse, fail to cite a source altogether.…

计算与语言 · 计算机科学 2024-05-21 Tong Zeng , Daniel E. Acuna

We present a corpus of 5,000 richly annotated abstracts of medical articles describing clinical randomized controlled trials. Annotations include demarcations of text spans that describe the Patient population enrolled, the Interventions…

计算与语言 · 计算机科学 2018-06-13 Benjamin Nye , Junyi Jessy Li , Roma Patel , Yinfei Yang , Iain J. Marshall , Ani Nenkova , Byron C. Wallace

We focus on electronic theses and dissertations (ETDs), aiming to improve access and expand their utility, since more than 6 million are publicly available, and they constitute an important corpus to aid research and education across…

计算机视觉与模式识别 · 计算机科学 2021-06-30 Sampanna Yashwant Kahu , William A. Ingram , Edward A. Fox , Jian Wu

Large Language Models (LLMs), such as GPT-3 and BERT, reshape how textual content is written and communicated. These models have the potential to generate scientific content that is indistinguishable from that written by humans. Hence, LLMs…

计算与语言 · 计算机科学 2024-11-19 Bushra Alhijawi , Rawan Jarrar , Aseel AbuAlRub , Arwa Bader

The exponential growth of online textual content across diverse domains has necessitated advanced methods for automated text classification. Large Language Models (LLMs) based on transformer architectures have shown significant success in…

计算与语言 · 计算机科学 2025-09-09 Zhyar Rzgar K Rostam , Gábor Kertész

Scientific publications are the primary means to communicate research discoveries, where the writing quality is of crucial importance. However, prior work studying the human editing process in this domain mainly focused on the abstract or…

计算与语言 · 计算机科学 2022-11-01 Chao Jiang , Wei Xu , Samuel Stevens

In this paper we present a phenomenological approach to describe a complex system: scientific research impact through Citation Mining. The novel concept of Citation Mining, a combination of citation bibliometrics and text mining, is used…

物理与社会 · 物理学 2007-05-23 J. A del Rio , R. N. Kostoff , E. O. Garcia , A. M. Ramirez , J. A. Humenik

Extensive efforts in the past have been directed toward the development of summarization datasets. However, a predominant number of these resources have been (semi)-automatically generated, typically through web data crawling, resulting in…

计算与语言 · 计算机科学 2024-03-11 Sotaro Takeshita , Tommaso Green , Ines Reinig , Kai Eckert , Simone Paolo Ponzetto

The explosion of scientific literature has made the efficient and accurate extraction of structured data a critical component for advancing scientific knowledge and supporting evidence-based decision-making. However, existing tools often…

人机交互 · 计算机科学 2025-11-06 Xingbo Wang , Samantha L. Huey , Rui Sheng , Saurabh Mehta , Fei Wang