中文
相关论文

相关论文: Scientific Statement Classification over arXiv.org

200 篇论文

With thousands of academic articles shared on a daily basis, it has become increasingly difficult to keep up with the latest scientific findings. To overcome this problem, we introduce a new task of disentangled paper summarization, which…

计算与语言 · 计算机科学 2020-11-10 Hiroaki Hayashi , Wojciech Kryściński , Bryan McCann , Nazneen Rajani , Caiming Xiong

Scientific texts often convey authority due to their technical language and complex data. However, this complexity can sometimes lead to the spread of misinformation. Non-experts are particularly susceptible to misleading claims based on…

The rapid growth of scientific literature makes it challenging for researchers to identify novel and impactful ideas, especially across disciplines. Modern artificial intelligence (AI) systems offer new approaches, potentially inspiring…

人工智能 · 计算机科学 2025-01-09 Xuemei Gu , Mario Krenn

The vast body of scientific publications presents an increasing challenge of finding those that are relevant to a given research question, and making informed decisions on their basis. This becomes extremely difficult without the use of…

计算与语言 · 计算机科学 2022-11-09 Piotr Sowinski , Katarzyna Wasielewska-Michniewska , Maria Ganzha , Marcin Paprzycki

This master thesis describes an algorithm for automated categorization of scientific documents using deep learning techniques and compares the results to the results of existing classification algorithms. As an additional goal a reusable…

信息检索 · 计算机科学 2017-06-20 Thomas Krause

Large Language Models (LLMs) have shown promise in assisting scientific discovery. However, such applications are currently limited by LLMs' deficiencies in understanding intricate scientific concepts, deriving symbolic equations, and…

计算与语言 · 计算机科学 2024-11-19 Dan Zhang , Ziniu Hu , Sining Zhoubian , Zhengxiao Du , Kaiyu Yang , Zihan Wang , Yisong Yue , Yuxiao Dong , Jie Tang

This study aimed to identify and analyze the characteristics of highly cited publications in the field of artificial intelligence within the Science Citation Index Expanded from 1991 to 2022. The assessment focused on documents that…

数字图书馆 · 计算机科学 2024-11-19 Yuh-Shan Ho , Juan-Jose Prieto-Gutierrez

Engaging the public with science is critical for a well-informed population. A popular method of scientific communication is documentaries. Once released, it can be difficult to assess the impact of such works on a large scale, due to the…

Multi-document summarization is a challenging task for which there exists little large-scale datasets. We propose Multi-XScience, a large-scale multi-document summarization dataset created from scientific articles. Multi-XScience introduces…

计算与语言 · 计算机科学 2020-10-28 Yao Lu , Yue Dong , Laurent Charlin

Scientific team dynamics are critical in determining the nature and impact of research outputs. However, existing methods for classifying author roles based on self-reports and clustering lack comprehensive contextual analysis of…

数字图书馆 · 计算机科学 2025-02-26 Wonduk Seo , Yi Bu

We introduce the STEM (Science, Technology, Engineering, and Medicine) Dataset for Scientific Entity Extraction, Classification, and Resolution, version 1.0 (STEM-ECR v1.0). The STEM-ECR v1.0 dataset has been developed to provide a…

信息检索 · 计算机科学 2020-07-29 Jennifer D'Souza , Anett Hoppe , Arthur Brack , Mohamad Yaser Jaradeh , Sören Auer , Ralph Ewerth

Scientific writing involves retrieving, summarizing, and citing relevant papers, which can be time-consuming processes in large and rapidly evolving fields. By making these processes inter-operable, natural language processing (NLP)…

计算与语言 · 计算机科学 2023-11-07 Nianlong Gu , Richard H. R. Hahnloser

As the Internet grows in size, so does the amount of text based information that exists. For many application spaces it is paramount to isolate and identify texts that relate to a particular topic. While one-class classification would be…

人工智能 · 计算机科学 2021-11-02 Sameer Khanna

The increasing capacities of large language models (LLMs) have been shown to present an unprecedented opportunity to scale up data analytics in the humanities and social sciences, by automating complex qualitative tasks otherwise typically…

计算与语言 · 计算机科学 2024-10-22 Andres Karjus

Detecting salient parts in text using natural language processing has been widely used to mitigate the effects of information overflow. Nevertheless, most of the datasets available for this task are derived mainly from academic…

计算与语言 · 计算机科学 2024-03-26 Andrés García-Silva , Cristian Berrío , José Manuel Gómez-Pérez

Predicting the future citation rates of academic papers is an important step toward the automation of research evaluation and the acceleration of scientific progress. We present $\textbf{ForeCite}$, a simple but powerful framework to append…

机器学习 · 计算机科学 2025-05-15 Gavin Hull , Alex Bihlo

We present an overview of the SCIDOCA 2025 Shared Task, which focuses on citation discovery and prediction in scientific documents. The task is divided into three subtasks: (1) Citation Discovery, where systems must identify relevant…

数字图书馆 · 计算机科学 2025-09-30 An Dao , Vu Tran , Le-Minh Nguyen , Yuji Matsumoto

Recent developments in machine learning have introduced models that approach human performance at the cost of increased architectural complexity. Efforts to make the rationales behind the models' predictions transparent have inspired an…

计算与语言 · 计算机科学 2020-09-29 Pepa Atanasova , Jakob Grue Simonsen , Christina Lioma , Isabelle Augenstein

We present SciRIFF (Scientific Resource for Instruction-Following and Finetuning), a dataset of 137K instruction-following instances for training and evaluation, covering 54 tasks. These tasks span five core scientific literature…

The exponential increase in scientific literature and online information necessitates efficient methods for extracting knowledge from textual data. Natural language processing (NLP) plays a crucial role in addressing this challenge,…

计算与语言 · 计算机科学 2025-10-22 Zhyar Rzgar K. Rostam , Gábor Kertész