中文
相关论文

相关论文: SciTweets -- A Dataset and Annotation Framework fo…

200 篇论文

During the COVID-19 pandemic, scientific knowledge evolved rapidly, accompanied by a surge of misinformation, labelled an infodemic by the WHO. In this context, we study the interaction between science and misinformation on Twitter (now X)…

We introduce a new classification task for scientific statements and release a large-scale dataset for supervised learning. Our resource is derived from a machine-readable representation of the arXiv.org collection of preprint articles. We…

计算与语言 · 计算机科学 2025-03-21 Deyan Ginev , Bruce R. Miller

In this paper, we present our work developed for the scientific web discourse detection task (Task 4a) of CheckThat! 2025. We propose a novel council debate method that simulates structured academic discussions among multiple large language…

计算与语言 · 计算机科学 2025-08-13 Tarık Saraç , Selin Mergen , Mucahid Kutlu

Communication networks, in general, and internet technology, in particular, is a fast-evolving area of research. While it is important to keep track of emerging trends in this domain, it is such a fast-growing area that it can be very…

数字图书馆 · 计算机科学 2017-08-03 Bisma S. Khan , Muaz A. Niazi

Scientific talks are a growing medium for disseminating research, and automatically identifying relevant literature that grounds or enriches a talk would be highly valuable for researchers and students alike. We introduce Reference…

计算与语言 · 计算机科学 2025-10-29 Frederik Broy , Maike Züfle , Jan Niehues

Many altmetric studies analyze which papers were mentioned how often in specific altmetrics sources. In order to study the potential policy relevance of tweets from another perspective, we investigate which tweets were cited in papers. If…

数字图书馆 · 计算机科学 2020-03-26 Robin Haunschild , Lutz Bornmann

The public interest in accurate scientific communication, underscored by recent public health crises, highlights how content often loses critical pieces of information as it spreads online. However, multi-platform analyses of this…

计算机与社会 · 计算机科学 2023-03-14 Sohyeon Hwang , Emőke-Ágnes Horvát , Daniel M. Romero

Publicly available social media archives facilitate research in the social sciences and provide corpora for training and testing a wide range of machine learning and natural language processing methods. With respect to the recent outbreak…

社会与信息网络 · 计算机科学 2020-08-18 Dimitar Dimitrov , Erdal Baran , Pavlos Fafalios , Ran Yu , Xiaofei Zhu , Matthäus Zloch , Stefan Dietze

Structured information extraction from scientific literature is crucial for capturing core concepts and emerging trends in specialized fields. While existing datasets aid model development, most focus on specific publication sections due to…

计算与语言 · 计算机科学 2026-04-06 Decheng Duan , Yingyi Zhang , Jitong Peng , Chengzhi Zhang

Scientific document understanding is challenging as the data is highly domain specific and diverse. However, datasets for tasks with scientific text require expensive manual annotation and tend to be small and limited to only one or a few…

计算与语言 · 计算机科学 2021-05-26 Dustin Wright , Isabelle Augenstein

Online discussions of science involve complex interactions among experts, news media, and social media users as they interpret and disseminate scientific findings. While prior work has examined these actors in isolation, their interplay in…

社会与信息网络 · 计算机科学 2026-03-19 Alexandros Efstratiou , Giuseppe Russo , Luca Luceri

Recent transformer-based approaches demonstrate promising results on relational scientific information extraction. Existing datasets focus on high-level description of how research is carried out. Instead we focus on the subtleties of how…

计算与语言 · 计算机科学 2021-09-23 Ian H. Magnusson , Scott E. Friedman

With the rise in popularity of public social media and micro-blogging services, most notably Twitter, the people have found a venue to hear and be heard by their peers without an intermediary. As a consequence, and aided by the public…

计算与语言 · 计算机科学 2016-06-21 Prashanth Vijayaraghavan , Soroush Vosoughi , Deb Roy

To analyse large numbers of texts, social science researchers are increasingly confronting the challenge of text classification. When manual labeling is not possible and researchers have to find automatized ways to classify texts, computer…

计算与语言 · 计算机科学 2023-10-10 Karina Shyrokykh , Maksym Girnyk , Lisa Dellmuth

Purpose: This paper explores some influencing factors of Twitter mentions of scientific research. The results can help to understand the relationships between various altmetrics. Design/methodology/approach: Data on research mentions in…

数字图书馆 · 计算机科学 2022-12-13 Pablo Dorta-González

In recent times, social media sites such as Twitter have been extensively used for debating politics and public policies. These debates span millions of tweets and numerous topics of public importance. Thus, it is imperative that this vast…

社会与信息网络 · 计算机科学 2014-04-11 Ashwin Rajadesingan , Huan Liu

Identifying suitable datasets for a research question remains challenging because existing dataset search engines rely heavily on metadata quality and keyword overlap, which often fail to capture the semantic intent of scientific…

数字图书馆 · 计算机科学 2026-01-09 Zhiyin Tan , Changxu Duan

The COVID-19 pandemic brought about an extraordinary rate of scientific papers on the topic that were discussed among the general public, although often in biased or misinformed ways. In this paper, we present a mixed-methods analysis aimed…

计算机与社会 · 计算机科学 2024-06-11 Alexandros Efstratiou , Marina Efstratiou , Satrio Yudhoatmojo , Jeremy Blackburn , Emiliano De Cristofaro

One of the major challenges in automatic hate speech detection is the lack of datasets that cover a wide range of biased and unbiased messages and that are consistently labeled. We propose a labeling procedure that addresses some of the…

计算与语言 · 计算机科学 2023-05-01 Gunther Jikeli , Sameer Karali , Daniel Miehling , Katharina Soemer

Existing Natural Language Inference (NLI) datasets, while being instrumental in the advancement of Natural Language Understanding (NLU) research, are not related to scientific text. In this paper, we introduce SciNLI, a large dataset for…

计算与语言 · 计算机科学 2022-03-16 Mobashir Sadat , Cornelia Caragea