中文
相关论文

相关论文: Constructing the CORD-19 Vaccine Dataset

200 篇论文

The COVID-19 pandemic led to 1.1 million deaths in the United States, despite the explosion of coronavirus research. These new findings are slow to translate to clinical interventions, leading to poorer patient outcomes and unnecessary…

计算与语言 · 计算机科学 2023-06-09 Yousuf A. Khan , Clarisse Hokia , Jennifer Xu , Ben Ehlert

Long COVID continues to challenge public health by affecting a significant segment of individuals who have recovered from acute SARS-CoV-2 infection yet endure prolonged and often debilitating symptoms. Social media has emerged as a vital…

社会与信息网络 · 计算机科学 2024-12-30 Nirmalya Thakur

Retrieval Augmented Generation enhances the response accuracy of Large Language Models (LLMs) by integrating retrieval and generation modules with external knowledge, demonstrating particular strength in real-time queries and Visual…

计算与语言 · 计算机科学 2025-09-08 Qixin Sun , Ziqin Wang , Hengyuan Zhao , Yilin Li , Kaiyou Song , Linjiang Huang , Xiaolin Hu , Qingpei Guo , Si Liu

The COVID-19 pandemic exposed significant weaknesses in the healthcare information system. The overwhelming volume of misinformation on social media and other socioeconomic factors created extraordinary challenges to motivate people to take…

计算机与社会 · 计算机科学 2024-07-30 Ashiqur Rahman , Ehsan Mohammadi , Hamed Alhoori

In this system paper we present our contribution to the Constraint 2021 COVID-19 Fake News Detection Shared Task, which poses the challenge of classifying COVID-19 related social media posts as either fake or real. In our system, we address…

计算与语言 · 计算机科学 2021-01-14 Thomas Felber

The COVID-19 pandemic is accompanied by a massive "infodemic" that makes it hard to identify concise and credible information for COVID-19-related questions, like incubation time, infection rates, or the effectiveness of vaccines. As a…

计算与语言 · 计算机科学 2022-04-20 Johannes Graf , Gino Lancho , Patrick Zschech , Kai Heinrich

The sharing of fake news and conspiracy theories on social media has wide-spread negative effects. By designing and applying different machine learning models, researchers have made progress in detecting fake news from text. However,…

计算与语言 · 计算机科学 2022-05-03 Haoming Guo , Tianyi Huang , Huixuan Huang , Mingyue Fan , Gerald Friedland

We present The Vault, a dataset of high-quality code-text pairs in multiple programming languages for training large language models to understand and generate code. We present methods for thoroughly extracting samples that use both…

计算与语言 · 计算机科学 2023-10-31 Dung Nguyen Manh , Nam Le Hai , Anh T. V. Dau , Anh Minh Nguyen , Khanh Nghiem , Jin Guo , Nghi D. Q. Bui

This study presents a data-driven analysis of COVID-19 discourse on YouTube, examining the sentiment, toxicity, and thematic patterns of video content published between January 2023 and October 2024. The analysis involved applying advanced…

社会与信息网络 · 计算机科学 2024-12-24 Vanessa Su , Nirmalya Thakur

Aim of this paper is the description of a new tool to support institutions in the implementation of targeted countermeasures, based on quantitative and multi-scale elements, for the fight and prevention of emergencies, such as the current…

计算机与社会 · 计算机科学 2020-11-12 A. Sebastianelli , F. Mauro , G. Di Cosmo , F. Passarini , M. Carminati , S. L. Ullo

Detecting semantic similarities between sentences is still a challenge today due to the ambiguity of natural languages. In this work, we propose a simple approach to identifying semantically similar questions by combining the strengths of…

计算与语言 · 计算机科学 2020-06-09 Yoan Dimitrov

Malicious accounts spreading misinformation has led to widespread false and misleading narratives in recent times, especially during the COVID-19 pandemic, and social media platforms struggle to eliminate these contents rapidly. This is…

社会与信息网络 · 计算机科学 2022-02-28 Karishma Sharma , Emilio Ferrara , Yan Liu

Previous work established skip-gram word2vec models could be used to mine knowledge in the materials science literature for the discovery of thermoelectrics. Recent transformer architectures have shown great progress in language modeling…

计算与语言 · 计算机科学 2020-12-14 Leo K. Tam , Xiaosong Wang , Daguang Xu

Background. After a year and half and over 4 million deaths, the COVID-19 pandemic continues to be widespread, and its related topics continue to dominate the global media. Although COVID-19 diagnoses have been well monitored, neither the…

计算机与社会 · 计算机科学 2021-08-19 Guangqing Chi , Junjun Yin , M. Luke Smith , Yosef Bodovski

This paper describes a large global dataset on people's discourse and responses to the COVID-19 pandemic over the Twitter platform. From 28 January 2020 to 1 June 2022, we collected and processed over 252 million Twitter posts from more…

计算与语言 · 计算机科学 2022-06-28 Raj Kumar Gupta , Ajay Vishwanath , Yinping Yang

Dialogue systems are widely used in AI to support timely and interactive communication with users. We propose a general-purpose dialogue system architecture that leverages computational argumentation to perform reasoning and provide…

计算与语言 · 计算机科学 2021-10-18 Bettina Fazzinga , Andrea Galassi , Paolo Torroni

With the tremendous growth in the number of scientific papers being published, searching for references while writing a scientific paper is a time-consuming process. A technique that could add a reference citation at the appropriate place…

计算与语言 · 计算机科学 2019-03-18 Chanwoo Jeong , Sion Jang , Hyuna Shin , Eunjeong Park , Sungchul Choi

In this paper, we present a first multilingual cross-domain dataset of 5182 fact-checked news articles for COVID-19, collected from 04/01/2020 to 15/05/2020. We have collected the fact-checked articles from 92 different fact-checking…

计算机与社会 · 计算机科学 2020-06-23 Gautam Kishore Shahi , Durgesh Nandini

We describe Mega-COV, a billion-scale dataset from Twitter for studying COVID-19. The dataset is diverse (covers 268 countries), longitudinal (goes as back as 2007), multilingual (comes in 100+ languages), and has a significant number of…

社会与信息网络 · 计算机科学 2021-02-09 Muhammad Abdul-Mageed , AbdelRahim Elmadany , El Moatez Billah Nagoudi , Dinesh Pabbi , Kunal Verma , Rannie Lin

We present a system that allows life-science researchers to search a linguistically annotated corpus of scientific texts using patterns over dependency graphs, as well as using patterns over token sequences and a powerful variant of boolean…

计算与语言 · 计算机科学 2020-06-09 Hillel Taub-Tabib , Micah Shlain , Shoval Sadde , Dan Lahav , Matan Eyal , Yaara Cohen , Yoav Goldberg