中文
相关论文

相关论文: Towards a corpus for credibility assessment in sof…

200 篇论文

We study the problem of finding fake online news. This is an important problem as news of questionable credibility have recently been proliferating in social media at an alarming scale. As this is an understudied problem, especially for…

计算与语言 · 计算机科学 2019-11-20 Momchil Hardalov , Ivan Koychev , Preslav Nakov

This paper simulates a low-resource setting across 17 languages in order to evaluate embedding similarity, stability, and reliability under different conditions. The goal is to use corpus similarity measures before training to predict…

计算与语言 · 计算机科学 2022-06-10 Jonathan Dunn , Haipeng Li , Damian Sastre

Misalignment between claims and their cited evidence is a common failure mode in reports generated by large language models, limiting their reliability in scientific and other high-stakes settings. We present DeepSciVerify, a two-stage…

人工智能 · 计算机科学 2026-05-28 Shaghayegh Sadeghi , Khashayar Khajavi , Rise Adhikari , Alexander Tessier

This paper presents a pipeline to detect and explain anomalous reviews in online platforms. The pipeline is made up of three modules and allows the detection of reviews that do not generate value for users due to either worthless or…

计算与语言 · 计算机科学 2024-02-29 David Novoa-Paradela , Oscar Fontenla-Romero , Bertha Guijarro-Berdiñas

The Forum for Information Retrieval (FIRE) started a shared task this year for classification of comments of different code segments. This is binary text classification task where the objective is to identify whether comments given for…

信息检索 · 计算机科学 2023-08-14 Sruthi S , Tanmay Basu

Fake information poses one of the major threats for society in the 21st century. Identifying misinformation has become a key challenge due to the amount of fake news that is published daily. Yet, no approach is established that addresses…

信息检索 · 计算机科学 2021-03-30 Vishwani Gupta , Katharina Beckh , Sven Giesselbach , Dennis Wegener , Tim Wirtz

In today's content-centric Internet, blogs are becoming increasingly popular and important from a data analysis perspective. According to Wikipedia, there were over 156 million public blogs on the Internet as of February 2011. Blogs are a…

信息检索 · 计算机科学 2014-05-16 Srayan Datta

This paper presents a tertiary review of software quality measurement research. To conduct this review, we examined an initial dataset of 7,811 articles and found 75 relevant and high-quality secondary analyses of software quality research.…

软件工程 · 计算机科学 2021-07-30 Kaylea Champion , Sejal Khatri , Benjamin Mako Hill

The concept of social trust has attracted an attention of information processors/data scientists and information consumers / business firms. One of the main reasons for acquiring the value of SBD is to provide frameworks and methodologies…

社会与信息网络 · 计算机科学 2021-04-20 Bilal Abu-Salih , Pornpit Wongthongtham , Dengya Zhu , Kit Yan Chan , Amit Rudra

With the spread of online social networks, it is more and more difficult to monitor all the user-generated content. Automating the moderation process of the inappropriate exchange content on Internet has thus become a priority task. Methods…

计算与语言 · 计算机科学 2021-01-19 Noé Cecillon , Vincent Labatut , Richard Dufour , Georges Linares

Media seems to have become more partisan, often providing a biased coverage of news catering to the interest of specific groups. It is therefore essential to identify credible information content that provides an objective narrative of an…

人工智能 · 计算机科学 2017-05-16 Subhabrata Mukherjee , Gerhard Weikum

Context: Quality requirements (QRs) are a topic of constant discussions both in industry and academia. Debates entwine around the definition of quality requirements, the way how to handle them, or their importance for project success. While…

软件工程 · 计算机科学 2020-02-10 Andreas Vogelsang , Jonas Eckhardt , Daniel Mendez , Moritz Berger

The premises of an argument give evidence or other reasons to support a conclusion. However, the amount of support required depends on the generality of a conclusion, the nature of the individual premises, and similar. An argument whose…

计算与语言 · 计算机科学 2021-10-27 Timon Gurcke , Milad Alshomary , Henning Wachsmuth

This study reviewed the use of Large Language Models (LLMs) in healthcare, focusing on their training corpora, customization techniques, and evaluation metrics. A systematic search of studies from 2021 to 2024 identified 61 articles. Four…

计算与语言 · 计算机科学 2025-02-18 Shuqi Yang , Mingrui Jing , Shuai Wang , Jiaxin Kou , Manfei Shi , Weijie Xing , Yan Hu , Zheng Zhu

We introduce the Self-Annotated Reddit Corpus (SARC), a large corpus for sarcasm research and for training and evaluating systems for sarcasm detection. The corpus has 1.3 million sarcastic statements -- 10 times more than any previous…

计算与语言 · 计算机科学 2018-03-26 Mikhail Khodak , Nikunj Saunshi , Kiran Vodrahalli

Text similarity detection aims at measuring the degree of similarity between a pair of texts. Corpora available for text similarity detection are designed to evaluate the algorithms to assess the paraphrase level among documents. In this…

信息检索 · 计算机科学 2017-03-14 Juan-Manuel Torres-Moreno , Gerardo Sierra , Peter Peinl

Semantic annotations have to satisfy quality constraints to be useful for digital libraries, which is particularly challenging on large and diverse datasets. Confidence scores of multi-label classification methods typically refer only to…

信息检索 · 计算机科学 2018-06-08 Martin Toepfer , Christin Seifert

Comparable corpus is a set of topic aligned documents in multiple languages, which are not necessarily translations of each other. These documents are useful for multilingual natural language processing when there is no parallel text…

计算与语言 · 计算机科学 2025-08-05 Motaz Saad , David Langlois , Kamel Smaili

Multi-document summarization (MDS) is the task of reflecting key points from any set of documents into a concise text paragraph. In the past, it has been used to aggregate news, tweets, product reviews, etc. from various sources. Owing to…

计算与语言 · 计算机科学 2020-10-06 Alvin Dey , Tanya Chowdhury , Yash Kumar Atri , Tanmoy Chakraborty

Maintaining high quality content is one of the foremost objectives of any web-based collaborative service that depends on a large number of users. In such systems, it is nearly impossible for automated scripts to judge semantics as it is to…

社会与信息网络 · 计算机科学 2012-02-15 Shreyas Sekar , Patrick Maillé