中文
相关论文

相关论文: Whose Language Counts as High Quality? Measuring L…

200 篇论文

As large language models (LLMs) are more frequently used in retrieval-augmented generation pipelines, it is increasingly relevant to study their behavior under knowledge conflicts. Thus far, the role of the source of the retrieved…

计算与语言 · 计算机科学 2026-04-20 Jakob Schuster , Vagrant Gautam , Katja Markert

With the advance of language models, privacy protection is receiving more attention. Training data extraction is therefore of great importance, as it can serve as a potential tool to assess privacy leakage. However, due to the difficulty of…

计算与语言 · 计算机科学 2023-06-02 Weichen Yu , Tianyu Pang , Qian Liu , Chao Du , Bingyi Kang , Yan Huang , Min Lin , Shuicheng Yan

Despite major advances in multilingual modeling, large quality disparities persist across languages. Besides the obvious impact of uneven training resources, typological properties have also been proposed to determine the intrinsic…

计算与语言 · 计算机科学 2026-02-04 Vitalii Hirak , Jaap Jumelet , Arianna Bisazza

Wikipedia is the world's largest online encyclopedia, but maintaining article quality through collaboration is challenging. Wikipedia designed a quality scale, but with such a manual assessment process, many articles remain unassessed. We…

计算与语言 · 计算机科学 2023-10-04 Pedro Miguel Moás , Carla Teixeira Lopes

Agents based on Large Language Models (LLMs) are increasingly being deployed as interfaces to information on online platforms. These agents filter, prioritize, and synthesize information retrieved from the platforms' back-end databases or…

We present improved models for the granular detection and sub-classification news media bias in English news articles. We compare the performance of zero-shot versus fine-tuned large pre-trained neural transformer language models, explore…

计算与语言 · 计算机科学 2026-01-08 Tim Menzner , Jochen L. Leidner

Social media feed algorithms infer user preferences from their past behaviors. Yet what drives engagement often diverges from what users value. We examine this gap between stated preferences (what users say they prefer) and revealed…

人机交互 · 计算机科学 2026-04-14 Do Won Kim , Cody Buntain , Giovanni Luca Ciampaglia

Over the last few years, Text classification is one of the fundamental tasks in natural language processing (NLP) in which the objective is to categorize text documents into one of the predefined classes. The news is full of our life.…

计算与语言 · 计算机科学 2022-01-26 Ke Yahan , Ruyi Qu , Lu Xiaoxia

As Large Language Models (LLMs) scale, data curation has shifted from maximizing volume to optimizing the signal-to-noise ratio by performing quality filtering. However, for many languages, native high quality data is insufficient to train…

计算与语言 · 计算机科学 2026-04-23 Yassine Turki , Vinko Sabolčec , Bettina Messmer , Martin Jaggi

This study presents a comparative analysis of 55 Wikipedia language editions employing a citation index alongside a synthetic quality measure. Specifically, we identified the most significant Wikipedia articles within distinct topical…

信息检索 · 计算机科学 2025-05-23 Włodzimierz Lewoniewski , Krzysztof Węcel , Witold Abramowicz

To seek reliable information sources for news events, we introduce a novel task of expert recommendation, which aims to identify trustworthy sources based on their previously quoted statements. To achieve this, we built a novel dataset,…

信息检索 · 计算机科学 2024-06-18 Wenjia Zhang , Lin Gui , Rob Procter , Yulan He

Language students are most engaged while reading texts at an appropriate difficulty level. However, existing methods of evaluating text difficulty focus mainly on vocabulary and do not prioritize grammatical features, hence they do not work…

计算与语言 · 计算机科学 2017-02-17 Shuhan Wang , Erik Andersen

Large Language Models (LLMs) are advancing quickly and impacting people's lives for better or worse. In higher education, concerns have emerged such as students' misuse of LLMs and degraded education outcomes. To unpack the ethical concerns…

Wikidata has been increasingly adopted by many communities for a wide variety of applications, which demand high-quality knowledge to deliver successful results. In this paper, we develop a framework to detect and analyze low-quality…

人工智能 · 计算机科学 2021-11-22 Kartik Shenoy , Filip Ilievski , Daniel Garijo , Daniel Schwabe , Pedro Szekely

Web content quality measurement is crucial to various web content processing applications. This paper will explore multi-scale features which may affect the quality of a host, and develop automatic statistical methods to evaluate the Web…

信息检索 · 计算机科学 2013-04-24 Guang-Gang Geng , Xiao-Bo Jin , Xin-Chang Zhang , De-Xian Zhang

Large language models have achieved impressive progress in multilingual translation, yet they continue to face challenges with certain language pairs-particularly those with limited training data or significant linguistic divergence from…

计算与语言 · 计算机科学 2025-07-01 Yumeng Lin , Xufeng Duan , David Haslett , Yige Chen , Zhenguang G. Cai

This publication describes the motivation and generation of $Q_{bias}$, a large dataset of Google and Bing search queries, a scraping tool and dataset for biased news articles, as well as language models for the investigation of bias in…

信息检索 · 计算机科学 2023-11-30 Fabian Haak , Philipp Schaer

This paper focuses on the development of an advanced intelligent article scoring system that not only assesses the overall quality of written work but also offers detailed feature-based scoring tailored to various article genres. By…

计算与语言 · 计算机科学 2024-10-21 Chihang Wang , Yuxin Dong , Zhenhong Zhang , Ruotong Wang , Shuo Wang , Jiajing Chen

Predicting the political bias and the factuality of reporting of entire news outlets are critical elements of media profiling, which is an understudied but an increasingly important research direction. The present level of proliferation of…

计算与语言 · 计算机科学 2020-05-12 Ramy Baly , Georgi Karadzhov , Jisun An , Haewoon Kwak , Yoan Dinkov , Ahmed Ali , James Glass , Preslav Nakov

Wikidata is one of the most important sources of structured data on the web, built by a worldwide community of volunteers. As a secondary source, its contents must be backed by credible references; this is particularly important as Wikidata…

人工智能 · 计算机科学 2021-09-21 Gabriel Amaral , Alessandro Piscopo , Lucie-Aimée Kaffee , Odinaldo Rodrigues , Elena Simperl
‹ 上一页 1 8 9 10 下一页 ›