中文
相关论文

相关论文: Fair multilingual vandalism detection system for W…

200 篇论文

The detection of offensive language in the context of a dialogue has become an increasingly important application of natural language processing. The detection of trolls in public forums (Gal\'an-Garc\'ia et al., 2016), and the deployment…

计算与语言 · 计算机科学 2019-08-20 Emily Dinan , Samuel Humeau , Bharath Chintagunta , Jason Weston

This paper replicates, extends, and refutes conclusions made in a study published in PLoS ONE ("Even Good Bots Fight"), which claimed to identify substantial levels of conflict between automated software agents (or bots) in Wikipedia using…

计算机与社会 · 计算机科学 2018-10-18 R. Stuart Geiger , Aaron Halfaker

With more than 11 times as many pageviews as the next largest edition, English Wikipedia dominates global knowledge access relative to other language editions. Readers are prone to assuming English Wikipedia as a superset of all language…

人机交互 · 计算机科学 2026-01-21 Zining Wang , Yuxuan Zhang , Dongwook Yoon , Nicholas Vincent , Farhan Samir , Vered Shwartz

In recent times, we have seen a rapid development of large Vision-Language Models (VLMs). They have shown impressive results on academic benchmarks, primarily in widely spoken languages but lack performance on low-resource languages and…

In this paper we present a novel method for retrieving information in languages other than that of the query. We use this technique in combination with existing traditional Cross Language Information Retrieval (CLIR) techniques to improve…

信息检索 · 计算机科学 2009-06-17 Mikhail Basilyan

This article analyzes one month of edits to Wikipedia in order to examine the role of users editing multiple language editions (referred to as multilingual users). Such multilingual users may serve an important function in diffusing…

计算机与社会 · 计算机科学 2014-05-13 Scott A. Hale

Wikipedia is the largest existing knowledge repository that is growing on a genuine crowdsourcing support. While the English Wikipedia is the most extensive and the most researched one with over five million articles, comparatively little…

数字图书馆 · 计算机科学 2017-10-20 Kristina Ban , Matjaz Perc , Zoran Levnajic

Wikipedia is an online encyclopedia available in 285 languages. It composes an extremely relevant Knowledge Base (KB), which could be leveraged by automatic systems for several purposes. However, the structure and organisation of such…

计算与语言 · 计算机科学 2021-05-13 Ruben Cardoso , Afonso Mendes , Andre Lamurias

With 60M articles in more than 300 language versions, Wikipedia is the largest platform for open and freely accessible knowledge. While the available content has been growing continuously at a rate of around 200K new articles each month,…

社会与信息网络 · 计算机科学 2024-10-08 Akhil Arora , Robert West , Martin Gerlach

We present an approach based on multilingual sentence embeddings to automatically extract parallel sentences from the content of Wikipedia articles in 85 languages, including several dialects or low-resource languages. We do not limit the…

计算与语言 · 计算机科学 2019-07-17 Holger Schwenk , Vishrav Chaudhary , Shuo Sun , Hongyu Gong , Francisco Guzmán

Early detection of relevant locations in a piece of news is especially important in extreme events such as environmental disasters, war conflicts, disease outbreaks, or political turmoils. Additionally, this detection also helps recommender…

计算与语言 · 计算机科学 2022-12-23 Víctor Suárez-Paniagua , Steven Derby , Tri Kurniawan Wijaya

One of the biggest challenges hindering progress in low-resource and multilingual machine translation is the lack of good evaluation benchmarks. Current evaluation benchmarks either lack good coverage of low-resource languages, consider…

The vast amount of online information today poses challenges for non-English speakers, as much of it is concentrated in high-resource languages such as English and French. Wikipedia reflects this imbalance, with content in low-resource…

计算与语言 · 计算机科学 2025-04-08 Siddharth Khincha , Tushar Kataria , Ankita Anand , Dan Roth , Vivek Gupta

With the increasing popularity of large language models, concerns over content authenticity have led to the development of myriad watermarking schemes. These schemes can be used to detect a machine-generated text via an appropriate key,…

机器学习 · 统计学 2025-09-26 Soham Bonnerjee , Sayar Karmakar , Subhrajyoty Roy

With the ability to watch Wikipedia and Wikidata edits in realtime, the online encyclopedia and the knowledge base have become increasingly used targets of research for the detection of breaking news events. In this paper, we present a case…

社会与信息网络 · 计算机科学 2014-03-19 Thomas Steiner

This paper introduces the Cross-lingual Fact Extraction and VERification (XFEVER) dataset designed for benchmarking the fact verification models across different languages. We constructed it by translating the claim and evidence texts of…

计算与语言 · 计算机科学 2023-10-26 Yi-Chen Chang , Canasai Kruengkrai , Junichi Yamagishi

The Internet has significantly expanded the potential for global collaboration, allowing millions of users to contribute to collective projects like Wikipedia. While prior work has assessed the success of online collaborations, most…

计算机与社会 · 计算机科学 2025-03-17 Abraham Israeli , David Jurgens , Daniel Romero

Software Engineering (SE) communities such as Stack Overflow have become unwelcoming, particularly through members' use of offensive language. Research has shown that offensive language drives users away from active engagement within these…

软件工程 · 计算机科学 2022-11-02 Jithin Cheriyan , Bastin Tony Roy Savarimuthu , Stephen Cranefield

Hate speech detection is a challenging problem with most of the datasets available in only one language: English. In this paper, we conduct a large scale analysis of multilingual hate speech in 9 languages from 16 different sources. We…

社会与信息网络 · 计算机科学 2020-12-10 Sai Saketh Aluru , Binny Mathew , Punyajoy Saha , Animesh Mukherjee

Advances in Natural Language Processing (NLP) have revolutionized the way researchers and practitioners address crucial societal problems. Large language models are now the standard to develop state-of-the-art solutions for text detection…

机器学习 · 计算机科学 2022-05-20 Gaurav Verma , Rohit Mujumdar , Zijie J. Wang , Munmun De Choudhury , Srijan Kumar