中文
相关论文

相关论文: Fair multilingual vandalism detection system for W…

200 篇论文

The presence of offensive language on social media platforms and the implications this poses is becoming a major concern in modern society. Given the enormous amount of content created every day, automatic methods are required to detect and…

计算与语言 · 计算机科学 2023-03-24 Gudbjartur Ingi Sigurbergsson , Leon Derczynski

MIRACL (Multilingual Information Retrieval Across a Continuum of Languages) is a multilingual dataset we have built for the WSDM 2023 Cup challenge that focuses on ad hoc retrieval across 18 different languages, which collectively encompass…

An edit summary is a succinct comment written by a Wikipedia editor explaining the nature of, and reasons for, an edit to a Wikipedia page. Edit summaries are crucial for maintaining the encyclopedia: they are the first thing seen by…

计算与语言 · 计算机科学 2024-08-20 Marija Šakota , Isaac Johnson , Guosheng Feng , Robert West

Recent years have witnessed a proliferation of valuable original natural language contents found in subscription-based media outlets, web novel platforms, and outputs of large language models. However, these contents are susceptible to…

计算与语言 · 计算机科学 2023-06-12 KiYoon Yoo , Wonhyuk Ahn , Jiho Jang , Nojun Kwak

Software plays a crucial role in our daily lives, and therefore the quality and security of software systems have become increasingly important. However, vulnerabilities in software still pose a significant threat, as they can have serious…

软件工程 · 计算机科学 2023-09-18 Chaozheng Wang , Zongjie Li , Yun Peng , Shuzheng Gao , Sirong Chen , Shuai Wang , Cuiyun Gao , Michael R. Lyu

Automated content moderation for collaborative knowledge hubs like Wikipedia or Wikidata is an important yet challenging task due to multiple factors. In this paper, we construct a database of discussions happening around articles marked…

计算与语言 · 计算机科学 2025-03-14 Hsuvas Borkakoty , Luis Espinosa-Anke

As one of the Web's primary multilingual knowledge sources, Wikipedia is read by millions of people across the globe every day. Despite this global readership, little is known about why users read Wikipedia's various language editions. To…

计算机与社会 · 计算机科学 2018-12-04 Florian Lemmerich , Diego Sáez-Trumper , Robert West , Leila Zia

Document Layout Parsing serves as a critical gateway for Artificial Intelligence (AI) to access and interpret the world's vast stores of structured knowledge. This process,which encompasses layout detection, text recognition, and relational…

计算机视觉与模式识别 · 计算机科学 2025-12-18 Yumeng Li , Guang Yang , Hao Liu , Bowen Wang , Colin Zhang

The increasing accessibility of the internet facilitated social media usage and encouraged individuals to express their opinions liberally. Nevertheless, it also creates a place for content polluters to disseminate offensive posts or…

计算与语言 · 计算机科学 2021-03-02 Omar Sharif , Eftekhar Hossain , Mohammed Moshiul Hoque

Wikipedia is a community-created encyclopedia that contains information about notable people from different countries, epochs and disciplines and aims to document the world's knowledge from a neutral point of view. However, the narrow…

计算机与社会 · 计算机科学 2015-03-25 Claudia Wagner , David Garcia , Mohsen Jadidi , Markus Strohmaier

A major challenge for many analyses of Wikipedia dynamics -- e.g., imbalances in content quality, geographic differences in what content is popular, what types of articles attract more editor discussion -- is grouping the very diverse range…

计算机与社会 · 计算机科学 2021-03-02 Isaac Johnson , Martin Gerlach , Diego Sáez-Trumper

The interest in offensive content identification in social media has grown substantially in recent years. Previous work has dealt mostly with post level annotations. However, identifying offensive spans is useful in many ways. To help…

计算与语言 · 计算机科学 2021-04-20 Tharindu Ranasinghe , Marcos Zampieri

In this article we address the problem of text passage alignment across interlingual article pairs in Wikipedia. We develop methods that enable the identification and interlinking of text passages written in different languages and…

计算与语言 · 计算机科学 2019-05-22 Simon Gottschalk , Elena Demidova

Social media platforms are critical spaces for public discourse, shaping opinions and community dynamics, yet their widespread use has amplified harmful content, particularly hate speech, threatening online safety and inclusivity. While…

计算与语言 · 计算机科学 2025-06-11 Muhammad Usman , Muhammad Ahmad , M. Shahiki Tash , Irina Gelbukh , Rolando Quintero Tellez , Grigori Sidorov

We introduce WikiLingua, a large-scale, multilingual dataset for the evaluation of crosslingual abstractive summarization systems. We extract article and summary pairs in 18 languages from WikiHow, a high quality, collaborative resource of…

计算与语言 · 计算机科学 2020-10-08 Faisal Ladhak , Esin Durmus , Claire Cardie , Kathleen McKeown

Content moderation is the process of flagging content based on pre-defined platform rules. There has been a growing need for AI moderators to safeguard users as well as protect the mental health of human moderators from traumatic content.…

计算与语言 · 计算机科学 2023-02-21 Meng Ye , Karan Sikka , Katherine Atwell , Sabit Hassan , Ajay Divakaran , Malihe Alikhani

This paper proposes to use distributed representation of words (word embeddings) in cross-language textual similarity detection. The main contributions of this paper are the following: (a) we introduce new cross-language similarity…

计算与语言 · 计算机科学 2017-02-13 J. Ferrero , F. Agnes , L. Besacier , D. Schwab

In order to effectively analyze information regarding ongoing events that impact local communities across language and country borders, researchers often need to perform multilingual data analysis. This analysis can be particularly…

计算机与社会 · 计算机科学 2018-01-23 Simon Gottschalk , Elena Demidova , Viola Bernacchi , Richard Rogers

This paper is a deep investigation of cross-language plagiarism detection methods on a new recently introduced open dataset, which contains parallel and comparable collections of documents with multiple characteristics (different genres,…

计算与语言 · 计算机科学 2017-05-25 Jeremy Ferrero , Laurent Besacier , Didier Schwab , Frederic Agnes

The usage of non-authoritative data for disaster management presents the opportunity of accessing timely information that might not be available through other means, as well as the challenge of dealing with several layers of biases.…

信息检索 · 计算机科学 2020-01-27 Valerio Lorini , Javier Rando , Diego Saez-Trumper , Carlos Castillo