中文
相关论文

相关论文: Fair multilingual vandalism detection system for W…

200 篇论文

Social movements use social computing systems to complement offline mobilizations, but prior literature has focused almost exclusively on movement actors' use of social media. In this paper, we analyze participation and attention to topics…

社会与信息网络 · 计算机科学 2016-11-07 Marlon Twyman , Brian C. Keegan , Aaron Shaw

This research showcases the innovative integration of Large Language Models into machine learning workflows for traffic incident management, focusing on the classification of incident severity using accident reports. By leveraging features…

机器学习 · 计算机科学 2024-05-01 Artur Grigorev , Khaled Saleh , Yuming Ou , Adriana-Simona Mihaita

Wikipedia, an open collaborative website, can be edited by anyone, even anonymously, thus becoming victim to ill-intentioned changes. Therefore, ranking Wikipedia authors by calculating impact measures based on the edit history can help to…

数字图书馆 · 计算机科学 2017-09-06 Sebastian Neef

This paper describes our participation in SemEval-2020 Task 12: Multilingual Offensive Language Detection. We jointly-trained a single model by fine-tuning Multilingual BERT to tackle the task across all the proposed languages: English,…

计算与语言 · 计算机科学 2020-08-17 Juan Manuel Pérez , Aymé Arango , Franco Luque

We present the results and the main findings of SemEval-2019 Task 6 on Identifying and Categorizing Offensive Language in Social Media (OffensEval). The task was based on a new dataset, the Offensive Language Identification Dataset (OLID),…

计算与语言 · 计算机科学 2019-04-30 Marcos Zampieri , Shervin Malmasi , Preslav Nakov , Sara Rosenthal , Noura Farra , Ritesh Kumar

Identifying claims requiring verification is a critical task in automated fact-checking, especially given the proliferation of misinformation on social media platforms. Despite notable progress, challenges remain-particularly in handling…

计算与语言 · 计算机科学 2025-07-22 Rrubaa Panchendrarajan , Arkaitz Zubiaga

We present the results of research with the goal of automatically creating a multilingual thesaurus based on the freely available resources of Wikipedia and WordNet. Our goal is to increase resources for natural language processing tasks…

计算与语言 · 计算机科学 2013-03-07 Jessica Ramírez , Masayuki Asahara , Yuji Matsumoto

This paper presents an automated adversarial mechanism called WikipediaBot. WikipediaBot allows an adversary to create and control a bot infrastructure for the purpose of adversarial edits of Wikipedia articles. The WikipediaBot is a…

密码学与安全 · 计算机科学 2020-06-26 Filipo Sharevski , Peter Jachim

Wikipedia is among the largest examples of collective intelligence on the Web with over 61 million articles covering over 320 languages. Although edited and maintained by an active workforce of human volunteers, Wikipedia is highly reliant…

人机交互 · 计算机科学 2025-09-29 Neal Reeves , Elena Simperl

This paper presents LOLA, a massively multilingual large language model trained on more than 160 languages using a sparse Mixture-of-Experts Transformer architecture. Our architectural and implementation choices address the challenge of…

The number of people accessing online services is increasing day by day, and with new users, comes a greater need for effective and responsive cyber-security. Our goal in this study was to find out if there are common patterns within the…

密码学与安全 · 计算机科学 2024-05-15 Gábor Antal , Balázs Mosolygó , Norbert Vándor , Péter Hegedüs

In this paper, we introduce HateBERT, a re-trained BERT model for abusive language detection in English. The model was trained on RAL-E, a large-scale dataset of Reddit comments in English from communities banned for being offensive,…

计算与语言 · 计算机科学 2021-02-05 Tommaso Caselli , Valerio Basile , Jelena Mitrović , Michael Granitzer

Text alignment and text quality are critical to the accuracy of Machine Translation (MT) systems, some NLP tools, and any other text processing tasks requiring bilingual data. This research proposes a language independent bi-sentence…

计算与语言 · 计算机科学 2015-10-16 Krzysztof Wołk

We propose a novel framework for cross-lingual content flagging with limited target-language data, which significantly outperforms prior work in terms of predictive performance. The framework is based on a nearest-neighbour architecture. It…

Every culture and language is unique. Our work expressly focuses on the uniqueness of culture and language in relation to human affect, specifically sentiment and emotion semantics, and how they manifest in social multimedia. We develop…

多媒体 · 计算机科学 2015-10-08 Brendan Jou , Tao Chen , Nikolaos Pappas , Miriam Redi , Mercan Topkara , Shih-Fu Chang

Massive web-crawled image-text datasets lay the foundation for recent progress in multimodal learning. These datasets are designed with the goal of training a model to do well on standard computer vision benchmarks, many of which, however,…

计算机视觉与模式识别 · 计算机科学 2025-09-16 Thao Nguyen , Matthew Wallingford , Sebastin Santy , Wei-Chiu Ma , Sewoong Oh , Ludwig Schmidt , Pang Wei Koh , Ranjay Krishna

Information presented in Wikipedia articles must be attributable to reliable published sources in the form of references. This study examines over 5 million Wikipedia articles to assess the reliability of references in multiple language…

计算机与社会 · 计算机科学 2023-09-06 Aitolkyn Baigutanova , Diego Saez-Trumper , Miriam Redi , Meeyoung Cha , Pablo Aragón

Combating online hate speech in multilingual settings requires approaches that go beyond English-centric models and capture the cultural and linguistic diversity of global online discourse. This paper presents a comprehensive survey and…

计算与语言 · 计算机科学 2026-03-23 Zahra Safdari Fesaghandis , Suman Kalyan Maity

Software vulnerabilities (SVs) pose a critical threat to safety-critical systems, driving the adoption of AI-based approaches such as machine learning and deep learning for software vulnerability detection. Despite promising results, most…

密码学与安全 · 计算机科学 2025-10-07 Van Nguyen , Surya Nepal , Xingliang Yuan , Tingmin Wu , Fengchao Chen , Carsten Rudolph

Existing research on fairness evaluation of document classification models mainly uses synthetic monolingual data without ground truth for author demographic attributes. In this work, we assemble and publish a multilingual Twitter corpus…

计算与语言 · 计算机科学 2020-03-04 Xiaolei Huang , Linzi Xing , Franck Dernoncourt , Michael J. Paul
‹ 上一页 1 8 9 10 下一页 ›