中文
相关论文

相关论文: KIT-TIP-NLP at MultiPride: Continual Learning with…

200 篇论文

This work investigates the use of large-scale, English-only pre-trained models (CLIP and HuBERT) for multilingual image-speech retrieval. For non-English image-speech retrieval, we outperform the current state-of-the-art performance by a…

计算与语言 · 计算机科学 2023-04-11 Layne Berry , Yi-Jen Shih , Hsuan-Fu Wang , Heng-Jui Chang , Hung-yi Lee , David Harwath

Multilingual Retrieval-Augmented Generation (mRAG) leverages cross-lingual evidence to ground Large Language Models (LLMs) in global knowledge. However, we show that current mRAG systems suffer from a language bias during reranking,…

计算与语言 · 计算机科学 2026-04-23 Dan Wang , Guozhao Mo , Yafei Shi , Cheng Zhang , Bo Zheng , Boxi Cao , Xuanang Chen , Yaojie Lu , Hongyu Lin , Ben He , Xianpei Han , Le Sun

Reasoning has emerged as the next major frontier for language models (LMs), with rapid advances from both academic and industrial labs. However, this progress often outpaces methodological rigor, with many evaluations relying on…

Combating online hate speech in multilingual settings requires approaches that go beyond English-centric models and capture the cultural and linguistic diversity of global online discourse. This paper presents a comprehensive survey and…

计算与语言 · 计算机科学 2026-03-23 Zahra Safdari Fesaghandis , Suman Kalyan Maity

Although state-of-the-art Speech Foundational Models can produce high-quality text pseudo-labels, applying Semi-Supervised Learning (SSL) for in-the-wild real-world data remains challenging due to its richer and more complex acoustics…

计算与语言 · 计算机科学 2026-03-16 Wen Ding , Fan Qian

With a surge in the usage of social media postings to express opinions, emotions, and ideologies, there has been a significant shift towards the calibration of social media as a rapid medium of conveying viewpoints and outlooks over the…

计算与语言 · 计算机科学 2023-09-26 Mohammad Kashif , Mohammad Zohair , Saquib Ali

In recent times, more and more people are posting about their mental states across various social media platforms. Leveraging this data, AI-based systems can be developed that help in assessing the mental health of individuals, such as…

人机交互 · 计算机科学 2024-12-20 Chayan Tank , Shaina Mehta , Sarthak Pol , Vinayak Katoch , Avinash Anand , Raj Jaiswal , Rajiv Ratn Shah

Language models (LMs) now excel at many tasks such as few-shot learning, question answering, reasoning, and dialog. However, they sometimes generate unsupported or misleading content. A user cannot easily determine whether their outputs are…

With the widespread online social networks, hate speeches are spreading faster and causing more damage than ever before. Existing hate speech detection methods have limitations in several aspects, such as handling data insufficiency,…

计算与语言 · 计算机科学 2024-09-27 Guanyi Mou , Kyumin Lee

The emergence of cross-modal foundation models has introduced numerous approaches grounded in text-image retrieval. However, on some domain-specific retrieval tasks, these models fail to focus on the key attributes required. To address this…

计算机视觉与模式识别 · 计算机科学 2023-06-13 Yuguang Yang , Yiming Wang , Shupeng Geng , Runqi Wang , Yimi Wang , Sheng Wu , Baochang Zhang

Offensive content is pervasive in social media and a reason for concern to companies and government organizations. Several studies have been recently published investigating methods to detect the various forms of such content (e.g. hate…

计算与语言 · 计算机科学 2020-10-13 Tharindu Ranasinghe , Marcos Zampieri

This paper describes our approach to the task of identifying offensive languages in a multilingual setting. We investigate two data augmentation strategies: using additional semi-supervised labels with different thresholds and cross-lingual…

计算与语言 · 计算机科学 2020-08-05 Hwijeen Ahn , Jimin Sun , Chan Young Park , Jungyun Seo

Hate speech is harmful content that directly attacks or promotes hatred against members of groups or individuals based on actual or perceived aspects of identity, such as racism, religion, or sexual orientation. This can affect social life…

计算与语言 · 计算机科学 2024-03-19 Arijit Das , Somashree Nandy , Rupam Saha , Srijan Das , Diganta Saha

This paper proposes an approach to cross-language sentence selection in a low-resource setting. It uses data augmentation and negative sampling techniques on noisy parallel sentence data to directly learn a cross-lingual embedding-based…

计算与语言 · 计算机科学 2021-06-07 Yanda Chen , Chris Kedzie , Suraj Nair , Petra Galuščáková , Rui Zhang , Douglas W. Oard , Kathleen McKeown

Masked language models (MLMs) are pre-trained with a denoising objective that is in a mismatch with the objective of downstream fine-tuning. We propose pragmatic masking and surrogate fine-tuning as two complementing strategies that exploit…

计算与语言 · 计算机科学 2022-06-02 Chiyu Zhang , Muhammad Abdul-Mageed

Pretrained multilingual models exhibit the same social bias as models processing English texts. This systematic review analyzes emerging research that extends bias evaluation and mitigation approaches into multilingual and non-English…

计算与语言 · 计算机科学 2025-09-08 Lance Calvin Lim Gamboa , Yue Feng , Mark Lee

Recent state-of-the-art language models utilize a two-phase training procedure comprised of (i) unsupervised pre-training on unlabeled text, and (ii) fine-tuning for a specific supervised task. More recently, many studies have been focused…

计算与语言 · 计算机科学 2019-11-15 Itzik Malkiel , Lior Wolf

This paper explores hate speech detection in Devanagari-scripted languages, focusing on Hindi and Nepali, for Subtask B of the CHIPSAL@COLING 2025 Shared Task. Using a range of transformer-based models such as XLM-RoBERTa, MURIL, and…

计算与语言 · 计算机科学 2024-12-13 Anmol Guragain , Nadika Poudel , Rajesh Piryani , Bishesh Khanal

Multilingual Large Language Models (LLMs) can process many languages, yet how they internally represent this diversity remains unclear. Do they form shared multilingual representations with language-specific decoding, and if so, why does…

计算与语言 · 计算机科学 2026-02-10 Abir Harrasse , Florent Draye , Punya Syon Pandey , Zhijing Jin , Bernhard Schölkopf

For multilingual sequence-to-sequence pretrained language models (multilingual Seq2Seq PLMs), e.g. mBART, the self-supervised pretraining task is trained on a wide range of monolingual languages, e.g. 25 languages from CommonCrawl, while…

计算与语言 · 计算机科学 2022-09-22 Changtong Zan , Liang Ding , Li Shen , Yu Cao , Weifeng Liu , Dacheng Tao