中文
相关论文

相关论文: SEACrowd: A Multilingual Multimodal Data Hub and B…

200 篇论文

Although numerous datasets have been developed to support dialogue systems, most existing chit-chat datasets overlook the cultural nuances inherent in natural human conversations. To address this gap, we introduce SEADialogues, a culturally…

Safeguard models help large language models (LLMs) detect and block harmful content, but most evaluations remain English-centric and overlook linguistic and cultural diversity. Existing multilingual safety benchmarks often rely on…

计算与语言 · 计算机科学 2025-12-08 Panuthep Tasawong , Jian Gang Ngui , Alham Fikri Aji , Trevor Cohn , Peerat Limkonchotiwat

Southeast Asia (SEA) is a region of extraordinary linguistic and cultural diversity, yet it remains significantly underrepresented in vision-language (VL) research. This often results in artificial intelligence (AI) models that fail to…

计算机视觉与模式识别 · 计算机科学 2025-03-19 Samuel Cahyawijaya , Holy Lovenia , Joel Ruben Antony Moniz , Tack Hwa Wong , Mohammad Rifqi Farhansyah , Thant Thiri Maung , Frederikus Hudi , David Anugraha , Muhammad Ravi Shulthan Habibi , Muhammad Reza Qorib , Amit Agarwal , Joseph Marvin Imperial , Hitesh Laxmichand Patel , Vicky Feliren , Bahrul Ilmi Nasution , Manuel Antonio Rufino , Genta Indra Winata , Rian Adam Rajagede , Carlos Rafael Catalan , Mohamed Fazli Imam , Priyaranjan Pattnayak , Salsabila Zahirah Pranida , Kevin Pratama , Yeshil Bangera , Adisai Na-Thalang , Patricia Nicole Monderin , Yueqi Song , Christian Simon , Lynnette Hui Xian Ng , Richardy Lobo' Sapan , Taki Hasan Rafi , Bin Wang , Supryadi , Kanyakorn Veerakanjana , Piyalitt Ittichaiwong , Matthew Theodore Roque , Karissa Vincentio , Takdanai Kreangphet , Phakphum Artkaew , Kadek Hendrawan Palgunadi , Yanzhi Yu , Rochana Prih Hastuti , William Nixon , Mithil Bangera , Adrian Xuan Wei Lim , Aye Hninn Khine , Hanif Muhammad Zhafran , Teddy Ferdinan , Audra Aurora Izzani , Ayushman Singh , Evan , Jauza Akbar Krito , Michael Anugraha , Fenal Ashokbhai Ilasariya , Haochen Li , John Amadeo Daniswara , Filbert Aurelian Tjiaranata , Eryawan Presma Yulianrifat , Can Udomcharoenchaikit , Fadil Risdian Ansori , Mahardika Krisna Ihsani , Giang Nguyen , Anab Maulana Barik , Dan John Velasco , Rifo Ahmad Genadi , Saptarshi Saha , Chengwei Wei , Isaiah Flores , Kenneth Ko Han Chen , Anjela Gail Santos , Wan Shen Lim , Kaung Si Phyo , Tim Santos , Meisyarah Dwiastuti , Jiayun Luo , Jan Christian Blaise Cruz , Ming Shan Hee , Ikhlasul Akmal Hanif , M. Alif Al Hakim , Muhammad Rizky Sya'ban , Kun Kerdthaisong , Lester James V. Miranda , Fajri Koto , Tirana Noor Fatyanosa , Alham Fikri Aji , Jostin Jerico Rosal , Jun Kevin , Robert Wijaya , Onno P. Kampman , Ruochen Zhang , Börje F. Karlsson , Peerat Limkonchotiwat

With the rapid emergence of novel capabilities in Large Language Models (LLMs), the need for rigorous multilingual and multicultural benchmarks that are integrated has become more pronounced. Though existing LLM benchmarks are capable of…

Multilingual text embeddings are often assumed to encode meaning in a perspective-independent semantic space, yielding stable similarity judgments across tasks and languages. Our results show that this assumption does not hold in practice.…

Culturally aware safeguards are crucial for AI alignment in real-world settings, where safety extends beyond common sense and encompasses diverse local values, norms, and region-specific regulations. However, building large-scale,…

计算与语言 · 计算机科学 2026-02-03 Panuthep Tasawong , Jian Gang Ngui , Alham Fikri Aji , Trevor Cohn , Peerat Limkonchotiwat

The rapid growth of the digital economy in South-East Asia (SEA) has amplified the risks of audio deepfakes, yet current datasets cover SEA languages only sparsely, leaving models poorly equipped to handle this critical region. This…

声音 · 计算机科学 2025-09-26 Jinyang Wu , Nana Hou , Zihan Pan , Qiquan Zhang , Sailor Hardik Bhupendra , Soumik Mondal

Recently, Large Language Models (LLMs) have dominated much of the artificial intelligence scene with their ability to process and generate natural languages. However, the majority of LLM research and development remains English-centric,…

Multilingual document and scene text understanding plays an important role in applications such as search, finance, and public services. However, most existing benchmarks focus on high-resource languages and fail to evaluate models in…

计算与语言 · 计算机科学 2026-03-17 Pengfei Yue , Xingran Zhao , Juntao Chen , Peng Hou , Wang Longchao , Jianghang Lin , Shengchuan Zhang , Anxiang Zeng , Liujuan Cao

In this paper, we introduce SailCompass, a reproducible and robust evaluation benchmark for assessing Large Language Models (LLMs) on Southeast Asian Languages (SEA). SailCompass encompasses three main SEA languages, eight primary tasks…

计算与语言 · 计算机科学 2024-12-03 Jia Guo , Longxu Dou , Guangtao Zeng , Stanley Kok , Wei Lu , Qian Liu

Large Language Models (LLMs) have shown remarkable abilities across various tasks, yet their development has predominantly centered on high-resource languages like English and Chinese, leaving low-resource languages underserved. To address…

Reliable automatic evaluation of summarization systems is challenging due to the multifaceted and subjective nature of the task. This is especially the case for languages other than English, where human evaluations are scarce. In this work,…

A major obstacle to the advancements of machine learning models in marine science, particularly in sonar imagery analysis, is the scarcity of AI-ready datasets. While there have been efforts to make AI-ready sonar image dataset publicly…

计算机视觉与模式识别 · 计算机科学 2024-11-08 Kien X. Nguyen , Fengchun Qiao , Arthur Trembanis , Xi Peng

We introduce SeaLLMs-Audio, the first large audio-language model (LALM) tailored for multiple Southeast Asian (SEA) languages-Indonesian (id), Thai (th), and Vietnamese (vi)-alongside English (en) and Chinese (zh). Trained on a large-scale…

计算与语言 · 计算机科学 2025-11-04 Chaoqun Liu , Mahani Aljunied , Guizhen Chen , Hou Pong Chan , Weiwen Xu , Yu Rong , Wenxuan Zhang

Despite the remarkable achievements of large language models (LLMs) in various tasks, there remains a linguistic bias that favors high-resource languages, such as English, often at the expense of low-resource and regional languages. To…

This study introduces two novel benchmarks, SeaExam and SeaBench, designed to evaluate the capabilities of Large Language Models (LLMs) in Southeast Asian (SEA) application scenarios. Unlike existing multilingual datasets primarily derived…

计算与语言 · 计算机科学 2025-02-11 Chaoqun Liu , Wenxuan Zhang , Jiahao Ying , Mahani Aljunied , Anh Tuan Luu , Lidong Bing

While generative multilingual models are rapidly being deployed, their safety and fairness evaluations are largely limited to resources collected in English. This is especially problematic for evaluations targeting inherently socio-cultural…

计算与语言 · 计算机科学 2024-03-12 Mukul Bhutani , Kevin Robinson , Vinodkumar Prabhakaran , Shachi Dave , Sunipa Dev

Artificial intelligence (AI) has the potential to transform healthcare, education, governance and socioeconomic equity, but its benefits remain concentrated in a small number of languages (Bender, 2019; Blasi et al., 2022; Joshi et al.,…

Recent breakthroughs in large language models (LLMs) have centered around a handful of data-rich languages. What does it take to broaden access to breakthroughs beyond first-class citizen languages? Our work introduces Aya, a massively…

‹ 上一页 1 2 3 10 下一页 ›