中文
相关论文

相关论文: JuICE: A Benchmark for Evaluating LLM-Judge in Ide…

200 篇论文

Large language models (LLMs) have demonstrated substantial commonsense understanding through numerous benchmark evaluations. However, their understanding of cultural commonsense remains largely unexamined. In this paper, we conduct a…

计算与语言 · 计算机科学 2024-05-09 Siqi Shen , Lajanugen Logeswaran , Moontae Lee , Honglak Lee , Soujanya Poria , Rada Mihalcea

Although Large Language Models (LLMs) demonstrate strong capabilities across various tasks, they exhibit significant performance discrepancies across languages. While prompting LLMs in English typically yields the highest general…

计算与语言 · 计算机科学 2026-05-26 Andrew Ivan Soegeng , Patrick Sutanto , Tan Sang Nguyen

Social biases reflected in language are inherently shaped by cultural norms, which vary significantly across regions and lead to diverse manifestations of stereotypes. Existing evaluations of social bias in large language models (LLMs) for…

计算与语言 · 计算机科学 2026-03-26 Taihei Shiotani , Masahiro Kaneko , Ayana Niwa , Yuki Maruyama , Daisuke Oba , Masanari Ohi , Naoaki Okazaki

Language Models (LMs) often encounter knowledge conflicts when parametric memory contradicts contextual knowledge. Previous works attribute this conflict to the interplay between "memory heads" and "context heads", attention heads assumed…

计算与语言 · 计算机科学 2025-06-09 Gaotang Li , Yuzhong Chen , Hanghang Tong

This pilot study explores the localisation capabilities of state-of-the-art multilingual AI models when translating figurative language, such as idioms and puns, from English into a diverse range of global languages. It expands on existing…

计算与语言 · 计算机科学 2025-10-08 Madison Van Doren , Cory Holland

Large Language Models (LLMs) are rapidly being adopted by users across the globe, who interact with them in a diverse range of languages. At the same time, there are well-documented imbalances in the training data and optimisation…

人工智能 · 计算机科学 2025-11-07 Bram Bulté , Ayla Rigouts Terryn

The potential of using Large Language Models (LLMs) themselves to evaluate LLM outputs offers a promising method for assessing model performance across various contexts. Previous research indicates that LLM-as-a-judge exhibits a strong…

Large Language Models (LLMs) often exhibit cultural biases due to training data dominated by high-resource languages like English and Chinese. This poses challenges for accurately representing and evaluating diverse cultural contexts,…

计算与语言 · 计算机科学 2025-08-11 Zhong Ken Hew , Jia Xin Low , Sze Jue Yang , Chee Seng Chan

Recent advances in large language models (LLMs) have opened the door to culture-aware language tasks. We introduce the novel problem of adapting wine reviews across Chinese and English, which goes beyond literal translation by incorporating…

计算与语言 · 计算机科学 2025-09-17 Chenye Zou , Xingyue Wen , Tianyi Hu , Qian Janice Wang , Daniel Hershcovich

As large language models (LLMs) become key advisors in various domains, their cultural sensitivity and reasoning skills are crucial in multicultural environments. We introduce Nunchi-Bench, a benchmark designed to evaluate LLMs' cultural…

计算与语言 · 计算机科学 2025-07-08 Kyuhee Kim , Sangah Lee

As large language models (LLMs) are employed worldwide, existing evaluation paradigms for their multilingual capabilities primarily focus on factual task performance, neglecting the ability to judge content's deep-level values across…

计算与语言 · 计算机科学 2026-05-12 Yukun Chen , Xinyu Zhang , Boyi Deng , Jialong Tang , Yu Wan , Fei Huang , Yuxi Zhou , Baosong Yang , Yiming Li

Large Language Models (LLMs) are increasingly being used to autonomously evaluate the quality of content in communication systems, e.g., to assess responses in telecom customer support chatbots. However, the impartiality of these AI…

人工智能 · 计算机科学 2026-03-03 Jiaxin Gao , Chen Chen , Yanwen Jia , Xueluan Gong , Kwok-Yan Lam , Qian Wang

LLM-as-a-Judge has emerged as a popular evaluation strategy, where advanced large language models assess generation results in alignment with human instructions. While these models serve as a promising alternative to human annotators, their…

计算与语言 · 计算机科学 2025-05-20 Xiyan Fu , Wei Liu

Clinical communication skills are critical in medical education, and practicing and assessing clinical communication skills on a scale is challenging. Although LLM-powered clinical scenario simulations have shown promise in enhancing…

人工智能 · 计算机科学 2025-06-16 Weibing Zheng , Laurah Turner , Jess Kropczynski , Murat Ozer , Tri Nguyen , Shane Halse

We present a large-scale human evaluation benchmark for assessing cultural localisation in machine translation produced by state-of-the-art multilingual large language models (LLMs). Existing MT benchmarks emphasise token-level and…

计算与语言 · 计算机科学 2026-05-29 Madison Van Doren , Casey Ford , Jennifer Barajas , Riley VanMeter , Cory Holland

As LLMs are increasingly deployed in global applications, the importance of cultural sensitivity becomes paramount, ensuring that users from diverse backgrounds feel respected and understood. Cultural harm can arise when these models fail…

Cross-cultural competence in large language models (LLMs) requires the ability to identify Culture-Specific Items (CSIs) and to adapt them appropriately across cultural contexts. Progress in evaluating this capability has been constrained…

计算与语言 · 计算机科学 2026-01-21 Mohsinul Kabir , Tasnim Ahmed , Md Mezbaur Rahman , Shaoxiong Ji , Hassan Alhuzali , Sophia Ananiadou

Large language models (LLMs) are evolving fast and are now frequently used as evaluators, in a process typically referred to as LLM-as-a-Judge, which provides quality assessments of model outputs. However, recent research points out…

计算与语言 · 计算机科学 2026-01-27 Hugo Silva , Mateus Mendes , Hugo Gonçalo Oliveira

Large language models (LLMs) are increasingly deployed in multicultural settings; however, systematic evaluation of cultural specificity at the sentence level remains underexplored. We propose the Conceptual Cultural Index (CCI), which…

计算与语言 · 计算机科学 2026-02-11 Takumi Ohashi , Hitoshi Iyatomi

Evaluating Large Language Models (LLMs) in open-ended scenarios is challenging because existing benchmarks and metrics can not measure them comprehensively. To address this problem, we propose to fine-tune LLMs as scalable judges (JudgeLM)…

计算与语言 · 计算机科学 2025-03-04 Lianghui Zhu , Xinggang Wang , Xinlong Wang