English
Related papers

Related papers: JuICE: A Benchmark for Evaluating LLM-Judge in Ide…

200 papers

There is a significant gap in evaluating cultural reasoning in LLMs using conversational datasets that capture culturally rich and dialectal contexts. Most Arabic benchmarks focus on short text snippets in Modern Standard Arabic (MSA),…

Large language models (LLMs), despite their impressive performance in various language tasks, are typically limited to processing texts within context-window size. This limitation has spurred significant research efforts to enhance LLMs'…

Computation and Language · Computer Science 2024-09-09 Jiaqi Li , Mengmeng Wang , Zilong Zheng , Muhan Zhang

With the recent appearance of LLMs in practical settings, having methods that can effectively detect factual inconsistencies is crucial to reduce the propagation of misinformation and improve trust in model outputs. When testing on existing…

Computation and Language · Computer Science 2023-05-25 Philippe Laban , Wojciech Kryściński , Divyansh Agarwal , Alexander R. Fabbri , Caiming Xiong , Shafiq Joty , Chien-Sheng Wu

Recent advancements in instruction fine-tuning, alignment methods such as reinforcement learning from human feedback (RLHF), and optimization techniques like direct preference optimization (DPO) have significantly enhanced the adaptability…

Computation and Language · Computer Science 2025-03-04 Samar M. Magdy , Sang Yun Kwon , Fakhraddin Alwajih , Safaa Abdelfadil , Shady Shehata , Muhammad Abdul-Mageed

Large language models (LLMs) are increasingly deployed in culturally sensitive real-world tasks. However, existing cultural alignment approaches fail to align LLMs' broad cultural values with the specific goals of downstream tasks and…

Computation and Language · Computer Science 2026-02-27 Binchi Zhang , Xujiang Zhao , Jundong Li , Haifeng Chen , Zhengzhang Chen

Large language models (LLMs) have performed remarkably well in various natural language processing tasks by benchmarking, including in the Western medical domain. However, the professional evaluation benchmarks for LLMs have yet to be…

Computation and Language · Computer Science 2024-06-04 Wenjing Yue , Xiaoling Wang , Wei Zhu , Ming Guan , Huanran Zheng , Pengfei Wang , Changzhi Sun , Xin Ma

How much large language models (LLMs) can aid scientific discovery, notably in assisting academic peer review, is in heated debate. Between a literature digest and a human-comparable research assistant lies their practical application…

Computation and Language · Computer Science 2025-08-19 Tianyi Li , Yu Qin , Olivia R. Liu Sheng

A large number of studies rely on closed-style multiple-choice surveys to evaluate cultural alignment in Large Language Models (LLMs). In this work, we challenge this constrained evaluation paradigm and explore more realistic, unconstrained…

Computation and Language · Computer Science 2025-11-14 Mohsinul Kabir , Ajwad Abrar , Sophia Ananiadou

The rapid rise in popularity of Large Language Models (LLMs) with emerging capabilities has spurred public curiosity to evaluate and compare different LLMs, leading many researchers to propose their own LLM benchmarks. Noticing preliminary…

Artificial Intelligence · Computer Science 2025-05-15 Timothy R. McIntosh , Teo Susnjak , Nalin Arachchilage , Tong Liu , Paul Watters , Malka N. Halgamuge

Rigorous evaluation of large language models (LLMs) relies on comparing models by the prevalence of desirable or undesirable behaviors, such as task pass rates or policy violations. These prevalence estimates are produced by a classifier,…

Machine Learning · Computer Science 2026-03-25 Stephane Collot , Colin Fraser , Justin Zhao , William F. Shen , Timon Willi , Ilias Leontiadis

As large language models (LLMs) become increasingly embedded in our daily lives, evaluating their quality and reliability across diverse contexts has become essential. While comprehensive benchmarks exist for assessing LLM performance in…

Sentiment analysis in low-resource, culturally nuanced contexts challenges conventional NLP approaches that assume fixed labels and universal affective expressions. We present a diagnostic framework that treats sentiment as a…

Computation and Language · Computer Science 2025-08-07 Millicent Ochieng , Anja Thieme , Ignatius Ezeani , Risa Ueno , Samuel Maina , Keshet Ronen , Javier Gonzalez , Jacki O'Neill

We present DocPuzzle, a rigorously constructed benchmark for evaluating long-context reasoning capabilities in large language models (LLMs). This benchmark comprises 100 expert-level QA problems requiring multi-step reasoning over long…

Artificial Intelligence · Computer Science 2025-02-26 Tianyi Zhuang , Chuqiao Kuang , Xiaoguang Li , Yihua Teng , Jihao Wu , Yasheng Wang , Lifeng Shang

Generative speech technologies are progressing rapidly, but evaluating the perceptual quality of synthetic speech remains a core challenge. Existing methods typically rely on scalar scores or binary decisions, which lack interpretability…

Mathematical reasoning remains one of the most challenging domains for large language models (LLMs), requiring not only linguistic understanding but also structured logical deduction and numerical precision. While recent LLMs demonstrate…

Computation and Language · Computer Science 2026-01-27 Mahbub E Sobhani , Md. Faiyaz Abdullah Sayeedi , Tasnim Mohiuddin , Md Mofijul Islam , Swakkhar Shatabda

NLP research has increasingly focused on subjective tasks such as emotion analysis. However, existing emotion benchmarks suffer from two major shortcomings: (1) they largely rely on keyword-based emotion recognition, overlooking crucial…

Computation and Language · Computer Science 2025-05-29 Tadesse Destaw Belay , Ahmed Haj Ahmed , Alvin Grissom , Iqra Ameer , Grigori Sidorov , Olga Kolesnikova , Seid Muhie Yimam

Although the cultural (mis)alignment of Large Language Models (LLMs) has attracted increasing attention -- often framed in terms of cultural bias -- until recently there has been limited work on the design and development of datasets for…

Large language models (LLMs) are widely used in various tasks and applications. However, despite their wide capabilities, they are shown to lack cultural alignment \citep{ryan-etal-2024-unintended, alkhamissi-etal-2024-investigating} and…

Computation and Language · Computer Science 2025-12-17 Pramit Sahoo , Maharaj Brahma , Maunendra Sankar Desarkar

Large Language Models (LLMs) have excelled at language understanding and generating human-level text. However, even with supervised training and human alignment, these LLMs are susceptible to adversarial attacks where malicious users can…

Computation and Language · Computer Science 2024-08-08 Shachi H Kumar , Saurav Sahay , Sahisnu Mazumder , Eda Okur , Ramesh Manuvinakurike , Nicole Beckage , Hsuan Su , Hung-yi Lee , Lama Nachman

LLM-as-a-Judge has been widely adopted across various research and practical applications, yet the robustness and reliability of its evaluation remain a critical issue. A core challenge it faces is bias, which has primarily been studied in…

Computation and Language · Computer Science 2026-02-11 Peng Lai , Zhihao Ou , Yong Wang , Longyue Wang , Jian Yang , Yun Chen , Guanhua Chen