中文
相关论文

相关论文: CLASH: Evaluating Language Models on Judging High-…

200 篇论文

Recent large language models (LLMs) have shown indications of mathematical reasoning ability on challenging competition-level problems, especially with self-generated verbalizations of intermediate reasoning steps (i.e., chain-of-thought…

计算与语言 · 计算机科学 2024-06-11 Yujun Mao , Yoon Kim , Yilun Zhou

Large Language Models (LLMs) have recently achieved impressive performance in math and reasoning benchmarks. However, they often struggle with logic problems and puzzles that are relatively easy for humans. To further investigate this, we…

人工智能 · 计算机科学 2025-09-16 Nasim Borazjanizadeh , Roei Herzig , Trevor Darrell , Rogerio Feris , Leonid Karlinsky

We introduce a multi-turn benchmark for evaluating personalised alignment in LLM-based AI assistants, focusing on their ability to handle user-provided safety-critical contexts. Our assessment of ten leading models across five scenarios…

人机交互 · 计算机科学 2025-01-31 Lize Alberts , Benjamin Ellis , Andrei Lupu , Jakob Foerster

Large language models (LLMs) can lead to undesired consequences when misaligned with human values, especially in scenarios involving complex and sensitive social biases. Previous studies have revealed the misalignment of LLMs with human…

计算与语言 · 计算机科学 2025-09-18 Yang Liu , Chenhui Chu

Large language models are moving beyond transactional question answering to act as companions, coaches, mediators, and curators that scaffold human growth, decision-making, and well-being. This paper proposes a role-based framework for…

人机交互 · 计算机科学 2026-01-27 Zhiyin Zhou

The current evaluation of mathematical skills in LLMs is limited, as existing benchmarks are either relatively small, primarily focus on elementary and high-school problems, or lack diversity in topics. Additionally, the inclusion of visual…

Large language model (LLM) agents with extended autonomy unlock new capabilities, but also introduce heightened challenges for LLM safety. In particular, an LLM agent may pursue objectives that deviate from human values and ethical norms, a…

计算与语言 · 计算机科学 2026-01-27 Chen Chen , Kim Young Il , Yuan Yang , Wenhao Su , Yilin Zhang , Xueluan Gong , Qian Wang , Yongsen Zheng , Ziyao Liu , Kwok-Yan Lam

Cooperative autonomous driving requires traffic scene understanding from both vehicle and infrastructure perspectives. While vision-language models (VLMs) show strong general reasoning capabilities, their performance in safety-critical…

计算机视觉与模式识别 · 计算机科学 2026-04-13 Rui Gan , Junyi Ma , Pei Li , Xingyou Yang , Kai Chen , Sikai Chen , Bin Ran

We introduce Blink, a new benchmark for multimodal language models (LLMs) that focuses on core visual perception abilities not found in other evaluations. Most of the Blink tasks can be solved by humans "within a blink" (e.g., relative…

计算机视觉与模式识别 · 计算机科学 2024-07-04 Xingyu Fu , Yushi Hu , Bangzheng Li , Yu Feng , Haoyu Wang , Xudong Lin , Dan Roth , Noah A. Smith , Wei-Chiu Ma , Ranjay Krishna

In finance, Large Language Models (LLMs) face frequent knowledge conflicts arising from discrepancies between their pre-trained parametric knowledge and real-time market data. These conflicts are especially problematic in real-world…

投资组合管理 · 定量金融 2025-10-20 Hoyoung Lee , Junhyuk Seo , Suhwan Park , Junhyeong Lee , Wonbin Ahn , Chanyeol Choi , Alejandro Lopez-Lira , Yongjae Lee

Multi-modal Large Language Models (MLLMs) exhibit impressive problem-solving abilities in various domains, but their visual comprehension and abstract reasoning skills remain under-evaluated. To this end, we present PolyMATH, a challenging…

Large language models (LLMs) are increasingly used as epistemic partners in everyday reasoning, yet their errors remain predominantly analyzed through predictive metrics rather than through their interpretive effects on human judgment. This…

Although explainability and interpretability have received significant attention in artificial intelligence (AI) and natural language processing (NLP) for mental health, reasoning has not been examined in the same depth. Addressing this gap…

计算与语言 · 计算机科学 2025-11-10 Sneha Oram , Pushpak Bhattacharyya

The growing complexity of construction management (CM) projects, coupled with challenges such as strict regulatory requirements and labor shortages, requires specialized analytical tools that streamline project workflow and enhance…

计算与语言 · 计算机科学 2025-04-15 Ruoxin Xiong , Yanyu Wang , Suat Gunhan , Yimin Zhu , Charles Berryman

As large language models (LLMs) grow more capable, they face increasingly diverse and complex tasks, making reliable evaluation challenging. The paradigm of LLMs as judges has emerged as a scalable solution, yet prior work primarily focuses…

计算与语言 · 计算机科学 2025-11-03 Weiyuan Li , Xintao Wang , Siyu Yuan , Rui Xu , Jiangjie Chen , Qingqing Dong , Yanghua Xiao , Deqing Yang

Ensuring that Large Language Models (LLMs) align with the diverse and evolving human values across different regions and cultures remains a critical challenge in AI ethics. Current alignment approaches often yield superficial conformity…

人工智能 · 计算机科学 2025-11-04 Jiahao Wang , Songkai Xue , Jinghui Li , Xiaozhen Wang

The evaluation of Large Language Models (LLMs) increasingly relies on other LLMs acting as judges. However, current evaluation paradigms typically yield a single score or ranking, answering which model is better but not why. While essential…

计算与语言 · 计算机科学 2025-07-25 Asaf Yehudai , Lilach Eden , Yotam Perlitz , Roy Bar-Haim , Michal Shmueli-Scheuer

Large reasoning models with reasoning capabilities achieve state-of-the-art performance on complex tasks, but their robustness under multi-turn adversarial pressure remains underexplored. We evaluate nine frontier reasoning models under…

人工智能 · 计算机科学 2026-03-13 Yubo Li , Ramayya Krishnan , Rema Padman

Conversational recommender systems (CRS) have advanced with large language models, showing strong results in domains like movies. These domains typically involve fixed content and passive consumption, where user preferences can be matched…

信息检索 · 计算机科学 2026-02-26 Zheng Hui , Xiaokai Wei , Yexi Jiang , Kevin Gao , Chen Wang , Frank Ong , Se-eun Yoon , Rachit Pareek , Michelle Gong

We present PLUGH (https://www.urbandictionary.com/define.php?term=plugh), a modern benchmark that currently consists of 5 tasks, each with 125 input texts extracted from 48 different games and representing 61 different (non-isomorphic)…

计算与语言 · 计算机科学 2024-08-12 Alexey Tikhonov
‹ 上一页 1 8 9 10 下一页 ›