中文
相关论文

相关论文: CoPA: Benchmarking Personalized Question Answering…

200 篇论文

Reliable uncertainty quantification (UQ) is essential when employing large language models (LLMs) in high-risk domains such as clinical question answering (QA). In this work, we evaluate uncertainty estimation methods for clinical QA…

计算与语言 · 计算机科学 2026-01-27 Alberto Testoni , Iacer Calixto

The rapid progress of large language models (LLMs) has seen them excel and frequently surpass human performance on standard benchmarks. This has enabled many downstream applications, such as LLM agents, to rely on their reasoning to address…

计算与语言 · 计算机科学 2025-02-17 Harsh Kohli , Sachin Kumar , Huan Sun

Current large language model (LLM) development treats task-solving and preference-alignment as separate challenges, optimizing first for objective correctness, then for alignment to aggregated human preferences. This paradigm fails in…

计算与语言 · 计算机科学 2026-03-06 Shuyue Stella Li , Avinandan Bose , Faeze Brahman , Simon Shaolei Du , Pang Wei Koh , Maryam Fazel , Yulia Tsvetkov

Large Language Model (LLM) has gained popularity and achieved remarkable results in open-domain tasks, but its performance in real industrial domain-specific scenarios is average due to its lack of specific domain knowledge. This issue has…

计算与语言 · 计算机科学 2023-10-17 Fangkai Yang , Pu Zhao , Zezhong Wang , Lu Wang , Jue Zhang , Mohit Garg , Qingwei Lin , Saravan Rajmohan , Dongmei Zhang

Advancement in Large Language Models (LLMs) reasoning capabilities enables them to solve scientific problems with enhanced efficacy. Thereby, a high-quality benchmark for comprehensive and appropriate assessment holds significance, while…

This study validates Large Language Models (LLMs) as a dynamic alternative to questionnaire-based personality assessment. Using a within-subjects experiment (N=33), we compared Big Five personality scores derived from guided LLM…

计算与语言 · 计算机科学 2026-02-19 Andrius Matšenas , Anet Lello , Tõnis Lees , Hans Peep , Kim Lilii Tamm

Large Language Models (LLMs) have emerged as personalized assistants for users across a wide range of tasks -- from offering writing support to delivering tailored recommendations or consultations. Over time, the interaction history between…

计算与语言 · 计算机科学 2025-10-28 Bowen Jiang , Zhuoqun Hao , Young-Min Cho , Bryan Li , Yuan Yuan , Sihao Chen , Lyle Ungar , Camillo J. Taylor , Dan Roth

Deep Research Agents (DRAs) can autonomously conduct complex investigations and generate comprehensive reports, demonstrating strong real-world potential. However, existing evaluations mostly rely on close-ended benchmarks, while open-ended…

Large Language Models (LLMs) show promise as data analysis agents, but existing benchmarks overlook the iterative nature of the field, where experts' decisions evolve with deeper insights of the dataset. To address this, we introduce…

计算与语言 · 计算机科学 2025-06-09 Hanyu Li , Haoyu Liu , Tingyu Zhu , Tianyu Guo , Zeyu Zheng , Xiaotie Deng , Michael I. Jordan

This paper highlights the importance of personalization in large language models and introduces the LaMP benchmark -- a novel benchmark for training and evaluating language models for producing personalized outputs. LaMP offers a…

计算与语言 · 计算机科学 2024-06-06 Alireza Salemi , Sheshera Mysore , Michael Bendersky , Hamed Zamani

Large language models (LLMs) may not equitably represent diverse global perspectives on societal issues. In this paper, we develop a quantitative framework to evaluate whose opinions model-generated responses are more similar to. We first…

While LLM agents have demonstrated remarkable task-oriented abilities such as planning, reasoning, and action, few works have treated them as complete human personalities where emotional dimensions hold equal importance. In this paper, we…

计算与语言 · 计算机科学 2026-05-29 Weihan Peng , Chenxu Zhang , Qianao Wang , Yuling Shi , Heng Lian , Qihong Mao , Jiahao Pang , Chunliang Feng , Bowen Li , Xiaodong Gu

Large Language Models (LLMs) hold promise in addressing complex medical problems. However, while most prior studies focus on improving accuracy and reasoning abilities, a significant bottleneck in developing effective healthcare agents lies…

计算与语言 · 计算机科学 2025-10-06 Weikang Qiu , Tinglin Huang , Ryan Rullo , Yucheng Kuang , Ali Maatouk , S. Raquel Ramos , Rex Ying

As the utilization of language models in interdisciplinary, human-centered studies grow, expectations of their capabilities continue to evolve. Beyond excelling at conventional tasks, models are now expected to perform well on user-centric…

计算与语言 · 计算机科学 2025-09-25 Yuxiang Zhou , Hainiu Xu , Desmond C. Ong , Maria Liakata , Petr Slovak , Yulan He

Personalized agents that interact with users over long periods must maintain persistent memory across sessions and update it as circumstances change. However, existing benchmarks predominantly frame long-term memory evaluation as fact…

计算与语言 · 计算机科学 2026-04-23 Md Nayem Uddin , Kumar Shubham , Eduardo Blanco , Chitta Baral , Gengyu Wang

AI-powered recruitment tools are increasingly adopted in personnel selection, yet they struggle to capture the requisition (req)-specific personal competencies (PCs) that distinguish successful candidates beyond job categories. We propose a…

LLM-based agent judges are an emerging approach to evaluating conversational AI, yet a fundamental uncertainty remains: can we trust their assessments, and if so, how many are needed? Through 960 sessions with two model pairs across 15…

人工智能 · 计算机科学 2026-04-02 HyunJoon Jung , William Na

With the rise in capabilities of large language models (LLMs) and their deployment in real-world tasks, evaluating LLM alignment with human preferences has become an important challenge. Current benchmarks average preferences across all…

人工智能 · 计算机科学 2026-04-22 Cristina Garbacea , Heran Wang , Chenhao Tan

Recently, Product Question Answering (PQA) on E-Commerce platforms has attracted increasing attention as it can act as an intelligent online shopping assistant and improve the customer shopping experience. Its key function, automatic answer…

计算与语言 · 计算机科学 2021-12-28 Yang Deng , Yaliang Li , Wenxuan Zhang , Bolin Ding , Wai Lam

Pretrained language models (PLMs) have achieved superhuman performance on many benchmarks, creating a need for harder tasks. We introduce CoDA21 (Context Definition Alignment), a challenging benchmark that measures natural language…

计算与语言 · 计算机科学 2022-03-15 Lütfi Kerem Senel , Timo Schick , Hinrich Schütze