中文
相关论文

相关论文: AI Playing Business Games: Benchmarking Large Lang…

200 篇论文

Automatic evaluation is an integral aspect of dialogue system research. The traditional reference-based NLG metrics are generally found to be unsuitable for dialogue assessment. Consequently, recent studies have suggested various unique,…

计算与语言 · 计算机科学 2024-01-23 Chen Zhang , Luis Fernando D'Haro , Yiming Chen , Malu Zhang , Haizhou Li

Game-playing ability serves as an indicator for evaluating the strategic reasoning capability of large language models (LLMs). While most existing studies rely on utility performance metrics, which are not robust enough due to variations in…

人工智能 · 计算机科学 2025-08-19 Hongtao Liu , Zhicheng Du , Zihe Wang , Weiran Shen

As generative AI becomes increasingly embedded in everyday workflows, it is important to evaluate its performance in ways that reflect real-world usage rather than abstract notions of intelligence. Unlike many existing benchmarks that…

人工智能 · 计算机科学 2025-05-14 Justin K Miller , Wenjia Tang

The rapid advancements in large Language models (LLMs) have significantly enhanced their reasoning capabilities, driven by various strategies such as multi-agent collaboration. However, unlike the well-established performance improvements…

人工智能 · 计算机科学 2026-04-23 Zihan Chen , Song Wang , Zhen Tan , Xingbo Fu , Zhenyu Lei , Peng Wang , Huan Liu , Cong Shen , Jundong Li

Instruction-tuned Large Language Models (LLMs) are increasingly deployed as AI Assistants in firms for support in cognitive tasks. These AI assistants carry embedded perspectives which influence factors across the firm including…

计算机与社会 · 计算机科学 2025-05-27 Noah Broestl , Benjamin Lange , Cristina Voinea , Geoff Keeling , Rachael Lam

We aim to evaluate Large Language Models (LLMs) for embodied decision making. While a significant body of work has been leveraging LLMs for decision making in embodied environments, we still lack a systematic understanding of their…

Large Language Models (LLMs) have emerged as a transformative AI paradigm, profoundly influencing daily life through their exceptional language understanding and contextual generation capabilities. Despite their remarkable performance, LLMs…

The current paper presents the development and validation of SelfScore, a novel benchmark designed to assess the performance of automated Large Language Model (LLM) agents on help desk and professional consultation tasks. Given the…

计算机与社会 · 计算机科学 2024-10-23 John Mavi , Nathan Summers , Sergio Coronado

While Large Language Models (LLMs) can exhibit impressive proficiency in isolated, short-term tasks, they often fail to maintain coherent performance over longer time horizons. In this paper, we present Vending-Bench, a simulated…

人工智能 · 计算机科学 2025-02-25 Axel Backlund , Lukas Petersson

Large language models (LLMs) show increasingly advanced emergent capabilities and are being incorporated across various societal domains. Understanding their behavior and reasoning abilities therefore holds significant importance. We argue…

Building effective machine learning (ML) workflows to address complex tasks is a primary focus of the Automatic ML (AutoML) community and a critical step toward achieving artificial general intelligence (AGI). Recently, the integration of…

机器学习 · 计算机科学 2024-12-30 Yang Gu , Hengyu You , Jian Cao , Muran Yu , Haoran Fan , Shiyou Qian

In recent years, Large Language Models (LLMs) have emerged as a transformative development in artificial intelligence (AI), drawing significant attention from industry and academia. Trained on vast datasets, these sophisticated AI systems…

人工智能 · 计算机科学 2025-01-15 Oudom Hean , Utsha Saha , Binita Saha

To some, the advent of artificial intelligence (AI) promises better decision-making and increased military effectiveness while reducing the influence of human error and emotions. However, there is still debate about how AI systems,…

计算机与社会 · 计算机科学 2024-10-04 Max Lamparth , Anthony Corso , Jacob Ganz , Oriana Skylar Mastro , Jacquelyn Schneider , Harold Trinkunas

This paper investigates the ability of large language models (LLMs) to solve statistical tasks, as well as their capacity to assess the quality of reasoning. While state-of-the-art LLMs have demonstrated remarkable performance in a range of…

计算与语言 · 计算机科学 2026-01-22 Crish Nagarkar , Leonid Bogachev , Serge Sharoff

Large Language Models (LLMs) are highly proficient in language-based tasks. Their language capabilities have positioned them at the forefront of the future AGI (Artificial General Intelligence) race. However, on closer inspection, Valmeekam…

计算与语言 · 计算机科学 2025-03-17 Dibyanayan Bandyopadhyay , Soham Bhattacharjee , Asif Ekbal

In the rapidly evolving landscape of artificial intelligence (AI), generative large language models (LLMs) stand at the forefront, revolutionizing how we interact with our data. However, the computational intensity and memory consumption of…

机器学习 · 计算机科学 2025-07-24 Xupeng Miao , Gabriele Oliaro , Zhihao Zhang , Xinhao Cheng , Hongyi Jin , Tianqi Chen , Zhihao Jia

As large language models (LLMs) increasingly engage in complex social interactions, ensuring that their behaviors align with human ethical principles and intentions, known as value alignment, has become a critical scientific challenge.…

计算工程、金融与科学 · 计算机科学 2026-05-29 Yu Lei , Hao Liu , Chengxing Xie , Songjia Liu , Zhiyu Yin , Canyu Chen , Guohao Li , Philip Torr , Zhen Wu

This paper introduces the Word Synchronization Challenge, a novel benchmark to evaluate large language models (LLMs) in Human-Computer Interaction (HCI). This benchmark uses a dynamic game-like framework to test LLMs ability to mimic human…

人机交互 · 计算机科学 2026-01-15 Tanguy Cazalets , Joni Dambre

Sales dialogues require multi-turn, goal-directed persuasion under asymmetric incentives, which makes them a challenging setting for large language models (LLMs). Yet existing dialogue benchmarks rarely measure deal progression and…

计算与语言 · 计算机科学 2026-04-10 Xuanbo Su , Wenhao Hu , Haibo Su , Yunzhang Chen , Le Zhan , Yanqi Yang , Leo Huang

In task-oriented conversational AI evaluation, unsupervised methods poorly correlate with human judgments, and supervised approaches lack generalization. Recent advances in large language models (LLMs) show robust zeroshot and few-shot…

计算与语言 · 计算机科学 2024-06-26 Jinghan Jia , Abi Komma , Timothy Leffel , Xujun Peng , Ajay Nagesh , Tamer Soliman , Aram Galstyan , Anoop Kumar