English
Related papers

Related papers: Sim4IA-Bench: A User Simulation Benchmark Suite fo…

200 papers

LLM-based user simulation is the primary mechanism for end-to-end agent evaluation, yet simulated users are poor proxies for real humans: unconstrained LLM defaults produce a Formalism Ceiling (style match rates of 6-8% against real users),…

Human-Computer Interaction · Computer Science 2026-05-21 Ming Zhu , Juntao Tan , Rithesh Murthy , Jielin Qiu , Liangwei Yang , Wenting Zhao , Silvio Savarese , Shelby Heinecke , Huan Wang

User simulators can rapidly generate a large volume of timely user behavior data, providing a testing platform for reinforcement learning-based recommender systems, thus accelerating their iteration and optimization. However, prevalent user…

Information Retrieval · Computer Science 2024-12-24 Zijian Zhang , Shuchang Liu , Ziru Liu , Rui Zhong , Qingpeng Cai , Xiangyu Zhao , Chunxu Zhang , Qidong Liu , Peng Jiang

As generative AI models are increasingly used to simulate real-world systems, quantifying the ``sim-to-real'' gap is critical. For each input setting of interest -- which we call a \emph{scenario}, such as a survey question or operating…

Methodology · Statistics 2026-04-17 Garud Iyengar , Yu-Shiou Willy Lin , Kaizheng Wang

Conversational Information Retrieval (CIR) is an emerging field of Information Retrieval (IR) at the intersection of interactive IR and dialogue systems for open domain information needs. In order to optimize these interactions and enhance…

Information Retrieval · Computer Science 2022-01-11 Pierre Erbacher , Laure Soulier , Ludovic Denoyer

Benchmarks are paramount for gauging progress in the domain of Mobile GUI Agents. In practical scenarios, users frequently fail to articulate precise directives containing full task details at the onset, and their expressions are typically…

Simulation models often have parameters as input and return outputs to understand the behavior of complex systems. Calibration is the process of estimating the values of the parameters in a simulation model in light of observed data from…

Methodology · Statistics 2024-11-15 Özge Sürer

We introduce SpreadsheetBench, a challenging spreadsheet manipulation benchmark exclusively derived from real-world scenarios, designed to immerse current large language models (LLMs) in the actual workflow of spreadsheet users. Unlike…

Computation and Language · Computer Science 2024-10-18 Zeyao Ma , Bohan Zhang , Jing Zhang , Jifan Yu , Xiaokang Zhang , Xiaohan Zhang , Sijia Luo , Xi Wang , Jie Tang

Existing benchmarks in e-commerce primarily focus on basic user intents, such as finding or purchasing products. However, real-world users often pursue more complex goals, such as applying vouchers, managing budgets, and finding…

Computation and Language · Computer Science 2025-12-11 Jiangyuan Wang , Kejun Xiao , Qi Sun , Huaipeng Zhao , Tao Luo , Jian Dong Zhang , Xiaoyi Zeng

Simulation is a powerful tool to easily generate annotated data, and a highly desirable feature, especially in those domains where learning models need large training datasets. Machine learning and deep learning solutions, have proven to be…

Computer Vision and Pattern Recognition · Computer Science 2023-12-06 Niccolò Bisagno , Nicola Garau , Antonio Luigi Stefani , Nicola Conci

Due to the excellent capacities of large language models (LLMs), it becomes feasible to develop LLM-based agents for reliable user simulation. Considering the scarcity and limit (e.g., privacy issues) of real user data, in this paper, we…

Information Retrieval · Computer Science 2024-02-28 Ruiyang Ren , Peng Qiu , Yingqi Qu , Jing Liu , Wayne Xin Zhao , Hua Wu , Ji-Rong Wen , Haifeng Wang

Recent advancements in image generation models have enabled the prediction of future Graphical User Interface (GUI) states based on user instructions. However, existing benchmarks primarily focus on general domain visual fidelity, leaving…

The growing capabilities of large language models (LLMs) in instruction-following and context-understanding lead to the era of agents with numerous applications. Among these, task planning agents have become especially prominent in…

Artificial Intelligence for Science (AI4S) is an emerging research field that utilizes machine learning advancements to tackle complex scientific computational issues, aiming to enhance computational efficiency and accuracy. However, the…

Machine Learning · Computer Science 2023-11-30 Yatao Li , Jianfeng Zhan

Formal theorem-proving benchmarks enable mechanically verifiable evaluation of mathematical reasoning in large language models. However, existing benchmarks mainly focus on Olympiad-style problems and algebraic domains, leaving…

Artificial Intelligence · Computer Science 2026-05-19 Wentao Long , Yunfei Zhang , Chenyi Li , Li Zhou , Chumin Sun , Zaiwen Wen

Simulation-based testing has become a crucial complement to road testing for ensuring the safety of cyber physical systems (CPS). As a result, significant research efforts have been directed toward identifying failure scenarios within…

Artificial Intelligence · Computer Science 2025-11-14 Edward Kim , Devan Shanker , Varun Bharadwaj , Hongbeen Park , Jinkyu Kim , Hazem Torfah , Daniel J Fremont , Sanjit A Seshia

Network operators need to continuosly upgrade their infrastructures in order to keep their customer satisfaction levels high. Crowdsourcing-based approaches are generally adopted, where customers are directly asked to answer surveys about…

Networking and Internet Architecture · Computer Science 2020-11-02 Andrea Pimpinella , Marianna Repossi , Alessandro Enrico Cesare Redondi

Travel planning is a realistic task for evaluating the planning and tool-use abilities of LLM agents. However, existing benchmarks typically assume only a single user, thereby avoiding one of the most challenging aspects of real-world…

Computation and Language · Computer Science 2026-05-26 Xiang Cheng , Yulan Hu , Lulu Zheng , Zheng Pan , Xin Li , Yong Liu

Large language models (LLMs) are essential tools that users employ across various scenarios, so evaluating their performance and guiding users in selecting the suitable service is important. Although many benchmarks exist, they mainly focus…

Computation and Language · Computer Science 2024-09-23 Jiayin Wang , Fengran Mo , Weizhi Ma , Peijie Sun , Min Zhang , Jian-Yun Nie

Current Graphical User Interface (GUI) agents operate primarily under a reactive paradigm: a user must provide an explicit instruction for the agent to execute a task. However, an intelligent AI assistant should be proactive, which is…

Artificial Intelligence · Computer Science 2026-03-10 Yuxiang Chai , Shunye Tang , Han Xiao , Rui Liu , Hongsheng Li

With the growing demand for intelligent in-vehicle experiences, vehicle-based agents are evolving from simple assistants to long-term companions. This evolution requires agents to continuously model multi-user preferences and make reliable…

Artificial Intelligence · Computer Science 2026-03-26 Yuhao Chen , Yi Xu , Xinyun Ding , Xiang Fang , Shuochen Liu , Luxi Lin , Qingyu Zhang , Ya Li , Quan Liu , Tong Xu
‹ Prev 1 4 5 6 7 8 10 Next ›