English
Related papers

Related papers: HORIZON: A Benchmark for In-the-wild User Behaviou…

200 papers

Current evaluation paradigms for emotional support conversations tend to reward generic empathetic responses, yet they fail to assess whether the support is genuinely personalized to users' unique psychological profiles and contextual…

Computation and Language · Computer Science 2026-01-06 Jing Ye , Lu Xiang , Yaping Zhang , Chengqing Zong

Large Language Models (LLMs) are increasingly used to simulate how specific users respond to a given context, enabling more user-centric applications that rely on user feedback. However, existing user simulators mostly imitate surface-level…

Computation and Language · Computer Science 2026-03-05 Shirley Wu , Evelyn Choi , Arpandeep Khatua , Zhanghan Wang , Joy He-Yueya , Tharindu Cyril Weerasooriya , Wei Wei , Diyi Yang , Jure Leskovec , James Zou

Reinforcement learning (RL), large language models (LLMs), and vision-language models (VLMs) have been widely studied in isolation. However, existing infrastructure lacks the ability to deploy agents from different decision-making paradigms…

To address the lack of comparative evaluation of Human-in-the-Loop Topic Modeling (HLTM) systems, we implement and evaluate three contrasting HLTM modeling approaches using simulation experiments. These approaches extend previously proposed…

Computation and Language · Computer Science 2019-10-07 Varun Kumar , Alison Smith-Renner , Leah Findlater , Kevin Seppi , Jordan Boyd-Graber

Existing human-robot interaction systems often lack mechanisms for sustained personalization and dynamic adaptation in multi-user environments, limiting their effectiveness in real-world deployments. We present HARMONI, a multimodal…

Autonomous agents powered by large language models (LLMs) are increasingly deployed in real-world applications requiring complex, long-horizon workflows. However, existing benchmarks predominantly focus on atomic tasks that are…

Computation and Language · Computer Science 2025-08-13 Weixuan Wang , Dongge Han , Daniel Madrigal Diaz , Jin Xu , Victor Rühle , Saravan Rajmohan

The rapid advancement of large language models (LLMs) has accelerated progress toward universal AI assistants. However, existing benchmarks for personalized assistants remain misaligned with real-world user-assistant interactions, failing…

Computation and Language · Computer Science 2026-03-13 Feiyu Duan , Xuanjing Huang , Zhongyu Wei

While agent evaluation has shifted toward long-horizon tasks, most benchmarks still emphasize local, step-level reasoning rather than the global constrained optimization (e.g., time and financial budgets) that demands genuine planning…

Artificial Intelligence · Computer Science 2026-01-27 Yinger Zhang , Shutong Jiang , Renhao Li , Jianhong Tu , Yang Su , Lianghao Deng , Xudong Guo , Chenxu Lv , Junyang Lin

Existing benchmarks for AI reasoning provide limited insight into how closely these capabilities resemble human reasoning in naturalistic contexts. We present an adaptation of the Watson & Holmes detective tabletop game as a new benchmark…

Artificial Intelligence · Computer Science 2026-02-24 Thatchawin Leelawat , Lewis D Griffin

Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse natural language processing tasks, yet they remain susceptible to hallucinations -- generating content that is factually incorrect, unfaithful to provided…

Computation and Language · Computer Science 2026-05-25 Ahmed Cherif

Online education platforms have experienced explosive growth over the past decade, generating massive volumes of user-generated content in the form of reviews, ratings, and behavioral logs. These heterogeneous signals provide unprecedented…

Graphics · Computer Science 2026-04-14 Arman Bekov , Azamat Nurgali

In a variety of online settings involving interaction with end-users it is critical for the systems to adapt to changes in user preferences. User preferences on items tend to change over time due to a variety of factors such as change in…

Information Retrieval · Computer Science 2019-05-17 Farzad Eskandanian , Bamshad Mobasher

Large Language Model (LLM)-based agents have achieved notable success on short-horizon and highly structured tasks. However, their ability to maintain coherent decision-making over long horizons in realistic and dynamic environments remains…

Artificial Intelligence · Computer Science 2026-03-18 Linghua Zhang , Jun Wang , Jingtong Wu , Zhisong Zhang

Real-time human perception is crucial for effective human-robot interaction (HRI). Large vision-language models (VLMs) offer promising generalizable perceptual capabilities but often suffer from high latency, which negatively impacts user…

Human-robot interaction (HRI) has long studied how agents and people coordinate to achieve shared goals. In this work, we formalize and benchmark the non-intrusive assistance as an independent paradigm of HRI, where a robot proactively…

Robotics · Computer Science 2026-05-05 Yuedi Zhang , Shuanghao Bai , Wanqi Zhou , Haoran Zhang , Qi Zhang , Zhirong Luan , Badong Chen

Robots need models of human behavior for both inferring human goals and preferences, and predicting what people will do. A common model is the Boltzmann noisily-rational decision model, which assumes people approximately optimize a reward…

Robotics · Computer Science 2020-01-14 Andreea Bobu , Dexter R. R. Scobee , Jaime F. Fisac , S. Shankar Sastry , Anca D. Dragan

Recommender systems continuously interact with users, creating feedback loops that shape both individual behavior and collective market dynamics. This paper introduces a simulation framework to model these loops in online retail…

Information Retrieval · Computer Science 2025-10-17 Gabriele Barlacchi , Margherita Lalli , Emanuele Ferragina , Fosca Giannotti , Luca Pappalardo

Measuring user satisfaction level is a challenging task, and a critical component in developing large-scale conversational agent systems serving the needs of real users. An widely used approach to tackle this is to collect human annotation…

Traditional approaches to training agents have generally involved a single, deterministic environment of minimal complexity to solve various tasks such as robot locomotion or computer vision. However, agents trained in static environments…

Robotics · Computer Science 2025-10-01 Kevin Godin-Dubois , Karine Miras , Anna V. Kononova

Vision-and-Language Navigation (VLN) has been studied mainly in either discrete or continuous settings, with little attention to dynamic, crowded environments. We present HA-VLN 2.0, a unified benchmark introducing explicit social-awareness…

Artificial Intelligence · Computer Science 2025-10-13 Yifei Dong , Fengyi Wu , Qi He , Zhi-Qi Cheng , Heng Li , Minghan Li , Zebang Cheng , Yuxuan Zhou , Jingdong Sun , Qi Dai , Alexander G Hauptmann