English
Related papers

Related papers: HORIZON: A Benchmark for In-the-wild User Behaviou…

200 papers

LLM-based agents score well on search benchmarks, yet real users consistently find results unsatisfying, revealing a persistent evaluation-experience gap. We attribute this gap to existing benchmarks' reliance on over-specified queries,…

Computation and Language · Computer Science 2026-05-28 Xiaohongshu Inc

Despite growing interest in active inference for robotic control, its application to complex, long-horizon tasks remains untested. We address this gap by introducing a fully hierarchical active inference architecture for goal-directed…

Robotics · Computer Science 2025-07-24 Corrado Pezzato , Ozan Çatal , Toon Van de Maele , Riddhi J. Pitliya , Tim Verbelen

In the evolving field of robotics, the challenge of Object Navigation (ON) in household environments has attracted significant interest. Existing ON benchmarks typically place objects in locations guided by general scene priors, without…

Robotics · Computer Science 2026-02-09 Hongcheng Wang , Jinyu Zhu , Hao Dong

Predicting future human behavior is an increasingly popular topic in computer vision, driven by the interest in applications such as autonomous vehicles, digital assistants and human-robot interactions. The literature on behavior prediction…

Computer Vision and Pattern Recognition · Computer Science 2024-10-21 Bolin Lai , Sam Toyer , Tushar Nagarajan , Rohit Girdhar , Shengxin Zha , James M. Rehg , Kris Kitani , Kristen Grauman , Ruta Desai , Miao Liu

From loco-motion to dextrous manipulation, humanoid robots have made remarkable strides in demonstrating complex full-body capabilities. However, the majority of current robot learning datasets and benchmarks mainly focus on stationary…

Human-Object Interaction (HOI) detection has seen substantial advances in recent years. However, existing works focus on the standard setting with ideal images and natural distribution, far from practical scenarios with inevitable…

Computer Vision and Pattern Recognition · Computer Science 2025-06-24 Chi Xie , Shuang Liang , Jie Li , Feng Zhu , Rui Zhao , Yichen Wei , Shengjie Zhao

Environments built for people are increasingly operated by a new class of economic actors: LLM-powered software agents making decisions on our behalf. These decisions range from our purchases to travel plans to medical treatment selection.…

Artificial Intelligence · Computer Science 2026-02-25 Manuel Cherep , Chengtian Ma , Abigail Xu , Maya Shaked , Pattie Maes , Nikhil Singh

Recent advances in large language models have highlighted their potential for personalized recommendation, where accurately capturing user preferences remains a key challenge. Leveraging their strong reasoning and generalization…

Navigating social robots in dense, dynamic crowds is challenging due to environmental uncertainty and complex human-robot interactions. While Model Predictive Control (MPC) offers strong real-time performance, its reliance on a fixed…

Robotics · Computer Science 2026-03-03 Jiamin Shi , Haolin Zhang , Yuchen Yan , Shitao Chen , Jingmin Xin , Nanning Zheng

Long-horizon household tasks demand robust high-level planning and sustained reasoning capabilities, which are largely overlooked by existing embodied AI benchmarks that emphasize short-horizon navigation or manipulation and rely on fixed…

Artificial Intelligence · Computer Science 2026-05-19 Zilin Zhu , Longteng Guo , Yanghong Mei , Bowen Pang , Zongxun Zhang , Xingjian He , Ruyi Ji , Jing Liu

Recent advances in high-fidelity virtual environments serve as one of the major driving forces for building intelligent embodied agents to perceive, reason and interact with the physical world. Typically, these environments remain unchanged…

Computer Vision and Pattern Recognition · Computer Science 2024-01-24 Qinhong Zhou , Sunli Chen , Yisong Wang , Haozhe Xu , Weihua Du , Hongxin Zhang , Yilun Du , Joshua B. Tenenbaum , Chuang Gan

This paper introduces a novel dataset REGEN (Reviews Enhanced with GEnerative Narratives), designed to benchmark the conversational capabilities of recommender Large Language Models (LLMs), addressing the limitations of existing datasets…

The pursuit of general-purpose embodied agents is hindered by fragmented evaluation protocols that isolate navigation skills and fixate on specific robot morphologies, failing to reflect real-world scenarios where agents must orchestrate…

Evaluating language models and AI agents remains fundamentally challenging because static benchmarks fail to capture real-world uncertainty, distribution shift, and the gap between isolated task accuracy and human-aligned decision-making…

Artificial Intelligence · Computer Science 2026-01-27 Shirin Shahabi , Spencer Graham , Haruna Isah

Sources of complementary information are connected when we link user accounts belonging to the same user across different platforms or devices. The expanded information promotes the development of a wide range of applications, such as…

Social and Information Networks · Computer Science 2022-01-11 Wei Chen , Weiqing Wang , Hongzhi Yin , Lei Zhao , Xiaofang Zhou

Understanding how humans evaluate robot behavior during human-robot interactions is crucial for developing socially aware robots that behave according to human expectations. While the traditional approach to capturing these evaluations is…

Robotics · Computer Science 2025-12-19 Qiping Zhang , Nathan Tsoi , Mofeed Nagib , Hao-Tien Lewis Chiang , Marynel Vázquez

The rapid evolution of the web has led to an exponential growth in content. Recommender systems play a crucial role in Human-Computer Interaction (HCI) by tailoring content based on individual preferences. Despite their importance,…

Information Retrieval · Computer Science 2023-10-18 Yubo Shu , Haonan Zhang , Hansu Gu , Peng Zhang , Tun Lu , Dongsheng Li , Ning Gu

Foundation models (FMs) are large neural networks trained on broad datasets, excelling in downstream tasks with minimal fine-tuning. Human activity recognition in video has advanced with FMs, driven by competition among different…

Computer Vision and Pattern Recognition · Computer Science 2024-07-23 Thinesh Thiyakesan Ponbagavathi , Kunyu Peng , Alina Roitberg

The evaluation of large language models faces significant challenges. Technical benchmarks often lack real-world relevance, while existing human preference evaluations suffer from unrepresentative sampling, superficial assessment depth, and…

Computation and Language · Computer Science 2026-03-06 Nora Petrova , Andrew Gordon , Enzo Blindow

Large language models (LLMs) are increasingly applied to socially grounded tasks, such as online community moderation, media content analysis, and social reasoning games. Success in these contexts depends on a model's social reasoning…

‹ Prev 1 3 4 5 6 7 10 Next ›