English
Related papers

Related papers: Evaluating Multimodal Interactive Agents

200 papers

Recent advances in large language models (LLMs) have enabled the development of AI agents that exhibit increasingly human-like behaviors, including planning, adaptation, and social dynamics across diverse, interactive, and open-ended…

Neurons and Cognition · Quantitative Biology 2025-06-13 Lin Chen , Yunke Zhang , Jie Feng , Haoye Chai , Honglin Zhang , Bingbing Fan , Yibo Ma , Shiyuan Zhang , Nian Li , Tianhui Liu , Nicholas Sukiennik , Keyu Zhao , Yu Li , Ziyi Liu , Fengli Xu , Yong Li

Multi-Agent Systems (MAS) built using AI agents fulfill a variety of user intents that may be used to design and build a family of related applications. However, the creation of such MAS currently involves manual composition of the plan,…

Artificial Intelligence · Computer Science 2026-05-06 Kishan Athrey , Ramin Pishehvar , Brian Riordan , Mahesh Viswanathan

The paper describes a flexible and modular platform to create multimodal interactive agents. The platform operates through an event-bus on which signals and interpretations are posted in a sequence in time. Different sensors and…

Artificial Intelligence · Computer Science 2022-06-02 Thomas Baier , Selene Baez Santamaria , Piek Vossen

Social identity theory (SIT) and social categorization theory (SCT) are two facets of the social identity approach (SIA) to understanding social phenomena. SIT and SCT are models that describe and explain how people interact with one…

Physics and Society · Physics 2025-08-26 Katie Seaborn

The emergence of agentic recommender systems powered by Large Language Models (LLMs) represents a paradigm shift in personalized recommendations, leveraging LLMs' advanced reasoning and role-playing capabilities to enable autonomous,…

Information Retrieval · Computer Science 2025-05-29 Yu Shang , Peijie Liu , Yuwei Yan , Zijing Wu , Leheng Sheng , Yuanqing Yu , Chumeng Jiang , An Zhang , Fengli Xu , Yu Wang , Min Zhang , Yong Li

AI agents hold growing promise for accelerating scientific discovery; yet, a lack of frontier evaluations hinders adoption into real workflows. Expert-written benchmarks have proven effective at measuring AI reasoning, but most at this…

Social simulation provides a compelling testbed for studying social intelligence, where agents interact through multi-turn dialogues under evolving contexts and strategically adapting opponents. Such environments are inherently…

Artificial Intelligence · Computer Science 2026-05-20 Xiang Li , Liping Yi , Mingze Kong , Min Zhang , Zhongxiang Dai , QingHua Hu

Algorithm audits have increased in recent years due to a growing need to independently assess the performance of automatically curated services that process, filter, and rank the large and dynamic amount of information available on the…

Computers and Society · Computer Science 2023-04-06 Roberto Ulloa , Mykola Makhortykh , Aleksandra Urman

In this paper, we present a new methodology that employs tester agents to automate video game testing. We introduce two types of agents -synthetic and human-like- and two distinct approaches to create them. Our agents are derived from…

Artificial Intelligence · Computer Science 2019-06-04 Sinan Ariyurek , Aysu Betin-Can , Elif Surer

Recent advances in large language models have sparked growing interest in AI agents capable of solving complex, real-world tasks. However, most existing agent systems rely on manually crafted configurations that remain static after…

Artificial Intelligence · Computer Science 2025-09-03 Jinyuan Fang , Yanwen Peng , Xi Zhang , Yingxu Wang , Xinhao Yi , Guibin Zhang , Yi Xu , Bin Wu , Siwei Liu , Zihao Li , Zhaochun Ren , Nikos Aletras , Xi Wang , Han Zhou , Zaiqiao Meng

Recent advances in large language models (LLMs) have sparked growing interest in building fully autonomous agents. However, fully autonomous LLM-based agents still face significant challenges, including limited reliability due to…

We compare three methods of familiarizing a human with an artificial intelligence (AI) teammate ("agent") prior to operation in a collaborative, fast-paced intelligence, surveillance, and reconnaissance (ISR) environment. In a…

Artificial Intelligence · Computer Science 2025-05-21 Ryan Bowers , Richard Agbeyibor , Jack Kolb , Karen Feigh

While rapid advances in large language models (LLMs) are reshaping data-driven intelligent education, accurately simulating students remains an important but challenging bottleneck for scalable educational data collection, evaluation, and…

Computers and Society · Computer Science 2025-12-05 Haoxuan Li , Jifan Yu , Xin Cong , Yang Dang , Daniel Zhang-li , Lu Mi , Yisi Zhan , Huiqin Liu , Zhiyuan Liu

As NLP evaluation shifts from static benchmarks to multi-turn interactive settings, LLM-based simulators have become widely used as user proxies, serving two roles: generating user turns and providing evaluation signals. Yet, these…

Artificial Intelligence · Computer Science 2026-03-13 Xuhui Zhou , Weiwei Sun , Qianou Ma , Yiqing Xie , Jiarui Liu , Weihua Du , Sean Welleck , Yiming Yang , Graham Neubig , Sherry Tongshuang Wu , Maarten Sap

AI agents increasingly excel at generating, testing, and refining code. However, they fall short on tasks requiring formal guarantees of full coverage that testing alone cannot provide. Distributed systems are a prime example: properties…

Voice directors often iteratively refine voice actors' performances by providing feedback to achieve the desired outcome. While this iterative feedback-based refinement process is important in actual recordings, it has been overlooked in…

Sound · Computer Science 2025-07-03 Hiroki Kanagawa , Kenichi Fujita , Aya Watanabe , Yusuke Ijima

The ruling method in modern Artificial Intelligence spanning from Deep Reinforcement Learning (DRL) to Large Language Models (LLMs) relies on a surge of static, externally defined reward functions. While this "extrinsic maximization"…

Machine Learning · Computer Science 2025-12-30 Dhruv Tiwari

Simultaneous speech translation (SST) can be evaluated on simulated online events where human evaluators watch subtitled videos and continuously express their satisfaction by pressing buttons (so called Continuous Rating). Continuous Rating…

Computation and Language · Computer Science 2024-11-15 Dávid Javorský , Dominik Macháček , Ondřej Bojar

Autonomous agents operating around human actors must consider how their behaviors might affect those humans, even when not directly interacting with them. To this end, it is often beneficial to be predictable and appear naturalistic.…

Multiagent Systems · Computer Science 2024-05-30 Hamzah I. Khan , Adam J. Thorpe , David Fridovich-Keil

The advent of multimodal agents facilitates effective interaction within graphical user interface (GUI), especially in ubiquitous GUI control. However, their inability to reliably execute toggle control instructions remains a key…

Artificial Intelligence · Computer Science 2026-03-19 Zongru Wu , Rui Mao , Zhiyuan Tian , Pengzhou Cheng , Tianjie Ju , Zheng Wu , Lingzhong Dong , Haiyue Sheng , Zhuosheng Zhang , Gongshen Liu
‹ Prev 1 8 9 10 Next ›