中文
相关论文

相关论文: How do Visual Attributes Influence Web Agents? A C…

200 篇论文

Autonomous agents powered by Large Language Models are transforming AI, creating an imperative for the visualization field to embrace agentic frameworks. However, our field's focus on a human in the sensemaking loop raises critical…

人机交互 · 计算机科学 2025-09-17 Vaishali Dhanoa , Anton Wolter , Gabriela Molina León , Hans-Jörg Schulz , Niklas Elmqvist

Artificial agents are increasingly central to complex interactions and decision-making tasks, yet aligning their behaviors with desired human values remains an open challenge. In this work, we investigate how human-like personality traits…

计算与语言 · 计算机科学 2025-06-03 Seungwon Lim , Seungbeen Lee , Dongjun Min , Youngjae Yu

Visual statistical inference is a way to determine significance of patterns found while exploring data. It is dependent on the evaluation of a lineup, of a data plot among a sample of null plots, by human observers. Each individual is…

应用统计 · 统计学 2014-08-12 Mahbubul Majumder , Heike Hofmann , Dianne Cook

Visually-aware recommendation on E-commerce platforms aims to leverage visual information of items to predict a user's preference. It is commonly observed that user's attention to visual features does not always reflect the real preference.…

信息检索 · 计算机科学 2021-07-14 Ruihong Qiu , Sen Wang , Zhi Chen , Hongzhi Yin , Zi Huang

For web agents to be practically useful, they must adapt to the continuously evolving web environment characterized by frequent updates to user interfaces and content. However, most existing benchmarks only capture the static aspects of the…

计算与语言 · 计算机科学 2024-07-17 Yichen Pan , Dehan Kong , Sida Zhou , Cheng Cui , Yifei Leng , Bing Jiang , Hangyu Liu , Yanyi Shang , Shuyan Zhou , Tongshuang Wu , Zhengyang Wu

The model is based on a vector representation of each agent. The components of the vector are the key continuous attributes that determine the social behavior of the agent. A simple mathematical force vector model is used to predict the…

综合物理 · 物理学 2021-12-15 G. Jordan Maclay , Moody Ahmad

AI agents -- systems that execute multi-step reasoning workflows with persistent state, tool access, and specialist skills -- represent a qualitative shift from prior automation technologies in social science. Unlike chatbots that respond…

人工智能 · 计算机科学 2026-03-10 Yongjun Zhang

Diversity in image generation is essential to ensure fair representations and support creativity in ideation. Hence, many text-to-image models have implemented diversification mechanisms. Yet, after a few iterations of generation, a lack of…

人机交互 · 计算机科学 2025-06-25 M. Michelessa , J. Ng , C. Hurter , B. Y. Lim

Information seeking is an interactive behaviour of the end users with information systems, which occurs in a real environment known as context. Context affects information-seeking behaviour in many different ways. The purpose of this paper…

信息检索 · 计算机科学 2018-05-18 Shahram Sedghi , Zeinab Shourmeij , Iman Tahamtan

To be successful, Vision-and-Language Navigation (VLN) agents must be able to ground instructions to actions based on their surroundings. In this work, we develop a methodology to study agent behavior on a skill-specific basis -- examining…

计算机视觉与模式识别 · 计算机科学 2023-07-21 Zijiao Yang , Arjun Majumdar , Stefan Lee

Vision-and-Language Navigation (VLN) requires an agent to follow natural-language instructions, explore the given environments, and reach the desired target locations. These step-by-step navigational instructions are crucial when the agent…

计算与语言 · 计算机科学 2020-05-08 Yubo Zhang , Hao Tan , Mohit Bansal

LLMs have recently demonstrated strong potential in simulating online shopper behavior. Prior work has improved action prediction by applying SFT on action traces with LLM-generated rationales, and by leveraging RL to further enhance…

计算机与社会 · 计算机科学 2025-10-23 Yimeng Zhang , Jiri Gesi , Ran Xue , Tian Wang , Ziyi Wang , Yuxuan Lu , Sinong Zhan , Huimin Zeng , Qingjun Cui , Yufan Guo , Jing Huang , Mubarak Shah , Dakuo Wang

The dominant paradigm of monolithic scaling in Vision-Language Models (VLMs) is failing for understanding and reasoning in documents, yielding diminishing returns as it struggles with the inherent need of this domain for document-based…

计算机视觉与模式识别 · 计算机科学 2025-11-17 Xinlei Yu , Chengming Xu , Zhangquan Chen , Yudong Zhang , Shilin Lu , Cheng Yang , Jiangning Zhang , Shuicheng Yan , Xiaobin Hu

With recent advances in multi-modal foundation models, the previously text-only large language models (LLM) have evolved to incorporate visual input, opening up unprecedented opportunities for various applications in visualization. Our work…

人机交互 · 计算机科学 2023-12-08 Shusen Liu , Haichao Miao , Zhimin Li , Matthew Olson , Valerio Pascucci , Peer-Timo Bremer

Visual persuasion, which uses visual elements to influence cognition and behaviors, is crucial in fields such as advertising and political communication. With recent advancements in artificial intelligence, there is growing potential to…

计算与语言 · 计算机科学 2025-10-29 Junseo Kim , Jongwook Han , Dongmin Choi , Jongwook Yoon , Eun-Ju Lee , Yohan Jo

Autonomous agents that navigate Graphical User Interfaces (GUIs) to automate tasks like document editing and file management can greatly enhance computer workflows. While existing research focuses on online settings, desktop environments,…

Voice Assistants (VAs) can assist users in various everyday tasks, but many users are reluctant to rely on VAs for intricate tasks like online shopping. This study aims to examine whether the vocal characteristics of VAs can serve as an…

人机交互 · 计算机科学 2024-06-14 Sabid Bin Habib Pias , Ran Huang , Donald Williamson , Minjeong Kim , Apu Kapadia

Visual attributes constitute a large portion of information contained in a scene. Objects can be described using a wide variety of attributes which portray their visual appearance (color, texture), geometry (shape, size, posture), and other…

计算机视觉与模式识别 · 计算机科学 2021-06-18 Khoi Pham , Kushal Kafle , Zhe Lin , Zhihong Ding , Scott Cohen , Quan Tran , Abhinav Shrivastava

Autonomous agents capable of planning, reasoning, and executing actions on the web offer a promising avenue for automating computer tasks. However, the majority of existing benchmarks primarily focus on text-based agents, neglecting many…

Vision-language model (VLM)-based web agents increasingly power high-stakes selection tasks like content recommendation or product ranking by combining multimodal perception with preference reasoning. Recent studies reveal that these agents…

人工智能 · 计算机科学 2025-10-07 Tanqiu Jiang , Min Bai , Nikolaos Pappas , Yanjun Qi , Sandesh Swamy