中文
相关论文

相关论文: Through the Lens of Character: Resolving Modality-…

200 篇论文

Recently, Multimodal Large Language Models (MLLMs) and Vision Language Models (VLMs) have shown great promise in language-guided perceptual tasks such as recognition, segmentation, and object detection. However, their effectiveness in…

计算机视觉与模式识别 · 计算机科学 2025-09-16 Xu Cao , Yifan Shen , Bolin Lai , Wenqian Ye , Yunsheng Ma , Joerg Heintz , Jintai Chen , Meihuan Huang , Jianguo Cao , Aidong Zhang , James M. Rehg

Most current audio-visual emotion recognition models lack the flexibility needed for deployment in practical applications. We envision a multimodal system that works even when only one modality is available and can be implemented…

机器学习 · 计算机科学 2026-01-13 Lucas Goncalves , Seong-Gyun Leem , Wei-Cheng Lin , Berrak Sisman , Carlos Busso

The integration of Artificial Intelligence (AI) in medical diagnostics is often hindered by model opacity, where high-accuracy systems function as "black boxes" without transparent reasoning. This limitation is critical in clinical…

图像与视频处理 · 电气工程与系统科学 2024-09-23 Pascal Passigan , Vayd Ramkumar

Research on continual learning in multi-modal tasks has been receiving increasing attention. However, most existing work overlooks the explicit cross-modal and cross-task interactions. In this paper, we innovatively propose the Low-rank…

计算机视觉与模式识别 · 计算机科学 2025-01-27 Weicai Yan , Ye Wang , Wang Lin , Zirun Guo , Zhou Zhao , Tao Jin

Recent advancements in Vision-Language (VL) research have sparked new benchmarks for complex visual reasoning, challenging models' advanced reasoning ability. Traditional Vision-Language Models (VLMs) perform well in visual perception tasks…

计算机视觉与模式识别 · 计算机科学 2024-09-24 Zhiyuan Li , Dongnan Liu , Chaoyi Zhang , Heng Wang , Tengfei Xue , Weidong Cai

Traditional multi-agent reinforcement learning (MARL) algorithms, such as independent Q-learning, struggle when presented with partially observable scenarios, and where agents are required to develop delicate action sequences. This is often…

机器学习 · 计算机科学 2022-11-21 F. Bredell , H. A. Engelbrecht , J. C. Schoeman

Vision model have gained increasing attention due to their simplicity and efficiency in Scene Text Recognition (STR) task. However, due to lacking the perception of linguistic knowledge and information, recent vision models suffer from two…

计算机视觉与模式识别 · 计算机科学 2023-05-11 Boqiang Zhang , Hongtao Xie , Yuxin Wang , Jianjun Xu , Yongdong Zhang

Human-centered dynamic scene understanding plays a pivotal role in enhancing the capability of robotic and autonomous systems, in which Video-based Human-Object Interaction (V-HOI) detection is a crucial task in semantic scene…

计算机视觉与模式识别 · 计算机科学 2024-07-22 Hang Zhang , Wenxiao Zhang , Haoxuan Qu , Jun Liu

In this paper, we advance the study of AI-augmented reasoning in the context of Human-Computer Interaction (HCI), psychology and cognitive science, focusing on the critical task of visual perception. Specifically, we investigate the…

人机交互 · 计算机科学 2025-04-18 Shravan Chaudhari , Trilokya Akula , Yoon Kim , Tom Blake

Vision-language models (VLMs) have shown remarkable advancements in multimodal reasoning tasks. However, they still often generate inaccurate or irrelevant responses due to issues like hallucinated image understandings or unrefined…

计算机视觉与模式识别 · 计算机科学 2025-04-24 Di Zhang , Junxian Li , Jingdi Lei , Xunzhi Wang , Yujie Liu , Zonglin Yang , Jiatong Li , Weida Wang , Suorong Yang , Jianbo Wu , Peng Ye , Wanli Ouyang , Dongzhan Zhou

Exploiting relationships between visual regions and question words have achieved great success in learning multi-modality features for Visual Question Answering (VQA). However, we argue that existing methods mostly model relations between…

计算机视觉与模式识别 · 计算机科学 2019-08-14 Peng Gao , Haoxuan You , Zhanpeng Zhang , Xiaogang Wang , Hongsheng Li

Cooperative multi-agent reinforcement learning (MARL) struggles with sample efficiency, interpretability, and generalization. While Large Language Models (LLMs) offer powerful planning capabilities, their application has been hampered by a…

人工智能 · 计算机科学 2026-05-06 Zhiyuan Li , Wenshuai Zhao , Joni Pajarinen

This paper presents an innovative large language model (LLM)-based robotic system for enhancing multi-modal human-robot interaction (HRI). Traditional HRI systems relied on complex designs for intent estimation, reasoning, and behavior…

Recent advances in robotics and large language models (LLMs) have sparked growing interest in human-robot collaboration and embodied intelligence. To enable the broader deployment of robots in human-populated environments, socially-aware…

机器人学 · 计算机科学 2025-03-14 Weizheng Wang , Ike Obi , Byung-Cheol Min

Modern video games are becoming richer and more complex in terms of game mechanics. This complexity allows for the emergence of a wide variety of ways to play the game across the players. From the point of view of the game designer, this…

人工智能 · 计算机科学 2022-11-30 Pierre Le Pelletier de Woillemont , Rémi Labory , Vincent Corruble

Visual recognition models have achieved unprecedented success in various tasks. While researchers aim to understand the underlying mechanisms of these models, the growing demand for deployment in safety-critical areas like autonomous…

计算机视觉与模式识别 · 计算机科学 2026-03-12 Qiyang Wan , Chengzhi Gao , Ruiping Wang , Xilin Chen

Large vision-language models (LVLMs) have emerged as a powerful paradigm for multimodal intelligence, but their growing deployment also expands the attack surface of prompt injection. Despite this growing concern, existing attacks still…

密码学与安全 · 计算机科学 2026-05-18 Hao Yang , Zhuo Ma , Yang Liu , Yilong Yang , Guancheng Wang , JianFeng Ma

We present a Collaborative Agent-Based Framework for Multi-Image Reasoning. Our approach tackles the challenge of interleaved multimodal reasoning across diverse datasets and task formats by employing a dual-agent system: a language-based…

When applied to autonomous vehicle (AV) settings, action recognition can enhance an environment model's situational awareness. This is especially prevalent in scenarios where traditional geometric descriptions and heuristics in AVs are…

计算机视觉与模式识别 · 计算机科学 2024-01-17 Eddy Zhou , Alex Zhuang , Alikasim Budhwani , Owen Leather , Rowan Dempster , Quanquan Li , Mohammad Al-Sharman , Derek Rayside , William Melek

Reinforcement learning (RL), large language models (LLMs), and vision-language models (VLMs) have been widely studied in isolation. However, existing infrastructure lacks the ability to deploy agents from different decision-making paradigms…