中文
相关论文

相关论文: ChatterBox: Multi-round Multimodal Referring and G…

200 篇论文

With the recent significant advancements in large multi-modal models (LMMs), the importance of their grounding capability in visual chat is increasingly recognized. Despite recent efforts to enable LMMs to support grounding, their…

计算机视觉与模式识别 · 计算机科学 2023-12-06 Hao Zhang , Hongyang Li , Feng Li , Tianhe Ren , Xueyan Zou , Shilong Liu , Shijia Huang , Jianfeng Gao , Lei Zhang , Chunyuan Li , Jianwei Yang

Visual grounding (VG) occupies a pivotal position in multi-modality vision-language models. In this study, we propose ViLaM, a large multi-modality model, that supports multi-tasks of VG using the cycle training strategy, with abundant…

计算机视觉与模式识别 · 计算机科学 2024-04-29 Xiaoyu Yang , Lijian Xu , Hao Sun , Hongsheng Li , Shaoting Zhang

Multimodal large language models (MLLMs) are increasingly used for real-world tasks involving multi-step reasoning and long-form generation, where reliability requires grounding model outputs in heterogeneous input sources and verifying…

计算与语言 · 计算机科学 2026-05-08 David Wan , Han Wang , Ziyang Wang , Elias Stengel-Eskin , Hyunji Lee , Mohit Bansal

In this paper, we propose a novel end-to-end model, namely Single-Stage Grounding network (SSG), to localize the referent given a referring expression within an image. Different from previous multi-stage models which rely on object…

计算机视觉与模式识别 · 计算机科学 2018-12-11 Xinpeng Chen , Lin Ma , Jingyuan Chen , Zequn Jie , Wei Liu , Jiebo Luo

In this paper, we present a neat yet effective transformer-based framework for visual grounding, namely TransVG, to address the task of grounding a language query to the corresponding region onto an image. The state-of-the-art methods,…

计算机视觉与模式识别 · 计算机科学 2022-01-17 Jiajun Deng , Zhengyuan Yang , Tianlang Chen , Wengang Zhou , Houqiang Li

Recent advancements in Large Language Models (LLMs) have shown outstanding potential for role-playing applications. Evaluating these capabilities is becoming crucial yet remains challenging. Existing benchmarks mostly adopt a…

计算与语言 · 计算机科学 2025-10-24 Hao Xiang , Tianyi Tang , Yang Su , Bowen Yu , An Yang , Fei Huang , Yichang Zhang , Yaojie Lu , Hongyu Lin , Xianpei Han , Jingren Zhou , Junyang Lin , Le Sun

Multimodal search-based dialogue is a challenging new task: It extends visually grounded question answering systems into multi-turn conversations with access to an external database. We address this new challenge by learning a neural…

计算与语言 · 计算机科学 2018-11-22 Shubham Agarwal , Ondrej Dusek , Ioannis Konstas , Verena Rieser

With the rapid development of multimodal large language models (MLLMs), especially their capabilities in visual chat through refer and ground functionalities, their significance is increasingly recognized. However, the biomedical field…

计算机视觉与模式识别 · 计算机科学 2024-07-01 Xiaoshuang Huang , Haifeng Huang , Lingdong Shen , Yehui Yang , Fangxin Shang , Junwei Liu , Jia Liu

Spatial reasoning plays a vital role in both human cognition and machine intelligence, prompting new research into language models' (LMs) capabilities in this regard. However, existing benchmarks reveal shortcomings in evaluating…

计算与语言 · 计算机科学 2024-05-27 Fangjun Li , David C. Hogg , Anthony G. Cohn

AI models have achieved state-of-the-art results in textual reasoning; however, their ability to reason over spatial and relational structures remains a critical bottleneck -- particularly in early-grade maths, which relies heavily on…

The ability to process information from multiple modalities and to reason through it step-by-step remains a critical challenge in advancing artificial intelligence. However, existing reasoning benchmarks focus on text-only reasoning, or…

人工智能 · 计算机科学 2025-07-01 Yulun Jiang , Yekun Chai , Maria Brbić , Michael Moor

This paper presents ConvBench, a novel multi-turn conversation evaluation benchmark tailored for Large Vision-Language Models (LVLMs). Unlike existing benchmarks that assess individual capabilities in single-turn dialogues, ConvBench adopts…

多媒体 · 计算机科学 2024-04-26 Shuo Liu , Kaining Ying , Hao Zhang , Yue Yang , Yuqi Lin , Tianle Zhang , Chuanhao Li , Yu Qiao , Ping Luo , Wenqi Shao , Kaipeng Zhang

Recent advances in vision-language models (VLMs) have enabled powerful multimodal reasoning, but state-of-the-art approaches typically rely on extremely large models with prohibitive computational and memory requirements. This makes their…

计算机视觉与模式识别 · 计算机科学 2025-09-22 Abdarahmane Traore , Éric Hervet , Andy Couturier

Multimodal large language models (MLLMs) hold promise for integrating diverse data modalities, but current medical adaptations such as LLaVA-Med often fail to fully exploit the synergy between color fundus photography (CFP) and optical…

We introduce Vocal Sandbox, a framework for enabling seamless human-robot collaboration in situated environments. Systems in our framework are characterized by their ability to adapt and continually learn at multiple levels of abstraction…

机器人学 · 计算机科学 2024-11-06 Jennifer Grannen , Siddharth Karamcheti , Suvir Mirchandani , Percy Liang , Dorsa Sadigh

With recent advances in Multimodal Large Language Models (MLLMs), grounding and referring capabilities have gained increasing attention for achieving detailed understanding and flexible user interaction. However, these capabilities still…

计算机视觉与模式识别 · 计算机科学 2025-08-07 Yinan Zhou , Yuxin Chen , Haokun Lin , Yichen Wu , Shuyu Yang , Zhongang Qi , Chen Ma , Li Zhu , Ying Shan

Referring multi-object tracking (RMOT) is an emerging cross-modal task that aims to localize an arbitrary number of targets based on a language expression and continuously track them in a video. This intricate task involves reasoning on…

计算机视觉与模式识别 · 计算机科学 2025-07-28 Wenjun Huang , Yang Ni , Hanning Chen , Yirui He , Ian Bryant , Yezi Liu , Mohsen Imani

Intelligent robots designed to interact with humans in real scenarios need to be able to refer to entities actively by natural language. In spatial referring expression generation, the ambiguity is unavoidable due to the diversity of…

机器人学 · 计算机科学 2022-04-05 Mingjiang Liu , Chengli Xiao , Chunlin Chen

Accurate evaluation of conversational retrieval is pivotal for advancing Retrieval-Augmented Generation (RAG) systems. However, existing conversational retrieval benchmarks suffer from costly, sparse human annotation or rigid, unnatural…

With enhanced capabilities and widespread applications, Multimodal Large Language Models (MLLMs) are increasingly required to process and reason over multiple images simultaneously. However, existing MLLM benchmarks focus either on…