中文
相关论文

相关论文: SPHERE: Unveiling Spatial Blind Spots in Vision-La…

200 篇论文

Robotic Foundation Models (RFMs) hold great promise as generalist, end-to-end systems for robot control. Yet their ability to generalize across new environments, tasks, and embodiments remains limited. We argue that a major bottleneck lies…

This paper investigates a fundamental problem of scene understanding: how to parse a scene image into a structured configuration (i.e., a semantic object hierarchy with object interaction relations). We propose a deep architecture…

计算机视觉与模式识别 · 计算机科学 2018-01-30 Ruimao Zhang , Liang Lin , Guangrun Wang , Meng Wang , Wangmeng Zuo

Meeting summarization with large language models (LLMs) remains error-prone, often producing outputs with hallucinations, omissions, and irrelevancies. We present FRAME, a modular pipeline that reframes summarization as a semantic…

计算与语言 · 计算机科学 2025-11-17 Frederic Kirstein , Sonu Kumar , Terry Ruas , Bela Gipp

Reliable image correspondences form the foundation of vision-based spatial perception, enabling recovery of 3D structure and camera poses. However, unconstrained feature matching across domains such as aerial, indoor, and outdoor scenes…

计算机视觉与模式识别 · 计算机科学 2025-12-30 Zhimin Shao , Abhay Yadav , Rama Chellappa , Cheng Peng

We introduce FloorplanQA, a diagnostic benchmark for evaluating spatial reasoning in large language models (LLMs). FloorplanQA is grounded in structured representations of indoor scenes, such as (e.g., kitchens, living rooms, bedrooms,…

人工智能 · 计算机科学 2026-05-26 Fedor Rodionov , Abdelrahman Eldesokey , Michael Birsak , John Femiani , Bernard Ghanem , Peter Wonka

Speech emotion recognition predicts a speaker's emotional state from speech signals using discrete labels or continuous dimensions such as arousal, valence, and dominance (VAD). We propose EmoSphere-SER, a joint model that integrates…

声音 · 计算机科学 2025-10-20 Deok-Hyeon Cho , Hyung-Seok Oh , Seung-Bin Kim , Seong-Whan Lee

Multimodal Small-to-Medium sized Language Models (MSLMs) have demonstrated strong capabilities in integrating visual and textual information but still face significant limitations in visual comprehension and mathematical reasoning,…

机器学习 · 计算机科学 2026-01-27 Ashutosh Bajpai , Akshat Bhandari , Akshay Nambi , Tanmoy Chakraborty

Service robots are expected to reliably make sense of complex, fast-changing environments. From a cognitive standpoint, they need the appropriate reasoning capabilities and background knowledge required to exhibit human-like Visual…

人工智能 · 计算机科学 2021-04-02 Agnese Chiatti , Gianluca Bardaro , Enrico Motta , Enrico Daga

Large language models have shown strong reasoning capabilities through chain-structured methods such as Chain-of-Thought. Recent studies optimize thought structures by generating parallel or tree-like structures, switching between long and…

计算与语言 · 计算机科学 2025-10-30 Jinghan Zhang , Fengran Mo , Tharindu Cyril Weerasooriya , Xinyue Ye , Dongjie Wang , Yanjie Fu , Kunpeng Liu

Reasoning about images/objects and their hierarchical interactions is a key concept for the next generation of computer vision approaches. Here we present a new framework to deal with it through a visual hierarchical context-based…

计算机视觉与模式识别 · 计算机科学 2019-09-04 Pedro H. Bugatti , Priscila T. M. Saito , Larry S. Davis

Lightweight vision-language models perform competitively on standard benchmarks yet fail systematically in dense-scene reasoning, where multiple objects, attributes, and relations must be jointly grounded and resolved through multi-step…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Xinrui Shi , Kai Liu , Ziqing Zhang , Jianze Li , Anqi Li , Yulun Zhang

Large Vision Language Models (VLMs) have long struggled with spatial reasoning tasks. Surprisingly, even simple spatial reasoning tasks, such as recognizing "under" or "behind" relationships between only two objects, pose significant…

计算与语言 · 计算机科学 2025-10-14 Shiqi Chen , Tongyao Zhu , Ruochen Zhou , Jinghan Zhang , Siyang Gao , Juan Carlos Niebles , Mor Geva , Junxian He , Jiajun Wu , Manling Li

Humans can intuitively compose and arrange scenes in the 3D space for photography. However, can advanced AI image generators plan scenes with similar 3D spatial awareness when creating images from text or image prompts? We present GenSpace,…

计算机视觉与模式识别 · 计算机科学 2025-06-09 Zehan Wang , Jiayang Xu , Ziang Zhang , Tianyu Pang , Chao Du , Hengshuang Zhao , Zhou Zhao

The Theory of Multiple Intelligences underscores the hierarchical nature of cognitive capabilities. To advance Spatial Artificial Intelligence, we pioneer a psychometric framework defining five Basic Spatial Abilities (BSAs) in Visual…

计算机视觉与模式识别 · 计算机科学 2025-08-06 Wenrui Xu , Dalin Lyu , Weihang Wang , Jie Feng , Chen Gao , Yong Li

Recently, large language models (LLMs) have been explored widely for 3D scene understanding. Among them, training-free approaches are gaining attention for their flexibility and generalization over training-based methods. However, they…

计算机视觉与模式识别 · 计算机科学 2025-11-12 Haida Feng , Hao Wei , Zewen Xu , Haolin Wang , Chade Li , Yihong Wu

We introduce Spatial Reasoning Models (SRMs), a framework to perform reasoning over sets of continuous variables via denoising generative models. SRMs infer continuous representations on a set of unobserved variables, given observations on…

计算机视觉与模式识别 · 计算机科学 2025-06-12 Christopher Wewer , Bart Pogodzinski , Bernt Schiele , Jan Eric Lenssen

Vision-language models (VLMs) have advanced rapidly, but their ability to capture spatial relationships remains a blindspot. Current VLMs are typically built with contrastive language-image pretraining (CLIP) style image encoders. The…

As large language models (LLMs) are applied to increasingly longer and more complex tasks, there is a growing need for realistic long-context benchmarks that require selective reading and integration of heterogeneous, multi-modal…

计算与语言 · 计算机科学 2026-02-06 Aadi Palnitkar , Mingyang Mao , Nicholas Waytowich , Vinicius G. Goecks , Xiaomin Lin

Long-context processing has become a fundamental capability for large language models~(LLMs). To assess model's long-context performance, numerous long-context evaluation benchmarks have been proposed. However, variations in evaluation…

计算与语言 · 计算机科学 2025-07-08 Zecheng Tang , Haitian Wang , Quantong Qiu , Baibei Ji , Ruoxi Sun , Keyan Zhou , Juntao Li , Min Zhang

Understanding mid-level road semantics, which capture the structural and contextual cues that link low-level perception to high-level planning, is essential for reliable autonomous driving and digital map construction. However, existing…

计算机视觉与模式识别 · 计算机科学 2025-12-01 Xiyan Liu , Han Wang , Yuhu Wang , Junjie Cai , Zhe Cao , Jianzhong Yang , Zhen Lu