English
Related papers

Related papers: ESCA: Contextualizing Embodied Agents via Scene-Gr…

200 papers

Prompt tuning based on Context Optimization (CoOp) effectively adapts visual-language models (VLMs) to downstream tasks by inferring additional learnable prompt tokens. However, these tokens are less discriminative as they are independent…

Computer Vision and Pattern Recognition · Computer Science 2024-11-25 Hantao Yao , Rui Zhang , Lu Yu , Yongdong Zhang , Changsheng Xu

Embodiment is an important characteristic for all intelligent agents (creatures and robots), while existing scene description tasks mainly focus on analyzing images passively and the semantic understanding of the scenario is separated from…

Robotics · Computer Science 2020-05-08 Sinan Tan , Huaping Liu , Di Guo , Xinyu Zhang , Fuchun Sun

The connection between our 3D surroundings and the descriptive language that characterizes them would be well-suited for localizing and generating human motion in context but for one problem. The complexity introduced by multiple modalities…

Computer Vision and Pattern Recognition · Computer Science 2024-05-30 Zoltán Á. Milacski , Koichiro Niinuma , Ryosuke Kawamura , Fernando de la Torre , László A. Jeni

Large language model-based (LLM-based) multi-agent systems (MAS) are increasingly used to extend agentic problem solving via role specialization and collaboration. MAS workflows can be naturally modeled as directed computation graphs, where…

Computation and Language · Computer Science 2026-05-21 Yang Liu , Jinxuan Cai , Yishen Li , Qi Meng , Zedi Liu , Xin Li , Chen Qian , Chuan Shi , Cheng Yang

The development of embodied agents that can communicate with humans in natural language has gained increasing interest over the last years, as it facilitates the diffusion of robotic platforms in human-populated environments. As a step…

Robotics · Computer Science 2024-04-16 Roberto Bigazzi , Marcella Cornia , Silvia Cascianelli , Lorenzo Baraldi , Rita Cucchiara

Situational Graphs (S-Graphs) merge geometric models of the environment generated by Simultaneous Localization and Mapping (SLAM) approaches with 3D scene graphs into a multi-layered jointly optimizable factor graph. As an advantage,…

Graph reasoning agents operating from natural-language inputs must solve a coupled problem: they must reconstruct a structured graph instance from text, decide whether existing computational assets are sufficient, interact with tools under…

Artificial Intelligence · Computer Science 2026-05-12 Zike Yuan , Yukun Cao , Han Zhang , Jianzhi Yan , Le Liu , Cai ke , Yue Yu , Hui Wang , Ming Liu , Bing Qin

We present Scene-Graph Based Multi-Modal Traffic Agent (SGTA), a modular framework for traffic video understanding that combines structured scene graphs with multi-modal reasoning. It constructs a traffic scene graph from roadside videos…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Xingcheng Zhou , Mingyu Liu , Walter Zimmer , Jiajie Zhang , Alois Knoll

Semantic mapping methods are increasingly used as intermediate scene representations for downstream robotic reasoning and manipulation, yet their evaluation is still largely tied to fixed benchmark datasets with limited coverage of…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Regina Kurkova , Maxim Popov , Sergey Kolyubin

Group-level emotion recognition (GER) aims to identify holistic emotions within a scene involving multiple individuals. Current existed methods underestimate the importance of visual scene contextual information in modeling individual…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 Qing Zhu , Wangdong Guo , Qirong Mao , Xiaohua Huang , Xiuyan Shao , Wenming Zheng

Current Scene Graph Generation (SGG) methods explore contextual information to predict relationships among entity pairs. However, due to the diverse visual appearance of numerous possible subject-object combinations, there is a large…

Computer Vision and Pattern Recognition · Computer Science 2023-03-14 Chaofan Zheng , Xinyu Lyu , Lianli Gao , Bo Dai , Jingkuan Song

Embodied intelligence fundamentally requires a capability to determine where to act in 3D space. We formalize this requirement as embodied localization -- the problem of predicting executable 3D points conditioned on visual observations and…

Robotics · Computer Science 2026-03-31 Qiming Zhu , Zhirui Fang , Tianming Zhang , Chuanxiu Liu , Xiaoke Jiang , Lei Zhang

Vision-Language-Action (VLA) models drive next-generation autonomous systems, but training them requires scalable, high-quality annotations from complex environments. Current cloud pipelines rely on generic vision-language models (VLMs)…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Kangan Qian , ChuChu Xie , Yang Zhong , Jingrui Pang , Siwen Jiao , Sicong Jiang , Zilin Huang , Yunlong Wang , Kun Jiang , Mengmeng Yang , Hao Ye , Guanghao Zhang , Hangjun Ye , Guang Chen , Long Chen , Diange Yang

Large vision-language models (VLMs) have achieved substantial progress in multimodal perception and reasoning. When integrated into an embodied agent, existing embodied VLM works either output detailed action sequences at the manipulation…

Computer Vision and Pattern Recognition · Computer Science 2024-10-22 Jingkang Yang , Yuhao Dong , Shuai Liu , Bo Li , Ziyue Wang , Chencheng Jiang , Haoran Tan , Jiamu Kang , Yuanhan Zhang , Kaiyang Zhou , Ziwei Liu

In Scene Graph Generation (SGG), structured representations are extracted from visual inputs as object nodes and connecting predicates, enabling image-based reasoning for diverse downstream tasks. While fully supervised SGG has improved…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Abdelrahman Elskhawy , Mengze Li , Nassir Navab , Benjamin Busam

Graph based representation has been widely used in modelling spatio-temporal relationships in video understanding. Although effective, existing graph-based approaches focus on capturing the human-object relationships while ignoring…

Computer Vision and Pattern Recognition · Computer Science 2025-01-14 Chinthani Sugandhika , Chen Li , Deepu Rajan , Basura Fernando

3D scene graphs have recently emerged as a powerful high-level representation of 3D environments. A 3D scene graph describes the environment as a layered graph where nodes represent spatial concepts at multiple levels of abstraction and…

Robotics · Computer Science 2022-06-22 Nathan Hughes , Yun Chang , Luca Carlone

Recent advances in large language models (LLMs) have significantly improved language-driven 3D content generation, but most existing approaches still treat scene generation and user interaction as separate processes, limiting the…

Computer Vision and Pattern Recognition · Computer Science 2026-05-08 Anh H. Vo , Sungyo Lee , Phil-Joong Kim , Soo-Mi Choi , Yong-Guk Kim

Zero-shot referring expression comprehension (REC) aims to locate target objects in images given natural language queries without relying on task-specific training data, demanding strong visual understanding capabilities. Existing…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Yike Wu , Necva Bolucu , Stephen Wan , Dadong Wang , Jiahao Xia , Jian Zhang

Surgical captioning plays an important role in surgical instruction prediction and report generation. However, the majority of captioning models still rely on the heavy computational object detector or feature extractor to extract regional…

Computer Vision and Pattern Recognition · Computer Science 2022-07-04 Mengya Xu , Mobarakol Islam , Hongliang Ren
‹ Prev 1 8 9 10 Next ›