English
Related papers

Related papers: BridgeEQA: Virtual Embodied Agents for Real Bridge…

200 papers

We introduce HY-Embodied-0.5, a family of foundation models specifically designed for real-world embodied agents. To bridge the gap between general Vision-Language Models (VLMs) and the demands of embodied agents, our models are developed…

Large language models perform well on many medical QA benchmarks, but real clinical reasoning often requires integrating evidence across multiple images rather than interpreting a single view. We introduce MedThinkVQA, an expert-annotated…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Zonghai Yao , Benlu Wang , Yifan Zhang , Junda Wang , Iris Xia , Zhipeng Tang , Shuo Han , Feiyun Ouyang , Zhichao Yang , Arman Cohan , Hong Yu

Embodied AI models often employ off the shelf vision backbones like CLIP to encode their visual observations. Although such general purpose representations encode rich syntactic and semantic information about the scene, much of this…

Computer Vision and Pattern Recognition · Computer Science 2024-03-12 Ainaz Eftekhar , Kuo-Hao Zeng , Jiafei Duan , Ali Farhadi , Ani Kembhavi , Ranjay Krishna

Evaluating vision-language models (VLMs) in urban driving contexts remains challenging, as existing benchmarks rely on open-ended responses that are ambiguous, annotation-intensive, and inconsistent to score. This lack of standardized…

Computer Vision and Pattern Recognition · Computer Science 2025-10-07 Boshra Khalili , Andrew W. Smyth

Vision-Language-Action (VLA) models drive next-generation autonomous systems, but training them requires scalable, high-quality annotations from complex environments. Current cloud pipelines rely on generic vision-language models (VLMs)…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Kangan Qian , ChuChu Xie , Yang Zhong , Jingrui Pang , Siwen Jiao , Sicong Jiang , Zilin Huang , Yunlong Wang , Kun Jiang , Mengmeng Yang , Hao Ye , Guanghao Zhang , Hangjun Ye , Guang Chen , Long Chen , Diange Yang

As embodied intelligence emerges as a core frontier in artificial intelligence research, simulation platforms must evolve beyond low-level physical interactions to capture complex, human-centered social behaviors. We introduce FreeAskWorld,…

Artificial Intelligence · Computer Science 2025-12-23 Yuhang Peng , Yizhou Pan , Xinning He , Jihaoyu Yang , Xinyu Yin , Han Wang , Xiaoji Zheng , Chao Gao , Jiangtao Gong

Knowledge-based visual question answering (KB-VQA) requires vision-language models to understand images and use external knowledge, especially for rare entities and long-tail facts. Most existing retrieval-augmented generation (RAG) methods…

Computer Vision and Pattern Recognition · Computer Science 2026-04-10 Zhuohong Chen , Zhenxian Wu , Yunyao Yu , Hangrui Xu , Zirui Liao , Zhifang Liu , Xiangwen Deng , Pen Jiao , Haoqian Wang

Embodied AI is widely recognized as a cornerstone of artificial general intelligence (AGI) because it involves controlling embodied agents to perform tasks in the physical world. Building on the success of large language models (LLMs) and…

Robotics · Computer Science 2026-05-04 Yueen Ma , Zixing Song , Yuzheng Zhuang , Jianye Hao , Irwin King

We contend that embodied learning is fundamentally a lifecycle problem rather than a single-stage optimization. Systems that optimize only one link (data collection, simulation, learning, or deployment) rarely sustain improvement or…

Visual question answering (VQA) is a challenging task to provide an accurate natural language answer given an image and a natural language question about the image. It involves multi-modal learning, i.e., computer vision (CV) and natural…

Computer Vision and Pattern Recognition · Computer Science 2022-11-29 Luoqian Jiang , Yifan He , Jian Chen

Passive visual systems typically fail to recognize objects in the amodal setting where they are heavily occluded. In contrast, humans and other embodied agents have the ability to move in the environment, and actively control the viewing…

Computer Vision and Pattern Recognition · Computer Science 2019-04-10 Jianwei Yang , Zhile Ren , Mingze Xu , Xinlei Chen , David Crandall , Devi Parikh , Dhruv Batra

Visual Question Answering (VQA) holds great promise for clinical support, particularly in ophthalmology, where retinal fundus photography is essential for diagnosis. However, ophthalmic VQA benchmarks primarily emphasize answer accuracy,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-22 Xingyue Wang , Bo Liu , Meng Wang , Zhixuan Zhang , Chengcheng Zhu , Huazhu Fu , Jiang Liu

Visual question answering (VQA) is known as an AI-complete task as it requires understanding, reasoning, and inferring about the vision and the language content. Over the past few years, numerous neural architectures have been suggested for…

Computer Vision and Pattern Recognition · Computer Science 2024-04-15 Övgü Özdemir , Erdem Akagündüz

Entity Alignment (EA) aims to detect descriptions of the same real-world entities among different Knowledge Graphs (KG). Several embedding methods have been proposed to rank potentially matching entities of two KGs according to their…

In this paper, we propose a novel end-to-end trainable Video Question Answering (VideoQA) framework with three major components: 1) a new heterogeneous memory which can effectively learn global context information from appearance and motion…

Computer Vision and Pattern Recognition · Computer Science 2019-04-10 Chenyou Fan , Xiaofan Zhang , Shu Zhang , Wensheng Wang , Chi Zhang , Heng Huang

Recent advances in multimodal vision and language modeling have predominantly focused on the English language, mostly due to the lack of multilingual multimodal datasets to steer modeling efforts. In this work, we address this gap and…

Computation and Language · Computer Science 2022-03-18 Jonas Pfeiffer , Gregor Geigle , Aishwarya Kamath , Jan-Martin O. Steitz , Stefan Roth , Ivan Vulić , Iryna Gurevych

We present ECG-Expert-QA, a comprehensive multimodal dataset for evaluating diagnostic capabilities in electrocardiogram (ECG) interpretation. It combines real-world clinical ECG data with systematically generated synthetic cases, covering…

Signal Processing · Electrical Eng. & Systems 2025-04-08 Xu Wang , Jiaju Kang , Puyu Han , Yubao Zhao , Qian Liu , Liwenfei He , Lingqiong Zhang , Lingyun Dai , Yongcheng Wang , Jie Tao

There is no limit to how much a robot might explore and learn, but all of that knowledge needs to be searchable and actionable. Within language research, retrieval augmented generation (RAG) has become the workhorse of large-scale…

Visual question answering (VQA) requires joint comprehension of images and natural language questions, where many questions can't be directly or clearly answered from visual content but require reasoning from structured human knowledge with…

Computer Vision and Pattern Recognition · Computer Science 2018-06-14 Zhou Su , Chen Zhu , Yinpeng Dong , Dongqi Cai , Yurong Chen , Jianguo Li

Existing AGIQA models typically estimate image quality by measuring and aggregating the similarities between image embeddings and text embeddings derived from multi-grade quality descriptions. Although effective, we observe that such…

Computer Vision and Pattern Recognition · Computer Science 2025-12-30 Zhicheng Liao , Baoliang Chen , Hanwei Zhu , Lingyu Zhu , Shiqi Wang , Weisi Lin
‹ Prev 1 8 9 10 Next ›