中文
相关论文

相关论文: CAT-ViL: Co-Attention Gated Vision-Language Embedd…

200 篇论文

Despite the availability of computer-aided simulators and recorded videos of surgical procedures, junior residents still heavily rely on experts to answer their queries. However, expert surgeons are often overloaded with clinical and…

计算机视觉与模式识别 · 计算机科学 2023-05-22 Long Bai , Mobarakol Islam , Lalithkumar Seenivasan , Hongliang Ren

Medical visual question answering (VQA) bridges the gap between visual information and clinical decision-making, enabling doctors to extract understanding from clinical images and videos. In particular, surgical VQA can enhance the…

计算机视觉与模式识别 · 计算机科学 2024-09-04 Long Bai , Guankun Wang , Mobarakol Islam , Lalithkumar Seenivasan , An Wang , Hongliang Ren

Recent advancements in Surgical Visual Question Answering (Surgical-VQA) and related region grounding have shown great promise for robotic and medical applications, addressing the critical need for automated methods in personalized surgical…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Guankun Wang , Long Bai , Wan Jun Nah , Jie Wang , Zhaoxi Zhang , Zhen Chen , Jinlin Wu , Mobarakol Islam , Hongbin Liu , Hongliang Ren

Vision-Language (VL) models have gained significant research focus, enabling remarkable advances in multimodal reasoning. These architectures typically comprise a vision encoder, a Large Language Model (LLM), and a projection module that…

计算机视觉与模式识别 · 计算机科学 2024-02-09 Roy Ganz , Yair Kittenplon , Aviad Aberdam , Elad Ben Avraham , Oren Nuriel , Shai Mazor , Ron Litman

Visual question answering (VQA) is crucial for promoting surgical education. In practice, the needs of trainees are constantly evolving, such as learning more surgical types, adapting to different robots, and learning new surgical…

The visual-question localized-answering (VQLA) system can serve as a knowledgeable assistant in surgical education. Except for providing text-based answers, the VQLA system can highlight the interested region for better surgical scene…

计算机视觉与模式识别 · 计算机科学 2023-07-25 Long Bai , Mobarakol Islam , Hongliang Ren

Visual question answering (VQA) in surgery is largely unexplored. Expert surgeons are scarce and are often overloaded with clinical and academic workloads. This overload often limits their time answering questionnaires from patients,…

计算机视觉与模式识别 · 计算机科学 2022-06-28 Lalithkumar Seenivasan , Mobarakol Islam , Adithya K Krishna , Hongliang Ren

The increasing availability of multimodal data across text, tables, and images presents new challenges for developing models capable of complex cross-modal reasoning. Existing methods for Multimodal Multi-hop Question Answering (MMQA) often…

计算机视觉与模式识别 · 计算机科学 2025-04-14 Qi Zhi Lim , Chin Poo Lee , Kian Ming Lim , Kalaiarasi Sonai Muthu Anbananthen

Advances in GPT-based large language models (LLMs) are revolutionizing natural language processing, exponentially increasing its use across various domains. Incorporating uni-directional attention, these autoregressive LLMs can generate…

计算机视觉与模式识别 · 计算机科学 2023-07-25 Lalithkumar Seenivasan , Mobarakol Islam , Gokul Kannan , Hongliang Ren

Recent advances in Vision-Language-Action (VLA) models have enabled robotic agents to integrate multimodal understanding with action execution. However, our empirical analysis reveals that current VLAs struggle to allocate visual attention…

Vision-language-action (VLA) models have significantly advanced robotic learning, enabling training on large-scale, cross-embodiment data and fine-tuning for specific robots. However, state-of-the-art autoregressive VLAs struggle with…

机器人学 · 计算机科学 2025-11-04 Chengmeng Li , Yaxin Peng

A hierarchical cross-modal fusion model is proposed for vision-language question answering (VLQA) in industrial robotics, targeting the challenges of semantic ambiguity, complex environmental layouts, and domain-specific terminology common…

计算机视觉与模式识别 · 计算机科学 2026-05-05 Ping Li , Bartlomiej Brzozka

Visual Question Answering (VQA) within the surgical domain, utilizing Large Language Models (LLMs), offers a distinct opportunity to improve intra-operative decision-making and facilitate intuitive surgeon-AI interaction. However, the…

计算机视觉与模式识别 · 计算机科学 2024-05-24 Runlong He , Mengya Xu , Adrito Das , Danyal Z. Khan , Sophia Bano , Hani J. Marcus , Danail Stoyanov , Matthew J. Clarkson , Mobarakol Islam

Visual Question Answering (VQA) becomes one of the most active research problems in the medical imaging domain. A well-known VQA challenge is the intrinsic diversity between the image and text modalities, and in the medical VQA task, there…

计算机视觉与模式识别 · 计算机科学 2023-02-28 Yuan Zhou , Jing Mei , Yiqin Yu , Tanveer Syeda-Mahmood

Recent advancements in multimodal fusion have witnessed the remarkable success of vision-language (VL) models, which excel in various multimodal applications such as image captioning and visual question answering. However, building VL…

计算机视觉与模式识别 · 计算机科学 2024-10-24 Zhiwei Hao , Jianyuan Guo , Li Shen , Yong Luo , Han Hu , Yonggang Wen

A great challenge in video-language (VidL) modeling lies in the disconnection between fixed video representations extracted from image/video understanding models and downstream VidL data. Recent studies try to mitigate this disconnection…

计算机视觉与模式识别 · 计算机科学 2022-04-19 Tsu-Jui Fu , Linjie Li , Zhe Gan , Kevin Lin , William Yang Wang , Lijuan Wang , Zicheng Liu

Visual Question Answering (VQA) has recently emerged as a potential research domain, captivating the interest of many in the field of artificial intelligence and computer vision. Despite the prevalence of approaches in English, there is a…

计算机视觉与模式识别 · 计算机科学 2024-08-01 Ngoc Son Nguyen , Van Son Nguyen , Tung Le

An important goal of computer vision is to build systems that learn visual representations over time that can be applied to many tasks. In this paper, we investigate a vision-language embedding as a core representation and show that it…

计算机视觉与模式识别 · 计算机科学 2017-10-17 Tanmay Gupta , Kevin Shih , Saurabh Singh , Derek Hoiem

The deployment of vision-language models (VLMs) in dermatology is hindered by the trilemma of high computational costs, extreme data scarcity, and the black-box nature of deep learning. To address these challenges, we present SkinCLIP-VL, a…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Zhixiang Lu , Shijie Xu , Kaicheng Yan , Xuyue Cai , Chong Zhang , Yulong Li , Angelos Stefanidis , Anh Nguyen , Jionglong Su

In this paper, we propose GTA-VLA(Guide, Think, Act), an interactive Vision-Language-Action (VLA) framework that enables spatially steerable embodied reasoning by allowing users to guide robot policies with explicit visual cues. Existing…

机器人学 · 计算机科学 2026-05-14 Yiran Ling , Qing Lian , Jinghang Li , Qing Jiang , Tianming Zhang , Xiaoke Jiang , Chuanxiu Liu , Jie Liu , Lei Zhang
‹ 上一页 1 2 3 10 下一页 ›