English
Related papers

Related papers: A vision-language model and platform for temporall…

200 papers

Automated analysis of surgical videos is crucial for improving surgical training, workflow optimization, and postoperative assessment. We introduce a CSMAE, Masked Autoencoder (MAE)-based pretraining approach, specifically developed for…

Computer Vision and Pattern Recognition · Computer Science 2025-02-14 Nisarg A. Shah , Wele Gedara Chaminda Bandara , Shameema Skider , S. Swaroop Vedula , Vishal M. Patel

Training Artificial Intelligence (AI) models on 3D images presents unique challenges compared to the 2D case: Firstly, the demand for computational resources is significantly higher, and secondly, the availability of large datasets for…

A more robust and holistic language-video representation is the key to pushing video understanding forward. Despite the improvement in training strategies, the quality of the language-video dataset is less attention to. The current plain…

Multimedia · Computer Science 2024-06-21 Yuchen Yang , Yingxuan Duan

Effective evaluation is critical for driving advancements in MLLM research. The surgical action planning (SAP) task, which aims to generate future action sequences from visual inputs, demands precise and sophisticated analytical…

Computer Vision and Pattern Recognition · Computer Science 2025-06-16 Mengya Xu , Zhongzhen Huang , Dillan Imans , Yiru Ye , Xiaofan Zhang , Qi Dou

Current invasive assistive technologies are designed to infer high-dimensional motor control signals from severely paralyzed patients. However, they face significant challenges, including public acceptance, limited longevity, and barriers…

Robotics · Computer Science 2025-05-19 Ali Rabiee , Sima Ghafoori , MH Farhadi , Robert Beyer , Xiangyu Bai , David J Lin , Sarah Ostadabbas , Reza Abiri

While 3D visual self-supervised learning (vSSL) shows promising results in capturing visual representations, it overlooks the clinical knowledge from radiology reports. Meanwhile, 3D medical vision-language pre-training (MedVLP) remains…

Computer Vision and Pattern Recognition · Computer Science 2025-02-26 Che Liu , Cheng Ouyang , Yinda Chen , Cesar César Quilodrán-Casas , Lei Ma , Jie Fu , Yike Guo , Anand Shah , Wenjia Bai , Rossella Arcucci

Robotic assisted (RA) surgery promises to transform surgical intervention. Intuitive Surgical is committed to fostering these changes and the machine learning models and algorithms that will enable them. With these goals in mind we have…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Aneeq Zia , Max Berniker , Rogerio Garcia Nespolo , Xiaorui Zhang , Conor Perreault , Kiran Bhattacharyya , Xi Liu , Ziheng Wang , Satoshi Kondo , Satoshi Kasai , Kousuke Hirasawa , Bo Liu , David Austin , Yiheng Wang , Michal Futrega , Jean-Francois Puget , Zhenqiang Li , Yoichi Sato , Ryo Fujii , Ryo Hachiuma , Mana Masuda , Hideo Saito , An Wang , Mengya Xu , Mobarakol Islam , Long Bai , Winnie Pang , Hongliang Ren , Chinedu Nwoye , Luca Sestini , Nicolas Padoy , Maximilian Nielsen , Samuel Schüttler , Thilo Sentker , Hümeyra Husseini , Ivo Baltruschat , Rüdiger Schmitz , René Werner , Aleksandr Matsun , Mugariya Farooq , Numan Saaed , Jose Renato Restom Viera , Mohammad Yaqub , Neil Getty , Fangfang Xia , Zixuan Zhao , Xiaotian Duan , Xing Yao , Ange Lou , Hao Yang , Jintong Han , Jack Noble , Jie Ying Wu , Tamer Abdulbaki Alshirbaji , Nour Aldeen Jalal , Herag Arabian , Ning Ding , Knut Moeller , Weiliang Chen , Quan He , Muhammad Bilal , Taofeek Akinosho , Adnan Qayyum , Massimo Caputo , Hunaid Vohra , Michael Loizou , Anuoluwapo Ajayi , Ilhem Berrou , Faatihah Niyi-Odumosu , Charlie Budd , Oluwatosin Alabi , Tom Vercauteren , Ruoxi Zhao , Ayberk Acar , John Han , Jumanh Atoum , Yinhong Qin , Surong Hua , Lu Ping , Wenming Wu , Rongfeng Wei , Jinlin Wu , You Pang , Zhen Chen , Tim Jaspers , Amine Yamlahi , Piotr Kalinowski , Dominik Michael , Tim Rädsch , Marco Hübner , Danail Stoyanov , Stefanie Speidel , Lena Maier-Hein , Jie Tian , Ruxin Zhang , Khang Hoang Nguyen , Anh Quoc Nguyen , Tam Minh Nguyen , Khoi Dinh Tran , Minh Nguyen Dang Nhat , Trinh Thi Doan Pham , Linh Van Nguyen , Chunyang Jiang , Dewei Yang , Haitao Li , Yannick Prudent , Thibaut Boissin , Mahmood Alam , Shazad Ashraf , Andrew D. Beggs , Lukman Akanbi , Manuel D. Delgado , Narain Gupta , Amir M. Hajiyavand , Iqbal Qasim , Hafiz A. Alaka , Junaid Qadir , Shu Yang , Yihui Wang , Hao Chen , Shin Paul , Yosuke Yamagishi , Zhang Dong , Hongyun Li , Hongyu Gu , Xiaoliu Ding , Xiaoyao Liu , Xingyu Zhao , Mariana Ribeiro , Tiago Jesus , André Ferreira , Guilherme Barbosa , João Carvalho , Leonardo Barroso , Nuno Gomes , Rafael Peixoto , Rodrigo Ralha , Victor Alves , Stephanie , Nattapat Ittikosil , Achita Chitrapan , Quan Huu Cap , Jiayuan Huang , Shreyas C Dhake , Sergi Kavtaradze , Mobarak I Hoque , Ka Young Kim , Su Yong Yun , Young Tae Kim , Hyeon Bae Kim , Seong Tae Kim , Zuxing Deng , Ling Li , Jieyu Zheng , Xiaojian Li , Anthony Jarc

As generative AI continues to evolve, Vision Language Models (VLMs) have emerged as promising tools in various healthcare applications. One area that remains relatively underexplored is their use in human activity recognition (HAR) for…

Computation and Language · Computer Science 2025-11-18 Abderrazek Abid , Thanh-Cong Ho , Fakhri Karray

Vision-Language-Action (VLA) models show promise for robotic control, yet performance in complex household environments remains sub-optimal. Mobile manipulation requires reasoning about global scene layout, fine-grained geometry, and…

Robotics · Computer Science 2026-03-25 Ruisen Tu , Arth Shukla , Sohyun Yoo , Xuanlin Li , Junxi Li , Jianwen Xie , Hao Su , Zhuowen Tu

We introduce a speech-guided embodied agent framework for video-guided skull base surgery that dynamically executes perception and image-guidance tasks in response to surgeon queries. The proposed system integrates natural language…

Computer Vision and Pattern Recognition · Computer Science 2026-04-17 Jecia Z. Y. Mao , Francis X. Creighton , Russell H. Taylor , Manish Sahu

Vision-and-language models (VLMs) have been increasingly explored in the medical domain, particularly following the success of CLIP in general domain. However, unlike the relatively straightforward pairing of 2D images and text, curating…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Ziyang Zhang , Yang Yu , Xulei Yang , Si Yong Yeo

A large labeled dataset is a key to the success of supervised deep learning, but for medical image segmentation, it is highly challenging to obtain sufficient annotated images for model training. In many scenarios, unannotated images are…

Computer Vision and Pattern Recognition · Computer Science 2021-07-13 Hao Zheng , Jun Han , Hongxiao Wang , Lin Yang , Zhuo Zhao , Chaoli Wang , Danny Z. Chen

Visual-Language-Action models (VLAs) have advanced generalist robot control by mapping multimodal observations and language instructions directly to actions, but sparse action supervision often encourages shortcut mappings rather than…

Robotics · Computer Science 2026-05-04 Hao Luo , Wanpeng Zhang , Yicheng Feng , Sipeng Zheng , Haiweng Xu , Chaoyi Xu , Ziheng Xi , Yuhui Fu , Zongqing Lu

Humans effortlessly interpret images by parsing them into part-whole hierarchies; deep learning excels in learning multi-level feature spaces, but they often lack explicit coding of part-whole relations, a prominent property of medical…

Computer Vision and Pattern Recognition · Computer Science 2024-04-25 Mohammad Reza Hosseinzadeh Taher , Michael B. Gotway , Jianming Liang

We introduce OG-VLA, a novel architecture and learning framework that combines the generalization strengths of Vision Language Action models (VLAs) with the robustness of 3D-aware policies. We address the challenge of mapping natural…

Robotics · Computer Science 2025-11-19 Ishika Singh , Ankit Goyal , Stan Birchfield , Dieter Fox , Animesh Garg , Valts Blukis

Minimally invasive surgery (MIS) has revolutionized many procedures and led to reduced recovery time and risk of patient injury. However, MIS poses additional complexity and burden on surgical teams. Data-driven surgical vision algorithms…

Computer Vision and Pattern Recognition · Computer Science 2025-02-11 Oluwatosin Alabi , Tom Vercauteren , Miaojing Shi

In clinical practice, segmenting specific lesions based on the needs of physicians can significantly enhance diagnostic accuracy and treatment efficiency. However, conventional lesion segmentation models lack the flexibility to distinguish…

Computer Vision and Pattern Recognition · Computer Science 2025-04-22 Shuyi Ouyang , Jinyang Zhang , Xiangye Lin , Xilai Wang , Qingqing Chen , Yen-Wei Chen , Lanfen Lin

Deep neural networks are increasingly applied in automated histopathology. Yet, whole-slide images (WSIs) are often acquired at gigapixel sizes, rendering them computationally infeasible to analyze entirely at high resolution. Diagnostic…

Image and Video Processing · Electrical Eng. & Systems 2026-02-10 Tarun G , Naman Malpani , Gugan Thoppe , Sridharan Devarajan

Medical image segmentation allows quantifying target structure size and shape, aiding in disease diagnosis, prognosis, surgery planning, and comprehension.Building upon recent advancements in foundation Vision-Language Models (VLMs) from…

Computer Vision and Pattern Recognition · Computer Science 2024-06-21 Kanchan Poudel , Manish Dhakal , Prasiddha Bhandari , Rabin Adhikari , Safal Thapaliya , Bishesh Khanal

The rapid progress of multimodal large language models (MLLM) has paved the way for Vision-Language-Action (VLA) paradigms, which integrate visual perception, natural language understanding, and control within a single policy. Researchers…