English
Related papers

Related papers: A vision-language model and platform for temporall…

200 papers

Achieving generalization in robotic manipulation remains a critical challenge, particularly for unseen scenarios and novel tasks. Current Vision-Language-Action (VLA) models, while building on top of general Vision-Language Models (VLMs),…

Robotics · Computer Science 2026-04-07 Yifu Yuan , Haiqin Cui , Yibin Chen , Zibin Dong , Fei Ni , Longxin Kou , Jinyi Liu , Pengyi Li , Yan Zheng , Jianye Hao

Prevailing Vision-Language-Action Models (VLAs) for robotic manipulation are built upon vision-language backbones pretrained on large-scale, but disconnected static web data. As a result, despite improved semantic generalization, the policy…

Robotics · Computer Science 2025-12-22 Jonas Pai , Liam Achenbach , Victoriano Montesinos , Benedek Forrai , Oier Mees , Elvis Nava

Vision-Language-Action (VLA) models extend vision-language models to embodied control by mapping natural-language instructions and visual observations to robot actions. Despite their capabilities, VLA systems face significant challenges due…

Robotics · Computer Science 2025-10-24 Weifan Guan , Qinghao Hu , Aosheng Li , Jian Cheng

Action-conditioned surgical video generation is a critical yet highly challenging problem for robotic surgery. The core difficulty is that low-dimensional control vectors must precisely govern complex image-space evolution. In this work, we…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Bohan Li , Shuojue Yang , Baorui Peng , Xianda Guo , Erli Zhang , Youqi Tao , Junfeng Duan , Daguang Xu , Qi Dou , Xin Jin , Wenjun Zeng , Hao Zhao , Yueming Jin

Surgical scenes convey crucial information about the quality of surgery. Pixel-wise localization of tools and anatomical structures is the first task towards deeper surgical analysis for microscopic or endoscopic surgical views. This is…

Computer Vision and Pattern Recognition · Computer Science 2024-09-13 Çağhan Köksal , Ghazal Ghazaei , Nassir Navab

Medical AI assistants support doctors in disease diagnosis, medical image analysis, and report generation. However, they still face significant challenges in clinical use, including limited accuracy with multimodal content and insufficient…

Computer Vision and Pattern Recognition · Computer Science 2025-05-07 Haonan Wang , Jiaji Mao , Lehan Wang , Qixiang Zhang , Marawan Elbatel , Yi Qin , Huijun Hu , Baoxun Li , Wenhui Deng , Weifeng Qin , Hongrui Li , Jialin Liang , Jun Shen , Xiaomeng Li

Searching through large volumes of medical data to retrieve relevant information is a challenging yet crucial task for clinical care. However the primitive and most common approach to retrieval, involving text in the form of keywords, is…

Image and Video Processing · Electrical Eng. & Systems 2023-06-21 Tong Yu , Pietro Mascagni , Juan Verde , Jacques Marescaux , Didier Mutter , Nicolas Padoy

Vision-language models, while effective in general domains and showing strong performance in diverse multi-modal applications like visual question-answering (VQA), struggle to maintain the same level of effectiveness in more specialized…

Computation and Language · Computer Science 2024-04-26 Cuong Nhat Ha , Shima Asaadi , Sanjeev Kumar Karn , Oladimeji Farri , Tobias Heimann , Thomas Runkler

While traditional computer vision models have historically struggled to generalize to endoscopic domains, the emergence of foundation models has shown promising cross-domain performance. In this work, we present the first large-scale study…

Computer Vision and Pattern Recognition · Computer Science 2025-07-09 Leon Mayer , Tim Rädsch , Dominik Michael , Lucas Luttner , Amine Yamlahi , Evangelia Christodoulou , Patrick Godau , Marcel Knopp , Annika Reinke , Fiona Kolbinger , Lena Maier-Hein

Following the successful paradigm shift of large language models, leveraging pre-training on a massive corpus of data and fine-tuning on different downstream tasks, generalist models have made their foray into computer vision. The…

Image and Video Processing · Electrical Eng. & Systems 2025-11-21 Andrea Moglia , Matteo Leccardi , Matteo Cavicchioli , Alice Maccarini , Marco Marcon , Luca Mainardi , Pietro Cerveri

Surgical scene perception via videos is critical for advancing robotic surgery, telesurgery, and AI-assisted surgery, particularly in ophthalmology. However, the scarcity of diverse and richly annotated video datasets has hindered the…

Computer Vision and Pattern Recognition · Computer Science 2024-07-22 Ming Hu , Peng Xia , Lin Wang , Siyuan Yan , Feilong Tang , Zhongxing Xu , Yimin Luo , Kaimin Song , Jurgen Leitner , Xuelian Cheng , Jun Cheng , Chi Liu , Kaijing Zhou , Zongyuan Ge

Accurate tracking of tissues and instruments in videos is crucial for Robotic-Assisted Minimally Invasive Surgery (RAMIS), as it enables the robot to comprehend the surgical scene with precise locations and interactions of tissues and…

Computer Vision and Pattern Recognition · Computer Science 2025-03-24 Bohan Zhan , Wang Zhao , Yi Fang , Bo Du , Francisco Vasconcelos , Danail Stoyanov , Daniel S. Elson , Baoru Huang

Language-augmented scene representations hold great promise for large-scale robotics applications such as search-and-rescue, smart cities, and mining. Many of these scenarios are time-sensitive, requiring rapid scene encoding while also…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Laszlo Szilagyi , Francis Engelmann , Jeannette Bohg

Semi-supervised learning (SSL) has emerged as an effective paradigm for medical image segmentation, reducing the reliance on extensive expert annotations. Meanwhile, vision-language models (VLMs) have demonstrated strong generalization and…

Computer Vision and Pattern Recognition · Computer Science 2025-11-27 Jiaqi Guo , Mingzhen Li , Hanyu Su , Santiago López , Lexiaozi Fan , Daniel Kim , Aggelos Katsaggelos

We present VISTA (Viewpoint-based Image selection with Semantic Task Awareness), an active exploration method for robots to plan informative trajectories that improve 3D map quality in areas most relevant for task completion. Given an…

Automatic pain intensity estimation plays a pivotal role in healthcare and medical fields. While many methods have been developed to gauge human pain using behavioral or physiological indicators, facial expressions have emerged as a…

Computer Vision and Pattern Recognition · Computer Science 2023-12-13 Issam Serraoui , Eric Granger , Abdenour Hadid , Abdelmalik Taleb-Ahmed

Surgical data science (SDS) is rapidly advancing, yet clinical adoption of artificial intelligence (AI) in surgery remains limited, with inadequate validation emerging as an important contributing factor. In fact, existing validation…

Other Quantitative Biology · Quantitative Biology 2026-04-22 Annika Reinke , Ziying O. Li , Minu D. Tizabi , Pascaline André , Marcel Knopp , Mika M. Rother , Ines P. Machado , Maria S. Altieri , Deepak Alapatt , Sophia Bano , Sebastian Bodenstedt , Oliver Burgert , Elvis C. S. Chen , Justin W. Collins , Olivier Colliot , Evangelia Christodoulou , Tobias Czempiel , Adrito Das , Reuben Docea , Daniel Donoho , Qi Dou , Jennifer Eckhoff , Sandy Engelhardt , Gabor Fichtinger , Philipp Fuernstahl , Pablo García Kilroy , Stamatia Giannarou , Stephen Gilbert , Ines Gockel , Patrick Godau , Jan Gödeke , Teodor P. Grantcharov , Tamas Haidegger , Alexander Hann , Makoto Hashizume , Charles Heitz , Rebecca Hisey , Hanna Hoffmann , Arnaud Huaulmé , Paul F. Jäger , Pierre Jannin , Anthony Jarc , Rohit Jena , Yueming Jin , Leo Joskowicz , Luc Joyeux , Max Kirchner , Axel Krieger , Gernot Kronreif , Kyle Lam , Shlomi Laufer , Joël L. Lavanchy , Gyusung I. Lee , Robert Lim , Peng Liu , Hani J. Marcus , Pietro Mascagni , Ozanan R. Meireles , Beat P. Mueller , Lars Mündermann , Hirenkumar Nakawala , Nassir Navab , Abdourahmane Ndong , Juliane Neumann , Felix Nickel , Marco Nolden , Chinedu Nwoye , Namkee Oh , Nicolas Padoy , Thomas Pausch , Micha Pfeiffer , Tim Rädsch , Hongliang Ren , Nicola Rieke , Dominik Rivoir , Duygu Sarikaya , Samuel Schmidgall , Matthias Seibold , Silvia Seidlitz , Alexander Seitel , Lalith Sharan , Jeffrey H. Siewerdsen , Vinkle Srivastav , Raphael Sznitman , Russell Taylor , Thuy N. Tran , Matthias Unberath , Fons van der Sommen , Martin Wagner , Amine Yamlahi , Shaohua K. Zhou , Aneeq Zia , Amin Madani , Danail Stoyanov , Stefanie Speidel , Daniel A. Hashimoto , Fiona R. Kolbinger , Lena Maier-Hein

Egocentric videos capture how humans manipulate objects and tools, providing diverse motion cues for learning object manipulation. Unlike the costly, expert-driven manual teleoperation commonly used in training Vision-Language-Action models…

Robotics · Computer Science 2025-09-29 Tomoya Yoshida , Shuhei Kurita , Taichi Nishimura , Shinsuke Mori

Maintaining situational awareness (SA) is critical in human-robot teams. Yet, under high workload and dynamic conditions, operators often experience SA gaps. Automated detection of SA gaps could provide timely assistance for operators.…

Despite the immense technology advancement in the surgeries the criteria of assessing the surgical skills still remains based on subjective standards. With the advent of robotic-assisted surgery, new opportunities for objective and…

Robotics · Computer Science 2016-11-15 Mahtab J. Fard , Sattar Ameri , R. Darin Ellis
‹ Prev 1 3 4 5 6 7 10 Next ›