中文
相关论文

相关论文: A vision-language model and platform for temporall…

200 篇论文

Achieving generalization in robotic manipulation remains a critical challenge, particularly for unseen scenarios and novel tasks. Current Vision-Language-Action (VLA) models, while building on top of general Vision-Language Models (VLMs),…

机器人学 · 计算机科学 2026-04-07 Yifu Yuan , Haiqin Cui , Yibin Chen , Zibin Dong , Fei Ni , Longxin Kou , Jinyi Liu , Pengyi Li , Yan Zheng , Jianye Hao

Prevailing Vision-Language-Action Models (VLAs) for robotic manipulation are built upon vision-language backbones pretrained on large-scale, but disconnected static web data. As a result, despite improved semantic generalization, the policy…

机器人学 · 计算机科学 2025-12-22 Jonas Pai , Liam Achenbach , Victoriano Montesinos , Benedek Forrai , Oier Mees , Elvis Nava

Vision-Language-Action (VLA) models extend vision-language models to embodied control by mapping natural-language instructions and visual observations to robot actions. Despite their capabilities, VLA systems face significant challenges due…

机器人学 · 计算机科学 2025-10-24 Weifan Guan , Qinghao Hu , Aosheng Li , Jian Cheng

Action-conditioned surgical video generation is a critical yet highly challenging problem for robotic surgery. The core difficulty is that low-dimensional control vectors must precisely govern complex image-space evolution. In this work, we…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Bohan Li , Shuojue Yang , Baorui Peng , Xianda Guo , Erli Zhang , Youqi Tao , Junfeng Duan , Daguang Xu , Qi Dou , Xin Jin , Wenjun Zeng , Hao Zhao , Yueming Jin

Surgical scenes convey crucial information about the quality of surgery. Pixel-wise localization of tools and anatomical structures is the first task towards deeper surgical analysis for microscopic or endoscopic surgical views. This is…

计算机视觉与模式识别 · 计算机科学 2024-09-13 Çağhan Köksal , Ghazal Ghazaei , Nassir Navab

Medical AI assistants support doctors in disease diagnosis, medical image analysis, and report generation. However, they still face significant challenges in clinical use, including limited accuracy with multimodal content and insufficient…

计算机视觉与模式识别 · 计算机科学 2025-05-07 Haonan Wang , Jiaji Mao , Lehan Wang , Qixiang Zhang , Marawan Elbatel , Yi Qin , Huijun Hu , Baoxun Li , Wenhui Deng , Weifeng Qin , Hongrui Li , Jialin Liang , Jun Shen , Xiaomeng Li

Searching through large volumes of medical data to retrieve relevant information is a challenging yet crucial task for clinical care. However the primitive and most common approach to retrieval, involving text in the form of keywords, is…

图像与视频处理 · 电气工程与系统科学 2023-06-21 Tong Yu , Pietro Mascagni , Juan Verde , Jacques Marescaux , Didier Mutter , Nicolas Padoy

Vision-language models, while effective in general domains and showing strong performance in diverse multi-modal applications like visual question-answering (VQA), struggle to maintain the same level of effectiveness in more specialized…

计算与语言 · 计算机科学 2024-04-26 Cuong Nhat Ha , Shima Asaadi , Sanjeev Kumar Karn , Oladimeji Farri , Tobias Heimann , Thomas Runkler

While traditional computer vision models have historically struggled to generalize to endoscopic domains, the emergence of foundation models has shown promising cross-domain performance. In this work, we present the first large-scale study…

Following the successful paradigm shift of large language models, leveraging pre-training on a massive corpus of data and fine-tuning on different downstream tasks, generalist models have made their foray into computer vision. The…

图像与视频处理 · 电气工程与系统科学 2025-11-21 Andrea Moglia , Matteo Leccardi , Matteo Cavicchioli , Alice Maccarini , Marco Marcon , Luca Mainardi , Pietro Cerveri

Surgical scene perception via videos is critical for advancing robotic surgery, telesurgery, and AI-assisted surgery, particularly in ophthalmology. However, the scarcity of diverse and richly annotated video datasets has hindered the…

计算机视觉与模式识别 · 计算机科学 2024-07-22 Ming Hu , Peng Xia , Lin Wang , Siyuan Yan , Feilong Tang , Zhongxing Xu , Yimin Luo , Kaimin Song , Jurgen Leitner , Xuelian Cheng , Jun Cheng , Chi Liu , Kaijing Zhou , Zongyuan Ge

Accurate tracking of tissues and instruments in videos is crucial for Robotic-Assisted Minimally Invasive Surgery (RAMIS), as it enables the robot to comprehend the surgical scene with precise locations and interactions of tissues and…

计算机视觉与模式识别 · 计算机科学 2025-03-24 Bohan Zhan , Wang Zhao , Yi Fang , Bo Du , Francisco Vasconcelos , Danail Stoyanov , Daniel S. Elson , Baoru Huang

Language-augmented scene representations hold great promise for large-scale robotics applications such as search-and-rescue, smart cities, and mining. Many of these scenarios are time-sensitive, requiring rapid scene encoding while also…

计算机视觉与模式识别 · 计算机科学 2025-08-19 Laszlo Szilagyi , Francis Engelmann , Jeannette Bohg

Semi-supervised learning (SSL) has emerged as an effective paradigm for medical image segmentation, reducing the reliance on extensive expert annotations. Meanwhile, vision-language models (VLMs) have demonstrated strong generalization and…

计算机视觉与模式识别 · 计算机科学 2025-11-27 Jiaqi Guo , Mingzhen Li , Hanyu Su , Santiago López , Lexiaozi Fan , Daniel Kim , Aggelos Katsaggelos

We present VISTA (Viewpoint-based Image selection with Semantic Task Awareness), an active exploration method for robots to plan informative trajectories that improve 3D map quality in areas most relevant for task completion. Given an…

Automatic pain intensity estimation plays a pivotal role in healthcare and medical fields. While many methods have been developed to gauge human pain using behavioral or physiological indicators, facial expressions have emerged as a…

计算机视觉与模式识别 · 计算机科学 2023-12-13 Issam Serraoui , Eric Granger , Abdenour Hadid , Abdelmalik Taleb-Ahmed

Surgical data science (SDS) is rapidly advancing, yet clinical adoption of artificial intelligence (AI) in surgery remains limited, with inadequate validation emerging as an important contributing factor. In fact, existing validation…

其他定量生物学 · 定量生物学 2026-04-22 Annika Reinke , Ziying O. Li , Minu D. Tizabi , Pascaline André , Marcel Knopp , Mika M. Rother , Ines P. Machado , Maria S. Altieri , Deepak Alapatt , Sophia Bano , Sebastian Bodenstedt , Oliver Burgert , Elvis C. S. Chen , Justin W. Collins , Olivier Colliot , Evangelia Christodoulou , Tobias Czempiel , Adrito Das , Reuben Docea , Daniel Donoho , Qi Dou , Jennifer Eckhoff , Sandy Engelhardt , Gabor Fichtinger , Philipp Fuernstahl , Pablo García Kilroy , Stamatia Giannarou , Stephen Gilbert , Ines Gockel , Patrick Godau , Jan Gödeke , Teodor P. Grantcharov , Tamas Haidegger , Alexander Hann , Makoto Hashizume , Charles Heitz , Rebecca Hisey , Hanna Hoffmann , Arnaud Huaulmé , Paul F. Jäger , Pierre Jannin , Anthony Jarc , Rohit Jena , Yueming Jin , Leo Joskowicz , Luc Joyeux , Max Kirchner , Axel Krieger , Gernot Kronreif , Kyle Lam , Shlomi Laufer , Joël L. Lavanchy , Gyusung I. Lee , Robert Lim , Peng Liu , Hani J. Marcus , Pietro Mascagni , Ozanan R. Meireles , Beat P. Mueller , Lars Mündermann , Hirenkumar Nakawala , Nassir Navab , Abdourahmane Ndong , Juliane Neumann , Felix Nickel , Marco Nolden , Chinedu Nwoye , Namkee Oh , Nicolas Padoy , Thomas Pausch , Micha Pfeiffer , Tim Rädsch , Hongliang Ren , Nicola Rieke , Dominik Rivoir , Duygu Sarikaya , Samuel Schmidgall , Matthias Seibold , Silvia Seidlitz , Alexander Seitel , Lalith Sharan , Jeffrey H. Siewerdsen , Vinkle Srivastav , Raphael Sznitman , Russell Taylor , Thuy N. Tran , Matthias Unberath , Fons van der Sommen , Martin Wagner , Amine Yamlahi , Shaohua K. Zhou , Aneeq Zia , Amin Madani , Danail Stoyanov , Stefanie Speidel , Daniel A. Hashimoto , Fiona R. Kolbinger , Lena Maier-Hein

Egocentric videos capture how humans manipulate objects and tools, providing diverse motion cues for learning object manipulation. Unlike the costly, expert-driven manual teleoperation commonly used in training Vision-Language-Action models…

机器人学 · 计算机科学 2025-09-29 Tomoya Yoshida , Shuhei Kurita , Taichi Nishimura , Shinsuke Mori

Maintaining situational awareness (SA) is critical in human-robot teams. Yet, under high workload and dynamic conditions, operators often experience SA gaps. Automated detection of SA gaps could provide timely assistance for operators.…

Despite the immense technology advancement in the surgeries the criteria of assessing the surgical skills still remains based on subjective standards. With the advent of robotic-assisted surgery, new opportunities for objective and…

机器人学 · 计算机科学 2016-11-15 Mahtab J. Fard , Sattar Ameri , R. Darin Ellis