中文
相关论文

相关论文: Multimodal Optimal Transport for Training-free Tem…

200 篇论文

Intra-operative ultrasound is an increasingly important imaging modality in neurosurgery. However, manual interaction with imaging data during the procedures, for example to select landmarks or perform segmentation, is difficult and can be…

计算机视觉与模式识别 · 计算机科学 2019-04-19 Julia Rackerseder , Rüdiger Göbl , Nassir Navab , Christoph Hennersperger

Training temporal action detection in videos requires large amounts of labeled data, yet such annotation is expensive to collect. Incorporating unlabeled or weakly-labeled data to train action detection model could help reduce annotation…

计算机视觉与模式识别 · 计算机科学 2021-02-19 Baifeng Shi , Qi Dai , Judy Hoffman , Kate Saenko , Trevor Darrell , Huijuan Xu

Real-time algorithms for automatically recognizing surgical phases are needed to develop systems that can provide assistance to surgeons, enable better management of operating room (OR) resources and consequently improve safety within the…

计算机视觉与模式识别 · 计算机科学 2018-05-23 Gaurav Yengera , Didier Mutter , Jacques Marescaux , Nicolas Padoy

Accurate nuclear instance segmentation is a pivotal task in computational pathology, supporting data-driven clinical insights and facilitating downstream translational applications. While large vision foundation models have shown promise…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Wen Zhang , Qin Ren , Wenjing Liu , Haibin Ling , Chenyu You

Text-to-video (T2V) generation aims to synthesize videos with high visual quality and temporal consistency that are semantically aligned with input text. Reward-based post-training has emerged as a promising direction to improve the quality…

计算机视觉与模式识别 · 计算机科学 2026-02-03 Minh-Quan Le , Gaurav Mittal , Cheng Zhao , David Gu , Dimitris Samaras , Mei Chen

Task-oriented grasping (TOG) is an essential preliminary step for robotic task execution, which involves predicting grasps on regions of target objects that facilitate intended tasks. Existing literature reveals there is a limited…

机器人学 · 计算机科学 2025-06-09 Valerija Holomjova , Jamie Grech , Dewei Yi , Bruno Yun , Andrew Starkey , Pascal Meißner

Promptable video object segmentation and tracking (VOST) has seen significant advances with the emergence of foundation models like Segment Anything Model 2 (SAM2); however, their application in surgical video analysis remains challenging…

计算机视觉与模式识别 · 计算机科学 2025-08-11 Guoping Xu , Hua-Chieh Shao , You Zhang

Purpose: Segmentation of surgical instruments in endoscopic videos is essential for automated surgical scene understanding and process modeling. However, relying on fully supervised deep learning for this task is challenging because manual…

计算机视觉与模式识别 · 计算机科学 2021-03-03 Manish Sahu , Anirban Mukhopadhyay , Stefan Zachow

Transferring linguistic knowledge from a pretrained language model (PLM) to an acoustic model has been shown to greatly improve the performance of automatic speech recognition (ASR). However, due to the heterogeneous feature distributions…

声音 · 计算机科学 2024-09-06 Xugang Lu , Peng Shen , Yu Tsao , Hisashi Kawai

Task-oriented dialog systems empower users to accomplish their goals by facilitating intuitive and expressive natural language interactions. State-of-the-art approaches in task-oriented dialog systems formulate the problem as a conditional…

计算与语言 · 计算机科学 2024-07-24 Adib Mosharrof , M. H. Maqbool , A. B. Siddique

Temporal action detection (TAD) involves the localization and classification of action instances within untrimmed videos. While standard TAD follows fully supervised learning with closed-set setting on large training data, recent zero-shot…

计算机视觉与模式识别 · 计算机科学 2023-11-08 Thinh Phan , Khoa Vo , Duy Le , Gianfranco Doretto , Donald Adjeroh , Ngan Le

In pursuit of the time-optimal path tracking (TOPT) trajectory of a robot manipulator along a preset path, a beforehand identified robot dynamic model is usually used to obtain the required optimal trajectory for perfect tracking. However,…

机器人学 · 计算机科学 2019-08-06 Jiadong Xiao , Lin Li , Tie Zhang , Yanbiao Zou

Accident prediction and timely preventive actions improve road safety by reducing the risk of injury to road users and minimizing property damage. Hence, they are critical components of advanced driver assistance systems (ADAS) and…

计算机视觉与模式识别 · 计算机科学 2025-12-30 Vipooshan Vipulananthan , Kumudu Mohottala , Kavindu Chinthana , Nimsara Paramulla , Charith D Chitraranjan

Purpose: Foundation models, trained on multitudes of public datasets, often require additional fine-tuning or re-prompting mechanisms to be applied to visually distinct target domains such as surgical videos. Further, without domain…

图像与视频处理 · 电气工程与系统科学 2025-07-02 Ssharvien Kumar Sivakumar , Yannik Frisch , Amin Ranem , Anirban Mukhopadhyay

Many motion-centric video analysis tasks, such as atomic actions, detecting atypical motor behavior in individuals with autism, or analyzing articulatory motion in real-time MRI of human speech, require efficient and interpretable temporal…

计算机视觉与模式识别 · 计算机科学 2025-06-03 Hong Nguyen , Dung Tran , Hieu Hoang , Phong Nguyen , Shrikanth Narayanan

In this paper, we present a concurrent and scalable trajectory optimization method to improve the quality of robot-assisted manufacturing. Our method simultaneously optimizes tool orientations, kinematic redundancy, and waypoint timing on…

机器人学 · 计算机科学 2024-12-23 Yongxue Chen , Tianyu Zhang , Yuming Huang , Tao Liu , Charlie C. L. Wang

Multimarginal optimal transport (MOT) is a powerful framework for modeling interactions between multiple distributions, yet its applicability is bottlenecked by a high computational overhead. Entropic regularization provides computational…

机器学习 · 计算机科学 2025-06-03 Dor Tsur , Ziv Goldfeld , Kristjan Greenewald , Haim Permuter

In the field of computer- and robot-assisted minimally invasive surgery, enormous progress has been made in recent years based on the recognition of surgical instruments in endoscopic images and videos. In particular, the determination of…

计算机视觉与模式识别 · 计算机科学 2024-01-24 Tobias Rueckert , Daniel Rueckert , Christoph Palm

Deep learning has demonstrated remarkable success in medical image segmentation and computer-aided diagnosis. In particular, numerous advanced methods have achieved state-of-the-art performance in brain tumor segmentation from MRI scans.…

计算机视觉与模式识别 · 计算机科学 2025-06-24 Xiaoyu Shi , Rahul Kumar Jain , Yinhao Li , Ruibo Hou , Jingliang Cheng , Jie Bai , Guohua Zhao , Lanfen Lin , Rui Xu , Yen-wei Chen

Surgical phase recognition is a fundamental task in computer-assisted surgery systems. Most existing works are under the supervision of expensive and time-consuming full annotations, which require the surgeons to repeat watching videos to…

计算机视觉与模式识别 · 计算机科学 2022-12-02 Xinpeng Ding , Xinjian Yan , Zixun Wang , Wei Zhao , Jian Zhuang , Xiaowei Xu , Xiaomeng Li