中文
相关论文

相关论文: Single- and Multi-Task Architectures for Surgical …

200 篇论文

Computer-assisted minimally invasive surgery has great potential in benefiting modern operating theatres. The video data streamed from the endoscope provides rich information to support context-awareness for next-generation intelligent…

计算机视觉与模式识别 · 计算机科学 2022-08-04 Ziyi Wang , Bo Lu , Yonghao Long , Fangxun Zhong , Tak-Hong Cheung , Qi Dou , Yunhui Liu

Robotic-assisted minimally invasive esophagectomy (RAMIE) is a recognized treatment for esophageal cancer, offering better patient outcomes compared to open surgery and traditional minimally invasive surgery. RAMIE is highly complex,…

Medical image segmentation has been significantly advanced by deep learning (DL) techniques, though the data scarcity inherent in medical applications poses a great challenge to DL-based segmentation methods. Self-supervised learning offers…

计算机视觉与模式识别 · 计算机科学 2024-02-13 Binyan Hu , A. K. Qin

Conventional video models rely on a single stream to capture the complex spatial-temporal features. Recent work on two-stream video models, such as SlowFast network and AssembleNet, prescribe separate streams to learn complementary…

计算机视觉与模式识别 · 计算机科学 2021-08-31 Xinyu Gong , Heng Wang , Zheng Shou , Matt Feiszli , Zhangyang Wang , Zhicheng Yan

Spatial understanding of the physical world from 2D visual inputs hinges on two complementary forms of geometric knowledge: holistic 3D structural perception and fine-grained metric scale estimation. Existing multimodal large language…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Yufei Zheng , Xuhan Zhu , Zide Liu , Chunpeng Zhou , Chenfeng Wang , Yongchao Xu , Yunnan Wang , Jiawei Liu , Pengfei Yu , Wei Zhai , Yang Cao , Zheng-Jun Zha

Balancing temporal resolution and spatial detail under limited compute budget remains a key challenge for video-based multi-modal large language models (MLLMs). Existing methods typically compress video representations using predefined…

计算机视觉与模式识别 · 计算机科学 2025-04-03 Min Shi , Shihao Wang , Chieh-Yun Chen , Jitesh Jain , Kai Wang , Junjun Xiong , Guilin Liu , Zhiding Yu , Humphrey Shi

Real-time image segmentation demands architectures that preserve fine spatial detail while capturing global context under tight latency and memory budgets. Image segmentation is one of the most fundamental problems in computer vision and…

计算机视觉与模式识别 · 计算机科学 2025-11-11 Hongyuan Yu , Cheng Wan , Xiyang Dai , Mengchen Liu , Dongdong Chen , Bin Xiao , Yan Huang , Yuan Lu , Liang Wang

Accurate and reliable lane detection is vital for the safe performance of lane-keeping assistance and lane departure warning systems. However, under certain challenging circumstances, it is difficult to get satisfactory performance in…

计算机视觉与模式识别 · 计算机科学 2022-07-21 Yongqi Dong , Sandeep Patil , Bart van Arem , Haneen Farah

We explore a solution for learning disease signatures from weakly, yet easily obtainable, annotated volumetric medical imaging data by analyzing 3D volumes as a sequence of 2D images. We demonstrate the performance of our solution in the…

计算机视觉与模式识别 · 计算机科学 2020-01-27 Nathaniel Braman , David Beymer , Ehsan Dehghan

Representation learning of the task-oriented attention while tracking instrument holds vast potential in image-guided robotic surgery. Incorporating cognitive ability to automate the camera control enables the surgeon to concentrate more on…

计算机视觉与模式识别 · 计算机科学 2021-12-16 Mobarakol Islam , Vibashan VS , Chwee Ming Lim , Hongliang Ren

Purpose: Surgical workflow recognition enables context-aware assistance and skill assessment in computer-assisted interventions. Despite recent advances, current methods suffer from two critical challenges: prediction jitter across…

计算机视觉与模式识别 · 计算机科学 2025-12-23 Yueyao Chen , Kai-Ni Wang , Dario Tayupo , Arnaud Huaulm'e , Krystel Nyangoh Timoh , Pierre Jannin , Qi Dou

Surgical triplet recognition, which involves identifying instrument, verb, target, and their combinations, is a complex surgical scene understanding challenge plagued by long-tailed data distribution. The mainstream multi-task learning…

计算机视觉与模式识别 · 计算机科学 2025-09-17 Yiyi Zhang , Yuchen Yuan , Ying Zheng , Jialun Pei , Jinpeng Li , Zheng Li , Pheng-Ann Heng

Large Language Model (LLM) agents are increasingly applied to engineering design tasks, yet existing evaluation frameworks do not adequately address multi-agent systems that combine simulation, retrieval, and manufacturing preparation. We…

人工智能 · 计算机科学 2026-05-28 Gioele Molinari , Florian Felten , Soheyl Massoudi , Mark Fuge

Motivated by the current research in data centers and cloud computing, we study the problem of scheduling a set of two-stage jobs on multiple two-stage flowshops. A new formulation for configurations of such scheduling is proposed, which…

数据结构与算法 · 计算机科学 2018-01-30 Guangwei Wu , Jianer Chen , Jianxin Wang

Activity recognition from long unstructured egocentric photo-streams has several applications in assistive technology such as health monitoring and frailty detection, just to name a few. However, one of its main technical challenges is to…

计算机视觉与模式识别 · 计算机科学 2017-08-29 Alejandro Cartas , Mariella Dimiccoli , Petia Radeva

Surgical phase recognition is crucial for enhancing the efficiency and safety of computer-assisted interventions. One of the fundamental challenges involves modeling the long-distance temporal relationships present in surgical videos.…

计算机视觉与模式识别 · 计算机科学 2024-07-12 Rui Cao , Jiangliu Wang , Yun-Hui Liu

Existing state-of-the-art methods for surgical phase recognition either rely on the extraction of spatial-temporal features at a short-range temporal resolution or adopt the sequential extraction of the spatial and temporal features across…

计算机视觉与模式识别 · 计算机科学 2024-08-08 Shu Yang , Luyang Luo , Qiong Wang , Hao Chen

Despite its exceptional soft tissue contrast, Magnetic Resonance Imaging (MRI) faces the challenge of long scanning times compared to other modalities like X-ray radiography. Shortening scanning times is crucial in clinical settings, as it…

机器学习 · 计算机科学 2023-12-08 Thomas Sanchez

The drastic variation of motion in spatial and temporal dimensions makes the video prediction task extremely challenging. Existing RNN models obtain higher performance by deepening or widening the model. They obtain the multi-scale features…

计算机视觉与模式识别 · 计算机科学 2024-02-19 Zhifeng Ma , Hao Zhang , Jie Liu

Modeling and recognition of surgical activities poses an interesting research problem. Although a number of recent works studied automatic recognition of surgical activities, generalizability of these works across different tasks and…

计算机视觉与模式识别 · 计算机科学 2020-08-17 Duygu Sarikaya , Pierre Jannin