中文
相关论文

相关论文: Towards Holistic Surgical Scene Understanding

200 篇论文

We propose Cross-Attention in Audio, Space, and Time (CA^2ST), a transformer-based method for holistic video recognition. Recognizing actions in videos requires both spatial and temporal understanding, yet most existing models lack a…

计算机视觉与模式识别 · 计算机科学 2025-11-18 Jongseo Lee , Joohyun Chang , Dongho Lee , Jinwoo Choi

Traditional control and task automation have been successfully demonstrated in a variety of structured, controlled environments through the use of highly specialized modeled robotic systems in conjunction with multiple sensors. However, the…

机器人学 · 计算机科学 2020-02-20 Yang Li , Florian Richter , Jingpei Lu , Emily K. Funk , Ryan K. Orosco , Jianke Zhu , Michael C. Yip

Scene rearrangement, like table tidying, is a challenging task in robotic manipulation due to the complexity of predicting diverse object arrangements. Web-scale trained generative models such as Stable Diffusion can aid by generating…

机器人学 · 计算机科学 2024-12-03 Shutong Jin , Ruiyu Wang , Kuangyi Chen , Florian T. Pokorny

Multiple instance learning (MIL) has emerged as a popular method for classifying histopathology whole slide images (WSIs). Existing approaches typically rely on frozen pre-trained models to extract instance features, neglecting the…

计算机视觉与模式识别 · 计算机科学 2024-03-20 Yi Lin , Zhengjie Zhu , Kwang-Ting Cheng , Hao Chen

Audio-visual segmentation (AVS) aims to segment sound sources in the video sequence, requiring a pixel-level understanding of audio-visual correspondence. As the Segment Anything Model (SAM) has strongly impacted extensive fields of dense…

计算机视觉与模式识别 · 计算机科学 2024-06-11 Juhyeong Seon , Woobin Im , Sebin Lee , Jumin Lee , Sung-Eui Yoon

Scene graphs (SGs) provide structured relational representations crucial for decoding complex, dynamic surgical environments. This PRISMA-ScR-guided scoping review systematically maps the evolving landscape of SG research in surgery,…

计算机视觉与模式识别 · 计算机科学 2025-09-26 Angelo Henriques , Korab Hoxha , Daniel Zapp , Peter C. Issa , Nassir Navab , M. Ali Nasseri

Whole slide imaging is fundamental to biomedical microscopy and computational pathology. Previously, learning representations for gigapixel-sized whole slide images (WSIs) has relied on multiple instance learning with weak labels, which do…

计算机视觉与模式识别 · 计算机科学 2024-05-27 Xinhai Hou , Cheng Jiang , Akhil Kondepudi , Yiwei Lyu , Asadur Chowdury , Honglak Lee , Todd C. Hollon

Following the technological advancements in medicine, the operation rooms are evolving into intelligent environments. The context-aware systems (CAS) can comprehensively interpret the surgical state, enable real-time warning, and support…

计算机视觉与模式识别 · 计算机科学 2023-12-12 Negin Ghamsarian

Purpose: Accurate detection and 6D pose estimation of surgical instruments are crucial for many computer-assisted interventions. However, supervised methods lack flexibility for new or unseen tools and require extensive annotated data. This…

计算机视觉与模式识别 · 计算机科学 2026-03-27 Jonas Hein , Lilian Calvet , Matthias Seibold , Siyu Tang , Marc Pollefeys , Philipp Fürnstahl

Understanding surgical tasks represents an important challenge for autonomy in surgical robotic systems. To achieve this, we propose an online task segmentation framework that uses hierarchical transition state clustering to activate…

机器人学 · 计算机科学 2024-06-17 Yutaro Yamada , Jacinto Colan , Ana Davila , Yasuhisa Hasegawa

Transformer-based models have achieved state-of-the-art performance in various computer vision tasks, including image and video analysis. However, Transformer's complex architecture and black-box nature pose challenges for explainability, a…

计算机视觉与模式识别 · 计算机科学 2024-11-04 Zerui Wang , Yan Liu

Visual emotion analysis (VEA) has attracted great attention recently, due to the increasing tendency of expressing and understanding emotions through images on social networks. Different from traditional vision tasks, VEA is inherently more…

计算机视觉与模式识别 · 计算机科学 2021-09-07 Jingyuan Yang , Jie Li , Xiumei Wang , Yuxuan Ding , Xinbo Gao

In this article, we introduce a novel problem of audio-visual autism behavior recognition, which includes social behavior recognition, an essential aspect previously omitted in AI-assisted autism screening research. We define the task at…

Human Action Recognition (HAR) is a challenging domain in computer vision, involving recognizing complex patterns by analyzing the spatiotemporal dynamics of individuals' movements in videos. These patterns arise in sequential data, such as…

计算机视觉与模式识别 · 计算机科学 2025-01-23 Ali K. AlShami , Ryan Rabinowitz , Khang Lam , Yousra Shleibik , Melkamu Mersha , Terrance Boult , Jugal Kalita

A longstanding challenge surrounding deep learning algorithms is unpacking and understanding how they make their decisions. Explainable Artificial Intelligence (XAI) offers methods to provide explanations of internal functions of algorithms…

人工智能 · 计算机科学 2022-08-16 Amin Nayebi , Sindhu Tipirneni , Brandon Foreman , Chandan K. Reddy , Vignesh Subbian

This paper studies how to introduce viewpoint-invariant feature representations that can help action recognition and detection. Although we have witnessed great progress of action recognition in the past decade, it remains challenging yet…

计算机视觉与模式识别 · 计算机科学 2020-12-07 Junwei Liang , Liangliang Cao , Xuehan Xiong , Ting Yu , Alexander Hauptmann

Advancing AI in computational pathology requires large, high-quality, and diverse datasets, yet existing public datasets are often limited in organ diversity, class coverage, or annotation quality. To bridge this gap, we introduce SPIDER…

图像与视频处理 · 电气工程与系统科学 2025-04-08 Dmitry Nechaev , Alexey Pchelnikov , Ekaterina Ivanova

In robot-assisted laparoscopic radical prostatectomy (RALP), the location of the instrument tip is important to register the ultrasound frame with the laparoscopic camera frame. A long-standing limitation is that the instrument tip position…

计算机视觉与模式识别 · 计算机科学 2025-04-15 Zijian Wu , Shuojue Yang , Yueming Jin , Septimiu E Salcudean

The rapid proliferation of video in applications such as autonomous driving, surveillance, and sports analytics necessitates robust methods for dynamic scene understanding. Despite advances in static scene graph generation and early…

计算机视觉与模式识别 · 计算机科学 2025-07-15 Trong-Thuan Nguyen , Pha Nguyen , Jackson Cothren , Alper Yilmaz , Minh-Triet Tran , Khoa Luu