中文
相关论文

相关论文: Visualization Framework for Colonoscopy Videos

200 篇论文

Visual explanation methods have an important role in the prognosis of the patients where the annotated data is limited or unavailable. There have been several attempts to use gradient-based attribution methods to localize pathology from…

图像与视频处理 · 电气工程与系统科学 2021-06-24 Ugur Demir , Ismail Irmakci , Elif Keles , Ahmet Topcu , Ziyue Xu , Concetto Spampinato , Sachin Jambawalikar , Evrim Turkbey , Baris Turkbey , Ulas Bagci

Video Captioning is considered to be one of the most challenging problems in the field of computer vision. Video Captioning involves the combination of different deep learning models to perform object detection, action detection, and…

计算机视觉与模式识别 · 计算机科学 2021-04-08 Soheyla Amirian , Abolfazl Farahani , Hamid R. Arabnia , Khaled Rasheed , Thiab R. Taha

Understanding objects in videos in terms of fine-grained localization masks and detailed semantic properties is a fundamental task in video understanding. In this paper, we propose VoCap, a flexible video model that consumes a video and a…

计算机视觉与模式识别 · 计算机科学 2025-09-01 Jasper Uijlings , Xingyi Zhou , Xiuye Gu , Arsha Nagrani , Anurag Arnab , Alireza Fathi , David Ross , Cordelia Schmid

In minimally invasive surgery, surgical instrument localization is a crucial task for endoscopic videos, which enables various applications for improving surgical outcomes. However, annotating the instrument localization in endoscopic…

图像与视频处理 · 电气工程与系统科学 2024-06-24 Rongfeng Wei , Jinlin Wu , Xuexue Bai , Ming Feng , Zhen Lei , Hongbin Liu , Zhen Chen

We present in this paper an intelligent video data visualization tool, based on semantic classification, for retrieving and exploring a large scale corpus of videos. Our work is based on semantic classification resulting from semantic…

信息检索 · 计算机科学 2012-09-07 Jamel Slimi , Anis Ben Ammar , Adel M. Alimi

Interactive segmentation is a crucial research area in medical image analysis aiming to boost the efficiency of costly annotations by incorporating human feedback. This feedback takes the form of clicks, scribbles, or masks and allows for…

图像与视频处理 · 电气工程与系统科学 2024-10-28 Zdravko Marinov , Paul F. Jäger , Jan Egger , Jens Kleesiek , Rainer Stiefelhagen

Video question answering is a challenging task that requires understanding jointly the language input, the visual information in individual video frames, as well as the temporal information about the events occurring in the video. In this…

计算机视觉与模式识别 · 计算机科学 2022-08-02 AJ Piergiovanni , Kairo Morton , Weicheng Kuo , Michael S. Ryoo , Anelia Angelova

Most natural videos contain numerous events. For example, in a video of a "man playing a piano", the video might also contain "another man dancing" or "a crowd clapping". We introduce the task of dense-captioning events, which involves both…

计算机视觉与模式识别 · 计算机科学 2017-05-03 Ranjay Krishna , Kenji Hata , Frederic Ren , Li Fei-Fei , Juan Carlos Niebles

Video summarization is a crucial research area that aims to efficiently browse and retrieve relevant information from the vast amount of video content available today. With the exponential growth of multimedia data, the ability to extract…

计算机视觉与模式识别 · 计算机科学 2024-04-09 Hai-Dang Huynh-Lam , Ngoc-Phuong Ho-Thi , Minh-Triet Tran , Trung-Nghia Le

Background: The clinical documentation of cystoscopy includes visual and textual materials. However, the secondary use of visual cystoscopic data for educational and research purposes remains limited due to inefficient data management in…

Animated data videos have gained significant popularity in recent years. However, authoring data videos remains challenging due to the complexity of creating and coordinating diverse components (e.g., visualization, animation, audio, etc.).…

人机交互 · 计算机科学 2025-02-10 Leixian Shen , Haotian Li , Yun Wang , Huamin Qu

Automatic annotation of images with descriptive words is a challenging problem with vast applications in the areas of image search and retrieval. This problem can be viewed as a label-assignment problem by a classifier dealing with a very…

计算机视觉与模式识别 · 计算机科学 2017-05-09 Amara Tariq , Hassan Foroosh

Surgical scene perception via videos is critical for advancing robotic surgery, telesurgery, and AI-assisted surgery, particularly in ophthalmology. However, the scarcity of diverse and richly annotated video datasets has hindered the…

计算机视觉与模式识别 · 计算机科学 2024-07-22 Ming Hu , Peng Xia , Lin Wang , Siyuan Yan , Feilong Tang , Zhongxing Xu , Yimin Luo , Kaimin Song , Jurgen Leitner , Xuelian Cheng , Jun Cheng , Chi Liu , Kaijing Zhou , Zongyuan Ge

Image annotation is one of the most essential tasks for guaranteeing proper treatment for patients and tracking progress over the course of therapy in the field of medical imaging and disease diagnosis. However, manually annotating a lot of…

计算机视觉与模式识别 · 计算机科学 2024-03-25 Md Abdul Kadir , Hasan Md Tusfiqur Alam , Pascale Maul , Hans-Jürgen Profitlich , Moritz Wolf , Daniel Sonntag

In this paper we introduce a new dataset for 360-degree video summarization: the transformation of 360-degree video content to concise 2D-video summaries that can be consumed via traditional devices, such as TV sets and smartphones. The…

计算机视觉与模式识别 · 计算机科学 2024-06-06 Ioannis Kontostathis , Evlampios Apostolidis , Vasileios Mezaris

With the broad growth of video capturing devices and applications on the web, it is more demanding to provide desired video content for users efficiently. Video summarization facilitates quickly grasping video content by creating a compact…

计算机视觉与模式识别 · 计算机科学 2022-10-24 Mayu Otani , Yale Song , Yang Wang

Multimodal vision-language (VL) learning has noticeably pushed the tendency toward generic intelligence owing to emerging large foundation models. However, tracking, as a fundamental vision problem, surprisingly enjoys less bonus from…

计算机视觉与模式识别 · 计算机科学 2023-07-20 Mingzhe Guo , Zhipeng Zhang , Liping Jing , Haibin Ling , Heng Fan

Using a collection of publicly available links to short form video clips of an average of 6 seconds duration each, 1,275 users manually annotated each video multiple times to indicate both long-term and short-term memorability of the…

Many recent machine learning approaches used in medical imaging are highly reliant on large amounts of image and ground truth data. In the context of object segmentation, pixel-wise annotations are extremely expensive to collect, especially…

计算机视觉与模式识别 · 计算机科学 2017-07-18 Laurent Lejeune , Mario Christoudias , Raphael Sznitman

Long-form clinical videos are central to visual evidence-based decision-making, with growing importance for applications such as surgical robotics and related settings. However, current multimodal large language models typically process…

计算机视觉与模式识别 · 计算机科学 2026-02-17 Wenjie Li , Yujie Zhang , Haoran Sun , Xingqi He , Hongcheng Gao , Chenglong Ma , Ming Hu , Guankun Wang , Shiyi Yao , Renhao Yang , Hongliang Ren , Lei Wang , Junjun He , Yankai Jiang