English
Related papers

Related papers: SurgTEMP: Temporal-Aware Surgical Video Question A…

200 papers

Video temporal understanding is crucial for multimodal large language models (MLLMs) to reason over events in videos. Despite recent advances in general video understanding, current MLLMs still struggle with fine-grained temporal reasoning.…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Fuwen Luo , Shengfeng Lou , Chi Chen , Ziyue Wang , Chenliang Li , Weizhou Shen , Jiyue Guo , Peng Li , Ming Yan , Ji Zhang , Fei Huang , Yang Liu

Automatic surgical phase recognition is one of the key technologies to support Video-Based Assessment (VBA) systems for surgical education. Utilizing temporal information is crucial for surgical phase recognition, hence various recent…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Bokai Zhang , Mohammad Hasan Sarhan , Bharti Goel , Svetlana Petculescu , Amer Ghanem

Medical Visual Question Answering (VQA) is an important challenge, as it would lead to faster and more accurate diagnoses and treatment decisions. Most existing methods approach it as a multi-class classification problem, which restricts…

Computer Vision and Pattern Recognition · Computer Science 2023-07-25 Tom van Sonsbeek , Mohammad Mahdi Derakhshani , Ivona Najdenkoska , Cees G. M. Snoek , Marcel Worring

Temporal understanding in autonomous driving (AD) remains a significant challenge, even for recent state-of-the-art (SoTA) Vision-Language Models (VLMs). Prior work has introduced datasets and benchmarks aimed at improving temporal…

Computer Vision and Pattern Recognition · Computer Science 2025-12-18 Kevin Cannons , Saeed Ranjbar Alvar , Mohammad Asiful Hossain , Ahmad Rezaei , Mohsen Gholami , Alireza Heidarikhazaei , Zhou Weimin , Yong Zhang , Mohammad Akbari

In the video-language domain, recent works in leveraging zero-shot Large Language Model-based reasoning for video understanding have become competitive challengers to previous end-to-end models. However, long video understanding presents…

Computer Vision and Pattern Recognition · Computer Science 2024-10-08 Ruotong Liao , Max Erler , Huiyu Wang , Guangyao Zhai , Gengyuan Zhang , Yunpu Ma , Volker Tresp

Large vision-language models (VLMs) have evolved from general-purpose applications to specialized use cases such as in the clinical domain, demonstrating potential for decision support in radiology. One promising application is assisting…

Computer Vision and Pattern Recognition · Computer Science 2026-02-25 Zhifan Jiang , Dong Yang , Vishwesh Nath , Abhijeet Parida , Nishad P. Kulkarni , Ziyue Xu , Daguang Xu , Syed Muhammad Anwar , Holger R. Roth , Marius George Linguraru

The temporal answering grounding in the video (TAGV) is a new task naturally derived from temporal sentence grounding in the video (TSGV). Given an untrimmed video and a text question, this task aims at locating the matching span from the…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Bin Li , Yixuan Weng , Bin Sun , Shutao Li

Surgical workflow analysis is essential in robot-assisted surgeries, yet the long duration of such procedures poses significant challenges for comprehensive video analysis. Recent approaches have predominantly relied on transformer models;…

Computer Vision and Pattern Recognition · Computer Science 2025-06-27 Haoyang Wu , Tsun-Hsuan Wang , Mathias Lechner , Ramin Hasani , Jennifer A. Eckhoff , Paul Pak , Ozanan R. Meireles , Guy Rosman , Yutong Ban , Daniela Rus

How to efficiently utilize temporal information to recover videos in a consistent way is the main issue for video inpainting problems. Conventional 2D CNNs have achieved good performance on image inpainting but often lead to temporally…

Computer Vision and Pattern Recognition · Computer Science 2025-05-14 Ya-Liang Chang , Zhe Yu Liu , Kuan-Ying Lee , Winston Hsu

The explosive growth of videos on streaming media platforms has underscored the urgent need for effective video quality assessment (VQA) algorithms to monitor and perceptually optimize the quality of streaming videos. However, VQA remains…

Computer Vision and Pattern Recognition · Computer Science 2024-08-27 Qihang Ge , Wei Sun , Yu Zhang , Yunhao Li , Zhongpeng Ji , Fengyu Sun , Shangling Jui , Xiongkuo Min , Guangtao Zhai

Video Temporal Grounding (VTG) strives to accurately pinpoint event timestamps in a specific video using linguistic queries, significantly impacting downstream tasks like video browsing and editing. Unlike traditional task-specific models,…

Computer Vision and Pattern Recognition · Computer Science 2025-02-04 Yongxin Guo , Jingyu Liu , Mingda Li , Dingxin Cheng , Xiaoying Tang , Dianbo Sui , Qingbin Liu , Xi Chen , Kevin Zhao

Automating crash video analysis is essential to leverage the growing availability of driving video data for traffic safety research and accountability attribution in autonomous driving. Crash video analysis is a challenging multitask…

Computer Vision and Pattern Recognition · Computer Science 2025-12-23 Kaidi Liang , Ke Li , Xianbiao Hu , Ruwen Qin

Video Large Language Models (Video LLMs) have achieved significant success by adopting the paradigm of large-scale pre-training followed by supervised fine-tuning (SFT). However, existing approaches struggle with temporal reasoning due to…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Shicheng Li , Lei Li , Kun Ouyang , Shuhuai Ren , Yuanxin Liu , Yuanxing Zhang , Fuzheng Zhang , Lingpeng Kong , Qi Liu , Xu Sun

While foundation models have advanced surgical video analysis, current approaches rely predominantly on pixel-level reconstruction objectives that waste model capacity on low-level visual details, such as smoke, specular reflections, and…

In the past year, video-based large language models (Video LLMs) have achieved impressive progress, particularly in their ability to process long videos through extremely extended context lengths. However, this comes at the cost of…

Computer Vision and Pattern Recognition · Computer Science 2026-05-04 Shangkun Sun , Ruyang Liu , Haoran Tang , Yixiao Ge , Haibo Lu , Wei Gao , Jiankun Yang , Chen Li

In the construction sector, workers often endure prolonged periods of high-intensity physical work and prolonged use of tools, resulting in injuries and illnesses primarily linked to postural ergonomic risks, a longstanding predominant…

Computer Vision and Pattern Recognition · Computer Science 2024-12-31 Chao Fan , Qipei Mei , Xiaonan Wang , Xinming Li

Large vision-language models (VLMs) typically process hundreds or thousands of visual tokens per image or video frame, incurring quadratic attention cost and substantial redundancy. Existing token reduction methods often ignore the textual…

Computer Vision and Pattern Recognition · Computer Science 2025-12-24 Kaitong Cai , Jusheng Zhang , Jing Yang , Yijia Fan , Pengtao Xie , Jian Wang , Keze Wang

Video Large Language Models (VideoLLMs) extend the capabilities of vision-language models to spatiotemporal inputs, enabling tasks such as video question answering (VideoQA). Despite recent advances in VideoLLMs, their internal mechanisms…

Computer Vision and Pattern Recognition · Computer Science 2026-03-04 Minji Kim , Taekyung Kim , Bohyung Han

Neuro-symbolic approaches to long-form video question answering (LVQA) have demonstrated significant accuracy improvements by grounding temporal reasoning in formal verification. However, existing methods incur prohibitive latency…

Computer Vision and Pattern Recognition · Computer Science 2026-03-02 Shawn Liang , Sahil Shah , Chengwei Zhou , SP Sharan , Harsh Goel , Arnab Sanyal , Sandeep Chinchali , Gourav Datta

Video large language models (Video-LLMs) have made strong progress in general video understanding, but their ability to maintain temporal object consistency remains underexplored. Existing benchmarks often emphasize event recognition,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Junzhe Chen , Siyuan Meng , Yuxi Chen , Man Zhao , Wenyao Gui , Xiaojie Guo
‹ Prev 1 8 9 10 Next ›