中文
相关论文

相关论文: USV: Towards Understanding the User-generated Shor…

200 篇论文

In this paper, we initiate an attempt of developing an end-to-end chat-centric video understanding system, coined as VideoChat. It integrates video foundation models and large language models via a learnable neural interface, excelling in…

计算机视觉与模式识别 · 计算机科学 2024-01-05 KunChang Li , Yinan He , Yi Wang , Yizhuo Li , Wenhai Wang , Ping Luo , Yali Wang , Limin Wang , Yu Qiao

We introduce TV show Retrieval (TVR), a new multimodal retrieval dataset. TVR requires systems to understand both videos and their associated subtitle (dialogue) texts, making it more realistic. The dataset contains 109K queries collected…

计算机视觉与模式识别 · 计算机科学 2020-08-19 Jie Lei , Licheng Yu , Tamara L. Berg , Mohit Bansal

Recent video generation models demonstrate impressive synthesis capabilities but remain limited by single-modality conditioning, constraining their holistic world understanding. This stems from insufficient cross-modal interaction and…

计算机视觉与模式识别 · 计算机科学 2025-12-09 Jiehui Huang , Yuechen Zhang , Xu He , Yuan Gao , Zhi Cen , Bin Xia , Yan Zhou , Xin Tao , Pengfei Wan , Jiaya Jia

We study multi-modal summarization for instructional videos, whose goal is to provide users an efficient way to learn skills in the form of text instructions and key video frames. We observe that existing benchmarks focus on generic…

计算机视觉与模式识别 · 计算机科学 2025-06-17 Yuan Zang , Hao Tan , Seunghyun Yoon , Franck Dernoncourt , Jiuxiang Gu , Kushal Kafle , Chen Sun , Trung Bui

The rapid growth of video on the internet has made searching for video content using natural language queries a significant challenge. Human-generated queries for video datasets `in the wild' vary a lot in terms of degree of specificity,…

计算机视觉与模式识别 · 计算机科学 2020-02-17 Yang Liu , Samuel Albanie , Arsha Nagrani , Andrew Zisserman

Long video understanding is a complex task that requires both spatial detail and temporal awareness. While Vision-Language Models (VLMs) obtain frame-level understanding capabilities through multi-frame input, they suffer from information…

计算机视觉与模式识别 · 计算机科学 2025-04-10 Ziyi Wang , Haoran Wu , Yiming Rong , Deyang Jiang , Yixin Zhang , Yunlong Zhao , Shuang Xu , Bo XU

The scalability of high-fidelity video diffusion models (VDMs) is constrained by two key sources of redundancy: the quadratic complexity of global spatio-temporal attention and the computational overhead of long iterative denoising…

计算机视觉与模式识别 · 计算机科学 2025-12-08 Xinjian Wu , Hongmei Wang , Yuan Zhou , Qinglin Lu

Short-form video poses new challenges to the quality assessment of user-generated content (UGC) due to its complex generation pipeline, rapid content variation, and mixed distortions. To address this challenge, we propose an end-to-end…

图像与视频处理 · 电气工程与系统科学 2026-05-20 Xinyi Wang , Angeliki Katsenou , Junxiao Shen , David Bull

Previous studies on question generation from videos have mostly focused on generating questions about common objects and attributes and hence are not entity-centric. In this work, we focus on the generation of entity-centric…

计算机视觉与模式识别 · 计算机科学 2024-10-15 Arpan Phukan , Manish Gupta , Asif Ekbal

Video classification has advanced tremendously over the recent years. A large part of the improvements in video classification had to do with the work done by the image classification community and the use of deep convolutional networks…

计算机视觉与模式识别 · 计算机科学 2015-05-26 Balakrishnan Varadarajan , George Toderici , Sudheendra Vijayanarasimhan , Apostol Natsev

Boosted by Multi-modal Large Language Models (MLLMs), text-guided universal segmentation models for the image and video domains have made rapid progress recently. However, these methods are often developed separately for specific domains,…

计算机视觉与模式识别 · 计算机科学 2024-12-19 Cong Wei , Yujie Zhong , Haoxian Tan , Yingsen Zeng , Yong Liu , Zheng Zhao , Yujiu Yang

This paper addresses automatic summarization and search in visual data comprising of videos, live streams and image collections in a unified manner. In particular, we propose a framework for multi-faceted summarization which extracts…

计算机视觉与模式识别 · 计算机科学 2017-04-06 Anurag Sahoo , Vishal Kaushal , Khoshrav Doctor , Suyash Shetty , Rishabh Iyer , Ganesh Ramakrishnan

Despite recent advances in Video Large Language Models (VideoLLMs), effectively understanding long-form videos remains a significant challenge. Perceiving lengthy videos containing thousands of frames poses substantial computational burden.…

计算机视觉与模式识别 · 计算机科学 2025-09-03 Linli Yao , Haoning Wu , Kun Ouyang , Yuanxing Zhang , Caiming Xiong , Bei Chen , Xu Sun , Junnan Li

Much of the delivery of University education is now by synchronous or asynchronous video. For students, one of the challenges is managing the sheer volume of such video material as video presentations of taught material are difficult to…

多媒体 · 计算机科学 2021-06-28 Hyowon Lee , Mingming Liu , Michael Scriney , Alan F. Smeaton

Automated surgical workflow analysis is crucial for education, research, and clinical decision-making, but the lack of annotated datasets hinders the development of accurate and comprehensive workflow analysis solutions. We introduce a…

计算机视觉与模式识别 · 计算机科学 2025-03-17 David Gastager , Ghazal Ghazaei , Constantin Patsch

Recent advances in text-to-video generation have achieved impressive performance on short clips, yet evaluating long-form generation under complex textual inputs remains a significant challenge. In response to this challenge, we present…

计算机视觉与模式识别 · 计算机科学 2026-05-29 Xiangqing Zheng , Chengyue Wu , Kehai Chen , Min Zhang

Long-form video content constitutes a significant portion of internet traffic, making automated video summarization an essential research problem. However, existing video summarization datasets are notably limited in their size,…

计算机视觉与模式识别 · 计算机科学 2024-04-05 Dawit Mureja Argaw , Seunghyun Yoon , Fabian Caba Heilbron , Hanieh Deilamsalehy , Trung Bui , Zhaowen Wang , Franck Dernoncourt , Joon Son Chung

Existing large video-language models (LVLMs) struggle to comprehend long videos correctly due to limited context. To address this problem, fine-tuning long-context LVLMs and employing GPT-based agents have emerged as promising solutions.…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Yongdong Luo , Xiawu Zheng , Guilin Li , Shukang Yin , Haojia Lin , Chaoyou Fu , Jinfa Huang , Jiayi Ji , Fei Chao , Jiebo Luo , Rongrong Ji

Despite the number of currently available datasets on video question answering, there still remains a need for a dataset involving multi-step and non-factoid answers. Moreover, relying on video transcripts remains an under-explored topic.…

计算与语言 · 计算机科学 2020-06-02 Anthony Colas , Seokhwan Kim , Franck Dernoncourt , Siddhesh Gupte , Daisy Zhe Wang , Doo Soon Kim

We introduce the Lecture Video Visual Objects (LVVO) dataset, a new benchmark for visual object detection in educational video content. The dataset consists of 4,000 frames extracted from 245 lecture videos spanning biology, computer…

计算机视觉与模式识别 · 计算机科学 2025-06-18 Dipayan Biswas , Shishir Shah , Jaspal Subhlok