中文
相关论文

相关论文: Building Scalable Video Understanding Benchmarks t…

200 篇论文

We make available to the community a new dataset to support action-recognition research. This dataset is different from prior datasets in several key ways. It is significantly larger. It contains streaming video with long segments…

计算机视觉与模式识别 · 计算机科学 2015-11-19 Daniel Paul Barrett , Ran Xu , Haonan Yu , Jeffrey Mark Siskind

Recent advances in multimodal large language models (MLLMs) have expanded research in video understanding, primarily focusing on high-level tasks such as video captioning and question-answering. Meanwhile, a smaller body of work addresses…

计算机视觉与模式识别 · 计算机科学 2025-04-04 Ali Athar , Xueqing Deng , Liang-Chieh Chen

Learning commonsense reasoning from visual contexts and scenes in real-world is a crucial step toward advanced artificial intelligence. However, existing video reasoning benchmarks are still inadequate since they were mainly designed for…

计算机视觉与模式识别 · 计算机科学 2024-05-20 Andong Wang , Bo Wu , Sunli Chen , Zhenfang Chen , Haotian Guan , Wei-Ning Lee , Li Erran Li , Chuang Gan

Multimodal large language models (MLLMs) have enabled open-world visual understanding by injecting visual input as extra tokens into large language models (LLMs) as contexts. However, when the visual input changes from a single image to a…

计算机视觉与模式识别 · 计算机科学 2025-03-03 Xi Tang , Jihao Qiu , Lingxi Xie , Yunjie Tian , Jianbin Jiao , Qixiang Ye

Using massive datasets to train large-scale models has emerged as a dominant approach for broad generalization in natural language and vision applications. In reinforcement learning, however, a key challenge is that available data of…

机器学习 · 计算机科学 2022-12-07 David Venuto , Sherry Yang , Pieter Abbeel , Doina Precup , Igor Mordatch , Ofir Nachum

Video depth estimation has long been hindered by the scarcity of consistent and scalable ground truth data, leading to inconsistent and unreliable results. In this paper, we introduce Depth Any Video, a model that tackles the challenge…

计算机视觉与模式识别 · 计算机科学 2025-03-13 Honghui Yang , Di Huang , Wei Yin , Chunhua Shen , Haifeng Liu , Xiaofei He , Binbin Lin , Wanli Ouyang , Tong He

Long video understanding remains a fundamental challenge for multimodal large language models (MLLMs), particularly in tasks requiring precise temporal reasoning and event localization. Existing approaches typically adopt uniform frame…

计算机视觉与模式识别 · 计算机科学 2025-09-12 Chao Yuan , Yang Yang , Yehui Yang , Zach Cheng

Large-scale data collection is essential for developing personalized training data, mitigating the shortage of training data, and fine-tuning specialized models. However, creating high-quality datasets quickly and accurately remains a…

Generative image models have emerged as a promising technology to produce realistic images. Despite potential benefits, concerns grow about its misuse, particularly in generating deceptive images that could raise significant ethical, legal,…

计算机视觉与模式识别 · 计算机科学 2024-04-05 Jinbin Huang , Chen Chen , Aditi Mishra , Bum Chul Kwon , Zhicheng Liu , Chris Bryan

While most modern video understanding models operate on short-range clips, real-world videos are often several minutes long with semantically consistent segments of variable length. A common approach to process long videos is applying a…

计算机视觉与模式识别 · 计算机科学 2023-09-22 Mohamed Afham , Satya Narayan Shukla , Omid Poursaeed , Pengchuan Zhang , Ashish Shah , Sernam Lim

Large multimodal models (LMMs) are processing increasingly longer and richer inputs. Albeit the progress, few public benchmark is available to measure such development. To mitigate this gap, we introduce LongVideoBench, a question-answering…

计算机视觉与模式识别 · 计算机科学 2024-07-23 Haoning Wu , Dongxu Li , Bei Chen , Junnan Li

Audio Description (AD) provides linguistic descriptions of movies and allows visually impaired people to follow a movie along with their peers. Such descriptions are by design mainly visual and thus naturally form an interesting data source…

计算机视觉与模式识别 · 计算机科学 2016-05-13 Anna Rohrbach , Atousa Torabi , Marcus Rohrbach , Niket Tandon , Christopher Pal , Hugo Larochelle , Aaron Courville , Bernt Schiele

While current video generation focuses on text or image conditions, practical applications like video editing and vlogging often need to seamlessly connect separate clips. In our work, we introduce Video Connecting, an innovative task that…

计算机视觉与模式识别 · 计算机科学 2026-01-28 Zhiyu Yin , Zhipeng Liu , Kehai Chen , Lemao Liu , Jin Liu , Hong-Dong Li , Yang Xiang , Min Zhang

We describe a protocol to study text-to-video retrieval training with unlabeled videos, where we assume (i) no access to labels for any videos, i.e., no access to the set of ground-truth captions, but (ii) access to labeled images in the…

计算机视觉与模式识别 · 计算机科学 2024-04-29 Lucas Ventura , Cordelia Schmid , Gül Varol

Vision-language models such as CLIP have shown impressive capabilities in aligning images and text, but they often struggle with lengthy and detailed text descriptions due to pre-training on short and concise captions. We present FAST-GOAL…

人工智能 · 计算机科学 2026-05-27 Hyungyu Choi , Young Kyun Jang , Chanho Eom

Soccer analytics is attracting increasing interest in academia and industry, thanks to the availability of data that describe all the spatio-temporal events that occur in each match. These events (e.g., passes, shots, fouls) are collected…

计算机视觉与模式识别 · 计算机科学 2020-07-14 Danilo Sorano , Fabio Carrara , Paolo Cintia , Fabrizio Falchi , Luca Pappalardo

Learning text-video embeddings usually requires a dataset of video clips with manually provided captions. However, such datasets are expensive and time consuming to create and therefore difficult to obtain on a large scale. In this work, we…

计算机视觉与模式识别 · 计算机科学 2019-08-01 Antoine Miech , Dimitri Zhukov , Jean-Baptiste Alayrac , Makarand Tapaswi , Ivan Laptev , Josef Sivic

Our goal in this paper is the adaptation of image-text models for long video retrieval. Recent works have demonstrated state-of-the-art performance in video retrieval by adopting CLIP, effectively hitchhiking on the image-text…

计算机视觉与模式识别 · 计算机科学 2022-05-18 Max Bain , Arsha Nagrani , Gül Varol , Andrew Zisserman

Real-world instructional videos are long, noisy, and often contain extended background segments, repeated actions, and execution variability that do not correspond to meaningful procedural steps. We propose **REMAP**, an unsupervised…

计算机视觉与模式识别 · 计算机科学 2026-05-08 Soumyadeep Chandra , Kaushik Roy
‹ 上一页 1 8 9 10 下一页 ›