中文
相关论文

相关论文: VidChapters-7M: Video Chapters at Scale

200 篇论文

Large-scale web-crawled datasets are fundamental for the success of pre-training vision-language models, such as CLIP. However, the inherent noise and potential irrelevance of web-crawled AltTexts pose challenges in achieving precise…

计算机视觉与模式识别 · 计算机科学 2024-03-15 Zhengfeng Lai , Haotian Zhang , Bowen Zhang , Wentao Wu , Haoping Bai , Aleksei Timofeev , Xianzhi Du , Zhe Gan , Jiulong Shan , Chen-Nee Chuah , Yinfei Yang , Meng Cao

Currently, no large-scale training data is available for the task of scientific paper summarization. In this paper, we propose a novel method that automatically generates summaries for scientific papers, by utilizing videos of talks at…

计算与语言 · 计算机科学 2019-06-14 Guy Lev , Michal Shmueli-Scheuer , Jonathan Herzig , Achiya Jerbi , David Konopnicki

This paper presents a new large-scale dataset for recognition and temporal localization of human actions collected from Web videos. We refer to it as HACS (Human Action Clips and Segments). We leverage both consensus and disagreement among…

计算机视觉与模式识别 · 计算机科学 2019-09-05 Hang Zhao , Antonio Torralba , Lorenzo Torresani , Zhicheng Yan

Short-video platforms show an increasing impact on people's daily lives nowadays, with billions of active users spending plenty of time each day. The interactions between users and online platforms give rise to many scientific problems…

多媒体 · 计算机科学 2025-02-11 Yu Shang , Chen Gao , Nian Li , Yong Li

Existing automatic captioning methods for visual content face challenges such as lack of detail, content hallucination, and poor instruction following. In this work, we propose VisualFactChecker (VFC), a flexible training-free pipeline that…

计算机视觉与模式识别 · 计算机科学 2024-05-01 Yunhao Ge , Xiaohui Zeng , Jacob Samuel Huffman , Tsung-Yi Lin , Ming-Yu Liu , Yin Cui

We present REM, a framework for segmenting a wide range of concepts in video that can be described through natural language. Our method leverages the universal visual-language mapping learned by video diffusion models on Internet-scale data…

计算机视觉与模式识别 · 计算机科学 2025-08-08 Anurag Bagchi , Zhipeng Bao , Yu-Xiong Wang , Pavel Tokmakov , Martial Hebert

This work proposes TimeChat, a time-sensitive multimodal large language model specifically designed for long video understanding. Our model incorporates two key architectural contributions: (1) a timestamp-aware frame encoder that binds…

计算机视觉与模式识别 · 计算机科学 2024-03-29 Shuhuai Ren , Linli Yao , Shicheng Li , Xu Sun , Lu Hou

Existing works on semantic segmentation typically consider a small number of labels, ranging from tens to a few hundreds. With a large number of labels, training and evaluation of such task become extremely challenging due to correlation…

计算机视觉与模式识别 · 计算机科学 2018-08-21 Yufei Wang , Zhe Lin , Xiaohui Shen , Jianming Zhang , Scott Cohen

In this paper, we propose VidLA, an approach for video-language alignment at scale. There are two major limitations of previous video-language alignment approaches. First, they do not capture both short-range and long-range temporal…

计算机视觉与模式识别 · 计算机科学 2024-03-25 Mamshad Nayeem Rizve , Fan Fei , Jayakrishnan Unnikrishnan , Son Tran , Benjamin Z. Yao , Belinda Zeng , Mubarak Shah , Trishul Chilimbi

Automatic video summarization is still an unsolved problem due to several challenges. The currently available datasets either have very short videos or have few long videos of only a particular type. We introduce a new benchmarking video…

计算机视觉与模式识别 · 计算机科学 2021-01-27 Vishal Kaushal , Suraj Kothawade , Anshul Tomar , Rishabh Iyer , Ganesh Ramakrishnan

For research results to be comparable, it is important to have common datasets for experimentation and evaluation. The size of such datasets, however, can be an obstacle to their use. The Vimeo Creative Commons Collection (V3C) is a video…

多媒体 · 计算机科学 2021-05-05 Luca Rossetto , Klaus Schoeffmann , Abraham Bernstein

The task of describing video content in natural language is commonly referred to as video captioning. Unlike conventional video captions, which are typically brief and widely available, long-form paragraph descriptions in natural language…

计算机视觉与模式识别 · 计算机科学 2025-07-08 Mihai Masala , Marius Leordeanu

Temporal localization remains an important challenge in video understanding. In this work, we present our solution to the 3rd YouTube-8M Video Understanding Challenge organized by Google Research. Participants were required to build a…

计算机视觉与模式识别 · 计算机科学 2019-11-19 Lijun Zhang , Srinath Nizampatnam , Ahana Gangopadhyay , Marcos V. Conde

Recent advances in Large Multi-modal Models (LMMs) are primarily focused on offline video understanding. Instead, streaming video understanding poses great challenges to recent models due to its time-sensitive, omni-modal and interactive…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Shenghao Fu , Qize Yang , Yuan-Ming Li , Yi-Xing Peng , Kun-Yu Lin , Xihan Wei , Jian-Fang Hu , Xiaohua Xie , Wei-Shi Zheng

Video classification problem has been studied many years. The success of Convolutional Neural Networks (CNN) in image recognition tasks gives a powerful incentive for researchers to create more advanced video classification approaches. As…

计算机视觉与模式识别 · 计算机科学 2017-06-15 Manuk Akopyan , Eshsou Khashba

The recognition of human activities is one of the key problems in video understanding. Action recognition is challenging even for specific categories of videos, such as sports, that contain only a small set of actions. Interestingly, sports…

多媒体 · 计算机科学 2017-09-28 Rahul Anand Sharma , Pramod Sankar K , CV Jawahar

Instruction-based video editing allows effective and interactive editing of videos using only instructions without extra inputs such as masks or attributes. However, collecting high-quality training triplets (source video, edited video,…

计算机视觉与模式识别 · 计算机科学 2025-07-14 Yuhui Wu , Liyi Chen , Ruibin Li , Shihao Wang , Chenxi Xie , Lei Zhang

We introduce CameraBench, a large-scale dataset and benchmark designed to assess and improve camera motion understanding. CameraBench consists of ~3,000 diverse internet videos, annotated by experts through a rigorous multi-stage quality…

Automatic video summarization is still an unsolved problem due to several challenges. We take steps towards making automatic video summarization more realistic by addressing them. Firstly, the currently available datasets either have very…

计算机视觉与模式识别 · 计算机科学 2020-08-26 Vishal Kaushal , Suraj Kothawade , Rishabh Iyer , Ganesh Ramakrishnan

Computer-use agents can operate computers and automate laborious tasks, but despite recent rapid progress, they still lag behind human users, especially when tasks require domain-specific procedural knowledge about particular applications,…

计算机视觉与模式识别 · 计算机科学 2025-11-07 Yujian Liu , Ze Wang , Hao Chen , Ximeng Sun , Xiaodong Yu , Jialian Wu , Jiang Liu , Emad Barsoum , Zicheng Liu , Shiyu Chang