English
Related papers

Related papers: M$^3$AV: A Multimodal, Multigenre, and Multipurpos…

200 papers

The convergence of 3D geometric perception and video synthesis has created an unprecedented demand for large-scale video data that is rich in both semantic and spatio-temporal information. While existing datasets have advanced either 3D…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Yunnan Wang , Kecheng Zheng , Jianyuan Wang , Minghao Chen , David Novotny , Christian Rupprecht , Yinghao Xu , Xing Zhu , Wenjun Zeng , Xin Jin , Yujun Shen

Video editing tools are widely used nowadays for digital design. Although the demand for these tools is high, the prior knowledge required makes it difficult for novices to get started. Systems that could follow natural language…

Computer Vision and Pattern Recognition · Computer Science 2022-03-22 Tsu-Jui Fu , Xin Eric Wang , Scott T. Grafton , Miguel P. Eckstein , William Yang Wang

An increasing number of datasets contain multiple views, such as video, sound and automatic captions. A basic challenge in representation learning is how to leverage multiple views to learn better representations. This is further…

Machine Learning · Computer Science 2019-03-04 Nils Holzenberger , Shruti Palaskar , Pranava Madhyastha , Florian Metze , Raman Arora

Recent methods for visual question answering rely on large-scale annotated datasets. Manual annotation of questions and answers for videos, however, is tedious, expensive and prevents scalability. In this work, we propose to avoid manual…

Computer Vision and Pattern Recognition · Computer Science 2022-05-12 Antoine Yang , Antoine Miech , Josef Sivic , Ivan Laptev , Cordelia Schmid

Vision Language Models (VLMs) have shown strong performance on multimodal reasoning tasks, yet most evaluations focus on short videos and assume unconstrained computational resources. In industrial settings such as pharmaceutical content…

Computer Vision and Pattern Recognition · Computer Science 2026-01-09 Suyash Mishra , Qiang Li , Srikanth Patil , Satyanarayan Pati , Baddu Narendra

We present HourVideo, a benchmark dataset for hour-long video-language understanding. Our dataset consists of a novel task suite comprising summarization, perception (recall, tracking), visual reasoning (spatial, temporal, predictive,…

Computer Vision and Pattern Recognition · Computer Science 2024-11-08 Keshigeyan Chandrasegaran , Agrim Gupta , Lea M. Hadzic , Taran Kota , Jimming He , Cristóbal Eyzaguirre , Zane Durante , Manling Li , Jiajun Wu , Li Fei-Fei

Automatically generating a natural language sentence to describe the content of an input video is a very challenging problem. It is an essential multimodal task in which auditory and visual contents are equally important. Although audio…

Computer Vision and Pattern Recognition · Computer Science 2018-12-10 Yapeng Tian , Chenxiao Guan , Justin Goodman , Marc Moore , Chenliang Xu

Evaluating the quality of slide-based multimedia instruction is challenging. Existing methods like manual assessment, reference-based metrics, and large language model evaluators face limitations in scalability, context capture, or bias. In…

Computation and Language · Computer Science 2025-05-06 Joy Lim Jia Yin , Daniel Zhang-Li , Jifan Yu , Haoxuan Li , Shangqing Tu , Yuanchun Wang , Zhiyuan Liu , Huiqin Liu , Lei Hou , Juanzi Li , Bin Xu

Visual saliency prediction for omnidirectional videos (ODVs) has shown great significance and necessity for omnidirectional videos to help ODV coding, ODV transmission, ODV rendering, etc.. However, most studies only consider visual…

Computer Vision and Pattern Recognition · Computer Science 2023-11-10 Yuxin Zhu , Xilei Zhu , Huiyu Duan , Jie Li , Kaiwei Zhang , Yucheng Zhu , Li Chen , Xiongkuo Min , Guangtao Zhai

In this work, we introduce a dataset of video annotated with high quality natural language phrases describing the visual content in a given segment of time. Our dataset is based on the Descriptive Video Service (DVS) that is now encoded on…

Computer Vision and Pattern Recognition · Computer Science 2015-03-04 Atousa Torabi , Christopher Pal , Hugo Larochelle , Aaron Courville

We present the Multiview Extended Video with Activities (MEVA) dataset, a new and very-large-scale dataset for human activity recognition. Existing security datasets either focus on activity counts by aggregating public video disseminated…

Computer Vision and Pattern Recognition · Computer Science 2020-12-03 Kellie Corona , Katie Osterdahl , Roderic Collins , Anthony Hoogs

Temporal video segmentation and classification have been advanced greatly by public benchmarks in recent years. However, such research still mainly focuses on human actions, failing to describe videos in a holistic view. In addition,…

Computer Vision and Pattern Recognition · Computer Science 2022-12-12 Jie Jiang , Zhimin Li , Jiangfeng Xiong , Rongwei Quan , Qinglin Lu , Wei Liu

In this paper, we present a novel dataset captured using a VR headset to record conversations between participants within a physics simulator (AI2-THOR). Our primary objective is to extend the field of co-speech gesture generation by…

Computer Vision and Pattern Recognition · Computer Science 2024-10-02 Anna Deichler , Jim O'Regan , Jonas Beskow

Large Language Models (LLMs) have shown immense potential in education, automating tasks like quiz generation and content summarization. However, generating effective presentation slides introduces unique challenges due to the complexity of…

Artificial Intelligence · Computer Science 2025-11-14 Eric Xie , Danielle Waterfield , Michael Kennedy , Aidong Zhang

Multi-modal Entity Linking (MEL) is a fundamental component for various downstream tasks. However, existing MEL datasets suffer from small scale, scarcity of topic types and limited coverage of tasks, making them incapable of effectively…

Information Retrieval · Computer Science 2024-10-25 Fang Wang , Shenglin Yin , Xiaoying Bai , Minghao Hu , Tianwei Yan , Yi Liang

Recently, the AI community has made significant strides in developing powerful foundation models, driven by large-scale multimodal datasets. However, for audio representation learning, existing datasets suffer from limitations in the…

Sound · Computer Science 2024-09-10 Luoyi Sun , Xuenan Xu , Mengyue Wu , Weidi Xie

Everyday news coverage has shifted from traditional broadcasts towards a wide range of presentation formats such as first-hand, unedited video footage. Datasets that reflect the diverse array of multimodal, multilingual news sources…

Information Retrieval · Computer Science 2023-07-07 Kate Sanders , David Etter , Reno Kriz , Benjamin Van Durme

Traditional lecture videos offer flexibility but lack mechanisms for real-time clarification, forcing learners to search externally when confusion arises. Recent advances in large language models and neural avatars provide new opportunities…

Computer Vision and Pattern Recognition · Computer Science 2025-12-25 Md Zabirul Islam , Md Motaleb Hossen Manik , Ge Wang

Thanks to the emerging of foundation models, the large language and vision models are integrated to acquire the multimodal ability of visual captioning, question answering, etc. Although existing multimodal models present impressive…

Computer Vision and Pattern Recognition · Computer Science 2023-12-29 Bo Zhao , Boya Wu , Muyang He , Tiejun Huang

We introduce TalkVerse, a large-scale, open corpus for single-person, audio-driven talking video generation designed to enable fair, reproducible comparison across methods. While current state-of-the-art systems rely on closed data or…

Computer Vision and Pattern Recognition · Computer Science 2025-12-18 Zhenzhi Wang , Jian Wang , Ke Ma , Dahua Lin , Bing Zhou