English
Related papers

Related papers: QuerYD: A video dataset with high-quality text and…

200 papers

Video description is the automatic generation of natural language sentences that describe the contents of a given video. It has applications in human-robot interaction, helping the visually impaired and video subtitling. The past few years…

Computer Vision and Pattern Recognition · Computer Science 2020-03-04 Nayyer Aafaq , Ajmal Mian , Wei Liu , Syed Zulqarnain Gilani , Mubarak Shah

Recent years have witnessed an increasing amount of dialogue/conversation on the web especially on social media. That inspires the development of dialogue-based retrieval, in which retrieving videos based on dialogue is of increasing…

Information Retrieval · Computer Science 2023-03-30 Chenyang Lyu , Manh-Duy Nguyen , Van-Tu Ninh , Liting Zhou , Cathal Gurrin , Jennifer Foster

Narrated ''how-to'' videos have emerged as a promising data source for a wide range of learning problems, from learning visual representations to training robot policies. However, this data is extremely noisy, as the narrations do not…

Computer Vision and Pattern Recognition · Computer Science 2023-07-20 Kumar Ashutosh , Rohit Girdhar , Lorenzo Torresani , Kristen Grauman

Live streaming plays a major role in today's digital platforms, supporting entertainment, education, social media, etc. However, research in this field is limited by the lack of large, publicly available datasets that capture real-time…

Retrieving target videos based on text descriptions is a task of great practical value and has received increasing attention over the past few years. Despite recent progress, imperfect annotations in existing video retrieval datasets have…

Computer Vision and Pattern Recognition · Computer Science 2022-07-22 Zeyu Wang , Yu Wu , Karthik Narasimhan , Olga Russakovsky

In this study, we introduce YODAS (YouTube-Oriented Dataset for Audio and Speech), a large-scale, multilingual dataset comprising currently over 500k hours of speech data in more than 100 languages, sourced from both labeled and unlabeled…

Computation and Language · Computer Science 2024-06-04 Xinjian Li , Shinnosuke Takamichi , Takaaki Saeki , William Chen , Sayaka Shiota , Shinji Watanabe

This paper introduces a new multi-modal dataset for visual and audio-visual speech recognition. It includes face tracks from over 400 hours of TED and TEDx videos, along with the corresponding subtitles and word alignment boundaries. The…

Computer Vision and Pattern Recognition · Computer Science 2018-10-30 Triantafyllos Afouras , Joon Son Chung , Andrew Zisserman

While existing video benchmarks largely consider specialized downstream tasks like retrieval or question-answering (QA), contemporary multimodal AI systems must be capable of well-rounded common-sense reasoning akin to human visual…

Computer Vision and Pattern Recognition · Computer Science 2024-06-17 Kate Sanders , Benjamin Van Durme

Multimodal counterfactual reasoning is a vital yet challenging ability for AI systems. It involves predicting the outcomes of hypothetical circumstances based on vision and language inputs, which enables AI models to learn from failures and…

Computer Vision and Pattern Recognition · Computer Science 2023-11-06 Te-Lin Wu , Zi-Yi Dou , Qingyuan Hu , Yu Hou , Nischal Reddy Chandra , Marjorie Freedman , Ralph M. Weischedel , Nanyun Peng

Educational recommenders have received much less attention in comparison to e-commerce and entertainment-related recommenders, even though efficient intelligent tutors have great potential to improve learning gains. One of the main…

Information Retrieval · Computer Science 2021-09-15 Sahan Bulathwela , Maria Perez-Ortiz , Erik Novak , Emine Yilmaz , John Shawe-Taylor

Multi-modal learning, particularly among imaging and linguistic modalities, has made amazing strides in many high-level fundamental visual understanding problems, ranging from language grounding to dense event captioning. However, much of…

Computer Vision and Pattern Recognition · Computer Science 2019-10-28 Tanzila Rahman , Bicheng Xu , Leonid Sigal

With the emergence of e-learning and personalised education, the production and distribution of digital educational resources have boomed. Video lectures have now become one of the primary modalities to impart knowledge to masses in the…

Computers and Society · Computer Science 2020-11-05 Sahan Bulathwela , Maria Perez-Ortiz , Emine Yilmaz , John Shawe-Taylor

Question-answering (QA) is a natural approach for humans to understand a piece of music audio. However, for machines, accessing a large-scale dataset covering diverse aspects of music is crucial, yet challenging, due to the scarcity of…

Sound · Computer Science 2025-08-28 Zhihao Ouyang , Ju-Chiang Wang , Daiyu Zhang , Bin Chen , Shangjie Li , Quan Lin

Composed video retrieval is a challenging task that strives to retrieve a target video based on a query video and a textual description detailing specific modifications. Standard retrieval frameworks typically struggle to handle the…

Computer Vision and Pattern Recognition · Computer Science 2025-08-20 Omkar Thawakar , Dmitry Demidov , Ritesh Thawkar , Rao Muhammad Anwer , Mubarak Shah , Fahad Shahbaz Khan , Salman Khan

Most existing datasets for sound event recognition (SER) are relatively small and/or domain-specific, with the exception of AudioSet, based on over 2M tracks from YouTube videos and encompassing over 500 sound classes. However, AudioSet is…

Sound · Computer Science 2022-04-26 Eduardo Fonseca , Xavier Favory , Jordi Pons , Frederic Font , Xavier Serra

Deep learning has shown remarkable progress in a wide range of problems. However, efficient training of such models requires large-scale datasets, and getting annotations for such datasets can be challenging and costly. In this work, we…

Multimedia · Computer Science 2021-10-14 Mohit Sharma , Raj Patra , Harshal Desai , Shruti Vyas , Yogesh Rawat , Rajiv Ratn Shah

Knowing how to construct text-based Search Queries (SQs) for use in Search Engines (SEs) such as Google or Wikipedia has become a fundamental skill. Though much data are available through such SEs, most structured datasets live outside…

Human-Computer Interaction · Computer Science 2022-08-03 Leonardo Christino , Martha D. Ferreira , Fernando V. Paulovich

The rapid growth of video on the internet has made searching for video content using natural language queries a significant challenge. Human-generated queries for video datasets `in the wild' vary a lot in terms of degree of specificity,…

Computer Vision and Pattern Recognition · Computer Science 2020-02-17 Yang Liu , Samuel Albanie , Arsha Nagrani , Andrew Zisserman

Descriptive video service (DVS) provides linguistic descriptions of movies and allows visually impaired people to follow a movie along with their peers. Such descriptions are by design mainly visual and thus naturally form an interesting…

Computer Vision and Pattern Recognition · Computer Science 2015-01-13 Anna Rohrbach , Marcus Rohrbach , Niket Tandon , Bernt Schiele

Recently, the AI community has made significant strides in developing powerful foundation models, driven by large-scale multimodal datasets. However, for audio representation learning, existing datasets suffer from limitations in the…

Sound · Computer Science 2024-09-10 Luoyi Sun , Xuenan Xu , Mengyue Wu , Weidi Xie