English
Related papers

Related papers: Multi-language Video Subtitle Dataset for Image-ba…

200 papers

Progress in AI is driven largely by the scale and quality of training data. Despite this, there is a deficit of empirical analysis examining the attributes of well-established datasets beyond text. In this work we conduct the largest and…

Script identification plays a vital role in applications that involve handwriting and document analysis within a multi-script and multi-lingual environment. Moreover, it exhibits a profound connection with human cognition. This paper…

Computer Vision and Pattern Recognition · Computer Science 2024-05-30 Miguel A. Ferrer , Abhijit Das , Moises Diaz , Aythami Morales , Cristina Carmona-Duarte , Umapada Pal

Natural scene text detection is a significant challenge in computer vision, with tremendous potential applications in multilingual, diverse, and complex text scenarios. We propose a multilingual text detection model to address the issues of…

Computer Vision and Pattern Recognition · Computer Science 2024-01-08 Tao Wang

Existing datasets for audio understanding primarily focus on single-turn interactions (i.e. audio captioning, audio question answering) for describing audio in natural language, thus limiting understanding audio via interactive dialogue. To…

Computation and Language · Computer Science 2024-04-12 Arushi Goel , Zhifeng Kong , Rafael Valle , Bryan Catanzaro

Joint understanding of video and language is an active research area with many applications. Prior work in this domain typically relies on learning text-video embeddings. One difficulty with this approach, however, is the lack of…

Computer Vision and Pattern Recognition · Computer Science 2020-01-17 Antoine Miech , Ivan Laptev , Josef Sivic

Currently, image-text-driven multi-modal deep learning models have demonstrated their outstanding potential in many fields. In practice, tasks centered around facial images have broad application prospects. This paper presents…

Computer Vision and Pattern Recognition · Computer Science 2024-07-15 Dawei Dai , YuTang Li , YingGe Liu , Mingming Jia , Zhang YuanHui , Guoyin Wang

This paper presents a task of audio-visual scene classification (SC) where input videos are classified into one of five real-life crowded scenes: 'Riot', 'Noise-Street', 'Firework-Event', 'Music-Event', and 'Sport-Atmosphere'. To this end,…

Computer Vision and Pattern Recognition · Computer Science 2021-12-20 Lam Pham , Dat Ngo , Phu X. Nguyen , Truong Hoang , Alexander Schindler

This paper introduces HarmonySet, a comprehensive dataset designed to advance video-music understanding. HarmonySet consists of 48,328 diverse video-music pairs, annotated with detailed information on rhythmic synchronization, emotional…

Computer Vision and Pattern Recognition · Computer Science 2025-03-05 Zitang Zhou , Ke Mei , Yu Lu , Tianyi Wang , Fengyun Rao

We describe a system for large-scale audiovisual translation and dubbing, which translates videos from one language to another. The source language's speech content is transcribed to text, translated, and automatically synthesized into…

Computer Vision and Pattern Recognition · Computer Science 2020-11-09 Yi Yang , Brendan Shillingford , Yannis Assael , Miaosen Wang , Wendi Liu , Yutian Chen , Yu Zhang , Eren Sezener , Luis C. Cobo , Misha Denil , Yusuf Aytar , Nando de Freitas

Multi-modal multi-party conversation (MMC) is a less studied yet important topic of research due to that it well fits real-world scenarios and thus potentially has more widely-used applications. Compared with the traditional multi-modal…

Computation and Language · Computer Science 2024-12-24 Yueqian Wang , Xiaojun Meng , Yuxuan Wang , Jianxin Liang , Qun Liu , Dongyan Zhao

With the rise of multimodal large language models, accurately extracting and understanding textual information from video content, referred to as video based optical character recognition (Video OCR), has become a crucial capability. This…

Computer Vision and Pattern Recognition · Computer Science 2024-12-31 Yulin Fei , Yuhui Gao , Xingyuan Xian , Xiaojin Zhang , Tao Wu , Wei Chen

This paper introduces a new multi-modal dataset for visual and audio-visual speech recognition. It includes face tracks from over 400 hours of TED and TEDx videos, along with the corresponding subtitles and word alignment boundaries. The…

Computer Vision and Pattern Recognition · Computer Science 2018-10-30 Triantafyllos Afouras , Joon Son Chung , Andrew Zisserman

Deep learning has shown remarkable progress in a wide range of problems. However, efficient training of such models requires large-scale datasets, and getting annotations for such datasets can be challenging and costly. In this work, we…

Multimedia · Computer Science 2021-10-14 Mohit Sharma , Raj Patra , Harshal Desai , Shruti Vyas , Yogesh Rawat , Rajiv Ratn Shah

Video description is the automatic generation of natural language sentences that describe the contents of a given video. It has applications in human-robot interaction, helping the visually impaired and video subtitling. The past few years…

Computer Vision and Pattern Recognition · Computer Science 2020-03-04 Nayyer Aafaq , Ajmal Mian , Wei Liu , Syed Zulqarnain Gilani , Mubarak Shah

Retrieving target videos based on text descriptions is a task of great practical value and has received increasing attention over the past few years. Despite recent progress, imperfect annotations in existing video retrieval datasets have…

Computer Vision and Pattern Recognition · Computer Science 2022-07-22 Zeyu Wang , Yu Wu , Karthik Narasimhan , Olga Russakovsky

This paper attacks the challenging problem of video retrieval by text. In such a retrieval paradigm, an end user searches for unlabeled videos by ad-hoc queries described exclusively in the form of a natural-language sentence, with no…

Computer Vision and Pattern Recognition · Computer Science 2021-02-19 Jianfeng Dong , Xirong Li , Chaoxi Xu , Xun Yang , Gang Yang , Xun Wang , Meng Wang

Massive web-crawled image-text datasets lay the foundation for recent progress in multimodal learning. These datasets are designed with the goal of training a model to do well on standard computer vision benchmarks, many of which, however,…

Computer Vision and Pattern Recognition · Computer Science 2025-09-16 Thao Nguyen , Matthew Wallingford , Sebastin Santy , Wei-Chiu Ma , Sewoong Oh , Ludwig Schmidt , Pang Wei Koh , Ranjay Krishna

In this paper, we introduce the ShopSign dataset, which is a newly developed natural scene text dataset of Chinese shop signs in street views. Although a few scene text datasets are already publicly available (e.g. ICDAR2015, COCO-Text),…

Computer Vision and Pattern Recognition · Computer Science 2019-03-26 Chongsheng Zhang , Guowen Peng , Yuefeng Tao , Feifei Fu , Wei Jiang , George Almpanidis , Ke Chen

Multilingual document and scene text understanding plays an important role in applications such as search, finance, and public services. However, most existing benchmarks focus on high-resource languages and fail to evaluate models in…

Computation and Language · Computer Science 2026-03-17 Pengfei Yue , Xingran Zhao , Juntao Chen , Peng Hou , Wang Longchao , Jianghang Lin , Shengchuan Zhang , Anxiang Zeng , Liujuan Cao

Memes are a widely popular tool for web users to express their thoughts using visual metaphors. Understanding memes requires recognizing and interpreting visual metaphors with respect to the text inside or around the meme, often while…

Computation and Language · Computer Science 2023-05-24 EunJeong Hwang , Vered Shwartz