English
Related papers

Related papers: Multi-language Video Subtitle Dataset for Image-ba…

200 papers

Textual overlays are often used in social media videos as people who watch them without the sound would otherwise miss essential information conveyed in the audio stream. This is why extraction of those overlays can serve as an important…

Computer Vision and Pattern Recognition · Computer Science 2018-05-02 Adam Słucki , Tomasz Trzcinski , Adam Bielski , Paweł Cyrta

Curation methods for massive vision-language datasets trade off between dataset size and quality. However, even the highest quality of available curated captions are far too short to capture the rich visual detail in an image. To show the…

Computer Vision and Pattern Recognition · Computer Science 2024-06-18 Jack Urbanek , Florian Bordes , Pietro Astolfi , Mary Williamson , Vasu Sharma , Adriana Romero-Soriano

An outstanding image-text retrieval model depends on high-quality labeled data. While the builders of existing image-text retrieval datasets strive to ensure that the caption matches the linked image, they cannot prevent a caption from…

Computer Vision and Pattern Recognition · Computer Science 2022-09-02 Xu Yan , Chunhui Ai , Ziqiang Cao , Min Cao , Sujian Li , Wenjie Li , Guohong Fu

Multi-modal retrieval is an important problem for many applications, such as recommendation and search. Current benchmarks and even datasets are often manually constructed and consist of mostly clean samples where all modalities are…

Computer Vision and Pattern Recognition · Computer Science 2022-10-21 Laura Hanu , James Thewlis , Yuki M. Asano , Christian Rupprecht

Video transcript summarization is a fundamental task for video understanding. Conventional approaches for transcript summarization are usually built upon the summarization data for written language such as news articles, while the domain…

Computation and Language · Computer Science 2021-07-16 Tengchao Lv , Lei Cui , Momcilo Vasilijevic , Furu Wei

Automatic generation of video descriptions in natural language, also called video captioning, aims to understand the visual content of the video and produce a natural language sentence depicting the objects and actions in the scene. This…

Computer Vision and Pattern Recognition · Computer Science 2020-12-15 Begum Citamak , Ozan Caglayan , Menekse Kuyu , Erkut Erdem , Aykut Erdem , Pranava Madhyastha , Lucia Specia

A key requirement for leveraging supervised deep learning methods is the availability of large, labeled datasets. Unfortunately, in the context of RGB-D scene understanding, very little data is available -- current datasets cover a small…

Computer Vision and Pattern Recognition · Computer Science 2017-04-12 Angela Dai , Angel X. Chang , Manolis Savva , Maciej Halber , Thomas Funkhouser , Matthias Nießner

Scene text recognition is essential in many applications, including automated translation, information retrieval, driving assistance, and enhancing accessibility for individuals with visual impairments. Much research has been done to…

Computer Vision and Pattern Recognition · Computer Science 2024-05-21 Fadila Wendigoundi Douamba , Jianjun Song , Ling Fu , Yuliang Liu , Xiang Bai

Neural networks trained on datasets such as ImageNet have led to major advances in visual object classification. One obstacle that prevents networks from reasoning more deeply about complex scenes and situations, and from integrating visual…

Nowadays document analysis and recognition remain challenging tasks. However, only a few datasets designed for text detection (TD) and optical character recognition (OCR) problems exist. In this paper we present Distorted Document Images…

Computer Vision and Pattern Recognition · Computer Science 2021-09-20 Ilia Zharikov , Filipp Nikitin , Ilia Vasiliev , Vladimir Dokholyan

As computers have become efficient at understanding visual information and transforming it into a written representation, research interest in tasks like automatic image captioning has seen a significant leap over the last few years. While…

Computation and Language · Computer Science 2022-05-31 Mohammad Faiyaz Khan , S. M. Sadiq-Ur-Rahman Shifath , Md Saiful Islam

While progress has been made in the domain of video-language understanding, current state-of-the-art algorithms are still limited in their ability to understand videos at high levels of abstraction, such as news-oriented videos.…

Computer Vision and Pattern Recognition · Computer Science 2024-01-24 Shih-Han Chou , Matthew Kowal , Yasmin Niknam , Diana Moyano , Shayaan Mehdi , Richard Pito , Cheng Zhang , Ian Knopke , Sedef Akinli Kocak , Leonid Sigal , Yalda Mohsenzadeh

Deepfakes have become a growing concern in recent years, prompting researchers to develop benchmark datasets and detection algorithms to tackle the issue. However, existing datasets suffer from significant drawbacks that hamper their…

Computers and Society · Computer Science 2023-09-07 Beomsang Cho , Binh M. Le , Jiwon Kim , Simon Woo , Shahroz Tariq , Alsharif Abuadbba , Kristen Moore

The surge of audiovisual content on streaming platforms and social media has heightened the demand for accurate and accessible subtitles. However, existing subtitle generation methods primarily speech-based transcription or OCR-based…

Machine Learning · Computer Science 2025-10-29 Arpita Kundu , Joyita Chakraborty , Anindita Desarkar , Aritra Sen , Srushti Anil Patil , Vishwanathan Raman

The quality of the video dataset (image quality, resolution, and fine-grained caption) greatly influences the performance of the video generation model. The growing demand for video applications sets higher requirements for high-quality…

Computer Vision and Pattern Recognition · Computer Science 2025-06-17 Zhucun Xue , Jiangning Zhang , Teng Hu , Haoyang He , Yinan Chen , Yuxuan Cai , Yabiao Wang , Chengjie Wang , Yong Liu , Xiangtai Li , Dacheng Tao

The increasing use of machine learning models has amplified the demand for high-quality, large-scale multimodal datasets. However, the availability of such datasets, especially those combining acoustic, visual and textual data, remains…

Multimedia · Computer Science 2025-09-09 Jorge E. León , Miguel Carrasco

A lot of research has been devoted to identity documents analysis and recognition on mobile devices. However, no publicly available datasets designed for this particular problem currently exist. There are a few datasets which are useful for…

Computer Vision and Pattern Recognition · Computer Science 2020-02-12 Vladimir V. Arlazarov , Konstantin Bulatov , Timofey Chernov , Vladimir L. Arlazarov

Translating visual data into natural language is essential for machines to understand the world and interact with humans. In this work, a comprehensive study is conducted on video paragraph captioning, with the goal to generate…

Computer Vision and Pattern Recognition · Computer Science 2022-03-15 Qinyu Li , Tengpeng Li , Hanli Wang , Chang Wen Chen

Significant progress has been made recently on challenging tasks in automatic sign language understanding, such as sign language recognition, translation and production. However, these works have focused on datasets with relatively few…

Computer Vision and Pattern Recognition · Computer Science 2022-09-07 Alvaro Budria , Laia Tarres , Gerard I. Gallego , Francesc Moreno-Noguer , Jordi Torres , Xavier Giro-i-Nieto

Recently, the AI community has made significant strides in developing powerful foundation models, driven by large-scale multimodal datasets. However, for audio representation learning, existing datasets suffer from limitations in the…

Sound · Computer Science 2024-09-10 Luoyi Sun , Xuenan Xu , Mengyue Wu , Weidi Xie