English
Related papers

Related papers: QuerYD: A video dataset with high-quality text and…

200 papers

Text-to-Speech (TTS) synthesis faces the inherent challenge of producing multiple speech outputs with varying prosody given a single text input. While previous research has addressed this by predicting prosodic information from both text…

Computation and Language · Computer Science 2025-08-19 Shumin Que , Anton Ragni

Large datasets of paired images and text have become increasingly popular for learning generic representations for vision and vision-and-language tasks. Such datasets have been built by querying search engines or collecting HTML alt-text --…

Computer Vision and Pattern Recognition · Computer Science 2021-11-23 Karan Desai , Gaurav Kaul , Zubin Aysola , Justin Johnson

Text-to-video generation has surged in interest since Sora, yet open-source models still face a data bottleneck: there is no large, high-quality, easily obtainable video-text corpus. Existing public datasets typically require manual YouTube…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Timing Yang , Sucheng Ren , Alan Yuille , Feng Wang

We present HowSumm, a novel large-scale dataset for the task of query-focused multi-document summarization (qMDS), which targets the use-case of generating actionable instructions from a set of sources. This use-case is different from the…

Computation and Language · Computer Science 2021-10-12 Odellia Boni , Guy Feigenblat , Guy Lev , Michal Shmueli-Scheuer , Benjamin Sznajder , David Konopnicki

The largest source of sound events is web videos. Most videos lack sound event labels at segment level, however, a significant number of them do respond to text queries, from a match found using metadata by search engines. In this paper we…

Sound · Computer Science 2018-04-05 Rohan Badlani , Ankit Shah , Benjamin Elizalde , Anurag Kumar , Bhiksha Raj

A question answering system that in addition to providing an answer provides an explanation of the reasoning that leads to that answer has potential advantages in terms of debuggability, extensibility and trust. To this end, we propose QED,…

Computation and Language · Computer Science 2020-09-15 Matthew Lamm , Jennimaria Palomaki , Chris Alberti , Daniel Andor , Eunsol Choi , Livio Baldini Soares , Michael Collins

Humans use context and scene knowledge to easily localize moving objects in conditions of complex illumination changes, scene clutter and occlusions. In this paper, we present a method to leverage human knowledge in the form of annotated…

Computer Vision and Pattern Recognition · Computer Science 2016-04-20 Archith J. Bency , S. Karthikeyan , Carter De Leo , Santhoshkumar Sunderrajan , B. S. Manjunath

Despite progress in speech-to-video synthesis, existing methods often struggle to capture cross-individual dependencies and provide fine-grained control over reactive behaviors in dyadic settings. To address these challenges, we propose…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Dongwei Pan , Longwei Guo , Jiazhi Guan , Luying Huang , Yiding Li , Haojie Liu , Haocheng Feng , Wei He , Kaisiyuan Wang , Hang Zhou

Automatic transcriptions of consumer-generated multi-media content such as "Youtube" videos still exhibit high word error rates. Such data typically occupies a very broad domain, has been recorded in challenging conditions, with cheap…

Computation and Language · Computer Science 2017-12-08 Abhinav Gupta , Yajie Miao , Leonardo Neves , Florian Metze

Effective real-time data presentation is essential in small-group interactive contexts, where discussions evolve dynamically and presenters must adapt visualizations to shifting audience interests. However, most existing interactive…

Human-Computer Interaction · Computer Science 2025-10-17 Kentaro Takahira , Yuki Ueno

We present a novel human annotated dataset for evaluating the ability for visual-language models to generate both short and long descriptions for real-world video clips, termed DeVAn (Dense Video Annotation). The dataset contains 8.5K…

Computer Vision and Pattern Recognition · Computer Science 2024-08-12 Tingkai Liu , Yunzhe Tao , Haogeng Liu , Qihang Fan , Ding Zhou , Huaibo Huang , Ran He , Hongxia Yang

The proliferation of datasets across open data portals and enterprise data lakes presents an opportunity for deriving data-driven insights. Widely-used dataset search systems rely on keyword search over dataset metadata, including…

Databases · Computer Science 2025-12-19 Haoxiang Zhang , Yurong Liu , Aécio Santos , Wei-Lun Hung , Juliana Freire

Recent years have seen a significant increase in video content creation and consumption. Crafting engaging content requires the careful curation of both visual and audio elements. While visual cue curation, through techniques like optimal…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Chao Huang , Ruohan Gao , J. M. F. Tsang , Jan Kurcius , Cagdas Bilen , Chenliang Xu , Anurag Kumar , Sanjeel Parekh

With the exponential growth of video content, the need for automated video highlight detection to extract key moments or highlights from lengthy videos has become increasingly pressing. This technology has the potential to enhance user…

Computer Vision and Pattern Recognition · Computer Science 2025-05-16 Zahidul Islam , Sujoy Paul , Mrigank Rochan

Educational question generation (EQG) is a crucial component of intelligent educational systems, significantly aiding self-assessment, active learning, and personalized education. While EQG systems have emerged, existing datasets typically…

Computation and Language · Computer Science 2025-04-30 Mengxia Yu , Bang Nguyen , Olivia Zino , Meng Jiang

While progress has been made in the domain of video-language understanding, current state-of-the-art algorithms are still limited in their ability to understand videos at high levels of abstraction, such as news-oriented videos.…

Computer Vision and Pattern Recognition · Computer Science 2024-01-24 Shih-Han Chou , Matthew Kowal , Yasmin Niknam , Diana Moyano , Shayaan Mehdi , Richard Pito , Cheng Zhang , Ian Knopke , Sedef Akinli Kocak , Leonid Sigal , Yalda Mohsenzadeh

We present the Moments in Time Dataset, a large-scale human-annotated collection of one million short videos corresponding to dynamic events unfolding within three seconds. Modeling the spatial-audio-temporal dynamics even for actions…

Computer Vision and Pattern Recognition · Computer Science 2019-02-19 Mathew Monfort , Alex Andonian , Bolei Zhou , Kandan Ramakrishnan , Sarah Adel Bargal , Tom Yan , Lisa Brown , Quanfu Fan , Dan Gutfruend , Carl Vondrick , Aude Oliva

A domain shift exists between the large-scale, internet data used to train a Vision-Language Model (VLM) and the raw image streams collected by a robot. Existing adaptation strategies require the definition of a closed-set of classes, which…

Robotics · Computer Science 2025-02-27 Nicolas Harvey Chapman , Feras Dayoub , Will Browne , Christopher Lehnert

Data videos are a powerful medium for visual data based storytelling, combining animated, chart-centric visualizations with synchronized narration. Widely used in journalism, education, and public communication, they help audiences…

Artificial Intelligence · Computer Science 2026-04-29 Ridwan Mahbub , Syem Aziz , Mahir Ahmed , Shadikur Rahman , Mizanur Rahman , Shafiq Joty , Enamul Hoque

A new methodology to measure coded image/video quality using the just-noticeable-difference (JND) idea was proposed. Several small JND-based image/video quality datasets were released by the Media Communications Lab at the University of…