English
Related papers

Related papers: QuerYD: A video dataset with high-quality text and…

200 papers

State-of-the-art conversational agents have advanced significantly in conjunction with the use of large transformer-based language models. However, even with these advancements, conversational agents still lack the ability to produce…

Computation and Language · Computer Science 2020-10-21 Sashank Santhanam , Wei Ping , Raul Puri , Mohammad Shoeybi , Mostofa Patwary , Bryan Catanzaro

Video storytelling is engaging multimedia content that utilizes video and its accompanying narration to attract the audience, where a key challenge is creating narrations for recorded visual scenes. Previous studies on dense video…

Multimedia · Computer Science 2024-12-31 Dingyi Yang , Chunru Zhan , Ziheng Wang , Biao Wang , Tiezheng Ge , Bo Zheng , Qin Jin

Deepfakes represent a growing concern across domains such as disinformation, fraud, and non-consensual media. In particular, the rise of video conference and identity-driven attacks in high-stakes scenarios--such as impostor hiring--demands…

Computer Vision and Pattern Recognition · Computer Science 2026-04-08 Sarah Barrington , Maty Bohacek , Hany Farid

Most natural videos contain numerous events. For example, in a video of a "man playing a piano", the video might also contain "another man dancing" or "a crowd clapping". We introduce the task of dense-captioning events, which involves both…

Computer Vision and Pattern Recognition · Computer Science 2017-05-03 Ranjay Krishna , Kenji Hata , Frederic Ren , Li Fei-Fei , Juan Carlos Niebles

Recent years have witnessed a resurgence of interest in video summarization. However, one of the main obstacles to the research on video summarization is the user subjectivity - users have various preferences over the summaries. The…

Computer Vision and Pattern Recognition · Computer Science 2017-07-18 Aidean Sharghi , Jacob S. Laurel , Boqing Gong

360{\deg} videos in recent years have experienced booming development. Compared to traditional videos, 360{\deg} videos are featured with uncertain user behaviors, bringing opportunities as well as challenges. Datasets are necessary for…

Multimedia · Computer Science 2022-08-09 Yili Jin , Junhua Liu , Fangxin Wang , Shuguang Cui

An ideal description for a given video should fix its gaze on salient and representative content, which is capable of distinguishing this video from others. However, the distribution of different words is unbalanced in video captioning…

Computer Vision and Pattern Recognition · Computer Science 2019-01-03 Jiarong Dong , Ke Gao , Xiaokai Chen , Junbo Guo , Juan Cao , Yongdong Zhang

Multimodal summarization with multimodal output (MSMO) has emerged as a promising research direction. Nonetheless, numerous limitations exist within existing public MSMO datasets, including insufficient maintenance, data inaccessibility,…

Computer Vision and Pattern Recognition · Computer Science 2023-11-21 Jielin Qiu , Jiacheng Zhu , William Han , Aditesh Kumar , Karthik Mittal , Claire Jin , Zhengyuan Yang , Linjie Li , Jianfeng Wang , Ding Zhao , Bo Li , Lijuan Wang

Existing datasets for audio understanding primarily focus on single-turn interactions (i.e. audio captioning, audio question answering) for describing audio in natural language, thus limiting understanding audio via interactive dialogue. To…

Computation and Language · Computer Science 2024-04-12 Arushi Goel , Zhifeng Kong , Rafael Valle , Bryan Catanzaro

This paper presents a novel approach for temporal and semantic segmentation of edited videos into meaningful segments, from the point of view of the storytelling structure. The objective is to decompose a long video into more manageable…

Computer Vision and Pattern Recognition · Computer Science 2016-11-11 Lorenzo Baraldi , Costantino Grana , Rita Cucchiara

We introduce a new large-scale data set of video URLs with densely-sampled object bounding box annotations called YouTube-BoundingBoxes (YT-BB). The data set consists of approximately 380,000 video segments about 19s long, automatically…

Computer Vision and Pattern Recognition · Computer Science 2017-03-28 Esteban Real , Jonathon Shlens , Stefano Mazzocchi , Xin Pan , Vincent Vanhoucke

The rapid development of large-scale models has catalyzed significant breakthroughs in the digital human domain. These advanced methodologies offer high-fidelity solutions for avatar driving and rendering, leading academia to focus on the…

Computer Vision and Pattern Recognition · Computer Science 2025-07-15 Youliang Zhang , Zhaoyang Li , Duomin Wang , Jiahe Zhang , Deyu Zhou , Zixin Yin , Xili Dai , Gang Yu , Xiu Li

Visual content memorability has intrigued the scientific community for decades, with applications ranging widely, from understanding nuanced aspects of human memory to enhancing content design. A significant challenge in progressing the…

Computer Vision and Pattern Recognition · Computer Science 2025-11-27 Sree Bhattacharyya , Yaman Kumar Singla , Sudhir Yarram , Somesh Kumar Singh , Harini S , James Z. Wang

The detection and localization of highly realistic deepfake audio-visual content are challenging even for the most advanced state-of-the-art methods. While most of the research efforts in this domain are focused on detecting high-quality…

Computer Vision and Pattern Recognition · Computer Science 2024-07-30 Zhixi Cai , Shreya Ghosh , Aman Pankaj Adatia , Munawar Hayat , Abhinav Dhall , Tom Gedeon , Kalin Stefanov

We introduce RoadSocial, a large-scale, diverse VideoQA dataset tailored for generic road event understanding from social media narratives. Unlike existing datasets limited by regional bias, viewpoint bias and expert-driven annotations,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-28 Chirag Parikh , Deepti Rawat , Rakshitha R. T. , Tathagata Ghosh , Ravi Kiran Sarvadevabhatla

Podcasts have become daily companions for half a billion users. Given the enormous amount of podcast content available, highlights provide a valuable signal that helps viewers get the gist of an episode and decide if they want to invest in…

Computation and Language · Computer Science 2025-09-09 Younghan Park , Anuj Diwan , David Harwath , Eunsol Choi

The goal of this paper is speaker diarisation of videos collected 'in the wild'. We make three key contributions. First, we propose an automatic audio-visual diarisation method for YouTube videos. Our method consists of active speaker…

Sound · Computer Science 2021-08-17 Joon Son Chung , Jaesung Huh , Arsha Nagrani , Triantafyllos Afouras , Andrew Zisserman

Speech activity detection (or endpointing) is an important processing step for applications such as speech recognition, language identification and speaker diarization. Both audio- and vision-based approaches have been used for this task in…

We introduce PointOdyssey, a large-scale synthetic dataset, and data generation framework, for the training and evaluation of long-term fine-grained tracking algorithms. Our goal is to advance the state-of-the-art by placing emphasis on…

Computer Vision and Pattern Recognition · Computer Science 2023-07-28 Yang Zheng , Adam W. Harley , Bokui Shen , Gordon Wetzstein , Leonidas J. Guibas

Audio event detection is a widely studied audio processing task, with applications ranging from self-driving cars to healthcare. In-the-wild datasets such as Audioset have propelled research in this field. However, many efforts typically…

Audio and Speech Processing · Electrical Eng. & Systems 2023-02-16 Rajat Hebbar , Digbalay Bose , Krishna Somandepalli , Veena Vijai , Shrikanth Narayanan
‹ Prev 1 3 4 5 6 7 10 Next ›