中文
相关论文

相关论文: M-VAD Names: a Dataset for Video Captioning with N…

200 篇论文

Story video-text alignment, a core task in computational story understanding, aims to align video clips with corresponding sentences in their descriptions. However, progress on the task has been held back by the scarcity of manually…

计算与语言 · 计算机科学 2024-10-04 Yidan Sun , Jianfei Yu , Boyang Li

Neural networks trained on datasets such as ImageNet have led to major advances in visual object classification. One obstacle that prevents networks from reasoning more deeply about complex scenes and situations, and from integrating visual…

Dense video captioning is a task of localizing interesting events from an untrimmed video and producing textual description (captions) for each localized event. Most of the previous works in dense video captioning are solely based on visual…

计算机视觉与模式识别 · 计算机科学 2020-05-07 Vladimir Iashin , Esa Rahtu

Recently, the AI community has made significant strides in developing powerful foundation models, driven by large-scale multimodal datasets. However, for audio representation learning, existing datasets suffer from limitations in the…

声音 · 计算机科学 2024-09-10 Luoyi Sun , Xuenan Xu , Mengyue Wu , Weidi Xie

Video captioning is a challenging task since it requires generating sentences describing various diverse and complex videos. Existing video captioning models lack adequate visual representation due to the neglect of the existence of gaps…

计算机视觉与模式识别 · 计算机科学 2021-10-14 Mingkang Tang , Zhanyu Wang , Zhenhua Liu , Fengyun Rao , Dian Li , Xiu Li

Generating Audio Description (AD) for movies is a challenging task that requires fine-grained visual understanding and an awareness of the characters and their names. Currently, visual language models for AD generation are limited by a lack…

计算机视觉与模式识别 · 计算机科学 2024-04-23 Tengda Han , Max Bain , Arsha Nagrani , Gül Varol , Weidi Xie , Andrew Zisserman

Identity documents recognition is an important sub-field of document analysis, which deals with tasks of robust document detection, type identification, text fields recognition, as well as identity fraud prevention and document authenticity…

Automatic generation of video descriptions in natural language, also called video captioning, aims to understand the visual content of the video and produce a natural language sentence depicting the objects and actions in the scene. This…

计算机视觉与模式识别 · 计算机科学 2020-12-15 Begum Citamak , Ozan Caglayan , Menekse Kuyu , Erkut Erdem , Aykut Erdem , Pranava Madhyastha , Lucia Specia

In this paper, we present the Multi-view Extended Videos with Identities (MEVID) dataset for large-scale, video person re-identification (ReID) in the wild. To our knowledge, MEVID represents the most-varied video person ReID dataset,…

Multi-modal retrieval is an important problem for many applications, such as recommendation and search. Current benchmarks and even datasets are often manually constructed and consist of mostly clean samples where all modalities are…

计算机视觉与模式识别 · 计算机科学 2022-10-21 Laura Hanu , James Thewlis , Yuki M. Asano , Christian Rupprecht

This paper introduces a video dataset of spatio-temporally localized Atomic Visual Actions (AVA). The AVA dataset densely annotates 80 atomic visual actions in 430 15-minute video clips, where actions are localized in space and time,…

Metaphors are a common communication tool used in our day-to-day life. The detection and generation of metaphors in textual form have been studied extensively but metaphors in other forms have been under-explored. Recent studies have shown…

计算机视觉与模式识别 · 计算机科学 2024-10-03 Abisek Rajakumar Kalarani , Pushpak Bhattacharyya , Sumit Shekhar

In real-world applications, e.g. law enforcement and video retrieval, one often needs to search a certain person in long videos with just one portrait. This is much more challenging than the conventional settings for person…

计算机视觉与模式识别 · 计算机科学 2018-07-30 Qingqiu Huang , Wentao Liu , Dahua Lin

Person identification in the wild is very challenging due to great variation in poses, face quality, clothes, makeup and so on. Traditional research, such as face recognition, person re-identification, and speaker recognition, often focuses…

计算机视觉与模式识别 · 计算机科学 2019-04-23 Yuanliu Liu , Bo Peng , Peipei Shi , He Yan , Yong Zhou , Bing Han , Yi Zheng , Chao Lin , Jianbin Jiang , Yin Fan , Tingwei Gao , Ganwen Wang , Jian Liu , Xiangju Lu , Danming Xie

Universal video understanding requires modeling fine-grained visual and audio information over time in diverse real-world scenarios. However, the performance of existing models is primarily constrained by video-instruction data that…

计算机视觉与模式识别 · 计算机科学 2026-02-16 Yunheng Li , Hengrui Zhang , Meng-Hao Guo , Wenzhao Gao , Shaoyong Jia , Shaohui Jiao , Qibin Hou , Ming-Ming Cheng

Movie Audio Description (AD) aims to narrate visual content during dialogue-free segments, particularly benefiting blind and visually impaired (BVI) audiences. Compared with general video captioning, AD demands plot-relevant narration with…

计算机视觉与模式识别 · 计算机科学 2025-05-13 Xiaojun Ye , Chun Wang , Yiren Song , Sheng Zhou , Liangcheng Li , Jiajun Bu

Vision and text have been fully explored in contemporary video-text foundational models, while other modalities such as audio and subtitles in videos have not received sufficient attention. In this paper, we resort to establish connections…

计算机视觉与模式识别 · 计算机科学 2023-10-10 Sihan Chen , Handong Li , Qunbo Wang , Zijia Zhao , Mingzhen Sun , Xinxin Zhu , Jing Liu

A more robust and holistic language-video representation is the key to pushing video understanding forward. Despite the improvement in training strategies, the quality of the language-video dataset is less attention to. The current plain…

多媒体 · 计算机科学 2024-06-21 Yuchen Yang , Yingxuan Duan

Video captioning which automatically translates video clips into natural language sentences is a very important task in computer vision. By virtue of recent deep learning technologies, e.g., convolutional neural networks (CNNs) and…

计算机视觉与模式识别 · 计算机科学 2016-11-18 Junbo Wang , Wei Wang , Yan Huang , Liang Wang , Tieniu Tan

Data visualization captions help readers understand the purpose of a visualization and are crucial for individuals with visual impairments. The prevalence of poor figure captions and the successful application of deep learning approaches to…

计算机视觉与模式识别 · 计算机科学 2022-07-18 Anita Mahinpei , Zona Kostic , Chris Tanner