中文
相关论文

相关论文: M$^3$AV: A Multimodal, Multigenre, and Multipurpos…

200 篇论文

A large number of annotated video-caption pairs are required for training video captioning models, resulting in high annotation costs. Active learning can be instrumental in reducing these annotation requirements. However, active learning…

计算机视觉与模式识别 · 计算机科学 2022-12-22 Gyanendra Das , Xavier Thomas , Anant Raj , Vikram Gupta

Many educational organizations are employing instructional video in their pedagogy, but there is limited understanding of the possible presentation styles. In practice, the presentation style of video lectures ranges from a direct recording…

计算机与社会 · 计算机科学 2020-01-01 Konstantinos Chorianopoulos

We introduce Replay, a collection of multi-view, multi-modal videos of humans interacting socially. Each scene is filmed in high production quality, from different viewpoints with several static cameras, as well as wearable action cameras,…

计算机视觉与模式识别 · 计算机科学 2023-07-25 Roman Shapovalov , Yanir Kleiman , Ignacio Rocco , David Novotny , Andrea Vedaldi , Changan Chen , Filippos Kokkinos , Ben Graham , Natalia Neverova

Multimodal learning is defined as learning over multiple heterogeneous input modalities such as video, audio, and text. In this work, we are concerned with understanding how models behave as the type of modalities differ between training…

机器学习 · 计算机科学 2023-04-12 Brandon McKinzie , Joseph Cheng , Vaishaal Shankar , Yinfei Yang , Jonathon Shlens , Alexander Toshev

Multi-modal learning, particularly among imaging and linguistic modalities, has made amazing strides in many high-level fundamental visual understanding problems, ranging from language grounding to dense event captioning. However, much of…

计算机视觉与模式识别 · 计算机科学 2019-10-28 Tanzila Rahman , Bicheng Xu , Leonid Sigal

We present a new large-scale multilingual video description dataset, VATEX, which contains over 41,250 videos and 825,000 captions in both English and Chinese. Among the captions, there are over 206,000 English-Chinese parallel translation…

计算机视觉与模式识别 · 计算机科学 2020-06-18 Xin Wang , Jiawei Wu , Junkun Chen , Lei Li , Yuan-Fang Wang , William Yang Wang

The way of understanding online higher education has greatly changed due to the worldwide pandemic situation. Teaching is undertaken remotely, and the faculty incorporate lecture audio recordings as part of the teaching material. This new…

计算与语言 · 计算机科学 2024-01-01 Oscar Sapena , Eva Onaindia

The increase in the availability of online videos has transformed the way we access information and knowledge. A growing number of individuals now prefer instructional videos as they offer a series of step-by-step procedures to accomplish…

计算与语言 · 计算机科学 2023-09-22 Deepak Gupta , Kush Attal , Dina Demner-Fushman

3D facial animation has attracted considerable attention due to its extensive applications in the multimedia field. Audio-driven 3D facial animation has been widely explored with promising results. However, multi-modal 3D facial animation,…

计算机视觉与模式识别 · 计算机科学 2024-10-11 Sijing Wu , Yunhao Li , Yichao Yan , Huiyu Duan , Ziwei Liu , Guangtao Zhai

Video aesthetic assessment, a vital area in multimedia computing, integrates computer vision with human cognition. Its progress is limited by the lack of standardized datasets and robust models, as the temporal dynamics of video and…

计算机视觉与模式识别 · 计算机科学 2025-11-14 Qianqian Qiao , DanDan Zheng , Yihang Bo , Bao Peng , Heng Huang , Longteng Jiang , Huaye Wang , Jingdong Chen , Jun Zhou , Xin Jin

Video descriptions are crucial for blind and low vision (BLV) users to access visual content. However, current artificial intelligence models for generating descriptions often fall short due to limitations in the quality of human…

计算机视觉与模式识别 · 计算机科学 2025-03-03 Chaoyu Li , Sid Padmanabhuni , Maryam Cheema , Hasti Seifi , Pooyan Fazli

As research on neural volumetric video reconstruction and compression flourishes, there is a need for diverse and realistic datasets, which can be used to develop and validate reconstruction and compression models. However, existing…

计算机视觉与模式识别 · 计算机科学 2026-04-23 Adrian Azzarelli , Ge Gao , Ho Man Kwan , Fan Zhang , Nantheera Anantrasirichai , Ollie Moolan-Feroze , David Bull

Referring audio-visual segmentation (RAVS) has recently seen significant advancements, yet challenges remain in integrating multimodal information and deeply understanding and reasoning about audiovisual content. To extend the boundaries of…

计算机视觉与模式识别 · 计算机科学 2025-08-01 Kaining Ying , Henghui Ding , Guangquan Jie , Yu-Gang Jiang

Multimodal embedding models aim to map heterogeneous inputs, such as text, images, videos, and audio, into a shared semantic space. However, existing methods and benchmarks remain largely limited to partial modality coverage, making it…

信息检索 · 计算机科学 2026-04-28 Haohang Huang , Xuan Lu , Mingyi Su , Xuan Zhang , Ziyan Jiang , Ping Nie , Kai Zou , Tomas Pfister , Wenhu Chen , Wei Zhang , Xiaoyu Shen , Rui Meng

Audio-language models (ALMs) generate linguistic descriptions of sound-producing events and scenes. Advances in dataset creation and computational power have led to significant progress in this domain. This paper surveys 69 datasets used to…

声音 · 计算机科学 2025-02-10 Gijs Wijngaard , Elia Formisano , Michele Esposito , Michel Dumontier

Understanding animal species from multimodal data poses an emerging challenge at the intersection of computer vision and ecology. While recent biological models, such as BioCLIP, have demonstrated strong alignment between images and textual…

计算机视觉与模式识别 · 计算机科学 2026-03-26 Risa Shinoda , Kaede Shiohara , Nakamasa Inoue , Kuniaki Saito , Hiroaki Santo , Fumio Okura

Vision and text have been fully explored in contemporary video-text foundational models, while other modalities such as audio and subtitles in videos have not received sufficient attention. In this paper, we resort to establish connections…

计算机视觉与模式识别 · 计算机科学 2023-10-10 Sihan Chen , Handong Li , Qunbo Wang , Zijia Zhao , Mingzhen Sun , Xinxin Zhu , Jing Liu

Multimodal learning, a rapidly evolving field in artificial intelligence, seeks to construct more versatile and robust systems by integrating and analyzing diverse types of data, including text, images, audio, and video. Inspired by the…

There has been a long-standing quest for a unified audio-visual-text model to enable various multimodal understanding tasks, which mimics the listening, seeing and reading process of human beings. Humans tends to represent knowledge using…

音频与语音处理 · 电气工程与系统科学 2024-02-22 Xianghu Yue , Xiaohai Tian , Lu Lu , Malu Zhang , Zhizheng Wu , Haizhou Li

The advancement of audio-language (AL) multimodal learning tasks has been significant in recent years. However, researchers face challenges due to the costly and time-consuming collection process of existing audio-language datasets, which…

音频与语音处理 · 电气工程与系统科学 2024-07-22 Xinhao Mei , Chutong Meng , Haohe Liu , Qiuqiang Kong , Tom Ko , Chengqi Zhao , Mark D. Plumbley , Yuexian Zou , Wenwu Wang