English
Related papers

Related papers: MCAD: Multimodal Context-Aware Audio Description G…

200 papers

Recent focus in video captioning has been on designing architectures that can consume both video and text modalities, and using large-scale video datasets with text transcripts for pre-training, such as HowTo100M. Though these approaches…

Computer Vision and Pattern Recognition · Computer Science 2023-06-23 Yuhan Shen , Linjie Yang , Longyin Wen , Haichao Yu , Ehsan Elhamifar , Heng Wang

Accurate audio quality estimation is essential for developing and evaluating audio generation, retrieval, and enhancement systems. Existing non-intrusive assessment models predict a single Mean Opinion Score (MOS) for speech, merging…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-13 Yi-Cheng Lin , Jia-Hung Chen , Hung-yi Lee

Computer-aided design (CAD) is the digital construction of 2D and 3D objects, and is central to a wide range of engineering and manufacturing applications like automobile and aviation. Despite its importance, CAD modeling remains largely a…

Graphics · Computer Science 2026-01-09 Prashant Govindarajan , Davide Baldelli , Jay Pathak , Quentin Fournier , Sarath Chandar

With the proliferation of various gaming technology, services, game styles, and platforms, multi-dimensional aesthetic assessment of the gaming contents is becoming more and more important for the gaming industry. Depending on the diverse…

Computer Vision and Pattern Recognition · Computer Science 2021-01-29 Zhenyu Lei , Yejing Xie , Suiyi Ling , Andreas Pastor , Junle Wang , Patrick Le Callet

Accurately understanding and deciding high-level meta-actions is essential for ensuring reliable and safe autonomous driving systems. While vision-language models (VLMs) have shown significant potential in various autonomous driving tasks,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-19 Yujin Wang , Quanfeng Liu , Zhengxin Jiang , Tianyi Wang , Junfeng Jiao , Hongqing Chu , Bingzhao Gao , Hong Chen

Video description involves the generation of the natural language description of actions, events, and objects in the video. There are various applications of video description by filling the gap between languages and vision for visually…

Computer Vision and Pattern Recognition · Computer Science 2020-12-01 Alok Singh , Thoudam Doren Singh , Sivaji Bandyopadhyay

Multimodal Large Language Models (MLLMs) have achieved impressive success in natural visual understanding, yet they consistently underperform in industrial anomaly detection (IAD). This is because MLLMs trained mostly on general web data…

Computer Vision and Pattern Recognition · Computer Science 2026-04-22 Xi Jiang , Yue Guo , Jian Li , Yong Liu , Bin-Bin Gao , Hanqiu Deng , Jun Liu , Heng Zhao , Chengjie Wang , Feng Zheng

As AI-generated video becomes increasingly pervasive across media platforms, the ability to reliably distinguish synthetic content from authentic footage has become both urgent and essential. Existing approaches have primarily treated this…

Computer Vision and Pattern Recognition · Computer Science 2025-06-19 Yifeng Gao , Yifan Ding , Hongyu Su , Juncheng Li , Yunhan Zhao , Lin Luo , Zixing Chen , Li Wang , Xin Wang , Yixu Wang , Xingjun Ma , Yu-Gang Jiang

Generating semantically and temporally aligned audio content in accordance with video input has become a focal point for researchers, particularly following the remarkable breakthrough in text-to-video generation. In this work, we aim to…

Sound · Computer Science 2025-03-12 Manjie Xu , Chenxing Li , Xinyi Tu , Yong Ren , Rilin Chen , Yu Gu , Wei Liang , Dong Yu

Real-time video commentary generation provides textual descriptions of ongoing events in videos. It supports accessibility and engagement in domains such as sports, esports, and livestreaming. Commentary generation involves two essential…

Computation and Language · Computer Science 2026-03-04 Anum Afzal , Yuki Saito , Hiroya Takamura , Katsuhito Sudoh , Shinnosuke Takamichi , Graham Neubig , Florian Matthes , Tatsuya Ishigaki

Conditional diffusion models are powerful generative models that can leverage various types of conditional information, such as class labels, segmentation masks, or text captions. However, in many real-world scenarios, conditional…

Computer Vision and Pattern Recognition · Computer Science 2025-02-19 Nicolas Dufour , Victor Besnier , Vicky Kalogeiton , David Picard

We consider the task of generating diverse and realistic videos guided by natural audio samples from a wide variety of semantic classes. For this task, the videos are required to be aligned both globally and temporally with the input audio:…

Machine Learning · Computer Science 2023-09-29 Guy Yariv , Itai Gat , Sagie Benaim , Lior Wolf , Idan Schwartz , Yossi Adi

Publishing open-source academic video recordings is an emergent and prevalent approach to sharing knowledge online. Such videos carry rich multimodal information including speech, the facial and body movements of the speakers, as well as…

Computation and Language · Computer Science 2024-06-05 Zhe Chen , Heyang Liu , Wenyi Yu , Guangzhi Sun , Hongcheng Liu , Ji Wu , Chao Zhang , Yu Wang , Yanfeng Wang

Anomaly detection (AD) is a fundamental task of critical importance across numerous domains. Current systems increasingly operate in rapidly evolving environments that generate diverse yet interconnected data modalities -- such as time…

Machine Learning · Computer Science 2025-12-02 Zhongyuan Wu , Jingyuan Wang , Zexuan Cheng , Yilong Zhou , Weizhi Wang , Juhua Pu , Chao Li , Changqing Ma

Dog emotion recognition plays a crucial role in enhancing human-animal interactions, veterinary care, and the development of automated systems for monitoring canine well-being. However, accurately interpreting dog emotions is challenging…

Computer Vision and Pattern Recognition · Computer Science 2026-02-18 Jinho Baek , Houwei Cao , Kate Blackwell

The intelligent dialogue system, aiming at communicating with humans harmoniously with natural language, is brilliant for promoting the advancement of human-machine interaction in the era of artificial intelligence. With the gradually…

Artificial Intelligence · Computer Science 2022-07-05 Hao Wang , Bin Guo , Yating Zeng , Yasan Ding , Chen Qiu , Ying Zhang , Lina Yao , Zhiwen Yu

With the rapid adoption of multimodal large language models (MLLMs) across diverse applications, there is a pressing need for task-centered, high-quality training data. A key limitation of current training datasets is their reliance on…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Xiaoyu Lin , Aniket Ghorpade , Hansheng Zhu , Justin Qiu , Dea Rrozhani , Monica Lama , Mick Yang , Zixuan Bian , Ruohan Ren , Alan B. Hong , Jiatao Gu , Chris Callison-Burch

Authors make their videos visually accessible by adding audio descriptions (AD), and auditorily accessible by adding closed captions (CC). However, creating AD and CC is challenging and tedious, especially for non-professional describers…

Human-Computer Interaction · Computer Science 2025-02-19 Xingyu "Bruce" Liu , Ruolin Wang , Dingzeyu Li , Xiang 'Anthony' Chen , Amy Pavel

Previous audio generation mainly focuses on specified sound classes such as speech or music, whose form and content are greatly restricted. In this paper, we go beyond specific audio generation by using natural language description as a…

Sound · Computer Science 2023-05-04 Guangwei Li , Xuenan Xu , Lingfeng Dai , Mengyue Wu , Kai Yu

Training-free video anomaly detection (VAD) has recently emerged as a scalable alternative to supervised approaches, yet existing methods largely rely on static prompting and geometry-agnostic feature fusion. As a result, anomaly inference…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Ali Zia , Usman Ali , Muhammad Umer Ramzan , Hamza Abid , Abdul Rehman , Wei Xiang