English
Related papers

Related papers: MCAD: Multimodal Context-Aware Audio Description G…

200 papers

Audio context determines which sound components and sources are relevant and which can be perceived as irrelevant (noise) by listeners. For example, traffic noise is informative in urban surveillance but noise for a phone call at the same…

Sound · Computer Science 2026-05-22 Diep Luong , Konstantinos Drossos , Mikko Heikkinen , Tuomas Virtanen

Existing automated dubbing methods are usually designed for Professionally Generated Content (PGC) production, which requires massive training data and training time to learn a person-specific audio-video mapping. In this paper, we…

Computer Vision and Pattern Recognition · Computer Science 2023-09-04 Linsen Song , Wayne Wu , Chaoyou Fu , Chen Change Loy , Ran He

Vision-language models (VLMs) have recently shown strong potential in soccer video understanding. However, given the high complexity of soccer videos due to large viewpoint variations, rapid shot transitions, and cluttered scenes, it…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Ismael Elsharkawi , Ahmed Sait , Silvio Giancola , Bernard Ghanem , Hossam Sharara , Abdelrahman Eldesokey

Recently, video captioning has been attracting an increasing amount of interest, due to its potential for improving accessibility and information retrieval. While existing methods rely on different kinds of visual features and model…

Computer Vision and Pattern Recognition · Computer Science 2016-12-02 Xiang Long , Chuang Gan , Gerard de Melo

Generative models have shown significant achievements in audio generation tasks. However, existing models struggle with complex and detailed prompts, leading to potential performance degradation. We hypothesize that this problem stems from…

Automated soccer commentary generation has evolved from template-based systems to advanced neural architectures, aiming to produce real-time descriptions of sports events. While frameworks like SoccerNet-Caption laid foundational work,…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Chidaksh Ravuru

Accurate dialogue description in audiovisual video captioning is crucial for downstream understanding and generation tasks. However, existing models generally struggle to produce faithful dialogue descriptions within audiovisual captions.…

Computation and Language · Computer Science 2026-01-28 Xinlong Chen , Weihong Lin , Jingyun Hua , Linli Yao , Yue Ding , Bozhou Li , Bohan Zeng , Yang Shi , Qiang Liu , Yuanxing Zhang , Pengfei Wan , Liang Wang , Tieniu Tan

How does textual representation of audio relate to the Large Language Model's (LLMs) learning about the audio world? This research investigates the extent to which LLMs can be prompted to generate audio, despite their primary training in…

The aesthetic quality assessment task is crucial for developing a human-aligned quantitative evaluation system for AIGC. However, its inherently complex nature, spanning visual perception, cognition, and emotion, poses fundamental…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Henglin Liu , Nisha Huang , Chang Liu , Jiangpeng Yan , Huijuan Huang , Jixuan Ying , Tong-Yee Lee , Pengfei Wan , Xiangyang Ji

Generating realistic audio for human actions is important for many applications, such as creating sound effects for films or virtual reality games. Existing approaches implicitly assume total correspondence between the video and audio…

Computer Vision and Pattern Recognition · Computer Science 2024-07-26 Changan Chen , Puyuan Peng , Ami Baid , Zihui Xue , Wei-Ning Hsu , David Harwath , Kristen Grauman

Anomaly detection (AD) is a machine learning task that identifies anomalies by learning patterns from normal training data. In many real-world scenarios, anomalies vary in severity, from minor anomalies with little risk to severe…

Machine Learning · Computer Science 2024-11-25 Tri Cao , Minh-Huy Trinh , Ailin Deng , Quoc-Nam Nguyen , Khoa Duong , Ngai-Man Cheung , Bryan Hooi

Generating coherent and useful image/video scenes from a free-form textual description is technically a very difficult problem to handle. Textual description of the same scene can vary greatly from person to person, or sometimes even for…

Computer Vision and Pattern Recognition · Computer Science 2020-12-01 Faria Huq , Nafees Ahmed , Anindya Iqbal

While large audio-language models have advanced open-ended audio understanding, they still fall short of nuanced human-level comprehension. This gap persists largely because current benchmarks, limited by data annotations and evaluation…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-12 Yadong Niu , Tianzi Wang , Heinrich Dinkel , Xingwei Sun , Jiahao Zhou , Gang Li , Jizhong Liu , Xunying Liu , Junbo Zhang , Jian Luan

Audio-visual deepfakes have reached a level of realism that makes perceptual detection unreliable, threatening media integrity and biometric security. While multimodal detection has shown promise, most approaches are binary classification…

Computer Vision and Pattern Recognition · Computer Science 2026-04-30 Wasim Ahmad , Wei Zhang , Xuerui Mao

The performance evaluation remains a complex challenge in audio separation, and existing evaluation metrics are often misaligned with human perception, course-grained, relying on ground truth signals. On the other hand, subjective listening…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-28 Helin Wang , Bowen Shi , Andros Tjandra , John Hoffman , Yi-Chiao Wu , Apoorv Vyas , Najim Dehak , Ann Lee , Wei-Ning Hsu

Automated audio captioning aims at generating textual descriptions for an audio clip. To evaluate the quality of generated audio captions, previous works directly adopt image captioning metrics like SPICE and CIDEr, without justifying their…

Sound · Computer Science 2022-01-28 Zelin Zhou , Zhiling Zhang , Xuenan Xu , Zeyu Xie , Mengyue Wu , Kenny Q. Zhu

Automated Audio Captioning is a multimodal task that aims to convert audio content into natural language. The assessment of audio captioning systems is typically based on quantitative metrics applied to text data. Previous studies have…

Sound · Computer Science 2024-03-28 Gijs Wijngaard , Elia Formisano , Bruno L. Giordano , Michel Dumontier

Automatic Video Dubbing (AVD) generates speech aligned with lip motion and facial emotion from scripts. Recent research focuses on modeling multimodal context to enhance prosody expressiveness but overlooks two key issues: 1) Multiscale…

Multimedia · Computer Science 2025-01-03 Yuan Zhao , Rui Liu , Gaoxiang Cong

Generating accurate descriptions for online fashion items is important not only for enhancing customers' shopping experiences, but also for the increase of online sales. Besides the need of correctly presenting the attributes of items, the…

Computer Vision and Pattern Recognition · Computer Science 2022-04-26 Xuewen Yang , Heming Zhang , Di Jin , Yingru Liu , Chi-Hao Wu , Jianchao Tan , Dongliang Xie , Jue Wang , Xin Wang

Lesion detection, symptom tracking, and visual explainability are central to real-world medical image analysis, yet current medical Vision-Language Models (VLMs) still lack mechanisms that translate their broad knowledge into clinically…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Woohyeon Park , Jaeik Kim , Sunghwan Steve Cho , Pa Hong , Wookyoung Jeong , Yoojin Nam , Namjoon Kim , Ginny Y. Wong , Ka Chun Cheung , Jaeyoung Do
‹ Prev 1 3 4 5 6 7 10 Next ›