中文
相关论文

相关论文: Can Hierarchical Cross-Modal Fusion Predict Human …

200 篇论文

Automatic mean opinion score (MOS) prediction provides a more perceptual alternative to objective metrics, offering deeper insights into the evaluated models. With the rapid progress of multimodal large language models (MLLMs), their…

声音 · 计算机科学 2025-09-23 Yuhang Jia , Xu Zhang , Yang Chen , Hui Wang , Enzhi Wang , Yong Qin

A major challenge for video captioning is to combine audio and visual cues. Existing multi-modal fusion methods have shown encouraging results in video understanding. However, the temporal structures of multiple modalities at different…

计算与语言 · 计算机科学 2018-04-17 Xin Wang , Yuan-Fang Wang , William Yang Wang

Given a piece of text, a video clip and a reference audio, the movie dubbing (also known as visual voice clone V2C) task aims to generate speeches that match the speaker's emotion presented in the video using the desired speaker voice as…

计算与语言 · 计算机科学 2023-04-05 Gaoxiang Cong , Liang Li , Yuankai Qi , Zhengjun Zha , Qi Wu , Wenyu Wang , Bin Jiang , Ming-Hsuan Yang , Qingming Huang

Cross-lingual dubbing of lecture videos requires the transcription of the original audio, correction and removal of disfluencies, domain term discovery, text-to-text translation into the target language, chunking of text using target…

Current movie dubbing technology can produce the desired speech using a reference voice and input video, maintaining perfect synchronization with the visuals while effectively conveying the intended emotions. However, crucial aspects of…

多媒体 · 计算机科学 2025-05-23 Junjie Zheng , Zihao Chen , Chaofan Ding , Yunming Liang , Yihan Fan , Huan Yang , Lei Xie , Xinhan Di

Multimodal sentiment analysis, a pivotal task in affective computing, seeks to understand human emotions by integrating cues from language, audio, and visual signals. While many recent approaches leverage complex attention mechanisms and…

计算与语言 · 计算机科学 2025-05-09 Nischal Mandal , Yang Li

High-quality, large-scale audio captioning is crucial for advancing audio understanding, yet current automated methods often generate captions that lack fine-grained detail and contextual accuracy, primarily due to their reliance on limited…

声音 · 计算机科学 2025-06-03 Shunian Chen , Xinyuan Xie , Zheshu Chen , Liyan Zhao , Owen Lee , Zhan Su , Qilin Sun , Benyou Wang

In this paper, we present a novel deep multimodal framework to predict human emotions based on sentence-level spoken language. Our architecture has two distinctive characteristics. First, it extracts the high-level features from both text…

计算与语言 · 计算机科学 2018-02-26 Yue Gu , Shuhong Chen , Ivan Marsic

In the last decade, video blogs (vlogs) have become an extremely popular method through which people express sentiment. The ubiquitousness of these videos has increased the importance of multimodal fusion models, which incorporate video and…

计算机视觉与模式识别 · 计算机科学 2018-07-04 Nathaniel Blanchard , Daniel Moreira , Aparna Bharati , Walter J. Scheirer

Automatic speech quality assessment aims to quantify subjective human perception of speech through computational models to reduce the need for labor-consuming manual evaluations. While models based on deep learning have achieved progress in…

声音 · 计算机科学 2025-04-30 Zhicheng Lian , Lizhi Wang , Hua Huang

Existing works on weakly-supervised audio-visual video parsing adopt hybrid attention network (HAN) as the multi-modal embedding to capture the cross-modal context. It embeds the audio and visual modalities with a shared network, where the…

计算机视觉与模式识别 · 计算机科学 2023-11-15 Yating Xu , Conghui Hu , Gim Hee Lee

Perceptually-inspired objective functions such as the perceptual evaluation of speech quality (PESQ), signal-to-distortion ratio (SDR), and short-time objective intelligibility (STOI), have recently been used to optimize performance of…

音频与语音处理 · 电气工程与系统科学 2023-03-27 Khandokar Md. Nayem , Donald S. Williamson

Multimodal sentiment analysis is a very actively growing field of research. A promising area of opportunity in this field is to improve the multimodal fusion mechanism. We present a novel feature fusion strategy that proceeds in a…

计算与语言 · 计算机科学 2018-06-19 N. Majumder , D. Hazarika , A. Gelbukh , E. Cambria , S. Poria

Existing objective evaluation metrics for voice conversion (VC) are not always correlated with human perception. Therefore, training VC models with such criteria may not effectively improve naturalness and similarity of converted speech. In…

声音 · 计算机科学 2022-03-01 Chen-Chou Lo , Szu-Wei Fu , Wen-Chin Huang , Xin Wang , Junichi Yamagishi , Yu Tsao , Hsin-Min Wang

Existing robotic manipulation methods primarily rely on visual and proprioceptive observations, which may struggle to infer contact-related interaction states in partially observable real-world environments. Acoustic cues, by contrast,…

机器人学 · 计算机科学 2026-02-17 Siyuan Li , Jiani Lu , Yu Song , Xianren Li , Bo An , Peng Liu

Multimodal affective computing, learning to recognize and interpret human affects and subjective information from multiple data sources, is still challenging because: (i) it is hard to extract informative features to represent human affects…

计算与语言 · 计算机科学 2018-05-23 Yue Gu , Kangning Yang , Shiyu Fu , Shuhong Chen , Xinyu Li , Ivan Marsic

Multimodal sentiment analysis (MSA) integrates various modalities, such as text, image, and audio, to provide a more comprehensive understanding of sentiment. However, effective MSA is challenged by alignment and fusion issues. Alignment…

计算机视觉与模式识别 · 计算机科学 2025-12-08 Yuhua Wen , Qifei Li , Yingying Zhou , Yingming Gao , Zhengqi Wen , Jianhua Tao , Ya Li

Automatic Video Dubbing (AVD) generates speech aligned with lip motion and facial emotion from scripts. Recent research focuses on modeling multimodal context to enhance prosody expressiveness but overlooks two key issues: 1) Multiscale…

多媒体 · 计算机科学 2025-01-03 Yuan Zhao , Rui Liu , Gaoxiang Cong

We describe a system for large-scale audiovisual translation and dubbing, which translates videos from one language to another. The source language's speech content is transcribed to text, translated, and automatically synthesized into…

计算机视觉与模式识别 · 计算机科学 2020-11-09 Yi Yang , Brendan Shillingford , Yannis Assael , Miaosen Wang , Wendi Liu , Yutian Chen , Yu Zhang , Eren Sezener , Luis C. Cobo , Misha Denil , Yusuf Aytar , Nando de Freitas

We propose MORAL (a multimodal reinforcement learning framework for decision making in autonomous laboratories) that enhances sequential decision-making in autonomous robotic laboratories through the integration of visual and textual…

机器学习 · 计算机科学 2025-04-07 Natalie Tirabassi , Sathish A. P. Kumar , Sumit Jha , Arvind Ramanathan
‹ 上一页 1 2 3 10 下一页 ›