English
Related papers

Related papers: Towards Expressive Video Dubbing with Multiscale M…

200 papers

Robust voice activity detection (VAD) is a challenging task in low signal-to-noise (SNR) environments. Recent studies show that speech enhancement is helpful to VAD, but the performance improvement is limited. To address this issue, here we…

Audio and Speech Processing · Electrical Eng. & Systems 2021-04-14 Xu Tan , Xiao-Lei Zhang

Dense video captioning jointly localizes and captions salient events in untrimmed videos. Recent methods primarily focus on leveraging additional prior knowledge and advanced multi-task architectures to achieve competitive performance.…

Computer Vision and Pattern Recognition · Computer Science 2025-11-14 Mingda Jia , Weiliang Meng , Zenghuang Fu , Yiheng Li , Qi Zeng , Yifan Zhang , Ju Xin , Rongtao Xu , Jiguang Zhang , Xiaopeng Zhang

Natural human conversation is full-duplex and audio-visual: people simultaneously speak and listen while continuously interpreting and producing nonverbal cues, such as nods, smiles, and gestures. To support successful human-agent…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Amrita Mazumdar , Seonwook Park , Rajarshi Roy , Nikhil Srihari , Shengze Wang , Yuhao Zhou , Julia Wang , Koki Nagano , Shalini De Mello

Previous studies have explored generating accurately lip-synced talking faces for arbitrary targets given audio conditions. However, most of them deform or generate the whole facial area, leading to non-realistic results. In this work, we…

Computer Vision and Pattern Recognition · Computer Science 2022-12-12 Yasheng Sun , Hang Zhou , Kaisiyuan Wang , Qianyi Wu , Zhibin Hong , Jingtuo Liu , Errui Ding , Jingdong Wang , Ziwei Liu , Hideki Koike

Prosody is an integral part of communication, but remains an open problem in state-of-the-art speech synthesis. There are two major issues faced when modelling prosody: (1) prosody varies at a slower rate compared with other content in the…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-15 Zack Hodari , Alexis Moinet , Sri Karlapati , Jaime Lorenzo-Trueba , Thomas Merritt , Arnaud Joly , Ammar Abbas , Penny Karanasou , Thomas Drugman

This paper presents an audio visual automatic speech recognition (AV-ASR) system using a Transformer-based architecture. We particularly focus on the scene context provided by the visual information, to ground the ASR. We extract…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-01 Georgios Paraskevopoulos , Srinivas Parthasarathy , Aparna Khare , Shiva Sundaram

In this work, we investigate the task of textual response generation in a multimodal task-oriented dialogue system. Our work is based on the recently released Multimodal Dialogue (MMD) dataset (Saha et al., 2017) in the fashion domain. We…

Computation and Language · Computer Science 2018-11-22 Shubham Agarwal , Ondrej Dusek , Ioannis Konstas , Verena Rieser

Multimodal emotion recognition (MER) aims to infer human affect by jointly modeling audio and visual cues; however, existing approaches often struggle with temporal misalignment, weakly discriminative feature representations, and suboptimal…

Multimedia · Computer Science 2026-01-21 Joe Dhanith P R , Shravan Venkatraman , Vigya Sharma , Santhosh Malarvannan

Multimodal sentiment analysis aims to recognize people's attitudes from multiple communication channels such as verbal content (i.e., text), voice, and facial expressions. It has become a vibrant and important research topic in natural…

Machine Learning · Computer Science 2022-02-23 Xingbo Wang , Jianben He , Zhihua Jin , Muqiao Yang , Yong Wang , Huamin Qu

The existing state-of-the-art method for audio-visual conditioned video prediction uses the latent codes of the audio-visual frames from a multimodal stochastic network and a frame encoder to predict the next visual frame. However, a direct…

Computer Vision and Pattern Recognition · Computer Science 2023-09-21 Yating Xu , Conghui Hu , Gim Hee Lee

Multimodal sentiment analysis is a fundamental problem in the field of affective computing. Although significant progress has been made in cross-modal interaction, it remains a challenge due to the insufficient reference context in…

Multimedia · Computer Science 2025-08-12 Xianbing Zhao , Shengzun Yang , Buzhou Tang , Ronghuan Jiang

Text-based speech editors expedite the process of editing speech recordings by permitting editing via intuitive cut, copy, and paste operations on a speech transcript. A major drawback of current systems, however, is that edited recordings…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-17 Max Morrison , Lucas Rencker , Zeyu Jin , Nicholas J. Bryan , Juan-Pablo Caceres , Bryan Pardo

Automated deception detection systems can enhance health, justice, and security in society by helping humans detect deceivers in high-stakes situations across medical and legal domains, among others. This paper presents a novel analysis of…

Computer Vision and Pattern Recognition · Computer Science 2020-10-27 Leena Mathur , Maja J Matarić

LLM-based translation agents have achieved highly human-like translation results and are capable of handling longer and more complex contexts with greater efficiency. However, they are typically limited to text-only inputs. In this paper,…

Artificial Intelligence · Computer Science 2025-07-11 Yichen Lu , Wei Dai , Jiaen Liu , Ching Wing Kwok , Zongheng Wu , Xudong Xiao , Ao Sun , Sheng Fu , Jianyuan Zhan , Yian Wang , Takatomo Saito , Sicheng Lai

In human communication, both verbal and non-verbal cues play a crucial role in conveying emotions, intentions, and meaning beyond words alone. These non-linguistic information, such as facial expressions, eye contact, voice tone, and pitch,…

Computer Vision and Pattern Recognition · Computer Science 2024-12-24 Se Jin Park , Yeonju Kim , Hyeongseop Rha , Bella Godiva , Yong Man Ro

Recent advancements in the field of Diffusion Transformers have substantially improved the generation of high-quality 2D images, 3D videos, and 3D shapes. However, the effectiveness of the Transformer architecture in the domain of co-speech…

Computer Vision and Pattern Recognition · Computer Science 2024-08-07 Xiaofeng Mao , Zhengkai Jiang , Qilin Wang , Chencan Fu , Jiangning Zhang , Jiafu Wu , Yabiao Wang , Chengjie Wang , Wei Li , Mingmin Chi

Detecting sarcasm effectively requires a nuanced understanding of context, including vocal tones and facial expressions. The progression towards multimodal computational methods in sarcasm detection, however, faces challenges due to the…

Computation and Language · Computer Science 2024-12-16 Xiyuan Gao , Shubhi Bansal , Kushaan Gowda , Zhu Li , Shekhar Nayak , Nagendra Kumar , Matt Coler

In computer vision, multi-label recognition are important tasks with many real-world applications, but classifying previously unseen labels remains a significant challenge. In this paper, we propose a novel algorithm, Aligned Dual moDality…

Computer Vision and Pattern Recognition · Computer Science 2023-10-10 Shichao Xu , Yikang Li , Jenhao Hsiao , Chiuman Ho , Zhu Qi

Audio-visual speech recognition (AVSR) is an extension of ASR that incorporates visual signals. Current AVSR approaches primarily focus on lip motion, largely overlooking rich context present in the video such as speaking scene and…

Multi-modal based speech separation has exhibited a specific advantage on isolating the target character in multi-talker noisy environments. Unfortunately, most of current separation strategies prefer a straightforward fusion based on…

Sound · Computer Science 2022-03-08 Junwen Xiong , Peng Zhang , Lei Xie , Wei Huang , Yufei Zha , Yanning Zhang