中文
相关论文

相关论文: Leveraging WaveNet for Dynamic Listening Head Mode…

200 篇论文

Visually-grounded spoken language datasets can enable models to learn cross-modal correspondences with very weak supervision. However, modern audio-visual datasets contain biases that undermine the real-world performance of models trained…

计算与语言 · 计算机科学 2021-10-15 Ian Palmer , Andrew Rouditchenko , Andrei Barbu , Boris Katz , James Glass

In the field of human-computer interaction and psychological assessment, speech emotion recognition (SER) plays an important role in deciphering emotional states from speech signals. Despite advancements, challenges persist due to system…

声音 · 计算机科学 2025-02-04 Alaa Nfissi , Wassim Bouachir , Nizar Bouguila , Brian Mishara

We consider the task of unsupervised extraction of meaningful latent representations of speech by applying autoencoding neural networks to speech waveforms. The goal is to learn a representation able to capture high level semantic content…

机器学习 · 计算机科学 2019-09-12 Jan Chorowski , Ron J. Weiss , Samy Bengio , Aäron van den Oord

Recent Speech Large Language Models~(LLMs) have achieved impressive capabilities in end-to-end speech interaction. However, the prevailing autoregressive paradigm imposes strict serial constraints, limiting generation efficiency and…

计算与语言 · 计算机科学 2026-02-10 Ziyang Cheng , Yuhao Wang , Heyang Liu , Ronghua Wu , Qunshan Gu , Yanfeng Wang , Yu Wang

Understanding the lip movement and inferring the speech from it is notoriously difficult for the common person. The task of accurate lip-reading gets help from various cues of the speaker and its contextual or environmental setting. Every…

计算机视觉与模式识别 · 计算机科学 2022-08-23 Munender Varshney , Ravindra Yadav , Vinay P. Namboodiri , Rajesh M Hegde

Continuous QoE prediction is crucial in the purpose of maximizing viewer satisfaction, by which video service providers could improve the revenue. Continuously predicting QoE is challenging since it requires QoE models that are capable of…

多媒体 · 计算机科学 2020-03-23 Phan Xuan Tan , Tho Nguyen Duc , Chanh Minh Tran , Eiji Kamioka

Wideband Audio Waveform Evaluation Networks (WAWEnets) are convolutional neural networks that operate directly on wideband audio waveforms in order to produce evaluations of those waveforms. In the present work these evaluations give…

音频与语音处理 · 电气工程与系统科学 2023-11-21 Andrew Catellier , Stephen Voran

Face super-resolution aims to reconstruct a high-resolution face image from a low-resolution face image. Previous methods typically employ an encoder-decoder structure to extract facial structural features, where the direct downsampling…

计算机视觉与模式识别 · 计算机科学 2024-07-31 Wenjie Li , Heng Guo , Xuannan Liu , Kongming Liang , Jiani Hu , Zhanyu Ma , Jun Guo

Unlike prevalent facial expressions, micro expressions have subtle, involuntary muscle movements which are short-lived in nature. These minute muscle movements reflect true emotions of a person. Due to the short duration and low intensity,…

计算机视觉与模式识别 · 计算机科学 2020-01-08 Monu Verma , Santosh Kumar Vipparthi , Girdhari Singh , Subrahmanyam Murala

Long-form video understanding remains challenging for Vision-Language Models (VLMs) due to the inherent tension between computational constraints and the need to capture information distributed across thousands of frames. Existing…

计算机视觉与模式识别 · 计算机科学 2026-02-05 Junbo Zou , Ziheng Huang , Shengjie Zhang , Liwen Zhang , Weining Shen

Generating accurate sounds for complex audio-visual scenes is challenging, especially in the presence of multiple objects and sound sources. In this paper, we propose an {\em interactive object-aware audio generation} model that grounds…

计算机视觉与模式识别 · 计算机科学 2025-06-05 Tingle Li , Baihe Huang , Xiaobin Zhuang , Dongya Jia , Jiawei Chen , Yuping Wang , Zhuo Chen , Gopala Anumanchipalli , Yuxuan Wang

We introduce a new approach for audio-visual speech separation. Given a video, the goal is to extract the speech associated with a face in spite of simultaneous background sounds and/or other human speakers. Whereas existing methods focus…

计算机视觉与模式识别 · 计算机科学 2021-04-07 Ruohan Gao , Kristen Grauman

In this paper, we propose an open source, production first, and production ready speech recognition toolkit called WeNet in which a new two-pass approach is implemented to unify streaming and non-streaming end-to-end (E2E) speech…

声音 · 计算机科学 2021-12-30 Zhuoyuan Yao , Di Wu , Xiong Wang , Binbin Zhang , Fan Yu , Chao Yang , Zhendong Peng , Xiaoyu Chen , Lei Xie , Xin Lei

The objective of face animation is to generate dynamic and expressive talking head videos from a single reference face, utilizing driving conditions derived from either video or audio inputs. Current approaches often require fine-tuning for…

计算机视觉与模式识别 · 计算机科学 2024-07-15 He Feng , Donglin Di , Yongjia Ma , Wei Chen , Tonghua Su

Speechreading is a notoriously difficult task for humans to perform. In this paper we present an end-to-end model based on a convolutional neural network (CNN) for generating an intelligible acoustic speech signal from silent video frames…

计算机视觉与模式识别 · 计算机科学 2017-01-10 Ariel Ephrat , Shmuel Peleg

Deepfake generation has witnessed remarkable progress, contributing to highly realistic generated images, videos, and audio. While technically intriguing, such progress has raised serious concerns related to the misuse of manipulated media.…

计算机视觉与模式识别 · 计算机科学 2025-12-01 Maheswar Bora , Tashvik Dhamija , Shukesh Reddy , Baptiste Chopin , Pranav Balaji , Abhijit Das , Antitza Dantcheva

We present AlignNet, a model that synchronizes videos with reference audios under non-uniform and irregular misalignments. AlignNet learns the end-to-end dense correspondence between each frame of a video and an audio. Our method is…

计算机视觉与模式识别 · 计算机科学 2020-02-13 Jianren Wang , Zhaoyuan Fang , Hang Zhao

Given an arbitrary audio clip, audio-driven 3D facial animation aims to generate lifelike lip motions and facial expressions for a 3D head. Existing methods typically rely on training their models using limited public 3D datasets that…

计算机视觉与模式识别 · 计算机科学 2023-06-21 Liying Lu , Tianke Zhang , Yunfei Liu , Xuangeng Chu , Yu Li

Video-to-video synthesis is a challenging problem aiming at learning a translation function between a sequence of semantic maps and a photo-realistic video depicting the characteristics of a driving video. We propose a head-to-head system…

计算机视觉与模式识别 · 计算机科学 2020-06-19 Mohammad Rami Koujan , Michail Christos Doukas , Anastasios Roussos , Stefanos Zafeiriou

Emotional expressions are the behaviors that communicate our emotional state or attitude to others. They are expressed through verbal and non-verbal communication. Complex human behavior can be understood by studying physical features from…

计算机视觉与模式识别 · 计算机科学 2021-09-15 Liam Schoneveld , Alice Othmani , Hazem Abdelkawy