中文
相关论文

相关论文: JoyHallo: Digital human model for Mandarin

200 篇论文

Cued Speech (CS) is a multi-modal visual coding system combining lip reading with several hand cues at the phonetic level to make the spoken language visible to the hearing impaired. Previous studies solved asynchronous problems between lip…

音频与语音处理 · 电气工程与系统科学 2023-06-06 Lufei Gao , Shan Huang , Li Liu

Understanding videos inherently requires reasoning over both visual and auditory information. To properly evaluate Omni-Large Language Models (Omni-LLMs), which are capable of processing multi-modal information including vision and audio,…

多媒体 · 计算机科学 2026-05-15 Jianghan Chao , Jianzhang Gao , Wenhui Tan , Yuchong Sun , Ruihua Song , Liyun Ru

Audio-driven cospeech video generation typically involves two stages: speech-to-gesture and gesture-to-video. While significant advances have been made in speech-to-gesture generation, synthesizing natural expressions and gestures remains…

计算机视觉与模式识别 · 计算机科学 2025-04-14 Renda Li , Xiaohua Qi , Qiang Ling , Jun Yu , Ziyi Chen , Peng Chang , Mei HanJing Xiao

The task of few-shot visual dubbing focuses on synchronizing the lip movements with arbitrary speech input for any talking head video. Albeit moderate improvements in current approaches, they commonly require high-quality homologous data…

计算机视觉与模式识别 · 计算机科学 2022-01-19 Tianyi Xie , Liucheng Liao , Cheng Bi , Benlai Tang , Xiang Yin , Jianfei Yang , Mingjie Wang , Jiali Yao , Yang Zhang , Zejun Ma

We propose Wav2CLIP, a robust audio representation learning method by distilling from Contrastive Language-Image Pre-training (CLIP). We systematically evaluate Wav2CLIP on a variety of audio tasks including classification, retrieval, and…

声音 · 计算机科学 2022-02-16 Ho-Hsiang Wu , Prem Seetharaman , Kundan Kumar , Juan Pablo Bello

This paper presents AMNet, an Acoustic Model Network designed to improve the performance of Mandarin speech synthesis by incorporating phrase structure annotation and local convolution modules. AMNet builds upon the FastSpeech 2…

声音 · 计算机科学 2025-04-15 Yubing Cao , Yinfeng Yu , Yongming Li , Liejun Wang

Traditionally, the performance of non-native mispronunciation verification systems relied on effective phone-level labelling of non-native corpora. In this study, a multi-view approach is proposed to incorporate discriminative feature…

音频与语音处理 · 电气工程与系统科学 2020-09-10 Zhenyu Wang , John H. L. Hansen , Yanlu Xie

We present MagicInfinite, a novel diffusion Transformer (DiT) framework that overcomes traditional portrait animation limitations, delivering high-fidelity results across diverse character types-realistic humans, full-body figures, and…

计算机视觉与模式识别 · 计算机科学 2025-03-11 Hongwei Yi , Tian Ye , Shitong Shao , Xuancheng Yang , Jiantong Zhao , Hanzhong Guo , Terrance Wang , Qingyu Yin , Zeke Xie , Lei Zhu , Wei Li , Michael Lingelbach , Daquan Zhou

Multimedia or spoken content presents more attractive information than plain text content, but the former is more difficult to display on a screen and be selected by a user. As a result, accessing large collections of the former is much…

计算与语言 · 计算机科学 2017-01-03 Wei Fang , Jui-Yang Hsu , Hung-yi Lee , Lin-Shan Lee

As multimodal content continues to expand at a rapid pace, audio retrieval has emerged as a key enabling technology for media search, content organization, and intelligent assistants. However, most existing benchmarks concentrate on…

人工智能 · 计算机科学 2026-05-07 Honglei Zhang , Yuting Chen , Chenpeng Hu , Siyue Zhang , Yilei Shi

In most cases, bilingual TTS needs to handle three types of input scripts: first language only, second language only, and second language embedded in the first language. In the latter two situations, the pronunciation and intonation of the…

声音 · 计算机科学 2022-12-08 Fengyu Yang , Jian Luan , Yujun Wang

In recent years, emotion recognition plays a critical role in applications such as human-computer interaction, mental health monitoring, and sentiment analysis. While datasets for emotion analysis in languages such as English have…

We propose StyleTalker, a novel audio-driven talking head generation model that can synthesize a video of a talking person from a single reference image with accurately audio-synced lip shapes, realistic head poses, and eye blinks.…

计算机视觉与模式识别 · 计算机科学 2024-03-18 Dongchan Min , Minyoung Song , Eunji Ko , Sung Ju Hwang

Existing video avatar models can produce fluid human animations, yet they struggle to move beyond mere physical likeness to capture a character's authentic essence. Their motions typically synchronize with low-level cues like audio rhythm,…

计算机视觉与模式识别 · 计算机科学 2025-08-27 Jianwen Jiang , Weihong Zeng , Zerong Zheng , Jiaqi Yang , Chao Liang , Wang Liao , Han Liang , Yuan Zhang , Mingyuan Gao

Talking head synthesis is a promising approach for the video production industry. Recently, a lot of effort has been devoted in this research area to improve the generation quality or enhance the model generalization. However, there are few…

计算机视觉与模式识别 · 计算机科学 2023-04-21 Shuai Shen , Wenliang Zhao , Zibin Meng , Wanhua Li , Zheng Zhu , Jie Zhou , Jiwen Lu

Real-world talking faces often accompany with natural head movement. However, most existing talking face video generation methods only consider facial animation with fixed head pose. In this paper, we address this problem by proposing a…

计算机视觉与模式识别 · 计算机科学 2020-03-06 Ran Yi , Zipeng Ye , Juyong Zhang , Hujun Bao , Yong-Jin Liu

In this work, we tackle the challenge of enhancing the realism and expressiveness in talking head video generation by focusing on the dynamic and nuanced relationship between audio cues and facial movements. We identify the limitations of…

计算机视觉与模式识别 · 计算机科学 2024-08-09 Linrui Tian , Qi Wang , Bang Zhang , Liefeng Bo

Automatic Cued Speech Recognition (ACSR) provides an intelligent human-machine interface for visual communications, where the Cued Speech (CS) system utilizes lip movements and hand gestures to code spoken language for hearing-impaired…

计算机视觉与模式识别 · 计算机科学 2023-02-28 Lei Liu , Li Liu

We introduce a novel method for joint expression and audio-guided talking face generation. Recent approaches either struggle to preserve the speaker identity or fail to produce faithful facial expressions. To address these challenges, we…

计算机视觉与模式识别 · 计算机科学 2024-09-19 Sai Tanmay Reddy Chakkera , Aggelina Chatziagapi , Dimitris Samaras

Even for better-studied sign languages like American Sign Language (ASL), data is the bottleneck for machine learning research. The situation is worse yet for the many other sign languages used by Deaf/Hard of Hearing communities around the…

计算与语言 · 计算机科学 2024-07-17 Garrett Tanzer , Biao Zhang
‹ 上一页 1 8 9 10 下一页 ›