中文
相关论文

相关论文: JoyHallo: Digital human model for Mandarin

200 篇论文

Significant progress has been made in talking-face video generation research; however, precise lip-audio synchronization and high visual quality remain challenging in editing lip shapes based on input audio. This paper introduces JoyGen, a…

计算机视觉与模式识别 · 计算机科学 2025-01-06 Qili Wang , Dajiang Wu , Zihang Xu , Junshi Huang , Jun Lv

Whisper speech recognition is crucial not only for ensuring privacy in sensitive communications but also for providing a critical communication bridge for patients under vocal restraint and enabling discrete interaction in noise-sensitive…

音频与语音处理 · 电气工程与系统科学 2025-09-30 Cancan Li , Fei Su , Juan Liu , Hui Bu , Yulong Wan , Hongbin Suo , Ming Li

Recent studies in speech-driven 3D talking head generation have achieved convincing results in verbal articulations. However, generating accurate lip-syncs degrades when applied to input speech in other languages, possibly due to the lack…

计算机视觉与模式识别 · 计算机科学 2024-06-21 Kim Sung-Bin , Lee Chae-Yeon , Gihun Son , Oh Hyun-Bin , Janghoon Ju , Suekyeong Nam , Tae-Hyun Oh

The development of multi-modal large language models (LLMs) leads to intelligent approaches capable of speech interactions. As one of the most widely spoken languages globally, Mandarin is supported by most models to enhance their…

计算与语言 · 计算机科学 2025-11-18 Heyang Liu , Ziyang Cheng , Yuhao Wang , Hongcheng Liu , Yiqi Li , Ronghua Wu , Qunshan Gu , Yanfeng Wang , Yu Wang

Large speech generation models are evolving from single-speaker, short sentence synthesis to multi-speaker, long conversation geneartion. Current long-form speech generation models are predominately constrained to dyadic, turn-based…

Automatic speech recognition (ASR) systems have advanced significantly with models like Whisper, Conformer, and self-supervised frameworks such as Wav2vec 2.0 and HuBERT. However, developing robust ASR models for young children's speech…

Existing video avatar models have demonstrated impressive capabilities in scenarios such as talking, public speaking, and singing. However, the majority of these methods exhibit limited alignment with respect to text instructions,…

计算机视觉与模式识别 · 计算机科学 2026-04-01 Ruikui Wang , Jinheng Feng , Lang Tian , Huaishao Luo , Chaochao Li , Liangbo Zhou , Huan Zhang , Youzheng Wu , Xiaodong He

In this paper, we present AISHELL-4, a sizable real-recorded Mandarin speech dataset collected by 8-channel circular microphone array for speech processing in conference scenario. The dataset consists of 211 recorded meeting sessions, each…

声音 · 计算机科学 2021-08-11 Yihui Fu , Luyao Cheng , Shubo Lv , Yukai Jv , Yuxiang Kong , Zhuo Chen , Yanxin Hu , Lei Xie , Jian Wu , Hui Bu , Xin Xu , Jun Du , Jingdong Chen

We introduce the Mandarin-English Language Interview (MELI) Corpus, an open-source resource of 29.8 hours of speech from 51 Mandarin-English bilingual speakers. MELI combines matched sessions in Mandarin and English with two speaking…

计算与语言 · 计算机科学 2026-05-18 Suyuan Liu , Molly Babel

Audio-driven talking head generation has drawn much attention in recent years, and many efforts have been made in lip-sync, expressive facial expressions, natural head pose generation, and high video quality. However, no model has yet led…

计算机视觉与模式识别 · 计算机科学 2023-12-08 Xusen Sun , Longhao Zhang , Hao Zhu , Peng Zhang , Bang Zhang , Xinya Ji , Kangneng Zhou , Daiheng Gao , Liefeng Bo , Xun Cao

Audio-visual speech recognition (AVSR) gains increasing attention from researchers as an important part of human-computer interaction. However, the existing available Mandarin audio-visual datasets are limited and lack the depth…

声音 · 计算机科学 2023-06-06 Jianrong Wang , Yuchen Huo , Li Liu , Tianyi Xu , Qi Li , Sen Li

Incorporating visual modalities to assist Automatic Speech Recognition (ASR) tasks has led to significant improvements. However, existing Audio-Visual Speech Recognition (AVSR) datasets and methods typically rely solely on lip-reading…

多媒体 · 计算机科学 2025-04-22 Jinghua Zhao , Yuhang Jia , Shiyao Wang , Jiaming Zhou , Hui Wang , Yong Qin

Modeling the reactive tempo of human conversation remains difficult because most audio-visual datasets portray isolated speakers delivering short monologues. We introduce \textbf{Face-to-Face with Jimmy Fallon (F2F-JF)}, a 70-hour, 14k-clip…

计算机视觉与模式识别 · 计算机科学 2026-04-01 Ernie Chu , Vishal M. Patel

Recently, there has been an increasing interest in neural speech synthesis. While the deep neural network achieves the state-of-the-art result in text-to-speech (TTS) tasks, how to generate a more emotional and more expressive speech is…

计算与语言 · 计算机科学 2021-06-24 Chenye Cui , Yi Ren , Jinglin Liu , Feiyang Chen , Rongjie Huang , Ming Lei , Zhou Zhao

With the rapid advancement of large language models (LLMs), foundational models (FMs) have seen significant advancements. Healthcare is one of the most crucial application areas for these FMs, given the significant time and effort required…

计算机视觉与模式识别 · 计算机科学 2024-11-18 Kaito Baba , Ryota Yagi , Junichiro Takahashi , Risa Kishikawa , Satoshi Kodera

Recent advances in diffusion-based video generation have enabled photo-realistic short clips, but current methods still struggle to achieve multi-modal consistency when jointly generating whole-body motion and natural speech. Current…

计算机视觉与模式识别 · 计算机科学 2025-07-30 Xinhan Di , Kristin Qi , Pengqian Yu

Singing, as a common facial movement second only to talking, can be regarded as a universal language across ethnicities and cultures, plays an important role in emotional communication, art, and entertainment. However, it is often…

计算机视觉与模式识别 · 计算机科学 2024-07-16 Sijing Wu , Yunhao Li , Weitian Zhang , Jun Jia , Yucheng Zhu , Yichao Yan , Guangtao Zhai , Xiaokang Yang

In this paper, we present WenetSpeech, a multi-domain Mandarin corpus consisting of 10000+ hours high-quality labeled speech, 2400+ hours weakly labeled speech, and about 10000 hours unlabeled speech, with 22400+ hours in total. We collect…

Mandarin Chinese is characterized by being a tonal language; the pitch (or $F_0$) of its utterances carries considerable linguistic information. However, speech samples from different individuals are subject to changes in amplitude and…

Audio-Visual Speech Recognition (AVSR) uses lip-based video to improve performance in noise. Since videos are harder to obtain than audio, the video training data of AVSR models is usually limited to a few thousand hours. In contrast,…

音频与语音处理 · 电气工程与系统科学 2024-11-21 Andrew Rouditchenko , Yuan Gong , Samuel Thomas , Leonid Karlinsky , Hilde Kuehne , Rogerio Feris , James Glass
‹ 上一页 1 2 3 10 下一页 ›