中文
相关论文

相关论文: JoyHallo: Digital human model for Mandarin

200 篇论文

An increasing number of Chinese people are troubled by different degrees of visual impairment, which has made the modal conversion between a single image or video frame in the visual field and the audio expressing the same information a…

声音 · 计算机科学 2024-07-22 Chun Xu , En-Wei Sun

Audio-driven portrait animation has made significant advances with diffusion-based models, improving video quality and lipsync accuracy. However, the increasing complexity of these models has led to inefficiencies in training and inference,…

计算机视觉与模式识别 · 计算机科学 2026-04-17 Xuyang Cao , Guoxin Wang , Sheng Shi , Jun Zhao , Yang Yao , Jintao Fei , Minyu Gao , Pei Xie

Real-time text-driven joint audio-video avatar generation requires jointly synthesizing portrait video and speech with high fidelity and precise synchronization, yet existing audio-visual diffusion models remain too slow for interactive use…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Chunyu Li , Jiaye Li , Ruiqiao Mei , Haoyuan Xia , Hao Zhu , Jingdong Wang , Siyu Zhu

In this paper, we present JoVA, a unified framework for joint video-audio generation. Despite recent encouraging advances, existing methods face two critical limitations. First, most existing approaches can only generate ambient sounds and…

计算机视觉与模式识别 · 计算机科学 2025-12-16 Xiaohu Huang , Hao Zhou , Qiangpeng Yang , Shilei Wen , Kai Han

Our contribution introduces a groundbreaking multimodal large language model designed to comprehend multi-images, multi-audio, and multi-images-multi-audio within a single multiturn session. Leveraging state-of-the-art models, we utilize…

计算与语言 · 计算机科学 2024-02-20 Husein Zolkepli , Aisyah Razak , Kamarul Adha , Ariff Nazhan

The field of portrait image animation, driven by speech audio input, has experienced significant advancements in the generation of realistic and dynamic portraits. This research delves into the complexities of synchronizing facial movements…

计算机视觉与模式识别 · 计算机科学 2024-06-18 Mingwang Xu , Hui Li , Qingkun Su , Hanlin Shang , Liwei Zhang , Ce Liu , Jingdong Wang , Yao Yao , Siyu Zhu

The aim of this work is to investigate the impact of crossmodal self-supervised pre-training for speech reconstruction (video-to-audio) by leveraging the natural co-occurrence of audio and visual streams in videos. We propose LipSound2…

声音 · 计算机科学 2022-09-13 Leyuan Qu , Cornelius Weber , Stefan Wermter

This paper outlines the methodology for modeling tonal learning in fully unsupervised models of human language acquisition. Tonal patterns are among the computationally most complex learning objectives in language. We argue that a realistic…

计算与语言 · 计算机科学 2025-09-23 Kai Schenck , Gašper Beguš

The tasks of automatic lyrics transcription and lyrics alignment have witnessed significant performance improvements in the past few years. However, most of the previous works only focus on English in which large-scale datasets are…

音频与语音处理 · 电气工程与系统科学 2023-11-22 Jun-You Wang , Chon-In Leong , Yu-Chen Lin , Li Su , Jyh-Shing Roger Jang

This study asks how self-supervised speech models represent suprasegmental categories like Mandarin lexical tone, English lexical stress, and English phrasal accents. Through a series of probing tasks, we make layer-wise comparisons of…

计算与语言 · 计算机科学 2024-08-27 Antón de la Fuente , Dan Jurafsky

AISHELL-1 is by far the largest open-source speech corpus available for Mandarin speech recognition research. It was released with a baseline system containing solid training and testing pipelines for Mandarin ASR. In AISHELL-2, 1000 hours…

计算与语言 · 计算机科学 2018-09-14 Jiayu Du , Xingyu Na , Xuechen Liu , Hui Bu

This paper introduces a new open-sourced Mandarin speech corpus, called DiDiSpeech. It consists of about 800 hours of speech data at 48kHz sampling rate from 6000 speakers and the corresponding texts. All speech data in the corpus is…

音频与语音处理 · 电气工程与系统科学 2021-02-09 Tingwei Guo , Cheng Wen , Dongwei Jiang , Ne Luo , Ruixiong Zhang , Shuaijiang Zhao , Wubo Li , Cheng Gong , Wei Zou , Kun Han , Xiangang Li

The global aging population faces considerable challenges, particularly in communication, due to the prevalence of hearing and speech impairments. To address these, we introduce the AVE speech, a comprehensive multi-modal dataset for speech…

声音 · 计算机科学 2025-07-08 Dongliang Zhou , Yakun Zhang , Jinghan Wu , Xingyu Zhang , Liang Xie , Erwei Yin

Cued Speech (CS) is a communication system developed for deaf people, which exploits hand cues to complement speechreading at the phonetic level. Currently, it is estimated that CS has been adapted to over 60 languages; however, no official…

音频与语音处理 · 电气工程与系统科学 2020-01-06 Liu Li , Feng Gang

Audio-driven human animation has attracted wide attention thanks to its practical applications. However, critical challenges remain in generating high-resolution, long-duration videos with consistent appearance and natural hand motions.…

计算机视觉与模式识别 · 计算机科学 2025-08-29 Xiaodi Li , Pan Xie , Yi Ren , Qijun Gan , Chen Zhang , Fangyuan Kong , Xiang Yin , Bingyue Peng , Zehuan Yuan

High-quality multilingual training data is essential for effectively pretraining large language models (LLMs). Yet, the availability of suitable open-source multilingual datasets remains limited. Existing state-of-the-art datasets mostly…

Previous works on expressive speech synthesis focus on modelling the mono-scale style embedding from the current sentence or context, but the multi-scale nature of speaking style in human speech is neglected. In this paper, we propose a…

声音 · 计算机科学 2022-07-06 Shun Lei , Yixuan Zhou , Liyang Chen , Jiankun Hu , Zhiyong Wu , Shiyin Kang , Helen Meng

We present HourVideo, a benchmark dataset for hour-long video-language understanding. Our dataset consists of a novel task suite comprising summarization, perception (recall, tracking), visual reasoning (spatial, temporal, predictive,…

计算机视觉与模式识别 · 计算机科学 2024-11-08 Keshigeyan Chandrasegaran , Agrim Gupta , Lea M. Hadzic , Taran Kota , Jimming He , Cristóbal Eyzaguirre , Zane Durante , Manling Li , Jiajun Wu , Li Fei-Fei

Recent advances in latent diffusion-based generative models for portrait image animation, such as Hallo, have achieved impressive results in short-duration video synthesis. In this paper, we present updates to Hallo, introducing several…

计算机视觉与模式识别 · 计算机科学 2024-10-15 Jiahao Cui , Hui Li , Yao Yao , Hao Zhu , Hanlin Shang , Kaihui Cheng , Hang Zhou , Siyu Zhu , Jingdong Wang

Creating spoken dialogue datasets is methodologically challenging, and these challenges are amplified when the goal is to build multilingual, multi-parallel datasets at scale. This work introduces HEALTHDIAL, a large-scale, multilingual,…

计算与语言 · 计算机科学 2026-05-29 Songbo Hu , Yinhong Liu , Ej Zhou , Evgeniia Razumovskaia , Xiaobin Wang , Alexander Fraser , Ivan Vulić , Anna Korhonen