中文
相关论文

相关论文: Speech-driven Personalized Gesture Synthetics: Har…

200 篇论文

Recent work has explored a range of model families for human motion generation, including Variational Autoencoders (VAEs), Generative Adversarial Networks (GANs), and diffusion-based models. Despite their differences, many methods rely on…

计算机视觉与模式识别 · 计算机科学 2025-10-15 David Björkstrand , Tiesheng Wang , Lars Bretzner , Josephine Sullivan

Although recent speech processing technologies have achieved significant improvements in objective metrics, there still remains a gap in human perceptual quality. This paper proposes Diffiner, a novel solution that utilizes the powerful…

音频与语音处理 · 电气工程与系统科学 2026-02-11 Masato Hirano , Ryosuke Sawata , Naoki Murata , Shusuke Takahashi , Yuki Mitsufuji

The recent advancements in image-text diffusion models have stimulated research interest in large-scale 3D generative models. Nevertheless, the limited availability of diverse 3D resources presents significant challenges to learning. In…

计算机视觉与模式识别 · 计算机科学 2023-06-01 Chi Zhang , Yiwen Chen , Yijun Fu , Zhenglin Zhou , Gang YU , Billzb Wang , Bin Fu , Tao Chen , Guosheng Lin , Chunhua Shen

Gestures are pivotal in enhancing co-speech communication. While recent works have mostly focused on point-level motion transformation or fully supervised motion representations through data-driven approaches, we explore the representation…

计算机视觉与模式识别 · 计算机科学 2024-09-27 Huan Yang , Jiahui Chen , Chaofan Ding , Runhua Shi , Siyu Xiong , Qingqi Hong , Xiaoqi Mo , Xinhan Di

This paper investigates a novel task of talking face video generation solely from speeches. The speech-to-video generation technique can spark interesting applications in entertainment, customer service, and human-computer-interaction…

声音 · 计算机科学 2021-07-15 Shijing Si , Jianzong Wang , Xiaoyang Qu , Ning Cheng , Wenqi Wei , Xinghua Zhu , Jing Xiao

In this paper, we introduce GaussianMotion, a novel human rendering model that generates fully animatable scenes aligned with textual descriptions using Gaussian Splatting. Although existing methods achieve reasonable text-to-3D generation…

计算机视觉与模式识别 · 计算机科学 2025-02-18 Gyumin Shim , Sangmin Lee , Jaegul Choo

As artificial intelligence-generated content (AIGC) continues to evolve, video-to-audio (V2A) generation has emerged as a key area with promising applications in multimedia editing, augmented reality, and automated content creation. While…

计算机视觉与模式识别 · 计算机科学 2025-03-17 Yuhuan You , Xihong Wu , Tianshu Qu

Creating realistic avatars from a single RGB image is an attractive yet challenging problem. Due to its ill-posed nature, recent works leverage powerful prior from 2D diffusion models pretrained on large datasets. Although 2D diffusion…

计算机视觉与模式识别 · 计算机科学 2024-12-17 Yuxuan Xue , Xianghui Xie , Riccardo Marin , Gerard Pons-Moll

Recent advancements in diffusion models have significantly improved the realism and generalizability of character-driven animation, enabling the synthesis of high-quality motion from just a single RGB image and a set of driving poses.…

Dysarthric speech reconstruction (DSR) aims to convert dysarthric speech into comprehensible speech while maintaining the speaker's identity. Despite significant advancements, existing methods often struggle with low speech intelligibility…

声音 · 计算机科学 2025-06-03 Xueyuan Chen , Dongchao Yang , Wenxuan Wu , Minglin Wu , Jing Xu , Xixin Wu , Zhiyong Wu , Helen Meng

Generating high-quality 360-degree views of human heads from single-view images is essential for enabling accessible immersive telepresence applications and scalable personalized content creation. While cutting-edge methods for full head…

计算机视觉与模式识别 · 计算机科学 2025-03-21 Yuming Gu , Phong Tran , Yujian Zheng , Hongyi Xu , Heyuan Li , Adilbek Karmanov , Hao Li

The development of increasingly realistic experimental stimuli and task environments is important for understanding behavior outside the laboratory. We report a process for generating 3D human model stimuli that combines commonly used…

神经元与认知 · 定量生物学 2016-06-17 Jean M. Vettel , Justin Kantner , Matthew Jaswa , Michael Miller

The field of portrait image animation, driven by speech audio input, has experienced significant advancements in the generation of realistic and dynamic portraits. This research delves into the complexities of synchronizing facial movements…

计算机视觉与模式识别 · 计算机科学 2024-06-18 Mingwang Xu , Hui Li , Qingkun Su , Hanlin Shang , Liwei Zhang , Ce Liu , Jingdong Wang , Yao Yao , Siyu Zhu

Online Speech Enhancement was mainly reserved for predictive models. A key advantage of these models is that for an incoming signal frame from a stream of data, the model is called only once for enhancement. In contrast, generative Speech…

音频与语音处理 · 电气工程与系统科学 2025-10-22 Bunlong Lay , Rostislav Makarov , Simon Welker , Maris Hillemann , Timo Gerkmann

Contemporary conversational systems often present a significant limitation: their responses lack the emotional depth and disfluent characteristic of human interactions. This absence becomes particularly noticeable when users seek more…

计算与语言 · 计算机科学 2024-04-03 Rohan Chaudhury , Mihir Godbole , Aakash Garg , Jinsil Hwaryoung Seo

Gesture recognition research, unlike NLP, continues to face acute data scarcity, with progress constrained by the need for costly human recordings or image processing approaches that cannot generate authentic variability in the gestures…

计算机视觉与模式识别 · 计算机科学 2026-04-17 Hassan Ali , Doreen Jirak , Luca Müller , Stefan Wermter

To the best of our knowledge, we first present a live system that generates personalized photorealistic talking-head animation only driven by audio signals at over 30 fps. Our system contains three stages. The first stage is a deep neural…

图形学 · 计算机科学 2021-09-27 Yuanxun Lu , Jinxiang Chai , Xun Cao

Generating natural, correct, and visually smooth 3D avatar sign language motion conditioned on the text inputs continues to be very challenging. In this work, we train a generative model of 3D body motion and explore the role of…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Rui Hong , Jana Kosecka

DiffusionAvatars synthesizes a high-fidelity 3D head avatar of a person, offering intuitive control over both pose and expression. We propose a diffusion-based neural renderer that leverages generic 2D priors to produce compelling images of…

计算机视觉与模式识别 · 计算机科学 2024-04-18 Tobias Kirschstein , Simon Giebenhain , Matthias Nießner

This paper aims to generate physically-layered 3D humans from text prompts. Existing methods either generate 3D clothed humans as a whole or support only tight and simple clothing generation, which limits their applications to virtual…

计算机视觉与模式识别 · 计算机科学 2024-08-22 Yi Wang , Jian Ma , Ruizhi Shao , Qiao Feng , Yu-kun Lai , Kun Li
‹ 上一页 1 8 9 10 下一页 ›