中文
相关论文

相关论文: Toward Phonology-Guided Sign Language Motion Gener…

200 篇论文

The automatic generation of stylized co-speech gestures has recently received increasing attention. Previous systems typically allow style control via predefined text labels or example motion clips, which are often not flexible enough to…

计算机视觉与模式识别 · 计算机科学 2023-10-17 Tenglong Ao , Zeyi Zhang , Libin Liu

In this paper, we show different fine-tuning methods for Stable Diffusion XL; this includes inference steps, and caption customization for each image to align with generating images in the style of a commercial 2D icon training set. We also…

计算机视觉与模式识别 · 计算机科学 2024-07-16 Youssef Sultan , Jiangqin Ma , Yu-Ying Liao

Audio-driven video generation aims to synthesize realistic videos that align with input audio recordings, akin to the human ability to visualize scenes from auditory input. However, existing approaches predominantly focus on exploring…

图形学 · 计算机科学 2026-03-17 Kien T. Pham , Yingqing He , Yazhou Xing , Qifeng Chen , Long Chen

Human motion synthesis conditioned on textual input has gained significant attention in recent years due to its potential applications in various domains such as gaming, film production, and virtual reality. Conditioned Motion synthesis…

计算机视觉与模式识别 · 计算机科学 2025-01-06 Avinash Amballa , Gayathri Akkinapalli , Vinitra Muralikrishnan

Text encoders in diffusion models have rapidly evolved, transitioning from CLIP to T5-XXL. Although this evolution has significantly enhanced the models' ability to understand complex prompts and generate text, it also leads to a…

计算机视觉与模式识别 · 计算机科学 2025-03-26 Lifu Wang , Daqing Liu , Xinchen Liu , Xiaodong He

Handwriting stroke generation is crucial for improving the performance of tasks such as handwriting recognition and writers order recovery. In handwriting stroke generation, it is significantly important to imitate the sample calligraphic…

计算机视觉与模式识别 · 计算机科学 2025-09-22 Sidra Hanif , Longin Jan Latecki

Sign Language Translation (SLT) is a challenging task that requires bridging the modality gap between visual and linguistic information while capturing subtle variations in hand shapes and movements. To address these challenges, we…

计算机视觉与模式识别 · 计算机科学 2025-09-03 Sobhan Asasi , Mohamed Ilyas Lakhal , Ozge Mercanoglu Sincan , Richard Bowden

We demonstrate that morphological pressure creates navigable gradients at multiple levels of the text-to-image generative pipeline. In Study~1, identity basins in Stable Diffusion 1.5 can be navigated using morphological descriptors --…

计算机视觉与模式识别 · 计算机科学 2026-02-24 Andrew Fraser

Language plays a vital role in the realm of human motion. Existing methods have largely depended on CLIP text embeddings for motion generation, yet they fall short in effectively aligning language and motion due to CLIP's pretraining on…

计算机视觉与模式识别 · 计算机科学 2025-03-11 Zhe Li , Weihao Yuan , Yisheng He , Lingteng Qiu , Shenhao Zhu , Xiaodong Gu , Weichao Shen , Yuan Dong , Zilong Dong , Laurence T. Yang

Text-guided molecule generation is a task where molecules are generated to match specific textual descriptions. Recently, most existing SMILES-based molecule generation methods rely on an autoregressive architecture. In this work, we…

机器学习 · 计算机科学 2024-02-21 Haisong Gong , Qiang Liu , Shu Wu , Liang Wang

Realistic, high-fidelity 3D facial animations are crucial for expressive avatar systems in human-computer interaction and accessibility. Although prior methods show promising quality, their reliance on the mesh domain limits their ability…

计算机视觉与模式识别 · 计算机科学 2025-07-22 Alexandre Symeonidis-Herzig , Özge Mercanoğlu Sincan , Richard Bowden

Text-guided domain adaptation and generation of 3D-aware portraits find many applications in various fields. However, due to the lack of training data and the challenges in handling the high variety of geometry and appearance, the existing…

计算机视觉与模式识别 · 计算机科学 2024-04-15 Biwen Lei , Kai Yu , Mengyang Feng , Miaomiao Cui , Xuansong Xie

Speech-driven 3D facial animation aims to generate realistic lip movements and facial expressions for 3D head models from arbitrary audio clips. Although existing diffusion-based methods are capable of producing natural motions, their slow…

计算机视觉与模式识别 · 计算机科学 2025-09-05 Xuangeng Chu , Nabarun Goswami , Ziteng Cui , Hanqin Wang , Tatsuya Harada

Denoising Diffusion models have shown remarkable performance in generating diverse, high quality images from text. Numerous techniques have been proposed on top of or in alignment with models like Stable Diffusion and Imagen that generate…

Sign Language (SL) enables two-way communication for the deaf and hard-of-hearing community, yet many sign languages remain under-resourced in the AI space. Sign Language Instruction Generation (SLIG) produces step-by-step textual…

人机交互 · 计算机科学 2025-08-27 Md Tariquzzaman , Md Farhan Ishmam , Saiyma Sittul Muna , Md Kamrul Hasan , Hasan Mahmud

Masked diffusion language models (MDLMs) have emerged as a promising alternative to dominant autoregressive approaches. Although they achieve competitive performance on several tasks, a substantial gap remains in open-ended text generation.…

计算与语言 · 计算机科学 2026-02-02 Mengyu Ye , Ryosuke Takahashi , Keito Kudo , Jun Suzuki

Continuously recognizing sign gestures and converting them to glosses plays a key role in bridging the gap between the hearing and hearing-impaired communities. This involves recognizing and interpreting the hands, face, and body gestures…

计算机视觉与模式识别 · 计算机科学 2025-07-29 Samuel Ebimobowei Johnny , Blessed Guda , Andrew Blayama Stephen , Assane Gueye

Attributes such as style, fine-grained text, and trajectory are specific conditions for describing motion. However, existing methods often lack precise user control over motion attributes and suffer from limited generalizability to unseen…

计算机视觉与模式识别 · 计算机科学 2025-03-17 Mingjie Wei , Xuemei Xie , Guangming Shi

We present SoundLoCD, a novel text-to-sound generation framework, which incorporates a LoRA-based conditional discrete contrastive latent diffusion model. Unlike recent large-scale sound generation models, our model can be efficiently…

声音 · 计算机科学 2024-05-27 Xinlei Niu , Jing Zhang , Christian Walder , Charles Patrick Martin

We address the problem of generating diverse 3D human motions from textual descriptions. This challenging task requires joint modeling of both modalities: understanding and extracting useful human-centric information from the text, and then…

计算机视觉与模式识别 · 计算机科学 2022-07-25 Mathis Petrovich , Michael J. Black , Gül Varol