English
Related papers

Related papers: Zero-Shot Text-to-Parameter Translation for Game C…

200 papers

In the paradigm of AI-generated content (AIGC), there has been increasing attention to transferring knowledge from pre-trained text-to-image (T2I) models to text-to-video (T2V) generation. Despite their effectiveness, these frameworks face…

Computer Vision and Pattern Recognition · Computer Science 2024-02-07 Susung Hong , Junyoung Seo , Heeseong Shin , Sunghwan Hong , Seungryong Kim

In this paper we present the Process-To-Text (P2T) framework for the automatic generation of textual descriptive explanations of processes. P2T integrates three AI paradigms: process mining for extracting temporal and structural information…

Computation and Language · Computer Science 2023-05-24 Yago Fontenla-Seco , Alberto Bugarín-Diz , Manuel Lama

Novel text-to-speech systems can generate entirely new voices that were not seen during training. However, it remains a difficult task to efficiently create personalized voices from a high-dimensional speaker space. In this work, we use…

This paper presents a novel method to manipulate the visual appearance (pose and attribute) of a person image according to natural language descriptions. Our method can be boiled down to two stages: 1) text guided pose generation and 2)…

Computer Vision and Pattern Recognition · Computer Science 2019-04-11 Xingran Zhou , Siyu Huang , Bin Li , Yingming Li , Jiachen Li , Zhongfei Zhang

Adaptive text to speech (TTS) can synthesize new voices in zero-shot scenarios efficiently, by using a well-trained source TTS model without adapting it on the speech data of new speakers. Considering seen and unseen speakers have diverse…

Audio and Speech Processing · Electrical Eng. & Systems 2022-04-04 Yihan Wu , Xu Tan , Bohan Li , Lei He , Sheng Zhao , Ruihua Song , Tao Qin , Tie-Yan Liu

We propose EditID, a training-free approach based on the DiT architecture, which achieves highly editable customized IDs for text to image generation. Existing text-to-image models for customized IDs typically focus more on ID consistency…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Guandong Li , Zhaobin Chu

Robot design has traditionally been costly and labor-intensive. Despite advancements in automated processes, it remains challenging to navigate a vast design space while producing physically manufacturable robots. We introduce Text2Robot, a…

Robotics · Computer Science 2025-02-27 Ryan P. Ringel , Zachary S. Charlick , Jiaxun Liu , Boxi Xia , Boyuan Chen

With the growing popularity of personalized human content creation and sharing, there is a rising demand for advanced techniques in customized human image generation. However, current methods struggle to simultaneously maintain the fidelity…

Graphics · Computer Science 2025-02-21 Ye Wang , Xuping Xie , Lanjun Wang , Zili Yi , Rui Ma

Text-to-Image (T2I) diffusion models have recently gained traction for their versatility and user-friendliness in 2D content generation and editing. However, training a diffusion model specifically for 3D scene editing is challenging due to…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Nazmul Karim , Hasan Iqbal , Umar Khalid , Jing Hua , Chen Chen

Speech synthesis models convert written text into natural-sounding audio. While earlier models were limited to a single speaker, recent advancements have led to the development of zero-shot systems that generate realistic speech from a wide…

Sound · Computer Science 2025-02-12 Łukasz Bondaruk , Jakub Kubiak

Both zero-shot and tuning-based customized text-to-image (CT2I) generation have made significant progress for storytelling content creation. In contrast, research on customized text-to-video (CT2V) generation remains relatively limited.…

Computer Vision and Pattern Recognition · Computer Science 2025-05-13 Panwen Hu , Jiehui Huang , Qiang Sun , Xiaodan Liang

We propose a weakly supervised approach for creating maps using free-form textual descriptions. We refer to this work of creating textual maps as zero-shot mapping. Prior works have approached mapping tasks by developing models that predict…

Computer Vision and Pattern Recognition · Computer Science 2024-04-15 Aayush Dhakal , Adeel Ahmad , Subash Khanal , Srikumar Sastry , Hannah Kerner , Nathan Jacobs

Subject-driven image generation aims to synthesize novel scenes that faithfully preserve subject identity from reference images while adhering to textual guidance. However, existing methods struggle with a critical trade-off between…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Zebin Yao , Lei Ren , Huixing Jiang , Wei Chen , Xiaojie Wang , Ruifan Li , Fangxiang Feng

Due to the increasing demand in films and games, synthesizing 3D avatar animation has attracted much attention recently. In this work, we present a production-ready text/speech-driven full-body animation synthesis system. Given the text and…

Graphics · Computer Science 2022-06-01 Wenlin Zhuang , Jinwei Qi , Peng Zhang , Bang Zhang , Ping Tan

Texturing 3D humans with semantic UV maps remains a challenge due to the difficulty of acquiring reasonably unfolded UV. Despite recent text-to-3D advancements in supervising multi-view renderings using large text-to-image (T2I) models,…

Computer Vision and Pattern Recognition · Computer Science 2024-03-20 Yufei Liu , Junwei Zhu , Junshu Tang , Shijie Zhang , Jiangning Zhang , Weijian Cao , Chengjie Wang , Yunsheng Wu , Dongjin Huang

Large-scale text-to-image generative models have shown their remarkable ability to synthesize diverse and high-quality images. However, it is still challenging to directly apply these models for editing real images for two reasons. First,…

Computer Vision and Pattern Recognition · Computer Science 2023-02-07 Gaurav Parmar , Krishna Kumar Singh , Richard Zhang , Yijun Li , Jingwan Lu , Jun-Yan Zhu

Current learning-based subject customization approaches, predominantly relying on U-Net architectures, suffer from limited generalization ability and compromised image quality. Meanwhile, optimization-based methods require subject-specific…

Computer Vision and Pattern Recognition · Computer Science 2025-04-18 Jiale Tao , Yanbing Zhang , Qixun Wang , Yiji Cheng , Haofan Wang , Xu Bai , Zhengguang Zhou , Ruihuang Li , Linqing Wang , Chunyu Wang , Qin Lin , Qinglin Lu

Character line drawing synthesis can be formulated as a special case of image-to-image translation problem that automatically manipulates the photo-to-line drawing style transformation. In this paper, we present the first generative…

Multimedia · Computer Science 2023-06-16 Cheng-Yu Fang , Xian-Feng Han

Generating speech from a face image is crucial for developing virtual humans capable of interacting using their unique voices, without relying on pre-recorded human speech. In this paper, we propose Face-StyleSpeech, a zero-shot…

Computer Vision and Pattern Recognition · Computer Science 2024-12-31 Minki Kang , Wooseok Han , Eunho Yang

Text-to-image models have made significant strides, producing impressive results in generating images from textual descriptions. However, creating a scalable pipeline for deploying these models in production remains a challenge. Achieving…

Computer Vision and Pattern Recognition · Computer Science 2026-02-17 Parmida Atighehchian , Henry Wang , Andrei Kapustin , Boris Lerner , Tiancheng Jiang , Taylor Jensen , Negin Sokhandan
‹ Prev 1 3 4 5 6 7 10 Next ›