English
Related papers

Related papers: MagicCraft: Natural Language-Driven Generation of …

200 papers

We introduce WorldGen, a system that enables the automatic creation of large-scale, interactive 3D worlds directly from text prompts. Our approach transforms natural language descriptions into traversable, fully textured environments that…

Recent remarkable advances in large-scale text-to-image diffusion models have inspired a significant breakthrough in text-to-3D generation, pursuing 3D content creation solely from a given text prompt. However, existing text-to-3D…

Computer Vision and Pattern Recognition · Computer Science 2023-11-10 Yang Chen , Yingwei Pan , Yehao Li , Ting Yao , Tao Mei

This paper introduces VisuCraft, a novel framework designed to significantly enhance the capabilities of Large Vision-Language Models (LVLMs) in complex visual-guided creative content generation. Existing LVLMs often exhibit limitations in…

Computer Vision and Pattern Recognition · Computer Science 2025-08-06 Rongxin Jiang , Robert Long , Chenghao Gu , Mingrui Yan

Generative models for 3D object synthesis have seen significant advancements with the incorporation of prior knowledge distilled from 2D diffusion models. Nevertheless, challenges persist in the form of multi-view geometric inconsistencies…

Computer Vision and Pattern Recognition · Computer Science 2023-11-20 Lincong Feng , Muyu Wang , Maoyu Wang , Kuo Xu , Xiaoli Liu

Traditional 3D content creation tools empower users to bring their imagination to life by giving them direct control over a scene's geometry, appearance, motion, and camera path. Creating computer-generated videos, however, is a tedious…

Computer Vision and Pattern Recognition · Computer Science 2023-12-05 Shengqu Cai , Duygu Ceylan , Matheus Gadelha , Chun-Hao Paul Huang , Tuanfeng Yang Wang , Gordon Wetzstein

Augmented reality (AR) has shown promise for supporting Deaf and hard-of-hearing (DHH) individuals by captioning speech and visualizing environmental sounds, yet existing systems do not allow users to create personalized sound…

Human-Computer Interaction · Computer Science 2025-11-25 Jaewook Lee , Davin Win Kyi , Leejun Kim , Jenny Peng , Gagyeom Lim , Jeremy Zhengqi Huang , Dhruv Jain , Jon E. Froehlich

Realistic trajectory generation with natural language control is pivotal for advancing autonomous vehicle technology. However, previous methods focus on individual traffic participant trajectory generation, thus failing to account for the…

Artificial Intelligence · Computer Science 2024-05-27 Junkai Xia , Chenxin Xu , Qingyao Xu , Chen Xie , Yanfeng Wang , Siheng Chen

Human motion generation is a significant pursuit in generative computer vision with widespread applications in film-making, video games, AR/VR, and human-robot interaction. Current methods mainly utilize either diffusion-based generative…

Computer Vision and Pattern Recognition · Computer Science 2025-02-03 Canxuan Gang

Online memes have emerged as powerful digital cultural artifacts in the age of social media, offering not only humor but also platforms for political discourse, social critique, and information dissemination. Their extensive reach and…

Computers and Society · Computer Science 2024-03-25 Han Wang , Roy Ka-Wei Lee

Talking head generation creates lifelike avatars from static portraits for virtual communication and content creation. However, current models do not yet convey the feeling of truly interactive communication, often generating one-way…

Machine Learning · Computer Science 2026-01-05 Taekyung Ki , Sangwon Jang , Jaehyeong Jo , Jaehong Yoon , Sung Ju Hwang

The generation of anchor-style product promotion videos presents promising opportunities in e-commerce, advertising, and consumer engagement. Despite advancements in pose-guided human video generation, creating product promotion videos…

Computer Vision and Pattern Recognition · Computer Science 2025-06-24 Ziyi Xu , Ziyao Huang , Juan Cao , Yong Zhang , Xiaodong Cun , Qing Shuai , Yuchen Wang , Linchao Bao , Jintao Li , Fan Tang

Artistic typography aims to stylize input characters with visual effects that are both creative and legible. Traditional approaches rely heavily on manual design, while recent generative models, particularly diffusion-based methods, have…

Computer Vision and Pattern Recognition · Computer Science 2025-07-15 Zhe Wang , Jingbo Zhang , Tianyi Wei , Wanchao Su , Can Wang

With the emergence of Cloud computing, Internet of Things-enabled Human-Computer Interfaces, Generative Artificial Intelligence, and high-accurate Machine and Deep-learning recognition and predictive models, along with the Post Covid-19…

Human-Computer Interaction · Computer Science 2023-08-25 Leila Ismail , Rajkumar Buyya

With the explosive growth of 3D content creation, there is an increasing demand for automatically converting static 3D models into articulation-ready versions that support realistic animation. Traditional approaches rely heavily on manual…

Computer Vision and Pattern Recognition · Computer Science 2025-02-19 Chaoyue Song , Jianfeng Zhang , Xiu Li , Fan Yang , Yiwen Chen , Zhongcong Xu , Jun Hao Liew , Xiaoyang Guo , Fayao Liu , Jiashi Feng , Guosheng Lin

Generative AI has made significant progress in recent years, with text-guided content generation being the most practical as it facilitates interaction between human instructions and AI-generated content (AIGC). Thanks to advancements in…

Computer Vision and Pattern Recognition · Computer Science 2024-10-28 Chenghao Li , Chaoning Zhang , Joseph Cho , Atish Waghwase , Lik-Hang Lee , Francois Rameau , Yang Yang , Sung-Ho Bae , Choong Seon Hong

There is increased interest in using generative AI to create 3D spaces for Virtual Reality (VR) applications. However, today's models produce artificial environments, falling short of supporting collaborative tasks that benefit from…

Artificial Intelligence · Computer Science 2024-09-24 Nels Numan , Shwetha Rajaram , Balasaravanan Thoravi Kumaravel , Nicolai Marquardt , Andrew D. Wilson

Generating articulated objects, such as laptops and microwaves, is a crucial yet challenging task with extensive applications in Embodied AI and AR/VR. Current image-to-3D methods primarily focus on surface geometry and texture, neglecting…

Computer Vision and Pattern Recognition · Computer Science 2025-07-09 Ruijie Lu , Yu Liu , Jiaxiang Tang , Junfeng Ni , Yuxiang Wang , Diwen Wan , Gang Zeng , Yixin Chen , Siyuan Huang

The AR 3D book has shown significant potential in enhancing students' learning outcomes. However, the creation process of 3D books requires a significant investment of time, effort, and specialized skills. Thus, in this paper, we first…

Human-Computer Interaction · Computer Science 2025-08-12 Yibo Wang , Yuanyuan Mao , Lik-Hang Lee , Shi-ting Ni , Zeyu Wang , Xiaole Gu , Pan Hui

The global metaverse development is facing a "cooldown moment", while the academia and industry attention moves drastically from the Metaverse to AI Generated Content (AIGC) in 2023. Nonetheless, the current discussion rarely considers the…

Human-Computer Interaction · Computer Science 2023-04-19 Lik-Hang Lee , Pengyuan Zhou , Chaoning Zhang , Simo Hosio

Generating coherent and useful image/video scenes from a free-form textual description is technically a very difficult problem to handle. Textual description of the same scene can vary greatly from person to person, or sometimes even for…

Computer Vision and Pattern Recognition · Computer Science 2020-12-01 Faria Huq , Nafees Ahmed , Anindya Iqbal
‹ Prev 1 3 4 5 6 7 10 Next ›