中文
相关论文

相关论文: Zero-Shot Text-to-Parameter Translation for Game C…

200 篇论文

Text-video prediction (TVP) is a downstream video generation task that requires a model to produce subsequent video frames given a series of initial video frames and text describing the required motion. In practice TVP methods focus on a…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Zheyuan Liu , Junyan Wang , Zicheng Duan , Cristian Rodriguez-Opazo , Anton van den Hengel

Personalized portrait synthesis, essential in domains like social entertainment, has recently made significant progress. Person-wise fine-tuning based methods, such as LoRA and DreamBooth, can produce photorealistic outputs but need…

计算机视觉与模式识别 · 计算机科学 2025-03-24 Mengtian Li , Jinshu Chen , Wanquan Feng , Bingchuan Li , Fei Dai , Songtao Zhao , Qian He

The rapid advancement of generative AI has democratized access to powerful tools such as Text-to-Image models. However, to generate high-quality images, users must still craft detailed prompts specifying scene, style, and context-often…

多智能体系统 · 计算机科学 2025-09-25 Dawei Xiang , Wenyan Xu , Kexin Chu , Tianqi Ding , Zixu Shen , Yiming Zeng , Jianchang Su , Wei Zhang

Text-to-video (T2V) generation has advanced rapidly, yet maintaining consistent character identities across scenes remains a major challenge. Existing personalization methods often focus on facial identity but fail to preserve broader…

计算机视觉与模式识别 · 计算机科学 2025-12-09 Ziyang Mai , Yu-Wing Tai

Text-driven motion generation offers a powerful and intuitive way to create human movements directly from natural language. By removing the need for predefined motion inputs, it provides a flexible and accessible approach to controlling…

计算机视觉与模式识别 · 计算机科学 2025-05-15 Ali Rida Sahili , Najett Neji , Hedi Tabia

Current video generation models usually convert signals indicating appearance and motion received from inputs (e.g., image, text) or latent spaces (e.g., noise vectors) into consecutive frames, fulfilling a stochastic generation process for…

计算机视觉与模式识别 · 计算机科学 2022-10-07 Xue Song , Jingjing Chen , Bin Zhu , Yu-Gang Jiang

There is a growing demand for the accessible creation of high-quality 3D avatars that are animatable and customizable. Although 3D morphable models provide intuitive control for editing and animation, and robustness for single-view face…

计算机视觉与模式识别 · 计算机科学 2023-05-05 Connor Z. Lin , Koki Nagano , Jan Kautz , Eric R. Chan , Umar Iqbal , Leonidas Guibas , Gordon Wetzstein , Sameh Khamis

Recent advances in scene-based video generation enable coherent visual narratives from structured prompts, yet a key aspect of storytelling -- character-driven dialogue and speech -- remains underexplored. We present a modular pipeline that…

计算机视觉与模式识别 · 计算机科学 2026-05-20 Taewon Kang , Ming C. Lin

The current text-to-video (T2V) generation has made significant progress in synthesizing realistic general videos, but it is still under-explored in identity-specific human video generation with customized ID images. The key challenge lies…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Hengjia Li , Haonan Qiu , Shiwei Zhang , Xiang Wang , Yujie Wei , Zekun Li , Yingya Zhang , Boxi Wu , Deng Cai

Talking face generation has gained immense popularity in the computer vision community, with various applications including AR, VR, teleconferencing, digital assistants, and avatars. Traditional methods are mainly audio-driven, which have…

计算机视觉与模式识别 · 计算机科学 2024-11-21 Xingjian Diao , Ming Cheng , Wayner Barrios , SouYoung Jin

Prompt-based or in-context learning has achieved high zero-shot performance on many natural language generation (NLG) tasks. Here we explore the performance of prompt-based learning for simultaneously controlling the personality and the…

计算与语言 · 计算机科学 2023-02-09 Angela Ramirez , Mamon Alsalihy , Kartik Aggarwal , Cecilia Li , Liren Wu , Marilyn Walker

Personalized text-to-image generation has attracted unprecedented attention in the recent few years due to its unique capability of generating highly-personalized images via using the input concept dataset and novel textual prompt. However,…

人工智能 · 计算机科学 2024-07-02 Shian Du , Xiaotian Cheng , Qi Qian , Henglu Wei , Yi Xu , Xiangyang Ji

Offline meta-RL usually tackles generalization by inferring task beliefs from high-quality samples or warmup explorations. The restricted form limits their generality and usability since these supervision signals are expensive and even…

人工智能 · 计算机科学 2025-11-25 Shilin Zhang , Zican Hu , Wenhao Wu , Xinyi Xie , Jianxiang Tang , Chunlin Chen , Daoyi Dong , Yu Cheng , Zhenhong Sun , Zhi Wang

In this work, we develop intuitive controls for editing the style of 3D objects. Our framework, Text2Mesh, stylizes a 3D mesh by predicting color and local geometric details which conform to a target text prompt. We consider a disentangled…

计算机视觉与模式识别 · 计算机科学 2021-12-07 Oscar Michel , Roi Bar-On , Richard Liu , Sagie Benaim , Rana Hanocka

Recent advancements in controllable human image generation have led to zero-shot generation using structural signals (e.g., pose, depth) or facial appearance. Yet, generating human images conditioned on multiple parts of human appearance…

计算机视觉与模式识别 · 计算机科学 2024-04-24 Zehuan Huang , Hongxing Fan , Lipeng Wang , Lu Sheng

Text-driven 3D scene generation is widely applicable to video gaming, film industry, and metaverse applications that have a large demand for 3D scenes. However, existing text-to-3D generation methods are limited to producing 3D objects with…

计算机视觉与模式识别 · 计算机科学 2024-02-01 Jingbo Zhang , Xiaoyu Li , Ziyu Wan , Can Wang , Jing Liao

Textual information in a captured scene plays an important role in scene interpretation and decision making. Though there exist methods that can successfully detect and interpret complex text regions present in a scene, to the best of our…

计算机视觉与模式识别 · 计算机科学 2025-02-19 Prasun Roy , Saumik Bhattacharya , Subhankar Ghosh , Umapada Pal

Leveraging the advances of natural language processing, most recent scene text recognizers adopt an encoder-decoder architecture where text images are first converted to representative features and then a sequence of characters via…

计算机视觉与模式识别 · 计算机科学 2022-12-20 Chuhui Xue , Jiaxing Huang , Wenqing Zhang , Shijian Lu , Changhu Wang , Song Bai

Procedural content generation uses algorithmic techniques to create large amounts of new content for games at much lower production costs. In newer approaches, procedural content generation utilizes machine learning. However, these methods…

人工智能 · 计算机科学 2024-07-01 Davor Hafnar , Jure Demšar

Parameter generation has long struggled to match the scale of today large vision and language models, curbing its broader utility. In this paper, we introduce Recurrent Diffusion for Large Scale Parameter Generation (RPG), a novel framework…

机器学习 · 计算机科学 2025-02-12 Kai Wang , Dongwen Tang , Wangbo Zhao , Konstantin Schürholt , Zhangyang Wang , Yang You