中文
相关论文

相关论文: OmniBooth: Learning Latent Control for Image Synth…

200 篇论文

Recent multimodal face generation models address the spatial control limitations of text-to-image diffusion models by augmenting text-based conditioning with spatial priors such as segmentation masks, sketches, or edge maps. This multimodal…

计算机视觉与模式识别 · 计算机科学 2026-04-01 Bharath Krishnamurthy , Ajita Rattani

In this work, we study Human-Object Interaction Video Generation (HOIVG), which aims to synthesize high-quality human-object interaction videos conditioned on text, reference images, audio, and pose. This task holds significant practical…

计算机视觉与模式识别 · 计算机科学 2026-04-20 Donghao Zhou , Guisheng Liu , Hao Yang , Jiatong Li , Jingyu Lin , Xiaohu Huang , Yichen Liu , Xin Gao , Cunjian Chen , Shilei Wen , Chi-Wing Fu , Pheng-Ann Heng

Research in vision-language models has seen rapid developments off-late, enabling natural language-based interfaces for image generation and manipulation. Many existing text guided manipulation techniques are restricted to specific classes…

计算机视觉与模式识别 · 计算机科学 2024-05-07 Paramanand Chandramouli , Kanchana Vaishnavi Gandikota

Most testbeds for omni-modal models assess multimodal understanding via textual outputs, leaving it unclear whether these models can properly speak their answers. To study this, we introduce OmniACBench, a benchmark for evaluating…

计算与语言 · 计算机科学 2026-03-26 Seunghee Kim , Bumkyu Park , Kyudan Jung , Joosung Lee , Soyoon Kim , Jeonghoon Kim , Taeuk Kim , Hwiyeol Jo

Conditional human motion generation remains a fundamental challenge in computer vision and robotics. Despite significant progress, current methods are often constrained by fixed modality configurations and task-specific architectures,…

计算机视觉与模式识别 · 计算机科学 2026-05-29 Yiheng Li , Zhuo Li , Ruibing Hou , Yingjie Chen , Hong Chang , Hao Liu , Shiguang Shan

Customized video generation aims to produce videos that faithfully preserve the subject's appearance from reference images while maintaining temporally consistent motion from reference videos. Existing methods struggle to ensure both…

计算机视觉与模式识别 · 计算机科学 2025-12-16 Xuancheng Xu , Yaning Li , Sisi You , Bing-Kun Bao

Filmmaking and animation production often require sophisticated techniques for coordinating camera transitions and object movements, typically involving labor-intensive real-world capturing. Despite advancements in generative AI for video…

计算机视觉与模式识别 · 计算机科学 2024-06-24 Yaowei Li , Xintao Wang , Zhaoyang Zhang , Zhouxia Wang , Ziyang Yuan , Liangbin Xie , Yuexian Zou , Ying Shan

We present JoyAI-Image, a unified multimodal foundation model for visual understanding, text-to-image generation, and instruction-guided image editing. JoyAI-Image couples a spatially enhanced Multimodal Large Language Model (MLLM) with a…

Recent progress in generative models has significantly advanced image editing capabilities, yet precise and intuitive user control remains difficult. Specifically, users often struggle to communicate both exact spatial layouts and specific…

计算机视觉与模式识别 · 计算机科学 2026-05-06 Anya Ji , George Ma , Téa Wright , Yiming Zhang , David M. Chan , Alane Suhr , Somayeh Sojoudi

Semantic image synthesis is a process for generating photorealistic images from a single semantic mask. To enrich the diversity of multimodal image synthesis, previous methods have controlled the global appearance of an output image by…

计算机视觉与模式识别 · 计算机科学 2021-06-30 Yuki Endo , Yoshihiro Kanamori

Embodied navigation presents a core challenge for intelligent robots, requiring the comprehension of visual environments, natural language instructions, and autonomous exploration. Existing models often fall short in offering a unified…

机器人学 · 计算机科学 2026-01-08 Xinda Xue , Junjun Hu , Minghua Luo , Shichao Xie , Jintao Chen , Zixun Xie , Kuichen Quan , Wei Guo , Mu Xu , Zedong Chu

Recent research on representation learning has proved the merits of multi-modal clues for robust semantic segmentation. Nevertheless, a flexible pretrain-and-finetune pipeline for multiple visual modalities remains unexplored. In this…

计算机视觉与模式识别 · 计算机科学 2025-09-19 Bo-Wen Yin , Jiao-Long Cao , Xuying Zhang , Yuming Chen , Ming-Ming Cheng , Qibin Hou

Recent advancements in foundation models have revolutionized joint audio-video generation. However, existing approaches typically treat human-centric tasks including reference-based audio-video generation (R2AV), video editing (RV2AV) and…

计算机视觉与模式识别 · 计算机科学 2026-02-13 Xu Guo , Fulong Ye , Qichao Sun , Liyang Chen , Bingchuan Li , Pengze Zhang , Jiawei Liu , Songtao Zhao , Qian He , Xiangwang Hou

Mitigating biases in generative AI and, particularly in text-to-image models, is of high importance given their growing implications in society. The biased datasets used for training pose challenges in ensuring the responsible development…

计算机视觉与模式识别 · 计算机科学 2024-06-11 Carolina Lopez Olmos , Alexandros Neophytou , Sunando Sengupta , Dim P. Papadopoulos

We present DreamBooth3D, an approach to personalize text-to-3D generative models from as few as 3-6 casually captured images of a subject. Our approach combines recent advances in personalizing text-to-image models (DreamBooth) with…

Text-to-image generation models represent the next step of evolution in image synthesis, offering a natural way to achieve flexible yet fine-grained control over the result. One emerging area of research is the fast adaptation of large…

计算机视觉与模式识别 · 计算机科学 2023-11-02 Anton Voronov , Mikhail Khoroshikh , Artem Babenko , Max Ryabinin

Text-editable and pose-controllable character video generation is a challenging but prevailing topic with practical applications. However, existing approaches mainly focus on single-object video generation with pose guidance, ignoring the…

计算机视觉与模式识别 · 计算机科学 2024-12-30 Beiyuan Zhang , Yue Ma , Chunlei Fu , Xinyang Song , Zhenan Sun , Ziqiang Li

Existing multi-turn image editing paradigms are often confined to isolated single-step execution. Due to a lack of context-awareness and closed-loop feedback mechanisms, they are prone to error accumulation and semantic drift during…

图形学 · 计算机科学 2026-04-01 Fei Shen , Chengyu Xie , Lihong Wang , Zhanyi Zhang , Xin Jiang , Xiaoyu Du , Jinhui Tang

Recent advancements in image-conditioned image generation have demonstrated substantial progress. However, foreground-conditioned image generation remains underexplored, encountering challenges such as compromised object integrity,…

计算机视觉与模式识别 · 计算机科学 2025-02-25 Tianyidan Xie , Rui Ma , Qian Wang , Xiaoqian Ye , Feixuan Liu , Ying Tai , Zhenyu Zhang , Lanjun Wang , Zili Yi

While language-guided image manipulation has made remarkable progress, the challenge of how to instruct the manipulation process faithfully reflecting human intentions persists. An accurate and comprehensive description of a manipulation…

计算机视觉与模式识别 · 计算机科学 2023-08-03 Yasheng Sun , Yifan Yang , Houwen Peng , Yifei Shen , Yuqing Yang , Han Hu , Lili Qiu , Hideki Koike