English
Related papers

Related papers: SpatialLock: Precise Spatial Control in Text-to-Im…

200 papers

Spatial Description Resolution, as a language-guided localization task, is proposed for target location in a panoramic street view, given corresponding language descriptions. Explicitly characterizing an object-level relationship while…

Computer Vision and Pattern Recognition · Computer Science 2020-10-28 Peiyao Wang , Weixin Luo , Yanyu Xu , Haojie Li , Shugong Xu , Jianyu Yang , Shenghua Gao

Controlling robots to perform tasks via natural language is one of the most challenging topics in human-robot interaction. In this work, we present a robot system that follows unconstrained language instructions to pick and place arbitrary…

Robotics · Computer Science 2021-02-17 Oier Mees , Wolfram Burgard

Text-to-image (T2I) diffusion models have revolutionized generative modeling by producing high-fidelity, diverse, and visually realistic images from textual prompts. Despite these advances, existing models struggle with complex prompts…

Computer Vision and Pattern Recognition · Computer Science 2024-11-27 Eric Hanchen Jiang , Yasi Zhang , Zhi Zhang , Yixin Wan , Andrew Lizarraga , Shufan Li , Ying Nian Wu

Text-to-Image (T2I) diffusion models have achieved remarkable success in image generation. Despite their progress, challenges remain in both prompt-following ability, image quality and lack of high-quality datasets, which are essential for…

Computer Vision and Pattern Recognition · Computer Science 2025-01-03 Jingkun An , Yinghao Zhu , Zongjian Li , Enshen Zhou , Haoran Feng , Xijie Huang , Bohua Chen , Yemin Shi , Chengwei Pan

Text-to-image generation intends to automatically produce a photo-realistic image, conditioned on a textual description. It can be potentially employed in the field of art creation, data augmentation, photo-editing, etc. Although many…

Computer Vision and Pattern Recognition · Computer Science 2022-03-01 Zhenxing Zhang , Lambert Schomaker

The automated generation of layouts is vital for embodied intelligence and autonomous systems, supporting applications from virtual environment construction to home robot deployment. Current approaches, however, suffer from spatial…

This paper introduces PoseLess, a novel framework for robot hand control that eliminates the need for explicit pose estimation by directly mapping 2D images to joint angles using projected representations. Our approach leverages synthetic…

Robotics · Computer Science 2025-03-12 Alan Dao , Dinh Bach Vu , Tuan Le Duc Anh , Bui Quang Huy

Current text-to-image generation models often struggle to follow textual instructions, especially the ones requiring spatial reasoning. On the other hand, Large Language Models (LLMs), such as GPT-4, have shown remarkable precision in…

Computer Vision and Pattern Recognition · Computer Science 2023-05-31 Tianjun Zhang , Yi Zhang , Vibhav Vineet , Neel Joshi , Xin Wang

Recent advances in diffusion models have enabled high-quality generation and manipulation of images guided by texts, as well as concept learning from images. However, naive applications of existing methods to editing tasks that require…

Computer Vision and Pattern Recognition · Computer Science 2025-12-29 Xudong Liu , Zikun Chen , Ruowei Jiang , Ziyi Wu , Kejia Yin , Han Zhao , Parham Aarabi , Igor Gilitschenski

Text-to-image (T2I) models can generate not-safe-for-work (NSFW) content, motivating multi-stage safety pipelines with both text and image filters. Newer LLM-based filters detect latent intent beyond keywords, making token-level…

Machine Learning · Computer Science 2026-05-26 Zixuan Chen , Hao Lin , Ke Xu , Xinghao Jiang , Tanfeng Sun

Precise spatial fidelity in Image-to-3D multi-instance generation is critical for downstream real-world applications. Recent work attempts to address this by fine-tuning pre-trained Image-to-3D (I23D) models on multi-instance datasets,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Xiao Cai , Lianli Gao , Pengpeng Zeng , Ji Zhang , Heng Tao Shen , Jingkuan Song

Integration of Large Language Models (LLMs) into visual domain tasks, resulting in visual-LLMs (V-LLMs), has enabled exceptional performance in vision-language tasks, particularly for visual question answering (VQA). However, existing…

Computer Vision and Pattern Recognition · Computer Science 2024-04-12 Kanchana Ranasinghe , Satya Narayan Shukla , Omid Poursaeed , Michael S. Ryoo , Tsung-Yu Lin

Pose-guided person image generation and animation aim to transform a source person image to target poses. These tasks require spatial manipulation of source data. However, Convolutional Neural Networks are limited by the lack of ability to…

Computer Vision and Pattern Recognition · Computer Science 2021-12-01 Yurui Ren , Ge Li , Shan Liu , Thomas H. Li

Recent advancements in video generation, particularly in diffusion models, have driven notable progress in text-to-video (T2V) and image-to-video (I2V) synthesis. However, challenges remain in effectively integrating dynamic motion signals…

Computer Vision and Pattern Recognition · Computer Science 2025-07-04 Ziye Li , Hao Luo , Xincheng Shuai , Henghui Ding

Recent text-to-image diffusion models have demonstrated an astonishing capacity to generate high-quality images. However, researchers mainly studied the way of synthesizing images with only text prompts. While some works have explored using…

Computer Vision and Pattern Recognition · Computer Science 2023-08-22 Jinheng Xie , Yuexiang Li , Yawen Huang , Haozhe Liu , Wentian Zhang , Yefeng Zheng , Mike Zheng Shou

Typical methods for text-to-image synthesis seek to design effective generative architecture to model the text-to-image mapping directly. It is fairly arduous due to the cross-modality translation. In this paper we circumvent this problem…

Computer Vision and Pattern Recognition · Computer Science 2020-07-14 Jiadong Liang , Wenjie Pei , Feng Lu

Text-to-image (T2I) diffusion models, when fine-tuned on a few personal images, can generate visuals with a high degree of consistency. However, such fine-tuned models are not robust; they often fail to compose with concepts of pretrained…

Computer Vision and Pattern Recognition · Computer Science 2024-12-13 Kyungmin Lee , Sangkyung Kwak , Kihyuk Sohn , Jinwoo Shin

How would a static scene react to a local poke? What are the effects on other parts of an object if you could locally push it? There will be distinctive movement, despite evident variations caused by the stochastic nature of our world.…

Computer Vision and Pattern Recognition · Computer Science 2021-10-07 Andreas Blattmann , Timo Milbich , Michael Dorkenwald , Björn Ommer

Multimodal large language models (MLLMs) have made significant advancements in vision understanding and reasoning. However, the autoregressive Transformer architecture used by MLLMs requries tokenization on input images, which limits their…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Xiangxuan Ren , Zhongdao Wang , Liping Hou , Pin Tang , Guoqing Wang , Chao Ma

In this paper, we conduct a study on the state-of-the-art methods for text-to-image synthesis and propose a framework to evaluate these methods. We consider syntheses where an image contains a single or multiple objects. Our study outlines…

Computer Vision and Pattern Recognition · Computer Science 2022-07-20 Tan M. Dinh , Rang Nguyen , Binh-Son Hua
‹ Prev 1 8 9 10 Next ›