English
Related papers

Related papers: simpleposter: a simple baseline for product poster…

200 papers

Sparse-view reconstruction models typically require precise camera poses, yet obtaining these parameters from sparse-view images remains challenging. We introduce FreeSplatter, a scalable feed-forward framework that generates high-quality…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Jiale Xu , Shenghua Gao , Ying Shan

Image inpainting is a fundamental research area between image editing and image generation. Recent state-of-the-art (SOTA) methods have explored novel attention mechanisms, lightweight architectures, and context-aware modeling,…

Computer Vision and Pattern Recognition · Computer Science 2025-05-01 Ziyang Xu , Kangsheng Duan , Xiaolei Shen , Zhifeng Ding , Wenyu Liu , Xiaohu Ruan , Xiaoxin Chen , Xinggang Wang

In this work, we propose a complete framework that generates visual art. Unlike previous stylization methods that are not flexible with style parameters (i.e., they allow stylization with only one style image, a single stylization text or…

Computer Vision and Pattern Recognition · Computer Science 2025-08-08 Marian Lupascu , Ryan Murdock , Ionut Mironica , Yijun Li

We propose a new method for object pose estimation without CAD models. The previous feature-matching-based method OnePose has shown promising results under a one-shot setting which eliminates the need for CAD models or object-specific…

Computer Vision and Pattern Recognition · Computer Science 2023-01-19 Xingyi He , Jiaming Sun , Yuang Wang , Di Huang , Hujun Bao , Xiaowei Zhou

The existing auto-encoder based face pose editing methods primarily focus on modeling the identity preserving ability during pose synthesis, but are less able to preserve the image style properly, which refers to the color, brightness,…

Computer Vision and Pattern Recognition · Computer Science 2021-06-16 Xiangnan Yin , Di Huang , Hongyu Yang , Zehua Fu , Yunhong Wang , Liming Chen

This paper presents a novel framework for converting 2D videos to immersive stereoscopic 3D, addressing the growing demand for 3D content in immersive experience. Leveraging foundation models as priors, our approach overcomes the…

Computer Vision and Pattern Recognition · Computer Science 2024-09-12 Sijie Zhao , Wenbo Hu , Xiaodong Cun , Yong Zhang , Xiaoyu Li , Zhe Kong , Xiangjun Gao , Muyao Niu , Ying Shan

Subject-driven text-to-image generation still struggles to preserve high-frequency identity details such as logos, patterns, and text. Existing methods typically operate directly in RGB space, which often leads to detail degradation under…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Hanzhong Guo , Yizhou Yu

Academic posters are vital for scholarly communication, yet their manual creation is time-consuming. However, automated academic poster generation faces significant challenges in preserving intricate scientific details and achieving…

Computation and Language · Computer Science 2025-05-26 Tao Sun , Enhao Pan , Zhengkai Yang , Kaixin Sui , Jiajun Shi , Xianfu Cheng , Tongliang Li , Wenhao Huang , Ge Zhang , Jian Yang , Zhoujun Li

Spatial control is a core capability in controllable image generation. Advancements in layout-guided image generation have shown promising results on in-distribution (ID) datasets with similar spatial configurations. However, it is unclear…

Computer Vision and Pattern Recognition · Computer Science 2024-07-23 Jaemin Cho , Linjie Li , Zhengyuan Yang , Zhe Gan , Lijuan Wang , Mohit Bansal

Rendering is the process of generating 2D images from 3D assets, simulated in a virtual environment, typically with a graphics pipeline. By inverting such renderer, one can think of a learning approach to predict a 3D shape from an input…

Computer Vision and Pattern Recognition · Computer Science 2019-01-24 Shichen Liu , Weikai Chen , Tianye Li , Hao Li

Pixel synthesis is a promising research paradigm for image generation, which can well exploit pixel-wise prior knowledge for generation. However, existing methods still suffer from excessive memory footprint and computation overhead. In…

Computer Vision and Pattern Recognition · Computer Science 2022-07-28 Jing He , Yiyi Zhou , Qi Zhang , Jun Peng , Yunhang Shen , Xiaoshuai Sun , Chao Chen , Rongrong Ji

Existing preference optimization objectives for language model alignment require additional hyperparameters that must be extensively tuned to achieve optimal performance, increasing both the complexity and time required for fine-tuning…

Machine Learning · Computer Science 2025-02-21 Teng Xiao , Yige Yuan , Zhengyu Chen , Mingxiao Li , Shangsong Liang , Zhaochun Ren , Vasant G Honavar

Recent years have witnessed significant progress in generative models for music, featuring diverse architectures that balance output quality, diversity, speed, and user control. This study explores a user-friendly graphical interface…

Sound · Computer Science 2024-07-02 Scott H. Hawley

This paper proposes a method for generating images of customized objects specified by users. The method is based on a general framework that bypasses the lengthy optimization required by previous approaches, which often employ a per-object…

Computer Vision and Pattern Recognition · Computer Science 2023-04-06 Xuhui Jia , Yang Zhao , Kelvin C. K. Chan , Yandong Li , Han Zhang , Boqing Gong , Tingbo Hou , Huisheng Wang , Yu-Chuan Su

In the realm of image generation, creating customized images from visual prompt with additional textual instruction emerges as a promising endeavor. However, existing methods, both tuning-based and tuning-free, struggle with interpreting…

Computer Vision and Pattern Recognition · Computer Science 2024-10-29 Jing He , Haodong Li , Yongzhe Hu , Guibao Shen , Yingjie Cai , Weichao Qiu , Ying-Cong Chen

To enhance controllability in text-to-image generation, ControlNet introduces image-based control signals, while ControlNet++ improves pixel-level cycle consistency between generated images and the input control signal. To avoid the…

Computer Vision and Pattern Recognition · Computer Science 2025-11-10 Zonglin Lyu , Ming Li , Xinxin Liu , Chen Chen

Layout plays a crucial role in graphic design and poster generation. Recently, the application of deep learning models for layout generation has gained significant attention. This paper focuses on using a GAN-based model conditioned on…

Machine Learning · Computer Science 2026-04-10 Chenchen Xu , Min Zhou , Tiezheng Ge , Weiwei Xu

Generating diverse VLSI layout patterns is essential for various downstream tasks in design for manufacturing, as design rules continually evolve during the development of new technology nodes. However, existing training-based methods for…

Computer Vision and Pattern Recognition · Computer Science 2025-04-29 Guanglei Zhou , Bhargav Korrapati , Gaurav Rajavendra Reddy , Chen-Chia Chang , Jingyu Pan , Jiang Hu , Yiran Chen , Dipto G. Thakurta

This paper introduces a novel approach to aesthetic quality improvement in pre-trained text-to-image diffusion models when given a simple prompt. Our method, dubbed Prompt Embedding Optimization (PEO), leverages a pre-trained text-to-image…

Computer Vision and Pattern Recognition · Computer Science 2025-10-06 Hovhannes Margaryan , Bo Wan , Tinne Tuytelaars

Diffusion-based text-to-image generation has advanced significantly, yet customizing scenes with multiple distinct subjects while maintaining fine-grained control over their interactions remains challenging. Existing methods often struggle…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Pengxiang Cai , Mengyang Li