English
Related papers

Related papers: Toffee: Efficient Million-Scale Dataset Constructi…

200 papers

In this paper, we present an empirical study introducing a nuanced evaluation framework for text-to-image (T2I) generative models, applied to human image synthesis. Our framework categorizes evaluations into two distinct groups: first,…

Computer Vision and Pattern Recognition · Computer Science 2024-10-29 Muxi Chen , Yi Liu , Jian Yi , Changran Xu , Qiuxia Lai , Hongliang Wang , Tsung-Yi Ho , Qiang Xu

Efficient image tokenization with high compression ratios remains a critical challenge for training generative models. We present SoftVQ-VAE, a continuous image tokenizer that leverages soft categorical posteriors to aggregate multiple…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Hao Chen , Ze Wang , Xiang Li , Ximeng Sun , Fangyi Chen , Jiang Liu , Jindong Wang , Bhiksha Raj , Zicheng Liu , Emad Barsoum

Instruction tuning has become a foundation for unlocking the capabilities of large-scale pretrained models and improving their performance on complex tasks. Thus, the construction of high-quality instruction datasets is crucial for…

Artificial Intelligence · Computer Science 2026-02-12 Li Du , Hanyu Zhao , Yiming Ju , Tengfei Pan

Subject-driven text-to-image generation models create novel renditions of an input subject based on text prompts. Existing models suffer from lengthy fine-tuning and difficulties preserving the subject fidelity. To overcome these…

Computer Vision and Pattern Recognition · Computer Science 2023-06-23 Dongxu Li , Junnan Li , Steven C. H. Hoi

In this work, we introduce Pixelsmith, a zero-shot text-to-image generative framework to sample images at higher resolutions with a single GPU. We are the first to show that it is possible to scale the output of a pre-trained diffusion…

Computer Vision and Pattern Recognition · Computer Science 2024-10-25 Athanasios Tragakis , Marco Aversa , Chaitanya Kaul , Roderick Murray-Smith , Daniele Faccio

We present Match-and-Fuse - a zero-shot, training-free method for consistent controlled generation of unstructured image sets - collections that share a common visual element, yet differ in viewpoint, time of capture, and surrounding…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Kate Feingold , Omri Kaduri , Tali Dekel

Current image generation models produce visually compelling but scientifically implausible images, exposing a fundamental gap between visual fidelity and physical realism. In this work, we introduce ScienceT2I, an expert-annotated dataset…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Jialuo Li , Wenhao Chai , Xingyu Fu , Haiyang Xu , Saining Xie

The continuous development of foundational models for video generation is evolving into various applications, with subject-consistent video generation still in the exploratory stage. We refer to this as Subject-to-Video, which extracts…

Computer Vision and Pattern Recognition · Computer Science 2025-04-11 Lijie Liu , Tianxiang Ma , Bingchuan Li , Zhuowei Chen , Jiawei Liu , Gen Li , Siyu Zhou , Qian He , Xinglong Wu

Text-to-image (T2I) generation has been actively studied using Diffusion Models and Autoregressive Models. Recently, Masked Generative Transformers have gained attention as an alternative to Autoregressive Models to overcome the inherent…

Computer Vision and Pattern Recognition · Computer Science 2025-08-08 Wonjun Kang , Byeongkeun Ahn , Minjae Lee , Kevin Galim , Seunghyuk Oh , Hyung Il Koo , Nam Ik Cho

Synthetic datasets are widely used for training urban scene recognition models, but even highly realistic renderings show a noticeable gap to real imagery. This gap is particularly pronounced when adapting to a specific target domain, such…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Denis Zavadski , Damjan Kalšan , Tim Küchler , Haebom Lee , Stefan Roth , Carsten Rother

The most advanced text-to-image (T2I) models require significant training costs (e.g., millions of GPU hours), seriously hindering the fundamental innovation for the AIGC community while increasing CO2 emissions. This paper introduces…

Computer Vision and Pattern Recognition · Computer Science 2024-01-01 Junsong Chen , Jincheng Yu , Chongjian Ge , Lewei Yao , Enze Xie , Yue Wu , Zhongdao Wang , James Kwok , Ping Luo , Huchuan Lu , Zhenguo Li

Customized text-to-video generation aims to generate text-guided videos with user-given subjects, which has gained increasing attention. However, existing works are primarily limited to single-subject oriented text-to-video generation,…

Computer Vision and Pattern Recognition · Computer Science 2025-04-15 Hong Chen , Xin Wang , Guanning Zeng , Yipeng Zhang , Yuwei Zhou , Feilin Han , Yaofei Wu , Wenwu Zhu

As scaling laws in generative AI push performance, they also simultaneously concentrate the development of these models among actors with large computational resources. With a focus on text-to-image (T2I) generative models, we aim to…

Computer Vision and Pattern Recognition · Computer Science 2024-07-23 Vikash Sehwag , Xianghao Kong , Jingtao Li , Michael Spranger , Lingjuan Lyu

Recent advancements in generative models have enabled high-fidelity text-to-image generation. However, open-source image-editing models still lag behind their proprietary counterparts, primarily due to limited high-quality data and…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Yang Ye , Xianyi He , Zongjian Li , Bin Lin , Shenghai Yuan , Zhiyuan Yan , Bohan Hou , Li Yuan

Subject-to-Video (S2V) generation aims to create videos that faithfully incorporate reference content, providing enhanced flexibility in the production of videos. To establish the infrastructure for S2V generation, we propose OpenS2V-Nexus,…

Computer Vision and Pattern Recognition · Computer Science 2025-06-04 Shenghai Yuan , Xianyi He , Yufan Deng , Yang Ye , Jinfa Huang , Bin Lin , Jiebo Luo , Li Yuan

Image captioning requires numerous annotated image-text pairs, resulting in substantial annotation costs. Recently, large models (e.g. diffusion models and large language models) have excelled in producing high-quality images and text. This…

Computer Vision and Pattern Recognition · Computer Science 2023-12-20 Feipeng Ma , Yizhou Zhou , Fengyun Rao , Yueyi Zhang , Xiaoyan Sun

Recent years have witnessed the strong power of 3D generation models, which offer a new level of creative flexibility by allowing users to guide the 3D content generation process through a single image or natural language. However, it…

Computer Vision and Pattern Recognition · Computer Science 2024-03-15 Fangfu Liu , Hanyang Wang , Weiliang Chen , Haowen Sun , Yueqi Duan

Despite recent advances in large language models, building dependable and deployable NLP models typically requires abundant, high-quality training data. However, task-specific data is not available for many use cases, and manually curating…

Computation and Language · Computer Science 2024-04-30 Saumya Gandhi , Ritu Gala , Vijay Viswanathan , Tongshuang Wu , Graham Neubig

Current instruction-based image editing (IBIE) methods struggle with challenging editing tasks, as both editing types and sample counts of existing datasets are limited. Moreover, traditional dataset construction often contains noisy…

Computer Vision and Pattern Recognition · Computer Science 2025-09-19 Mingsong Li , Lin Liu , Hongjun Wang , Haoxing Chen , Xijun Gu , Shizhan Liu , Dong Gong , Junbo Zhao , Zhenzhong Lan , Jianguo Li

Training-free consistent text-to-image generation depicting the same subjects across different images is a topic of widespread recent interest. Existing works in this direction predominantly rely on cross-frame self-attention; which…

Computer Vision and Pattern Recognition · Computer Science 2025-04-09 Jaskirat Singh , Junshen Kevin Chen , Jonas Kohler , Michael Cohen