English
Related papers

Related papers: Kandinsky 3.0 Technical Report

200 papers

Recent breakthroughs in text-to-image synthesis have been driven by diffusion models trained on billions of image-text pairs. Adapting this approach to 3D synthesis would require large-scale datasets of labeled 3D data and efficient…

Computer Vision and Pattern Recognition · Computer Science 2022-09-30 Ben Poole , Ajay Jain , Jonathan T. Barron , Ben Mildenhall

The rapid advancement in image generation models has predominantly been driven by diffusion models, which have demonstrated unparalleled success in generating high-fidelity, diverse images from textual prompts. Despite their success,…

Computer Vision and Pattern Recognition · Computer Science 2024-03-29 Yusuf Dalva , Hidir Yesiltepe , Pinar Yanardag

In this work, we introduce Prometheus, a 3D-aware latent diffusion model for text-to-3D generation at both object and scene levels in seconds. We formulate 3D scene generation as multi-view, feed-forward, pixel-aligned 3D Gaussian…

Computer Vision and Pattern Recognition · Computer Science 2025-01-03 Yuanbo Yang , Jiahao Shao , Xinyang Li , Yujun Shen , Andreas Geiger , Yiyi Liao

We introduce Imagen 3, a latent diffusion model that generates high quality images from text prompts. We describe our quality and responsibility evaluations. Imagen 3 is preferred over other state-of-the-art (SOTA) models at the time of…

Computer Vision and Pattern Recognition · Computer Science 2024-12-24 Imagen-Team-Google , : , Jason Baldridge , Jakob Bauer , Mukul Bhutani , Nicole Brichtova , Andrew Bunner , Lluis Castrejon , Kelvin Chan , Yichang Chen , Sander Dieleman , Yuqing Du , Zach Eaton-Rosen , Hongliang Fei , Nando de Freitas , Yilin Gao , Evgeny Gladchenko , Sergio Gómez Colmenarejo , Mandy Guo , Alex Haig , Will Hawkins , Hexiang Hu , Huilian Huang , Tobenna Peter Igwe , Christos Kaplanis , Siavash Khodadadeh , Yelin Kim , Ksenia Konyushkova , Karol Langner , Eric Lau , Rory Lawton , Shixin Luo , Soňa Mokrá , Henna Nandwani , Yasumasa Onoe , Aäron van den Oord , Zarana Parekh , Jordi Pont-Tuset , Hang Qi , Rui Qian , Deepak Ramachandran , Poorva Rane , Abdullah Rashwan , Ali Razavi , Robert Riachi , Hansa Srinivasan , Srivatsan Srinivasan , Robin Strudel , Benigno Uria , Oliver Wang , Su Wang , Austin Waters , Chris Wolff , Auriel Wright , Zhisheng Xiao , Hao Xiong , Keyang Xu , Marc van Zee , Junlin Zhang , Katie Zhang , Wenlei Zhou , Konrad Zolna , Ola Aboubakar , Canfer Akbulut , Oscar Akerlund , Isabela Albuquerque , Nina Anderson , Marco Andreetto , Lora Aroyo , Ben Bariach , David Barker , Sherry Ben , Dana Berman , Courtney Biles , Irina Blok , Pankil Botadra , Jenny Brennan , Karla Brown , John Buckley , Rudy Bunel , Elie Bursztein , Christina Butterfield , Ben Caine , Viral Carpenter , Norman Casagrande , Ming-Wei Chang , Solomon Chang , Shamik Chaudhuri , Tony Chen , John Choi , Dmitry Churbanau , Nathan Clement , Matan Cohen , Forrester Cole , Mikhail Dektiarev , Vincent Du , Praneet Dutta , Tom Eccles , Ndidi Elue , Ashley Feden , Shlomi Fruchter , Frankie Garcia , Roopal Garg , Weina Ge , Ahmed Ghazy , Bryant Gipson , Andrew Goodman , Dawid Górny , Sven Gowal , Khyatti Gupta , Yoni Halpern , Yena Han , Susan Hao , Jamie Hayes , Jonathan Heek , Amir Hertz , Ed Hirst , Emiel Hoogeboom , Tingbo Hou , Heidi Howard , Mohamed Ibrahim , Dirichi Ike-Njoku , Joana Iljazi , Vlad Ionescu , William Isaac , Reena Jana , Gemma Jennings , Donovon Jenson , Xuhui Jia , Kerry Jones , Xiaoen Ju , Ivana Kajic , Christos Kaplanis , Burcu Karagol Ayan , Jacob Kelly , Suraj Kothawade , Christina Kouridi , Ira Ktena , Jolanda Kumakaw , Dana Kurniawan , Dmitry Lagun , Lily Lavitas , Jason Lee , Tao Li , Marco Liang , Maggie Li-Calis , Yuchi Liu , Javier Lopez Alberca , Matthieu Kim Lorrain , Peggy Lu , Kristian Lum , Yukun Ma , Chase Malik , John Mellor , Thomas Mensink , Inbar Mosseri , Tom Murray , Aida Nematzadeh , Paul Nicholas , Signe Nørly , João Gabriel Oliveira , Guillermo Ortiz-Jimenez , Michela Paganini , Tom Le Paine , Roni Paiss , Alicia Parrish , Anne Peckham , Vikas Peswani , Igor Petrovski , Tobias Pfaff , Alex Pirozhenko , Ryan Poplin , Utsav Prabhu , Yuan Qi , Matthew Rahtz , Cyrus Rashtchian , Charvi Rastogi , Amit Raul , Ali Razavi , Sylvestre-Alvise Rebuffi , Susanna Ricco , Felix Riedel , Dirk Robinson , Pankaj Rohatgi , Bill Rosgen , Sarah Rumbley , Moonkyung Ryu , Anthony Salgado , Tim Salimans , Sahil Singla , Florian Schroff , Candice Schumann , Tanmay Shah , Eleni Shaw , Gregory Shaw , Brendan Shillingford , Kaushik Shivakumar , Dennis Shtatnov , Zach Singer , Evgeny Sluzhaev , Valerii Sokolov , Thibault Sottiaux , Florian Stimberg , Brad Stone , David Stutz , Yu-Chuan Su , Eric Tabellion , Shuai Tang , David Tao , Kurt Thomas , Gregory Thornton , Andeep Toor , Cristian Udrescu , Aayush Upadhyay , Cristina Vasconcelos , Alex Vasiloff , Andrey Voynov , Amanda Walker , Luyu Wang , Miaosen Wang , Simon Wang , Stanley Wang , Qifei Wang , Yuxiao Wang , Ágoston Weisz , Olivia Wiles , Chenxia Wu , Xingyu Federico Xu , Andrew Xue , Jianbo Yang , Luo Yu , Mete Yurtoglu , Ali Zand , Han Zhang , Jiageng Zhang , Catherine Zhao , Adilet Zhaxybay , Miao Zhou , Shengqi Zhu , Zhenkai Zhu , Dawn Bloxwich , Mahyar Bordbar , Luis C. Cobo , Eli Collins , Shengyang Dai , Tulsee Doshi , Anca Dragan , Douglas Eck , Demis Hassabis , Sissie Hsiao , Tom Hume , Koray Kavukcuoglu , Helen King , Jack Krawczyk , Yeqing Li , Kathy Meier-Hellstern , Andras Orban , Yury Pinsky , Amar Subramanya , Oriol Vinyals , Ting Yu , Yori Zwols

Recent advances in image editing with diffusion models have achieved impressive results, offering fine-grained control over the generation process. However, these methods are computationally intensive because of their iterative nature.…

Computer Vision and Pattern Recognition · Computer Science 2025-06-25 Ilia Beletskii , Andrey Kuznetsov , Aibek Alanov

This paper introduces a pioneering 3D volumetric encoder designed for text-to-3D generation. To scale up the training data for the diffusion model, a lightweight network is developed to efficiently acquire feature volumes from multi-view…

Computer Vision and Pattern Recognition · Computer Science 2024-08-14 Zhicong Tang , Shuyang Gu , Chunyu Wang , Ting Zhang , Jianmin Bao , Dong Chen , Baining Guo

Most 3D generation research focuses on up-projecting 2D foundation models into the 3D space, either by minimizing 2D Score Distillation Sampling (SDS) loss or fine-tuning on multi-view datasets. Without explicit 3D priors, these methods…

Computer Vision and Pattern Recognition · Computer Science 2023-12-11 Lihe Ding , Shaocong Dong , Zhanpeng Huang , Zibin Wang , Yiyuan Zhang , Kaixiong Gong , Dan Xu , Tianfan Xue

Text-to-Image (T2I) generation methods based on diffusion model have garnered significant attention in the last few years. Although these image synthesis methods produce visually appealing results, they frequently exhibit spelling errors…

Computer Vision and Pattern Recognition · Computer Science 2023-12-27 Yiming Zhao , Zhouhui Lian

Text-to-Image synthesis is the task of generating an image according to a specific text description. Generative Adversarial Networks have been considered the standard method for image synthesis virtually since their introduction. Denoising…

Computer Vision and Pattern Recognition · Computer Science 2023-05-18 Konstantina Nikolaidou , George Retsinas , Vincent Christlein , Mathias Seuret , Giorgos Sfikas , Elisa Barney Smith , Hamam Mokayed , Marcus Liwicki

Large-scale diffusion-based generative models have led to breakthroughs in text-conditioned high-resolution image synthesis. Starting from random noise, such text-to-image diffusion models gradually synthesize images in an iterative fashion…

Computer Vision and Pattern Recognition · Computer Science 2023-03-15 Yogesh Balaji , Seungjun Nah , Xun Huang , Arash Vahdat , Jiaming Song , Qinsheng Zhang , Karsten Kreis , Miika Aittala , Timo Aila , Samuli Laine , Bryan Catanzaro , Tero Karras , Ming-Yu Liu

One highly promising direction for enabling flexible real-time on-device image editing is utilizing data distillation by leveraging large-scale text-to-image diffusion models to generate paired datasets used for training generative…

Computer Vision and Pattern Recognition · Computer Science 2024-06-04 Yifan Gong , Zheng Zhan , Qing Jin , Yanyu Li , Yerlan Idelbayev , Xian Liu , Andrey Zharkov , Kfir Aberman , Sergey Tulyakov , Yanzhi Wang , Jian Ren

In recent times, the generation of 3D assets from text prompts has shown impressive results. Both 2D and 3D diffusion models can help generate decent 3D objects based on prompts. 3D diffusion models have good 3D consistency, but their…

Computer Vision and Pattern Recognition · Computer Science 2024-05-14 Taoran Yi , Jiemin Fang , Junjie Wang , Guanjun Wu , Lingxi Xie , Xiaopeng Zhang , Wenyu Liu , Qi Tian , Xinggang Wang

Text-to-image diffusion models have demonstrated an impressive ability to produce high-quality outputs. However, they often struggle to accurately follow fine-grained spatial information in an input text. To this end, we propose a…

Computer Vision and Pattern Recognition · Computer Science 2024-10-22 Ran Galun , Sagie Benaim

Distilling 3D representations from pretrained 2D diffusion models is essential for 3D creative applications across gaming, film, and interior design. Current SDS-based methods are hindered by inefficient information distillation from…

Computer Vision and Pattern Recognition · Computer Science 2025-03-13 Haoran Li , Yuli Tian , Yonghui Wang , Yong Liao , Lin Wang , Yuyang Wang , Peng Yuan Zhou

The recent wave of large-scale text-to-image diffusion models has dramatically increased our text-based image generation abilities. These models can generate realistic images for a staggering variety of prompts and exhibit impressive…

Machine Learning · Computer Science 2023-09-14 Alexander C. Li , Mihir Prabhudesai , Shivam Duggal , Ellis Brown , Deepak Pathak

We introduce RealmDreamer, a technique for generating forward-facing 3D scenes from text descriptions. Our method optimizes a 3D Gaussian Splatting representation to match complex text prompts using pretrained diffusion models. Our key…

Computer Vision and Pattern Recognition · Computer Science 2025-03-12 Jaidev Shriram , Alex Trevithick , Lingjie Liu , Ravi Ramamoorthi

Personalized text-to-image models allow users to generate varied styles of images (specified with a sentence) for an object (specified with a set of reference images). While remarkable results have been achieved using diffusion-based…

Computer Vision and Pattern Recognition · Computer Science 2024-07-19 Fanyue Wei , Wei Zeng , Zhenyang Li , Dawei Yin , Lixin Duan , Wen Li

Text-to-image generative models have made remarkable advancements in generating high-quality images. However, generated images often contain undesirable artifacts or other errors due to model limitations. Existing techniques to fine-tune…

Computer Vision and Pattern Recognition · Computer Science 2023-10-31 Peyman Gholami , Robert Xiao

Diffusion models have emerged as a dominant paradigm for generative modeling across a wide range of domains, including prompt-conditional generation. The vast majority of samplers, however, rely on forward discretization of the reverse…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Zhenghan Fang , Jian Zheng , Qiaozi Gao , Xiaofeng Gao , Jeremias Sulam

The growing adoption of generative AI in real-world applications has exposed a critical bottleneck in the computational demands of diffusion-based text-to-image models. In this work, we propose KDC-Diff, a novel and scalable generative…

Computer Vision and Pattern Recognition · Computer Science 2025-10-01 Md. Naimur Asif Borno , Md Sakib Hossain Shovon , Asmaa Soliman Al-Moisheer , Mohammad Ali Moni