English
Related papers

Related papers: Safe Latent Diffusion: Mitigating Inappropriate De…

200 papers

Diffusion models have recently emerged as powerful generative priors for solving inverse problems. However, training diffusion models in the pixel space are both data-intensive and computationally demanding, which restricts their…

Computer Vision and Pattern Recognition · Computer Science 2024-04-17 Bowen Song , Soo Min Kwon , Zecheng Zhang , Xinyu Hu , Qing Qu , Liyue Shen

Diffusion models are a new class of generative models, and have dramatically promoted image generation with unprecedented quality and diversity. Existing diffusion models mainly try to reconstruct input image from a corrupted one with a…

Computer Vision and Pattern Recognition · Computer Science 2024-06-05 Ling Yang , Jingwei Liu , Shenda Hong , Zhilong Zhang , Zhilin Huang , Zheming Cai , Wentao Zhang , Bin Cui

There is an increasing interest in using image-generating diffusion models for deep data augmentation and image morphing. In this context, it is useful to interpolate between latents produced by inverting a set of input images, in order to…

Computer Vision and Pattern Recognition · Computer Science 2025-05-13 Erik Landolsi , Fredrik Kahl

The rapid adoption of text-to-image diffusion models in society underscores an urgent need to address their biases. Without interventions, these biases could propagate a skewed worldview and restrict opportunities for minority groups. In…

Machine Learning · Computer Science 2024-03-18 Xudong Shen , Chao Du , Tianyu Pang , Min Lin , Yongkang Wong , Mohan Kankanhalli

Diffusion-based models have achieved state-of-the-art performance on text-to-image synthesis tasks. However, one critical limitation of these models is the low fidelity of generated images with respect to the text description, such as…

Computer Vision and Pattern Recognition · Computer Science 2023-04-11 Qiucheng Wu , Yujian Liu , Handong Zhao , Trung Bui , Zhe Lin , Yang Zhang , Shiyu Chang

While latent diffusion models (LDMs), such as Stable Diffusion, are designed for high-resolution (HR) image generation, they often struggle with significant structural distortions when generating images at resolutions higher than their…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Boyuan Cao , Jiaxin Ye , Yujie Wei , Hongming Shan

Text-to-image diffusion models, e.g. Stable Diffusion (SD), lately have shown remarkable ability in high-quality content generation, and become one of the representatives for the recent wave of transformative AI. Nevertheless, such advance…

Computation and Language · Computer Science 2026-01-15 Zhi-Yi Chin , Chieh-Ming Jiang , Ching-Chun Huang , Pin-Yu Chen , Wei-Chen Chiu

Text-to-image diffusion models are a class of deep generative models that have demonstrated an impressive capacity for high-quality image generation. However, these models are susceptible to implicit biases that arise from web-scale…

Computer Vision and Pattern Recognition · Computer Science 2024-01-24 Yinan Zhang , Eric Tzeng , Yilun Du , Dmitry Kislyuk

State-of-the-art Diffusion Models (DMs) produce highly realistic images. While prior work has successfully mitigated Not Safe For Work (NSFW) content in the visual domain, we identify a novel threat: the generation of NSFW text embedded…

Computer Vision and Pattern Recognition · Computer Science 2026-01-16 Aditya Kumar , Tom Blanchard , Adam Dziedzic , Franziska Boenisch

Subject-driven image generation (SDIG) aims to manipulate specific subjects within images while adhering to textual instructions, a task crucial for advancing text-to-image diffusion models. SDIG requires reconciling the tension between…

Computer Vision and Pattern Recognition · Computer Science 2025-09-19 Jibai Lin , Bo Ma , Yating Yang , Xi Zhou , Rong Ma , Turghun Osman , Ahtamjan Ahmat , Rui Dong , Lei Wang

The tremendous progress in neural image generation, coupled with the emergence of seemingly omnipotent vision-language models has finally enabled text-based interfaces for creating and editing images. Handling generic images requires a…

Computer Vision and Pattern Recognition · Computer Science 2023-07-27 Omri Avrahami , Ohad Fried , Dani Lischinski

Recent text-to-image diffusion models can generate striking visuals from text prompts, but they often fail to maintain subject consistency across generations and contexts. One major limitation of current fine-tuning approaches is the…

Computer Vision and Pattern Recognition · Computer Science 2025-09-23 Gordon Chen , Ziqi Huang , Cheston Tan , Ziwei Liu

Large-scale text-to-image diffusion models have achieved great success in synthesizing high-quality and diverse images given target text prompts. Despite the revolutionary image generation ability, current state-of-the-art models still…

Computer Vision and Pattern Recognition · Computer Science 2025-01-20 Jingyuan Zhu , Huimin Ma , Jiansheng Chen , Jian Yuan

Recent progress in text-to-image (TTI) systems, such as StableDiffusion, Imagen, and DALL-E 2, have made it possible to create realistic images with simple text prompts. It is tempting to use these systems to eliminate the manual task of…

Computer Vision and Pattern Recognition · Computer Science 2023-11-02 David Marwood , Shumeet Baluja , Yair Alon

Latent diffusion models (LDMs) dominate high-quality image generation, yet integrating representation learning with generative modeling remains a challenge. We introduce a novel generative image modeling framework that seamlessly bridges…

Computer Vision and Pattern Recognition · Computer Science 2026-01-23 Theodoros Kouzelis , Efstathios Karypidis , Ioannis Kakogeorgiou , Spyros Gidaris , Nikos Komodakis

Classifier guidance -- using the gradients of an image classifier to steer the generations of a diffusion model -- has the potential to dramatically expand the creative control over image generation and editing. However, currently…

Computer Vision and Pattern Recognition · Computer Science 2023-06-02 Bram Wallace , Akash Gokul , Stefano Ermon , Nikhil Naik

Sign language production (SLP) aims to translate spoken language sentences into a sequence of pose frames in a sign language, bridging the communication gap and promoting digital inclusion for deaf and hard-of-hearing communities. Existing…

Computation and Language · Computer Science 2025-09-16 Liqian Feng , Lintao Wang , Kun Hu , Dehui Kong , Zhiyong Wang

We introduce Imagen 3, a latent diffusion model that generates high quality images from text prompts. We describe our quality and responsibility evaluations. Imagen 3 is preferred over other state-of-the-art (SOTA) models at the time of…

Computer Vision and Pattern Recognition · Computer Science 2024-12-24 Imagen-Team-Google , : , Jason Baldridge , Jakob Bauer , Mukul Bhutani , Nicole Brichtova , Andrew Bunner , Lluis Castrejon , Kelvin Chan , Yichang Chen , Sander Dieleman , Yuqing Du , Zach Eaton-Rosen , Hongliang Fei , Nando de Freitas , Yilin Gao , Evgeny Gladchenko , Sergio Gómez Colmenarejo , Mandy Guo , Alex Haig , Will Hawkins , Hexiang Hu , Huilian Huang , Tobenna Peter Igwe , Christos Kaplanis , Siavash Khodadadeh , Yelin Kim , Ksenia Konyushkova , Karol Langner , Eric Lau , Rory Lawton , Shixin Luo , Soňa Mokrá , Henna Nandwani , Yasumasa Onoe , Aäron van den Oord , Zarana Parekh , Jordi Pont-Tuset , Hang Qi , Rui Qian , Deepak Ramachandran , Poorva Rane , Abdullah Rashwan , Ali Razavi , Robert Riachi , Hansa Srinivasan , Srivatsan Srinivasan , Robin Strudel , Benigno Uria , Oliver Wang , Su Wang , Austin Waters , Chris Wolff , Auriel Wright , Zhisheng Xiao , Hao Xiong , Keyang Xu , Marc van Zee , Junlin Zhang , Katie Zhang , Wenlei Zhou , Konrad Zolna , Ola Aboubakar , Canfer Akbulut , Oscar Akerlund , Isabela Albuquerque , Nina Anderson , Marco Andreetto , Lora Aroyo , Ben Bariach , David Barker , Sherry Ben , Dana Berman , Courtney Biles , Irina Blok , Pankil Botadra , Jenny Brennan , Karla Brown , John Buckley , Rudy Bunel , Elie Bursztein , Christina Butterfield , Ben Caine , Viral Carpenter , Norman Casagrande , Ming-Wei Chang , Solomon Chang , Shamik Chaudhuri , Tony Chen , John Choi , Dmitry Churbanau , Nathan Clement , Matan Cohen , Forrester Cole , Mikhail Dektiarev , Vincent Du , Praneet Dutta , Tom Eccles , Ndidi Elue , Ashley Feden , Shlomi Fruchter , Frankie Garcia , Roopal Garg , Weina Ge , Ahmed Ghazy , Bryant Gipson , Andrew Goodman , Dawid Górny , Sven Gowal , Khyatti Gupta , Yoni Halpern , Yena Han , Susan Hao , Jamie Hayes , Jonathan Heek , Amir Hertz , Ed Hirst , Emiel Hoogeboom , Tingbo Hou , Heidi Howard , Mohamed Ibrahim , Dirichi Ike-Njoku , Joana Iljazi , Vlad Ionescu , William Isaac , Reena Jana , Gemma Jennings , Donovon Jenson , Xuhui Jia , Kerry Jones , Xiaoen Ju , Ivana Kajic , Christos Kaplanis , Burcu Karagol Ayan , Jacob Kelly , Suraj Kothawade , Christina Kouridi , Ira Ktena , Jolanda Kumakaw , Dana Kurniawan , Dmitry Lagun , Lily Lavitas , Jason Lee , Tao Li , Marco Liang , Maggie Li-Calis , Yuchi Liu , Javier Lopez Alberca , Matthieu Kim Lorrain , Peggy Lu , Kristian Lum , Yukun Ma , Chase Malik , John Mellor , Thomas Mensink , Inbar Mosseri , Tom Murray , Aida Nematzadeh , Paul Nicholas , Signe Nørly , João Gabriel Oliveira , Guillermo Ortiz-Jimenez , Michela Paganini , Tom Le Paine , Roni Paiss , Alicia Parrish , Anne Peckham , Vikas Peswani , Igor Petrovski , Tobias Pfaff , Alex Pirozhenko , Ryan Poplin , Utsav Prabhu , Yuan Qi , Matthew Rahtz , Cyrus Rashtchian , Charvi Rastogi , Amit Raul , Ali Razavi , Sylvestre-Alvise Rebuffi , Susanna Ricco , Felix Riedel , Dirk Robinson , Pankaj Rohatgi , Bill Rosgen , Sarah Rumbley , Moonkyung Ryu , Anthony Salgado , Tim Salimans , Sahil Singla , Florian Schroff , Candice Schumann , Tanmay Shah , Eleni Shaw , Gregory Shaw , Brendan Shillingford , Kaushik Shivakumar , Dennis Shtatnov , Zach Singer , Evgeny Sluzhaev , Valerii Sokolov , Thibault Sottiaux , Florian Stimberg , Brad Stone , David Stutz , Yu-Chuan Su , Eric Tabellion , Shuai Tang , David Tao , Kurt Thomas , Gregory Thornton , Andeep Toor , Cristian Udrescu , Aayush Upadhyay , Cristina Vasconcelos , Alex Vasiloff , Andrey Voynov , Amanda Walker , Luyu Wang , Miaosen Wang , Simon Wang , Stanley Wang , Qifei Wang , Yuxiao Wang , Ágoston Weisz , Olivia Wiles , Chenxia Wu , Xingyu Federico Xu , Andrew Xue , Jianbo Yang , Luo Yu , Mete Yurtoglu , Ali Zand , Han Zhang , Jiageng Zhang , Catherine Zhao , Adilet Zhaxybay , Miao Zhou , Shengqi Zhu , Zhenkai Zhu , Dawn Bloxwich , Mahyar Bordbar , Luis C. Cobo , Eli Collins , Shengyang Dai , Tulsee Doshi , Anca Dragan , Douglas Eck , Demis Hassabis , Sissie Hsiao , Tom Hume , Koray Kavukcuoglu , Helen King , Jack Krawczyk , Yeqing Li , Kathy Meier-Hellstern , Andras Orban , Yury Pinsky , Amar Subramanya , Oriol Vinyals , Ting Yu , Yori Zwols

Diffusion models generate images with an unprecedented level of quality, but how can we freely rearrange image layouts? Recent works generate controllable scenes via learning spatially disentangled latent codes, but these methods do not…

Computer Vision and Pattern Recognition · Computer Science 2024-04-11 Jiawei Ren , Mengmeng Xu , Jui-Chieh Wu , Ziwei Liu , Tao Xiang , Antoine Toisoul

Text-to-Image (T2I) diffusion models are widely recognized for their ability to generate high-quality and diverse images based on text prompts. However, despite recent advances, these models are still prone to generating unsafe images…

Computer Vision and Pattern Recognition · Computer Science 2024-12-06 Jiangweizhi Peng , Zhiwei Tang , Gaowen Liu , Charles Fleming , Mingyi Hong
‹ Prev 1 3 4 5 6 7 10 Next ›