English
Related papers

Related papers: L3GO: Language Agents with Chain-of-3D-Thoughts fo…

200 papers

Recent advances in imitation learning for 3D robotic manipulation have shown promising results with diffusion-based policies. However, achieving human-level dexterity requires seamless integration of geometric precision and semantic…

Generating high-quality 3D objects from textual descriptions remains a challenging problem due to computational cost, the scarcity of 3D data, and complex 3D representations. We introduce Geometry Image Diffusion (GIMDiffusion), a novel…

Computer Vision and Pattern Recognition · Computer Science 2024-09-06 Slava Elizarov , Ciara Rowles , Simon Donné

Generative models have achieved success in producing semantically plausible 2D images, but it remains challenging in 3D generation due to the absence of spatial geometry constraints. Typically, existing methods utilize geometric features as…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Haonan Wang , Hanyu Zhou , Haoyue Liu , Tao Gu , Luxin Yan

Multimodal learning has rapidly advanced visual understanding, largely via multimodal large language models (MLLMs) that use powerful LLMs as cognitive cores. In visual generation, however, these powerful core models are typically reduced…

Computer Vision and Pattern Recognition · Computer Science 2025-12-15 Han Lin , Xichen Pan , Ziqi Huang , Ji Hou , Jialiang Wang , Weifeng Chen , Zecheng He , Felix Juefei-Xu , Junzhe Sun , Zhipeng Fan , Ali Thabet , Mohit Bansal , Chu Wang

Interleaved text-image generation aims to jointly produce coherent visual frames and aligned textual descriptions within a single sequence, enabling tasks such as style transfer, compositional synthesis, and procedural tutorials. We present…

Computer Vision and Pattern Recognition · Computer Science 2025-12-23 Mingcheng Ye , Jiaming Liu , Yiren Song

While diffusion models excel at generating high-quality images, they often struggle with accurate counting, attributes, and spatial relationships in complex multi-object scenes. One potential solution involves employing Multimodal Large…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Jiayang Sun , Hongbo Wang , Jie Cao , Huaibo Huang , Ran He

Although perception systems have made remarkable advancements in recent years, particularly in 2D reasoning segmentation, these systems still rely on explicit human instruction or pre-defined categories to identify target objects before…

Computer Vision and Pattern Recognition · Computer Science 2025-06-06 Kunshen Zhang

We tackle the task of text-to-3D creation with pre-trained latent-based NeRFs (NeRFs that generate 3D objects given input latent code). Recent works such as DreamFusion and Magic3D have shown great success in generating 3D content using…

Computer Vision and Pattern Recognition · Computer Science 2023-12-21 Yu-Jhe Li , Tao Xu , Ji Hou , Bichen Wu , Xiaoliang Dai , Albert Pumarola , Peizhao Zhang , Peter Vajda , Kris Kitani

While 2D diffusion models have achieved remarkable success in identity-preserving personalization, extending this capability to 3D assets remains a significant challenge due to the complexities of multi-view consistency and spatial control.…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Jinxin Ai , Matthias Nießner , Ziya Erkoç

Compositing an object into an image involves multiple non-trivial sub-tasks such as object placement and scaling, color/lighting harmonization, viewpoint/geometry adjustment, and shadow/reflection generation. Recent generative image…

Computer Vision and Pattern Recognition · Computer Science 2024-09-12 Gemma Canet Tarrés , Zhe Lin , Zhifei Zhang , Jianming Zhang , Yizhi Song , Dan Ruta , Andrew Gilbert , John Collomosse , Soo Ye Kim

Automatically generating multiview illusions is a compelling challenge, where a single piece of visual content offers distinct interpretations from different viewing perspectives. Traditional methods, such as shadow art and wire art, create…

Computer Vision and Pattern Recognition · Computer Science 2025-06-09 Yue Feng , Vaibhav Sanjay , Spencer Lutz , Badour AlBahar , Songwei Ge , Jia-Bin Huang

Automatic 3D generation has recently attracted widespread attention. Recent methods have greatly accelerated the generation speed, but usually produce less-detailed objects due to limited model capacity or 3D data. Motivated by recent…

Computer Vision and Pattern Recognition · Computer Science 2024-03-12 Zilong Chen , Yikai Wang , Feng Wang , Zhengyi Wang , Huaping Liu

We introduce Imagen 3, a latent diffusion model that generates high quality images from text prompts. We describe our quality and responsibility evaluations. Imagen 3 is preferred over other state-of-the-art (SOTA) models at the time of…

Computer Vision and Pattern Recognition · Computer Science 2024-12-24 Imagen-Team-Google , : , Jason Baldridge , Jakob Bauer , Mukul Bhutani , Nicole Brichtova , Andrew Bunner , Lluis Castrejon , Kelvin Chan , Yichang Chen , Sander Dieleman , Yuqing Du , Zach Eaton-Rosen , Hongliang Fei , Nando de Freitas , Yilin Gao , Evgeny Gladchenko , Sergio Gómez Colmenarejo , Mandy Guo , Alex Haig , Will Hawkins , Hexiang Hu , Huilian Huang , Tobenna Peter Igwe , Christos Kaplanis , Siavash Khodadadeh , Yelin Kim , Ksenia Konyushkova , Karol Langner , Eric Lau , Rory Lawton , Shixin Luo , Soňa Mokrá , Henna Nandwani , Yasumasa Onoe , Aäron van den Oord , Zarana Parekh , Jordi Pont-Tuset , Hang Qi , Rui Qian , Deepak Ramachandran , Poorva Rane , Abdullah Rashwan , Ali Razavi , Robert Riachi , Hansa Srinivasan , Srivatsan Srinivasan , Robin Strudel , Benigno Uria , Oliver Wang , Su Wang , Austin Waters , Chris Wolff , Auriel Wright , Zhisheng Xiao , Hao Xiong , Keyang Xu , Marc van Zee , Junlin Zhang , Katie Zhang , Wenlei Zhou , Konrad Zolna , Ola Aboubakar , Canfer Akbulut , Oscar Akerlund , Isabela Albuquerque , Nina Anderson , Marco Andreetto , Lora Aroyo , Ben Bariach , David Barker , Sherry Ben , Dana Berman , Courtney Biles , Irina Blok , Pankil Botadra , Jenny Brennan , Karla Brown , John Buckley , Rudy Bunel , Elie Bursztein , Christina Butterfield , Ben Caine , Viral Carpenter , Norman Casagrande , Ming-Wei Chang , Solomon Chang , Shamik Chaudhuri , Tony Chen , John Choi , Dmitry Churbanau , Nathan Clement , Matan Cohen , Forrester Cole , Mikhail Dektiarev , Vincent Du , Praneet Dutta , Tom Eccles , Ndidi Elue , Ashley Feden , Shlomi Fruchter , Frankie Garcia , Roopal Garg , Weina Ge , Ahmed Ghazy , Bryant Gipson , Andrew Goodman , Dawid Górny , Sven Gowal , Khyatti Gupta , Yoni Halpern , Yena Han , Susan Hao , Jamie Hayes , Jonathan Heek , Amir Hertz , Ed Hirst , Emiel Hoogeboom , Tingbo Hou , Heidi Howard , Mohamed Ibrahim , Dirichi Ike-Njoku , Joana Iljazi , Vlad Ionescu , William Isaac , Reena Jana , Gemma Jennings , Donovon Jenson , Xuhui Jia , Kerry Jones , Xiaoen Ju , Ivana Kajic , Christos Kaplanis , Burcu Karagol Ayan , Jacob Kelly , Suraj Kothawade , Christina Kouridi , Ira Ktena , Jolanda Kumakaw , Dana Kurniawan , Dmitry Lagun , Lily Lavitas , Jason Lee , Tao Li , Marco Liang , Maggie Li-Calis , Yuchi Liu , Javier Lopez Alberca , Matthieu Kim Lorrain , Peggy Lu , Kristian Lum , Yukun Ma , Chase Malik , John Mellor , Thomas Mensink , Inbar Mosseri , Tom Murray , Aida Nematzadeh , Paul Nicholas , Signe Nørly , João Gabriel Oliveira , Guillermo Ortiz-Jimenez , Michela Paganini , Tom Le Paine , Roni Paiss , Alicia Parrish , Anne Peckham , Vikas Peswani , Igor Petrovski , Tobias Pfaff , Alex Pirozhenko , Ryan Poplin , Utsav Prabhu , Yuan Qi , Matthew Rahtz , Cyrus Rashtchian , Charvi Rastogi , Amit Raul , Ali Razavi , Sylvestre-Alvise Rebuffi , Susanna Ricco , Felix Riedel , Dirk Robinson , Pankaj Rohatgi , Bill Rosgen , Sarah Rumbley , Moonkyung Ryu , Anthony Salgado , Tim Salimans , Sahil Singla , Florian Schroff , Candice Schumann , Tanmay Shah , Eleni Shaw , Gregory Shaw , Brendan Shillingford , Kaushik Shivakumar , Dennis Shtatnov , Zach Singer , Evgeny Sluzhaev , Valerii Sokolov , Thibault Sottiaux , Florian Stimberg , Brad Stone , David Stutz , Yu-Chuan Su , Eric Tabellion , Shuai Tang , David Tao , Kurt Thomas , Gregory Thornton , Andeep Toor , Cristian Udrescu , Aayush Upadhyay , Cristina Vasconcelos , Alex Vasiloff , Andrey Voynov , Amanda Walker , Luyu Wang , Miaosen Wang , Simon Wang , Stanley Wang , Qifei Wang , Yuxiao Wang , Ágoston Weisz , Olivia Wiles , Chenxia Wu , Xingyu Federico Xu , Andrew Xue , Jianbo Yang , Luo Yu , Mete Yurtoglu , Ali Zand , Han Zhang , Jiageng Zhang , Catherine Zhao , Adilet Zhaxybay , Miao Zhou , Shengqi Zhu , Zhenkai Zhu , Dawn Bloxwich , Mahyar Bordbar , Luis C. Cobo , Eli Collins , Shengyang Dai , Tulsee Doshi , Anca Dragan , Douglas Eck , Demis Hassabis , Sissie Hsiao , Tom Hume , Koray Kavukcuoglu , Helen King , Jack Krawczyk , Yeqing Li , Kathy Meier-Hellstern , Andras Orban , Yury Pinsky , Amar Subramanya , Oriol Vinyals , Ting Yu , Yori Zwols

Recent research has demonstrated that the combination of pretrained diffusion models with neural radiance fields (NeRFs) has emerged as a promising approach for text-to-3D generation. Simply coupling NeRF with diffusion models will result…

Computer Vision and Pattern Recognition · Computer Science 2023-06-19 Lu Yu , Wei Xiang , Kang Han

With the rising industrial attention to 3D virtual modeling technology, generating novel 3D content based on specified conditions (e.g. text) has become a hot issue. In this paper, we propose a new generative 3D modeling framework called…

Computer Vision and Pattern Recognition · Computer Science 2023-05-09 Muheng Li , Yueqi Duan , Jie Zhou , Jiwen Lu

Controllable generation is a fundamental task in NLP with many applications, providing a basis for function calling to agentic communication. However, even state-of-the-art autoregressive Large Language Models (LLMs) today exhibit…

Computation and Language · Computer Science 2025-09-29 Zhen Xiong , Yujun Cai , Zhecheng Li , Yiwei Wang

Motivated by discrete diffusion's success in language-vision modeling, we explore its potential for multi-view generation, a task dominated by continuous approaches. We introduce ViewMask-1-to-3, formulating multi-view synthesis as a…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Ruishu Zhu , Zhihao Huang , Jiacheng Sun , Ping Luo , Hongyuan Zhang , Xuelong Li

Animating virtual avatars to make co-speech gestures facilitates various applications in human-machine interaction. The existing methods mainly rely on generative adversarial networks (GANs), which typically suffer from notorious mode…

Computer Vision and Pattern Recognition · Computer Science 2023-03-21 Lingting Zhu , Xian Liu , Xuanyu Liu , Rui Qian , Ziwei Liu , Lequan Yu

Fashion content generation is an emerging area at the intersection of artificial intelligence and creative design, with applications ranging from virtual try-on to culturally diverse design prototyping. Existing methods often struggle with…

Computation and Language · Computer Science 2025-01-28 Spencer Ramsey , Amina Grant , Jeffrey Lee

This paper presents DiffSurf, a transformer-based denoising diffusion model for generating and reconstructing 3D surfaces. Specifically, we design a diffusion transformer architecture that predicts noise from noisy 3D surface vertices and…

Computer Vision and Pattern Recognition · Computer Science 2024-08-28 Yusuke Yoshiyasu , Leyuan Sun