English
Related papers

Related papers: Instructive3D: Editing Large Reconstruction Models…

200 papers

Creating machines capable of understanding the world in 3D is essential in assisting designers that build and edit 3D environments and robots navigating and interacting within a three-dimensional space. Inspired by advances in language and…

Computer Vision and Pattern Recognition · Computer Science 2026-01-07 Aadarsh Sahoo , Vansh Tibrewal , Georgia Gkioxari

We propose a novel deep reinforcement learning-based approach for 3D object reconstruction from monocular images. Prior works that use mesh representations are template based. Thus, they are limited to the reconstruction of objects that…

Computer Vision and Pattern Recognition · Computer Science 2021-09-27 Tarek Ben Charrada , Hedi Tabia , Aladine Chetouani , Hamid Laga

With the advent of depth-to-image diffusion models, text-guided generation, editing, and transfer of realistic textures are no longer difficult. However, due to the limitations of pre-trained diffusion models, they can only create…

Computer Vision and Pattern Recognition · Computer Science 2023-05-11 Zhibin Tang , Tiantong He

Text-to-3D generation is to craft a 3D object according to a natural language description. This can significantly reduce the workload for manually designing 3D models and provide a more natural way of interaction for users. However, this…

Computer Vision and Pattern Recognition · Computer Science 2023-09-27 Han Yi , Zhedong Zheng , Xiangyu Xu , Tat-seng Chua

Text-to-3D generation has attracted much attention from the computer vision community. Existing methods mainly optimize a neural field from scratch for each text prompt, relying on heavy and repetitive training cost which impedes their…

Computer Vision and Pattern Recognition · Computer Science 2024-04-30 Ming Li , Pan Zhou , Jia-Wei Liu , Jussi Keppo , Min Lin , Shuicheng Yan , Xiangyu Xu

Recent works have explored text-guided image editing using diffusion models and generated edited images based on text prompts. However, the models struggle to accurately locate the regions to be edited and faithfully perform precise edits.…

Computer Vision and Pattern Recognition · Computer Science 2023-05-30 Qian Wang , Biao Zhang , Michael Birsak , Peter Wonka

Learning-based 3D reconstruction methods have shown impressive results. However, most methods require 3D supervision which is often hard to obtain for real-world datasets. Recently, several works have proposed differentiable rendering…

Computer Vision and Pattern Recognition · Computer Science 2020-03-24 Michael Niemeyer , Lars Mescheder , Michael Oechsle , Andreas Geiger

The remarkable potential of multi-modal large language models (MLLMs) in comprehending both vision and language information has been widely acknowledged. However, the scarcity of 3D scenes-language pairs in comparison to their 2D…

Computer Vision and Pattern Recognition · Computer Science 2024-01-17 Zeju Li , Chao Zhang , Xiaoyan Wang , Ruilong Ren , Yifan Xu , Ruifei Ma , Xiangde Liu

Enhancing AI systems to perform tasks following human instructions can significantly boost productivity. In this paper, we present InstructP2P, an end-to-end framework for 3D shape editing on point clouds, guided by high-level textual…

Computer Vision and Pattern Recognition · Computer Science 2023-06-13 Jiale Xu , Xintao Wang , Yan-Pei Cao , Weihao Cheng , Ying Shan , Shenghua Gao

Animatable 3D human reconstruction from a single image is a challenging problem due to the ambiguity in decoupling geometry, appearance, and deformation. Recent advances in 3D human reconstruction mainly focus on static human modeling, and…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Lingteng Qiu , Xiaodong Gu , Peihao Li , Qi Zuo , Weichao Shen , Junfei Zhang , Kejie Qiu , Weihao Yuan , Guanying Chen , Zilong Dong , Liefeng Bo

3D object reconstruction from single-view image is a fundamental task in computer vision with wide-ranging applications. Recent advancements in Large Reconstruction Models (LRMs) have shown great promise in leveraging multi-view images…

Computer Vision and Pattern Recognition · Computer Science 2025-03-13 Zhiyuan Wu , Xibin Song , Senbo Wang , Weizhe Liu , Jiayu Yang , Ziang Cheng , Shenzhou Chen , Taizhang Shang , Weixuan Sun , Shan Luo , Pan Ji

We present MeshLLM, a novel framework that leverages large language models (LLMs) to understand and generate text-serialized 3D meshes. Our approach addresses key limitations in existing methods, including the limited dataset scale when…

Large Language Model (LLM)-driven digital humans have sparked a series of recent studies on co-speech gesture generation systems. However, existing approaches struggle with real-time synthesis and long-text comprehension. This paper…

Graphics · Computer Science 2025-06-03 Yueqian Guo , Tianzhao Li , Xin Lyu , Jiehaolin Chen , Zhaohan Wang , Sirui Xiao , Yurun Chen , Yezi He , Helin Li , Fan Zhang

In recent years, efforts have been made to use text information for better user profiling and item characterization in recommendations. However, text information can sometimes be of low quality, hindering its effectiveness for real-world…

Artificial Intelligence · Computer Science 2024-02-15 Yingpeng Du , Ziyan Wang , Zhu Sun , Haoyan Chua , Hongzhi Liu , Zhonghai Wu , Yining Ma , Jie Zhang , Youchen Sun

Despite rapid advancements in the capabilities of generative models, pretrained text-to-image models still struggle in capturing the semantics conveyed by complex prompts that compound multiple objects and instance-level attributes.…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Etai Sella , Yanir Kleiman , Hadar Averbuch-Elor

Data augmentation plays a crucial role in deep learning, enhancing the generalization and robustness of learning-based models. Standard approaches involve simple transformations like rotations and flips for generating extra data. However,…

Computer Vision and Pattern Recognition · Computer Science 2024-08-27 Shichao Dong , Ze Yang , Guosheng Lin

Multistep instructions, such as recipes and how-to guides, greatly benefit from visual aids, such as a series of images that accompany the instruction steps. While Large Language Models (LLMs) have become adept at generating coherent…

Computer Vision and Pattern Recognition · Computer Science 2024-05-17 João Bordalo , Vasco Ramos , Rodrigo Valério , Diogo Glória-Silva , Yonatan Bitton , Michal Yarom , Idan Szpektor , Joao Magalhaes

Recent advancements in 3D Large Language Models (3DLLMs) have highlighted their potential in building general-purpose agents in the 3D real world, yet challenges remain due to the lack of high-quality robust instruction-following data,…

Artificial Intelligence · Computer Science 2025-02-21 Weitai Kang , Haifeng Huang , Yuzhang Shang , Mubarak Shah , Yan Yan

We propose MeshLRM, a novel LRM-based approach that can reconstruct a high-quality mesh from merely four input images in less than one second. Different from previous large reconstruction models (LRMs) that focus on NeRF-based…

Computer Vision and Pattern Recognition · Computer Science 2025-01-24 Xinyue Wei , Kai Zhang , Sai Bi , Hao Tan , Fujun Luan , Valentin Deschaintre , Kalyan Sunkavalli , Hao Su , Zexiang Xu

Creating and editing the shape and color of 3D objects require tremendous human effort and expertise. Compared to direct manipulation in 3D interfaces, 2D interactions such as sketches and scribbles are usually much more natural and…

Computer Vision and Pattern Recognition · Computer Science 2022-07-26 Zezhou Cheng , Menglei Chai , Jian Ren , Hsin-Ying Lee , Kyle Olszewski , Zeng Huang , Subhransu Maji , Sergey Tulyakov
‹ Prev 1 4 5 6 7 8 10 Next ›