中文
相关论文

相关论文: VecSet-Edit: Unleashing Pre-trained LRM for Mesh E…

200 篇论文

3D object detection and dense depth estimation are one of the most vital tasks in autonomous driving. Multiple sensor modalities can jointly attribute towards better robot perception, and to that end, we introduce a method for jointly…

计算机视觉与模式识别 · 计算机科学 2021-09-16 Shubham Shrivastava

Recent works have explored text-guided image editing using diffusion models and generated edited images based on text prompts. However, the models struggle to accurately locate the regions to be edited and faithfully perform precise edits.…

计算机视觉与模式识别 · 计算机科学 2023-05-30 Qian Wang , Biao Zhang , Michael Birsak , Peter Wonka

Single-view 3D object reconstruction is a fundamental and challenging computer vision task that aims at recovering 3D shapes from single-view RGB images. Most existing deep learning based reconstruction methods are trained and evaluated on…

计算机视觉与模式识别 · 计算机科学 2023-06-06 Xianghui Yang , Guosheng Lin , Luping Zhou

Recent advances in deep learning have significantly pushed the state-of-the-art in photorealistic video animation given a single image. In this paper, we extrapolate those advances to the 3D domain, by studying 3D image-to-video translation…

计算机视觉与模式识别 · 计算机科学 2020-07-22 Rolandos Alexandros Potamias , Jiali Zheng , Stylianos Ploumpis , Giorgos Bouritsas , Evangelos Ververas , Stefanos Zafeiriou

Visual autoregressive (VAR) models have recently emerged as a promising family of generative models, enabling a wide range of downstream vision tasks such as text-guided image editing. By shifting the editing paradigm from noise…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Tao Xia , Jiawei Liu , Yukun Zhang , Ting Liu , Wei Wang , Lei Zhang

Model editing aims to correct errors in large, pretrained models without altering unrelated behaviors. While some recent works have edited vision-language models (VLMs), no existing editors tackle reasoning-heavy tasks, which typically…

计算机视觉与模式识别 · 计算机科学 2026-05-13 Jiaxing Qiu , Kaihua Hou , Roxana Daneshjou , Ahmed Alaa , Thomas Hartvigsen

Large language models (LLMs) are pretrained on corpora containing trillions of tokens and, therefore, inevitably memorize sensitive information. Locate-then-edit methods, as a mainstream paradigm of model editing, offer a promising solution…

密码学与安全 · 计算机科学 2026-05-19 Zhiyu Sun , Minrui Luo , Yu Wang , Zhili Chen , Tianxing He

We present LlamaSeg, a visual autoregressive framework that unifies multiple image segmentation tasks via natural language instructions. We reformulate image segmentation as a visual generation problem, representing masks as "visual" tokens…

计算机视觉与模式识别 · 计算机科学 2025-05-27 Jiru Deng , Tengjin Weng , Tianyu Yang , Wenhan Luo , Zhiheng Li , Wenhao Jiang

We introduce Motion2VecSets, a 4D diffusion model for dynamic surface reconstruction from point cloud sequences. While existing state-of-the-art methods have demonstrated success in reconstructing non-rigid objects using neural field…

计算机视觉与模式识别 · 计算机科学 2024-04-16 Wei Cao , Chang Luo , Biao Zhang , Matthias Nießner , Jiapeng Tang

Learning radiance fields (NeRF) with powerful 2D diffusion models has garnered popularity for text-to-3D generation. Nevertheless, the implicit 3D representations of NeRF lack explicit modeling of meshes and textures over surfaces, and such…

计算机视觉与模式识别 · 计算机科学 2024-09-12 Haibo Yang , Yang Chen , Yingwei Pan , Ting Yao , Zhineng Chen , Zuxuan Wu , Yu-Gang Jiang , Tao Mei

Recent volumetric 3D reconstruction methods can produce very accurate results, with plausible geometry even for unobserved surfaces. However, they face an undesirable trade-off when it comes to multi-view fusion. They can fuse all available…

计算机视觉与模式识别 · 计算机科学 2021-12-02 Noah Stier , Alexander Rich , Pradeep Sen , Tobias Höllerer

We propose VecGAN, an image-to-image translation framework for facial attribute editing with interpretable latent directions. Facial attribute editing task faces the challenges of precise attribute editing with controllable strength and…

计算机视觉与模式识别 · 计算机科学 2022-07-08 Yusuf Dalva , Said Fahri Altindis , Aysegul Dundar

Text-guided 3D editing aims to modify existing 3D assets using natural-language instructions. Current methods struggle to jointly understand complex prompts, automatically localize edits in 3D, and preserve unedited content. We introduce…

计算机视觉与模式识别 · 计算机科学 2026-02-24 Yankuan Chi , Xiang Li , Zixuan Huang , James M. Rehg

Previous efforts have managed to generate production-ready 3D assets from text or images. However, these methods primarily employ NeRF or 3D Gaussian representations, which are not adept at producing smooth, high-quality geometries required…

图形学 · 计算机科学 2024-10-15 Rengan Xie , Wenting Zheng , Kai Huang , Yizheng Chen , Qi Wang , Qi Ye , Wei Chen , Yuchi Huo

Determining the relative pose of a previously unseen object between two images is pivotal to the success of generalizable object pose estimation. Existing approaches typically predict 3D translation utilizing the ground-truth object…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Chen Zhao , Tong Zhang , Zheng Dang , Mathieu Salzmann

Score Distillation Sampling (SDS) enables high-quality text-to-3D generation by supervising 3D models through the denoising of multi-view 2D renderings, using a pretrained text-to-image diffusion model to align with the input prompt and…

计算机视觉与模式识别 · 计算机科学 2025-09-22 Weimin Bai , Yubo Li , Weijian Luo , Wenzheng Chen , He Sun

We present the first text-based image editing approach for object parts based on pre-trained diffusion models. Diffusion-based image editing approaches capitalized on the deep understanding of diffusion models of image semantics to perform…

计算机视觉与模式识别 · 计算机科学 2025-06-30 Aleksandar Cvejic , Abdelrahman Eldesokey , Peter Wonka

3D Human Body Reconstruction from a monocular image is an important problem in computer vision with applications in virtual and augmented reality platforms, animation industry, en-commerce domain, etc. While several of the existing works…

计算机视觉与模式识别 · 计算机科学 2019-08-20 Abbhinav Venkat , Chaitanya Patel , Yudhik Agrawal , Avinash Sharma

Utilizing patch-based transformers for unstructured geometric data such as polygon meshes presents significant challenges, primarily due to the absence of a canonical ordering and variations in input sizes. Prior approaches to handling 3D…

计算机视觉与模式识别 · 计算机科学 2024-11-04 Mohammad Farazi , Yalin Wang

Despite the remarkable capabilities of text-to-image (T2I) generation models, real-world applications often demand fine-grained, iterative image editing that existing methods struggle to provide. Key challenges include granular instruction…

计算机视觉与模式识别 · 计算机科学 2025-08-26 Zihan Liang , Jiahao Sun , Haoran Ma