中文
相关论文

相关论文: SAR3D: Autoregressive 3D Object Generation and Und…

200 篇论文

We present AMB3R, a multi-view feed-forward model for dense 3D reconstruction on a metric-scale that addresses diverse 3D vision tasks. The key idea is to leverage a sparse, yet compact, volumetric scene representation as our backend,…

计算机视觉与模式识别 · 计算机科学 2025-11-26 Hengyi Wang , Lourdes Agapito

Most image-based 3D object reconstructors assume that objects are fully visible, ignoring occlusions that commonly occur in real-world scenarios. In this paper, we introduce Amodal3R, a conditional 3D generative model designed to…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Tianhao Wu , Chuanxia Zheng , Frank Guan , Andrea Vedaldi , Tat-Jen Cham

Recent advances in vision foundation models have revolutionized geometry reconstruction and semantic understanding. Yet, most of the existing approaches treat these capabilities in isolation, leading to redundant pipelines and compounded…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Chaoyi Zhou , Run Wang , Feng Luo , Mert D. Pesé , Zhiwen Fan , Yiqi Zhong , Siyu Huang

We consider the problem of scaling deep generative shape models to high-resolution. Drawing motivation from the canonical view representation of objects, we introduce a novel method for the fast up-sampling of 3D objects in voxel space…

计算机视觉与模式识别 · 计算机科学 2018-11-16 Edward Smith , Scott Fujimoto , David Meger

Recent advancements in 3D robotic manipulation have improved grasping of everyday objects, but transparent and specular materials remain challenging due to depth sensing limitations. While several 3D reconstruction and depth completion…

机器人学 · 计算机科学 2025-06-23 Mingxu Zhang , Xiaoqi Li , Jiahui Xu , Kaichen Zhou , Hojin Bae , Yan Shen , Chuyan Xiong , Hao Dong

Medical vision-and-language pre-training provides a feasible solution to extract effective vision-and-language representations from medical images and texts. However, few studies have been dedicated to this field to facilitate medical…

计算机视觉与模式识别 · 计算机科学 2022-09-16 Zhihong Chen , Yuhao Du , Jinpeng Hu , Yang Liu , Guanbin Li , Xiang Wan , Tsung-Hui Chang

We study the problem of 3D object generation. We propose a novel framework, namely 3D Generative Adversarial Network (3D-GAN), which generates 3D objects from a probabilistic space by leveraging recent advances in volumetric convolutional…

计算机视觉与模式识别 · 计算机科学 2017-01-05 Jiajun Wu , Chengkai Zhang , Tianfan Xue , William T. Freeman , Joshua B. Tenenbaum

Non-autoregressive (NAR) models simultaneously generate multiple outputs in a sequence, which significantly reduces the inference speed at the cost of accuracy drop compared to autoregressive baselines. Showing great potential for real-time…

音频与语音处理 · 电气工程与系统科学 2021-10-12 Yosuke Higuchi , Nanxin Chen , Yuya Fujita , Hirofumi Inaguma , Tatsuya Komatsu , Jaesong Lee , Jumon Nozaki , Tianzi Wang , Shinji Watanabe

360 panoramic images are increasingly used in virtual reality, autonomous driving, and robotics for holistic scene understanding. However, current Vision-Language Models (VLMs) struggle with 3D spatial reasoning on Equirectangular…

计算机视觉与模式识别 · 计算机科学 2026-02-26 Zekai Lin , Xu Zheng

Autoregressive (AR) models have achieved unified and strong performance across both visual understanding and image generation tasks. However, removing undesired concepts from AR models while maintaining overall generation quality remains an…

计算机视觉与模式识别 · 计算机科学 2025-06-26 Haipeng Fan , Shiyuan Zhang , Baohunesitu , Zihang Guo , Huaiwen Zhang

We present a novel approach for enhancing the resolution and geometric fidelity of 3D Gaussian Splatting (3DGS) beyond native training resolution. Current 3DGS methods are fundamentally limited by their input resolution, producing…

图形学 · 计算机科学 2025-06-10 Shuja Khalid , Mohamed Ibrahim , Yang Liu

Studies on the automatic processing of 3D human pose data have flourished in the recent past. In this paper, we are interested in the generation of plausible and diverse future human poses following an observed 3D pose sequence. Current…

计算机视觉与模式识别 · 计算机科学 2022-04-05 Xiaoyu Bie , Wen Guo , Simon Leglaive , Lauren Girin , Francesc Moreno-Noguer , Xavier Alameda-Pineda

We present Gen3R, a method that bridges the strong priors of foundational reconstruction models and video diffusion models for scene-level 3D generation. We repurpose the VGGT reconstruction model to produce geometric latents by training an…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Jiaxin Huang , Yuanbo Yang , Bangbang Yang , Lin Ma , Yuewen Ma , Yiyi Liao

Recent advances in large language models (LLMs) provide new opportunities for context understanding in virtual reality (VR). However, VR contexts are often highly localized and personalized, limiting the effectiveness of general-purpose…

信息检索 · 计算机科学 2025-04-15 Shiyi Ding , Ying Chen

Auto-regressive models have made significant progress in the realm of language generation, yet they do not perform on par with diffusion models in the domain of image synthesis. In this work, we introduce MARS, a novel framework for T2I…

计算机视觉与模式识别 · 计算机科学 2024-07-12 Wanggui He , Siming Fu , Mushui Liu , Xierui Wang , Wenyi Xiao , Fangxun Shu , Yi Wang , Lei Zhang , Zhelun Yu , Haoyuan Li , Ziwei Huang , LeiLei Gan , Hao Jiang

3D assets are essential in the digital age. While automatic 3D generation, such as image-to-3d, has made significant strides in recent years, it often struggles to achieve fast, detailed, and high-fidelity generation simultaneously. In this…

计算机视觉与模式识别 · 计算机科学 2025-11-26 Huanning Dong , Yinuo Huang , Fan Li , Ping Kuang

Recent multimodal models such as DALL-E and CM3 have achieved remarkable progress in text-to-image and image-to-text generation. However, these models store all learned knowledge (e.g., the appearance of the Eiffel Tower) in the model…

计算机视觉与模式识别 · 计算机科学 2023-06-07 Michihiro Yasunaga , Armen Aghajanyan , Weijia Shi , Rich James , Jure Leskovec , Percy Liang , Mike Lewis , Luke Zettlemoyer , Wen-tau Yih

While Vision-Language Models (VLMs) exhibit exceptional 2D visual understanding, their ability to comprehend and reason about 3D space--a cornerstone of spatial intelligence--remains superficial. Current methodologies attempt to bridge this…

计算机视觉与模式识别 · 计算机科学 2026-02-25 Haoyi Jiang , Liu Liu , Xinjie Wang , Yonghao He , Wei Sui , Zhizhong Su , Wenyu Liu , Xinggang Wang

This article proposes a data-driven methodology to achieve a fast design support, in order to generate or develop novel designs covering multiple object categories. This methodology implements two state-of-the-art Variational Autoencoder…

计算机视觉与模式识别 · 计算机科学 2018-08-07 Zhangsihao Yang , Haoliang Jiang , Zou Lan

We present a novel time series anomaly detection method that achieves excellent detection accuracy while offering a superior level of explainability. Our proposed method, TimeVQVAE-AD, leverages masked generative modeling adapted from the…

机器学习 · 计算机科学 2024-08-01 Daesoo Lee , Sara Malacarne , Erlend Aune
‹ 上一页 1 8 9 10 下一页 ›