中文
相关论文

相关论文: TALO: Pushing 3D Vision Foundation Models Towards …

200 篇论文

Existing 3D foundation models typically align point clouds to frozen vision-language spaces like CLIP, which achieve strong cross-modal retrieval by compressing 3D shape into a global vector. However, this global-only alignment cannot…

计算机视觉与模式识别 · 计算机科学 2026-05-20 Zebin He , Mingxin Yang , Shuhui Yang , Hanxiao Sun , Xintong Han , Chunchao Guo , Wenhan Luo

Generalizable 3D Gaussian Splatting reconstruction showcases advanced Image-to-3D content creation but requires substantial computational resources and large datasets, posing challenges to training models from scratch. Current methods…

计算机视觉与模式识别 · 计算机科学 2026-01-05 Xiufeng Huang , Ka Chun Cheung , Runmin Cong , Simon See , Renjie Wan

On-the-fly 3D reconstruction from monocular image sequences is a long-standing challenge in computer vision, critical for applications such as real-to-sim, AR/VR, and robotics. Existing methods face a major tradeoff: per-scene optimization…

计算机视觉与模式识别 · 计算机科学 2025-10-10 Guanghao Li , Kerui Ren , Linning Xu , Zhewen Zheng , Changjian Jiang , Xin Gao , Bo Dai , Jian Pu , Mulin Yu , Jiangmiao Pang

Creating 3D maps on robots and other mobile devices has become a reality in recent years. Online 3D reconstruction enables many exciting applications in robotics and AR/VR gaming. However, the reconstructions are noisy and generally…

计算机视觉与模式识别 · 计算机科学 2017-03-29 Maksym Dzitsiuk , Jürgen Sturm , Robert Maier , Lingni Ma , Daniel Cremers

This work introduces alternating latent topologies (ALTO) for high-fidelity reconstruction of implicit 3D surfaces from noisy point clouds. Previous work identifies that the spatial arrangement of latent encodings is important to recover…

计算机视觉与模式识别 · 计算机科学 2022-12-09 Zhen Wang , Shijie Zhou , Jeong Joon Park , Despoina Paschalidou , Suya You , Gordon Wetzstein , Leonidas Guibas , Achuta Kadambi

Video-based 3D human pose and shape estimations are evaluated by intra-frame accuracy and inter-frame smoothness. Although these two metrics are responsible for different ranges of temporal consistency, existing state-of-the-art methods…

计算机视觉与模式识别 · 计算机科学 2023-03-28 Xiaolong Shen , Zongxin Yang , Xiaohan Wang , Jianxin Ma , Chang Zhou , Yi Yang

Neural implicit representations have recently demonstrated compelling results on dense Simultaneous Localization And Mapping (SLAM) but suffer from the accumulation of errors in camera tracking and distortion in the reconstruction.…

计算机视觉与模式识别 · 计算机科学 2023-09-06 Youmin Zhang , Fabio Tosi , Stefano Mattoccia , Matteo Poggi

Novel view synthesis from sparse-view inputs poses a significant challenge in 3D computer vision, particularly for achieving high-quality scene reconstructions with limited viewpoints. We introduce TWINGS, a framework that enhances 3D…

计算机视觉与模式识别 · 计算机科学 2026-05-29 Hyeseong Kim , Geonhui Son , Deukhee Lee , Dosik Hwang

Scene reconstruction has emerged as a central challenge in computer vision, with approaches such as Neural Radiance Fields (NeRF) and Gaussian Splatting achieving remarkable progress. While Gaussian Splatting demonstrates strong performance…

计算机视觉与模式识别 · 计算机科学 2025-11-11 Alexander Valverde , Brian Xu , Yuyin Zhou , Meng Xu , Hongyun Wang

Recent advances in 3D Gaussian Splatting (3DGS) have enabled generalizable, on-the-fly reconstruction of sequential input views. However, existing methods often predict per-pixel Gaussians and combine Gaussians from all views as the scene…

计算机视觉与模式识别 · 计算机科学 2025-10-20 Jiaxin Guo , Tongfan Guan , Wenzhen Dong , Wenzhao Zheng , Wenting Wang , Yue Wang , Yeung Yam , Yun-Hui Liu

Diffusion models excel at short-horizon robot planning, yet scaling them to long-horizon tasks remains challenging due to computational constraints and limited training data. Existing compositional approaches stitch together short segments…

机器人学 · 计算机科学 2026-03-04 Yixin Zhang , Yunhao Luo , Utkarsh Aashu Mishra , Woo Chul Shin , Yongxin Chen , Danfei Xu

The deepfake threats to society and cybersecurity have provoked significant public apprehension, driving intensified efforts within the realm of deepfake video detection. Current video-level methods are mostly based on {3D CNNs} resulting…

计算机视觉与模式识别 · 计算机科学 2024-03-21 Yuting Xu , Jian Liang , Lijun Sheng , Xiao-Yu Zhang

This paper introduces a 3D point cloud sequence learning model based on inconsistent spatio-temporal propagation for LiDAR odometry, termed DSLO. It consists of a pyramid structure with a spatial information reuse strategy, a sequential…

计算机视觉与模式识别 · 计算机科学 2024-09-04 Huixin Zhang , Guangming Wang , Xinrui Wu , Chenfeng Xu , Mingyu Ding , Masayoshi Tomizuka , Wei Zhan , Hesheng Wang

Real-time, high-quality, 3D scanning of large-scale scenes is key to mixed reality and robotic applications. However, scalability brings challenges of drift in pose estimation, introducing significant errors in the accumulated model.…

图形学 · 计算机科学 2017-02-09 Angela Dai , Matthias Nießner , Michael Zollhöfer , Shahram Izadi , Christian Theobalt

Dimensionality reduction algorithms map high-dimensional data into visualizable 2D or 3D spaces, but traditionally rely on a discrete point-cloud paradigm. This discrete abstraction is susceptible to visual occlusion and artificial…

图形学 · 计算机科学 2026-05-19 João Paulo Gois , Luis Gustavo Nonato

Reconstructing accurate 3D models of large-scale real-world scenes from unstructured, in-the-wild imagery remains a core challenge in computer vision, especially when the input views have little or no overlap. In such cases, existing…

计算机视觉与模式识别 · 计算机科学 2026-04-06 Tamir Cohen , Leo Segre , Shay Shomer-Chai , Shai Avidan , Hadar Averbuch-Elor

Embedding 3D morphable basis functions into deep neural networks opens great potential for models with better representation power. However, to faithfully learn those models from an image collection, it requires strong regularization to…

计算机视觉与模式识别 · 计算机科学 2019-04-11 Luan Tran , Feng Liu , Xiaoming Liu

Accurate and robust 3D scene reconstruction from casual, in-the-wild videos can significantly simplify robot deployment to new environments. However, reliable camera pose estimation and scene reconstruction from such unconstrained videos…

计算机视觉与模式识别 · 计算机科学 2025-04-30 Shuo Sun , Torsten Sattler , Malcolm Mielle , Achim J. Lilienthal , Martin Magnusson

Three dimensional (3D) object recognition is becoming a key desired capability for many computer vision systems such as autonomous vehicles, service robots and surveillance drones to operate more effectively in unstructured environments.…

计算机视觉与模式识别 · 计算机科学 2021-08-25 Chenxi Xiao , Juan Wachs

Face Anti-Spoofing (FAS) remains challenging due to the requirement for robust domain generalization across unseen environments. While recent trends leverage Vision-Language Models (VLMs) for semantic supervision, these multimodal…

计算机视觉与模式识别 · 计算机科学 2026-04-22 Mika Feng , Pierre Gallin-Martel , Koichi Ito , Takafumi Aoki
‹ 上一页 1 2 3 10 下一页 ›