English
Related papers

Related papers: Pose-Free Generalizable Rendering Transformer

200 papers

Image matching that finding robust and accurate correspondences across images is a challenging task under extreme conditions. Capturing local and global features simultaneously is an important way to mitigate such an issue but recent…

Computer Vision and Pattern Recognition · Computer Science 2023-11-30 Wenhao Zhong , Jie Jiang

Novel view synthesis (NVS) aims to generate images at arbitrary viewpoints using multi-view images, and recent insights from neural radiance fields (NeRF) have contributed to remarkable improvements. Recently, studies on generalizable NeRF…

Computer Vision and Pattern Recognition · Computer Science 2024-10-02 Youngho Yoon , Hyun-Kurl Jang , Kuk-Jin Yoon

We introduce native-resolution image synthesis, a novel generative modeling paradigm that enables the synthesis of images at arbitrary resolutions and aspect ratios. This approach overcomes the limitations of conventional fixed-resolution,…

Computer Vision and Pattern Recognition · Computer Science 2025-06-04 Zidong Wang , Lei Bai , Xiangyu Yue , Wanli Ouyang , Yiyuan Zhang

Accurate state estimation is a fundamental problem for autonomous robots. To achieve locally accurate and globally drift-free state estimation, multiple sensors with complementary properties are usually fused together. Local sensors…

Computer Vision and Pattern Recognition · Computer Science 2019-01-14 Tong Qin , Shaozu Cao , Jie Pan , Shaojie Shen

Localized subject-driven image editing aims to seamlessly integrate user-specified objects into target scenes. As generative models continue to scale, training becomes increasingly costly in terms of memory and computation, highlighting the…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Peilin Xiong , Junwen Chen , Honghui Yuan , Keiji Yanai

Capturing and faithfully rendering photo-realistic humans from novel views is a fundamental problem for AR/VR applications. While prior work has shown impressive performance capture results in laboratory settings, it is non-trivial to…

Computer Vision and Pattern Recognition · Computer Science 2022-08-03 Phong Nguyen-Ha , Nikolaos Sarafianos , Christoph Lassner , Janne Heikkila , Tony Tung

We present RayZer, a self-supervised multi-view 3D Vision model trained without any 3D supervision, i.e., camera poses and scene geometry, while exhibiting emerging 3D awareness. Concretely, RayZer takes unposed and uncalibrated images as…

Computer Vision and Pattern Recognition · Computer Science 2025-05-02 Hanwen Jiang , Hao Tan , Peng Wang , Haian Jin , Yue Zhao , Sai Bi , Kai Zhang , Fujun Luan , Kalyan Sunkavalli , Qixing Huang , Georgios Pavlakos

While generalizable 3D Gaussian splatting enables efficient, high-quality rendering of unseen scenes, it heavily depends on precise camera poses for accurate geometry. In real-world scenarios, obtaining accurate poses is challenging,…

Computer Vision and Pattern Recognition · Computer Science 2025-10-22 Youngju Na , Taeyeon Kim , Jumin Lee , Kyu Beom Han , Woo Jae Kim , Sung-eui Yoon

Face recognition under extreme head poses is a challenging task. Ideally, a face recognition system should perform well across different head poses, which is known as pose-invariant face recognition. To achieve pose invariance, current…

Computer Vision and Pattern Recognition · Computer Science 2025-06-06 Patrik Mesec , Alan Jović

Pose-guided video generation refers to controlling the motion of subjects in generated video through a sequence of poses. It enables precise control over subject motion and has important applications in animation. However, current…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Ruiyan Wang , Teng Hu , Kaihui Huang , Zihan Su , Ran Yi , Lizhuang Ma

Scene view synthesis, which generates novel views from limited perspectives, is increasingly vital for applications like virtual reality, augmented reality, and robotics. Unlike object-based tasks, such as generating 360{\deg} views of a…

Computer Vision and Pattern Recognition · Computer Science 2025-04-29 Xiaofeng Jin , Yan Fang , Matteo Frosi , Jianfei Ge , Jiangjian Xiao , Matteo Matteucci

Neural radiance fields (NeRF) and 3D Gaussian Splatting (3DGS) are popular techniques to reconstruct and render photo-realistic images. However, the pre-requisite of running Structure-from-Motion (SfM) to get camera poses limits their…

Computer Vision and Pattern Recognition · Computer Science 2024-11-26 Yu Chen , Rolandos Alexandros Potamias , Evangelos Ververas , Jifei Song , Jiankang Deng , Gim Hee Lee

We propose SelfSplat, a novel 3D Gaussian Splatting model designed to perform pose-free and 3D prior-free generalizable 3D reconstruction from unposed multi-view images. These settings are inherently ill-posed due to the lack of…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Gyeongjin Kang , Jisang Yoo , Jihyeon Park , Seungtae Nam , Hyeonsoo Im , Sangheon Shin , Sangpil Kim , Eunbyung Park

Differentiable rendering techniques have recently shown promising results for free-viewpoint video synthesis of characters. However, such methods, either Gaussian Splatting or neural implicit rendering, typically necessitate per-subject…

Computer Vision and Pattern Recognition · Computer Science 2024-11-19 Boyao Zhou , Shunyuan Zheng , Hanzhang Tu , Ruizhi Shao , Boning Liu , Shengping Zhang , Liqiang Nie , Yebin Liu

Recent advancements in autonomous driving (AD) systems have highlighted the potential of world models in achieving robust and generalizable performance across both ordinary and challenging driving conditions. However, a key challenge…

Computer Vision and Pattern Recognition · Computer Science 2025-07-21 Bu Jin , Weize Li , Baihan Yang , Zhenxin Zhu , Junpeng Jiang , Huan-ang Gao , Haiyang Sun , Kun Zhan , Hengtong Hu , Xueyang Zhang , Peng Jia , Hao Zhao

We present a framework for end-to-end joint quantization of Vision Transformers trained on ImageNet for the purpose of image classification. Unlike prior post-training or block-wise reconstruction methods, we jointly optimize over the…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Shile Li , Markus Karmann , Onay Urfalioglu

Image generation today can produce somewhat realistic images from text prompts. However, if one asks the generator to synthesize a specific camera setting such as creating different fields of view using a 24mm lens versus a 70mm lens, the…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Yu Yuan , Xijun Wang , Yichen Sheng , Prateek Chennuri , Xingguang Zhang , Stanley Chan

Explicitly modeling room background depth as a geometric constraint has proven effective for panoramic depth estimation. However, reconstructing this background depth for regular enclosed regions in a complex indoor scene without external…

Computer Vision and Pattern Recognition · Computer Science 2026-02-17 Kanglin Ning , Ruzhao Chen , Penghong Wang , Xingtao Wang , Ruiqin Xiong , Xiaopeng Fan

Multi-focus image fusion (MFIF) and super-resolution (SR) are the inverse problem of imaging model, purposes of MFIF and SR are obtaining all-in-focus and high-resolution 2D mapping of targets. Though various MFIF and SR methods have been…

Computer Vision and Pattern Recognition · Computer Science 2023-01-18 Yuanjie Gu , Yinghan Guan , Zhibo Xiao , Haoran Dai , Cheng Liu , Shouyu Wang

We introduce GROOT, an imitation learning method for learning robust policies with object-centric and 3D priors. GROOT builds policies that generalize beyond their initial training conditions for vision-based manipulation. It constructs…

Robotics · Computer Science 2023-10-24 Yifeng Zhu , Zhenyu Jiang , Peter Stone , Yuke Zhu