English
Related papers

Related papers: MVPbev: Multi-view Perspective Image Generation fr…

200 papers

As a cornerstone technique for autonomous driving, Bird's Eye View (BEV) segmentation has recently achieved remarkable progress with pinhole cameras. However, it is non-trivial to extend the existing methods to fisheye cameras with severe…

Computer Vision and Pattern Recognition · Computer Science 2025-09-18 Hang Li , Dianmo Sheng , Qiankun Dong , Zichun Wang , Zhiwei Xu , Tao Li

Existing multi-view image generation methods often make invasive modifications to pre-trained text-to-image (T2I) models and require full fine-tuning, leading to (1) high computational costs, especially with large base models and…

Computer Vision and Pattern Recognition · Computer Science 2024-12-06 Zehuan Huang , Yuan-Chen Guo , Haoran Wang , Ran Yi , Lizhuang Ma , Yan-Pei Cao , Lu Sheng

An accurate understanding of a self-driving vehicle's surrounding environment is crucial for its navigation system. To enhance the effectiveness of existing algorithms and facilitate further research, it is essential to provide…

Computer Vision and Pattern Recognition · Computer Science 2023-11-14 Abtin Mahyar , Hossein Motamednia , Dara Rahmati

Birds-eye-view (BEV) semantic segmentation is critical for autonomous driving for its powerful spatial representation ability. It is challenging to estimate the BEV semantic maps from monocular images due to the spatial gap, since it is…

Computer Vision and Pattern Recognition · Computer Science 2022-07-22 Shi Gong , Xiaoqing Ye , Xiao Tan , Jingdong Wang , Errui Ding , Yu Zhou , Xiang Bai

This paper proposes an efficient multi-camera to Bird's-Eye-View (BEV) view transformation method for 3D perception, dubbed MatrixVT. Existing view transformers either suffer from poor transformation efficiency or rely on device-specific…

Computer Vision and Pattern Recognition · Computer Science 2022-11-22 Hongyu Zhou , Zheng Ge , Zeming Li , Xiangyu Zhang

Bird's-eye-view (BEV) is a powerful and widely adopted representation for road scenes that captures surrounding objects and their spatial locations, along with overall context in the scene. In this work, we focus on bird's eye semantic…

Computer Vision and Pattern Recognition · Computer Science 2020-06-24 Mong H. Ng , Kaahan Radia , Jianfei Chen , Dequan Wang , Ionel Gog , Joseph E. Gonzalez

BEV-based 3D perception has emerged as a focal point of research in end-to-end autonomous driving. However, existing BEV approaches encounter significant challenges due to the large feature space, complicating efficient modeling and…

Computer Vision and Pattern Recognition · Computer Science 2025-09-24 Feng Li , Zhaoyue Wang , Enyuan Zhang , Mohammad Masum Billah , Yunduan Cui , Kun Xu

Existing Multi-Plane Image (MPI) based view-synthesis methods generate an MPI aligned with the input view using a fixed number of planes in one forward pass. These methods produce fast, high-quality rendering of novel views, but rely on…

Computer Vision and Pattern Recognition · Computer Science 2021-07-13 Sushobhan Ghosh , Zhaoyang Lv , Nathan Matsuda , Lei Xiao , Andrew Berkovich , Oliver Cossairt

Semantic segmentation in bird's eye view (BEV) is an important task for autonomous driving. Though this task has attracted a large amount of research efforts, it is still challenging to flexibly cope with arbitrary (single or multiple)…

Computer Vision and Pattern Recognition · Computer Science 2022-08-19 Lang Peng , Zhirong Chen , Zhangjie Fu , Pengpeng Liang , Erkang Cheng

With the attention gained by camera-only 3D object detection in autonomous driving, methods based on Bird-Eye-View (BEV) representation especially derived from the forward view transformation paradigm, i.e., lift-splat-shoot (LSS), have…

Computer Vision and Pattern Recognition · Computer Science 2024-11-20 Weijie Ma , Jingwei Jiang , Yang Yang , Zehui Chen , Hao Chen

Recent advances in unified multimodal models (UMMs) have enabled impressive progress in visual comprehension and generation. However, existing datasets and benchmarks focus primarily on single-turn interactions, failing to capture the…

Computer Vision and Pattern Recognition · Computer Science 2025-11-17 Wei Chow , Jiachun Pan , Yongyuan Liang , Mingze Zhou , Xue Song , Liyu Jia , Saining Zhang , Siliang Tang , Juncheng Li , Fengda Zhang , Weijia Wu , Hanwang Zhang , Tat-Seng Chua

Multi-view clustering (MvC) aims to integrate information from different views to enhance the capability of the model in capturing the underlying data structures. The widely used joint training paradigm in MvC is potentially not fully…

Computer Vision and Pattern Recognition · Computer Science 2025-02-05 Zhenglai Li , Jun Wang , Chang Tang , Xinzhong Zhu , Wei Zhang , Xinwang Liu

Bird's-Eye-View (BEV) semantic segmentation provides comprehensive environmental perception for autonomous driving but suffers multi-modal misalignment and sensor noise. We propose RESAR-BEV, a progressive refinement framework that advances…

Computer Vision and Pattern Recognition · Computer Science 2026-03-06 Zhiwen Zeng , Yunfei Yin , Zheng Yuan , Argho Dey , Xianjian Bao

We introduce a framework that enables both multi-view character consistency and 3D camera control in video diffusion models through a novel customization data pipeline. We train the character consistency component with recorded volumetric…

Computer Vision and Pattern Recognition · Computer Science 2025-10-17 Yuancheng Xu , Wenqi Xian , Li Ma , Julien Philip , Ahmet Levent Taşel , Yiwei Zhao , Ryan Burgert , Mingming He , Oliver Hermann , Oliver Pilarski , Rahul Garg , Paul Debevec , Ning Yu

Recent months have witnessed rapid progress in 3D generation based on diffusion models. Most advances require fine-tuning existing 2D Stable Diffsuions into multi-view settings or tedious distilling operations and hence fall short of 3D…

Computer Vision and Pattern Recognition · Computer Science 2023-12-19 Suyi Jiang , Haimin Luo , Haoran Jiang , Ziyu Wang , Jingyi Yu , Lan Xu

Recent advances in text-to-video diffusion models have enabled the generation of high-quality videos conditioned on textual descriptions. However, most existing text-to-video models rely solely on textual conditions, lacking general…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Yuheng Chen , Teng Hu , Jiangning Zhang , Zhucun Xue , Ran Yi , Lizhuang Ma

Bird's-eye-view (BEV) representations are the dominant paradigm for 3D perception in autonomous driving, providing a unified spatial canvas where detection and segmentation features are geometrically registered to the same physical…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Ahmet İnanç , Özgür Erkent

Generative video modeling has made significant strides, yet ensuring structural and temporal consistency over long sequences remains a challenge. Current methods predominantly rely on RGB signals, leading to accumulated errors in object…

Computer Vision and Pattern Recognition · Computer Science 2025-08-22 Zhiheng Liu , Xueqing Deng , Shoufa Chen , Angtian Wang , Qiushan Guo , Mingfei Han , Zeyue Xue , Mengzhao Chen , Ping Luo , Linjie Yang

A unified diffusion framework for multi-modal generation and understanding has the transformative potential to achieve seamless and controllable image diffusion and other cross-modal tasks. In this paper, we introduce MMGen, a unified…

Computer Vision and Pattern Recognition · Computer Science 2025-03-27 Jiepeng Wang , Zhaoqing Wang , Hao Pan , Yuan Liu , Dongdong Yu , Changhu Wang , Wenping Wang

Three-dimensional object detection is one of the key tasks in autonomous driving. To reduce costs in practice, low-cost multi-view cameras for 3D object detection are proposed to replace the expansive LiDAR sensors. However, relying solely…

Computer Vision and Pattern Recognition · Computer Science 2024-03-26 Zhiwei Lin , Zhe Liu , Zhongyu Xia , Xinhao Wang , Yongtao Wang , Shengxiang Qi , Yang Dong , Nan Dong , Le Zhang , Ce Zhu
‹ Prev 1 8 9 10 Next ›