中文
相关论文

相关论文: 3D-LMVIC: Learning-based Multi-View Image Coding w…

200 篇论文

Novel-view synthesis and 3D reconstruction from sparse posed images are central to robotics and AR/VR. Yet, feed-forward 3D Gaussian reconstruction fails under lowlight due to noise, color shifts, and unreliable correspondence. We propose…

计算机视觉与模式识别 · 计算机科学 2026-05-27 Fuzhen Jiang , Zengtian Xie , Zhuoran Li

Novel view synthesis has shown rapid progress recently, with methods capable of producing increasingly photorealistic results. 3D Gaussian Splatting has emerged as a promising method, producing high-quality renderings of scenes and enabling…

计算机视觉与模式识别 · 计算机科学 2024-07-19 Richard Shaw , Michal Nazarczuk , Jifei Song , Arthur Moreau , Sibi Catley-Chandar , Helisa Dhamo , Eduardo Perez-Pellitero

Recent open-vocabulary 3D scene understanding approaches mainly focus on training 3D networks through contrastive learning with point-text pairs or by distilling 2D features into 3D models via point-pixel alignment. While these methods show…

计算机视觉与模式识别 · 计算机科学 2025-07-01 Xingyilang Yin , Jiale Wang , Xi Yang , Mutian Xu , Xu Gu , Nannan Wang

Recent advances in Gaussian Splatting-based inverse rendering extend Gaussian primitives with shading parameters and physically grounded light transport, enabling high-quality material recovery from dense multi-view captures. However, these…

计算机视觉与模式识别 · 计算机科学 2025-12-11 Patrick Noras , Jun Myeong Choi , Didier Stricker , Pieter Peers , Roni Sengupta

Modeling latent variables with priors and hyperpriors is an essential problem in variational image compression. Formally, trade-off between rate and distortion is handled well if priors and hyperpriors precisely describe latent variables.…

计算机视觉与模式识别 · 计算机科学 2022-03-23 Xiaosu Zhu , Jingkuan Song , Lianli Gao , Feng Zheng , Heng Tao Shen

Conventional multi-view re-ranking methods usually perform asymmetrical matching between the region of interest (ROI) in the query image and the whole target image for similarity computation. Due to the inconsistency in the visual…

计算机视觉与模式识别 · 计算机科学 2018-08-15 Jun Li , Chang Xu , Wankou Yang , Changyin Sun , Dacheng Tao , Hong Zhang

Free-Viewpoint Video (FVV) enables immersive 3D experiences, but efficient compression of dynamic 3D representation remains a major challenge. Existing dynamic 3D Gaussian Splatting methods couple reconstruction with optimization-dependent…

计算机视觉与模式识别 · 计算机科学 2025-12-30 Wenkang Zhang , Yan Zhao , Qiang Wang , Zhixin Xu , Li Song , Zhengxue Cheng

Autonomous driving sensors generate an enormous amount of data. In this paper, we explore learned multimodal compression for autonomous driving, specifically targeted at 3D object detection. We focus on camera and LiDAR modalities and…

图像与视频处理 · 电气工程与系统科学 2024-08-16 Hadi Hadizadeh , Ivan V. Bajić

Multi-view Clustering (MVC) has achieved significant progress, with many efforts dedicated to learn knowledge from multiple views. However, most existing methods are either not applicable or require additional steps for incomplete MVC. Such…

计算机视觉与模式识别 · 计算机科学 2024-08-19 Junjie Liu , Junlong Liu , Rongxin Jiang , Yaowu Chen , Chen Shen , Jieping Ye

Perceptual optimization is widely recognized as essential for neural compression, yet balancing the rate-distortion-perception tradeoff remains challenging. This difficulty is especially pronounced in video compression, where frame-wise…

图像与视频处理 · 电气工程与系统科学 2025-10-14 Zongyu Guo , Zhaoyang Jia , Jiahao Li , Xiaoyi Zhang , Bin Li , Yan Lu

Self-supervised detection and segmentation of foreground objects aims for accuracy without annotated training data. However, existing approaches predominantly rely on restrictive assumptions on appearance and motion. For scenes with dynamic…

计算机视觉与模式识别 · 计算机科学 2021-08-20 Isinsu Katircioglu , Helge Rhodin , Jörg Spörri , Mathieu Salzmann , Pascal Fua

Deep learning has revolutionized many computer vision fields in the last few years, including learning-based image compression. In this paper, we propose a deep semantic segmentation-based layered image compression (DSSLIC) framework in…

计算机视觉与模式识别 · 计算机科学 2019-04-22 Mohammad Akbari , Jie Liang , Jingning Han

Cross-view geo-spatial learning consists of two important tasks: Cross-View Geo-Localization (CVGL) and Cross-View Image Synthesis (CVIS), both of which rely on establishing geometric correspondences between ground and aerial views. Recent…

计算机视觉与模式识别 · 计算机科学 2026-03-30 Yancheng Zhang , Xiaohan Zhang , Guangyu Sun , Zonglin Lyu , Safwan Wshah , Chen Chen

Depictions of similar human body configurations can vary with changing viewpoints. Using only 2D information, we would like to enable vision algorithms to recognize similarity in human body poses across multiple views. This ability is…

计算机视觉与模式识别 · 计算机科学 2020-10-26 Jennifer J. Sun , Jiaping Zhao , Liang-Chieh Chen , Florian Schroff , Hartwig Adam , Ting Liu

Compared with previous 3D reconstruction methods like Nerf, recent Generalizable 3D Gaussian Splatting (G-3DGS) methods demonstrate impressive efficiency even in the sparse-view setting. However, the promising reconstruction performance of…

计算机视觉与模式识别 · 计算机科学 2024-08-27 Chuanrui Zhang , Yingshuang Zou , Zhuoling Li , Minmin Yi , Haoqian Wang

Visual encoding followed by token condensing has become the standard architectural paradigm in multi-modal large language models (MLLMs). Many recent MLLMs increasingly favor global native- resolution visual encoding over slice-based…

计算机视觉与模式识别 · 计算机科学 2025-11-27 Shichu Sun , Yichen Zhang , Haolin Song , Zonghao Guo , Chi Chen , Yidan Zhang , Yuan Yao , Zhiyuan Liu , Maosong Sun

We develop an approach to learning visual representations that embraces multimodal data, driven by a combination of intra- and inter-modal similarity preservation objectives. Unlike existing visual pre-training methods, which solve a proxy…

计算机视觉与模式识别 · 计算机科学 2021-04-28 Xin Yuan , Zhe Lin , Jason Kuen , Jianming Zhang , Yilin Wang , Michael Maire , Ajinkya Kale , Baldo Faieta

Recent MLLMs have shown emerging visual understanding and reasoning abilities after being pre-trained on large-scale multimodal datasets. Unlike pre-training, where MLLMs receive rich visual-text alignment, instruction-tuning is often…

We introduce MVSplat, an efficient model that, given sparse multi-view images as input, predicts clean feed-forward 3D Gaussians. To accurately localize the Gaussian centers, we build a cost volume representation via plane sweeping, where…

计算机视觉与模式识别 · 计算机科学 2024-10-29 Yuedong Chen , Haofei Xu , Chuanxia Zheng , Bohan Zhuang , Marc Pollefeys , Andreas Geiger , Tat-Jen Cham , Jianfei Cai

Multi-camera 3D object detection for autonomous driving is a challenging problem that has garnered notable attention from both academia and industry. An obstacle encountered in vision-based techniques involves the precise extraction of…

计算机视觉与模式识别 · 计算机科学 2023-04-10 Linyan Huang , Huijie Wang , Jia Zeng , Shengchuan Zhang , Liujuan Cao , Junchi Yan , Hongyang Li