中文
相关论文

相关论文: Elite360M: Efficient 360 Multi-task Learning via B…

200 篇论文

Although equirectangular projection (ERP) is a convenient form to store omnidirectional images (also known as 360-degree images), it is neither equal-area nor conformal, thus not friendly to subsequent visual communication. In the context…

图像与视频处理 · 电气工程与系统科学 2021-12-28 Mu Li , Kede Ma , Jinxing Li , David Zhang

In this paper, we introduce Vox-Fusion++, a multi-maps-based robust dense tracking and mapping system that seamlessly fuses neural implicit representations with traditional volumetric fusion techniques. Building upon the concept of implicit…

计算机视觉与模式识别 · 计算机科学 2024-03-20 Hongjia Zhai , Hai Li , Xingrui Yang , Gan Huang , Yuhang Ming , Hujun Bao , Guofeng Zhang

We study multi-sensor fusion for 3D semantic segmentation that is important to scene understanding for many applications, such as autonomous driving and robotics. Existing fusion-based methods, however, may not achieve promising performance…

计算机视觉与模式识别 · 计算机科学 2024-09-10 Mingkui Tan , Zhuangwei Zhuang , Sitao Chen , Rong Li , Kui Jia , Qicheng Wang , Yuanqing Li

This paper presents VGGT-360, a novel training-free framework for zero-shot, geometry-consistent panoramic depth estimation. Unlike prior view-independent training-free approaches, VGGT-360 reformulates the task as panoramic reprojection…

计算机视觉与模式识别 · 计算机科学 2026-05-15 Jiayi Yuan , Haobo Jiang , De Wen Soh , Na Zhao

Implicit neural representations for videos (NeRV) have shown strong potential for video compression. However, applying NeRV to high-resolution 360-degree videos causes high memory usage and slow decoding, making real-time applications…

计算机视觉与模式识别 · 计算机科学 2026-02-17 Daichi Arai , Kyohei Unno , Yasuko Sugito , Yuichi Kusakabe

Depth estimation from a single image represents a very exciting challenge in computer vision. While other image-based depth sensing techniques leverage on the geometry between different viewpoints (e.g., stereo or structure from motion),…

计算机视觉与模式识别 · 计算机科学 2018-10-29 Pierluigi Zama Ramirez , Matteo Poggi , Fabio Tosi , Stefano Mattoccia , Luigi Di Stefano

Humans excel at constructing panoramic mental models of their surroundings, maintaining object permanence and inferring scene structure beyond visible regions. In contrast, current artificial vision systems struggle with persistent,…

计算机视觉与模式识别 · 计算机科学 2025-12-16 Finlay G. C. Hudson , James A. D. Gardner , William A. P. Smith

The field of 360-degree omnidirectional understanding has been receiving increasing attention for advancing spatial intelligence. However, the lack of large-scale and diverse data remains a major limitation. In this work, we propose…

计算机视觉与模式识别 · 计算机科学 2025-12-02 Xian Ge , Yuling Pan , Yuhang Zhang , Xiang Li , Weijun Zhang , Dizhe Zhang , Zhaoliang Wan , Xin Lin , Xiangkai Zhang , Juntao Liang , Jason Li , Wenjie Jiang , Bo Du , Ming-Hsuan Yang , Lu Qi

Vision-language models enable the understanding and reasoning of complex traffic scenarios through multi-source information fusion, establishing it as a core technology for autonomous driving. However, existing vision-language models are…

计算机视觉与模式识别 · 计算机科学 2025-12-17 Minghui Hou , Wei-Hsing Huang , Shaofeng Liang , Daizong Liu , Tai-Hao Wen , Gang Wang , Runwei Guan , Weiping Ding

3D semantic scene completion (SSC) is an ill-posed perception task that requires inferring a dense 3D scene from limited observations. Previous camera-based methods struggle to predict accurate semantic scenes due to inherent geometric…

计算机视觉与模式识别 · 计算机科学 2024-05-07 Bohan Li , Yasheng Sun , Zhujin Liang , Dalong Du , Zhuanghui Zhang , Xiaofeng Wang , Yunnan Wang , Xin Jin , Wenjun Zeng

Geometry problem-solving remains a significant challenge for Large Multimodal Models (LMMs), requiring not only global shape recognition but also attention to intricate local relationships related to geometric theory. To address this, we…

计算机视觉与模式识别 · 计算机科学 2026-02-10 Linger Deng , Yuliang Liu , Wenwen Yu , Zujia Zhang , Jianzhong Ju , Zhenbo Luo , Xiang Bai

Panoramic cameras, capable of capturing a 360-degree field of view, are crucial in robotic vision, particularly in environments with sparse features. However, non-upright panoramas due to unstable robot postures hinder downstream tasks.…

计算机视觉与模式识别 · 计算机科学 2025-12-02 Yuhao Shan , Qianyi Yuan , Jingguo Liu , Shigang Li , Jianfeng Li , Tong Chen

Omnidirectional and 360{\deg} images are becoming widespread in industry and in consumer society, causing omnidirectional computer vision to gain attention. Their wide field of view allows the gathering of a great amount of information…

计算机视觉与模式识别 · 计算机科学 2024-01-31 Bruno Berenguel-Baeta , Jesus Bermudez-Cameo , Jose J. Guerrero

Masked Image Modeling (MIM) has recently been established as a potent pre-training paradigm. A pretext task is constructed by masking patches in an input image, and this masked content is then predicted by a neural network using visible…

Building robots that can automate labor-intensive tasks has long been the core motivation behind the advancements in computer vision and the robotics community. Recent interest in leveraging 3D algorithms, particularly neural fields, has…

计算机视觉与模式识别 · 计算机科学 2023-12-13 Litian Liang , Liuyu Bian , Caiwei Xiao , Jialin Zhang , Linghao Chen , Isabella Liu , Fanbo Xiang , Zhiao Huang , Hao Su

Global perception is essential for embodied agents in 360{\deg} spaces, yet current affordance grounding remains largely object-centric and restricted to perspective views. To bridge this gap, we introduce a novel task: Holistic Affordance…

计算机视觉与模式识别 · 计算机科学 2026-03-11 Guoliang Zhu , Wanjun Jia , Caoyang Shao , Yuheng Zhang , Zhiyong Li , Kailun Yang

Structured light has proven instrumental in 3D imaging, LiDAR, and holographic light projection. Metasurfaces, comprised of sub-wavelength-sized nanostructures, facilitate 180$^\circ$ field-of-view (FoV) structured light, circumventing the…

光学 · 物理学 2023-06-28 Eunsue Choi , Gyeongtae Kim , Jooyeong Yun , Yujin Jeon , Junsuk Rho , Seung-Hwan Baek

This paper presents an investigation of vision transformer learning for multi-view geometry tasks, such as optical flow estimation, by fine-tuning video foundation models. Unlike previous methods that involve custom architectural designs…

计算机视觉与模式识别 · 计算机科学 2025-12-23 Huimin Wu , Kwang-Ting Cheng , Stephen Lin , Zhirong Wu

This report serves as a supplementary document for TaskPrompter, detailing its implementation on a new joint 2D-3D multi-task learning benchmark based on Cityscapes-3D. TaskPrompter presents an innovative multi-task prompting framework that…

计算机视觉与模式识别 · 计算机科学 2023-04-07 Hanrong Ye , Dan Xu

Visual localization plays an important role for intelligent robots and autonomous driving, especially when the accuracy of GNSS is unreliable. Recently, camera localization in LiDAR maps has attracted more and more attention for its low…

计算机视觉与模式识别 · 计算机科学 2024-10-28 Zhipeng Zhao , Huai Yu , Chenwei Lyv , Wen Yang , Sebastian Scherer