English
Related papers

Related papers: Point Primitive Transformer for Long-Term 4D Point…

200 papers

Understanding and reconstructing the complex geometry and motion of dynamic scenes from video remains a formidable challenge in computer vision. This paper introduces D4RT, a simple yet powerful feedforward model designed to efficiently…

Transform and entropy models are the two core components in deep image compression neural networks. Most existing learning-based image compression methods utilize convolutional-based transform, which lacks the ability to model long-range…

Image and Video Processing · Electrical Eng. & Systems 2023-09-20 Atefeh Khoshkhahtinat , Ali Zafari , Piyush M. Mehta , Mohammad Akyash , Hossein Kashiani , Nasser M. Nasrabadi

Learning accurate and parsimonious point cloud representations of scene surfaces from scratch remains a challenge in 3D representation learning. Existing point-based methods often suffer from the vanishing gradient problem or require a…

Computer Vision and Pattern Recognition · Computer Science 2023-12-08 Yanshu Zhang , Shichong Peng , Alireza Moazeni , Ke Li

The dominant paradigm in 3D human pose estimation that lifts a 2D pose sequence to 3D heavily relies on long-term temporal clues (i.e., using a daunting number of video frames) for improved accuracy, which incurs performance saturation,…

Computer Vision and Pattern Recognition · Computer Science 2023-11-10 Qitao Zhao , Ce Zheng , Mengyuan Liu , Chen Chen

With the increased interest in immersive experiences, point cloud came to birth and was widely adopted as the first choice to represent 3D media. Besides several distortions that could affect the 3D content spanning from acquisition to…

Computer Vision and Pattern Recognition · Computer Science 2023-03-16 Marouane Tliba , Aladine Chetouani , Giuseppe Valenzise , Frederic Dufaux

In this paper, we propose PointRCNN for 3D object detection from raw point cloud. The whole framework is composed of two stages: stage-1 for the bottom-up 3D proposal generation and stage-2 for refining proposals in the canonical…

Computer Vision and Pattern Recognition · Computer Science 2019-05-17 Shaoshuai Shi , Xiaogang Wang , Hongsheng Li

Over the years, scene understanding has attracted a growing interest in computer vision, providing the semantic and physical scene information necessary for robots to complete some particular tasks autonomously. In 3D scenes, rich spatial…

Computer Vision and Pattern Recognition · Computer Science 2024-04-24 Gang Ma , Hui Wei

Visual Place Recognition (VPR) localizes a query image by matching it against a database of geo-tagged reference images, making it essential for navigation and mapping in robotics. Although Vision Transformer (ViT) solutions deliver high…

Computer Vision and Pattern Recognition · Computer Science 2025-03-05 Oliver Grainge , Michael Milford , Indu Bodala , Sarvapali D. Ramchurn , Shoaib Ehsan

Efficient storage of large-scale point cloud data has become increasingly challenging due to advancements in scanning technology. Recent deep learning techniques have revolutionized this field; However, most existing approaches rely on…

Computer Vision and Pattern Recognition · Computer Science 2024-09-20 Guoqing Zhang , Wenbo Zhao , Jian Liu , Yuanchao Bai , Junjun Jiang , Xianming Liu

Point cloud stands as the most widely adopted format for representing 3D shapes and scenes due to its simplicity and geometric fidelity. However, its inherent unordered and irregular nature, exacerbated by sensor noise and occlusions,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Minhas Kamal , Hiranya Garbha Kumar , Balakrishnan Prabhakaran

Effectively representing 3D scenes for Multimodal Large Language Models (MLLMs) is crucial yet challenging. Existing approaches commonly only rely on 2D image features and use varied tokenization approaches. This work presents a rigorous…

Computer Vision and Pattern Recognition · Computer Science 2025-06-09 Hugues Thomas , Chen Chen , Jian Zhang

In this paper, we present a new method that reformulates point cloud completion as a set-to-set translation problem and design a new model, called PoinTr, which adopts a Transformer encoder-decoder architecture for point cloud completion.…

Computer Vision and Pattern Recognition · Computer Science 2023-01-12 Xumin Yu , Yongming Rao , Ziyi Wang , Jiwen Lu , Jie Zhou

This paper presents a neural network built upon Transformers, namely PlaneTR, to simultaneously detect and reconstruct planes from a single image. Different from previous methods, PlaneTR jointly leverages the context information and the…

Computer Vision and Pattern Recognition · Computer Science 2021-07-29 Bin Tan , Nan Xue , Song Bai , Tianfu Wu , Gui-Song Xia

Point clouds provide a compact and expressive representation of 3D objects, and have recently been integrated into multimodal large language models (MLLMs). However, existing methods primarily focus on static objects, while understanding…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Xindan Zhang , Weilong Yan , Yufei Shi , Xuerui Qiu , Tao He , Ying Li , Ming Li , Hehe Fan

Place recognition, an algorithm to recognize the re-visited places, plays the role of back-end optimization trigger in a full SLAM system. Many works equipped with deep learning tools, such as MLP, CNN, and transformer, have achieved great…

Computer Vision and Pattern Recognition · Computer Science 2023-03-03 Zhixing Hou , Yuzhang Shang , Tian Gao , Yan Yan

Recent advances in multi-modal pre-training methods have shown promising effectiveness in learning 3D representations by aligning multi-modal features between 3D shapes and their corresponding 2D counterparts. However, existing multi-modal…

Computer Vision and Pattern Recognition · Computer Science 2025-07-24 Liwen Liu , Weidong Yang , Lipeng Ma , Ben Fei

Powerful 3D representations such as DUSt3R invariant point maps, which encode 3D shape and camera parameters, have significantly advanced feed forward 3D reconstruction. While point maps assume static scenes, Dynamic Point Maps (DPMs)…

Computer Vision and Pattern Recognition · Computer Science 2026-01-15 Edgar Sucar , Eldar Insafutdinov , Zihang Lai , Andrea Vedaldi

Understanding dynamic 3D environment is crucial for robotic agents and many other applications. We propose a novel neural network architecture called $MeteorNet$ for learning representations for dynamic 3D point cloud sequences. Different…

Computer Vision and Pattern Recognition · Computer Science 2020-02-11 Xingyu Liu , Mengyuan Yan , Jeannette Bohg

This paper focuses on the task of 4D shape reconstruction from a sequence of point clouds. Despite the recent success achieved by extending deep implicit representations into 4D space, it is still a great challenge in two respects, i.e. how…

Computer Vision and Pattern Recognition · Computer Science 2021-03-31 Jiapeng Tang , Dan Xu , Kui Jia , Lei Zhang

The spatio-temporal relationship between the pixels of a video carries critical information for low-level 4D perception tasks. A single model that reasons about it should be able to solve several such tasks well. Yet, most state-of-the-art…

Computer Vision and Pattern Recognition · Computer Science 2025-10-03 Abhishek Badki , Hang Su , Bowen Wen , Orazio Gallo
‹ Prev 1 4 5 6 7 8 10 Next ›