English
Related papers

Related papers: Manboformer: Learning Gaussian Representations via…

200 papers

Multi-agent trajectory prediction is a fundamental problem in autonomous driving. The key challenges in prediction are accurately anticipating the behavior of surrounding agents and understanding the scene context. To address these…

Computer Vision and Pattern Recognition · Computer Science 2022-03-04 Elmira Amirloo , Amir Rasouli , Peter Lakner , Mohsen Rohani , Jun Luo

3D Gaussian Splatting has recently gained traction for its efficient training and real-time rendering. While its vanilla representation is mainly designed for view synthesis, recent works extended it to scene understanding with language…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Siyun Liang , Sen Wang , Kunyi Li , Michael Niemeyer , Stefano Gasperini , Hendrik P. A. Lensch , Nassir Navab , Federico Tombari

Occupancy is crucial for autonomous driving, providing essential geometric priors for perception and planning. However, existing methods predominantly rely on LiDAR-based occupancy annotations, which limits scalability and prevents…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Baijun Ye , Minghui Qin , Saining Zhang , Moonjun Gong , Shaoting Zhu , Zebang Shen , Luan Zhang , Lu Zhang , Hao Zhao , Hang Zhao

Reconstructing dynamic 3D scenes with photorealistic detail and strong temporal coherence remains a significant challenge. Existing Gaussian splatting approaches for dynamic scene modeling often rely on per-frame optimization, which can…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Tingxuan Huang , Haowei Zhu , Jun-hai Yong , Hao Pan , Bin Wang

We present latentSplat, a method to predict semantic Gaussians in a 3D latent space that can be splatted and decoded by a light-weight generative 2D architecture. Existing methods for generalizable 3D reconstruction either do not scale to…

Computer Vision and Pattern Recognition · Computer Science 2024-07-31 Christopher Wewer , Kevin Raj , Eddy Ilg , Bernt Schiele , Jan Eric Lenssen

Reliable multimodal sensor fusion algorithms require accurate spatiotemporal calibration. Recently, targetless calibration techniques based on implicit neural representations have proven to provide precise and robust results. Nevertheless,…

Computer Vision and Pattern Recognition · Computer Science 2024-09-19 Quentin Herau , Moussab Bennehar , Arthur Moreau , Nathan Piasco , Luis Roldao , Dzmitry Tsishkou , Cyrille Migniot , Pascal Vasseur , Cédric Demonceaux

We present the first unified framework for rate-distortion-optimized compression and segmentation of 3D Gaussian Splatting (3DGS). While 3DGS has proven effective for both real-time rendering and semantic scene understanding, prior works…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Yu-Jen Tseng , Chia-Hao Kao , Jing-Zhong Chen , Alessandro Gnutti , Shao-Yuan Lo , Yen-Yu Lin , Wen-Hsiao Peng

The advent of neural 3D Gaussians has recently brought about a revolution in the field of neural rendering, facilitating the generation of high-quality renderings at real-time speeds. However, the explicit and discrete representation…

Computer Vision and Pattern Recognition · Computer Science 2023-12-01 Yingwenqi Jiang , Jiadong Tu , Yuan Liu , Xifeng Gao , Xiaoxiao Long , Wenping Wang , Yuexin Ma

Diffusion Transformer, the backbone of Sora for video generation, successfully scales the capacity of diffusion models, pioneering new avenues for high-fidelity sequential data generation. Unlike static data such as images, sequential data…

Machine Learning · Computer Science 2025-02-05 Hengyu Fu , Zehao Dou , Jiawei Guo , Mengdi Wang , Minshuo Chen

Recent advancements in 3D reconstruction methods and vision-language models have propelled the development of multi-modal 3D scene understanding, which has vital applications in robotics, autonomous driving, and virtual/augmented reality.…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Qucheng Peng , Benjamin Planche , Zhongpai Gao , Meng Zheng , Anwesa Choudhuri , Terrence Chen , Chen Chen , Ziyan Wu

Recent advancements in dynamic 3D scene reconstruction have shown promising results, enabling high-fidelity 3D novel view synthesis with improved temporal consistency. Among these, 4D Gaussian Splatting (4DGS) has emerged as an appealing…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Seungjun Oh , Younggeun Lee , Hyejin Jeon , Eunbyung Park

Semantic-aware 3D scene reconstruction is essential for autonomous robots to perform complex interactions. Semantic SLAM, an online approach, integrates pose tracking, geometric reconstruction, and semantic mapping into a unified framework,…

Robotics · Computer Science 2025-05-20 Zuxing Lu , Xin Yuan , Shaowen Yang , Jingyu Liu , Changyin Sun

Transformers have recently gained attention in the computer vision domain due to their ability to model long-range dependencies. However, the self-attention mechanism, which is the core part of the Transformer model, usually suffers from…

Computer Vision and Pattern Recognition · Computer Science 2023-07-28 Reza Azad , René Arimond , Ehsan Khodapanah Aghdam , Amirhossein Kazerouni , Dorit Merhof

Gaze is an essential prompt for analyzing human behavior and attention. Recently, there has been an increasing interest in determining gaze direction from facial videos. However, video gaze estimation faces significant challenges, such as…

Computer Vision and Pattern Recognition · Computer Science 2024-04-11 Swati Jindal , Mohit Yadav , Roberto Manduchi

To automatically localize a target object in an image is crucial for many computer vision applications. To represent the 2D object, ellipse labels have recently been identified as a promising alternative to axis-aligned bounding boxes. This…

Computer Vision and Pattern Recognition · Computer Science 2023-08-03 Vincent Gaudillière , Leo Pauly , Arunkumar Rathinam , Albert Garcia Sanchez , Mohamed Adel Musallam , Djamila Aouada

A self-driving vehicle must understand its environment to determine the appropriate action. Traditional autonomy systems rely on object detection to find the agents in the scene. However, object detection assumes a discrete set of objects…

Robotics · Computer Science 2024-04-03 Sourav Biswas , Sergio Casas , Quinlan Sykora , Ben Agro , Abbas Sadat , Raquel Urtasun

Recent advances in 3D Gaussian Splatting (3DGS) have enabled significant progress in photorealistic novel view synthesis. However, traditional 3DGS relies on a slow, iterative optimization process, which limits its use in scenarios…

Computer Vision and Pattern Recognition · Computer Science 2026-05-14 Can Wang , Lei Liu , Wei Jiang , Dong Xu

Generating robust and reliable correspondences across images is a fundamental task for a diversity of applications. To capture context at both global and local granularity, we propose ASpanFormer, a Transformer-based detector-free matcher…

Computer Vision and Pattern Recognition · Computer Science 2022-08-31 Hongkai Chen , Zixin Luo , Lei Zhou , Yurun Tian , Mingmin Zhen , Tian Fang , David Mckinnon , Yanghai Tsin , Long Quan

Human activity intensity prediction is crucial to many location-based services. Despite tremendous progress in modeling dynamics of human activity, most existing methods overlook physical constraints of spatial interaction, leading to…

Machine Learning · Computer Science 2025-10-27 Yi Wang , Zhenghong Wang , Fan Zhang , Chaogui Kang , Sijie Ruan , Di Zhu , Chengling Tang , Zhongfu Ma , Weiyu Zhang , Yu Zheng , Philip S. Yu , Yu Liu

Graph-based next-step prediction models have recently been very successful in modeling complex high-dimensional physical systems on irregular meshes. However, due to their short temporal attention span, these models suffer from error…

Machine Learning · Computer Science 2022-05-27 Xu Han , Han Gao , Tobias Pfaff , Jian-Xun Wang , Li-Ping Liu