English
Related papers

Related papers: BEVPoolv2: A Cutting-edge Implementation of BEVDet…

200 papers

Understanding point cloud has recently gained huge interests following the development of 3D scanning devices and the accumulation of large-scale 3D data. Most point cloud processing algorithms can be classified as either point-based or…

Computer Vision and Pattern Recognition · Computer Science 2022-02-07 Pyunghwan Ahn , Juyoung Yang , Eojindl Yi , Chanho Lee , Junmo Kim

The recent introduction of powerful embedded graphics processing units (GPUs) has allowed for unforeseen improvements in real-time computer vision applications. It has enabled algorithms to run onboard, well above the standard video rates,…

Computer Vision and Pattern Recognition · Computer Science 2020-08-04 Balazs Nagy , Philipp Foehn , Davide Scaramuzza

We present VGGT-SLAM 2.0, a real-time RGB feed-forward SLAM system which substantially improves upon VGGT-SLAM for incrementally aligning submaps created from VGGT. Firstly, we remove high-dimensional 15-degree-of-freedom drift and planar…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Dominic Maggio , Luca Carlone

We present Vchitect-2.0, a parallel transformer architecture designed to scale up video diffusion models for large-scale text-to-video generation. The overall Vchitect-2.0 system has several key designs. (1) By introducing a novel…

Computer Vision and Pattern Recognition · Computer Science 2025-01-16 Weichen Fan , Chenyang Si , Junhao Song , Zhenyu Yang , Yinan He , Long Zhuo , Ziqi Huang , Ziyue Dong , Jingwen He , Dongwei Pan , Yi Wang , Yuming Jiang , Yaohui Wang , Peng Gao , Xinyuan Chen , Hengjie Li , Dahua Lin , Yu Qiao , Ziwei Liu

Using synthesized images to boost the performance of perception models is a long-standing research challenge in computer vision. It becomes more eminent in visual-centric autonomous driving systems with multi-view cameras as some long-tail…

Computer Vision and Pattern Recognition · Computer Science 2023-09-26 Kairui Yang , Enhui Ma , Jibin Peng , Qing Guo , Di Lin , Kaicheng Yu

We present a new video compression framework (ViSTRA2) which exploits adaptation of spatial resolution and effective bit depth, down-sampling these parameters at the encoder based on perceptual criteria, and up-sampling at the decoder using…

Image and Video Processing · Electrical Eng. & Systems 2021-06-22 Fan Zhang , Mariana Afonso , David R. Bull

Depth estimation is a cornerstone of perception in autonomous driving and robotic systems. The considerable cost and relatively sparse data acquisition of LiDAR systems have led to the exploration of cost-effective alternatives, notably,…

Computer Vision and Pattern Recognition · Computer Science 2023-06-21 Yucheng Mao , Ruowen Zhao , Tianbao Zhang , Hang Zhao

Implicit neural representations for videos (NeRV) have shown strong potential for video compression. However, applying NeRV to high-resolution 360-degree videos causes high memory usage and slow decoding, making real-time applications…

Computer Vision and Pattern Recognition · Computer Science 2026-02-17 Daichi Arai , Kyohei Unno , Yasuko Sugito , Yuichi Kusakabe

Most downstream adaptation methods tune all or part of the parameters of pre-trained models (PTMs) through gradient descent, where the tuning cost increases linearly with the growth of the model size. By contrast, gradient-free methods only…

Computation and Language · Computer Science 2022-10-17 Tianxiang Sun , Zhengfu He , Hong Qian , Yunhua Zhou , Xuanjing Huang , Xipeng Qiu

Inspired by the long-range modeling ability of ViTs, large-kernel convolutions are widely studied and adopted recently to enlarge the receptive field and improve model performance, like the remarkable work ConvNeXt which employs 7x7…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Weihao Yu , Pan Zhou , Shuicheng Yan , Xinchao Wang

We introduce Qwen2.5-VL, the latest flagship model of Qwen vision-language series, which demonstrates significant advancements in both foundational capabilities and innovative functionalities. Qwen2.5-VL achieves a major leap forward in…

High-resolution dense prediction enables many appealing real-world applications, such as computational photography, autonomous driving, etc. However, the vast computational cost makes deploying state-of-the-art high-resolution dense…

Computer Vision and Pattern Recognition · Computer Science 2024-02-07 Han Cai , Junyan Li , Muyan Hu , Chuang Gan , Song Han

Video-to-video translation aims to generate video frames of a target domain from an input video. Despite its usefulness, the existing networks require enormous computations, necessitating their model compression for wide use. While there…

Computer Vision and Pattern Recognition · Computer Science 2023-10-05 Chaeyeon Chung , Yeojeong Park , Seunghwan Choi , Munkhsoyol Ganbat , Jaegul Choo

Generating a detailed near-field perceptual model of the environment is an important and challenging problem in both self-driving vehicles and autonomous mobile robotics. A Bird Eye View (BEV) map, providing a panoptic representation, is a…

Computer Vision and Pattern Recognition · Computer Science 2022-06-01 Pramit Dutta , Ganesh Sistu , Senthil Yogamani , Edgar Galván , John McDonald

This work focuses on developing parameter-efficient and lightweight models for dense predictions while trading off parameters, FLOPs, and performance. Our goal is to set up the new frontier of the 5M magnitude lightweight model on various…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Jiangning Zhang , Teng Hu , Haoyang He , Zhucun Xue , Yabiao Wang , Chengjie Wang , Yong Liu , Xiangtai Li , Dacheng Tao

In the era of large foundation models, the quality of embeddings has become a central determinant of downstream task performance and overall system capability. Yet widely used dense embeddings are often extremely high-dimensional, incurring…

Machine Learning · Computer Science 2026-03-03 Lixuan Guo , Yifei Wang , Tiansheng Wen , Yifan Wang , Aosong Feng , Bo Chen , Stefanie Jegelka , Chenyu You

We introduce ProVideLLM, an end-to-end framework for real-time procedural video understanding. ProVideLLM integrates a multimodal cache configured to store two types of tokens - verbalized text tokens, which provide compressed textual…

Computer Vision and Pattern Recognition · Computer Science 2025-04-22 Dibyadip Chatterjee , Edoardo Remelli , Yale Song , Bugra Tekin , Abhay Mittal , Bharat Bhatnagar , Necati Cihan Camgöz , Shreyas Hampali , Eric Sauser , Shugao Ma , Angela Yao , Fadime Sener

With the advancement of deep learning technologies, specialized neural processing hardware such as Brain Processing Units (BPUs) have emerged as dedicated platforms for CNN acceleration, offering optimized INT8 computation capabilities for…

Computer Vision and Pattern Recognition · Computer Science 2026-02-09 Jinchi Tang , Yan Guo

We describe a simple, low-level approach for embedding probabilistic programming in a deep learning ecosystem. In particular, we distill probabilistic programming down to a single abstraction---the random variable. Our lightweight…

Large Language Models (LLMs) are widely used across various domains, processing millions of daily requests. This surge in demand poses significant challenges in optimizing throughput and latency while keeping costs manageable. The Key-Value…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-07-23 Jiale Xu , Rui Zhang , Cong Guo , Weiming Hu , Zihan Liu , Feiyang Wu , Yu Feng , Shixuan Sun , Changxu Shao , Yuhong Guo , Junping Zhao , Ke Zhang , Minyi Guo , Jingwen Leng