中文
相关论文

相关论文: UHR-BAT: Budget-Aware Token Compression Vision-Lan…

200 篇论文

Object detection in Ultra High-Resolution (UHR) images has long been a challenging problem in computer vision due to the varying scales of the targeted objects. When it comes to barcode detection, resizing UHR input images to smaller sizes…

计算机视觉与模式识别 · 计算机科学 2021-10-18 Jerome Quenum , Kehan Wang , Avideh Zakhor

Multimodal Large Language Models (MLLMs) have demonstrated immense potential in Earth observation. However, the massive visual tokens generated when processing Ultra-High-Resolution (UHR) imagery introduce prohibitive computational…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Yueying Li , Fengxiang Wang , Yan Li , Mingshuo Chen , Mengying Zhao , Long Lan

Ultra-High-Resolution (UHR) imagery has become essential for modern remote sensing, offering unprecedented spatial coverage. However, detecting small objects in such vast scenes presents a critical dilemma: retaining the original resolution…

计算机视觉与模式识别 · 计算机科学 2026-04-24 Jingfang Li , Haoran Zhu , Wen Yang , Jinrui Zhang , Fang Xu , Haijian Zhang , Gui-Song Xia

Ultra High Resolution (UHR) remote sensing imagery (RSI) (e.g. 100,000 $\times$ 100,000 pixels or more) poses a significant challenge for current Remote Sensing Multimodal Large Language Models (RSMLLMs). If choose to resize the UHR image…

计算机视觉与模式识别 · 计算机科学 2025-05-29 Zilun Zhang , Haozhan Shen , Tiancheng Zhao , Zian Guan , Bin Chen , Yuhao Wang , Xu Jia , Yuxiang Cai , Yongheng Shang , Jianwei Yin

Segmentation of ultra-high resolution (UHR) images is a critical task with numerous applications, yet it poses significant challenges due to high spatial resolution and rich fine details. Recent approaches adopt a dual-branch architecture,…

计算机视觉与模式识别 · 计算机科学 2024-12-24 Haopeng Sun , Yingwei Zhang , Lumin Xu , Sheng Jin , Yiqiang Chen

Semantic segmentation in remote sensing is commonly addressed using classical deep learning architectures such as U-Net, which require a large number of parameters to model complex spatial relationships. Quantum machine learning (QML)…

计算机视觉与模式识别 · 计算机科学 2026-05-01 Md Aminur Hossain , Ayush V. Patel , Ikshwaku Vanani , Biplab Banerjee

With the rapid development of ultra-high resolution (UHR) remote sensing technology, the demand for accurate and efficient semantic segmentation has increased significantly. However, existing methods face challenges in computational…

计算机视觉与模式识别 · 计算机科学 2025-06-25 Chen Yi , Shan LianLei

Unified models aim to support both understanding and generation by encoding images into discrete tokens and processing them alongside text within a single autoregressive framework. This unified design offers architectural simplicity and…

计算机视觉与模式识别 · 计算机科学 2026-03-13 Ziyao Wang , Chen Chen , Jingtao Li , Weiming Zhuang , Jiabo Huang , Ang Li , Lingjuan Lyu

Accurate lesion segmentation is crucial for clinical diagnosis and treatment planning. However, lesions often resemble surrounding tissues and exhibit ill-defined boundaries, leading to unstable predictions in boundary/transition regions.…

计算机视觉与模式识别 · 计算机科学 2026-05-01 Shuokun Cheng , Jinghao Shi , Kun Sun

The exponential growth of Large Multimodal Models (LMMs) has driven advancements in cross-modal reasoning but at significant computational costs. In this work, we focus on visual language models. We highlight the redundancy and inefficiency…

计算机视觉与模式识别 · 计算机科学 2025-04-28 Yasmine Omri , Parth Shroff , Thierry Tambe

With the increasing interest and rapid development of methods for Ultra-High Resolution (UHR) segmentation, a large-scale benchmark covering a wide range of scenes with full fine-grained dense annotations is urgently needed to facilitate…

计算机视觉与模式识别 · 计算机科学 2023-05-19 Deyi Ji , Feng Zhao , Hongtao Lu , Mingyuan Tao , Jieping Ye

A key challenge for large language models is token cost per query and overall deployment cost. Clinical inputs are long, heterogeneous, and often redundant, while downstream tasks are short and high stakes. We study budgeted context…

计算与语言 · 计算机科学 2026-05-04 Khizar Qureshi , Geoffrey Martin , Yifan Peng

A typical image retrieval pipeline starts with the comparison of global descriptors from a large database to find a short list of candidate matches. A good image descriptor is key to the retrieval pipeline and should reconcile two…

信息检索 · 计算机科学 2015-11-11 Jie Lin , Olivier Morère , Julie Petta , Vijay Chandrasekhar , Antoine Veillard

Recent Vision-Language Models (VLMs) have demonstrated remarkable multimodal understanding capabilities, yet the redundant visual tokens incur prohibitive computational overhead and degrade inference efficiency. Prior studies typically…

计算机视觉与模式识别 · 计算机科学 2026-02-24 Qiankun Ma , Ziyao Zhang , Haofei Wang , Jie Chen , Zhen Song , Hairong Zheng

Ultra-high-resolution (UHR) remote sensing (RS) images offer rich fine-grained information but also present challenges in effective processing. Existing dynamic resolution and token pruning methods are constrained by a passive perception…

计算机视觉与模式识别 · 计算机科学 2025-11-18 Ruixun Liu , Bowen Fu , Jiayi Song , Kaiyu Li , Wanchen Li , Lanxuan Xue , Hui Qiao , Weizhan Zhang , Deyu Meng , Xiangyong Cao

Refining visual representations by eliminating their internal feature-level redundancy is crucial for simultaneously optimizing the performance and computational cost of models in visual tracking. To enhance their performance, many…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Weijing Wu , Qihua Liang , Bineng Zhong , Haiying Xia , Zhiyi Mo , Shuxiang Song

We introduce a multi-scale Image Super Resolution (ISR) method building on recent advances in Visual Auto-Regressive (VAR) modeling. VAR models break image tokenization into additive, gradually increasing scales, using Residual Quantization…

计算机视觉与模式识别 · 计算机科学 2026-05-15 Isma Hadji , Enrique Sanchez , Adrian Bulat , Brais Martinez , Georgios Tzimiropoulos

Token compression techniques have recently emerged as powerful tools for accelerating Vision Transformer (ViT) inference in computer vision. Due to the quadratic computational complexity with respect to the token sequence length, these…

计算机视觉与模式识别 · 计算机科学 2025-07-15 Phat Nguyen , Ngai-Man Cheung

Real-world applications are stretching context windows to hundreds of thousand of tokens while Large Language Models (LLMs) swell from billions to trillions of parameters. This dual expansion send compute and memory costs skyrocketing,…

计算与语言 · 计算机科学 2025-12-12 Ling Xing , Alex Jinpeng Wang , Rui Yan , Xiangbo Shu , Jinhui Tang

Ultra-low bitrate image compression faces a critical challenge: preserving small-font scene text while maintaining overall visual quality. Region-of-interest (ROI) bit allocation can prioritize text but often degrades global fidelity,…

计算机视觉与模式识别 · 计算机科学 2026-03-05 Bingxin Wang , Yuan Lan , Zhaoyi Sun , Yang Xiang , Jie Sun
‹ 上一页 1 2 3 10 下一页 ›