English
Related papers

Related papers: TC-SSA: Token Compression via Semantic Slot Aggreg…

200 papers

Compressed sensing (CS) exploits the sparsity of a signal in order to integrate acquisition and compression. CS theory enables exact reconstruction of a sparse signal from relatively few linear measurements via a suitable nonlinear…

Information Theory · Computer Science 2014-09-04 Shmuel Friedland , Qun Li , Dan Schonfeld , Edgar A. Bernal

Compositional Zero-Shot Learning (CZSL) aims to recognize novel state-object compositions by leveraging the shared knowledge of their primitive components. Despite considerable progress, effectively calibrating the bias between semantically…

Computer Vision and Pattern Recognition · Computer Science 2025-01-28 Miaoge Li , Jingcai Guo , Richard Yi Da Xu , Dongsheng Wang , Xiaofeng Cao , Zhijie Rao , Song Guo

Deep neural networks (DNNs) are the de facto standard for essential use cases, such as image classification, computer vision, and natural language processing. As DNNs and datasets get larger, they require distributed training on…

Machine Learning · Computer Science 2024-03-07 Minghao Li , Ran Ben Basat , Shay Vargaftik , ChonLam Lao , Kevin Xu , Michael Mitzenmacher , Minlan Yu

In this study, we introduce a novel method called group-wise \textbf{VI}sual token \textbf{S}election and \textbf{A}ggregation (VISA) to address the issue of inefficient inference stemming from excessive visual tokens in multimoal large…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Pengfei Jiang , Hanjun Li , Linglan Zhao , Fei Chao , Ke Yan , Shouhong Ding , Rongrong Ji

Digital whole slide images (WSIs) are generally captured at microscopic resolution and encompass extensive spatial data. Directly feeding these images to deep learning models is computationally intractable due to memory constraints, while…

Image and Video Processing · Electrical Eng. & Systems 2024-11-22 Manahil Raza , Ruqayya Awan , Raja Muhammad Saad Bashir , Talha Qaiser , Nasir M. Rajpoot

Implementation of digital pathology leads to an increased number of whole slide images (WSIs). The large size of WSIs is challenging. Today, WSIs are compressed with codecs like JPEG resulting in several gigabytes per WSI, and large amounts…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Maren Høibø , Etienne Gaucher , Ingerid Reinertsen , Marit Valla , Erik Smistad

Large language models (LLMs) can memorize and reproduce training sequences verbatim -- a tendency that undermines both generalization and privacy. Existing mitigation methods apply interventions uniformly, degrading performance on the…

Machine Learning · Computer Science 2026-02-10 Xuanqi Zhang , Haoyang Shang , Xiaoxiao Li

Multimodal models have achieved remarkable success in natural image segmentation, yet they often underperform when applied to the medical domain. Through extensive study, we attribute this performance gap to the challenges of multimodal…

Computer Vision and Pattern Recognition · Computer Science 2025-09-11 Wenjun Yu , Yinchen Zhou , Jia-Xuan Jiang , Shubin Zeng , Yuee Li , Zhong Wang

Recently, end-to-end automatic speech recognition has become the mainstream approach in both industry and academia. To optimize system performance in specific scenarios, the Weighted Finite-State Transducer (WFST) is extensively used to…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-08 Wei Zhang , Tian-Hao Zhang , Chao Luo , Hui Zhou , Chao Yang , Xinyuan Qian , Xu-Cheng Yin

Recent advances in vision-language models (VLMs) have garnered substantial attention in open-vocabulary semantic and part segmentation (OSPS). However, existing methods extract image-text alignment cues from cost volumes through a serial…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Jianjian Yin , Tao Chen , Yi Chen , Gensheng Pei , Xiangbo Shu , Yazhou Yao , Fumin Shen

A major obstacle to building models for effective semantic segmentation, and particularly video semantic segmentation, is a lack of large and well annotated datasets. This bottleneck is particularly prohibitive in highly specialized and…

Deep-learning techniques have been used widely to alleviate the labour-intensive and time-consuming manual annotation required for pixel-level tissue characterization. Our previous study introduced an efficient single dynamic network -…

Image and Video Processing · Electrical Eng. & Systems 2023-05-25 Haoju Leng , Ruining Deng , Zuhayr Asad , R. Michael Womick , Haichun Yang , Lipeng Wan , Yuankai Huo

Network data are commonly collected in a variety of applications, representing either directly measured or statistically inferred connections between features of interest. In an increasing number of domains, these networks are collected…

Machine Learning · Statistics 2022-09-05 Michael Weylandt , George Michailidis

The advent of large-scale self-supervised learning (SSL) has produced a vast zoo of medical foundation models. However, selecting optimal medical foundation models for specific segmentation tasks remains a computational bottleneck. Existing…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Jiaqi Tang , Shaoyang Zhang , Xiaoqi Wang , Jiaying Zhou , Yang Liu , Qingchao Chen

Existing KV cache compression methods generally operate on discrete tokens or non-semantic chunks. However, such approaches often lead to semantic fragmentation, where linguistically coherent units are disrupted, causing irreversible…

Computation and Language · Computer Science 2026-03-17 Shunlong Wu , Hai Lin , Shaoshen Chen , Tingwei Lu , Yongqin Zeng , Shaoxiong Zhan , Hai-Tao Zheng , Hong-Gee Kim

Linear attention Transformers and their gated variants, celebrated for enabling parallel training and efficient recurrent inference, still fall short in recall-intensive tasks compared to traditional Transformers and demand significant…

Computation and Language · Computer Science 2024-11-01 Yu Zhang , Songlin Yang , Ruijie Zhu , Yue Zhang , Leyang Cui , Yiqiao Wang , Bolun Wang , Freda Shi , Bailin Wang , Wei Bi , Peng Zhou , Guohong Fu

Histopathological whole slide images (WSIs) classification has become a foundation task in medical microscopic imaging processing. Prevailing approaches involve learning WSIs as instance-bag representations, emphasizing significant…

Computer Vision and Pattern Recognition · Computer Science 2024-03-13 Jiawen Li , Yuxuan Chen , Hongbo Chu , Qiehe Sun , Tian Guan , Anjia Han , Yonghong He

More accurate machine learning models often demand more computation and memory at test time, making them difficult to deploy on CPU- or memory-constrained devices. Teacher-student compression (TSC), also known as distillation, alleviates…

Machine Learning · Computer Science 2020-03-24 Ruishan Liu , Nicolo Fusi , Lester Mackey

This paper presents an autoencoder-based neural network architecture to compress histopathological images while retaining the denser and more meaningful representation of the original images. Current research into improving compression…

Image and Video Processing · Electrical Eng. & Systems 2023-05-15 Agnes Barsi , Suvendu Chandan Nayak , Sasmita Parida , Raj Mani Shukla

Whole Slide Image (WSI) classification is often formulated as a Multiple Instance Learning (MIL) problem. Recently, Vision-Language Models (VLMs) have demonstrated remarkable performance in WSI classification. However, existing methods…

Computer Vision and Pattern Recognition · Computer Science 2024-04-08 Hao Li , Ying Chen , Yifei Chen , Wenxian Yang , Bowen Ding , Yuchen Han , Liansheng Wang , Rongshan Yu
‹ Prev 1 8 9 10 Next ›