English
Related papers

Related papers: mTREE: Multi-Level Text-Guided Representation End-…

200 papers

Whole Slide Images (WSIs) are critical for various clinical applications, including histopathological analysis. However, current deep learning approaches in this field predominantly focus on individual tumor types, limiting model…

Image and Video Processing · Electrical Eng. & Systems 2024-09-18 Sharon Peled , Yosef E. Maruvka , Moti Freiman

We introduce ENTIRE, a novel deep learning-based approach for fast and accurate volume rendering time prediction. Predicting rendering time is inherently challenging due to its dependence on multiple factors, including volume data…

Graphics · Computer Science 2026-04-21 Zikai Yin , Hamid Gadirov , Jiri Kosinka , Steffen Frey

Whole Slide Imaging (WSI) has become a gold standard in cancer diagnosis, inspecting multi-scale information from cellular to tissue levels. Processing an entire WSI directly is infeasible due to GPU memory constraints; thus, Multiple…

Image and Video Processing · Electrical Eng. & Systems 2026-05-08 Tianyi Zhang , Sicheng Chen , Borui Kang , Dankai Liao , Qiaochu Xue , Bochong Zhang , Fei Xia , Zeyu Liu , Yueming Jin

In computational pathology, multiple instance learning (MIL) is widely used to circumvent the computational impasse in giga-pixel whole slide image (WSI) analysis. It usually consists of two stages: patch-level feature extraction and…

Image and Video Processing · Electrical Eng. & Systems 2023-09-21 Beidi Zhao , Wenlong Deng , Zi Han , Li , Chen Zhou , Zuhua Gao , Gang Wang , Xiaoxiao Li

Accurate survival prediction from histopathology whole-slide images (WSIs) remains challenging due to their gigapixel resolution, strong spatial heterogeneity, and complex survival distributions. We introduce a comprehensive computational…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Ardhendu Sekhar , Vasu Soni , Keshav Aske , Shivam Madnoorkar , Pranav Jeevan , Amit Sethi

Unifying text detection and text recognition in an end-to-end training fashion has become a new trend for reading text in the wild, as these two tasks are highly relevant and complementary. In this paper, we investigate the problem of scene…

Computer Vision and Pattern Recognition · Computer Science 2019-08-23 Minghui Liao , Pengyuan Lyu , Minghang He , Cong Yao , Wenhao Wu , Xiang Bai

A mainstream type of current self-supervised learning methods pursues a general-purpose representation that can be well transferred to downstream tasks, typically by optimizing on a given pretext task such as instance discrimination. In…

Computer Vision and Pattern Recognition · Computer Science 2022-10-21 Xin Liu , Zhongdao Wang , Yali Li , Shengjin Wang

Whole slide image (WSI) analysis has become increasingly important in the medical imaging community, enabling automated and objective diagnosis, prognosis, and therapeutic-response prediction. However, in clinical practice, the…

Computer Vision and Pattern Recognition · Computer Science 2023-08-28 Yanyan Huang , Weiqin Zhao , Shujun Wang , Yu Fu , Yuming Jiang , Lequan Yu

Building scalable models to learn from diverse, multimodal data remains an open challenge. For vision-language data, the dominant approaches are based on contrastive learning objectives that train a separate encoder for each modality. While…

Computer Vision and Pattern Recognition · Computer Science 2022-10-24 Xinyang Geng , Hao Liu , Lisa Lee , Dale Schuurmans , Sergey Levine , Pieter Abbeel

The multi-scale information among the whole slide images (WSIs) is essential for cancer diagnosis. Although the existing multi-scale vision Transformer has shown its effectiveness for learning multi-scale image representation, it still…

Computer Vision and Pattern Recognition · Computer Science 2023-05-26 Saisai Ding , Juncheng Li , Jun Wang , Shihui Ying , Jun Shi

Missing input sequences are common in medical imaging data, posing a challenge for deep learning models reliant on complete input data. In this work, inspired by MultiMAE [2], we develop a masked autoencoder (MAE) paradigm for multi-modal,…

Computer Vision and Pattern Recognition · Computer Science 2026-02-04 Ayhan Can Erdur , Christian Beischl , Daniel Scholz , Jiazhen Pan , Benedikt Wiestler , Daniel Rueckert , Jan C Peeken

Multiple instance learning (MIL) has been increasingly used in the classification of histopathology whole slide images (WSIs). However, MIL approaches for this specific classification problem still face unique challenges, particularly those…

Computer Vision and Pattern Recognition · Computer Science 2022-03-24 Hongrun Zhang , Yanda Meng , Yitian Zhao , Yihong Qiao , Xiaoyun Yang , Sarah E. Coupland , Yalin Zheng

Whole Slide Image (WSI) classification poses unique challenges due to the vast image size and numerous non-informative regions, which introduce noise and cause data imbalance during feature aggregation. To address these issues, we propose…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Jianwei Zhao , Xin Li , Fan Yang , Qiang Zhai , Ao Luo , Yang Zhao , Hong Cheng , Huazhu Fu

Due to the large size and lack of fine-grained annotation, Whole Slide Images (WSIs) analysis is commonly approached as a Multiple Instance Learning (MIL) problem. However, previous studies only learn from training data, posing a stark…

Computer Vision and Pattern Recognition · Computer Science 2024-11-28 Weiqin Zhao , Ziyu Guo , Yinshuang Fan , Yuming Jiang , Maximus Yeung , Lequan Yu

The current state-of-the-art for image annotation and image retrieval tasks is obtained through deep neural networks, which combine an image representation and a text representation into a shared embedding space. In this paper we evaluate…

Computer Vision and Pattern Recognition · Computer Science 2017-08-10 Armand Vilalta , Dario Garcia-Gasulla , Ferran Parés , Eduard Ayguadé , Jesus Labarta , Ulises Cortés , Toyotaro Suzumura

Multimodality Representation Learning, as a technique of learning to embed information from different modalities and their correlations, has achieved remarkable success on a variety of applications, such as Visual Question Answering (VQA),…

Artificial Intelligence · Computer Science 2024-03-04 Muhammad Arslan Manzoor , Sarah Albarri , Ziting Xian , Zaiqiao Meng , Preslav Nakov , Shangsong Liang

The field of advanced text-to-image generation is witnessing the emergence of unified frameworks that integrate powerful text encoders, such as CLIP and T5, with Diffusion Transformer backbones. Although there have been efforts to control…

Computer Vision and Pattern Recognition · Computer Science 2025-02-28 Liang Chen , Shuai Bai , Wenhao Chai , Weichu Xie , Haozhe Zhao , Leon Vinci , Junyang Lin , Baobao Chang

Image-text matching is a key multimodal task that aims to model the semantic association between images and text as a matching relationship. With the advent of the multimedia information age, image, and text data show explosive growth, and…

Machine Learning · Computer Science 2024-06-24 Jinyin Wang , Haijing Zhang , Yihao Zhong , Yingbin Liang , Rongwei Ji , Yiru Cang

Accurate prediction of placental diseases via whole slide images (WSIs) is critical for preventing severe maternal and fetal complications. However, WSI analysis presents significant computational challenges due to the massive data volume.…

Computer Vision and Pattern Recognition · Computer Science 2025-08-06 Hang Guo , Qing Zhang , Zixuan Gao , Siyuan Yang , Shulin Peng , Xiang Tao , Ting Yu , Yan Wang , Qingli Li

Remote sensing image interpretation plays a critical role in environmental monitoring, urban planning, and disaster assessment. However, acquiring high-quality labeled data is often costly and time-consuming. To address this challenge, we…

Computer Vision and Pattern Recognition · Computer Science 2026-01-27 Tong Wang , Guanzhou Chen , Xiaodong Zhang , Chenxi Liu , Jiaqi Wang , Xiaoliang Tan , Wenchao Guo , Qingyuan Yang , Kaiqi Zhang