English
Related papers

Related papers: CPathAgent: An Agent-based Foundation Model for In…

200 papers

We present HippoCamp, a new benchmark designed to evaluate agents' capabilities on multimodal file management. Unlike existing agent benchmarks that focus on tasks like web interaction, tool use, or software automation in generic settings,…

Artificial Intelligence · Computer Science 2026-04-02 Zhe Yang , Shulin Tian , Kairui Hu , Shuai Liu , Hoang-Nhat Nguyen , Yichi Zhang , Zujin Guo , Mengying Yu , Zinan Zhang , Jingkang Yang , Chen Change Loy , Ziwei Liu

Photo retouching is integral to photographic art, extending far beyond simple technical fixes to heighten emotional expression and narrative depth. While artists leverage expertise to create unique visual effects through deliberate…

Computer Vision and Pattern Recognition · Computer Science 2025-05-30 Haoyu Chen , Keda Tao , Yizao Wang , Xinlei Wang , Lei Zhu , Jinjin Gu

Histopathology remains the gold standard for cancer diagnosis because it provides detailed cellular-level assessment of tissue morphology. However, manual histopathological examination is time-consuming, labour-intensive, and subject to…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Ravi Mosalpuri , Mohammed Abdelsamea , Ahmed Karam Eldaly

Recent advancements in large foundation models have remarkably enhanced our understanding of sensory information in open-world environments. In leveraging the power of foundation models, it is crucial for AI research to pivot away from…

Omnimodal large language models have made significant strides in unifying audio and visual modalities; however, they often face challenges in fine-grained cross-modal understanding and have difficulty with multimodal alignment. To address…

Computer Vision and Pattern Recognition · Computer Science 2026-02-06 Keda Tao , Wenjie Du , Bohan Yu , Weiqiang Wang , Jian Liu , Huan Wang

Large language models show potential for scalable mental-health support by simulating Cognitive Behavioral Therapy (CBT) counselors. However, existing methods often rely on static cognitive profiles and omniscient single-agent simulation,…

Computation and Language · Computer Science 2026-04-09 Chang Liu , Changsheng Ma , Yongfeng Tao , Bin Hu , Minqiang Yang

The emergence of foundation models in computational pathology has transformed histopathological image analysis, with whole slide imaging (WSI) diagnosis being a core application. Traditionally, weakly supervised fine-tuning via multiple…

Computer Vision and Pattern Recognition · Computer Science 2025-03-03 Jiawen Li , Jiali Hu , Qiehe Sun , Renao Yan , Minxi Ouyang , Tian Guan , Anjia Han , Chao He , Yonghong He

We introduce PhysicalAgent, an agentic framework for robotic manipulation that integrates iterative reasoning, diffusion-based video generation, and closed-loop execution. Given a textual instruction, our method generates short video…

Designing realistic multi-object scenes requires not only generating images, but also planning spatial layouts that respect semantic relations and physical plausibility. On one hand, while recent advances in diffusion models have enabled…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Zezhong Fan , Xiaohan Li , Luyi Ma , Kai Zhao , Liang Peng , Topojoy Biswas , Evren Korpeoglu , Kaushiki Nag , Kannan Achan

Advances in optical microscopy scanning have significantly contributed to computational pathology (CPath) by converting traditional histopathological slides into whole slide images (WSIs). This development enables comprehensive digital…

Image and Video Processing · Electrical Eng. & Systems 2024-11-19 Xitong Ling , Yuanyuan Lei , Jiawen Li , Junru Cheng , Wenting Huang , Tian Guan , Jian Guan , Yonghong He

In pathological research, education, and clinical practice, the decision-making process based on pathological images is critically important. This significance extends to digital pathology image analysis: its adequacy is demonstrated by the…

Image and Video Processing · Electrical Eng. & Systems 2024-08-19 Zhi-Bo Liu , Xiaobo Pang , Jizhao Wang , Shuai Liu , Chen Li

Diagnosing lung cancer typically involves physicians identifying lung nodules in Computed tomography (CT) scans and generating diagnostic reports based on their morphological features and medical expertise. Although advancements have been…

Computer Vision and Pattern Recognition · Computer Science 2025-11-27 Cheng Yang , Hui Jin , Xinlei Yu , Zhipeng Wang , Yaoqun Liu , Fenglei Fan , Dajiang Lei , Gangyong Jia , Changmiao Wang , Ruiquan Ge

The previous advancements in pathology image understanding primarily involved developing models tailored to specific tasks. Recent studies has demonstrated that the large vision-language model can enhance the performance of various…

Artificial Intelligence · Computer Science 2024-08-20 Dawei Dai , Yuanhui Zhang , Long Xu , Qianlan Yang , Xiaojing Shen , Shuyin Xia , Guoyin Wang

Chest X-ray (CXR) plays a pivotal role in clinical diagnosis, and a variety of task-specific and foundation models have been developed for automatic CXR interpretation. However, these models often struggle to adapt to new diagnostic tasks…

Artificial Intelligence · Computer Science 2025-10-27 Jinhui Lou , Yan Yang , Zhou Yu , Zhenqi Fu , Weidong Han , Qingming Huang , Jun Yu

In histopathology, tissue samples are often larger than a standard microscope slide, making stitching of multiple fragments necessary to process entire structures such as tumors. Automated stitching is a prerequisite for scaling analysis,…

Computer Vision and Pattern Recognition · Computer Science 2025-08-06 Stefan Brandstätter , Maximilian Köller , Philipp Seeböck , Alissa Blessing , Felicitas Oberndorfer , Svitlana Pochepnia , Helmut Prosch , Georg Langs

Radiology reporting generative AI holds significant potential to alleviate clinical workloads and streamline medical care. However, achieving high clinical accuracy is challenging, as radiological images often feature subtle lesions and…

Computer Vision and Pattern Recognition · Computer Science 2024-11-19 Yijian Gao , Dominic Marshall , Xiaodan Xing , Junzhi Ning , Giorgos Papanastasiou , Guang Yang , Matthieu Komorowski

Pretraining on large-scale, in-domain datasets grants histopathology foundation models (FM) the ability to learn task-agnostic data representations, enhancing transfer learning on downstream tasks. In computational pathology, automated…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Pablo Meseguer , Rocío del Amor , Valery Naranjo

In recent years, a standard computational pathology workflow has emerged where whole slide images are cropped into tiles, these tiles are processed using a foundation model, and task-specific models are built using the resulting…

Computer Vision and Pattern Recognition · Computer Science 2026-02-26 Eric Zimmermann , Julian Viret , Michal Zelechowski , James Brian Hall , Neil Tenenholtz , Adam Casson , George Shaikovski , Eugene Vorontsov , Siqi Liu , Kristen A Severson

We propose a CNN based technique that aggregates feature maps from its multiple layers that can localize abnormalities with greater details as well as predict pathology under consideration. Existing class activation mapping (CAM) techniques…

Computer Vision and Pattern Recognition · Computer Science 2019-10-01 Sumeet Shinde , Tanay Chougule , Jitender Saini , Madhura Ingalhalikar

Amodal completion, generating invisible parts of occluded objects, is vital for applications like image editing and AR. Prior methods face challenges with data needs, generalization, or error accumulation in progressive pipelines. We…

Computer Vision and Pattern Recognition · Computer Science 2025-09-23 Hongxing Fan , Lipeng Wang , Haohua Chen , Zehuan Huang , Jiangtao Wu , Lu Sheng