中文
相关论文

相关论文: HiPath: Hierarchical Vision-Language Alignment for…

200 篇论文

Pathology foundation models (PFMs) have enabled robust generalization in computational pathology through large-scale datasets and expansive architectures, but their substantial computational cost, particularly for gigapixel whole slide…

The condition monitoring (CM) of synthetic fibre ropes (SFRs) used in offshore, maritime, and industrial settings demands more than a classifier: inspectors need continuous severity estimates, maintenance recommendations, anomaly flags,…

计算机视觉与模式识别 · 计算机科学 2026-05-07 Anju Rani , Daniel Ortiz-Arroyo , Petar Durdevic

The scalability of current language-image pre-training for 3D medical imaging, such as CT and MRI, is constrained by the need for radiologists to manually curate raw clinical studies. In this work, we pioneer pre-training directly on…

计算机视觉与模式识别 · 计算机科学 2026-02-20 Chenhui Zhao , Yiwei Lyu , Asadur Chowdury , Edward Harake , Akhil Kondepudi , Akshay Rao , Xinhai Hou , Honglak Lee , Todd Hollon

Training Long-Context Large Language Models (LLMs) is challenging, as hybrid training with long-context and short-context data often leads to workload imbalances. Existing works mainly use data packing to alleviate this issue, but fail to…

机器学习 · 计算机科学 2025-10-14 Yongqiang Yao , Jingru Tan , Kaihuan Liang , Feizhao Zhang , Jiahao Hu , Shuo Wu , Yazhe Niu , Ruihao Gong , Dahua Lin , Ningyi Xu

Navigating quadruped robots in unstructured 3D environments poses significant challenges, requiring goal-directed motion, effective exploration to escape from local minima, and posture adaptation to traverse narrow, height-constrained…

机器人学 · 计算机科学 2026-04-30 Jeil Jeong , Minsung Yoon , Seokryun Choi , Heechan Shin , Taegeun Yang , Sung-eui Yoon

This report provides an architecture-led analysis of two modern vision-language models (VLMs), Qwen2.5-VL-7B-Instruct and Llama-4-Scout-17B-16E-Instruct, and explains how their architectural properties map to a practical video-to-artifact…

计算机视觉与模式识别 · 计算机科学 2025-12-30 Thomson Tong , Diba Darooneh

Real-world datasets often exhibit class imbalance across multiple categories, manifesting as long-tailed distributions and few-shot scenarios. This is especially challenging in Class-Imbalanced Multi-Label Image Classification (CI-MLIC)…

计算机视觉与模式识别 · 计算机科学 2025-09-23 Sheng Huang , Jiexuan Yan , Beiyan Liu , Bo Liu , Richang Hong

Scarcity of labeled histopathology data limits the applicability of deep learning methods to under-profiled cancer types and labels. Transfer learning allows researchers to overcome the limitations of small datasets by pre-training machine…

计算机视觉与模式识别 · 计算机科学 2022-04-21 Anna Yeaton , Rahul G. Krishnan , Rebecca Mieloszyk , David Alvarez-Melis , Grace Huynh

Histopathology images; microscopy images of stained tissue biopsies contain fundamental prognostic information that forms the foundation of pathological analysis and diagnostic medicine. However, diagnostics from histopathology images…

计算机视觉与模式识别 · 计算机科学 2019-10-30 Aïcha BenTaieb , Ghassan Hamarneh

To achieve high-quality results, diffusion models must be trained on large datasets. This can be notably prohibitive for models in specialized domains, such as computational pathology. Conditioning on labeled data is known to help in…

计算机视觉与模式识别 · 计算机科学 2023-12-04 Srikar Yellapragada , Alexandros Graikos , Prateek Prasanna , Tahsin Kurc , Joel Saltz , Dimitris Samaras

Multimodal learning has shown promise in medical imaging, combining complementary modalities like images and text. Vision-language models (VLMs) capture rich diagnostic cues but often require large paired datasets and prompt- or text-based…

计算机视觉与模式识别 · 计算机科学 2025-11-26 Banafsheh Karimian , Giulia Avanzato , Soufian Belharbi , Alexis Guichemerre , Luke McCaffrey , Mohammadhadi Shateri , Eric Granger

Radiology report generation (RRG) has emerged as a promising approach to alleviate radiologists' workload and reduce human errors by automatically generating diagnostic reports from medical images. A key challenge in RRG is achieving…

计算机视觉与模式识别 · 计算机科学 2026-05-01 Yucheng Chen , Yang Yu , Yufei Shi , Conghao Xiong , Xulei Yang , Si Yong Yeo

Reproducibility remains a critical challenge in foundation model training for histopathology, often hindered by software randomness, hardware non-determinism, and inconsistent hyperparameter reporting. To investigate these issues, we…

计算机视觉与模式识别 · 计算机科学 2025-11-04 Usman Afzaal , Ziyu Su , Usama Sajjad , Hao Lu , Mostafa Rezapour , Metin Nafi Gurcan , Muhammad Khalid Khan Niazi

Current Large Language Models (LLMs) exhibit significant limitations, notably in structured, interpretable, and verifiable medical reasoning, alongside practical deployment challenges related to computational resources and data privacy.…

Cancer progression arises from interactions across multiple biological layers, especially beyond morphological and across molecular layers that remain invisible to image-only models. To capture this broader biological landscape, we present…

机器学习 · 计算机科学 2025-12-17 Juseung Yun , Sunwoo Yu , Sumin Ha , Jonghyun Kim , Janghyeon Lee , Jongseong Jang , Soonyoung Lee

Medical report interpretation plays a crucial role in healthcare, enabling both patient-facing explanations and effective information flow across clinical systems. While recent vision-language models (VLMs) and large language models (LLMs)…

计算机视觉与模式识别 · 计算机科学 2025-08-26 Fangxin Shang , Yuan Xia , Dalu Yang , Yahui Wang , Binglin Yang

Deep learning for histopathology has been successfully used for disease classification, image segmentation and more. However, combining image and text modalities using current state-of-the-art (SOTA) methods has been a challenge due to the…

计算机视觉与模式识别 · 计算机科学 2024-03-18 Saurav Sengupta , Donald E. Brown

As data requirements continue to grow, efficient learning increasingly depends on the curation and distillation of high-value data rather than brute-force scaling of model sizes. In the case of a hyperspectral image (HSI), the challenge is…

计算机视觉与模式识别 · 计算机科学 2025-09-30 Abhiroop Chatterjee , Susmita Ghosh

Topology is critical in tubular structures such as blood vessels, nerve fibers, and road networks, where connectivity and loop structure govern downstream functional analysis. Vision-Language Models (VLMs) are promising candidates for…

计算机视觉与模式识别 · 计算机科学 2026-05-14 Meilong Xu , Qingqiao Hu , Xiaoling Hu , Shahira Abousamra , Xin Yu , Weimin Lyu , Kehan Qi , Dimitris Samaras , Chao Chen

Deep learning for histopathology has been successfully used for disease classification, image segmentation and more. However, combining image and text modalities using current state-of-the-art methods has been a challenge due to the high…

计算机视觉与模式识别 · 计算机科学 2023-11-15 Saurav Sengupta , Donald E. Brown