English
Related papers

Related papers: Unleashing Video Language Models for Fine-grained …

200 papers

Visual abductive reasoning (VAR) is a challenging task that requires AI systems to infer the most likely explanation for incomplete visual observations. While recent MLLMs develop strong general-purpose multimodal reasoning capabilities,…

Computer Vision and Pattern Recognition · Computer Science 2026-01-07 Boyu Chang , Qi Wang , Xi Guo , Zhixiong Nan , Yazhou Yao , Tianfei Zhou

Mammography report generation is a critical yet underexplored task in medical AI, characterized by challenges such as multiview image reasoning, high-resolution visual cues, and unstructured radiologic language. In this work, we introduce…

Image and Video Processing · Electrical Eng. & Systems 2025-08-14 Nak-Jun Sung , Donghyun Lee , Bo Hwa Choi , Chae Jung Park

Automated interpretation of CT images-particularly localizing and describing abnormal findings across multi-plane and whole-body scans-remains a significant challenge in clinical radiology. This work aims to address this challenge through…

Image and Video Processing · Electrical Eng. & Systems 2025-11-18 Ziheng Zhao , Lisong Dai , Ya Zhang , Yanfeng Wang , Weidi Xie

Vision-Language Models (VLMs) increasingly power high-stakes applications, from medical imaging to autonomous systems, yet they routinely hallucinate, confidently describing content not present in the input. We investigate the root causes…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 Harshvardhan Saini , Samyak Jha , Yiming Tang , Dianbo Liu

The widespread use of chest X-rays (CXRs), coupled with a shortage of radiologists, has driven growing interest in automated CXR analysis and AI-assisted reporting. While existing vision-language models (VLMs) show promise in specific tasks…

Due to the advantages of hypergraphs in modeling high-order relationships in complex systems, they have been applied to higher-order clustering, hypergraph neural networks and computer vision. These applications rely heavily on access to…

Social and Information Networks · Computer Science 2025-10-15 Bingqiao Gu , Jiale Zeng , Xingqin Qi , Dong Li

This research introduces a transformative framework for integrating Vision-Enhanced Large Language Models (LLMs) with advanced transformer-based architectures to tackle challenges in high-resolution image synthesis and multimodal data…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Karthikeya KV

Evaluating AI-generated medical image segmentations for clinical acceptability poses a significant challenge, as traditional pixelagreement metrics often fail to capture true diagnostic utility. This paper introduces Hierarchical Clinical…

Image and Video Processing · Electrical Eng. & Systems 2025-06-03 Jiaxi Sheng , Leyi Yu , Haoyue Li , Yifan Gao , Xin Gao

Multimodal large language models (MLLMs) have achieved significant success in the general field of image processing. Their emerging task generalization and freeform conversational capabilities can greatly facilitate medical diagnostic…

Image and Video Processing · Electrical Eng. & Systems 2024-09-17 Youzhu Jin , Yichen Zhang

Vision--language models (VLMs) for radiology report generation (RRG) can produce long-form chest CT reports from volumetric scans and show strong potential to improve radiology workflow efficiency and consistency. However, existing methods…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Chenyu Wang , Weicheng Dai , Han Liu , Wenchao Li , Kayhan Batmanghelich

Recent advances in reasoning models have shown remarkable progress in text-based domains, but transferring those capabilities to multimodal settings, e.g., to allow reasoning over audio-visual data, still remains a challenge, in part…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Edson Araujo , Saurabhchand Bhati , M. Jehanzeb Mirza , Brian Kingsbury , Samuel Thomas , Rogerio Feris , James R. Glass , Hilde Kuehne

The recent release of RadGenome-Chest CT has significantly advanced CT-based report generation. However, existing methods primarily focus on global features, making it challenging to capture region-specific details, which may cause certain…

Image and Video Processing · Electrical Eng. & Systems 2025-07-01 Sunggu Kyung , Jinyoung Seo , Hyunseok Lim , Dongyeong Kim , Hyungbin Park , Jimin Sung , Jihyun Kim , Wooyoung Jo , Yoojin Nam , Namkug Kim

Video generation has witnessed remarkable progress with the advent of deep generative models, particularly diffusion models. While existing methods excel in generating high-quality videos from text prompts or single images, personalized…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Yufan Deng , Xun Guo , Yizhi Wang , Jacob Zhiyuan Fang , Angtian Wang , Shenghai Yuan , Yiding Yang , Bo Liu , Haibin Huang , Chongyang Ma

Rapid progress in video models has largely focused on visual quality, leaving their reasoning capabilities underexplored. Video reasoning grounds intelligence in spatiotemporally consistent visual environments that go beyond what text can…

Recent studies have demonstrated the effectiveness of Large Language Models (LLMs) as reasoning modules that can deconstruct complex tasks into more manageable sub-tasks, particularly when applied to visual reasoning tasks for images. In…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Ahmad Mahmood , Ashmal Vayani , Muzammal Naseer , Salman Khan , Fahad Shahbaz Khan

While large language models (LLMs) have advanced CT report generation, existing methods typically encode 3D volumes holistically, failing to distinguish informative cues from redundant anatomical background. Inspired by radiological…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Yuhe Tian , Kun Zhang , Haoran Ma , Rui Yan , Yingtai Li , Rongsheng Wang , Shaohua Kevin Zhou

Automated radiology report generation using vision-language models (VLMs) is limited by the risk of prior-comparison hallucination, where the model generates historical findings unsupported by the current study. We address this challenge…

Computer Vision and Pattern Recognition · Computer Science 2026-03-02 Ao Li , Rui Liu , Mingjie Li , Sheng Liu , Lei Wang , Xiaodan Liang , Lina Yao , Xiaojun Chang , Lei Xing

Automatic generation of radiology reports seeks to reduce clinician workload while improving documentation consistency. Existing methods that adopt encoder-decoder or retrieval-augmented pipelines achieve progress in fluency but remain…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Rong Fu , Yiqing Lyu , Chunlei Meng , Muge Qi , Yabin Jin , Qi Zhao , Li Bao , Juntao Gao , Fuqian Shi , Nilanjan Dey , Wei Luo , Simon Fong

Hallucination poses a challenge to the deployment of large vision-language models (LVLMs) in applications. Unlike in large language models (LLMs), hallucination in LVLMs often arises from misalignments between visual inputs and textual…

Computer Vision and Pattern Recognition · Computer Science 2024-10-23 Sheng Liu , Haotian Ye , Lei Xing , James Zou

Video Large Language Models (Video LLMs) have recently exhibited remarkable capabilities in general video understanding. However, they mainly focus on holistic comprehension and struggle with capturing fine-grained spatial and temporal…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Yuqian Yuan , Hang Zhang , Wentong Li , Zesen Cheng , Boqiang Zhang , Long Li , Xin Li , Deli Zhao , Wenqiao Zhang , Yueting Zhuang , Jianke Zhu , Lidong Bing