中文
相关论文

相关论文: Unleashing Video Language Models for Fine-grained …

200 篇论文

Radiology report generation represents a significant application within medical AI, and has achieved impressive results. Concurrently, large language models (LLMs) have demonstrated remarkable performance across various domains. However,…

计算机视觉与模式识别 · 计算机科学 2025-07-08 Haifeng Zhao , Yufei Zhang , Leilei Ma , Shuo Xu , Dengdi Sun

Hallucinations in Large Language Models (LLMs), defined as the generation of content inconsistent with facts or context, represent a core obstacle to their reliable deployment in critical domains. Current research primarily focuses on…

计算与语言 · 计算机科学 2026-03-20 Yanyi Liu , Qingwen Yang , Tiezheng Guo , Feiyu Qu , Jun Liu , Yingyou Wen

Vision-language modeling (VLM) aims to bridge the information gap between images and natural language. Under the new paradigm of first pre-training on massive image-text pairs and then fine-tuning on task-specific data, VLM in the remote…

计算机视觉与模式识别 · 计算机科学 2025-06-11 Xingxing Weng , Chao Pang , Gui-Song Xia

Reasoning segmentation aims to segment target objects in complex scenes based on human intent and spatial reasoning. While recent multimodal large language models (MLLMs) have demonstrated impressive 2D image reasoning segmentation,…

计算机视觉与模式识别 · 计算机科学 2025-11-11 Jiaxin Huang , Runnan Chen , Ziwen Li , Zhengqing Gao , Xiao He , Yandong Guo , Mingming Gong , Tongliang Liu

Multidimensional human understanding is essential for real-world applications such as film analysis and virtual digital humans, yet current LVLM benchmarks largely focus on single-task settings and lack fine-grained, human-centric…

计算机视觉与模式识别 · 计算机科学 2026-05-06 Kangkang Wang , Qinting Jiang , Wanping Zhang , Bowen Ren , Shengzhao Wen

Large Vision-Language Models (LVLMs) have advanced considerably, intertwining visual recognition and language understanding to generate content that is not only coherent but also contextually attuned. Despite their success, LVLMs still…

计算机视觉与模式识别 · 计算机科学 2023-11-29 Sicong Leng , Hang Zhang , Guanzheng Chen , Xin Li , Shijian Lu , Chunyan Miao , Lidong Bing

Medical AI systems face two fundamental limitations. First, conventional vision-language models (VLMs) perform single-pass inference, yielding black-box predictions that cannot be audited or explained in clinical terms. Second, iterative…

计算机视觉与模式识别 · 计算机科学 2026-03-26 Nicanor Mayumu , Zeenath Khan , Melodena Stephens , Patrick Mukala , Farhad Oroumchian

Automated radiology report generation from 3D CT volumes often suffers from incomplete pathology coverage. We provide empirical evidence that this limitation stems from a representational bottleneck: contrastive 3D CT embeddings encode…

计算机视觉与模式识别 · 计算机科学 2026-03-18 Renjie Liang , Yiling Ma , Yang Xing , Zhengkang Fan , Jinqian Pan , Chengkun Sun , Li Li , Kuang Gong , Jie Xu

Video Large Language Models (VideoLLMs) have recently demonstrated remarkable progress in general video understanding. However, existing models primarily focus on high-level comprehension and are limited to text-only responses, restricting…

计算机视觉与模式识别 · 计算机科学 2025-07-29 Haochen Wang , Qirui Chen , Cilin Yan , Jiayin Cai , Xiaolong Jiang , Yao Hu , Weidi Xie , Stratis Gavves

Despite significant progress in video-language modeling, hallucinations remain a persistent challenge in Video Large Language Models (Vid-LLMs), referring to outputs that appear plausible yet contradict the content of the input video. This…

计算机视觉与模式识别 · 计算机科学 2026-04-15 Yiyang Huang , Yitian Zhang , Yizhou Wang , Mingyuan Zhang , Liang Shi , Huimin Zeng , Yun Fu

While large audio-language models (LALMs) have demonstrated state-of-the-art audio understanding, their reasoning capability in complex soundscapes still falls behind large vision-language models (LVLMs). Compared to the visual domain, one…

声音 · 计算机科学 2025-09-22 Qiaolin Wang , Xilin Jiang , Linyang He , Junkai Wu , Nima Mesgarani

Prompting has emerged as a practical way to adapt frozen vision-language models (VLMs) for video anomaly detection (VAD). Yet, existing prompts are often overly abstract, overlooking the fine-grained human-object interactions or action…

计算机视觉与模式识别 · 计算机科学 2025-10-03 Shu Zou , Xinyu Tian , Lukas Wesemann , Fabian Waschkowski , Zhaoyuan Yang , Jing Zhang

Transformer architectures have achieved state-of-the-art performance across natural language tasks, yet they fundamentally misrepresent the hierarchical nature of human language by processing text as flat token sequences. This results in…

计算与语言 · 计算机科学 2025-09-26 Ayan Sar , Sampurna Roy , Kanav Gupta , Anurag Kaushish , Tanupriya Choudhury , Abhijit Kumar

Vision-language models (VLMs) have demonstrated exceptional generalization capabilities for downstream tasks. Due to its efficiency, prompt learning has gradually become a more effective and efficient method for transferring VLMs to…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Chenhao Ding , Xinyuan Gao , Songlin Dong , Jizhou Han , Qiang Wang , Zhengdong Zhou , Yuhang He , Yihong Gong

Despite significant advancements, large multimodal models (LMMs) still struggle to bridge the gap between low-level visual perception -- focusing on shapes, sizes, and layouts -- and high-level language reasoning, such as semantics and…

计算与语言 · 计算机科学 2025-06-13 Zhenhailong Wang , Joy Hsu , Xingyao Wang , Kuan-Hao Huang , Manling Li , Jiajun Wu , Heng Ji

Vision Language Models (VLMs) have achieved impressive progress in multimodal reasoning; yet, they remain vulnerable to hallucinations, where outputs are not grounded in visual evidence. In this paper, we investigate a previously overlooked…

计算机视觉与模式识别 · 计算机科学 2025-10-15 Sifan Li , Hongkai Chen , Yujun Cai , Qingwen Ye , Liyang Chen , Junsong Yuan , Yiwei Wang

Vision-language models (VLMs) have recently emerged as a promising paradigm for video anomaly detection (VAD) due to their strong visual reasoning ability and natural language-based explainability. In this paper, we aim to address a key…

计算机视觉与模式识别 · 计算机科学 2026-05-15 Mitchell Piehl , Muchao Ye

General-purpose large Vision-Language Models (VLMs) demonstrate strong capabilities in generating detailed descriptions for natural images. However, their performance in the medical domain remains suboptimal, even for relatively…

计算机视觉与模式识别 · 计算机科学 2025-10-23 Yifan Li , Fenghe Tang , Yingtai Li , Shaohua Kevin Zhou

Multimodal large language models (MLLMs) exhibit strong visual-language reasoning, yet cannot process structured, non-visual data such as human skeletons. Existing methods either compress skeleton dynamics into lossy feature vectors for…

计算机视觉与模式识别 · 计算机科学 2026-05-22 Ziyi Wang , Peiming Li , Xinshun Wang , Yang Tang , Kai-Kuang Ma , Mengyuan Liu

Existing Video Large Language Models (Video LLMs) struggle with complex video understanding, exhibiting limited reasoning capabilities and potential hallucinations. In particular, these methods tend to perform reasoning solely relying on…

计算机视觉与模式识别 · 计算机科学 2026-04-23 Qizhong Tan , Zhuotao Tian , Guangming Lu , Jun Yu , Wenjie Pei