English
Related papers

Related papers: A High-Quality Dataset and Reliable Evaluation for…

200 papers

The creation of high-quality human-labeled image-caption datasets presents a significant bottleneck in the development of Visual-Language Models (VLMs). In this work, we investigate an approach that leverages the strengths of Large Language…

Computer Vision and Pattern Recognition · Computer Science 2024-06-10 Sahand Sharifzadeh , Christos Kaplanis , Shreya Pathak , Dharshan Kumaran , Anastasija Ilic , Jovana Mitrovic , Charles Blundell , Andrea Banino

Based on powerful Large Language Models (LLMs), recent generative Multimodal Large Language Models (MLLMs) have gained prominence as a pivotal research area, exhibiting remarkable capability for both comprehension and generation. In this…

Computation and Language · Computer Science 2023-08-03 Bohao Li , Rui Wang , Guangzhi Wang , Yuying Ge , Yixiao Ge , Ying Shan

General-domain large multimodal models (LMMs) have achieved significant advances in various image-text tasks. However, their performance in the Intelligent Traffic Surveillance (ITS) domain remains limited due to the absence of dedicated…

Computer Vision and Pattern Recognition · Computer Science 2025-09-15 Kaikai Zhao , Zhaoxiang Liu , Peng Wang , Xin Wang , Zhicheng Ma , Yajun Xu , Wenjing Zhang , Yibing Nan , Kai Wang , Shiguo Lian

Multimodal Large Language Model (MLLM) relies on the powerful LLM to perform multimodal tasks, showing amazing emergent abilities in recent studies, such as writing poems based on an image. However, it is difficult for these case studies to…

Computer Vision and Pattern Recognition · Computer Science 2025-10-27 Chaoyou Fu , Peixian Chen , Yunhang Shen , Yulei Qin , Mengdan Zhang , Xu Lin , Jinrui Yang , Xiawu Zheng , Ke Li , Xing Sun , Yunsheng Wu , Rongrong Ji , Caifeng Shan , Ran He

Comprehending text-rich visual content is paramount for the practical application of Multimodal Large Language Models (MLLMs), since text-rich scenarios are ubiquitous in the real world, which are characterized by the presence of extensive…

Computer Vision and Pattern Recognition · Computer Science 2024-04-26 Bohao Li , Yuying Ge , Yi Chen , Yixiao Ge , Ruimao Zhang , Ying Shan

Multimodal Large Language Models (MLLMs) have facilitated Multimodal Summarization with Multimodal Output (MSMO), wherein systems generate concise textual summaries accompanied by salient visuals from multimodal sources. However, current…

Artificial Intelligence · Computer Science 2026-05-13 Abid Ali , Diego Molla-Aliod , Usman Naseem

We present Meissonic, which elevates non-autoregressive masked image modeling (MIM) text-to-image to a level comparable with state-of-the-art diffusion models like SDXL. By incorporating a comprehensive suite of architectural innovations,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Jinbin Bai , Tian Ye , Wei Chow , Enxin Song , Xiangtai Li , Zhen Dong , Lei Zhu , Shuicheng Yan

In this paper, we address a fundamental gap between pre-training and fine-tuning of deep neural networks: while pre-training has shifted from unimodal to multimodal learning with enhanced visual understanding, fine-tuning predominantly…

Computer Vision and Pattern Recognition · Computer Science 2026-01-30 Shohei Enomoto , Shin'ya Yamaguchi

This paper introduces VideoMind, a video-centric omni-modal dataset designed for deep video content cognition and enhanced multi-modal feature representation. The dataset comprises 103K video samples (3K reserved for testing), each paired…

Computer Vision and Pattern Recognition · Computer Science 2025-07-25 Baoyao Yang , Wanyun Li , Dixin Chen , Junxiang Chen , Wenbin Yao , Haifeng Lin

Despite the effectiveness of vision-language supervised fine-tuning in enhancing the performance of Vision Large Language Models (VLLMs). However, existing visual instruction tuning datasets include the following limitations: (1)…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Yangzhou Liu , Yue Cao , Zhangwei Gao , Weiyun Wang , Zhe Chen , Wenhai Wang , Hao Tian , Lewei Lu , Xizhou Zhu , Tong Lu , Yu Qiao , Jifeng Dai

Reward models (RMs) are essential for training large language models (LLMs), but remain underexplored for omni models that handle interleaved image and text sequences. We introduce Multimodal RewardBench 2 (MMRB2), the first comprehensive…

Computation and Language · Computer Science 2026-01-21 Yushi Hu , Reyhane Askari-Hemmat , Melissa Hall , Emily Dinan , Luke Zettlemoyer , Marjan Ghazvininejad

Automatically evaluating the quality of language generation is critical. Although recent learned metrics show high correlation with human judgement, these metrics can not explain their verdict or associate the scores with defects in…

Computation and Language · Computer Science 2023-10-30 Wenda Xu , Danqing Wang , Liangming Pan , Zhenqiao Song , Markus Freitag , William Yang Wang , Lei Li

Thanks to the emerging of foundation models, the large language and vision models are integrated to acquire the multimodal ability of visual captioning, question answering, etc. Although existing multimodal models present impressive…

Computer Vision and Pattern Recognition · Computer Science 2023-12-29 Bo Zhao , Boya Wu , Muyang He , Tiejun Huang

Despite the promising performance of Large Vision Language Models (LVLMs) in visual understanding, they occasionally generate incorrect outputs. While reward models (RMs) with reinforcement learning or test-time scaling offer the potential…

Computer Vision and Pattern Recognition · Computer Science 2025-05-21 Yuhang Zang , Xiaoyi Dong , Pan Zhang , Yuhang Cao , Ziyu Liu , Shengyuan Ding , Shenxi Wu , Yubo Ma , Haodong Duan , Wenwei Zhang , Kai Chen , Dahua Lin , Jiaqi Wang

While text-to-image models like DALLE-3 and Stable Diffusion are rapidly proliferating, they often encounter challenges such as hallucination, bias, and the production of unsafe, low-quality output. To effectively address these issues, it…

The rapid advancement of educational applications, artistic creation, and AI-generated content (AIGC) technologies has substantially increased practical requirements for comprehensive Image Aesthetics Assessment (IAA), particularly…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Shuo Cao , Nan Ma , Jiayang Li , Xiaohui Li , Lihao Shao , Kaiwen Zhu , Yu Zhou , Yuandong Pu , Jiarui Wu , Jiaquan Wang , Bo Qu , Wenhai Wang , Yu Qiao , Dajuin Yao , Yihao Liu

With the growing demands of AI-generated content (AIGC), the need for high-quality, diverse, and scalable data has become increasingly crucial. However, collecting large-scale real-world data remains costly and time-consuming, hindering the…

Computer Vision and Pattern Recognition · Computer Science 2025-08-08 Kunyu Feng , Yue Ma , Xinhua Zhang , Boshi Liu , Yikuang Yuluo , Yinhan Zhang , Runtao Liu , Hongyu Liu , Zhiyuan Qin , Shanhui Mo , Qifeng Chen , Zeyu Wang

In this paper, a self-guiding multimodal LSTM (sg-LSTM) image captioning model is proposed to handle uncontrolled imbalanced real-world image-sentence dataset. We collect FlickrNYC dataset from Flickr as our testbed with 306,165 images and…

Computer Vision and Pattern Recognition · Computer Science 2017-09-18 Yang Xian , Yingli Tian

Deep learning algorithms require extensive data to achieve robust performance. However, data availability is often restricted in the medical domain due to patient privacy concerns. Synthetic data presents a possible solution to these…

Recent advancements in large language models (LLMs) with extended context windows have significantly improved tasks such as information extraction, question answering, and complex planning scenarios. In order to achieve success in long…

Computation and Language · Computer Science 2025-05-20 Zhi Chen , Qiguang Chen , Libo Qin , Qipeng Guo , Haijun Lv , Yicheng Zou , Wanxiang Che , Hang Yan , Kai Chen , Dahua Lin
‹ Prev 1 4 5 6 7 8 10 Next ›