English
Related papers

Related papers: PUMGPT: A Large Vision-Language Model for Product …

200 papers

Large Vision-Language Models (LVLMs) show significant strides in general-purpose multimodal applications such as visual dialogue and embodied navigation. However, existing multimodal evaluation benchmarks cover a limited number of…

Goal-oriented script planning, or the ability to devise coherent sequences of actions toward specific goals, is commonly employed by humans to plan for typical activities. In e-commerce, customers increasingly seek LLM-based assistants to…

The emergence of multimodal large language models (MLLMs) advances multimodal emotion recognition (MER) to the next level, from naive discriminative tasks to complex emotion understanding with advanced video understanding abilities and…

Human-Computer Interaction · Computer Science 2025-05-08 Zheng Lian , Haoyu Chen , Lan Chen , Haiyang Sun , Licai Sun , Yong Ren , Zebang Cheng , Bin Liu , Rui Liu , Xiaojiang Peng , Jiangyan Yi , Jianhua Tao

With the boom of e-commerce and web applications, recommender systems have become an important part of our daily lives, providing personalized recommendations based on the user's preferences. Although deep neural networks (DNNs) have made…

Artificial Intelligence · Computer Science 2024-03-13 Xiaonan Xu , Yichao Wu , Penghao Liang , Yuhang He , Han Wang

Large Vision Language Models exhibit remarkable capabilities but struggle with hallucinations inconsistencies between images and their descriptions. Previous hallucination evaluation studies on LVLMs have identified hallucinations in terms…

Artificial Intelligence · Computer Science 2024-11-11 Chaoya Jiang , Hongrui Jia , Wei Ye , Mengfan Dong , Haiyang Xu , Ming Yan , Ji Zhang , Shikun Zhang

In this work, we present Valley3, an omni multimodal large language model (MLLM) developed for diverse global e-commerce tasks, with unified understanding and reasoning capabilities across text, images, video, and audio. A key feature of…

Artificial Intelligence · Computer Science 2026-05-07 Zeyu Chen , Guanghao Zhou , Qixiang Yin , Ziwang Zhao , Huanjin Yao , Pengjiu Xia , Min Yang , Cen Chen , Minghui Qiu

Vision-Language Models (VLMs) building upon the foundation of powerful large language models have made rapid progress in reasoning across visual and textual data. While VLMs perform well on vision tasks that they are trained on, our results…

Computer Vision and Pattern Recognition · Computer Science 2025-06-26 Zixuan Wu , Yoolim Kim , Carolyn Jane Anderson

Reliable interpretation of multimodal data in dentistry is essential for automated oral healthcare, yet current multimodal large language models (MLLMs) struggle to capture fine-grained dental visual details and lack sufficient reasoning…

Modern large language models become multimodal, analyzing various data formats like text and images. While fine-tuning is effective for adapting these multimodal language models (MLMs) to downstream tasks, full fine-tuning is…

Computation and Language · Computer Science 2025-12-01 Alexander Sergeev , Evgeny Kotelnikov

Vision-Language Models (VLMs) are becoming increasingly popular in the medical domain, bridging the gap between medical images and clinical language. Existing VLMs demonstrate an impressive ability to comprehend medical images and text…

Computer Vision and Pattern Recognition · Computer Science 2025-06-24 Bidur Khanal , Sandesh Pokhrel , Sanjay Bhandari , Ramesh Rana , Nikesh Shrestha , Ram Bahadur Gurung , Cristian Linte , Angus Watson , Yash Raj Shrestha , Binod Bhattarai

MLLMs (Multimodal Large Language Models) have showcased remarkable capabilities, but their performance in high-stakes, domain-specific scenarios like surgical settings, remains largely under-explored. To address this gap, we develop…

Computer Vision and Pattern Recognition · Computer Science 2025-10-03 Gui Wang , Yang Wennuo , Xusen Ma , Zehao Zhong , Zhuoru Wu , Ende Wu , Rong Qu , Wooi Ping Cheah , Jianfeng Ren , Linlin Shen

Large Vision-Language Models (LVLMs) have shown impressive capabilities across a range of tasks that integrate visual and textual understanding, such as image captioning and visual question answering. These models are trained on large-scale…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Xiaomei Zhang , Hanyu Zheng , Xiangyu Zhu , Jinghuan Wei , Junhong Zou , Zhen Lei , Zhaoxiang Zhang

Molecule discovery plays a crucial role in various scientific fields, advancing the design of tailored materials and drugs. However, most of the existing methods heavily rely on domain experts, require excessive computational cost, or…

Computation and Language · Computer Science 2024-05-10 Jiatong Li , Yunqing Liu , Wenqi Fan , Xiao-Yong Wei , Hui Liu , Jiliang Tang , Qing Li

Evaluating production-level retrieval systems at scale is a crucial yet challenging task due to the limited availability of a large pool of well-trained human annotators. Large Language Models (LLMs) have the potential to address this…

Information Retrieval · Computer Science 2024-09-19 Kasra Hosseini , Thomas Kober , Josip Krapac , Roland Vollgraf , Weiwei Cheng , Ana Peleteiro Ramallo

Multi-modal large language models (MLLMs) have demonstrated remarkable success in vision and visual-language tasks within the natural image domain. Owing to the significant diversities between the natural and remote sensing (RS) images, the…

Computer Vision and Pattern Recognition · Computer Science 2024-03-11 Wei Zhang , Miaoxin Cai , Tong Zhang , Yin Zhuang , Xuerui Mao

Large Language Models (LLMs) have allowed recent LLM-based approaches to achieve excellent performance on long-video understanding benchmarks. We investigate how extensive world knowledge and strong reasoning skills of underlying LLMs…

Computer Vision and Pattern Recognition · Computer Science 2025-06-12 Kanchana Ranasinghe , Xiang Li , Kumara Kahatapitiya , Michael S. Ryoo

The growth of social media, characterized by its multimodal nature, has led to the emergence of diverse phenomena and challenges, which calls for an effective approach to uniformly solve automated tasks. The powerful Large Vision Language…

Computation and Language · Computer Science 2024-10-11 Xinnong Zhang , Haoyu Kuang , Xinyi Mou , Hanjia Lyu , Kun Wu , Siming Chen , Jiebo Luo , Xuanjing Huang , Zhongyu Wei

Large models such as Large Language Models (LLMs) and Vision Language Models (VLMs) have transformed artificial intelligence, powering applications in natural language processing, computer vision, and multimodal learning. However, fully…

Large language models (LLMs) have recently been extended to the vision-language realm, obtaining impressive general multi-modal capabilities. However, the exploration of multi-modal large language models (MLLMs) for remote sensing (RS) data…

Computer Vision and Pattern Recognition · Computer Science 2024-01-19 Yang Zhan , Zhitong Xiong , Yuan Yuan

Large Language Models (LLMs) have demonstrated remarkable success as general-purpose task solvers across various fields. However, their capabilities remain limited when addressing domain-specific problems, particularly in downstream NLP…

Computation and Language · Computer Science 2025-02-28 Mohamed Bayan Kmainasi , Ali Ezzat Shahroor , Maram Hasanain , Sahinur Rahman Laskar , Naeemul Hassan , Firoj Alam