English
Related papers

Related papers: Multi-modal Learnable Queries for Image Aesthetics…

200 papers

The proliferation of machine learning (ML) has drawn unprecedented interest in the study of various multimedia contents such as text, image, audio and video, among others. Consequently, understanding and learning ML-based representations…

Multimedia · Computer Science 2023-05-02 Lei Gao , Ling Guan

One of the key goals of artificial intelligence (AI) is the development of a multimodal system that facilitates communication with the visual world (image and video) using a natural language query. Earlier works on medical question…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Deepak Gupta , Dina Demner-Fushman

The deep learning revolution has strongly impacted low-level image processing tasks such as style/domain transfer, enhancement/restoration, and visual quality assessments. Despite often being treated separately, the aforementioned tasks…

Image and Video Processing · Electrical Eng. & Systems 2025-08-26 Abhinau K. Venkataramanan , Cosmin Stejerean , Ioannis Katsavounidis , Hassene Tmar , Alan C. Bovik

Existing datasets for tabular question answering typically focus exclusively on text within cells. However, real-world data is inherently multimodal, often blending images such as symbols, faces, icons, patterns, and charts with textual…

In this paper we propose to learn a multimodal image and text embedding from Web and Social Media data, aiming to leverage the semantic knowledge learnt in the text domain and transfer it to a visual model for semantic image retrieval. We…

Computer Vision and Pattern Recognition · Computer Science 2018-08-21 Raul Gomez , Lluis Gomez , Jaume Gibert , Dimosthenis Karatzas

Recently, Visual Question Answering (VQA) has emerged as one of the most significant tasks in multimodal learning as it requires understanding both visual and textual modalities. Existing methods mainly rely on extracting image and question…

Computer Vision and Pattern Recognition · Computer Science 2018-07-23 Pan Lu , Lei Ji , Wei Zhang , Nan Duan , Ming Zhou , Jianyong Wang

Multimodal pre-training breaks down the modality barriers and allows the individual modalities to be mutually augmented with information, resulting in significant advances in representation learning. However, graph modality, as a very…

Multimedia · Computer Science 2022-11-01 Xuan Yang , Quanjin Tao , Xiao Feng , Donghong Cai , Xiang Ren , Yang Yang

Given a query photo issued by a user (q-user), the landmark retrieval is to return a set of photos with their landmarks similar to those of the query, while the existing studies on the landmark retrieval focus on exploiting geometries of…

Computer Vision and Pattern Recognition · Computer Science 2017-04-05 Yang Wang , Xuemin Lin , Lin Wu , Wenjie Zhang

The recent advancements in generative language models have demonstrated their ability to memorize knowledge from documents and recall knowledge to respond to user queries effectively. Building upon this capability, we propose to enable…

Multimedia · Computer Science 2024-02-19 Yongqi Li , Wenjie Wang , Leigang Qu , Liqiang Nie , Wenjie Li , Tat-Seng Chua

AI-based image enhancement techniques have been widely adopted in various visual applications, significantly improving the perceptual quality of user-generated content (UGC). However, the lack of specialized quality assessment models has…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Shushi Wang , Chunyi Li , Zicheng Zhang , Han Zhou , Wei Dong , Jun Chen , Guangtao Zhai , Xiaohong Liu

The recent emergence of Multi-modal Large Language Models (MLLMs) has introduced a new dimension to the Text-rich Image Understanding (TIU) field, with models demonstrating impressive and inspiring performance. However, their rapid…

Computer Vision and Pattern Recognition · Computer Science 2025-02-25 Pei Fu , Tongkun Guan , Zining Wang , Zhentao Guo , Chen Duan , Hao Sun , Boming Chen , Jiayao Ma , Qianyi Jiang , Kai Zhou , Junfeng Luo

There are two main lines of research on visual question answering (VQA): compositional model with explicit multi-hop reasoning, and monolithic network with implicit reasoning in the latent feature space. The former excels in…

Computer Vision and Pattern Recognition · Computer Science 2020-10-13 Ruixue Tang , Chao Ma

The aesthetic quality of an image is defined as the measure or appreciation of the beauty of an image. Aesthetics is inherently a subjective property but there are certain factors that influence it such as, the semantic content of the…

Computer Vision and Pattern Recognition · Computer Science 2022-08-25 Luigi Celona , Marco Leonardi , Paolo Napoletano , Alessandro Rozza

Multiple-choice questions (MCQs) are a widely used educational tool, particularly in domains such as visualization literacy that require broad conceptual coverage and support diverse real-world applications. However, designing high-quality…

Human-Computer Interaction · Computer Science 2026-03-03 Zixin Chen , Yuhang Zeng , Sicheng Song , Yanna Lin , Xian Xu , Huamin Qu , Meng Xia

Text-to-image retrieval is a fundamental task in vision-language learning, yet in real-world scenarios it is often challenged by short and underspecified user queries. Such queries are typically only one or two words long, rendering them…

Computer Vision and Pattern Recognition · Computer Science 2026-02-25 Jianglin Lu , Simon Jenni , Kushal Kafle , Jing Shi , Handong Zhao , Yun Fu

The highly abstract nature of image aesthetics perception (IAP) poses significant challenge for current multimodal large language models (MLLMs). The lack of human-annotated multi-modality aesthetic data further exacerbates this dilemma,…

Computer Vision and Pattern Recognition · Computer Science 2024-07-25 Yipo Huang , Xiangfei Sheng , Zhichao Yang , Quan Yuan , Zhichao Duan , Pengfei Chen , Leida Li , Weisi Lin , Guangming Shi

Pre-trained large multi-modal models (LMMs) exploit fine-tuning to adapt diverse user applications. Nevertheless, fine-tuning may face challenges due to deactivated sensors (e.g., cameras turned off for privacy or technical issues),…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Shu Zhao , Xiaohan Zou , Tan Yu , Huijuan Xu

Hateful meme detection is a new multimodal task that has gained significant traction in academic and industry research communities. Recently, researchers have applied pre-trained visual-linguistic models to perform the multimodal…

Computer Vision and Pattern Recognition · Computer Science 2022-04-07 Ming Shan Hee , Roy Ka-Wei Lee , Wen-Haw Chong

Multimodal Large Language Models (MLLMs) have shown impressive capabilities in jointly understanding text, images, and videos, often evaluated via Visual Question Answering (VQA). However, even state-of-the-art MLLMs struggle with…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Alberto Compagnoni , Marco Morini , Sara Sarto , Federico Cocchi , Davide Caffagni , Marcella Cornia , Lorenzo Baraldi , Rita Cucchiara

Multimodal Large Language Models (MLLMs) offer an opportunity to support multimedia learning through conversational systems grounded in educational content. However, while conversational AI is known to boost engagement, its impact on…

Human-Computer Interaction · Computer Science 2026-04-03 Karan Taneja , Anjali Singh , Ashok K. Goel
‹ Prev 1 8 9 10 Next ›