English
Related papers

Related papers: BriMA: Bridged Modality Adaptation for Multi-Modal…

200 papers

Can performance on the task of action quality assessment (AQA) be improved by exploiting a description of the action and its quality? Current AQA and skills assessment approaches propose to learn features that serve only one task -…

Computer Vision and Pattern Recognition · Computer Science 2019-06-17 Paritosh Parmar , Brendan Tran Morris

Image Quality Assessment (IQA) models benefit significantly from semantic information, which allows them to treat different types of objects distinctly. Currently, leveraging semantic information to enhance IQA is a crucial research…

Computer Vision and Pattern Recognition · Computer Science 2024-05-21 Wensheng Pan , Timin Gao , Yan Zhang , Runze Hu , Xiawu Zheng , Enwei Zhang , Yuting Gao , Yutao Liu , Yunhang Shen , Ke Li , Shengchuan Zhang , Liujuan Cao , Rongrong Ji

Foundation models for vision have transformed visual recognition with powerful pretrained representations and strong zero-shot capabilities, yet their potential for data-efficient learning remains largely untapped. Active Learning (AL) aims…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Huy Hoang Nguyen , Cédric Jung , Shirin Salehi , Tobias Glück , Anke Schmeink , Andreas Kugi

Understanding human intentions (e.g., emotions) from videos has received considerable attention recently. Video streams generally constitute a blend of temporal data stemming from distinct modalities, including natural language, facial…

Computer Vision and Pattern Recognition · Computer Science 2024-10-01 Dingkang Yang , Mingcheng Li , Linhao Qu , Kun Yang , Peng Zhai , Song Wang , Lihua Zhang

Image Quality Assessment (IQA) remains an unresolved challenge in computer vision due to complex distortions, diverse image content, and limited data availability. Existing Blind IQA (BIQA) methods largely rely on extensive human…

Computer Vision and Pattern Recognition · Computer Science 2025-09-23 Xudong Li , Zihao Huang , Yan Zhang , Yunhang Shen , Ke Li , Xiawu Zheng , Liujuan Cao , Rongrong Ji

Action Quality Assessment(AQA) is important for action understanding and resolving the task poses unique challenges due to subtle visual differences. Existing state-of-the-art methods typically rely on the holistic video representations for…

Computer Vision and Pattern Recognition · Computer Science 2022-07-20 Yang Bai , Desen Zhou , Songyang Zhang , Jian Wang , Errui Ding , Yu Guan , Yang Long , Jingdong Wang

Leveraging complementary relationships across modalities has recently drawn a lot of attention in multimodal emotion recognition. Most of the existing approaches explored cross-attention to capture the complementary relationships across the…

Computer Vision and Pattern Recognition · Computer Science 2024-07-02 G Rajasekhar , Jahangir Alam

Multimodal fusion is susceptible to modality imbalance, where dominant modalities overshadow weak ones, easily leading to biased learning and suboptimal fusion, especially for incomplete modality conditions. To address this problem, we…

Machine Learning · Computer Science 2026-03-20 Xiang Shi , Rui Zhang , Jiawei Liu , Yinpeng Liu , Qikai Cheng , Wei Lu

Emotion recognition in real-world environments is hindered by partial occlusions, missing modalities, and severe class imbalance. To address these issues, particularly for the Affective Behavior Analysis in-the-wild (ABAW) Expression…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Jun Yu , Naixiang Zheng , Guoyuan Wang , Yunxiang Zhang , Lingsi Zhu , Jiaen Liang , Wei Huang , Shengping Liu

Detecting deception by human behaviors is vital in many fields such as custom security and multimedia anti-fraud. Recently, audio-visual deception detection attracts more attention due to its better performance than using only a single…

Computer Vision and Pattern Recognition · Computer Science 2023-02-14 Zhaoxu Li , Zitong Yu , Nithish Muthuchamy Selvaraj , Xiaobao Guo , Bingquan Shen , Adams Wai-Kin Kong , Alex Kot

Cross-modal alignment is crucial for multimodal representation fusion due to the inherent heterogeneity between modalities. While Transformer-based methods have shown promising results in modeling inter-modal relationships, their quadratic…

Computer Vision and Pattern Recognition · Computer Science 2024-12-03 Yan Li , Yifei Xing , Xiangyuan Lan , Xin Li , Haifeng Chen , Dongmei Jiang

Human trajectory forecasting requires capturing the multimodal nature of pedestrian behavior. However, existing approaches suffer from prior misalignment. Their learned or fixed priors often fail to capture the full distribution of…

Computer Vision and Pattern Recognition · Computer Science 2026-04-15 Chao Li , Rui Zhang , Siyuan Huang , Xian Zhong , Hongbo Jiang

Video quality assessment (VQA) has attracted growing attention in recent years. While the great expense of annotating large-scale VQA datasets has become the main obstacle for current deep-learning methods. To surmount the constraint of…

Computer Vision and Pattern Recognition · Computer Science 2023-08-03 Hongbo Liu , Mingda Wu , Kun Yuan , Ming Sun , Yansong Tang , Chuanchuan Zheng , Xing Wen , Xiu Li

As multimedia data flourishes on the Internet, quality assessment (QA) of multimedia data becomes paramount for digital media applications. Since multimedia data includes multiple modalities including audio, image, video, and audio-visual…

Image and Video Processing · Electrical Eng. & Systems 2024-07-30 Yuqin Cao , Xiongkuo Min , Yixuan Gao , Wei Sun , Weisi Lin , Guangtao Zhai

While VideoQA Transformer models demonstrate competitive performance on standard benchmarks, the reasons behind their success are not fully understood. Do these models capture the rich multimodal structures and dynamics from video and text…

Computer Vision and Pattern Recognition · Computer Science 2024-06-10 Ishaan Singh Rawal , Alexander Matyasko , Shantanu Jaiswal , Basura Fernando , Cheston Tan

In this paper, we reveal that most current efficient multimodal fine-tuning methods are hindered by a key limitation: they are directly borrowed from LLMs, often neglecting the intrinsic differences of multimodal scenarios and even…

Computer Vision and Pattern Recognition · Computer Science 2025-12-30 Yake Wei , Yu Miao , Dongzhan Zhou , Di Hu

Learning from multiple modalities, such as audio and video, offers opportunities for leveraging complementary information, enhancing robustness, and improving contextual understanding and performance. However, combining such modalities…

Multimedia · Computer Science 2024-10-15 Konstantinos Kontras , Christos Chatzichristos , Matthew Blaschko , Maarten De Vos

Full-Reference image quality assessment (FR IQA) is important for image compression, restoration and generative modeling, yet current neural metrics remain slow and vulnerable to adversarial perturbations. We present BiRQA, a compact FR IQA…

Computer Vision and Pattern Recognition · Computer Science 2026-02-25 Aleksandr Gushchin , Dmitriy S. Vatolin , Anastasia Antsiferova

Multi-modal learning, which focuses on utilizing various modalities to improve the performance of a model, is widely used in video recognition. While traditional multi-modal learning offers excellent recognition results, its computational…

Computer Vision and Pattern Recognition · Computer Science 2021-05-13 Rameswar Panda , Chun-Fu Chen , Quanfu Fan , Ximeng Sun , Kate Saenko , Aude Oliva , Rogerio Feris

The research in image quality assessment (IQA) has a long history, and significant progress has been made by leveraging recent advances in deep neural networks (DNNs). Despite high correlation numbers on existing IQA datasets, DNN-based…

Image and Video Processing · Electrical Eng. & Systems 2021-04-09 Zhihua Wang , Kede Ma