中文
相关论文

相关论文: PIDNet: Progressive Implicit Decouple Network for …

200 篇论文

Pre-trained video large language models excel at visual reasoning. However, they struggle when videos arrive with auxiliary streams, such as audio, depth map, or dense temporal evidence. In such a scenario, uniform fusion induces modality…

计算机视觉与模式识别 · 计算机科学 2026-05-27 Bonan Ding , Umair Nawaz , Ufaq Khan , Abdelrahman M. Shaker , Muhammad Haris Khan , Jiale Cao , Jin Xie , Fahad Shahbaz Khan

With the fast development of artificial intelligence and short videos, emotion recognition in short videos has become one of the most important research topics in human-computer interaction. At present, most emotion recognition methods…

计算机视觉与模式识别 · 计算机科学 2024-12-10 Xuecheng Wu , Mengmeng Tian , Lanhang Zhai

Existing RGBT tracking methods often design various interaction models to perform cross-modal fusion of each layer, but can not execute the feature interactions among all layers, which plays a critical role in robust multimodal…

计算机视觉与模式识别 · 计算机科学 2024-12-31 Andong Lu , Wanyu Wang , Chenglong Li , Jin Tang , Bin Luo

Full-Reference image quality assessment (FR IQA) is important for image compression, restoration and generative modeling, yet current neural metrics remain slow and vulnerable to adversarial perturbations. We present BiRQA, a compact FR IQA…

计算机视觉与模式识别 · 计算机科学 2026-02-25 Aleksandr Gushchin , Dmitriy S. Vatolin , Anastasia Antsiferova

Diffusion-based models have recently revolutionized image generation, achieving unprecedented levels of fidelity. However, consistent generation of high-quality images remains challenging partly due to the lack of conditioning mechanisms…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Khaled Abud , Sergey Lavrushkin , Alexey Kirillov , Dmitriy Vatolin

Accurate air quality prediction is essential for public health, environmental monitoring, and industrial safety. However, most existing approaches rely on centralized learning paradigms, which introduce challenges related to scalability,…

机器学习 · 计算机科学 2026-05-19 Manjil Nepal , Kimsie Phan , Tamoghna Ojha , Aritra Dutta , M Krishna Siva Prasad

Dance improvisation is an active research topic in the arts. Motion analysis of improvised dance can be challenging due to its unique dynamics. Data-driven dance motion analysis, including recognition and generation, is often limited to…

计算机视觉与模式识别 · 计算机科学 2023-10-11 Jia Fu , Jiarui Tan , Wenjie Yin , Sepideh Pashami , Mårten Björkman

Audio question answering (AQA), acting as a widely used proxy task to explore scene understanding, has got more attention. The AQA is challenging for it requires comprehensive temporal reasoning from different scales' events of an audio…

声音 · 计算机科学 2023-05-30 Guangyao Li , Yixin Xu , Di Hu

Mortgage risk assessment traditionally relies on structured financial data, which is often proprietary, confidential, and costly. In this study, we propose a novel multimodal deep learning framework that uses cost-free, publicly available,…

计算工程、金融与科学 · 计算机科学 2025-10-28 Mahsa Tavakoli , Rohitash Chandra , Cristian Bravo

The rapid development of diffusion models has greatly advanced AI-generated videos in terms of length and consistency recently, yet assessing AI-generated videos still remains challenging. Previous approaches have often focused on…

计算机视觉与模式识别 · 计算机科学 2025-01-07 Jiaze Li , Haoran Xu , Shiding Zhu , Junwei He , Haozhao Wang

Image quality assessment (IQA) is an important element of a broad spectrum of applications ranging from automatic video streaming to display technology. Furthermore, the measurement of image quality requires a balanced investigation of…

计算机视觉与模式识别 · 计算机科学 2020-12-02 Domonkos Varga

Multi-modality image fusion (MMIF) aims to integrate complementary information from different modalities into a single fused image to represent the imaging scene and facilitate downstream visual tasks comprehensively. In recent years,…

计算机视觉与模式识别 · 计算机科学 2024-04-15 Zhe Li , Haiwei Pan , Kejia Zhang , Yuhua Wang , Fengming Yu

BIQA (Blind Image Quality Assessment) is an important field of study that evaluates images automatically. Although significant progress has been made, blind image quality assessment remains a difficult task since images vary in content and…

计算机视觉与模式识别 · 计算机科学 2024-10-10 Muhammad Azeem Aslam , Xu Wei , Hassan Khalid , Nisar Ahmed , Zhu Shuangtong , Xin Liu , Yimei Xu

Deepfakes are AI-synthesized multimedia data that may be abused for spreading misinformation. Deepfake generation involves both visual and audio manipulation. To detect audio-visual deepfakes, previous studies commonly employ two relatively…

声音 · 计算机科学 2025-06-10 Kuiyuan Zhang , Wenjie Pei , Rushi Lan , Yifang Guo , Zhongyun Hua

Integration of multimodal information from various sources has been shown to boost the performance of machine learning models and thus has received increased attention in recent years. Often such models use deep modality-specific networks…

机器学习 · 计算机科学 2022-11-22 Shiv Shankar , Laure Thompson , Madalina Fiterau

Technological advances in medical data collection, such as high-throughput genomic sequencing and digital high-resolution histopathology, have contributed to the rising requirement for multimodal biomedical modelling, specifically for…

机器学习 · 计算机科学 2024-10-29 Konstantin Hemker , Nikola Simidjievski , Mateja Jamnik

Image classification models often demonstrate unstable performance in real-world applications due to variations in image information, driven by differing visual perspectives of subject objects and lighting discrepancies. To mitigate these…

计算机视觉与模式识别 · 计算机科学 2024-07-29 Yuze Zheng , Zixuan Li , Xiangxian Li , Jinxing Liu , Yuqing Wang , Xiangxu Meng , Lei Meng

Multi-stage architectures have exhibited efficacy in image dehazing, which usually decomposes a challenging task into multiple more tractable sub-tasks and progressively estimates latent hazy-free images. Despite the remarkable progress,…

计算机视觉与模式识别 · 计算机科学 2023-08-15 Hao Shen , Zhong-Qiu Zhao , Yulun Zhang , Zhao Zhang

Multimodal systems have great potential to assist humans in procedural activities, where people follow instructions to achieve their goals. Despite diverse application scenarios, systems are typically evaluated on traditional classification…

In recent years, there has been growing interest in the video-based action quality assessment (AQA). Most existing methods typically solve AQA problem by considering the entire video yet overlooking the inherent stage-level characteristics…

计算机视觉与模式识别 · 计算机科学 2024-01-08 Qi An , Mengshi Qi , Huadong Ma