English
Related papers

Related papers: A Discriminative Vectorial Framework for Multi-mod…

200 papers

As artificial intelligence systems increasingly operate in Real-world environments, the integration of multi-modal data sources such as vision, language, and audio presents both unprecedented opportunities and critical challenges for…

Machine Learning · Computer Science 2025-07-01 Sree Bhargavi Balija

Accurate beam prediction is essential for mitigating signalling overhead and latency in integrated sensing and communication-enabled massive multi-input multi-output systems. With the aid of multimodal learning, the prediction accuracy can…

Signal Processing · Electrical Eng. & Systems 2026-05-15 Zijian Zheng , Wenqiang Yi , Hyundong Shin , Arumugam Nallanathan

Understanding the semantic shifts of multimodal information is only possible with models that capture cross-modal interactions over time. Under this paradigm, a new embedding is needed that structures visual-textual interactions according…

Multimedia · Computer Science 2019-10-01 David Semedo , João Magalhães

Matrix factorization has been recently utilized for the task of multi-modal hashing for cross-modality visual search, where basis functions are learned to map data from different modalities to the same Hamming embedding. In this paper, we…

Information Retrieval · Computer Science 2016-04-19 Hong Liu , Rongrong Ji , Yongjian Wu , Gang Hua

Facial expression recognition is a challenging task when neural network is applied to pattern recognition. Most of the current recognition research is based on single source facial data, which generally has the disadvantages of low accuracy…

Computer Vision and Pattern Recognition · Computer Science 2021-09-28 Yi Han , Xubin Wang , Zhengyu Lu

Detailed phenotype information is fundamental to accurate diagnosis and risk estimation of diseases. As a rich source of phenotype information, electronic health records (EHRs) promise to empower diagnostic variant interpretation. However,…

Machine Learning · Computer Science 2023-04-28 Shenghan Zhang , Haoxuan Li , Ruixiang Tang , Sirui Ding , Laila Rasmy , Degui Zhi , Na Zou , Xia Hu

Feature modeling of different modalities is a basic problem in current research of cross-modal information retrieval. Existing models typically project texts and images into one embedding space, in which semantically similar information…

Multimedia · Computer Science 2019-06-13 Jing Yu , Chenghao Yang , Zengchang Qin , Zhuoqian Yang , Yue Hu , Weifeng Zhang

We propose a novel framework, called Disjoint Mapping Network (DIMNet), for cross-modal biometric matching, in particular of voices and faces. Different from the existing methods, DIMNet does not explicitly learn the joint relationship…

Computer Vision and Pattern Recognition · Computer Science 2018-07-17 Yandong Wen , Mahmoud Al Ismail , Weiyang Liu , Bhiksha Raj , Rita Singh

Differential diagnosis of mental disorders remains a fundamental challenge in real-world clinical practice, where multiple conditions often exhibit overlapping symptoms. However, most existing public datasets are developed under…

Feature extraction is a key step in image processing for pattern recognition and machine learning processes. Its purpose lies in reducing the dimensionality of the input data through the computing of features which accurately describe the…

Computer Vision and Pattern Recognition · Computer Science 2020-05-14 Thomas Lacombe , Hugues Favreliere , Maurice Pillet

In deep metric learning (DML), high-level input data are represented in a lower-level representation (embedding) space, such that samples from the same class are mapped close together, while samples from disparate classes are mapped further…

Computer Vision and Pattern Recognition · Computer Science 2023-04-03 Ryan Furlong , Vincent O'Brien , James Garland , Daniel Palacios-Alonso , Francisco Dominguez-Mateos

Current movie dubbing technology can produce the desired speech using a reference voice and input video, maintaining perfect synchronization with the visuals while effectively conveying the intended emotions. However, crucial aspects of…

Multimedia · Computer Science 2025-05-23 Junjie Zheng , Zihao Chen , Chaofan Ding , Yunming Liang , Yihan Fan , Huan Yang , Lei Xie , Xinhan Di

Multimodal sentiment analysis is a fundamental problem in the field of affective computing. Although significant progress has been made in cross-modal interaction, it remains a challenge due to the insufficient reference context in…

Multimedia · Computer Science 2025-08-12 Xianbing Zhao , Shengzun Yang , Buzhou Tang , Ronghuan Jiang

This paper presents a new deep learning approach for video-based scene classification. We design a Heterogeneous Deep Discriminative Model (HDDM) whose parameters are initialized by performing an unsupervised pre-training in a layer-wise…

Computer Vision and Pattern Recognition · Computer Science 2018-07-24 Mohammad Tavakolian , Abdenour Hadid

Recently, multi-view learning (MVL) has garnered significant attention due to its ability to fuse discriminative information from multiple views. However, real-world multi-view datasets are often heterogeneous and imperfect, which usually…

Computer Vision and Pattern Recognition · Computer Science 2025-07-25 Jie Xu , Na Zhao , Gang Niu , Masashi Sugiyama , Xiaofeng Zhu

This paper introduces a deep learning enabled generative sensing framework which integrates low-end sensors with computational intelligence to attain a high recognition accuracy on par with that attained with high-end sensors. The proposed…

Computer Vision and Pattern Recognition · Computer Science 2018-01-10 Lina Karam , Tejas Borkar , Yu Cao , Junseok Chae

State-of-the-art deep learning algorithms generally require large amounts of data for model training. Lack thereof can severely deteriorate the performance, particularly in scenarios with fine-grained boundaries between categories. To this…

Computer Vision and Pattern Recognition · Computer Science 2018-06-15 Frederik Pahde , Patrick Jähnichen , Tassilo Klein , Moin Nabi

Deep learning-based techniques for the analysis of multimodal remote sensing data have become popular due to their ability to effectively integrate complementary spatial, spectral, and structural information from different sensors.…

Computer Vision and Pattern Recognition · Computer Science 2026-04-10 Hao Liu , Yongjie Zheng , Yuhan Kang , Mingyang Zhang , Maoguo Gong , Lorenzo Bruzzone

Visible-to-thermal face image matching is a challenging variate of cross-modality recognition. The challenge lies in the large modality gap and low correlation between visible and thermal modalities. Existing approaches employ image…

Computer Vision and Pattern Recognition · Computer Science 2021-11-30 Usman Cheema , Mobeen Ahmad , Dongil Han , Seungbin Moon

The 3D shapes of faces are well known to be discriminative. Yet despite this, they are rarely used for face recognition and always under controlled viewing conditions. We claim that this is a symptom of a serious but often overlooked…

Computer Vision and Pattern Recognition · Computer Science 2016-12-16 Anh Tuan Tran , Tal Hassner , Iacopo Masi , Gerard Medioni