English
Related papers

Related papers: RFOP: Rethinking Fusion and Orthogonal Projection …

200 papers

We study the problem of learning association between face and voice, which is gaining interest in the computer vision community lately. Prior works adopt pairwise or triplet loss formulations to learn an embedding space amenable for…

Computer Vision and Pattern Recognition · Computer Science 2021-12-21 Muhammad Saad Saeed , Muhammad Haris Khan , Shah Nawaz , Muhammad Haroon Yousaf , Alessio Del Bue

Recent years have seen an increased interest in establishing association between faces and voices of celebrities leveraging audio-visual information from YouTube. Prior works adopt metric learning methods to learn an embedding space that is…

Computer Vision and Pattern Recognition · Computer Science 2022-08-23 Muhammad Saad Saeed , Shah Nawaz , Muhammad Haris Khan , Sajid Javed , Muhammad Haroon Yousaf , Alessio Del Bue

This paper presents Team Xaiofei's innovative approach to exploring Face-Voice Association in Multilingual Environments (FAME) at ACM Multimedia 2024. We focus on the impact of different languages in face-voice matching by building upon…

Computer Vision and Pattern Recognition · Computer Science 2024-07-30 Jiehui Tang , Xiaofei Wang , Zhen Xiao , Jiayi Liu , Xueliang Liu , Richang Hong

The advancements of technology have led to the use of multimodal systems in various real-world applications. Among them, audio-visual systems are among the most widely used multimodal systems. In the recent years, associating face and voice…

We study the task of learning association between faces and voices, which is gaining interest in the multimodal community lately. These methods suffer from the deliberate crafting of negative mining procedures as well as the reliance on the…

Computer Vision and Pattern Recognition · Computer Science 2025-05-29 Abdul Hannan , Muhammad Arslan Manzoor , Shah Nawaz , Muhammad Irzam Liaqat , Markus Schedl , Mubashir Noman

The FAME 2026 challenge comprises two demanding tasks: training face-voice associations combined with a multilingual setting that includes testing on languages on which the model was not trained. Our approach consists of separate uni-modal…

Sound · Computer Science 2025-12-05 Christopher Simic , Korbinian Riedhammer , Tobias Bocklet

Over half of the world's population is bilingual and people often communicate under multilingual scenarios. The Face-Voice Association in Multilingual Environments (FAME) 2026 Challenge, held at ICASSP 2026, focuses on developing methods…

The advancements of technology have led to the use of multimodal systems in various real-world applications. Among them, the audio-visual systems are one of the widely used multimodal systems. In the recent years, associating face and voice…

Exploring proper way to conduct multi-speech feature fusion for cross-corpus speech emotion recognition is crucial as different speech features could provide complementary cues reflecting human emotion status. While most previous approaches…

Sound · Computer Science 2024-06-14 Xueyu Liu , Jie Lin , Chao Wang

While deep learning technologies are now capable of generating realistic images confusing humans, the research efforts are turning to the synthesis of images for more concrete and application-specific purposes. Facial image generation based…

Computer Vision and Pattern Recognition · Computer Science 2020-06-11 Yeqi Bai , Tao Ma , Lipo Wang , Zhenjie Zhang

Today, there have been many achievements in learning the association between voice and face. However, most previous work models rely on cosine similarity or L2 distance to evaluate the likeness of voices and faces following contrastive…

Computer Vision and Pattern Recognition · Computer Science 2024-04-16 Chong Peng , Liqiang He , Dan Su

Training SER models in natural, spontaneous speech is especially challenging due to the subtle expression of emotions and the unpredictable nature of real-world audio. In this paper, we present a robust system for the INTERSPEECH 2025…

Speaker verification has been widely explored using speech signals, which has shown significant improvement using deep models. Recently, there has been a surge in exploring faces and voices as they can offer more complementary and…

Sound · Computer Science 2023-09-29 R. Gnana Praveen , Jahangir Alam

Multimodal speech emotion recognition aims to detect speakers' emotions from audio and text. Prior works mainly focus on exploiting advanced networks to model and fuse different modality information to facilitate performance, while…

Computation and Language · Computer Science 2023-04-11 Zhen Wu , Yizhe Lu , Xinyu Dai

In this paper, we present our solutions for emotion recognition in the sub-challenges of Multimodal Emotion Recognition Challenge (MER2024). To mitigate the modal competition issue between audio and text, we adopt an early fusion strategy…

Multimedia · Computer Science 2024-10-01 Mengying Ge , Mingyang Li , Dongkai Tang , Pengbo Li , Kuo Liu , Shuhao Deng , Songbai Pu , Long Liu , Yang Song , Tao Zhang

Face-voice association is widely studied in multimodal learning and is approached representing faces and voices with embeddings that are close for a same person and well separated from those of others. Previous work achieved this with loss…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Marta Moscati , Oleksandr Kats , Mubashir Noman , Muhammad Zaigham Zaheer , Yufang Hou , Markus Schedl , Shah Nawaz

Previous works on voice-face matching and voice-guided face synthesis demonstrate strong correlations between voice and face, but mainly rely on coarse semantic cues such as gender, age, and emotion. In this paper, we aim to investigate the…

Computer Vision and Pattern Recognition · Computer Science 2023-07-27 Xiang Li , Yandong Wen , Muqiao Yang , Jinglu Wang , Rita Singh , Bhiksha Raj

In this paper we present a technique for fusion of optical and thermal face images based on image pixel fusion approach. Out of several factors, which affect face recognition performance in case of visual images, illumination changes are a…

Computer Vision and Pattern Recognition · Computer Science 2010-07-06 Mrinal Kanti Bhowmik , Debotosh Bhattacharjee , Mita Nasipuri , Dipak Kumar Basu , Mahantapas Kundu

Fusion of scores is a cornerstone of multimodal biometric systems composed of independent unimodal parts. In this work, we focus on quality-dependent fusion for speaker-face verification. To this end, we propose a universal model which can…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-17 Grigory Antipov , Nicolas Gengembre , Olivier Le Blouch , Gaël Le Lan

In this paper, we study the associations between human faces and voices. Audiovisual integration, specifically the integration of facial and vocal information is a well-researched area in neuroscience. It is shown that the overlapping…

Computer Vision and Pattern Recognition · Computer Science 2018-11-05 Changil Kim , Hijung Valentina Shin , Tae-Hyun Oh , Alexandre Kaspar , Mohamed Elgharib , Wojciech Matusik
‹ Prev 1 2 3 10 Next ›