English
Related papers

Related papers: Shared Multi-modal Embedding Space for Face-Voice …

200 papers

Face Recognition (FR) tasks have made significant progress with the advent of Deep Neural Networks, particularly through margin-based triplet losses that embed facial images into high-dimensional feature spaces. During training, these…

Computer Vision and Pattern Recognition · Computer Science 2025-07-16 Pierrick Leroy , Antonio Mastropietro , Marco Nurisso , Francesco Vaccarino

Fusion of scores is a cornerstone of multimodal biometric systems composed of independent unimodal parts. In this work, we focus on quality-dependent fusion for speaker-face verification. To this end, we propose a universal model which can…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-17 Grigory Antipov , Nicolas Gengembre , Olivier Le Blouch , Gaël Le Lan

With the advance in self-supervised learning for audio and visual modalities, it has become possible to learn a robust audio-visual speech representation. This would be beneficial for improving the audio-visual speech recognition (AVSR)…

Image and Video Processing · Electrical Eng. & Systems 2022-07-12 Zi-Qiang Zhang , Jie Zhang , Jian-Shu Zhang , Ming-Hui Wu , Xin Fang , Li-Rong Dai

This paper presents our submission to the Expression Classification Challenge of the fifth Affective Behavior Analysis in-the-wild (ABAW) Competition. In our method, multimodal feature combinations extracted by several different pre-trained…

Computer Vision and Pattern Recognition · Computer Science 2023-03-20 Chuanhe Liu , Xinjie Zhang , Xiaolong Liu , Tenggan Zhang , Liyu Meng , Yuchen Liu , Yuanyuan Deng , Wenqiang Jiang

Face Anti-Spoofing (FAS) is essential for ensuring the security and reliability of facial recognition systems. Most existing FAS methods are formulated as binary classification tasks, providing confidence scores without interpretation. They…

Computer Vision and Pattern Recognition · Computer Science 2025-01-27 Guosheng Zhang , Keyao Wang , Haixiao Yue , Ajian Liu , Gang Zhang , Kun Yao , Errui Ding , Jingdong Wang

Speaker verification, as a biometric authentication mechanism, has been widely used due to the pervasiveness of voice control on smart devices. However, the task of "in-the-wild" speaker verification is still challenging, considering the…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-27 Jianwei Tai , Xiaoqi Jia , Qingjia Huang , Weijuan Zhang , Haichao Du , Shengzhi Zhang

The performance of deep learning models is critically dependent on sophisticated optimization strategies. While existing optimizers have shown promising results, many rely on first-order Exponential Moving Average (EMA) techniques, which…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Roi Peleg , Yair Smadar , Teddy Lazebnik , Assaf Hoogi

The prevalent approach in speech emotion recognition (SER) involves integrating both audio and textual information to comprehensively identify the speaker's emotion, with the text generally obtained through automatic speech recognition…

Computation and Language · Computer Science 2024-05-29 Jiajun He , Xiaohan Shi , Xingfeng Li , Tomoki Toda

Most of the existing work on automatic facial expression analysis focuses on discrete emotion recognition, or facial action unit detection. However, facial expressions do not always fall neatly into pre-defined semantic categories. Also,…

Computer Vision and Pattern Recognition · Computer Science 2019-01-11 Raviteja Vemulapalli , Aseem Agarwala

Deep metric learning aims to learn an embedding function, modeled as deep neural network. This embedding function usually puts semantically similar images close while dissimilar images far from each other in the learned embedding space.…

Computer Vision and Pattern Recognition · Computer Science 2018-09-03 Wonsik Kim , Bhavya Goyal , Kunal Chawla , Jungmin Lee , Keunjoo Kwon

End-to-end acoustic-to-word speech recognition models have recently gained popularity because they are easy to train, scale well to large amounts of training data, and do not require a lexicon. In addition, word models may also be easier to…

Computation and Language · Computer Science 2019-02-20 Shruti Palaskar , Vikas Raunak , Florian Metze

Audiovisual active speaker detection (ASD) in egocentric recordings is challenged by frequent occlusions, motion blur, and audio interference, which undermine the discernability of temporal synchrony between lip movement and speech.…

Multimedia · Computer Science 2025-08-15 Jason Clarke , Yoshihiko Gotoh , Stefan Goetze

Learning shared representations is a primary area of multimodal representation learning. The current approaches to achieve a shared embedding space rely heavily on paired samples from each modality, which are significantly harder to obtain…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Amitai Yacobi , Nir Ben-Ari , Ronen Talmon , Uri Shaham

Humans possess a remarkable ability to integrate auditory and visual information, enabling a deeper understanding of the surrounding environment. This early fusion of audio and visual cues, demonstrated through cognitive psychology and…

Computer Vision and Pattern Recognition · Computer Science 2023-12-05 Shentong Mo , Pedro Morgado

Audio-visual speech recognition (AVSR) attracts a surge of research interest recently by leveraging multimodal signals to understand human speech. Mainstream approaches addressing this task have developed sophisticated architectures and…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-21 Yuchen Hu , Chen Chen , Ruizhe Li , Heqing Zou , Eng Siong Chng

Emotional Mimicry Intensity (EMI) estimation plays a pivotal role in understanding human social behavior and advancing human-computer interaction. The core challenges lie in dynamic correlation modeling and robust fusion of multimodal…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Jun Yu , Lingsi Zhu , Yanjun Chi , Yunxiang Zhang , Yang Zheng , Yongqi Wang , Xilong Lu

In this paper we address the problems of modeling the acoustic space generated by a full-spectrum sound source and of using the learned model for the localization and separation of multiple sources that simultaneously emit sparse-spectrum…

Sound · Computer Science 2015-02-06 Antoine Deleforge , Florence Forbes , Radu Horaud

Radiotherapists require accurate registration of MR/CT images to effectively use information from both modalities. In a typical registration pipeline, rigid or affine transformations are applied to roughly align the fixed and moving images…

Computer Vision and Pattern Recognition · Computer Science 2023-07-10 Xiaoyu Bai , Fan Bai , Xiaofei Huo , Jia Ge , Tony C. W. Mok , Zi Li , Minfeng Xu , Jingren Zhou , Le Lu , Dakai Jin , Xianghua Ye , Jingjing Lu , Ke Yan

This paper describes the UZH-CL system submitted to the FAME2026 Challenge. The challenge focuses on cross-modal verification under unique multilingual conditions, specifically unseen and unheard languages. Our approach investigates two…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-03 Aref Farhadipour , Teodora Vukovic , Volker Dellwo

Understanding audio-visual content and the ability to have an informative conversation about it have both been challenging areas for intelligent systems. The Audio Visual Scene-aware Dialog (AVSD) challenge, organized as a track of the…

Computation and Language · Computer Science 2018-12-19 Dat Tien Nguyen , Shikhar Sharma , Hannes Schulz , Layla El Asri
‹ Prev 1 3 4 5 6 7 10 Next ›