English
Related papers

Related papers: Learning Audio-Visual embedding for Person Verific…

200 papers

Incremental improvements in accuracy of Convolutional Neural Networks are usually achieved through use of deeper and more complex models trained on larger datasets. However, enlarging dataset and models increases the computation and storage…

Audio and Speech Processing · Electrical Eng. & Systems 2018-07-24 Mahdi Hajibabaei , Dengxin Dai

The FAME 2026 challenge comprises two demanding tasks: training face-voice associations combined with a multilingual setting that includes testing on languages on which the model was not trained. Our approach consists of separate uni-modal…

Sound · Computer Science 2025-12-05 Christopher Simic , Korbinian Riedhammer , Tobias Bocklet

State-of-the-art transformer models for Speech Emotion Recognition (SER) rely on temporal feature aggregation, yet advanced pooling methods remain underexplored. We systematically benchmark pooling strategies, including Multi-Query…

In this paper, we propose self-supervised speaker representation learning strategies, which comprise of a bootstrap equilibrium speaker representation learning in the front-end and an uncertainty-aware probabilistic speaker embedding…

Audio and Speech Processing · Electrical Eng. & Systems 2021-12-28 Sung Hwan Mun , Min Hyun Han , Dongjune Lee , Jihwan Kim , Nam Soo Kim

Visual recognition inside the vehicle cabin leads to safer driving and more intuitive human-vehicle interaction but such systems face substantial obstacles as they need to capture different granularities of driver behaviour while dealing…

Computer Vision and Pattern Recognition · Computer Science 2022-04-12 Alina Roitberg , Kunyu Peng , Zdravko Marinov , Constantin Seibold , David Schneider , Rainer Stiefelhagen

In video-based emotion recognition (ER), it is important to effectively leverage the complementary relationship among audio (A) and visual (V) modalities, while retaining the intra-modal characteristics of individual modalities. In this…

Computer Vision and Pattern Recognition · Computer Science 2023-04-18 R Gnana Praveen , Eric Granger , Patrick Cardinal

We propose a new method for speaker diarization that can handle overlapping speech with 2+ people. Our method is based on compositional embeddings [1]: Like standard speaker embedding methods such as x-vector [2], compositional embedding…

Sound · Computer Science 2021-02-11 Zeqian Li , Jacob Whitehill

This paper proposes an improved approach for open-set speaker identification based on pretrained speaker foundation models. Building upon the previous Speaker Reciprocal Points Learning framework (V1), we first introduce an enhanced…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-16 Zhiyong Chen , Shuhang Wu , Yingjie Duan , Xinkang Xu , Xinhui Hu

Predicting brain activity in response to naturalistic, multimodal stimuli is a key challenge in computational neuroscience. While encoding models are becoming more powerful, their ability to generalize to truly novel contexts remains a…

Computer Vision and Pattern Recognition · Computer Science 2025-07-28 Hamid Abdollahi , Amir Hossein Mansouri Majoumerd , Amir Hossein Bagheri Baboukani , Amir Abolfazl Suratgar , Mohammad Bagher Menhaj

We participated in the 10th ABAW Challenge, focusing on the Emotional Mimicry Intensity (EMI) Estimation track on the Hume-Vidmimic2 dataset. This task aims to predict six continuous emotion dimensions: Admiration, Amusement, Determination,…

Artificial Intelligence · Computer Science 2026-03-17 Jiawen Huang , Chenxi Huang , Zhuofan Wen , Hailiang Yao , Shun Chen , Longjiang Yang , Cong Yu , Fengyu Zhang , Ran Liu , Bin Liu

Contrastive language-image pre-training aligns the features of text-image pairs in a common latent space via distinct encoders for each modality. While this approach achieves impressive performance in several zero-shot tasks, it cannot…

Computer Vision and Pattern Recognition · Computer Science 2025-06-04 Christian Schlarmann , Francesco Croce , Nicolas Flammarion , Matthias Hein

Person re-identification is an important task in video surveillance that aims to associate people across camera views at different locations and time. View variability is always a challenging problem seriously degrading person…

Computer Vision and Pattern Recognition · Computer Science 2019-10-10 Fangyi Liu , Lei Zhang

Many of the existing methods for learning joint embedding of images and text use only supervised information from paired images and its textual attributes. Taking advantage of the recent success of unsupervised learning in deep neural…

Computer Vision and Pattern Recognition · Computer Science 2017-03-21 Yao-Hung Hubert Tsai , Liang-Kang Huang , Ruslan Salakhutdinov

While promising performance for speaker verification has been achieved by deep speaker embeddings, the advantage would reduce in the case of speaking-style variability. Speaking rate mismatch is often observed in practical speaker…

Audio and Speech Processing · Electrical Eng. & Systems 2022-05-31 Fuchuan Tong , Siqi Zheng , Haodong Zhou , Xingjia Xie , Qingyang Hong , Lin Li

In this paper, we propose a novel bidirectional multiscale feature aggregation (BMFA) network with attentional fusion modules for text-independent speaker verification. The feature maps from different stages of the backbone network are…

Audio and Speech Processing · Electrical Eng. & Systems 2021-04-02 Jiajun Qi , Wu Guo , Bin Gu

This paper aims to learn a compact representation of a video for video face recognition task. We make the following contributions: first, we propose a meta attention-based aggregation scheme which adaptively and fine-grained weighs the…

Computer Vision and Pattern Recognition · Computer Science 2019-09-13 Zhaoxiang Liu , Huan Hu , Jinqiang Bai , Shaohua Li , Shiguo Lian

Audio-Visual Target Speaker Extraction (AVTSE) is crucial for cocktail party scenarios. Leveraging multiple cues --such as utterance-level speaker embeddings or steady face images, and frame-level lip motion or facial expression features…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-12 Zhan Jin , Bang Zeng , Peijun Yang , Jiarong Du , Wei Ju , Yao Tian , Juan Liu , Ming Li

Video-based person recognition achieves robust identification by integrating face, body, and gait. However, current systems waste computational resources by processing all modalities with fixed heavyweight ensembles regardless of input…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Yuyang Ji , Yixuan Shen , Kien Nguyen , Lifeng Zhou , Feng Liu

This paper proposes a serialized multi-layer multi-head attention for neural speaker embedding in text-independent speaker verification. In prior works, frame-level features from one layer are aggregated to form an utterance-level…

Sound · Computer Science 2021-07-15 Hongning Zhu , Kong Aik Lee , Haizhou Li

Existing open-vocabulary object detection (OVD) develops methods for testing unseen categories by aligning object region embeddings with corresponding VLM features. A recent study leverages the idea that VLMs implicitly learn compositional…

Computer Vision and Pattern Recognition · Computer Science 2024-12-25 Hojun Choi , Junsuk Choe , Hyunjung Shim
‹ Prev 1 8 9 10 Next ›