English
Related papers

Related papers: Audio-Visual Kinship Verification

200 papers

Visual and audio events simultaneously occur and both attract attention. However, most existing saliency prediction works ignore the influence of audio and only consider vision modality. In this paper, we propose a multitask learning method…

Computer Vision and Pattern Recognition · Computer Science 2021-11-17 Minglang Qiao , Yufan Liu , Mai Xu , Xin Deng , Bing Li , Weiming Hu , Ali Borji

Automatic speech recognition can potentially benefit from the lip motion patterns, complementing acoustic speech to improve the overall recognition performance, particularly in noise. In this paper we propose an audio-visual fusion strategy…

Audio and Speech Processing · Electrical Eng. & Systems 2019-05-02 George Sterpu , Christian Saam , Naomi Harte

Person Re-Identification (ReID) requires comparing two images of person captured under different conditions. Existing work based on neural networks often computes the similarity of feature maps from one single convolutional layer. In this…

Computer Vision and Pattern Recognition · Computer Science 2018-04-03 Yiluan Guo , Ngai-Man Cheung

A complex visual navigation task puts an agent in different situations which call for a diverse range of visual perception abilities. For example, to "go to the nearest chair", the agent might need to identify a chair in a living room using…

Computer Vision and Pattern Recognition · Computer Science 2021-08-05 Bokui Shen , Danfei Xu , Yuke Zhu , Leonidas J. Guibas , Li Fei-Fei , Silvio Savarese

Face recognition in real life situations like low illumination condition is still an open challenge in biometric security. It is well established that the state-of-the-art methods in face recognition provide low accuracy in the case of poor…

Computer Vision and Pattern Recognition · Computer Science 2019-02-26 Sumit Agarwal , Harshit S. Sikchi , Suparna Rooj , Shubhobrata Bhattacharya , Aurobinda Routray

Face videos accompanied by audio have become integral to our daily lives, while they often suffer from complex degradations. Most face video restoration methods neglect the intrinsic correlations between the visual and audio features,…

Computer Vision and Pattern Recognition · Computer Science 2025-08-07 Yuqin Cao , Yixuan Gao , Wei Sun , Xiaohong Liu , Yulun Zhang , Xiongkuo Min

Realistic fake videos are a potential tool for spreading harmful misinformation given our increasing online presence and information intake. This paper presents a multimodal learning-based method for detection of real and fake videos. The…

Computer Vision and Pattern Recognition · Computer Science 2022-07-19 Kalin Stefanov , Bhawna Paliwal , Abhinav Dhall

Visual speech recognition is a challenging research problem with a particular practical application of aiding audio speech recognition in noisy scenarios. Multiple camera setups can be beneficial for the visual speech recognition systems in…

Computer Vision and Pattern Recognition · Computer Science 2018-06-29 Marina Zimmermann , Mostafa Mehdipour Ghazi , Hazım Kemal Ekenel , Jean-Philippe Thiran

Speaker verification systems have been used in many production scenarios in recent years. Unfortunately, they are still highly prone to different kinds of spoofing attacks such as voice conversion and speech synthesis, etc. In this paper,…

Audio and Speech Processing · Electrical Eng. & Systems 2021-09-07 Junxiao Xue , Hao Zhou , Yabo Wang

Classical person re-identification approaches assume that a person of interest has appeared across different cameras and can be queried by one of the existing images. However, in real-world surveillance scenarios, frequently no visual…

Computer Vision and Pattern Recognition · Computer Science 2020-03-03 Ammarah Farooq , Muhammad Awais , Fei Yan , Josef Kittler , Ali Akbari , Syed Safwan Khalid

Searching persons in large-scale image databases with the query of natural language description is a more practical important applications in video surveillance. Intuitively, for person search, the core issue should be visual-textual…

Computer Vision and Pattern Recognition · Computer Science 2019-12-09 Jing Ge , Guangyu Gao , Zhen Liu

The innate correlation between a person's face and voice has recently emerged as a compelling area of study, especially within the context of multilingual environments. This paper introduces our novel solution to the Face-Voice Association…

Sound · Computer Science 2024-08-20 Wuyang Chen , Yanjie Sun , Kele Xu , Yong Dou

One of the most pressing challenges for the detection of face-manipulated videos is generalising to forgery methods not seen during training while remaining effective under common corruptions such as compression. In this paper, we examine…

Computer Vision and Pattern Recognition · Computer Science 2022-10-24 Alexandros Haliassos , Rodrigo Mira , Stavros Petridis , Maja Pantic

Deep learning-based methods have pushed the limits of the state-of-the-art in face analysis. However, despite their success, these models have raised concerns regarding their bias towards certain demographics. This bias is inflicted both by…

Computer Vision and Pattern Recognition · Computer Science 2020-09-10 Markos Georgopoulos , Yannis Panagakis , Maja Pantic

Humans possess a remarkable ability to integrate auditory and visual information, enabling a deeper understanding of the surrounding environment. This early fusion of audio and visual cues, demonstrated through cognitive psychology and…

Computer Vision and Pattern Recognition · Computer Science 2023-12-05 Shentong Mo , Pedro Morgado

This thesis combines audio-analysis with computer vision to approach Music Information Retrieval (MIR) tasks from a multi-modal perspective. This thesis focuses on the information provided by the visual layer of music videos and how it can…

Multimedia · Computer Science 2020-02-04 Alexander Schindler

Skin tone recognition and generation play important roles in model fairness, healthcare, and generative AI, yet they remain challenging due to the lack of comprehensive datasets and robust methodologies. Compared to other human image…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Haoming Lu

Singing voice transcription converts recorded singing audio to musical notation. Sound contamination (such as accompaniment) and lack of annotated data make singing voice transcription an extremely difficult task. We take two approaches to…

Sound · Computer Science 2023-04-25 Xiangming Gu , Wei Zeng , Jianan Zhang , Longshen Ou , Ye Wang

In this paper we present a technique for fusion of optical and thermal face images based on image pixel fusion approach. Out of several factors, which affect face recognition performance in case of visual images, illumination changes are a…

Computer Vision and Pattern Recognition · Computer Science 2010-07-06 Mrinal Kanti Bhowmik , Debotosh Bhattacharjee , Mita Nasipuri , Dipak Kumar Basu , Mahantapas Kundu

Common and important applications of person identification occur at distances and viewpoints in which the face is not visible or is not sufficiently resolved to be useful. We examine body shape as a biometric across distance and viewpoint…

Computer Vision and Pattern Recognition · Computer Science 2023-05-31 Blake A. Myers , Lucas Jaggernauth , Thomas M. Metz , Matthew Q. Hill , Veda Nandan Gandi , Carlos D. Castillo , Alice J. O'Toole
‹ Prev 1 4 5 6 7 8 10 Next ›