English
Related papers

Related papers: Multilingual Audio-Visual Smartphone Dataset And E…

200 papers

The growing sophistication of speech generated by Artificial Intelligence (AI) has introduced new challenges in audio deepfake detection. Text-to-speech (TTS) and voice conversion (VC) technologies can create highly convincing synthetic…

Sound · Computer Science 2026-03-17 Vamshi Nallaguntla , Aishwarya Fursule , Shruti Kshirsagar , Anderson R. Avila

Among user authentication methods, behavioural biometrics has proven to be effective against identity theft as well as user-friendly and unobtrusive. One of the most popular traits in the literature is keystroke dynamics due to the large…

Computer Vision and Pattern Recognition · Computer Science 2022-10-05 Giuseppe Stragapede , Paula Delgado-Santos , Ruben Tolosana , Ruben Vera-Rodriguez , Richard Guest , Aythami Morales

In recent years, an association is established between faces and voices of celebrities leveraging large scale audio-visual information from YouTube. The availability of large scale audio-visual datasets is instrumental in developing speaker…

Sound · Computer Science 2023-02-28 Saqlain Hussain Shah , Muhammad Saad Saeed , Shah Nawaz , Muhammad Haroon Yousaf

In this paper, we study an approach to multimodal person verification using audio, visual, and thermal modalities. The combination of audio and visual modalities has already been shown to be effective for robust person verification. From…

Computer Vision and Pattern Recognition · Computer Science 2022-03-07 Madina Abdrakhmanova , Saniya Abushakimova , Yerbolat Khassanov , Huseyin Atakan Varol

Eavesdropping from the user's smartphone is a well-known threat to the user's safety and privacy. Existing studies show that loudspeaker reverberation can inject speech into motion sensor readings, leading to speech eavesdropping. While…

Sound · Computer Science 2022-12-26 Ahmed Tanvir Mahdad , Cong Shi , Zhengkun Ye , Tianming Zhao , Yan Wang , Yingying Chen , Nitesh Saxena

Expanding the language coverage of speech technology has the potential to improve access to information for many more people. However, current speech technology is restricted to about one hundred languages which is a small fraction of the…

Automatic speech recognition systems are part of people's daily lives, embedded in personal assistants and mobile phones, helping as a facilitator for human-machine interaction while allowing access to information in a practically intuitive…

Sound · Computer Science 2021-10-05 Julio Cesar Duarte , Sérgio Colcher

We present a large, multilingual study into how vision constrains linguistic choice, covering four languages and five linguistic properties, such as verb transitivity or use of numerals. We propose a novel method that leverages existing…

Computation and Language · Computer Science 2023-02-10 Uri Berger , Lea Frermann , Gabriel Stanovsky , Omri Abend

In this study, we introduce a novel multi-modal biometric authentication system that integrates facial, vocal, and signature data to enhance security measures. Utilizing a combination of Convolutional Neural Networks (CNNs) and Recurrent…

Computer Vision and Pattern Recognition · Computer Science 2024-11-05 Vatchala S , Yogesh C , Yeshwanth Govindarajan , Krithik Raja M , Vishal Pramav Amirtha Ganesan , Aashish Vinod A , Dharun Ramesh

Multi-modal learning in the audio-language domain has seen significant advancements in recent years. However, audio-language learning faces challenges due to limited and lower-quality data compared to image-language tasks. Existing…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-10 David Xu

Significant advancements made in the generation of deepfakes have caused security and privacy issues. Attackers can easily impersonate a person's identity in an image by replacing his face with the target person's face. Moreover, a new…

Computer Vision and Pattern Recognition · Computer Science 2021-09-08 Hasam Khalid , Minha Kim , Shahroz Tariq , Simon S. Woo

Audio-driven 3D facial animation has been widely explored, but achieving realistic, human-like performance is still unsolved. This is due to the lack of available 3D datasets, models, and standard evaluation metrics. To address this, we…

Computer Vision and Pattern Recognition · Computer Science 2019-05-09 Daniel Cudeiro , Timo Bolkart , Cassidy Laidlaw , Anurag Ranjan , Michael J. Black

With the rapid advancement of technology, different biometric user authentication, and identification systems are emerging. Traditional biometric systems like face, fingerprint, and iris recognition, keystroke dynamics, etc. are prone to…

The lack of standard input interfaces in the Internet of Things (IoT) ecosystems presents a challenge in securing such infrastructures. To tackle this challenge, we introduce a novel behavioral biometric system based on naturally occurring…

Cryptography and Security · Computer Science 2022-02-22 Klaudia Krawiecka , Simon Birnbach , Simon Eberz , Ivan Martinovic

Integrating vision and language has long been a dream in work on artificial intelligence (AI). In the past two years, we have witnessed an explosion of work that brings together vision and language from images to videos and beyond. The…

Computation and Language · Computer Science 2021-08-23 Francis Ferraro , Nasrin Mostafazadeh , Ting-Hao , Huang , Lucy Vanderwende , Jacob Devlin , Michel Galley , Margaret Mitchell

One of the many tasks facing the typically-developing child language learner is learning to discriminate between the distinctive sounds that make up words in their native language. Here we investigate whether multimodal…

Computation and Language · Computer Science 2024-07-24 Sophia Zhi , Roger P. Levy , Stephan C. Meylan

Recent works demonstrate that voice assistants do not perform equally well for everyone, but research on demographic robustness of speech technologies is still scarce. This is mainly due to the rarity of large datasets with controlled…

The need for reliably determining the identity of a person is critical in a number of different domains ranging from personal smartphones to border security; from autonomous vehicles to e-voting; from tracking child vaccinations to…

Computer Vision and Pattern Recognition · Computer Science 2019-05-14 Arun Ross , Sudipta Banerjee , Cunjian Chen , Anurag Chowdhury , Vahid Mirjalili , Renu Sharma , Thomas Swearingen , Shivangi Yadav

The rapid advancement of deepfake technology poses a significant threat to digital media integrity. Deepfakes, synthetic media created using AI, can convincingly alter videos and audio to misrepresent reality. This creates risks of…

Computer Vision and Pattern Recognition · Computer Science 2024-10-21 Kashish Gandhi , Prutha Kulkarni , Taran Shah , Piyush Chaudhari , Meera Narvekar , Kranti Ghag

Audio is a rich sensing modality that is useful for a variety of human activity recognition tasks. However, the ubiquitous nature of smartphones and smart speakers with always-on microphones has led to numerous privacy concerns and a lack…