English
Related papers

Related papers: Multilingual Audio-Visual Smartphone Dataset And E…

200 papers

Multilingual models can improve language processing, particularly for low resource situations, by sharing parameters across languages. Multilingual acoustic models, however, generally ignore the difference between phonemes (sounds that can…

Biometrics are widely used for authentication in consumer devices and business settings as they provide sufficiently strong security, instant verification and convenience for users. However, biometrics are hard to keep secret, stolen…

Cryptography and Security · Computer Science 2018-08-22 Mozhgan Azimpourkivi , Umut Topkara , Bogdan Carbunar

Speech recognition is very challenging in student learning environments that are characterized by significant cross-talk and background noise. To address this problem, we present a bilingual speech recognition system that uses an…

Spoofing attacks posed by generating artificial speech can severely degrade the performance of a speaker verification system. Recently, many anti-spoofing countermeasures have been proposed for detecting varying types of attacks from…

Audio and Speech Processing · Electrical Eng. & Systems 2020-12-08 Yuanjun Zhao , Roberto Togneri , Victor Sreeram

Deepfakes represent a growing concern across domains such as disinformation, fraud, and non-consensual media. In particular, the rise of video conference and identity-driven attacks in high-stakes scenarios--such as impostor hiring--demands…

Computer Vision and Pattern Recognition · Computer Science 2026-04-08 Sarah Barrington , Maty Bohacek , Hany Farid

Recent years has witnessed an increase in technologies that use speech for the sensing of the health of the talker. This survey paper proposes a general taxonomy of the technologies and a broad overview of current progress and challenges.…

Neurons and Cognition · Quantitative Biology 2024-08-12 Aki Härmä , Bert den Brinker , Ulf Grossekathofer , Okke Ouweltjes , Srikanth Nallanthighal , Sidharth Abrol , Vibhu Sharma

We review current solutions and technical challenges for automatic speech recognition, keyword spotting, device arbitration, speech enhancement, and source localization in multidevice home environments to provide context for the INTERSPEECH…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-01 Gregory Ciccarelli , Jarred Barber , Arun Nair , Israel Cohen , Tao Zhang

Automatic speaker verification (ASV) is one of the most natural and convenient means of biometric person recognition. Unfortunately, just like all other biometric systems, ASV is vulnerable to spoofing, also referred to as "presentation…

While the significant advancements have made in the generation of deepfakes using deep learning technologies, its misuse is a well-known issue now. Deepfakes can cause severe security and privacy issues as they can be used to impersonate a…

Computer Vision and Pattern Recognition · Computer Science 2022-03-02 Hasam Khalid , Shahroz Tariq , Minha Kim , Simon S. Woo

Biometric systems involve security assurance to make our system highly secured and robust. Nowadays, biometric technology has been fixed into new systems with the aim of enforcing strong privacy and security. Several innovative system have…

Cryptography and Security · Computer Science 2022-05-24 Jones Yeboah , Victor Adewopo , Sylvia Azumah , Izunna Okpala

We consider the task of animating 3D facial geometry from speech signal. Existing works are primarily deterministic, focusing on learning a one-to-one mapping from speech signal to 3D face meshes on small datasets with limited speakers.…

Computer Vision and Pattern Recognition · Computer Science 2023-12-01 Karren D. Yang , Anurag Ranjan , Jen-Hao Rick Chang , Raviteja Vemulapalli , Oncel Tuzel

Recent studies have shown how motion-based biometrics can be used as a form of user authentication and identification without requiring any human cooperation. This category of behavioural biometrics deals with the features we learn in our…

Machine Learning · Computer Science 2021-10-08 Akriti Verma , Valeh Moghaddam , Adnan Anwar

The Multi-language Video Subtitle Dataset is a comprehensive collection designed to support research in text recognition across multiple languages. This dataset includes 4,224 subtitle images extracted from 24 videos sourced from online…

Computer Vision and Pattern Recognition · Computer Science 2024-11-11 Thanadol Singkhornart , Olarik Surinta

Active speaker detection requires a solid integration of multi-modal cues. While individual modalities can approximate a solution, accurate predictions can only be achieved by explicitly fusing the audio and visual features and modeling…

Computer Vision and Pattern Recognition · Computer Science 2021-10-06 Juan León-Alcázar , Fabian Caba Heilbron , Ali Thabet , Bernard Ghanem

Large Multimodal Models (LMMs) are typically trained on vast corpora of image-text data but are often limited in linguistic coverage, leading to biased and unfair outputs across languages. While prior work has explored multimodal…

Computer Vision and Pattern Recognition · Computer Science 2025-07-11 Ananya Raval , Aravind Narayanan , Vahid Reza Khazaie , Shaina Raza

Nowadays, mobile smart devices are widely used in daily life. It is increasingly important to prevent malicious users from accessing private data, thus a secure and convenient authentication method is urgently needed. Compared with common…

Cryptography and Security · Computer Science 2025-04-02 Yadong Xie , Fan Li , Yu Wang

Dialog systems need to understand dynamic visual scenes in order to have conversations with users about the objects and events around them. Scene-aware dialog systems for real-world applications could be developed by integrating…

Prevailing user authentication schemes on smartphones rely on explicit user interaction, where a user types in a passcode or presents a biometric cue such as face, fingerprint, or iris. In addition to being cumbersome and obtrusive to the…

Computer Vision and Pattern Recognition · Computer Science 2019-01-17 Debayan Deb , Arun Ross , Anil K. Jain , Kwaku Prakah-Asante , K. Venkatesh Prasad

This paper provides a comprehensive evaluation of demographic and linguistic biases in omnimodal language models that process text, images, audio, and video within a single framework. Although these models are being widely deployed, their…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Alaa Elobaid

This paper deals with the identification of machines in a smart city environment. The concept of machine biometrics is proposed in this work for the first time, as a way to authenticate machine identities interacting with humans in everyday…

Machine Learning · Computer Science 2021-03-01 G. K. Sidiropoulos , G. A. Papakostas