English
Related papers

Related papers: WhisperWand: Simultaneous Voice and Gesture Tracki…

200 papers

Human identification plays an important role in human-computer interaction. There have been numerous methods proposed for human identification (e.g., face recognition, gait recognition, fingerprint identification, etc.). While these methods…

Human-Computer Interaction · Computer Science 2016-08-12 Tong Xin , Bin Guo , Zhu Wang , Mingyang Li , Zhiwen Yu

This report presents the technical details of our submission on the EGO4D Audio-Visual (AV) Automatic Speech Recognition Challenge 2023 from the OxfordVGG team. We present WhisperX, a system for efficient speech transcription of long-form…

Sound · Computer Science 2023-07-19 Jaesung Huh , Max Bain , Andrew Zisserman

Long-range Human-Robot Interaction (HRI) remains underexplored. Within it, Command Source Identification (CSI) - determining who issued a command - is especially challenging due to multi-user and distance-induced sensor ambiguity. We…

Human-Computer Interaction · Computer Science 2026-03-26 Chengwen Zhang , Chun Yu , Borong Zhuang , Haopeng Jin , Qingyang Wan , Zhuojun Li , Zhe He , Zhoutong Ye , Yu Mei , Chang Liu , Weinan Shi , Yuanchun Shi

Automatic speech synthesis is a challenging task that is becoming increasingly important as edge devices begin to interact with users through speech. Typical text-to-speech pipelines include a vocoder, which translates intermediate audio…

Sound · Computer Science 2020-01-17 Bohan Zhai , Tianren Gao , Flora Xue , Daniel Rothchild , Bichen Wu , Joseph E. Gonzalez , Kurt Keutzer

Overlapping Speech Detection (OSD) aims to identify regions where multiple speakers overlap in a conversation, a critical challenge in multi-party speech processing. This work proposes a speaker-aware progressive OSD model that leverages a…

Sound · Computer Science 2025-05-30 Zhaokai Sun , Li Zhang , Qing Wang , Pan Zhou , Lei Xie

Recent breakthroughs in zero-shot voice synthesis have enabled imitating a speaker's voice using just a few seconds of recording while maintaining a high level of realism. Alongside its potential benefits, this powerful technology…

Sound · Computer Science 2024-01-09 Guangyu Chen , Yu Wu , Shujie Liu , Tao Liu , Xiaoyong Du , Furu Wei

Speaker identification in multilingual settings presents unique challenges, particularly when conventional models are predominantly trained on English data. In this paper, we propose WSI (Whisper Speaker Identification), a framework that…

Sound · Computer Science 2025-03-14 Jakaria Islam Emon , Md Abu Salek , Kazi Tamanna Alam

The emergence of voice-assistant devices ushers in delightful user experiences not just on the smart home front, but also in diverse educational environments from classrooms to personalized-learning/tutoring. However, the use of voice as an…

Audio and Speech Processing · Electrical Eng. & Systems 2021-04-23 Mohammad Niknazar , Aditya Vempaty , Ravi Kokku

Gesture tracking technology provides users with a hands free interactive experience without the need to hold or touch devices. However, current gesture tracking research has primarily focused on tracking accuracy while neglecting issues of…

Cryptography and Security · Computer Science 2024-12-09 Bojun Zhang

We present HeadText, a hands-free technique on a smart earpiece for text entry by motion sensing. Users input text utilizing only 7 head gestures for key selection, word selection, word commitment and word cancelling tasks. Head gesture…

Human-Computer Interaction · Computer Science 2022-05-24 Songlin Xu , Guanjie Wang , Ziyuan Fang , Guangwei Zhang , Guangzhu Shang , Rongde Lu , Liqun He

Small footprint embedded devices require keyword spotters (KWS) with small model size and detection latency for enabling voice assistants. Such a keyword is often referred to as \textit{wake word} as it is used to wake up voice assistant…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-16 Christin Jose , Yuriy Mishchenko , Thibaud Senechal , Anish Shah , Alex Escott , Shiv Vitaladevuni

Achieving super-human performance in recognizing human speech has been a goal for several decades, as researchers have worked on increasingly challenging tasks. In the 1990's it was discovered, that conversational speech between two humans…

Computer Vision and Pattern Recognition · Computer Science 2021-07-28 Thai-Son Nguyen , Sebastian Stueker , Alex Waibel

With the rise of voice-enabled technologies, loudspeaker playback has become widespread, posing increasing risks to speech privacy. Traditional eavesdropping methods often require invasive access or line-of-sight, limiting their…

Sound · Computer Science 2025-11-11 Dachao Han , Teng Huang , Han Ding , Cui Zhao , Fei Wang , Ge Wang , Wei Xi

Gestures performed accompanying the voice are essential for voice interaction to convey complementary semantics for interaction purposes such as wake-up state and input modality. In this paper, we investigated voice-accompanying…

Human-Computer Interaction · Computer Science 2023-03-21 Zisu Li , Cheng Liang , Yuntao Wang , Yue Qin , Chun Yu , Yukang Yan , Mingming Fan , Yuanchun Shi

Reliability on cloud providers for ASR inference to support child-centered voice-based applications is becoming challenging due to regulatory and privacy challenges. Motivated by a privacy-preserving design, this study aims to develop a…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-22 Satwik Dutta , Shruthigna Chandupatla , John Hansen

The rapid advancement of AI has enabled highly realistic speech synthesis and voice cloning, posing serious risks to voice authentication, smart assistants, and telecom security. While most prior work frames spoof detection as a binary…

Sound · Computer Science 2025-09-10 Bin Hu , Kunyang Huang , Daehan Kwak , Meng Xu , Kuan Huang

Deepfake speech utterances can be forged by replacing one or more words in a bona fide utterance with semantically different words synthesized with speech-generative models. While a dedicated synthetic word detector could be developed, we…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-03 Hoan My Tran , Xin Wang , Wanying Ge , Xuechen Liu , Junichi Yamagishi

Wireless sensing systems, particularly those using mmWave technology, offer distinct advantages over traditional vision-based approaches, such as enhanced privacy and effectiveness in poor lighting conditions. These systems, leveraging FMCW…

Computer Vision and Pattern Recognition · Computer Science 2025-03-28 Teng Huang , Han Ding , Wenxin Sun , Cui Zhao , Ge Wang , Fei Wang , Kun Zhao , Zhi Wang , Wei Xi

Tracking beats of singing voices without the presence of musical accompaniment can find many applications in music production, automatic song arrangement, and social media interaction. Its main challenge is the lack of strong rhythmic and…

Audio and Speech Processing · Electrical Eng. & Systems 2022-09-01 Mojtaba Heydari , Zhiyao Duan

In recent years, deep learning techniques have been used to develop sign language recognition systems, potentially serving as a communication tool for millions of hearing-impaired individuals worldwide. However, there are inherent…

Computer Vision and Pattern Recognition · Computer Science 2024-08-15 Alvaro Leandro Cavalcante Carneiro , Denis Henrique Pinheiro Salvadeo , Lucas de Brito Silva
‹ Prev 1 4 5 6 7 8 10 Next ›