中文
相关论文

相关论文: Lip-Siri: Contactless Open-Sentence Silent Speech …

200 篇论文

Deepspeech was very useful for development IoT devices that need voice recognition. One of the voice recognition systems is deepspeech from Mozilla. Deepspeech is an open-source voice recognition that was using a neural network to convert…

音频与语音处理 · 电气工程与系统科学 2020-03-02 Muhammad Hafidh Firmansyah , Anand Paul , Deblina Bhattacharya , Gul Malik Urfa

In the coming decade, artificial intelligence systems will continue to improve and revolutionise every industry and facet of human life. Designing effective, seamless and symbiotic communication paradigms between humans and AI agents is…

人机交互 · 计算机科学 2024-08-13 Suyi Zhang , Ekram Alam , Jack Baber , Francesca Bianco , Edward Turner , Maysam Chamanzar , Hamid Dehghani

We present HighSync, an end-to-end diffusion-based framework for high-fidelity lip synchronization that generates photorealistic talking-face videos aligned with arbitrary input audio. Existing approaches consistently struggle to reconcile…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Saeed Firouzi Daghigh , Majid Iranpour Mobarekeh , Mostafa Alavi , Mehdi Bagheri

The ability to speak is an inherent part of human nature and fundamental to our existence as a social species. Unfortunately, this ability can be restricted in certain situations, such as for individuals who have lost their voice or in…

人机交互 · 计算机科学 2025-08-26 Zhao Ren , Simon Pistrosch , Buket Coşkun , Kevin Scheck , Anton Batliner , Björn W. Schuller , Tanja Schultz

Several sensing techniques have been proposed for silent speech recognition (SSR); however, many of these methods require invasive processes or sensor attachment to the skin using adhesive tape or glue, rendering them unsuitable for…

音频与语音处理 · 电气工程与系统科学 2023-12-22 Sunghwa Lee , Younghoon Shin , Myungjong Kim , Jiwon Seo

The contact-free sensing nature of Wi-Fi has been leveraged to achieve privacy breaches, yet existing attacks relying on Wi-Fi CSI (channel state information) demand hacking Wi-Fi hardware to obtain desired CSIs. Since such hacking has…

密码学与安全 · 计算机科学 2023-09-08 Jingyang Hu , Hongbo Wang , Tianyue Zheng , Jingzhi Hu , Zhe Chen , Hongbo Jiang , Jun Luo

Existing lip-sync deepfake detectors rely on pixel artifacts or audio-visual correspondence, and both fail under generator or language shift because the features they learn are tied to the training distribution. We take a different…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Hao Chen , Junnan Xu

Speaking aloud to a wearable AR assistant in public can be socially awkward, and re-articulating the same requests every day creates unnecessary effort. We present SpeechLess, a wearable AR assistant that introduces a speech-based intent…

人机交互 · 计算机科学 2026-04-14 Yoonsang Kim , Devshree Jadeja , Divyansh Pradhan , Yalong Yang , Arie Kaufman

This paper investigates the in-context learning abilities of the Whisper automatic speech recognition (ASR) models released by OpenAI. A novel speech-based in-context learning (SICL) approach is proposed for test-time adaptation, which can…

音频与语音处理 · 电气工程与系统科学 2024-03-21 Siyin Wang , Chao-Han Huck Yang , Ji Wu , Chao Zhang

Robust speech recognition systems rely on cloud service providers for inference. It needs to ensure that an untrustworthy provider cannot deduce the sensitive content in speech. Sanitization can be done on speech content keeping in mind…

音频与语音处理 · 电气工程与系统科学 2025-12-02 Afsara Benazir , Felix Xiaozhu Lin

Human lip-reading is a challenging task. It requires not only knowledge of underlying language but also visual clues to predict spoken words. Experts need certain level of experience and understanding of visual expressions learning to…

计算机视觉与模式识别 · 计算机科学 2018-02-16 M Faisal , Sanaullah Manzoor

Real-time Automatic Speech Recognition (ASR) is a fundamental building block for many commercial applications of ML, including live captioning, dictation, meeting transcriptions, and medical scribes. Accuracy and latency are the most…

声音 · 计算机科学 2025-07-16 Atila Orhon , Arda Okan , Berkin Durmus , Zach Nagengast , Eduardo Pacheco

In recent years, developing a speech understanding system that classifies a waveform to structured data, such as intents and slots, without first transcribing the speech to text has emerged as an interesting research problem. This work…

计算机视觉与模式识别 · 计算机科学 2020-11-11 Mohamed Mhiri , Samuel Myer , Vikrant Singh Tomar

Privacy-preserving semantic understanding of human activities is important for indoor sensing, yet existing Wi-Fi CSI-based systems mainly focus on pose estimation or predefined action classification rather than fine-grained language…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Tzu-Ti Wei , Chu-Yu Huang , Yu-Chee Tseng , Jen-Jee Chen

Audio-driven lip sync has recently drawn significant attention due to its widespread application in the multimedia domain. Individuals exhibit distinct lip shapes when speaking the same utterance, attributed to the unique speaking styles of…

计算机视觉与模式识别 · 计算机科学 2025-06-19 Weizhi Zhong , Jichang Li , Yinqi Cai , Ming Li , Feng Gao , Liang Lin , Guanbin Li

There is growing interest in enabling wireless sensing systems to interpret human motion from unsegmented wireless signals; however, existing CSI-based applications rely heavily on accurate signal segmentation and predefined action labels,…

网络与互联网体系结构 · 计算机科学 2026-05-15 Mahmuda Keya , Sneh Pillai , Jiawei Yuan , Kai Zeng , Long Jiao

Lip-to-Speech (Lip2Speech) synthesis, which predicts corresponding speech from talking face images, has witnessed significant progress with various models and training strategies in a series of independent studies. However, existing studies…

多媒体 · 计算机科学 2023-05-25 Zheng-Yan Sheng , Yang Ai , Zhen-Hua Ling

Current front-ends for robust automatic speech recognition(ASR) include masking- and mapping-based deep learning approaches to speech enhancement. A recently proposed deep learning approach toa prioriSNR estimation, called DeepXi, was able…

音频与语音处理 · 电气工程与系统科学 2020-01-29 Aaron Nicolson , Kuldip K. Paliwal

Voice disorders affect millions of people worldwide. Surface electromyography-based Silent Speech Interfaces (sEMG-based SSIs) have been explored as a potential solution for decades. However, previous works were limited by small…

音频与语音处理 · 电气工程与系统科学 2023-08-15 Wenqiang Lai , Qihan Yang , Ye Mao , Endong Sun , Jiangnan Ye

Hearing-impaired individuals often face significant barriers in daily communication due to the inherent challenges of producing clear speech. To address this, we introduce the Omni-Model paradigm into assistive technology and present…