中文
相关论文

相关论文: Towards Lipreading Sentences with Active Appearanc…

200 篇论文

Multi-task learning (MTL) and attention mechanism have been proven to effectively extract robust acoustic features for various speech-related tasks in noisy environments. In this study, we propose an attention-based MTL (ATM) approach that…

音频与语音处理 · 电气工程与系统科学 2021-02-23 Chiang-Jen Peng , Yun-Ju Chan , Cheng Yu , Syu-Siang Wang , Yu Tsao , Tai-Shih Chi

Audio deepfake detection (ADD) is crucial to combat the misuse of speech synthesized from generative AI models. Existing ADD models suffer from generalization issues, with a large performance discrepancy between in-domain and out-of-domain…

声音 · 计算机科学 2024-07-29 Yi Zhu , Surya Koppisetti , Trang Tran , Gaurav Bharaj

Speaker Recognition is a challenging task with essential applications such as authentication, automation, and security. The SincNet is a new deep learning based model which has produced promising results to tackle the mentioned task. To…

音频与语音处理 · 电气工程与系统科学 2019-10-15 João Antônio Chagas Nunes , David Macêdo , Cleber Zanchettin

Recently, the AI community has made significant strides in developing powerful foundation models, driven by large-scale multimodal datasets. However, for audio representation learning, existing datasets suffer from limitations in the…

声音 · 计算机科学 2024-09-10 Luoyi Sun , Xuenan Xu , Mengyue Wu , Weidi Xie

Audiovisual active speaker detection (ASD) is conventionally performed by modelling the temporal synchronisation of acoustic and visual speech cues. In egocentric recordings, however, the efficacy of synchronisation-based methods is…

多媒体 · 计算机科学 2025-06-24 Jason Clarke , Yoshihiko Gotoh , Stefan Goetze

The audio-visual speech fusion strategy AV Align has shown significant performance improvements in audio-visual speech recognition (AVSR) on the challenging LRS2 dataset. Performance improvements range between 7% and 30% depending on the…

音频与语音处理 · 电气工程与系统科学 2020-05-20 George Sterpu , Christian Saam , Naomi Harte

We apply topological data analysis (TDA) to speech classification problems and to the introspection of a pretrained speech model, HuBERT. To this end, we introduce a number of topological and algebraic features derived from Transformer…

Attention-based encoder-decoder architectures such as Listen, Attend, and Spell (LAS), subsume the acoustic, pronunciation and language model components of a traditional automatic speech recognition (ASR) system into a single neural…

Voice biometric tasks, such as age estimation require modeling the often complex relationship between voice features and the biometric variable. While deep learning models can handle such complexity, they typically require large amounts of…

机器学习 · 计算机科学 2025-01-29 Dareen Alharthi , Mahsa Zamani , Bhiksha Raj , Rita Singh

In this work, we present a hybrid CTC/Attention model based on a ResNet-18 and Convolution-augmented transformer (Conformer), that can be trained in an end-to-end manner. In particular, the audio and visual encoders learn to extract…

计算机视觉与模式识别 · 计算机科学 2021-02-15 Pingchuan Ma , Stavros Petridis , Maja Pantic

Diffusion Transformers (DiT) trained with flow matching in a VAE latent space have unified visual generation across images and videos. A natural next step toward a single architecture for both generation (visual synthesis) and understanding…

Active speaker detection (ASD) seeks to detect who is speaking in a visual scene of one or more speakers. The successful ASD depends on accurate interpretation of short-term and long-term audio and visual information, as well as…

音频与语音处理 · 电气工程与系统科学 2021-07-27 Ruijie Tao , Zexu Pan , Rohan Kumar Das , Xinyuan Qian , Mike Zheng Shou , Haizhou Li

Vision-based deep learning models can be promising for speech-and-hearing-impaired and secret communications. While such non-verbal communications are primarily investigated with hand-gestures and facial expressions, no research endeavour…

计算机视觉与模式识别 · 计算机科学 2022-01-19 Abtahi Ishmam , Mahmudul Hasan , Md. Saif Hassan Onim , Koushik Roy , Md. Akiful Haque Akif , Hussain Nyeem

Interactions involving children span a wide range of important domains from learning to clinical diagnostic and therapeutic contexts. Automated analyses of such interactions are motivated by the need to seek accurate insights and offer…

音频与语音处理 · 电气工程与系统科学 2025-06-13 Anfeng Xu , Kevin Huang , Tiantian Feng , Helen Tager-Flusberg , Shrikanth Narayanan

Speech recognition is the technology that enables machines to interpret and process human speech, converting spoken language into text or commands. This technology is essential for applications such as virtual assistants, transcription…

音频与语音处理 · 电气工程与系统科学 2025-01-09 Xinyu Wang , Haotian Jiang , Haolin Huang , Yu Fang , Mengjie Xu , Qian Wang

Lipreading or visually recognizing speech from the mouth movements of a speaker is a challenging and mentally taxing task. Unfortunately, multiple medical conditions force people to depend on this skill in their day-to-day lives for…

计算机视觉与模式识别 · 计算机科学 2021-11-04 Bipasha Sen , Aditya Agarwal , Rudrabha Mukhopadhyay , Vinay Namboodiri , C V Jawahar

Vision model have gained increasing attention due to their simplicity and efficiency in Scene Text Recognition (STR) task. However, due to lacking the perception of linguistic knowledge and information, recent vision models suffer from two…

计算机视觉与模式识别 · 计算机科学 2023-05-11 Boqiang Zhang , Hongtao Xie , Yuxin Wang , Jianjun Xu , Yongdong Zhang

Learning an effective speaker representation is crucial for achieving reliable performance in speaker verification tasks. Speech signals are high-dimensional, long, and variable-length sequences containing diverse information at each…

音频与语音处理 · 电气工程与系统科学 2023-08-25 Wei Xia , John H. L. Hansen

In this article, we introduce a novel problem of audio-visual autism behavior recognition, which includes social behavior recognition, an essential aspect previously omitted in AI-assisted autism screening research. We define the task at…

Articulatory distinctive features, as well as phonetic transcription, play important role in speech-related tasks: computer-assisted pronunciation training, text-to-speech conversion (TTS), studying speech production mechanisms, speech…

音频与语音处理 · 电气工程与系统科学 2019-07-04 Ievgen Karaulov , Dmytro Tkanov