中文
相关论文

相关论文: A large-scale multimodal dataset of human speech r…

200 篇论文

Silent speech interfaces (SSI) are being actively developed to assist individuals with communication impairments who have long suffered from daily hardships and a reduced quality of life. However, silent sentences are difficult to segment…

人机交互 · 计算机科学 2025-09-19 Yudong Xie , Zhifeng Han , Qinfan Xiao , Liwei Liang , Lu-Qi Tao , Tian-Ling Ren

Recent advances in speech synthesis and voice conversion have greatly improved the naturalness and authenticity of generated audio. Meanwhile, evolving encoding, compression, and transmission mechanisms on social media platforms further…

声音 · 计算机科学 2026-03-09 Daixian Li , Jun Xue , Yanzhen Ren , Zhuolin Yi , Yihuan Huang , Guanxiang Feng , Yi Chai

This paper delineates AISHELL-5, the first open-source in-car multi-channel multi-speaker Mandarin automatic speech recognition (ASR) dataset. AISHLL-5 includes two parts: (1) over 100 hours of multi-channel speech data recorded in an…

声音 · 计算机科学 2025-05-30 Yuhang Dai , He Wang , Xingchen Li , Zihan Zhang , Shuiyuan Wang , Lei Xie , Xin Xu , Hongxiao Guo , Shaoji Zhang , Hui Bu , Wei Chen

We introduce MMIS, a novel dataset designed to advance MultiModal Interior Scene generation and recognition. MMIS consists of nearly 160,000 images. Each image within the dataset is accompanied by its corresponding textual description and…

计算机视觉与模式识别 · 计算机科学 2024-07-09 Hozaifa Kassab , Ahmed Mahmoud , Mohamed Bahaa , Ammar Mohamed , Ali Hamdi

This paper introduces Multilingual LibriSpeech (MLS) dataset, a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages, including about 44.5K hours of…

音频与语音处理 · 电气工程与系统科学 2020-12-22 Vineel Pratap , Qiantong Xu , Anuroop Sriram , Gabriel Synnaeve , Ronan Collobert

This study presents the development and testing of a conversational speech system designed for robots to detect speech biomarkers indicative of cognitive impairments in people living with dementia (PLwD). The system integrates a backend…

Machine learning tools are finding interesting applications in millimeter wave (mmWave) and massive MIMO systems. This is mainly thanks to their powerful capabilities in learning unknown models and tackling hard optimization problems. To…

信息论 · 计算机科学 2019-02-19 Ahmed Alkhateeb

To let the state-of-the-art end-to-end ASR model enjoy data efficiency, as well as much more unpaired text data by multi-modal training, one needs to address two problems: 1) the synchronicity of feature sampling rates between speech and…

音频与语音处理 · 电气工程与系统科学 2022-11-02 Yuhang Yang , Haihua Xu , Hao Huang , Eng Siong Chng , Sheng Li

Detecting speech from biosignals is gaining increasing attention due to the potential to develop human-computer interfaces that are noise-robust, privacy-preserving, and scalable for both clinical applications and daily use. However, most…

In this project, we worked on speech recognition, specifically predicting individual words based on both the video frames and audio. Empowered by convolutional neural networks, the recent speech recognition and lip reading models are…

计算机视觉与模式识别 · 计算机科学 2018-12-27 Devesh Walawalkar , Yihui He , Rohit Pillai

Natural and efficient interaction remains a critical challenge for virtual reality and augmented reality (VR/AR) systems. Vision-based gesture recognition suffers from high computational cost, sensitivity to lighting conditions, and privacy…

人机交互 · 计算机科学 2025-11-11 Xijie Zhang , Fengliang He , Hong-Ning Dai

New techniques in cross-layer wireless networks are building demand for ubiquitous channel sounding, that is, the capability to measure channel impulse response (CIR) with any standard wireless network and node. Towards that goal, we…

其他计算机科学 · 计算机科学 2013-12-10 Mohammad H. Firooz , Dustin Maas , Junxing Zhang , Neal Patwari , Sneha K. Kasera

The aim of this work is to investigate the impact of crossmodal self-supervised pre-training for speech reconstruction (video-to-audio) by leveraging the natural co-occurrence of audio and visual streams in videos. We propose LipSound2…

声音 · 计算机科学 2022-09-13 Leyuan Qu , Cornelius Weber , Stefan Wermter

Recent research on word-level confidence estimation for speech recognition systems has primarily focused on lightweight models known as Confidence Estimation Modules (CEMs), which rely on hand-engineered features derived from Automatic…

音频与语音处理 · 电气工程与系统科学 2025-02-20 Vaibhav Aggarwal , Shabari S Nair , Yash Verma , Yash Jogi

Recent years have seen a surge in finding association between faces and voices within a cross-modal biometric application along with speaker recognition. Inspired from this, we introduce a challenging task in establishing association…

计算机视觉与模式识别 · 计算机科学 2021-04-23 Muhammad Saad Saeed , Shah Nawaz , Pietro Morerio , Arif Mahmood , Ignazio Gallo , Muhammad Haroon Yousaf , Alessio Del Bue

A new impulse response (IR) dataset called "MeshRIR" is introduced. Currently available datasets usually include IRs at an array of microphones from several source positions under various room conditions, which are basically designed for…

音频与语音处理 · 电气工程与系统科学 2021-07-26 Shoichi Koyama , Tomoya Nishida , Keisuke Kimura , Takumi Abe , Natsuki Ueno , Jesper Brunnström

This paper presents a new multimodal interventional radiology dataset, called PoCaP (Port Catheter Placement) Corpus. This corpus consists of speech and audio signals in German, X-ray images, and system commands collected from 31 PoCaP…

Lip-to-speech involves generating a natural-sounding speech synchronized with a soundless video of a person talking. Despite recent advances, current methods still cannot produce high-quality speech with high levels of intelligibility for…

音频与语音处理 · 电气工程与系统科学 2024-03-29 Yochai Yemini , Aviv Shamsian , Lior Bracha , Sharon Gannot , Ethan Fetaya

Visual cues, like lip motion, have been shown to improve the performance of Automatic Speech Recognition (ASR) systems in noisy environments. We propose LipGER (Lip Motion aided Generative Error Correction), a novel framework for leveraging…

音频与语音处理 · 电气工程与系统科学 2024-06-10 Sreyan Ghosh , Sonal Kumar , Ashish Seth , Purva Chiniya , Utkarsh Tyagi , Ramani Duraiswami , Dinesh Manocha

Visual speaker recognition based on lip motion offers a silent, hands-free, and behavior-driven biometric solution that remains effective even when acoustic cues are unavailable. Compared to traditional methods that rely heavily on…

计算机视觉与模式识别 · 计算机科学 2026-04-20 Junguang Yao , Wenye Liu , Stjepan Picek , Yue Zheng