中文
相关论文

相关论文: A large-scale multimodal dataset of human speech r…

200 篇论文

Human lip-reading is a challenging task. It requires not only knowledge of underlying language but also visual clues to predict spoken words. Experts need certain level of experience and understanding of visual expressions learning to…

计算机视觉与模式识别 · 计算机科学 2018-02-16 M Faisal , Sanaullah Manzoor

Human activity recognition (HAR) is essential in healthcare, elder care, security, and human-computer interaction. The use of precise sensor data to identify activities passively and continuously makes HAR accessible and ubiquitous.…

人机交互 · 计算机科学 2024-08-01 Argha Sen , Anirban Das , Swadhin Pradhan , Sandip Chakraborty

We introduce a novel dataset for multi-robot activity recognition (MRAR) using two robotic arms integrating WiFi channel state information (CSI), video, and audio data. This multimodal dataset utilizes signals of opportunity, leveraging…

机器人学 · 计算机科学 2025-02-18 Kian Behzad , Rojin Zandi , Elaheh Motamedi , Hojjat Salehinejad , Milad Siami

Silent speech interfaces (SSI) has been an exciting area of recent interest. In this paper, we present a non-invasive silent speech interface that uses inaudible acoustic signals to capture people's lip movements when they speak. We exploit…

音频与语音处理 · 电气工程与系统科学 2020-11-24 Jian Luo , Jianzong Wang , Ning Cheng , Guilin Jiang , Jing Xiao

Lip motion reflects behavior characteristics of speakers, and thus can be used as a new kind of biometrics in speaker recognition. In the literature, lots of works used two-dimensional (2D) lip images to recognize speaker in a textdependent…

计算机视觉与模式识别 · 计算机科学 2020-10-14 Jianrong Wang , Tong Wu , Shanyu Wang , Mei Yu , Qiang Fang , Ju Zhang , Li Liu

Synthesising 3D facial motion from speech is a crucial problem manifesting in a multitude of applications such as computer games and movies. Recently proposed methods tackle this problem in controlled conditions of speech. In this paper, we…

计算机视觉与模式识别 · 计算机科学 2019-04-16 Panagiotis Tzirakis , Athanasios Papaioannou , Alexander Lattas , Michail Tarasiou , Björn Schuller , Stefanos Zafeiriou

Recently reported state-of-the-art results in visual speech recognition (VSR) often rely on increasingly large amounts of video data, while the publicly available transcribed video datasets are limited in size. In this paper, for the first…

Lip reading has witnessed unparalleled development in recent years thanks to deep learning and the availability of large-scale datasets. Despite the encouraging results achieved, the performance of lip reading, unfortunately, remains…

计算机视觉与模式识别 · 计算机科学 2019-11-27 Ya Zhao , Rui Xu , Xinchao Wang , Peng Hou , Haihong Tang , Mingli Song

This paper presents a comprehensive dataset intended to evaluate passive Human Activity Recognition (HAR) and localization techniques with measurements obtained from synchronized Radio-Frequency (RF) devices and vision-based sensors. The…

We present a Multi-Window Data Augmentation (MWA-SER) approach for speech emotion recognition. MWA-SER is a unimodal approach that focuses on two key concepts; designing the speech augmentation method and building the deep learning model to…

声音 · 计算机科学 2022-02-17 Sarala Padi , Dinesh Manocha , Ram D. Sriram

Interactions with virtual assistants typically start with a trigger phrase followed by a command. In this work, we explore the possibility of making these interactions more natural by eliminating the need for a trigger phrase. Our goal is…

The goal of this project is to develop a limited lip reading algorithm for a subset of the English language. We consider a scenario in which no audio information is available. The raw video is processed and the position of the lips in each…

计算机视觉与模式识别 · 计算机科学 2017-08-04 Jithin Donny George , Ronan Keane , Conor Zellmer

In recent years, significant progress has been made in automatic lip reading. But these methods require large-scale datasets that do not exist for many low-resource languages. In this paper, we have presented a new multipurpose audio-visual…

We present a novel approach to improve the performance of learning-based speech dereverberation using accurate synthetic datasets. Our approach is designed to recover the reverb-free signal from a reverberant speech signal. We show that…

音频与语音处理 · 电气工程与系统科学 2022-12-13 Rohith Aralikatti , Zhenyu Tang , Dinesh Manocha

The global aging population faces considerable challenges, particularly in communication, due to the prevalence of hearing and speech impairments. To address these, we introduce the AVE speech, a comprehensive multi-modal dataset for speech…

声音 · 计算机科学 2025-07-08 Dongliang Zhou , Yakun Zhang , Jinghan Wu , Xingyu Zhang , Liang Xie , Erwei Yin

Automatic speech recognition systems are part of people's daily lives, embedded in personal assistants and mobile phones, helping as a facilitator for human-machine interaction while allowing access to information in a practically intuitive…

声音 · 计算机科学 2021-10-05 Julio Cesar Duarte , Sérgio Colcher

The training of deep learning-based multichannel speech enhancement and source localization systems relies heavily on the simulation of room impulse response and multichannel diffuse noise, due to the lack of large-scale real-recorded…

声音 · 计算机科学 2024-10-02 Bing Yang , Changsheng Quan , Yabo Wang , Pengyu Wang , Yujie Yang , Ying Fang , Nian Shao , Hui Bu , Xin Xu , Xiaofei Li

Today's Automatic Speech Recognition systems only rely on acoustic signals and often don't perform well under noisy conditions. Performing multi-modal speech recognition - processing acoustic speech signals and lip-reading video…

计算机视觉与模式识别 · 计算机科学 2018-03-14 Matthijs Van keirsbilck , Bert Moons , Marian Verhelst

Large datasets as required for deep learning of lip reading do not exist in many languages. In this paper we present the dataset GLips (German Lips) consisting of 250,000 publicly available videos of the faces of speakers of the Hessian…

计算机视觉与模式识别 · 计算机科学 2022-07-12 Gerald Schwiebert , Cornelius Weber , Leyuan Qu , Henrique Siqueira , Stefan Wermter

For conversational large-vocabulary continuous speech recognition (LVCSR) tasks, up to about two thousand hours of audio is commonly used to train state of the art models. Collection of labeled conversational audio however, is prohibitively…

计算与语言 · 计算机科学 2017-05-30 Shane Walker , Morten Pedersen , Iroro Orife , Jason Flaks