English
Related papers

Related papers: A large-scale multimodal dataset of human speech r…

200 papers

This work presents a scalable solution to open-vocabulary visual speech recognition. To achieve this, we constructed the largest existing visual speech recognition dataset, consisting of pairs of text and video clips of faces speaking…

This paper introduces mmWave-Whisper, a system that demonstrates the feasibility of full-corpus automated speech recognition (ASR) on phone calls eavesdropped remotely using off-the-shelf frequency modulated continuous wave (FMCW)…

Sound · Computer Science 2024-10-24 Suryoday Basak , Abhijeeth Padarthi , Mahanth Gowda

This paper presents Multimodal-Wireless, a large-scale open-source dataset for multimodal sensing and communication research. The dataset is generated through an integrated and customizable data pipeline built upon the CARLA simulator and…

Signal Processing · Electrical Eng. & Systems 2026-02-12 Tianhao Mao , Le Liang , Jie Yang , Hao Ye , Shi Jin , Geoffrey Ye Li

Several sensing techniques have been proposed for silent speech recognition (SSR); however, many of these methods require invasive processes or sensor attachment to the skin using adhesive tape or glue, rendering them unsuitable for…

Audio and Speech Processing · Electrical Eng. & Systems 2023-12-22 Sunghwa Lee , Younghoon Shin , Myungjong Kim , Jiwon Seo

Multimodal vital sign monitoring and speech detection hold significant importance in medical health, public safety, and other fields. This study proposes a broadband tunable microwave photonic radar system that can simultaneously monitor…

Silent speech interfaces (SSIs) enable silent interaction in noise-sensitive or privacy-sensitive settings. However, existing SSIs face practical deployment trade-offs among privacy, user experience, and energy consumption, and most remain…

Human-Computer Interaction · Computer Science 2026-01-27 Ye Tian , Haohua Du , Chao Gu , Junyang Zhang , Shanyue Wang , Hao Zhou , Jiahui Hou , Xiang-Yang Li

Millimeter Wave (mmWave) radar has emerged as a promising modality for speech sensing, offering advantages over traditional microphones. Prior works have demonstrated that radar captures motion signals related to vocal vibrations, but there…

Audio and Speech Processing · Electrical Eng. & Systems 2025-03-21 Isabella Lenz , Yu Rong , Daniel Bliss , Julie Liss , Visar Berisha

This paper presents DISC, a dataset of millimeter-wave channel impulse response measurements for integrated human activity sensing and communication. This is the first dataset collected with a software-defined radio testbed that transmits…

Signal Processing · Electrical Eng. & Systems 2025-02-13 Jacopo Pegoraro , Pablo Saucedo , Jesus Omar Lacruz , Michele Rossi , Joerg Widmer

Silent speech recognition (SSR) is a technology that recognizes speech content from non-acoustic speech-related biosignals. This paper utilizes an attention-enhanced temporal convolutional network architecture for contactless IR-UWB…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-01 Sunghwa Lee , Jaewon Yu

The goal of this work is to recognise phrases and sentences being spoken by a talking face, with or without the audio. Unlike previous works that have focussed on recognising a limited number of words or phrases, we tackle lip reading as an…

Computer Vision and Pattern Recognition · Computer Science 2018-12-27 Triantafyllos Afouras , Joon Son Chung , Andrew Senior , Oriol Vinyals , Andrew Zisserman

Recognizing speaking in humans is a central task towards understanding social interactions. Ideally, speaking would be detected from individual voice recordings, as done previously for meeting scenarios. However, individual voice recordings…

Computer Vision and Pattern Recognition · Computer Science 2024-03-05 Jose Vargas Quiros , Chirag Raman , Stephanie Tan , Ekin Gedik , Laura Cabrera-Quiros , Hayley Hung

This paper introduces BIRD, the Big Impulse Response Dataset. This open dataset consists of 100,000 multichannel room impulse responses (RIRs) generated from simulations using the Image Method, making it the largest multichannel open…

We present the Geometric-Wave Acoustic (GWA) dataset, a large-scale audio dataset of about 2 million synthetic room impulse responses (IRs) and their corresponding detailed geometric and simulation configurations. Our dataset samples…

Sound · Computer Science 2022-06-22 Zhenyu Tang , Rohith Aralikatti , Anton Ratnarajah , Dinesh Manocha

The audio data is increasing day by day throughout the globe with the increase of telephonic conversations, video conferences and voice messages. This research provides a mechanism for identifying a speaker in an audio file, based on the…

Sound · Computer Science 2022-05-31 Syeda Rabia Arshad , Syed Mujtaba Haider , Abdul Basit Mughal

Recent research has demonstrated the complementary nature of camera-based and inertial data for modeling human gestures, activities, and sentiment. Yet, despite its growing importance for environmental sensing as well as the advance of…

Databases · Computer Science 2025-11-11 Si Zuo , Yuqing Song , Sahar Golipoor , Ying Liu , Xujun Ma , Stephan Sigg

Automatic sign language recognition (SLR) has become a key enabler of inclusive human-computer interaction, fostering seamless communication between deaf individuals and hearing communities. Despite significant advances in multimodal…

Human-Computer Interaction · Computer Science 2026-05-08 Xiaofang Xiao , Guangchao Li , Guangrong Zhao , Qi Lin , Wen Ma , Hongkai Wen , Yanxiang Wang , Yiran Shen

Real-time magnetic resonance imaging (RT-MRI) of human speech production is enabling significant advances in speech science, linguistics, bio-inspired speech technology development, and clinical applications. Easy access to RT-MRI is…

In the last years, several machine learning-based techniques have been proposed to monitor human movements from Wi-Fi channel readings. However, the development of domain-adaptive algorithms that robustly work across different environments…

Signal Processing · Electrical Eng. & Systems 2024-08-27 Francesca Meneghello , Nicolò Dal Fabbro , Domenico Garlisi , Ilenia Tinnirello , Michele Rossi

The volumetric representation of human interactions is one of the fundamental domains in the development of immersive media productions and telecommunication applications. Particularly in the context of the rapid advancement of Extended…

Computer Vision and Pattern Recognition · Computer Science 2024-02-15 Fatemeh Ghorbani Lohesara , Davi Rabbouni Freitas , Christine Guillemot , Karen Eguiazarian , Sebastian Knorr

In this work, we present TalkCuts, a large-scale dataset designed to facilitate the study of multi-shot human speech video generation. Unlike existing datasets that focus on single-shot, static viewpoints, TalkCuts offers 164k clips…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Jiaben Chen , Zixin Wang , Ailing Zeng , Yang Fu , Xueyang Yu , Siyuan Cen , Julian Tanke , Yihang Chen , Koichi Saito , Yuki Mitsufuji , Chuang Gan
‹ Prev 1 2 3 10 Next ›