中文
相关论文

相关论文: A large-scale multimodal dataset of human speech r…

200 篇论文

This work presents a scalable solution to open-vocabulary visual speech recognition. To achieve this, we constructed the largest existing visual speech recognition dataset, consisting of pairs of text and video clips of faces speaking…

This paper introduces mmWave-Whisper, a system that demonstrates the feasibility of full-corpus automated speech recognition (ASR) on phone calls eavesdropped remotely using off-the-shelf frequency modulated continuous wave (FMCW)…

声音 · 计算机科学 2024-10-24 Suryoday Basak , Abhijeeth Padarthi , Mahanth Gowda

This paper presents Multimodal-Wireless, a large-scale open-source dataset for multimodal sensing and communication research. The dataset is generated through an integrated and customizable data pipeline built upon the CARLA simulator and…

信号处理 · 电气工程与系统科学 2026-02-12 Tianhao Mao , Le Liang , Jie Yang , Hao Ye , Shi Jin , Geoffrey Ye Li

Several sensing techniques have been proposed for silent speech recognition (SSR); however, many of these methods require invasive processes or sensor attachment to the skin using adhesive tape or glue, rendering them unsuitable for…

音频与语音处理 · 电气工程与系统科学 2023-12-22 Sunghwa Lee , Younghoon Shin , Myungjong Kim , Jiwon Seo

Multimodal vital sign monitoring and speech detection hold significant importance in medical health, public safety, and other fields. This study proposes a broadband tunable microwave photonic radar system that can simultaneously monitor…

光学 · 物理学 2025-12-29 Lei Gao , Dingding Liang , Jiawei Gao , Chulun Lin , Zhiqiang Huang , Taixia Shi , Yang Chen

Silent speech interfaces (SSIs) enable silent interaction in noise-sensitive or privacy-sensitive settings. However, existing SSIs face practical deployment trade-offs among privacy, user experience, and energy consumption, and most remain…

人机交互 · 计算机科学 2026-01-27 Ye Tian , Haohua Du , Chao Gu , Junyang Zhang , Shanyue Wang , Hao Zhou , Jiahui Hou , Xiang-Yang Li

Millimeter Wave (mmWave) radar has emerged as a promising modality for speech sensing, offering advantages over traditional microphones. Prior works have demonstrated that radar captures motion signals related to vocal vibrations, but there…

音频与语音处理 · 电气工程与系统科学 2025-03-21 Isabella Lenz , Yu Rong , Daniel Bliss , Julie Liss , Visar Berisha

This paper presents DISC, a dataset of millimeter-wave channel impulse response measurements for integrated human activity sensing and communication. This is the first dataset collected with a software-defined radio testbed that transmits…

信号处理 · 电气工程与系统科学 2025-02-13 Jacopo Pegoraro , Pablo Saucedo , Jesus Omar Lacruz , Michele Rossi , Joerg Widmer

Silent speech recognition (SSR) is a technology that recognizes speech content from non-acoustic speech-related biosignals. This paper utilizes an attention-enhanced temporal convolutional network architecture for contactless IR-UWB…

音频与语音处理 · 电气工程与系统科学 2025-10-01 Sunghwa Lee , Jaewon Yu

The goal of this work is to recognise phrases and sentences being spoken by a talking face, with or without the audio. Unlike previous works that have focussed on recognising a limited number of words or phrases, we tackle lip reading as an…

计算机视觉与模式识别 · 计算机科学 2018-12-27 Triantafyllos Afouras , Joon Son Chung , Andrew Senior , Oriol Vinyals , Andrew Zisserman

Recognizing speaking in humans is a central task towards understanding social interactions. Ideally, speaking would be detected from individual voice recordings, as done previously for meeting scenarios. However, individual voice recordings…

计算机视觉与模式识别 · 计算机科学 2024-03-05 Jose Vargas Quiros , Chirag Raman , Stephanie Tan , Ekin Gedik , Laura Cabrera-Quiros , Hayley Hung

This paper introduces BIRD, the Big Impulse Response Dataset. This open dataset consists of 100,000 multichannel room impulse responses (RIRs) generated from simulations using the Image Method, making it the largest multichannel open…

We present the Geometric-Wave Acoustic (GWA) dataset, a large-scale audio dataset of about 2 million synthetic room impulse responses (IRs) and their corresponding detailed geometric and simulation configurations. Our dataset samples…

声音 · 计算机科学 2022-06-22 Zhenyu Tang , Rohith Aralikatti , Anton Ratnarajah , Dinesh Manocha

The audio data is increasing day by day throughout the globe with the increase of telephonic conversations, video conferences and voice messages. This research provides a mechanism for identifying a speaker in an audio file, based on the…

声音 · 计算机科学 2022-05-31 Syeda Rabia Arshad , Syed Mujtaba Haider , Abdul Basit Mughal

Recent research has demonstrated the complementary nature of camera-based and inertial data for modeling human gestures, activities, and sentiment. Yet, despite its growing importance for environmental sensing as well as the advance of…

数据库 · 计算机科学 2025-11-11 Si Zuo , Yuqing Song , Sahar Golipoor , Ying Liu , Xujun Ma , Stephan Sigg

Automatic sign language recognition (SLR) has become a key enabler of inclusive human-computer interaction, fostering seamless communication between deaf individuals and hearing communities. Despite significant advances in multimodal…

人机交互 · 计算机科学 2026-05-08 Xiaofang Xiao , Guangchao Li , Guangrong Zhao , Qi Lin , Wen Ma , Hongkai Wen , Yanxiang Wang , Yiran Shen

Real-time magnetic resonance imaging (RT-MRI) of human speech production is enabling significant advances in speech science, linguistics, bio-inspired speech technology development, and clinical applications. Easy access to RT-MRI is…

In the last years, several machine learning-based techniques have been proposed to monitor human movements from Wi-Fi channel readings. However, the development of domain-adaptive algorithms that robustly work across different environments…

信号处理 · 电气工程与系统科学 2024-08-27 Francesca Meneghello , Nicolò Dal Fabbro , Domenico Garlisi , Ilenia Tinnirello , Michele Rossi

The volumetric representation of human interactions is one of the fundamental domains in the development of immersive media productions and telecommunication applications. Particularly in the context of the rapid advancement of Extended…

计算机视觉与模式识别 · 计算机科学 2024-02-15 Fatemeh Ghorbani Lohesara , Davi Rabbouni Freitas , Christine Guillemot , Karen Eguiazarian , Sebastian Knorr

In this work, we present TalkCuts, a large-scale dataset designed to facilitate the study of multi-shot human speech video generation. Unlike existing datasets that focus on single-shot, static viewpoints, TalkCuts offers 164k clips…

计算机视觉与模式识别 · 计算机科学 2025-10-14 Jiaben Chen , Zixin Wang , Ailing Zeng , Yang Fu , Xueyang Yu , Siyuan Cen , Julian Tanke , Yihang Chen , Koichi Saito , Yuki Mitsufuji , Chuang Gan
‹ 上一页 1 2 3 10 下一页 ›