English
Related papers

Related papers: A large-scale multimodal dataset of human speech r…

200 papers

Compared with an extensive list of automotive radar datasets that support autonomous driving, indoor radar datasets are scarce at a smaller scale in the format of low-resolution radar point clouds and usually under an open-space single-room…

Computer Vision and Pattern Recognition · Computer Science 2024-07-18 M. Mahbubur Rahman , Ryoma Yataka , Sorachi Kato , Pu Perry Wang , Peizhao Li , Adriano Cardace , Petros Boufounos

Speech is considered as a multi-modal process where hearing and vision are two fundamentals pillars. In fact, several studies have demonstrated that the robustness of Automatic Speech Recognition systems can be improved when audio and…

Computer Vision and Pattern Recognition · Computer Science 2023-11-22 David Gimeno-Gómez , Carlos-D. Martínez-Hinarejos

The CHiME challenge series aims to advance robust automatic speech recognition (ASR) technology by promoting research at the interface of speech and language processing, signal processing , and machine learning. This paper introduces the…

Sound · Computer Science 2018-03-29 Jon Barker , Shinji Watanabe , Emmanuel Vincent , Jan Trmal

Millimeter wave (mmWave) based gesture recognition technology provides a good human computer interaction (HCI) experience. Prior works focus on the close-range gesture recognition, but fall short in range extension, i.e., they are unable to…

Human-Computer Interaction · Computer Science 2020-02-10 Yu Liu , Yuheng Wang , Haipeng Liu , Anfu Zhou , Jianhua Liu , Ning Yang

Lipreading, i.e. speech recognition from visual-only recordings of a speaker's face, can be achieved with a processing pipeline based solely on neural networks, yielding significantly better accuracy than conventional methods. Feed-forward…

Computer Vision and Pattern Recognition · Computer Science 2016-02-01 Michael Wand , Jan Koutník , Jürgen Schmidhuber

Audio-visual speech recognition (AVSR) gains increasing attention from researchers as an important part of human-computer interaction. However, the existing available Mandarin audio-visual datasets are limited and lack the depth…

Sound · Computer Science 2023-06-06 Jianrong Wang , Yuchen Huo , Li Liu , Tianyi Xu , Qi Li , Sen Li

In Speech Emotion Recognition (SER), emotional characteristics often appear in diverse forms of energy patterns in spectrograms. Typical attention neural network classifiers of SER are usually optimized on a fixed attention granularity. In…

Sound · Computer Science 2021-02-04 Mingke Xu , Fan Zhang , Xiaodong Cui , Wei Zhang

Continuous detection of human activities and presence is essential for developing a pervasive interactive smart space. Existing literature lacks robust wireless sensing mechanisms capable of continuously monitoring multiple users'…

Human-Computer Interaction · Computer Science 2023-09-22 Argha Sen , Anirban Das , Swadhin Pradhan , Sandip Chakraborty

The objective of this paper is speaker recognition under noisy and unconstrained conditions. We make two key contributions. First, we introduce a very large-scale audio-visual speaker recognition dataset collected from open-source media.…

Sound · Computer Science 2020-11-05 Joon Son Chung , Arsha Nagrani , Andrew Zisserman

This paper presents a Multi-modal Emotion Recognition (MER) system designed to enhance emotion recognition accuracy in challenging acoustic conditions. Our approach combines a modified and extended Hierarchical Token-semantic Audio…

Sound · Computer Science 2025-07-30 Ohad Cohen , Gershon Hazan , Sharon Gannot

Language Identification (LID) is a challenging task, especially when the input texts are short and noisy such as posts and statuses on social media or chat logs on gaming forums. The task has been tackled by either designing a feature set…

Computation and Language · Computer Science 2019-10-16 Duy Tin Vo , Richard Khoury

This paper introduces GigaSpeech, an evolving, multi-domain English speech recognition corpus with 10,000 hours of high quality labeled audio suitable for supervised training, and 40,000 hours of total audio suitable for semi-supervised and…

The sixth generation (6G) of mobile communication system is witnessing a new paradigm shift, i.e., integrated sensing-communication system. A comprehensive dataset is a prerequisite for 6G integrated sensing-communication research. This…

Signal Processing · Electrical Eng. & Systems 2023-06-27 Xiang Cheng , Ziwei Huang , Lu Bai , Haotian Zhang , Mingran Sun , Boxun Liu , Sijiang Li , Jianan Zhang , Minson Lee

Voice assistants have become an essential tool for people with various disabilities because they enable complex phone- or tablet-based interactions without the need for fine-grained motor control, such as with touchscreens. However, these…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-17 Colin Lea , Zifang Huang , Dhruv Jain , Lauren Tooley , Zeinab Liaghat , Shrinath Thelapurath , Leah Findlater , Jeffrey P. Bigham

Recently, the AI community has made significant strides in developing powerful foundation models, driven by large-scale multimodal datasets. However, for audio representation learning, existing datasets suffer from limitations in the…

Sound · Computer Science 2024-09-10 Luoyi Sun , Xuenan Xu , Mengyue Wu , Weidi Xie

This paper presents a complete hardware and software pipeline for real-time speech enhancement in noisy and reverberant conditions. The device consists of a microphone array and a camera mounted on eyeglasses, connected to an embedded…

Training data cleaning is a new application for generative model-based speech restoration (SR). This paper introduces Miipher-2, an SR model designed for million-hour scale data, for training data cleaning for large-scale generative models…

Low latency models are critical for real-time speech enhancement applications, such as hearing aids and hearables. However, the sub-millisecond latency space for resource-constrained hearables remains underexplored. We demonstrate speech…

Speech representation and modelling in high-dimensional spaces of acoustic waveforms, or a linear transformation thereof, is investigated with the aim of improving the robustness of automatic speech recognition to additive noise. The…

Computation and Language · Computer Science 2015-03-31 Matthew Ager , Zoran Cvetkovic , Peter Sollich

Large-scale in-the-wild speech datasets have become more prevalent in recent years due to increased interest in models that can learn useful features from unlabelled data for tasks such as speech recognition or synthesis. These datasets…

‹ Prev 1 4 5 6 7 8 10 Next ›