中文
相关论文

相关论文: Towards Automatic Speech Identification from Vocal…

200 篇论文

It is increasingly considered that human speech perception and production both rely on articulatory representations. In this paper, we investigate whether this type of representation could improve the performances of a deep generative model…

声音 · 计算机科学 2021-04-08 Marc-Antoine Georges , Laurent Girin , Jean-Luc Schwartz , Thomas Hueber

We present a model for capturing musical features and creating novel sequences of music, called the Convolutional Variational Recurrent Neural Network. To generate sequential data, the model uses an encoder-decoder architecture with latent…

声音 · 计算机科学 2018-10-09 Eunjeong Stella Koh , Shlomo Dubnov , Dustin Wright

We present Deep Voice, a production-quality text-to-speech system constructed entirely from deep neural networks. Deep Voice lays the groundwork for truly end-to-end neural speech synthesis. The system comprises five major building blocks:…

A novel speech feature fusion algorithm with independent vector analysis (IVA) and parallel convolutional neural network (PCNN) is proposed for text-independent speaker recognition. Firstly, some different feature types, such as the time…

音频与语音处理 · 电气工程与系统科学 2022-12-02 Biao Ma , Chengben Xu , Ye Zhang

Sound events often occur in unstructured environments where they exhibit wide variations in their frequency content and temporal structure. Convolutional neural networks (CNN) are able to extract higher level features that are invariant to…

机器学习 · 计算机科学 2017-05-31 Emre Çakır , Giambattista Parascandolo , Toni Heittola , Heikki Huttunen , Tuomas Virtanen

We develop a Deep-Text Recurrent Network (DTRN) that regards scene text reading as a sequence labelling problem. We leverage recent advances of deep convolutional neural networks to generate an ordered high-level sequence from a whole word…

计算机视觉与模式识别 · 计算机科学 2015-12-22 Pan He , Weilin Huang , Yu Qiao , Chen Change Loy , Xiaoou Tang

Voice Activity Detection (VAD) refers to the task of identification of regions of human speech in digital signals such as audio and video. While VAD is a necessary first step in many speech processing systems, it poses challenges when there…

机器学习 · 计算机科学 2020-08-24 Arnab Kumar Mondal , Prathosh A. P

Languages have long been described according to their perceived rhythmic attributes. The associated typologies are of interest in psycholinguistics as they partly predict newborns' abilities to discriminate between languages and provide…

音频与语音处理 · 电气工程与系统科学 2024-01-29 François Deloche , Laurent Bonnasse-Gahot , Judit Gervain

In recent years, using raw waveforms as input for deep networks has been widely explored for the speaker verification system. For example, RawNet and RawNet2 extracted speaker's feature embeddings from waveforms automatically for…

音频与语音处理 · 电气工程与系统科学 2021-10-08 Jin Li , Nan Yan , Lan Wang

Self-supervised-learning-based pre-trained models for speech data, such as Wav2Vec 2.0 (W2V2), have become the backbone of many speech tasks. In this paper, to achieve speaker diarisation and speech recognition using a single model, a…

音频与语音处理 · 电气工程与系统科学 2022-07-11 Xianrui Zheng , Chao Zhang , Philip C. Woodland

Anomalous audio in speech recordings is often caused by speaker voice distortion, external noise, or even electric interferences. These obstacles have become a serious problem in some fields, such as high-quality music mixing and speech…

音频与语音处理 · 电气工程与系统科学 2021-02-11 Qiang Huang , Thomas Hain

The visual relationship recognition (VRR) task aims at understanding the pairwise visual relationships between interacting objects in an image. These relationships typically have a long-tail distribution due to their compositional nature.…

计算机视觉与模式识别 · 计算机科学 2022-03-30 Jun Chen , Aniket Agarwal , Sherif Abdelkarim , Deyao Zhu , Mohamed Elhoseiny

Multi-resolution spectro-temporal features of a speech signal represent how the brain perceives sounds by tuning cortical cells to different spectral and temporal modulations. These features produce a higher dimensional representation of…

音频与语音处理 · 电气工程与系统科学 2022-06-28 Rahil Parikh , Nadee Seneviratne , Ganesh Sivaraman , Shihab Shamma , Carol Espy-Wilson

This paper explores the use of multi-view features and their discriminative transforms in a convolutional deep neural network (CNN) architecture for a continuous large vocabulary speech recognition task. Mel-filterbank energies and…

计算与语言 · 计算机科学 2018-02-19 Vikramjit Mitra , Wen Wang , Chris Bartels , Horacio Franco , Dimitra Vergyri

This contribution gives an overview of face recogni-tion algorithms, their implementation and practical uses. First, a training set of different persons' faces has to be collected and used to train a face recognizer. The resulting face…

计算机视觉与模式识别 · 计算机科学 2017-07-05 Johannes Reschke , Armin Sehr

We present a multilinear statistical model of the human tongue that captures anatomical and tongue pose related shape variations separately. The model is derived from 3D magnetic resonance imaging data of 11 speakers sustaining speech…

计算机视觉与模式识别 · 计算机科学 2018-04-18 Alexander Hewer , Stefanie Wuhrer , Ingmar Steiner , Korin Richmond

Speaker identification in noisy audio recordings, specifically those from collaborative learning environments, can be extremely challenging. There is a need to identify individual students talking in small groups from other students talking…

音频与语音处理 · 电气工程与系统科学 2022-07-05 Antonio Gomez

While Word2Vec represents words (in text) as vectors carrying semantic information, audio Word2Vec was shown to be able to represent signal segments of spoken words as vectors carrying phonetic structure information. Audio Word2Vec can be…

计算与语言 · 计算机科学 2018-08-08 Yu-Hsuan Wang , Hung-yi Lee , Lin-shan Lee

Speech separation models are used for isolating individual speakers in many speech processing applications. Deep learning models have been shown to lead to state-of-the-art (SOTA) results on a number of speech separation benchmarks. One…

声音 · 计算机科学 2023-03-13 William Ravenscroft , Stefan Goetze , Thomas Hain

This paper proposes a voice conversion (VC) method based on a sequence-to-sequence (S2S) learning framework, which enables simultaneous conversion of the voice characteristics, pitch contour, and duration of input speech. We previously…

音频与语音处理 · 电气工程与系统科学 2020-11-10 Hirokazu Kameoka , Wen-Chin Huang , Kou Tanaka , Takuhiro Kaneko , Nobukatsu Hojo , Tomoki Toda