中文
相关论文

相关论文: Multi-input Multi-output Beta Wavelet Network: Mod…

200 篇论文

Auscultatory analysis using an electronic stethoscope has attracted increasing attention in the clinical diagnosis of respiratory diseases. Recently, neural networks have been applied to assist in respiratory sound classification with…

声音 · 计算机科学 2025-04-25 Jiadong Xie , Yunlian Zhou , Mingsheng Xu

Voice activity detection (VAD) makes a distinction between speech and non-speech and its performance is of crucial importance for speech based services. Recently, deep neural network (DNN)-based VADs have achieved better performance than…

音频与语音处理 · 电气工程与系统科学 2020-08-14 Zhenpeng Zheng , Jianzong Wang , Ning Cheng , Jian Luo , Jing Xiao

Classroom environments are particularly challenging for children with hearing impairments, where background noise, multiple talkers, and reverberation degrade speech perception. These difficulties are greater for children than adults, yet…

Recently, neural vocoders have been widely used in speech synthesis tasks, including text-to-speech and voice conversion. However, when encountering data distribution mismatch between training and inference, neural vocoders trained on real…

声音 · 计算机科学 2020-08-21 Po-chun Hsu , Chun-hsuan Wang , Andy T. Liu , Hung-yi Lee

Acoustic emotion recognition aims to categorize the affective state of the speaker and is still a difficult task for machine learning models. The difficulties come from the scarcity of training data, general subjectivity in emotion…

计算与语言 · 计算机科学 2018-04-02 Egor Lakomkin , Cornelius Weber , Sven Magg , Stefan Wermter

We learn about the world from a diverse range of sensory information. Automated systems lack this ability as investigation has centred on processing information presented in a single form. Adapting architectures to learn from multiple…

机器学习 · 计算机科学 2020-10-27 Jason Armitage , Shramana Thakur , Rishi Tripathi , Jens Lehmann , Maria Maleshkova

Even though convolutional neural networks have become the method of choice in many fields of computer vision, they still lack interpretability and are usually designed manually in a cumbersome trial-and-error process. This paper aims at…

Often the best performing deep neural models are ensembles of multiple base-level networks, nevertheless, ensemble learning with respect to domain adaptive person re-ID remains unexplored. In this paper, we propose a multiple expert…

计算机视觉与模式识别 · 计算机科学 2020-07-14 Yunpeng Zhai , Qixiang Ye , Shijian Lu , Mengxi Jia , Rongrong Ji , Yonghong Tian

Most current audio-visual emotion recognition models lack the flexibility needed for deployment in practical applications. We envision a multimodal system that works even when only one modality is available and can be implemented…

机器学习 · 计算机科学 2026-01-13 Lucas Goncalves , Seong-Gyun Leem , Wei-Cheng Lin , Berrak Sisman , Carlos Busso

Recently, fully recurrent neural network (RNN) based end-to-end models have been proven to be effective for multi-speaker speech recognition in both the single-channel and multi-channel scenarios. In this work, we explore the use of…

音频与语音处理 · 电气工程与系统科学 2020-02-14 Xuankai Chang , Wangyou Zhang , Yanmin Qian , Jonathan Le Roux , Shinji Watanabe

In the last decade, the sound quality of electric induction motors is a hot topic in the research field. Specially, due to its high number of applications, the population is exposed to physical and psychological discomfort caused by the…

机器学习 · 计算机科学 2024-01-30 F. J. Jimenez-Romero , D. Guijo-Rubio , F. R. Lara-Raya , A. Ruiz-Gonzalez , C. Hervas-Martinez

Word-piece models (WPMs) are commonly used subword units in state-of-the-art end-to-end automatic speech recognition (ASR) systems. For multilingual ASR, due to the differences in written scripts across languages, multilingual WPMs bring…

音频与语音处理 · 电气工程与系统科学 2023-02-23 Chao Zhang , Bo Li , Tara N. Sainath , Trevor Strohman , Shuo-yiin Chang

The performance of speaker diarization is strongly affected by its clustering algorithm at the test stage. However, it is known that clustering algorithms are sensitive to random noises and small variations, particularly when the clustering…

音频与语音处理 · 电气工程与系统科学 2019-10-25 Meng-Zhen Li , Xiao-Lei Zhang

Speech bandwidth extension (BWE) has demonstrated promising performance in enhancing the perceptual speech quality in real communication systems. Most existing BWE researches primarily focus on fixed upsampling ratios, disregarding the fact…

声音 · 计算机科学 2023-12-22 Guochen Yu , Xiguang Zheng , Nan Li , Runqiang Han , Chengshi Zheng , Chen Zhang , Chao Zhou , Qi Huang , Bing Yu

Future wireless multiple-input multiple-output (MIMO) systems will integrate both sub-6 GHz and millimeter wave (mmWave) frequency bands to meet the growing demands for high data rates. MIMO link establishment typically requires accurate…

信号处理 · 电气工程与系统科学 2025-10-16 Faruk Pasic , Lukas Eller , Stefan Schwarz , Markus Rupp , Christoph F. Mecklenbräuker

In automatic speech processing systems, speaker diarization is a crucial front-end component to separate segments from different speakers. Inspired by the recent success of deep neural networks (DNNs) in semantic inferencing, triplet…

音频与语音处理 · 电气工程与系统科学 2018-08-07 Huan Song , Megan Willi , Jayaraman J. Thiagarajan , Visar Berisha , Andreas Spanias

Initially developed for natural language processing (NLP), Transformer model is now widely used for speech processing tasks such as speaker recognition, due to its powerful sequence modeling capabilities. However, conventional…

音频与语音处理 · 电气工程与系统科学 2022-01-28 Rui Wang , Junyi Ao , Long Zhou , Shujie Liu , Zhihua Wei , Tom Ko , Qing Li , Yu Zhang

Deep learning (DL) based channel estimation (CE) and multiple input and multiple output detection (MIMODet), as two separate research topics, have provided convinced evidence to demonstrate the effectiveness and robustness of artificial…

信号处理 · 电气工程与系统科学 2024-01-30 Xiangzhao Qin , Sha Hu , Jiankun Zhang , Jing Qian , Hao Wang

Covert channels can be used to circumvent system and network policies by establishing communications that have not been considered in the design of the computing system. We construct a covert channel between different computing systems that…

密码学与安全 · 计算机科学 2014-06-06 Michael Hanspach , Michael Goetz

Recently, GAN-based neural vocoders such as Parallel WaveGAN, MelGAN, HiFiGAN, and UnivNet have become popular due to their lightweight and parallel structure, resulting in a real-time synthesized waveform with high fidelity, even on a CPU.…

声音 · 计算机科学 2022-06-22 Yi Wang , Yi Si