中文
相关论文

相关论文: Multi-channel Replay Speech Detection using an Ada…

200 篇论文

We present a novel speaker-independent acoustic-to-articulatory inversion (AAI) model, overcoming the limitations observed in conventional AAI models that rely on acoustic features derived from restricted datasets. To address these…

信号处理 · 电气工程与系统科学 2024-06-26 Woo-Jin Chung , Hong-Goo Kang

While the performance of face recognition systems has improved significantly in the last decade, they are proved to be highly vulnerable to presentation attacks (spoofing). Most of the research in the field of face presentation attack…

计算机视觉与模式识别 · 计算机科学 2019-07-10 Olegs Nikisins , Anjith George , Sebastien Marcel

This work introduces the Cleanformer, a streaming multichannel neural based enhancement frontend for automatic speech recognition (ASR). This model has a conformer-based architecture which takes as inputs a single channel each of raw and…

音频与语音处理 · 电气工程与系统科学 2023-05-05 Joseph Caroselli , Arun Narayanan , Nathan Howard , Tom O'Malley

Despite their immense popularity, deep learning-based acoustic systems are inherently vulnerable to adversarial attacks, wherein maliciously crafted audios trigger target systems to misbehave. In this paper, we present SirenAttack, a new…

密码学与安全 · 计算机科学 2019-07-25 Tianyu Du , Shouling Ji , Jinfeng Li , Qinchen Gu , Ting Wang , Raheem Beyah

Large Audio Language Models (LALMs) have significantly advanced audio understanding but introduce critical security risks, particularly through audio jailbreaks. While prior work has focused on English-centric attacks, we expose a far more…

声音 · 计算机科学 2025-04-03 Jaechul Roh , Virat Shejwalkar , Amir Houmansadr

Recently, fully recurrent neural network (RNN) based end-to-end models have been proven to be effective for multi-speaker speech recognition in both the single-channel and multi-channel scenarios. In this work, we explore the use of…

音频与语音处理 · 电气工程与系统科学 2020-02-14 Xuankai Chang , Wangyou Zhang , Yanmin Qian , Jonathan Le Roux , Shinji Watanabe

Millimeter-wave (mmWave) communication systems face increasing susceptibility to advanced beam-stealing attacks, posing a significant physical layer security threat. This paper introduces a novel framework employing an advanced Deep…

信号处理 · 电气工程与系统科学 2025-08-06 Seyed Bagher Hashemi Natanzi , Hossein Mohammadi , Bo Tang , Vuk Marojevic

Multi-channel multi-talker speech recognition presents formidable challenges in the realm of speech processing, marked by issues such as background noise, reverberation, and overlapping speech. Overcoming these complexities requires…

音频与语音处理 · 电气工程与系统科学 2023-10-09 Yiwen Shao

Many people are suffering from voice disorders, which can adversely affect the quality of their lives. In response, some researchers have proposed algorithms for automatic assessment of these disorders, based on voice signals. However,…

机器学习 · 计算机科学 2018-12-04 Yi-Te Hsu , Zining Zhu , Chi-Te Wang , Shih-Hau Fang , Frank Rudzicz , Yu Tsao

We present our system submission to the ASVspoof 2019 Challenge Physical Access (PA) task. The objective for this challenge was to develop a countermeasure that identifies speech audio as either bona fide or intercepted and replayed. The…

计算与语言 · 计算机科学 2019-09-27 Jennifer Williams , Joanna Rownicka

A person tends to generate dynamic attention towards speech under complicated environments. Based on this phenomenon, we propose a framework combining dynamic attention and recursive learning together for monaural speech enhancement. Apart…

声音 · 计算机科学 2020-04-02 Andong Li , Chengshi Zheng , Cunhang Fan , Renhua Peng , Xiaodong Li

Reinforcement learning (RL) is a promising data-driven approach for adaptive traffic signal control (ATSC) in complex urban traffic networks, and deep neural networks further enhance its learning power. However, centralized RL is infeasible…

机器学习 · 计算机科学 2019-03-13 Tianshu Chu , Jie Wang , Lara Codecà , Zhaojian Li

The growing popularity of multi-channel wearable devices, such as smart glasses, has led to a surge of applications such as targeted speech recognition and enhanced hearing. However, current approaches to solve these tasks use independently…

音频与语音处理 · 电气工程与系统科学 2024-09-19 Yufeng Yang , Desh Raj , Ju Lin , Niko Moritz , Junteng Jia , Gil Keren , Egor Lakomkin , Yiteng Huang , Jacob Donley , Jay Mahadeokar , Ozlem Kalinli

In this paper, we study several microphone channel selection and weighting methods for robust automatic speech recognition (ASR) in noisy conditions. For channel selection, we investigate two methods based on the maximum likelihood (ML)…

声音 · 计算机科学 2016-10-04 Zhaofeng Zhang , Xiong Xiao , Longbiao Wang , EngSiong Chng , Haizhou Li

Communication in multi-agent reinforcement learning (MARL) has been proven to effectively promote cooperation among agents recently. Since communication in real-world scenarios is vulnerable to noises and adversarial attacks, it is crucial…

多智能体系统 · 计算机科学 2023-12-20 Lebin Yu , Yunbo Qiu , Quanming Yao , Yuan Shen , Xudong Zhang , Jian Wang

Deep networks allow to obtain outstanding results in semantic segmentation, however they need to be trained in a single shot with a large amount of data. Continual learning settings where new classes are learned in incremental steps and…

计算机视觉与模式识别 · 计算机科学 2021-09-21 Andrea Maracani , Umberto Michieli , Marco Toldo , Pietro Zanuttigh

In real-life applications, the performance of speaker recognition systems always degrades when there is a mismatch between training and evaluation data. Many domain adaptation methods have been successfully used for eliminating the domain…

声音 · 计算机科学 2020-11-18 Qing Wang , Wei Rao , Pengcheng Guo , Lei Xie

Replay attacks comprise replaying previously recorded sensor measurements and injecting malicious signals into a physical plant, causing great damage to cyber-physical systems. Replay attack detection has been widely studied for linear…

系统与控制 · 电气工程与系统科学 2026-03-31 Tao Chen , Andreu Cecilia , Lei Wang , Daniele Astolfi , Zhitao Liu

Automatic speech recognition (ASR) of multi-channel multi-speaker overlapped speech remains one of the most challenging tasks to the speech community. In this paper, we look into this challenge by utilizing the location information of…

声音 · 计算机科学 2021-11-23 Yiwen Shao , Shi-Xiong Zhang , Dong Yu

Neural multi-channel speech enhancement models, in particular those based on the U-Net architecture, demonstrate promising performance and generalization potential. These models typically encode input channels independently, and integrate…

声音 · 计算机科学 2024-10-08 Ibrahim Aldarmaki , Thamar Solorio , Bhiksha Raj , Hanan Aldarmaki