中文
相关论文

相关论文: MMW: Side Talk Rejection Multi-Microphone Whisper …

200 篇论文

State-of-the-art vocal separation models like Mel-Band-Roformer rely on full temporal self-attention mechanisms, where each temporal frame interacts with every other frames. This incurs heavy computational costs that scales quadratically…

声音 · 计算机科学 2025-10-30 Christodoulos Benetatos , Yongyi Zang , Randal Leistikow

Multichannel convolutive blind speech source separation refers to the problem of separating different speech sources from the observed multichannel mixtures without much a priori information about the mixing system. Multichannel nonnegative…

声音 · 计算机科学 2024-01-04 Jianyu Wang , Shanzheng Guan

In this paper, we explore an improved framework to train a monoaural neural enhancement model for robust speech recognition. The designed training framework extends the existing mixture invariant training criterion to exploit both unpaired…

声音 · 计算机科学 2022-09-21 Jisi Zhang , Catalin Zorila , Rama Doddipatla , Jon Barker

Recent multimodal large language models (MLLMs) such as GPT-4o and Qwen3-Omni show strong perception but struggle in multi-speaker, dialogue-centric settings that demand agentic reasoning tracking who speaks, maintaining roles, and…

Weakly supervised semantic segmentation (WSSS) models relying on class activation maps (CAMs) have achieved desirable performance comparing to the non-CAMs-based counterparts. However, to guarantee WSSS task feasible, we need to generate…

计算机视觉与模式识别 · 计算机科学 2023-06-07 Tao Chen , Yazhou Yao , Jinhui Tang

The performance of millimeter-wave (mmWave) and sub-terahertz (sub-THz) communication systems is significantly impaired by sensitivity to sudden blockages. In this work, we employ machine learning (ML) and our physics-based simulation tool…

信号处理 · 电气工程与系统科学 2025-09-03 Roghieh Mahdavihaji , Alexandra Duel-Hallen , Hans Hallen

Dialogue state tracking models play an important role in a task-oriented dialogue system. However, most of them model the slot types conditionally independently given the input. We discover that it may cause the model to be confused by slot…

计算与语言 · 计算机科学 2021-11-16 Ting-Rui Chiang , Yi-Ting Yeh

User-defined keyword spotting (KWS) without resorting to domain-specific pre-labeled training data is of fundamental importance in building adaptable and personalized voice interfaces. However, such systems are still faced with arduous…

音频与语音处理 · 电气工程与系统科学 2026-04-07 Lo-Ya Li , Tien-Hong Lo , Jeih-Weih Hung , Shih-Chieh Huang , Berlin Chen

Recently, the state space model (SSM) represented by Mamba has shown remarkable performance in long-term sequence modeling tasks, including speech enhancement. However, due to substantial differences in sub-band features, applying the same…

声音 · 计算机科学 2025-02-25 Jizhen Li , Weiping Tu , Yuhong Yang , Xinmeng Xu , Yiqun Zhang , Yanzhen Ren

This research proposes the interaction loop model "ASR-LLMs-Smart Glasses", which model combines automatic speech recognition, large language model and smart glasses to facilitate seamless human-computer interaction. And the methodology of…

人机交互 · 计算机科学 2025-01-24 Libo Wang

Millimeter Wave (mmWave) band provides a large spectrum to meet the high-demand capacity by the 5th generation (5G) wireless networks. However, to fully exploit the available spectrum, obstacles such as high path loss, channel sparsity, and…

信息论 · 计算机科学 2018-07-16 Mojtaba Ahmadi Almasi , Roohollah Amiri , Hani Mehrpouyan

Speech enhancement and separation have been a long-standing problem, especially with the recent advances using a single microphone. Although microphones perform well in constrained settings, their performance for speech separation decreases…

音频与语音处理 · 电气工程与系统科学 2022-04-15 Muhammed Zahid Ozturk , Chenshu Wu , Beibei Wang , Min Wu , K. J. Ray Liu

Despite the remarkable quality of LLM-based text-to-speech systems, their reliance on autoregressive Transformers leads to quadratic computational complexity, which severely limits practical applications. Linear-time alternatives, notably…

音频与语音处理 · 电气工程与系统科学 2026-03-16 Tan Dat Nguyen , Sangmin Bae , Joon Son Chung , Ji-Hoon Kim

Multi-speaker automatic speech recognition (MS-ASR) faces significant challenges in transcribing overlapped speech, a task critical for applications like meeting transcription and conversational analysis. While serialized output training…

音频与语音处理 · 电气工程与系统科学 2025-06-09 Yuke Lin , Ming Cheng , Ze Li , Beilong Tang , Ming Li

This paper introduces the system submitted by the DKU-SMIIP team for the Auto-KWS 2021 Challenge. Our implementation consists of a two-stage keyword spotting system based on query-by-example spoken term detection and a speaker verification…

音频与语音处理 · 电气工程与系统科学 2021-04-13 Yechen Wang , Yan Jia , Murong Ma , Zexin Cai , Ming Li

This paper describes a system that gives a mobile robot the ability to perform automatic speech recognition with simultaneous speakers. A microphone array is used along with a real-time implementation of Geometric Source Separation and a…

Speech enhancement performance degrades significantly in noisy environments, limiting the deployment of speech-controlled technologies in industrial settings, such as manufacturing plants. Existing speech enhancement solutions primarly rely…

机器人学 · 计算机科学 2026-02-23 Zachary Turcotte , François Grondin

The multichannel Wiener filter (MWF) and its variations have been extensively applied to binaural hearing aids. However, its major drawback is the distortion of the binaural cues of the residual noise, changing the original acoustic…

音频与语音处理 · 电气工程与系统科学 2019-11-12 Johnny Werner , Marcio H. Costa

Since the first speech recognition systems were built more than 30 years ago, improvement in voice technology has enabled applications such as smart assistants and automated customer support. However, conversation intelligence of the future…

音频与语音处理 · 电气工程与系统科学 2024-02-15 Desh Raj

Having knowledge on the room acoustic properties, e.g., the location of acoustic reflectors, allows to better reproduce the sound field as intended. Current state-of-the-art methods for room boundary detection using microphone measurements…

音频与语音处理 · 电气工程与系统科学 2022-06-09 Ellen Riemens , Pablo Martínez-Nuevo , Jorge Martinez , Martin Møller , Richard C. Hendriks