中文
相关论文

相关论文: Rethinking the Separation Layers in Speech Separat…

200 篇论文

High quality speech capture has been widely studied for both voice communication and human computer interface reasons. To improve the capture performance, we can often find multi-microphone speech enhancement techniques deployed on various…

音频与语音处理 · 电气工程与系统科学 2023-03-15 Yang Yang , Shao-Fu Shih , Hakan Erdogan , Jamie Menjay Lin , Chehung Lee , Yunpeng Li , George Sung , Matthias Grundmann

This paper investigates a joint beamforming design in a multiuser multiple-input single-output (MISO) communication network aided with an intelligent reflecting surface (IRS) panel. The symbol-level precoding (SLP) is adopted to enhance the…

信息论 · 计算机科学 2021-09-01 Guangyang Zhang , Chao Shen , Bo Ai , Zhangdui Zhong

Conditional sound separation in multi-source audio mixtures without having access to single source sound data during training is a long standing challenge. Existing mix-and-separate based methods suffer from significant performance drop…

声音 · 计算机科学 2024-04-03 Tanvir Mahmud , Saeed Amizadeh , Kazuhito Koishida , Diana Marculescu

Real-time single-channel speech separation aims to unmix an audio stream captured from a single microphone that contains multiple people talking at once, environmental noise, and reverberation into multiple de-reverberated and noise-free…

音频与语音处理 · 电气工程与系统科学 2023-04-18 Julian Neri , Sebastian Braun

Modern large-scale language model pre-training relies heavily on the single program multiple data (SPMD) paradigm, which requires tight coupling across accelerators. Due to this coupling, transient slowdowns, hardware failures, and…

This paper examines the implications of using the Scale-Invariant Signal-to-Distortion Ratio (SI-SDR) as both evaluation and training objective in supervised speech separation, when the training references contain noise, as is the case with…

音频与语音处理 · 电气工程与系统科学 2025-08-21 Simon Dahl Jepsen , Mads Græsbøll Christensen , Jesper Rindom Jensen

Speech Foundation Models have gained significant attention recently. Prior works have shown that the fusion of representations from multiple layers of the same model or the fusion of multiple models can improve performance on downstream…

音频与语音处理 · 电气工程与系统科学 2025-11-12 Yi-Jen Shih , David Harwath

General audio source separation is a key capability for multimodal AI systems that can perceive and reason about sound. Despite substantial progress in recent years, existing separation models are either domain-specific, designed for fixed…

Multiple-input multiple-output (MIMO) has been a key technology of wireless communications for decades. A typical MIMO system employs antenna arrays with the inter-antenna spacing being half of the signal wavelength, which we term as…

信息论 · 计算机科学 2024-06-19 Xinrui Li , Hongqi Min , Yong Zeng , Shi Jin , Linglong Dai , Yifei Yuan , Rui Zhang

This paper addresses the challenge of speaker separation, which remains an active research topic despite the promising results achieved in recent years. These results, however, often degrade in real recording conditions due to the presence…

声音 · 计算机科学 2024-11-14 Rawad Melhem , Assef Jafar , Oumayma Al Dakkak

We consider the problem of audio voice separation for binaural applications, such as earphones and hearing aids. While today's neural networks perform remarkably well (separating $4+$ sources with 2 microphones) they assume a known or fixed…

声音 · 计算机科学 2022-07-18 Zhongweiyang Xu , Romit Roy Choudhury

Although fully end-to-end speaker diarization systems have made significant progress in recent years, modular systems often achieve superior results in real-world scenarios due to their greater adaptability and robustness. Historically,…

音频与语音处理 · 电气工程与系统科学 2024-09-26 Ruoyu Wang , Shutong Niu , Gaobin Yang , Jun Du , Shuangqing Qian , Tian Gao , Jia Pan

In this paper, the design of binary sequences exhibiting low values of aperiodic/periodic correlation functions, in terms of Integrated Sidelobe Level (ISL), is pursued via a learning-inspired method. Specifcally, the synthesis of either a…

信号处理 · 电气工程与系统科学 2023-05-17 Omid Rezaei , Mahdi Ahmadi , Mohammad Mahdi Naghsh , Augusto Aubry , Mohammad Mahdi Nayebi , Antonio De Maio

Multimodal large language models (MLLMs) achieve strong performance by jointly processing inputs from multiple modalities, such as vision, audio, and language. However, building such models or extending them to new modalities often requires…

机器学习 · 计算机科学 2026-03-24 Md Kaykobad Reza , Ameya Patil , Edward Ayrapetian , M. Salman Asif

Neural networks have led to tremendous performance gains for single-task speech enhancement, such as noise suppression and acoustic echo cancellation (AEC). In this work, we evaluate whether it is more useful to use a single joint or…

音频与语音处理 · 电气工程与系统科学 2022-07-14 Sebastian Braun , Maria Luis Valero

This study considers a dual-polarized intelligent reflecting surface (DP-IRS)-assisted multiple-input multiple-output (MIMO) single-user wireless communication system. The transmitter and receiver are equipped with DP antennas, and each…

信号处理 · 电气工程与系统科学 2023-07-13 Muteen Munawar , Kyungchun Lee

The performance of single channel source separation algorithms has improved greatly in recent times with the development and deployment of neural networks. However, many such networks continue to operate on the magnitude spectrogram of a…

音频与语音处理 · 电气工程与系统科学 2018-10-08 Shrikant Venkataramani , Paris Smaragdis

Large Language Models (LLMs) have enabled dynamic reasoning in automated data analytics, yet recent multi-agent systems remain limited by rigid, single-path workflows that restrict strategic exploration and often lead to suboptimal…

人工智能 · 计算机科学 2026-01-08 Wonduk Seo , Juhyeon Lee , Yanjun Shao , Qingshan Zhou , Seunghyun Lee , Yi Bu

Flexible intelligent metasurface (FIM) technology holds immense potential for increasing the spectral efficiency and energy efficiency of wireless networks. In contrast to traditional rigid reconfigurable intelligent surfaces (RIS), an FIM…

信号处理 · 电气工程与系统科学 2025-10-10 Hanwen Hu , Jiancheng An , Lu Gan , Naofal Al-Dhahir

Existing CNN-based speech separation models face local receptive field limitations and cannot effectively capture long time dependencies. Although LSTM and Transformer-based speech separation models can avoid this problem, their high…

声音 · 计算机科学 2024-09-11 Kai Li , Guo Chen , Runxuan Yang , Xiaolin Hu