English
Related papers

Related papers: ELEGANCE: Efficient LLM Guidance for Audio-Visual …

200 papers

Target speaker extraction (TSE) aims to isolate a specific speaker's voice from multi-speaker mixtures. Despite strong benchmark results, real-world performance often degrades due to different interacting factors. Previous curriculum…

Sound · Computer Science 2026-03-06 Yun Liu , Xuechen Liu , Xiaoxiao Miao , Junichi Yamagishi

Recent advancements in Vision-Language Models (VLMs) enable large language models (LLMs) to process high-resolution images, significantly improving real-world multimodal understanding. However, this capability introduces a large number of…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Yuna Lee , Kyoungho Min , Yulhwa Kim

Target speaker extraction (TSE) aims to isolate a desired speaker's voice from a multi-speaker mixture using auxiliary information such as a reference utterance. Although recent advances in diffusion and flow-matching models have improved…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-23 Riki Shimizu , Xilin Jiang , Nima Mesgarani

Research on audio clue-based target speaker extraction (TSE) has focused on modeling mixtures and reference speech, achieving strong results in English due to abundant datasets. However, cross-lingual properties remain underexplored, as…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-10 The Hieu Pham , Phuong Thanh Tran Nguyen , Xuan Tho Nguyen , Tan Dat Nguyen , Duc Dung Nguyen

Social media enables data-driven analysis of public opinion on contested issues. Target-Stance Extraction (TSE) is the task of identifying the target discussed in a document and the document's stance towards that target. Many works classify…

Computation and Language · Computer Science 2025-10-28 Ethan Mines , Bonnie Dorr

We propose a Beamformer-guided Target Speaker Extraction (BG-TSE) method to extract a target speaker's voice from a multi-channel recording informed by the direction of arrival of the target. The proposed method employs a front-end…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-16 Mohamed Elminshawi , Srikanth Raj Chetupalli , Emanuël A. P. Habets

We present the first neural target speech extraction (TSE) system that uses human feedback for iterative refinement. Our approach allows users to mark specific segments of the TSE output, generating an edit mask. The refinement system then…

Sound · Computer Science 2025-08-06 Malek Itani , Ashton Graves , Sefik Emre Eskimez , Shyamnath Gollakota

Language-queried target sound extraction (TSE) aims to extract specific sounds from mixtures based on language queries. Traditional fully-supervised training schemes require extensively annotated parallel audio-text data, which are…

Audio and Speech Processing · Electrical Eng. & Systems 2025-03-24 Hao Ma , Zhiyuan Peng , Xu Li , Yukai Li , Mingjie Shao , Qiuqiang Kong , Ju Liu

Target speaker extraction (TSE) is a technique for isolating a target speaker's voice from mixed speech using auxiliary features associated with the target speaker. It is another attempt at addressing the cocktail party problem and is…

Sound · Computer Science 2024-11-26 Chang Sun , Bo Qin

This paper proposes the target speaker enhancement based speaker verification network (TASE-SVNet), an all neural model that couples target speaker enhancement and speaker embedding extraction for robust speaker verification (SV).…

Audio and Speech Processing · Electrical Eng. & Systems 2021-03-17 Chunlei Zhang , Meng Yu , Chao Weng , Dong Yu

State-of-the-art target speaker extraction (TSE) systems are typically designed to generalize to any given mixing environment, necessitating a model with a large enough capacity as a generalist. Personalized speech enhancement could be a…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-06 Tsun-An Hsieh , Minje Kim

Target speaker extraction (TSE) is essential in speech processing applications, particularly in scenarios with complex acoustic environments. Current TSE systems face challenges in limited data diversity and a lack of robustness in…

Sound · Computer Science 2024-12-18 Yun Liu , Xuechen Liu , Xiaoxiao Miao , Junichi Yamagishi

Multimodal vision language models (VLMs) have made significant progress with the support of continuously increasing model sizes and data volumes. Running VLMs on edge devices has become a challenge for their widespread application. There…

Computer Vision and Pattern Recognition · Computer Science 2025-01-24 Miao Rang , Zhenni Bi , Chuanjian Liu , Yehui Tang , Kai Han , Yunhe Wang

Large Language Models (LLMs) have achieved impressive progress across a wide range of tasks, yet their heavy reliance on English-centric training data leads to significant performance degradation in non-English languages. While existing…

Computation and Language · Computer Science 2025-10-20 Hamin Koo , Jaehyung Kim

Large Language Models (LLMs) have shown remarkable capabilities as AI agents. However, existing methods for enhancing LLM-agent abilities often lack a focus on data quality, leading to inefficiencies and suboptimal results in both…

Machine Learning · Computer Science 2025-02-19 Yunxiao Zhang , Guanming Xiong , Haochen Li , Wen Zhao

Generative target speaker extraction (TSE) methods often produce more natural outputs than predictive models. Recent work based on diffusion or flow matching (FM) typically relies on a small, fixed number of reverse steps with a fixed step…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-21 Tsun-An Hsieh , Minje Kim

Emotion recognition from speech is a challenging task that requires capturing both linguistic and paralinguistic cues, with critical applications in human-computer interaction and mental health monitoring. Recent works have highlighted the…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-21 Hugo Thimonier , Antony Perzo , Renaud Seguier

Large language models with long context windows can answer complex questions directly from full-length academic, technical, and policy documents, but passing entire documents is often costly, slow, and can degrade answer quality while…

Previously, Target Speaker Extraction (TSE) has yielded outstanding performance in certain application scenarios for speech enhancement and source separation. However, obtaining auxiliary speaker-related information is still challenging in…

We propose LauraTSE, an Auto-Regressive Decoder-Only Language Model for Target Speaker Extraction built upon the LauraGPT backbone. LauraTSE employs a small-scale auto-regressive decoder-only language model that generates the initial layers…

Machine Learning · Computer Science 2025-08-19 Beilong Tang , Bang Zeng , Ming Li