English
Related papers

Related papers: Closed-Form Successive Relative Transfer Function …

200 papers

This study introduces an online target sound extraction (TSE) process using the similarity-and-independence-aware beamformer (SIBF) derived from an iterative batch algorithm. The study aimed to reduce latency while maintaining extraction…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-07 Atsuo Hiroe

We propose a natural way to generalize relative transfer functions (RTFs) to more than one source. We first prove that such a generalization is not possible using a single multichannel spectro-temporal observation, regardless of the number…

Sound · Computer Science 2015-07-02 Antoine Deleforge , Sharon Gannot , Walter Kellermann

The remote microphone technique (RMT) is often used in active noise control (ANC) applications to overcome design constraints in microphone placements by estimating the acoustic pressure at inconvenient locations using a pre-calibrated…

Signal Processing · Electrical Eng. & Systems 2023-07-04 Chung Kwan Lai , Bhan Lam , Dongyuan Shi , Woon-Seng Gan

This paper introduces a multi-microphone method for extracting a desired speaker from a mixture involving multiple speakers and directional noise in a reverberant environment. In this work, we propose leveraging the instantaneous relative…

Sound · Computer Science 2025-02-11 Aviad Eisenberg , Sharon Gannot , Shlomo E. Chazan

This paper proposes an efficient parameterization of the Room Transfer Function (RTF). Typically, the RTF rapidly varies with varying source and receiver positions, hence requires an impractical number of point to point measurements to…

Sound · Computer Science 2015-05-19 Prasanga Samarasinghe , Thushara Abhayapala , Mark Poletti , Terence Betlehem

This paper proposes a blind estimation method based on the modulation transfer function and Schroeder model for estimating reverberation time in seven-octave bands. Therefore, the speech transmission index and five room-acoustic parameters…

Sound · Computer Science 2021-03-16 Suradej Duangpummet , Jessada Karnjana , Waree Kongprawechnon , Masashi Unoki

To estimate the direction of arrival (DOA) of multiple speakers, subspace-based prototype transfer function matching methods such as multiple signal classification (MUSIC) or relative transfer function (RTF) vector matching are commonly…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-11 Daniel Fejgin , Simon Doclo

We consider the problem of estimating the sparse time-varying parameter vectors of a point process model in an online fashion, where the observations and inputs respectively consist of binary and continuous time series. We construct a novel…

Neural and Evolutionary Computing · Computer Science 2016-04-20 Alireza Sheikhattar , Jonathan B. Fritz , Shihab A. Shamma , Behtash Babadi

Speaker verification is to judge the similarity between two unknown voices in an open set, where the ideal speaker embedding should be able to condense discriminant information into a compact utterance-level representation that has small…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-10 Hongyu Wang , Hui Li , Bo Li

This work addresses the problem of block-online processing for multi-channel speech enhancement. Such processing is vital in scenarios with moving speakers and/or when very short utterances are processed, e.g., in voice assistant scenarios.…

Sound · Computer Science 2020-05-27 Jiri Malek , Zbynek Koldovsky , Marek Bohac

When using an electron microscope for imaging of particles embedded in vitreous ice, the objective lens will inevitably corrupt the projection images. This corruption manifests as a band-pass filter on the micrograph. In addition, it causes…

Image and Video Processing · Electrical Eng. & Systems 2020-01-29 Ayelet Heimowitz , Joakim Andén , Amit Singer

Dereverberation of a moving speech source in the presence of other directional interferers, is a harder problem than that of stationary source and interference cancellation. We explore joint multi channel linear prediction (MCLP) and…

Audio and Speech Processing · Electrical Eng. & Systems 2019-10-23 Srikanth Raj Chetupalli , Thippur V. Sreenivas

Recent years have seen an increased interest in establishing association between faces and voices of celebrities leveraging audio-visual information from YouTube. Prior works adopt metric learning methods to learn an embedding space that is…

Computer Vision and Pattern Recognition · Computer Science 2022-08-23 Muhammad Saad Saeed , Shah Nawaz , Muhammad Haris Khan , Sajid Javed , Muhammad Haroon Yousaf , Alessio Del Bue

Estimating Head-Related Transfer Functions (HRTFs) of arbitrary source points is essential in immersive binaural audio rendering. Computing each individual's HRTFs is challenging, as traditional approaches require expensive time and…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-04 Jin Woo Lee , Sungho Lee , Kyogu Lee

In this paper we present a new robust sound source localization and tracking method using an array of eight microphones (US patent pending) . The method uses a steered beamformer based on the reliability-weighted phase transform (RWPHAT)…

Robotics · Computer Science 2016-04-07 Jean-Marc Valin , François Michaud , Jean Rouat

Voice Type Discrimination (VTD) refers to discrimination between regions in a recording where speech was produced by speakers that are physically within proximity of the recording device ("Live Speech") from speech and other types of audio…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-12 Tyler Vuong , Yangyang Xia , Richard Stern

This paper presents a Head-Related Transfer Function (HRTF)-guided framework for binaural Target Speaker Extraction (TSE) from mixtures of concurrent sources. Unlike conventional TSE methods based on Direction of Arrival (DOA) estimation or…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-18 Yoav Ellinson , Sharon Gannot

In room acoustic environments, the Relative Transfer Functions (RTFs) are controlled by few underlying modes of variability. Accordingly, they are confined to a low-dimensional manifold. In this letter, we investigate a RTF inverse…

Sound · Computer Science 2017-10-26 Ziteng Wang , Emmanuel Vincent , Yonghong Yan

Visual Object Tracking (VOT) has synchronous needs for both robustness and accuracy. While most existing works fail to operate simultaneously on both, we investigate in this work the problem of conflicting performance between accuracy and…

Computer Vision and Pattern Recognition · Computer Science 2021-03-19 Jinghao Zhou , Bo Li , Lei Qiao , Peng Wang , Weihao Gan , Wei Wu , Junjie Yan , Wanli Ouyang

We study the problem of learning association between face and voice, which is gaining interest in the computer vision community lately. Prior works adopt pairwise or triplet loss formulations to learn an embedding space amenable for…

Computer Vision and Pattern Recognition · Computer Science 2021-12-21 Muhammad Saad Saeed , Muhammad Haris Khan , Shah Nawaz , Muhammad Haroon Yousaf , Alessio Del Bue