English
Related papers

Related papers: Deep Learning Based Audio-Visual Multi-Speaker DOA…

200 papers

An objective understanding of media depictions, such as inclusive portrayals of how much someone is heard and seen on screen such as in film and television, requires the machines to discern automatically who, when, how, and where someone is…

Computer Vision and Pattern Recognition · Computer Science 2022-12-26 Rahul Sharma , Krishna Somandepalli , Shrikanth Narayanan

Audio representation learning based on deep neural networks (DNNs) emerged as an alternative approach to hand-crafted features. For achieving high performance, DNNs often need a large amount of annotated data which can be difficult and…

Machine Learning · Computer Science 2020-07-09 Xavier Favory , Konstantinos Drossos , Tuomas Virtanen , Xavier Serra

Conventional direction of arrival (DOA) estimation algorithms suffer from performance degradation due to antenna pattern distortion and substantial computational complexity in real-time execution. The support vector regression (SVR)…

Signal Processing · Electrical Eng. & Systems 2021-10-20 Md Imrul Hasan , Mohammad Saquib

At hybrid analog-digital (HAD) transceiver, an improved HAD rotational invariance techniques (ESPRIT), called I-HAD-ESPRIT, is proposed to measure the direction of arrival (DOA) of desired user, where the phase ambiguity due to HAD…

Information Theory · Computer Science 2018-12-04 Feng Shu , Ling Xu , Jinyong Lin , Jingsong Hu , Lingqing Gui , Jun Li , Jiangzhou Wang

This study demonstrates that a Multimodal Large Language Model (MLLM) adapted via Low-Rank Adaptation (LoRA) can perform both Automatic Pronunciation Assessment (APA) and Mispronunciation Detection and Diagnosis (MDD) simultaneously.…

Computation and Language · Computer Science 2025-09-04 Taekyung Ahn , Hosung Nam

In this paper, we propose a novel reduced-rank algorithm for direction of arrival (DOA) estimation based on the minimum variance (MV) power spectral evaluation. It is suitable to DOA estimation with large arrays and can be applied to…

Information Theory · Computer Science 2013-03-07 Lei Wang , Rodrigo C. de Lamare

We introduce DIVE, an end-to-end speaker diarization algorithm. Our neural algorithm presents the diarization task as an iterative process: it repeatedly builds a representation for each speaker before predicting the voice activity of each…

Sound · Computer Science 2021-05-31 Neil Zeghidour , Olivier Teboul , David Grangier

Subjective evaluations are critical for assessing the perceptual realism of sounds in audio-synthesis driven technologies like augmented and virtual reality. However, they are challenging to set up, fatiguing for users, and expensive. In…

Audio and Speech Processing · Electrical Eng. & Systems 2021-12-22 Pranay Manocha , Anurag Kumar , Buye Xu , Anjali Menon , Israel D. Gebru , Vamsi K. Ithapu , Paul Calamia

In this paper, a deep learning approach is presented for direction of arrival estimation using automotive-grade ultrasonic sensors which are used for driving assistance systems such as automatic parking. A study and implementation of the…

Signal Processing · Electrical Eng. & Systems 2022-02-28 Mohamed Shawki Elamir , Heinrich Gotzig , Raoul Zoellner , Patrick Maeder

Image-text retrieval has developed rapidly in recent years. However, it is still a challenge in remote sensing due to visual-semantic imbalance, which leads to incorrect matching of non-semantic visual and textual features. To solve this…

Computer Vision and Pattern Recognition · Computer Science 2024-11-27 Qing Ma , Jiancheng Pan , Cong Bai

In multi-speaker applications is common to have pre-computed models from enrolled speakers. Using these models to identify the instances in which these speakers intervene in a recording is the task of speaker tracking. In this paper, we…

In this paper, we propose a novel end-to-end neural-network-based speaker diarization method. Unlike most existing methods, our proposed method does not have separate modules for extraction and clustering of speaker representations.…

Audio and Speech Processing · Electrical Eng. & Systems 2019-09-16 Yusuke Fujita , Naoyuki Kanda , Shota Horiguchi , Kenji Nagamatsu , Shinji Watanabe

Probabilistic Linear Discriminant Analysis (PLDA) has become state-of-the-art method for modeling $i$-vector space in speaker recognition task. However the performance degradation is observed if enrollment data size differs from one speaker…

Computation and Language · Computer Science 2016-02-24 Danila Doroshin , Nikolay Lubimov , Marina Nastasenko , Mikhail Kotov

Reverberation not only degrades the quality of speech for human perception, but also severely impacts the accuracy of automatic speech recognition. Prior work attempts to remove reverberation based on the audio modality only. Our idea is to…

Sound · Computer Science 2023-03-15 Changan Chen , Wei Sun , David Harwath , Kristen Grauman

Deep learning has shown a great potential for speech separation, especially for speech and non-speech separation. However, it encounters permutation problem for multi-speaker separation where both target and interference are speech.…

Sound · Computer Science 2021-03-29 Hao Li , Xueliang Zhang , Guanglai Gao

Deep neural networks (DNNs) have greatly benefited direction of arrival (DoA) estimation methods for speech source localization in noisy environments. However, their localization accuracy is still far from satisfactory due to the…

Audio and Speech Processing · Electrical Eng. & Systems 2023-02-21 Kuan-Lin Chen , Ching-Hua Lee , Bhaskar D. Rao , Harinath Garudadri

This paper presents a novel method for estimating the direction of arrival (DOA) for a non-uniform and sparse linear sensor array using the weighted lifted structure low-rank matrix completion. The proposed method uses a single snapshot…

Signal Processing · Electrical Eng. & Systems 2025-07-29 Saeed Razavikia , Mohammad Bokaei , Arash Amini , Stefano Rini , Carlo Fischione

Wide-band Direction of Arrival (DOA) estimation with sensor arrays is an essential task in sonar, radar, acoustics, biomedical and multimedia applications. Many state of the art wide-band DOA estimators coherently process frequency binned…

Applications · Statistics 2018-03-02 Elio D. Di Claudio , Raffaele Parisi , Giovanni Jacovitti

In this paper, an accurate direction-of-arrival (DOA) estimator is developed based on the real-valued singular value decomposition (SVD) of covariance matrix. Unitary transform on the complex-valued covariance matrix is first applied, and…

Signal Processing · Electrical Eng. & Systems 2020-04-13 Hui Cao , Qi Liu

Active speaker detection requires a solid integration of multi-modal cues. While individual modalities can approximate a solution, accurate predictions can only be achieved by explicitly fusing the audio and visual features and modeling…

Computer Vision and Pattern Recognition · Computer Science 2021-10-06 Juan León-Alcázar , Fabian Caba Heilbron , Ali Thabet , Bernard Ghanem
‹ Prev 1 8 9 10 Next ›