中文
相关论文

相关论文: On HRTF Notch Frequency Prediction Using Anthropom…

200 篇论文

We describe a novel pipeline to automatically discover hierarchies of repeated sections in musical audio. The proposed method uses similarity network fusion (SNF) to combine different frame-level features into clean affinity matrices, which…

信息检索 · 计算机科学 2019-02-05 Christopher J. Tralie , Brian McFee

Recent advances in Neural Radiance Fields (NeRF) have demonstrated promising results in 3D scene representations, including 3D human representations. However, these representations often lack crucial information on the underlying human pose…

计算机视觉与模式识别 · 计算机科学 2024-04-10 Arnab Dey , Di Yang , Rohith Agaram , Antitza Dantcheva , Andrew I. Comport , Srinath Sridhar , Jean Martinet

An important aspect of a humanoid robot is audition. Previous work has presented robot systems capable of sound localization and source segregation based on microphone arrays with various configurations. However, no theoretical framework…

音频与语音处理 · 电气工程与系统科学 2024-01-09 Vladimir Tourbabin , Boaz Rafaely

The identification of siren sounds in urban soundscapes is a crucial safety aspect for smart vehicles and has been widely addressed by means of neural networks that ensure robustness to both the diversity of siren signals and the strong and…

音频与语音处理 · 电气工程与系统科学 2024-09-16 Stefano Damiano , Thomas Dietzen , Toon van Waterschoot

Machine learning models differ in terms of accuracy, computational/memory complexity, training time, and adaptability among other characteristics. For example, neural networks (NNs) are well-known for their high accuracy due to the quality…

机器学习 · 计算机科学 2020-08-05 Mahdi Nazemi , Amirhossein Esmaili , Arash Fayyazi , Massoud Pedram

This paper presents a Head-Related Transfer Function (HRTF)-guided framework for binaural Target Speaker Extraction (TSE) from mixtures of concurrent sources. Unlike conventional TSE methods based on Direction of Arrival (DOA) estimation or…

音频与语音处理 · 电气工程与系统科学 2026-03-18 Yoav Ellinson , Sharon Gannot

The challenge of creating domain-centric embeddings arises from the abundance of unstructured data and the scarcity of domain-specific structured data. Conventional embedding techniques often rely on either modality, limiting their…

机器学习 · 计算机科学 2024-10-29 Sharadind Peddiraju , Srini Rajagopal

Convolutional neural networks (CNNs) and transformer architectures offer strengths for modeling temporal data: CNNs excel at capturing local patterns and translational invariances, while transformers effectively model long-range…

机器学习 · 计算机科学 2025-10-09 Stefano F. Stefenon , João P. Matos-Carvalho , Valderi R. Q. Leithardt , Kin-Choong Yow

The Frequency Following Response (FFR) reflects the brain's neural encoding of auditory stimuli including speech. Because the fundamental frequency (F0), a physical correlate of pitch, is one of the essential features of speech, there has…

In this paper, we attempt to study the conditioning of the Spherical Harmonic Matrix (SHM), which is widely used in the discrete, limited order orthogonal representation of sound fields. SHM's has been widely used in the audio applications…

音频与语音处理 · 电气工程与系统科学 2018-03-07 C Sandeep Reddy , Rajesh M Hegde

A primary challenge in developing synthetic spatial hearing systems, particularly underwater, is accurately modeling sound scattering. Biological organisms achieve 3D spatial hearing by exploiting sound scattering off their bodies to…

声音 · 计算机科学 2026-03-03 Siminfar Samakoush Galougah , Pranav Pulijala , Ramani Duraiswami

In recent years, there has been increasing interest in applying stylization on 3D scenes from a reference style image, in particular onto neural radiance fields (NeRF). While performing stylization directly on NeRF guarantees appearance…

计算机视觉与模式识别 · 计算机科学 2023-09-20 Hong-Wing Pang , Binh-Son Hua , Sai-Kit Yeung

Beamforming with desired directivity patterns using compact microphone arrays is essential in many audio applications. Directivity patterns achievable using traditional beamformers depend on the number of microphones and the array aperture.…

音频与语音处理 · 电气工程与系统科学 2026-03-24 Weilong Huang , Srikanth Raj Chetupalli , Mhd Modar Halimeh , Oliver Thiergart , Emanuël A. P. Habets

Time-Frequency Distributions (TFDs) support the heart sound characterisation and classification in early cardiac screening. However, despite the frequent use of TFDs in signal analysis, no study comprehensively compared their performances…

信号处理 · 电气工程与系统科学 2022-08-08 Xinqi Bao , Yujia Xu , Hak-Keung Lam , Mohamed Trabelsi , Ines Chihi , Lilia Sidhom , Ernest N. Kamavuako

In this paper we developed a hierarchical network model, called Hierarchical Prediction Network (HPNet), to understand how spatiotemporal memories might be learned and encoded in the recurrent circuits in the visual cortical hierarchy for…

神经与进化计算 · 计算机科学 2021-10-04 Jielin Qiu , Ge Huang , Tai Sing Lee

Ensuring the trustworthiness and interpretability of machine learning models is critical to their deployment in real-world applications. Feature attribution methods have gained significant attention, which provide local explanations of…

机器学习 · 计算机科学 2023-09-20 Md Abdul Kadir , Gowtham Krishna Addluri , Daniel Sonntag

This paper addresses the challenge of Neural Field (NeF) generalization, where models must efficiently adapt to new signals given only a few observations. To tackle this, we propose Geometric Neural Process Fields (G-NPF), a probabilistic…

计算机视觉与模式识别 · 计算机科学 2025-02-05 Wenzhe Yin , Zehao Xiao , Jiayi Shen , Yunlu Chen , Cees G. M. Snoek , Jan-Jakob Sonke , Efstratios Gavves

In deep learning research, many melody extraction models rely on redesigning neural network architectures to improve performance. In this paper, we propose an input feature modification and a training objective modification based on two…

声音 · 计算机科学 2023-08-08 Keren Shao , Ke Chen , Taylor Berg-Kirkpatrick , Shlomo Dubnov

We propose a method of head-related transfer function (HRTF) interpolation from sparsely measured HRTFs using an autoencoder with source position conditioning. The proposed method is drawn from an analogy between an HRTF interpolation…

声音 · 计算机科学 2022-07-25 Yuki Ito , Tomohiko Nakamura , Shoichi Koyama , Hiroshi Saruwatari

This paper presents ER-NeRF, a novel conditional Neural Radiance Fields (NeRF) based architecture for talking portrait synthesis that can concurrently achieve fast convergence, real-time rendering, and state-of-the-art performance with…

计算机视觉与模式识别 · 计算机科学 2023-08-25 Jiahe Li , Jiawei Zhang , Xiao Bai , Jun Zhou , Lin Gu