中文
相关论文

相关论文: Binaural recording methods with analysis on inter-…

200 篇论文

Binaural audio delivers spatial cues essential for immersion, yet most consumer videos are monaural due to capture constraints. We introduce SIREN, a visually guided mono to binaural framework that explicitly predicts left and right…

声音 · 计算机科学 2026-04-01 Mingyeong Song , Seoyeon Ko , Junhyug Noh

Speaker localization for binaural microphone arrays has been widely studied for applications such as speech communication, video conferencing, and robot audition. Many methods developed for this task, including the direct path dominance…

音频与语音处理 · 电气工程与系统科学 2023-11-01 Yanir Maymon , Israel Nelken , Boaz Rafaely

Audiovisual representation learning typically relies on the correspondence between sight and sound. However, there are often multiple audio tracks that can correspond with a visual scene. Consider, for example, different conversations on…

声音 · 计算机科学 2024-06-11 Nikhil Singh , Chih-Wei Wu , Iroro Orife , Mahdi Kalayeh

Non-interactive and linear experiences like cinema film offer high quality surround sound audio to enhance immersion, however the listener's experience is usually fixed to a single acoustic perspective. With the rise of virtual reality,…

音频与语音处理 · 电气工程与系统科学 2024-10-30 Lachlan Birnie , Thushara Abhayapala , Vladimir Tourbabin , Prasanga Samarasinghe

In the last few years, steganography has attracted increasing attention from a large number of researchers since its applications are expanding further than just the field of information security. The most traditional method is based on…

密码学与安全 · 计算机科学 2021-02-19 Quang Pham Huu , Thoi Hoang Dinh , Ngoc N. Tran , Toan Pham Van , Thanh Ta Minh

Cochlear implants(CIs) are arguably the most successful neural implant, having restored hearing to over one million people worldwide. While CI research has focused on modeling the cochlear activations in response to low-level acoustic…

神经与进化计算 · 计算机科学 2024-07-31 Cynthia R. Steinhardt , Menoua Keshishian , Nima Mesgarani , Kim Stachenfeld

This article presents a review of typical techniques used in three distinct aspects of deep learning model development for audio generation. In the first part of the article, we provide an explanation of audio representations, beginning…

声音 · 计算机科学 2024-06-04 Matej Božić , Marko Horvat

Many hearables contain an in-ear microphone, which may be used to capture the own voice of its user in noisy environments. Since the in-ear microphone mostly records body-conducted speech due to ear canal occlusion, it suffers from…

音频与语音处理 · 电气工程与系统科学 2024-03-25 Mattes Ohlenbusch , Christian Rollwage , Simon Doclo

This paper proposes a framework of explaining anomalous machine sounds in the context of anomalous sound detection~(ASD). While ASD has been extensively explored, identifying how anomalous sounds differ from normal sounds is also beneficial…

音频与语音处理 · 电气工程与系统科学 2024-10-30 Tomoya Nishida , Harsh Purohit , Kota Dohi , Takashi Endo , Yohei Kawaguchi

Only a few studies have been reported regarding human ear recognition in long wave infrared band. Thus, we have created ear database based on long wave infrared band. We have called that the database is long wave infrared band MIDAS…

计算机视觉与模式识别 · 计算机科学 2018-01-30 Umit Kacar , Murvet Kirci

Auditory display is concerned with the use of non-speech sound to communicate information. If the term seems at first oxymoronic, then consider auditory display as an activity of perceptualization, that is, the process of making perceptible…

人机交互 · 计算机科学 2013-11-25 Paul Vickers

The neural mechanisms underlying the comprehension of meaningful sounds are yet to be fully understood. While previous research has shown that the auditory cortex can classify auditory stimuli into distinct semantic categories, the specific…

神经元与认知 · 定量生物学 2023-09-20 Kumar Neelabh , Vishnu Sreekumar

Humans are highly dependent on the ability to process audio in order to interact through conversation and navigate from sound. For this, the shape of the ear acts as a mechanical audio filter. The anatomy of the outer human ear canal to…

计算机视觉与模式识别 · 计算机科学 2018-11-12 Sune Darkner , Stefan Sommer , Andreas Schuhmacher , Henrik Ingerslev Anders O. Baandrup , Carsten Thomsen , Søren Jønsson

Binaural audio generation (BAG) aims to convert monaural audio to stereo audio using visual prompts, requiring a deep understanding of spatial and semantic information. However, current models risk overfitting to room environments and lose…

Text-to-audio (TTA) systems have recently demonstrated strong performance in synthesizing monaural audio from text. However, the task of generating binaural spatial audio from text, which provides a more immersive auditory experience by…

音频与语音处理 · 电气工程与系统科学 2025-02-18 Linfeng Feng , Lei Zhao , Boyu Zhu , Xiao-Lei Zhang , Xuelong Li

We investigated the relationship among neural representations of vocalized, mimed, and imagined speech recorded using publicly available stereotactic EEG recordings. Most prior studies have focused on decoding speech responses within each…

声音 · 计算机科学 2026-02-27 Maryam Maghsoudi , Rupesh Chillale , Shihab A. Shamma

Human identification has always been a topic that interested researchers around the world. Biometric methods are found to be more effective and much easier for the users than the traditional identification methods like keys, smart cards and…

计算机视觉与模式识别 · 计算机科学 2013-09-30 Bijeesh T. , Nimmi I. P

Advanced auditory models are useful in designing signal-processing algorithms for hearing-loss compensation or speech enhancement. Such auditory models provide rich and detailed descriptions of the auditory pathway, and might allow for…

音频与语音处理 · 电气工程与系统科学 2024-03-18 Peter Leer , Jesper Jensen , Zheng-Hua Tan , Jan Østergaard , Lars Bramsløw

Current assistive hearing devices, such as hearing aids and cochlear implants, lack the ability to adapt to the listener's focus of auditory attention, limiting their effectiveness in complex acoustic environments like cocktail party…

信号处理 · 电气工程与系统科学 2025-10-29 Simon Geirnaert , Simon L. Kappel , Preben Kidmose

Acoustic identification of individual animals (AIID) is closely related to audio-based species classification but requires a finer level of detail to distinguish between individual animals within the same species. In this work, we frame…

声音 · 计算机科学 2024-09-16 Ines Nolasco , Ilyass Moummad , Dan Stowell , Emmanouil Benetos