中文
相关论文

相关论文: A comparative study of eight human auditory models…

200 篇论文

Over the years, many researchers have seemingly made the same observation: Brain and language model activations exhibit some structural similarities, enabling linear partial mappings between features extracted from neural recordings and…

计算与语言 · 计算机科学 2023-06-09 Antonia Karamolegkou , Mostafa Abdou , Anders Søgaard

Do machines and humans process language in similar ways? Recent research has hinted at the affirmative, showing that human neural activity can be effectively predicted using the internal representations of language models (LMs). Although…

计算与语言 · 计算机科学 2025-01-15 Yuchen Zhou , Emmy Liu , Graham Neubig , Michael J. Tarr , Leila Wehbe

In the first year of life, infants' speech perception becomes attuned to the sounds of their native language. Many accounts of this early phonetic learning exist, but computational models predicting the attunement patterns observed in…

计算与语言 · 计算机科学 2020-08-10 Yevgen Matusevych , Thomas Schatz , Herman Kamper , Naomi H. Feldman , Sharon Goldwater

This research work demonstrates that current AI systems fail catastrophically on auditory tasks that humans perform effortlessly. Drawing inspiration from Moravec's paradox (i.e., tasks simple for humans often prove difficult for machines,…

人工智能 · 计算机科学 2025-08-01 David Noever , Forrest McKee

Recent literature suggests that the bigger the model, the more likely it is to converge to similar, ``universal'' representations, despite different training objectives, datasets, or modalities. While this literature shows that there is an…

计算机视觉与模式识别 · 计算机科学 2026-01-30 Matéo Mahaut , Marco Baroni

ODAQ (Open Dataset of Audio Quality) provides a comprehensive framework for exploring both monaural and binaural audio quality degradations across a range of distortion classes and signals, accompanied by subjective quality ratings. A…

音频与语音处理 · 电气工程与系统科学 2025-12-12 Pablo M. Delgado , Sascha Dick , Christoph Thompson , Chih-Wei Wu , Phillip A. Williams

Predictive coding is the leading algorithmic framework to understand how expectations shape our experience of reality. Its main tenet is that sensory neurons encode prediction error: the residuals between a generative model of the sensory…

神经元与认知 · 定量生物学 2022-01-20 Alejandro Tabas , Katharina von Kriegstein

While both the data volume and heterogeneity of the digital music content is huge, it has become increasingly important and convenient to build a recommendation or search system to facilitate surfacing these content to the user or consumer…

Recently, neural approaches to coherence modeling have achieved state-of-the-art results in several evaluation tasks. However, we show that most of these models often fail on harder tasks with more realistic application scenarios. In…

计算与语言 · 计算机科学 2019-09-04 Han Cheol Moon , Tasnim Mohiuddin , Shafiq Joty , Xu Chi

Non-native speakers show difficulties with spoken word processing. Many studies attribute these difficulties to imprecise phonological encoding of words in the lexical memory. We test an alternative hypothesis: that some of these…

计算与语言 · 计算机科学 2021-03-12 Yevgen Matusevych , Herman Kamper , Thomas Schatz , Naomi H. Feldman , Sharon Goldwater

Understanding the internal mechanisms of large audio-language models (LALMs) is crucial for interpreting their behavior and improving performance. This work presents the first in-depth analysis of how LALMs internally perceive and recognize…

计算与语言 · 计算机科学 2025-08-26 Chih-Kai Yang , Neo Ho , Yi-Jyun Lee , Hung-yi Lee

In this paper, we present an audio analyzer assistant tool designed for a wide range of audio-based surveillance applications (This work is a part of our DEFAME FAKES and EUCINF projects). The proposed tool, refered to as Aud-Sur, comprises…

音频与语音处理 · 电气工程与系统科学 2025-04-01 Phat Lam , Lam Pham , Dat Tran , Alexander Schindler , Silvia Poletti , Marcel Hasenbalg , David Fischinger , Martin Boyer

Hair cells, the sensory receptors of the internal ear, subserve different functions in various receptor organs: they detect oscillatory stimuli in the auditory system, but transduce constant and step stimuli in the vestibular and…

神经元与认知 · 定量生物学 2017-03-29 Joshua D. Salvi , Daibhid O Maoileidigh , Brian A. Fabella , Melanie Tobin , A. J. Hudspeth

We present a framework to model the perceived quality of audio signals by combining convolutional architectures, with ideas from classical signal processing, and describe an approach to enhancing perceived acoustical quality. We demonstrate…

声音 · 计算机科学 2019-12-13 Prateek Verma , Jonathan Berger

Incoming sound is in cochlea and auditory nerve encoded into spike trains. At the third neuron of the auditory pathway, spike trains of the left and right sides are processed in brainstem nuclei to yield sound localization information. Two…

神经元与认知 · 定量生物学 2020-07-02 Petr Marsalek , Pavel Sanda , Zbynek Bures

When songs are composed or performed, there is often an intent by the singer/songwriter of expressing feelings or emotions through it. For humans, matching the emotiveness in a musical composition or performance with the subjective…

Earable acoustic sensing offers a powerful and non-invasive modality for capturing fine-grained auditory and physiological signals directly from the ear canal, enabling continuous and context-aware monitoring of cognitive states. As earable…

人机交互 · 计算机科学 2025-12-23 Xijia Wei , Ting Dang , Khaldoon Al-Naimi , Yang Liu , Fahim Kawsar , Alessandro Montanari

Self-supervised language and audio models effectively predict brain responses to speech. However, traditional prediction models rely on linear mappings from unimodal features, despite the complex integration of auditory signals with…

计算与语言 · 计算机科学 2025-02-19 Danny Dongyeop Han , Yunju Cho , Jiook Cha , Jay-Yoon Lee

The ability to modulate vocal sounds and generate speech is one of the features which set humans apart from other living beings. The human voice can be characterized by several attributes such as pitch, timbre, loudness, and vocal tone. It…

计算机视觉与模式识别 · 计算机科学 2017-10-30 Poorna Banerjee Dasgupta

Efficient audio quality assessment is vital for streamlining audio codec development. Objective assessment tools have been developed over time to algorithmically predict quality ratings from subjective assessments, the gold standard for…

音频与语音处理 · 电气工程与系统科学 2024-11-28 Pablo M. Delgado , Jürgen Herre