中文
相关论文

相关论文: A Data-Driven Exploration of Elevation Cues in HRT…

200 篇论文

The various speech sounds of a language are obtained by varying the shape and position of the articulators surrounding the vocal tract. Analyzing their variations is crucial for understanding speech production, diagnosing speech disorders…

图像与视频处理 · 电气工程与系统科学 2020-02-04 Mohammad Eslami , Christiane Neuschaefer-Rube , Antoine Serrurier

Feature attributions are post-training analysis methods that assess how various input features of a machine learning model contribute to an output prediction. Their interpretation is straightforward when features act independently, but it…

机器学习 · 计算机科学 2026-01-29 Kurt Butler , Guanchao Feng , Petar Djuric

This paper provides an overview of recent progress in non-intrusive speech intelligibility prediction for hearing aids (HA). We summarize developments in robust acoustic feature extraction, hearing loss modeling, and the use of emerging…

音频与语音处理 · 电气工程与系统科学 2025-09-04 Ryandhimas E. Zezario

Recognizing patterns in lung sounds is crucial to detecting and monitoring respiratory diseases. Current techniques for analyzing respiratory sounds demand domain experts and are subject to interpretation. Hence an accurate and automatic…

音频与语音处理 · 电气工程与系统科学 2022-08-31 Zizhao Chen , Hongliang Wang , Chia-Hui Yeh , Xilin Liu

A deep feature based saliency model (DeepFeat) is developed to leverage the understanding of the prediction of human fixations. Traditional saliency models often predict the human visual attention relying on few level image cues. Although…

计算机视觉与模式识别 · 计算机科学 2017-09-11 Ali Mahdi , Jun Qin

Despite there being clear evidence for top-down (e.g., attentional) effects in biological spatial hearing, relatively few machine hearing systems exploit top-down model-based knowledge in sound localisation. This paper addresses this issue…

音频与语音处理 · 电气工程与系统科学 2019-04-08 Ning Ma , Jose A. Gonzalez , Guy J. Brown

Speech language models align with human brain responses to natural language to an impressive degree. However, current models rely heavily on low-level speech features, indicating they lack brain-relevant semantics which limits their utility…

计算与语言 · 计算机科学 2025-03-05 Omer Moussa , Dietrich Klakow , Mariya Toneva

Large language models are able to exploit in-context learning to access external knowledge beyond their training data through retrieval-augmentation. While promising, its inner workings remain unclear. In this work, we shed light on the…

计算与语言 · 计算机科学 2025-10-28 Patrick Kahardipraja , Reduan Achtibat , Thomas Wiegand , Wojciech Samek , Sebastian Lapuschkin

Emotion recognition is the task of classifying perceived emotions in people. Previous works have utilized various nonverbal cues to extract features from images and correlate them to emotions. Of these cues, situational context is…

计算机视觉与模式识别 · 计算机科学 2023-05-08 Willams de Lima Costa , Estefania Talavera Martinez , Lucas Silva Figueiredo , Veronica Teichrieb

We conduct a large-scale study of language models for chord prediction. Specifically, we compare N-gram models to various flavours of recurrent neural networks on a comprehensive dataset comprising all publicly available datasets of…

机器学习 · 计算机科学 2018-04-06 Filip Korzeniowski , David R. W. Sears , Gerhard Widmer

Transfer learning is a crucial concept within deep learning that allows artificial neural networks to benefit from a large pre-training data basis when confronted with a task of limited data. Despite its ubiquitous use and clear benefits,…

Audio impairment recognition is based on finding noise in audio files and categorising the impairment type. Recently, significant performance improvement has been obtained thanks to the usage of advanced deep learning models. However,…

音频与语音处理 · 电气工程与系统科学 2021-10-28 Alessandro Ragano , Emmanouil Benetos , Andrew Hines

Understanding uncertainty plays a critical role in achieving common ground (Clark et al.,1983). This is especially important for multimodal AI systems that collaborate with users to solve a problem or guide the user through a challenging…

计算与语言 · 计算机科学 2024-10-21 Qi Cheng , Mert İnan , Rahma Mbarki , Grace Grmek , Theresa Choi , Yiming Sun , Kimele Persaud , Jenny Wang , Malihe Alikhani

Synthesizing medical images while preserving their structural information is crucial in medical research. In such scenarios, the preservation of anatomical content becomes especially important. Although recent advances have been made by…

图像与视频处理 · 电气工程与系统科学 2024-11-14 Ziqi Yu , Botao Zhao , Shengjie Zhang , Xiang Chen , Jianfeng Feng , Tingying Peng , Xiao-Yong Zhang

The purpose of this paper is to compare different learnable frontends in medical acoustics tasks. A framework has been implemented to classify human respiratory sounds and heartbeats in two categories, i.e. healthy or affected by…

声音 · 计算机科学 2026-01-21 Alessandro Maria Poirè , Federico Simonetta , Stavros Ntalampiras

We consider the problem of audio voice separation for binaural applications, such as earphones and hearing aids. While today's neural networks perform remarkably well (separating $4+$ sources with 2 microphones) they assume a known or fixed…

声音 · 计算机科学 2022-07-18 Zhongweiyang Xu , Romit Roy Choudhury

Speaker recognition performance in emotional talking environments is not as high as it is in neutral talking environments. This work focuses on proposing, implementing, and evaluating a new approach to enhance the performance in emotional…

声音 · 计算机科学 2017-06-30 Ismail Shahin

In this paper, we summarize recent progresses made in deep learning based acoustic models and the motivation and insights behind the surveyed techniques. We first discuss acoustic models that can effectively exploit variable-length…

音频与语音处理 · 电气工程与系统科学 2018-04-30 Dong Yu , Jinyu Li

Facial expression recognition is vital for human behavior analysis, and deep learning has enabled models that can outperform humans. However, it is unclear how closely they mimic human processing. This study aims to explore the similarity…

计算机视觉与模式识别 · 计算机科学 2024-09-04 F. Xavier Gaya-Morey , Silvia Ramis-Guarinos , Cristina Manresa-Yee , Jose M. Buades-Rubio

Speech disorders such as stuttering disrupt the normal fluency of speech by involuntary repetitions, prolongations and blocking of sounds and syllables. In addition to these disruptions to speech fluency, most adults who stutter (AWS) also…

计算机视觉与模式识别 · 计算机科学 2020-10-06 Arun Das , Jeffrey Mock , Henry Chacon , Farzan Irani , Edward Golob , Peyman Najafirad
‹ 上一页 1 8 9 10 下一页 ›