English
Related papers

Related papers: A Wide Dataset of Ear Shapes and Pinna-Related Tra…

200 papers

Sound source localization relies on spatial cues such as interaural time differences (ITD), interaural level differences (ILD), and monaural spectral cues. Individually measured Head-Related Transfer Functions (HRTFs) facilitate precise…

Audio and Speech Processing · Electrical Eng. & Systems 2025-03-17 Nils Marggraf-Turley , Martha Shiell , Niels Pontoppidan , Drew Cappotto , Lorenzo Picinali

In the fields of brain-computer interaction and cognitive neuroscience, effective decoding of auditory signals from task-based functional magnetic resonance imaging (fMRI) is key to understanding how the brain processes complex auditory…

Neurons and Cognition · Quantitative Biology 2024-06-05 Wanli Ma , Xuegang Tang , Jin Gu , Ying Wang , Yuling Xia

We present a head-related transfer function (HRTF) estimation method which relies on a data-driven prior given by a score-based diffusion model. The HRTF is estimated in reverberant environments using natural excitation signals, e.g. human…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-03 Etienne Thuillier , Jean-Marie Lemercier , Eloi Moliner , Timo Gerkmann , Vesa Välimäki

Data-centric artificial intelligence (AI) has remarkably advanced medical imaging, with emerging methods using synthetic data to address data scarcity while introducing synthetic-to-real gaps. Unsupervised domain adaptation (UDA) shows…

Image and Video Processing · Electrical Eng. & Systems 2025-09-03 Linyu Fan , Che Wang , Ming Ye , Qizhi Yang , Zejun Wu , Xinghao Ding , Yue Huang , Jianfeng Bao , Shuhui Cai , Congbo Cai

Deep neural networks have shown exemplary performance on semantic scene understanding tasks on source domains, but due to the absence of style diversity during training, enhancing performance on unseen target domains using only single…

Computer Vision and Pattern Recognition · Computer Science 2024-03-29 Sumanth Udupa , Prajwal Gurunath , Aniruddh Sikdar , Suresh Sundaram

Speech deepfake detection is a well-established research field with different models, datasets, and training strategies. However, the lack of standardized implementations and evaluation protocols limits reproducibility, benchmarking, and…

In this paper, we abstract the process of people hearing speech, extracting meaningful cues, and creating various dynamically audio-consistent talking faces, termed Listening and Imagining, into the task of high-fidelity diverse talking…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Chao Xu , Yang Liu , Jiazheng Xing , Weida Wang , Mingze Sun , Jun Dan , Tianxin Huang , Siyuan Li , Zhi-Qi Cheng , Ying Tai , Baigui Sun

Recent advances in deep learning have enabled the creation of natural-sounding synthesised speech. However, attackers have also utilised these tech-nologies to conduct attacks such as phishing. Numerous public datasets have been created to…

Sound · Computer Science 2024-04-30 Abdulazeez AlAli , George Theodorakopoulos

Whole brain neuroanatomy using tera-voxel light-microscopic data sets is of much current interest. A fundamental problem in this field is the mapping of individual brain data sets to a reference space. Previous work has not rigorously…

Neurons and Cognition · Quantitative Biology 2019-04-18 Brian C. Lee , Meng Kuan Lin , Yan Fu , Junichi Hata , Michael I. Miller , Partha P. Mitra

Audio-driven talking face generation, which aims to synthesize talking faces with realistic facial animations (including accurate lip movements, vivid facial expression details and natural head poses) corresponding to the audio, has…

Computer Vision and Pattern Recognition · Computer Science 2023-04-19 Rongliang Wu , Yingchen Yu , Fangneng Zhan , Jiahui Zhang , Xiaoqin Zhang , Shijian Lu

This paper presents a novel high-fidelity and low-latency universal neural vocoder framework based on multiband WaveRNN with data-driven linear prediction for discrete waveform modeling (MWDLP). MWDLP employs a coarse-fine bit WaveRNN…

Sound · Computer Science 2021-07-06 Patrick Lumban Tobing , Tomoki Toda

In MR fingerprinting (MRF) reconstruction, measured data is pattern-matched to simulated signals to extract quantitative tissue parameters. A critical drawback to this approach is the exponentially increasing compute time for mapping of…

Medical Physics · Physics 2024-07-17 Victoria Y. Yu , Kathryn R. Tringale , Ricardo Otazo , Ouri Cohen

Hearing loss is a major health problem and psychological burden in humans. Mouse models offer a possibility to elucidate genes involved in the underlying developmental and pathophysiological mechanisms of hearing impairment. To this end,…

We propose a data-driven design method of perfect-reconstruction filterbank (PRFB) for sound-source enhancement (SSE) based on deep neural network (DNN). DNNs have been used to estimate a time-frequency (T-F) mask in the short-time Fourier…

Audio and Speech Processing · Electrical Eng. & Systems 2019-03-22 Daiki Takeuchi , Kohei Yatabe , Yuma Koizumi , Yasuhiro Oikawa , Noboru Harada

Speech-driven three-dimensional (3D) facial animation synthesis aims to build a mapping from one-dimensional (1D) speech signals to time-varying 3D facial motion signals. Current methods still face challenges in maintaining lip-sync…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Bin Liu , Zhixiang Xiong , Zhifen He , Bo Li

Ear recognition system has been widely studied whereas there are just a few ear presentation attack detection methods for ear recognition systems, consequently, there is no publicly available ear presentation attack detection (PAD)…

Computer Vision and Pattern Recognition · Computer Science 2021-12-13 Jalil Nourmohammadi Khiarak

Reconstructing a 3D sound field from sparse microphone measurements is a fundamental yet ill-posed problem, which we address through Acoustic Transfer Function (ATF) magnitude estimation. ATF magnitude encapsulates key perceptual and…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-12 Ege Erdem , Shoichi Koyama , Tomohiko Nakamura , Orchisama Das , Zoran Cvetković

This paper presents a simulation-based approach to own voice detection (OVD) in hearing aids using a single microphone. While OVD can significantly improve user comfort and speech intelligibility, existing solutions often rely on multiple…

This paper presents a multimodal biometric system of fingerprint and ear biometrics. Scale Invariant Feature Transform (SIFT) descriptor based feature sets extracted from fingerprint and ear are fused. The fused set is encoded by K-medoids…

Computer Vision and Pattern Recognition · Computer Science 2010-02-03 Dakshina Ranjan Kisku , Phalguni Gupta , Jamuna Kanta Sing

This paper introduces a shoebox room simulator able to systematically generate synthetic datasets of binaural room impulse responses (BRIRs) given an arbitrary set of head-related transfer functions (HRTFs). The evaluation of machine…

Sound · Computer Science 2021-06-25 Roberto Barumerli , Daniele Bianchi , Michele Geronazzo , Federico Avanzini
‹ Prev 1 4 5 6 7 8 10 Next ›