English
Related papers

Related papers: A comparative study of two-dimensional vocal tract…

200 papers

Time domain simulations of electromagnetic problems are highly valuable in engineering applications, as they allow for the analysis of transient behavior and broadband responses. These simulations utilize time stepping schemes, where each…

Computational Physics · Physics 2024-10-23 Ruth Medeiros , Valentin de la Rubia

Acoustic to articulatory inversion has often been limited to a small part of the vocal tract because the data are generally EMA (ElectroMagnetic Articulography) data requiring sensors to be glued to easily accessible articulators. The…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-16 Sofiane Azzouz , Pierre-André Vuissoz , Yves Laprie

Two-stage pipeline is popular in speech enhancement tasks due to its superiority over traditional single-stage methods. The current two-stage approaches usually enhance the magnitude spectrum in the first stage, and further modify the…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-22 Yuewei Zhang , Huanbin Zou , Jie Zhu

Large-scale latent diffusion models (LDMs) excel in content generation across various modalities, but their reliance on phonemes and durations in text-to-speech (TTS) limits scalability and access from other fields. While recent studies…

Audio and Speech Processing · Electrical Eng. & Systems 2025-02-18 Keon Lee , Dong Won Kim , Jaehyeon Kim , Seungjun Chung , Jaewoong Cho

Simulation of 3D low-frequency electromagnetic fields propagating in the Earth is computationally expensive. We present a fictitious wave domain high-order finite-difference time-domain (FDTD) modelling method on nonuniform grids to compute…

Numerical Analysis · Mathematics 2023-02-13 Pengliang Yang , Rune Mittet

The finite difference time domain (FDTD) method has been successfully applied to obtain energies and wave functions for two electrons in a quantum dot modeled by a three dimensional harmonic potential. The FDTD method uses the…

Computational Physics · Physics 2017-06-12 I Wayan Sudiarta , Lily Maysari Angraini

Voice timbre attribute detection (vTAD) is the task of determining the relative intensity of timbre attributes between speech utterances. Voice timbre is a crucial yet inherently complex component of speech perception. While deep neural…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-06 Aemon Yat Fei Chiu , Yujia Xiao , Qiuqiang Kong , Tan Lee

The numerical investigation of acoustic damping materials, such as foams, constitutes a valuable enhancement to experimental testing. Typically, such materials are modeled in a homogenized way in order to reduce the computational effort and…

Numerical Analysis · Mathematics 2023-10-04 Lars Radtke , Paul Marter , Fabian Duvigneau , Sascha Eisenträger , Daniel Juhre , Alexander Düster

A sound source was proposed for acoustic measurements of physical models of the human vocal tract. The physical models are produced by Fast Prototyping, based on Magnetic Resonance Imaging during prolonged vowel production. The sound…

Instrumentation and Detectors · Physics 2017-11-22 Antti Hannukainen , Juha Kuortti , Jarmo Malinen , Antti Ojalammi

Vocal Percussion Transcription (VPT) is concerned with the automatic detection and classification of vocal percussion sound events, allowing music creators and producers to sketch drum lines on the fly. Classifier algorithms in VPT systems…

Sound · Computer Science 2022-04-12 Alejandro Delgado , Emir Demirel , Vinod Subramanian , Charalampos Saitis , Mark Sandler

In autonomous driving, dynamic environment and corner cases pose significant challenges to the robustness of ego vehicle's decision-making. To address these challenges, commencing with the representation of state-action mapping in the…

Computer Vision and Pattern Recognition · Computer Science 2025-03-04 Ziang Guo , Konstantin Gubernatorov , Selamawit Asfaw , Zakhar Yagudin , Dzmitry Tsetserukou

In this work, we present a new Vector Space Model (VSM) of speech utterances for the task of spoken dialect identification. Generally, DID systems are built using two sets of features that are extracted from speech utterances; acoustic and…

Computation and Language · Computer Science 2016-09-20 Sameer Khurana , Ahmed Ali , Steve Renals

Acoustic-to-articulatory inversion (AAI) is to convert audio into articulator movements, such as ultrasound tongue imaging (UTI) data. An issue of existing AAI methods is only using the personalized acoustic information to derive the…

Sound · Computer Science 2024-03-13 Yudong Yang , Rongfeng Su , Xiaokang Liu , Nan Yan , Lan Wang

Objective: To achieve accurate 3-D reconstruction and quantitative analysis of human retinal vasculature from a single optical coherence tomography angiography (OCTA) scan. Methods: We introduce Freqformer, a novel Transformer-based model…

Image and Video Processing · Electrical Eng. & Systems 2025-09-30 Lingyun Wang , Bingjie Wang , Jay Chhablani , Jose Alain Sahel , Shaohua Pi

Previous research in speech enhancement has mostly focused on modeling time or time-frequency domain information alone, with little consideration given to the potential benefits of simultaneously modeling both domains. Since these domains…

Sound · Computer Science 2023-05-16 Feng Dang , Qi Hu , Pengyuan Zhang , Yonghong Yan

Audio-visual multi-modal modeling has been demonstrated to be effective in many speech related tasks, such as speech recognition and speech enhancement. This paper introduces a new time-domain audio-visual architecture for target speaker…

Audio and Speech Processing · Electrical Eng. & Systems 2019-09-24 Jian Wu , Yong Xu , Shi-Xiong Zhang , Lian-Wu Chen , Meng Yu , Lei Xie , Dong Yu

Changing the vocal tract shape is one of the techniques which can be used by the players of wind instruments to modify the quality of the sound. It has been intensely studied in the case of reed instruments but has received only little…

Classical Physics · Physics 2016-01-22 R Auvray , Augustin Ernoult , S Terrien , B Fabre , C Vergez

A time-domain numerical modeling of transversely isotropic Biot poroelastic waves is proposed in two dimensions. The viscous dissipation occurring in the pores is described using the dynamic permeability model developed by…

Computational Physics · Physics 2015-06-24 Emilie Blanc , Guillaume Chiavassa , Bruno Lombard

Speaker verification (SV) aims to determine whether the speaker's identity of a test utterance is the same as the reference speech. In the past few years, extracting speaker embeddings using deep neural networks for SV systems has gone…

Sound · Computer Science 2022-05-27 Nan Zhang , Jianzong Wang , Zhenhou Hong , Chendong Zhao , Xiaoyang Qu , Jing Xiao

We introduce the joint time-frequency scattering transform, a time shift invariant descriptor of time-frequency structure for audio classification. It is obtained by applying a two-dimensional wavelet transform in time and log-frequency to…

Sound · Computer Science 2018-08-06 Joakim Andén , Vincent Lostanlen , Stéphane Mallat