English
Related papers

Related papers: Automatic vocal tract landmark localization from m…

200 papers

Segmenting vocal tract articulators in real-time MRI (rtMRI) is a challenging dynamic image segmentation problem characterized by low contrast, rapid motion, and limited spatial resolution. However, while rtMRI acquisitions may provide…

At present, deep neural network methods have played a dominant role in face alignment field. However, they generally use predefined network structures to predict landmarks, which tends to learn general features and leads to mediocre…

Computer Vision and Pattern Recognition · Computer Science 2024-12-03 Jun Wan , He Liu , Yujia Wu , Zhihui Lai , Wenwen Min , Jun Liu

We propose SpeakerNet - a new neural architecture for speaker recognition and speaker verification tasks. It is composed of residual blocks with 1D depth-wise separable convolutions, batch-normalization, and ReLU layers. This architecture…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-27 Nithin Rao Koluguri , Jason Li , Vitaly Lavrukhin , Boris Ginsburg

Vocal tract articulation is a natural, grounded control space of speech production. The spatiotemporal coordination of articulators combined with the vocal source shapes intelligible speech sounds to enable effective spoken communication.…

Audio and Speech Processing · Electrical Eng. & Systems 2025-03-04 Cheol Jun Cho , Peter Wu , Tejas S. Prabhune , Dhruv Agarwal , Gopala K. Anumanchipalli

Deep learning architectures have made significant progress in terms of performance in many research areas. The automatic speech recognition (ASR) field has thus benefited from these scientific and technological advances, particularly for…

Sound · Computer Science 2024-03-01 Quentin Raymondaud , Mickael Rouvier , Richard Dufour

Rapid advances in 3D model scanning have enabled the mass digitization of dental clay models. However, most clinicians and researchers continue to use manual morphometric analysis methods on these models such as landmarking. This is a…

Image and Video Processing · Electrical Eng. & Systems 2025-01-28 Artur Agaronyan , HyeRan Choo , Marius Linguraru , Syed Muhammad Anwar

We address the problem of vehicle self-localization from multi-modal sensor information and a reference map. The map is generated off-line by extracting landmarks from the vehicle's field of view, while the measurements are collected…

Robotics · Computer Science 2019-07-22 Nico Engel , Stefan Hoermann , Markus Horn , Vasileios Belagiannis , Klaus Dietmayer

Different transformer architectures implement identical linguistic computations via distinct connectivity patterns, yielding model imprinted ``computational fingerprints'' detectable through spectral analysis. Using graph signal processing…

Computation and Language · Computer Science 2025-10-23 Valentin Noël

The study of speech disorders can benefit greatly from time-aligned data. However, audio-text mismatches in disfluent speech cause rapid performance degradation for modern speech aligners, hindering the use of automatic approaches. In this…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-05 Theodoros Kouzelis , Georgios Paraskevopoulos , Athanasios Katsamanis , Vassilis Katsouros

Segmentation of the Left ventricle (LV) is a crucial step for quantitative measurements such as area, volume, and ejection fraction. However, the automatic LV segmentation in 2D echocardiographic images is a challenging task due to…

Image and Video Processing · Electrical Eng. & Systems 2019-12-24 Shakiba Moradi , Mostafa Ghelich-Oghli , Azin Alizadehasl , Isaac Shiri , Niki Oveisi , Mehrdad Oveisi , Majid Maleki , Jan Dhooge

Multi-channel multi-talker speech recognition presents formidable challenges in the realm of speech processing, marked by issues such as background noise, reverberation, and overlapping speech. Overcoming these complexities requires…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-09 Yiwen Shao

Dysarthric speech reconstruction (DSR) aims to transform dysarthric speech into normal speech by improving the intelligibility and naturalness. This is a challenging task especially for patients with severe dysarthria and speaking in…

Sound · Computer Science 2024-02-01 Xueyuan Chen , Yuejiao Wang , Xixin Wu , Disong Wang , Zhiyong Wu , Xunying Liu , Helen Meng

Dysarthric speech severity assessment typically requires trained clinicians or supervised models built from labelled pathological speech, limiting scalability across languages and clinical settings. We present a training-free method that…

Computation and Language · Computer Science 2026-04-14 Bernard Muller , Antonio Armando Ortiz Barrañón , LaVonne Roberts

3D ultrasound (US) can facilitate detailed prenatal examinations for fetal growth monitoring. To analyze a 3D US volume, it is fundamental to identify anatomical landmarks of the evaluated organs accurately. Typical deep learning methods…

Computer Vision and Pattern Recognition · Computer Science 2020-04-02 Chaoyu Chen , Xin Yang , Ruobing Huang , Wenlong Shi , Shengfeng Liu , Mingrong Lin , Yuhao Huang , Yong Yang , Yuanji Zhang , Huanjia Luo , Yankai Huang , Yi Xiong , Dong Ni

Dysarthria speech contains the pathological characteristics of vocal tract and vocal fold, but so far, they have not yet been included in traditional acoustic feature sets. Moreover, the nonlinearity and non-stationarity of speech have been…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-02 Ting Zhu , Shufei Duan , Camille Dingam , Huizhi Liang , Wei Zhang

Deep learning networks have shown promising performance for accurate object localization in medial images, but require large amount of annotated data for supervised training, which is expensive and expertise burdensome. To address this…

Computer Vision and Pattern Recognition · Computer Science 2021-05-26 Wenhui Lei , Wei Xu , Ran Gu , Hao Fu , Shaoting Zhang , Guotai Wang

Landmark detection is a critical component of the image processing pipeline for automated aortic size measurements. Given that the thoracic aorta has a relatively conserved topology across the population and that a human annotator with…

Image and Video Processing · Electrical Eng. & Systems 2023-04-18 Zhangxing Bian , Jiayang Zhong , Yanglong Lu , Charles R. Hatt , Nicholas S. Burris

Speech sound disorders are a common communication impairment in childhood. Because speech disorders can negatively affect the lives and the development of children, clinical intervention is often recommended. To help with diagnosis and…

Audio and Speech Processing · Electrical Eng. & Systems 2021-03-02 Manuel Sam Ribeiro , Joanne Cleland , Aciel Eshky , Korin Richmond , Steve Renals

Statistical shape analysis is a very useful tool in a wide range of medical and biological applications. However, it typically relies on the ability to produce a relatively small number of features that can capture the relevant variability…

Computer Vision and Pattern Recognition · Computer Science 2020-06-16 Riddhish Bhalodia , Ladislav Kavan , Ross Whitaker

Pattern recognition from audio signals is an active research topic encompassing audio tagging, acoustic scene classification, music classification, and other areas. Spectrogram and mel-frequency cepstral coefficients (MFCC) are among the…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-18 Md. Istiaq Ansari , Taufiq Hasan