中文
相关论文

相关论文: Automatic vocal tract landmark localization from m…

200 篇论文

Segmenting vocal tract articulators in real-time MRI (rtMRI) is a challenging dynamic image segmentation problem characterized by low contrast, rapid motion, and limited spatial resolution. However, while rtMRI acquisitions may provide…

At present, deep neural network methods have played a dominant role in face alignment field. However, they generally use predefined network structures to predict landmarks, which tends to learn general features and leads to mediocre…

计算机视觉与模式识别 · 计算机科学 2024-12-03 Jun Wan , He Liu , Yujia Wu , Zhihui Lai , Wenwen Min , Jun Liu

We propose SpeakerNet - a new neural architecture for speaker recognition and speaker verification tasks. It is composed of residual blocks with 1D depth-wise separable convolutions, batch-normalization, and ReLU layers. This architecture…

音频与语音处理 · 电气工程与系统科学 2020-10-27 Nithin Rao Koluguri , Jason Li , Vitaly Lavrukhin , Boris Ginsburg

Vocal tract articulation is a natural, grounded control space of speech production. The spatiotemporal coordination of articulators combined with the vocal source shapes intelligible speech sounds to enable effective spoken communication.…

音频与语音处理 · 电气工程与系统科学 2025-03-04 Cheol Jun Cho , Peter Wu , Tejas S. Prabhune , Dhruv Agarwal , Gopala K. Anumanchipalli

Deep learning architectures have made significant progress in terms of performance in many research areas. The automatic speech recognition (ASR) field has thus benefited from these scientific and technological advances, particularly for…

声音 · 计算机科学 2024-03-01 Quentin Raymondaud , Mickael Rouvier , Richard Dufour

Rapid advances in 3D model scanning have enabled the mass digitization of dental clay models. However, most clinicians and researchers continue to use manual morphometric analysis methods on these models such as landmarking. This is a…

图像与视频处理 · 电气工程与系统科学 2025-01-28 Artur Agaronyan , HyeRan Choo , Marius Linguraru , Syed Muhammad Anwar

We address the problem of vehicle self-localization from multi-modal sensor information and a reference map. The map is generated off-line by extracting landmarks from the vehicle's field of view, while the measurements are collected…

机器人学 · 计算机科学 2019-07-22 Nico Engel , Stefan Hoermann , Markus Horn , Vasileios Belagiannis , Klaus Dietmayer

Different transformer architectures implement identical linguistic computations via distinct connectivity patterns, yielding model imprinted ``computational fingerprints'' detectable through spectral analysis. Using graph signal processing…

计算与语言 · 计算机科学 2025-10-23 Valentin Noël

The study of speech disorders can benefit greatly from time-aligned data. However, audio-text mismatches in disfluent speech cause rapid performance degradation for modern speech aligners, hindering the use of automatic approaches. In this…

音频与语音处理 · 电气工程与系统科学 2023-06-05 Theodoros Kouzelis , Georgios Paraskevopoulos , Athanasios Katsamanis , Vassilis Katsouros

Segmentation of the Left ventricle (LV) is a crucial step for quantitative measurements such as area, volume, and ejection fraction. However, the automatic LV segmentation in 2D echocardiographic images is a challenging task due to…

图像与视频处理 · 电气工程与系统科学 2019-12-24 Shakiba Moradi , Mostafa Ghelich-Oghli , Azin Alizadehasl , Isaac Shiri , Niki Oveisi , Mehrdad Oveisi , Majid Maleki , Jan Dhooge

Multi-channel multi-talker speech recognition presents formidable challenges in the realm of speech processing, marked by issues such as background noise, reverberation, and overlapping speech. Overcoming these complexities requires…

音频与语音处理 · 电气工程与系统科学 2023-10-09 Yiwen Shao

Dysarthric speech reconstruction (DSR) aims to transform dysarthric speech into normal speech by improving the intelligibility and naturalness. This is a challenging task especially for patients with severe dysarthria and speaking in…

声音 · 计算机科学 2024-02-01 Xueyuan Chen , Yuejiao Wang , Xixin Wu , Disong Wang , Zhiyong Wu , Xunying Liu , Helen Meng

Dysarthric speech severity assessment typically requires trained clinicians or supervised models built from labelled pathological speech, limiting scalability across languages and clinical settings. We present a training-free method that…

计算与语言 · 计算机科学 2026-04-14 Bernard Muller , Antonio Armando Ortiz Barrañón , LaVonne Roberts

3D ultrasound (US) can facilitate detailed prenatal examinations for fetal growth monitoring. To analyze a 3D US volume, it is fundamental to identify anatomical landmarks of the evaluated organs accurately. Typical deep learning methods…

计算机视觉与模式识别 · 计算机科学 2020-04-02 Chaoyu Chen , Xin Yang , Ruobing Huang , Wenlong Shi , Shengfeng Liu , Mingrong Lin , Yuhao Huang , Yong Yang , Yuanji Zhang , Huanjia Luo , Yankai Huang , Yi Xiong , Dong Ni

Dysarthria speech contains the pathological characteristics of vocal tract and vocal fold, but so far, they have not yet been included in traditional acoustic feature sets. Moreover, the nonlinearity and non-stationarity of speech have been…

音频与语音处理 · 电气工程与系统科学 2024-01-02 Ting Zhu , Shufei Duan , Camille Dingam , Huizhi Liang , Wei Zhang

Deep learning networks have shown promising performance for accurate object localization in medial images, but require large amount of annotated data for supervised training, which is expensive and expertise burdensome. To address this…

计算机视觉与模式识别 · 计算机科学 2021-05-26 Wenhui Lei , Wei Xu , Ran Gu , Hao Fu , Shaoting Zhang , Guotai Wang

Landmark detection is a critical component of the image processing pipeline for automated aortic size measurements. Given that the thoracic aorta has a relatively conserved topology across the population and that a human annotator with…

图像与视频处理 · 电气工程与系统科学 2023-04-18 Zhangxing Bian , Jiayang Zhong , Yanglong Lu , Charles R. Hatt , Nicholas S. Burris

Speech sound disorders are a common communication impairment in childhood. Because speech disorders can negatively affect the lives and the development of children, clinical intervention is often recommended. To help with diagnosis and…

音频与语音处理 · 电气工程与系统科学 2021-03-02 Manuel Sam Ribeiro , Joanne Cleland , Aciel Eshky , Korin Richmond , Steve Renals

Statistical shape analysis is a very useful tool in a wide range of medical and biological applications. However, it typically relies on the ability to produce a relatively small number of features that can capture the relevant variability…

计算机视觉与模式识别 · 计算机科学 2020-06-16 Riddhish Bhalodia , Ladislav Kavan , Ross Whitaker

Pattern recognition from audio signals is an active research topic encompassing audio tagging, acoustic scene classification, music classification, and other areas. Spectrogram and mel-frequency cepstral coefficients (MFCC) are among the…

音频与语音处理 · 电气工程与系统科学 2022-11-18 Md. Istiaq Ansari , Taufiq Hasan