中文
相关论文

相关论文: Robust Phonetic Segmentation Using Spectral Transi…

200 篇论文

Mismatched transcriptions have been proposed as a mean to acquire probabilistic transcriptions from non-native speakers of a language.Prior work has demonstrated the value of these transcriptions by successfully adapting cross-lingual ASR…

计算与语言 · 计算机科学 2017-01-16 Xiang Kong , Preethi Jyothi , Mark Hasegawa-Johnson

Traditionally, the performance of non-native mispronunciation verification systems relied on effective phone-level labelling of non-native corpora. In this study, a multi-view approach is proposed to incorporate discriminative feature…

音频与语音处理 · 电气工程与系统科学 2020-09-10 Zhenyu Wang , John H. L. Hansen , Yanlu Xie

In this paper, we analyse the error patterns of the raw waveform acoustic models in TIMIT's phone recognition task. Our analysis goes beyond the conventional phone error rate (PER) metric. We categorise the phones into three groups:…

声音 · 计算机科学 2024-06-04 Erfan Loweimi , Andrea Carmantini , Peter Bell , Steve Renals , Zoran Cvetkovic

A major challenge in image segmentation is classifying object boundaries. Recent efforts propose to refine the segmentation result with boundary masks. However, models are still prone to misclassifying boundary pixels even when they…

计算机视觉与模式识别 · 计算机科学 2022-03-15 Han Zhang , Zihao Zhang , Wenhao Zheng , Wei Xu

In this paper, we propose a neural-based coding scheme in which an artificial neural network is exploited to automatically compress and decompress speech signals by a trainable approach. Having a two-stage training phase, the system can be…

声音 · 计算机科学 2016-01-25 Mahmood Yousefi-Azar , Farbod Razzazi

Partial deepfake speech detection requires identifying manipulated regions that may occur within short temporal portions of an otherwise bona fide utterance, making the task particularly challenging for conventional utterance-level…

声音 · 计算机科学 2026-04-06 Inbal Rimon , Oren Gal , Haim Permuter

Connectionist temporal classification (CTC) provides an end-to-end acoustic model (AM) training strategy. CTC learns accurate AMs without time-aligned phonetic transcription, but sometimes fails to converge, especially in…

音频与语音处理 · 电气工程与系统科学 2019-02-28 Di He , Xuesong Yang , Boon Pang Lim , Yi Liang , Mark Hasegawa-Johnson , Deming Chen

The goal of Audio-Visual Segmentation (AVS) is to localize and segment the sounding source objects from video frames. Research on AVS suffers from data scarcity due to the high cost of fine-grained manual annotations. Recent works attempt…

计算机视觉与模式识别 · 计算机科学 2025-05-30 Kyungbok Lee , You Zhang , Zhiyao Duan

Speech processing systems rely on robust feature extraction to handle phonetic and semantic variations found in natural language. While techniques exist for desensitizing features to common noise patterns produced by Speech-to-Text (STT)…

音频与语音处理 · 电气工程与系统科学 2019-06-14 Chris Larson , Tarek Lahlou , Diana Mingels , Zachary Kulis , Erik Mueller

Existing Machine Translation (MT) research often suggests a single, fixed set of hyperparameters for word segmentation models, symmetric Byte Pair Encoding (BPE), which applies the same number of merge operations (NMO) to train tokenizers…

计算与语言 · 计算机科学 2026-02-16 Saumitra Yadav , Manish Shrivastava

This work introduces sequential neural beamforming, which alternates between neural network based spectral separation and beamforming based spatial separation. Our neural networks for separation use an advanced convolutional architecture…

Existing speech enhancement methods mainly separate speech from noises at the signal level or in the time-frequency domain. They seldom pay attention to the semantic information of a corrupted signal. In this paper, we aim to bridge this…

音频与语音处理 · 电气工程与系统科学 2021-04-09 Yajing Liu , Xiulian Peng , Zhiwei Xiong , Yan Lu

Objective: Speech tests aim to estimate discrimination loss or speech recognition threshold (SRT). This paper investigates the potential to estimate SRTs from clinical data that target at characterizing the discrimination loss. Knowledge…

声音 · 计算机科学 2025-01-16 Mareike Buhl , Eugen Kludt , Lena Schell-Majoor , Paul Avan , Marta Campi

Phonetic segmentation is the process of splitting speech into distinct phonetic units. Human experts routinely perform this task manually by analyzing auditory and visual cues using analysis software, which is an extremely time-consuming…

人机交互 · 计算机科学 2018-05-14 Arif Khan , Ingmar Steiner , Yusuke Sugano , Andreas Bulling , Ross Macdonald

This paper proposes a new pitch estimator and a novel pitch tracker for speakers. We first decompose the sound signal into subbands using an auditory filterbank, assuming time-frequency sparsity of human speech. Instead of directly…

音频与语音处理 · 电气工程与系统科学 2026-04-03 Shoufeng Lin

In this work, a new Physics laboratory experiment on Acoustics beats is presented. We have designed a simple experimental setup to study superposition of sound waves of slightly different frequencies (acoustic beat). The microphone of a…

Recent works have introduced methods to estimate segmentation performance without ground truth, relying solely on neural network softmax outputs. These techniques hold potential for intuitive output quality control. However, such…

图像与视频处理 · 电气工程与系统科学 2024-08-30 Anna M. Wundram , Paul Fischer , Michael Muehlebach , Lisa M. Koch , Christian F. Baumgartner

Spoken content, such as online videos and podcasts, often spans multiple topics, which makes automatic topic segmentation essential for user navigation and downstream applications. However, current methods do not fully leverage acoustic…

计算与语言 · 计算机科学 2026-02-09 Steffen Freisinger , Philipp Seeberger , Tobias Bocklet , Korbinian Riedhammer

Learned speech representations can drastically improve performance on tasks with limited labeled data. However, due to their size and complexity, learned representations have limited utility in mobile settings where run-time performance can…

声音 · 计算机科学 2022-12-20 Jacob Peplinski , Joel Shor , Sachin Joglekar , Jake Garrison , Shwetak Patel

The end-to-end architecture has made promising progress in speech translation (ST). However, the ST task is still challenging under low-resource conditions. Most ST models have shown unsatisfactory results, especially in the absence of word…

计算与语言 · 计算机科学 2022-03-31 Yao-Fei Cheng , Hung-Shin Lee , Hsin-Min Wang