English
Related papers

Related papers: VOP Detection for Read and Conversation Speech usi…

200 papers

We investigate recent transformer networks pre-trained for automatic speech recognition for their ability to detect speaker and language changes in speech. We do this by simply adding speaker (change) or language targets to the labels. For…

Audio and Speech Processing · Electrical Eng. & Systems 2023-02-21 Tijn Berns , Nik Vaessen , David A. van Leeuwen

Change detection is a fundamental task in remote sensing, aiming to quantify the impacts of human activities and ecological dynamics on land-cover changes. Existing change detection methods are limited to predefined classes in training…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 You Su , Yonghong Song , Jingqi Chen , Zehan Wen

In this paper, for real time enhancement of noisy speech, a method of threshold determination based on modeling of Teager energy (TE) operated perceptual wavelet packet (PWP) coefficients of the noisy speech and noise by an Erlang-2 PDF is…

Audio and Speech Processing · Electrical Eng. & Systems 2018-02-13 Md Tauhidul Islam , Celia Shahnaz , Wei-Ping Zhu , M. Omair Ahmad

Automated detection of voice disorders with computational methods is a recent research area in the medical domain since it requires a rigorous endoscopy for the accurate diagnosis. Efficient screening methods are required for the diagnosis…

Quantitative Methods · Quantitative Biology 2018-12-06 Vibhuti Gupta

Audio-visual semantic segmentation (AVSS) represents an extension of the audio-visual segmentation (AVS) task, necessitating a semantic understanding of audio-visual scenes beyond merely identifying sound-emitting objects at the visual…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Yujian Lee , Peng Gao , Yongqi Xu , Wentao Fan

Current approaches in pulse detection use domain transformations so as to concentrate frequency related information that can be distinguishable from noise. In real cases we do not know when the pulse will begin, so we need a time search…

Information Retrieval · Computer Science 2007-05-23 Jaime Gomez , Ignacio Melgar , Juan Seijas

We introduce an unsupervised approach for correcting highly imperfect speech transcriptions based on a decision-level fusion of stemming and two-way phoneme pruning. Transcripts are acquired from videos by extracting audio using Ffmpeg…

Computation and Language · Computer Science 2021-07-28 Sunakshi Mehra , Seba Susan

Target-speaker voice activity detection is currently a promising approach for speaker diarization in complex acoustic environments. This paper presents a novel Sequence-to-Sequence Target-Speaker Voice Activity Detection (Seq2Seq-TSVAD)…

Audio and Speech Processing · Electrical Eng. & Systems 2023-02-21 Ming Cheng , Weiqing Wang , Yucong Zhang , Xiaoyi Qin , Ming Li

Traditional object detection systems are typically constrained to predefined categories, limiting their applicability in dynamic environments. In contrast, open-vocabulary object detection (OVD) enables the identification of objects from…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Tianyi Zhang , Antoine Simoulin , Kai Li , Sana Lakdawala , Shiqing Yu , Arpit Mittal , Hongyu Fu , Yu Lin

A demonstration of a real-time and continuous turn-taking prediction system is presented. The system is based on a voice activity projection (VAP) model, which directly maps dialogue stereo audio to future voice activities. The VAP model…

Computation and Language · Computer Science 2024-01-11 Koji Inoue , Bing'er Jiang , Erik Ekstedt , Tatsuya Kawahara , Gabriel Skantze

In this paper a new method for recognition of consonant-vowel phonemes combination on a new Persian speech dataset titled as PCVC (Persian Consonant-Vowel Combination) is proposed which is used to recognize Persian phonemes. In PCVC…

Sound · Computer Science 2018-12-18 Saber Malekzadeh , Mohammad Hossein Gholizadeh , Seyed Naser Razavi

Open-Vocabulary Segmentation (OVS) aims to segment classes that are not present in the training dataset. However, most existing studies assume that the training data is fixed in advance, overlooking more practical scenarios where new…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Dongjun Hwang , Yejin Kim , Minyoung Lee , Seong Joon Oh , Junsuk Choe

Audio Word2Vec offers vector representations of fixed dimensionality for variable-length audio segments using Sequence-to-sequence Autoencoder (SA). These vector representations are shown to describe the sequential phonetic structures of…

Computation and Language · Computer Science 2018-02-20 Chia-Hao Shen , Janet Y. Sung , Hung-Yi Lee

Rapid advances in speech synthesis and audio editing have made realistic forgeries increasingly accessible, yet existing detection methods remain vulnerable to tampering or depend on visual/wearable sensors. In this paper, we present…

Human-Computer Interaction · Computer Science 2026-03-31 Mingda Han , Huanqi Yang , Chaoqun Li , Wenhao Li , Guoming Zhang , Yanni Yang , Yetong Cao , Weitao Xu , Pengfei Hu

Body-worn video (BWV) cameras are increasingly utilized by police departments to provide a record of police-public interactions. However, large-scale BWV deployment produces terabytes of data per week, necessitating the development of…

Computer Vision and Pattern Recognition · Computer Science 2016-10-21 Stephanie Allen , David Madras , Ye Ye , Greg Zanotti

Vision-language models (VLMs), such as CLIP, have demonstrated impressive zero-shot capabilities for various downstream tasks. Their performance can be further enhanced through few-shot prompt tuning methods. However, current studies…

Computer Vision and Pattern Recognition · Computer Science 2024-06-04 Zhi Zhou , Ming Yang , Jiang-Xin Shi , Lan-Zhe Guo , Yu-Feng Li

The purpose of this study is to detect the mismatch between text script and voice-over. For this, we present a novel utterance verification (UV) method, which calculates the degree of correspondence between a voice-over and the phoneme…

Audio and Speech Processing · Electrical Eng. & Systems 2020-03-23 Yoonjae Jeong , Hoon-Young Cho

Answering questions related to audio-visual scenes, i.e., the AVQA task, is becoming increasingly popular. A critical challenge is accurately identifying and tracking sounding objects related to the question along the timeline. In this…

Multimedia · Computer Science 2024-12-17 Zhangbin Li , Jinxing Zhou , Jing Zhang , Shengeng Tang , Kun Li , Dan Guo

Whispered speech is produced when the vocal folds are not used, either intentionally, or due to a temporary or permanent voice condition. The essential difference between natural speech and whispered speech is that periodic signal…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-10 Aníbal J. S. Ferreira , Luis M. T. Jesus , Laurentino M. M. Leal , Jorge E. F. Spratley

In this work, we address the task of voice conversion (VC) using a vector-based interface. To align audio embeddings across speakers, we employ discrete optimal transport (OT) and approximate the transport map using the barycentric…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-02 Anton Selitskiy , Maitreya Kocharekar