中文
相关论文

相关论文: Multilingual Bottleneck Features for Query by Exam…

200 篇论文

This paper focuses on the problem of query by example spoken term detection (QbE-STD) in zero-resource scenario. State-of-the-art approaches primarily rely on dynamic time warping (DTW) based template matching techniques using phone…

音频与语音处理 · 电气工程与系统科学 2019-11-20 Dhananjay Ram , Lesly Miculicich , Hervé Bourlard

Query-by-example spoken term detection (QbE-STD) is typically constrained by transcribed data scarcity and language specificity. This paper introduces a novel, language-agnostic QbE-STD model leveraging image processing techniques and…

机器学习 · 计算机科学 2024-10-08 Allahdadi Fatemeh , Mahdian Toroghi Rahil , Zareian Hassan

Query-by-example spoken term detection (QbE-STD) searches for matching words or phrases in an audio dataset using a sample spoken query. When annotated data is limited or unavailable, QbE-STD is often done using template matching methods…

音频与语音处理 · 电气工程与系统科学 2025-06-23 Akanksha Singh , Yi-Ping Phoebe Chen , Vipul Arora

Retrieving spoken content with spoken queries, or query-by- example spoken term detection (STD), is attractive because it makes possible the matching of signals directly on the acoustic level without transcribing them into text. Here, we…

计算与语言 · 计算机科学 2018-04-30 Chia-Wei Ao , Hung-yi Lee

Query-by-example (QbE) speech search is the task of matching spoken queries to utterances within a search collection. In low- or zero-resource settings, QbE search is often addressed with approaches based on dynamic time warping (DTW).…

计算与语言 · 计算机科学 2020-11-25 Yushi Hu , Shane Settle , Karen Livescu

A number of recent studies have started to investigate how speech systems can be trained on untranscribed speech by leveraging accompanying images at training time. Examples of tasks include keyword prediction and within- and across-mode…

计算与语言 · 计算机科学 2019-04-16 Herman Kamper , Aristotelis Anastassiou , Karen Livescu

How can we effectively develop speech technology for languages where no transcribed data is available? Many existing approaches use no annotated resources at all, yet it makes sense to leverage information from large annotated corpora in…

计算与语言 · 计算机科学 2018-11-12 Enno Hermann , Sharon Goldwater

We propose to learn acoustic word embeddings with temporal context for query-by-example (QbE) speech search. The temporal context includes the leading and trailing word sequences of a word. We assume that there exist spoken word pairs in…

计算与语言 · 计算机科学 2018-06-19 Yougen Yuan , Cheung-Chi Leung , Lei Xie , Hongjie Chen , Bin Ma , Haizhou Li

Pre-trained speech representations like wav2vec 2.0 are a powerful tool for automatic speech recognition (ASR). Yet many endangered languages lack sufficient data for pre-training such models, or are predominantly oral vernaculars without a…

Fast and accurate spoken content retrieval is vital for applications such as voice search. Query-by-Example Spoken Term Detection (STD) involves retrieving matching segments from an audio database given a spoken query. Token-based STD…

音频与语音处理 · 电气工程与系统科学 2026-02-19 Anup Singh , Vipul Arora , Kris Demuynck

Traditional Query-by-Example (QbE) speech search approaches usually use methods based on frame-level features, while state-of-the-art approaches tend to use models based on acoustic word embeddings (AWEs) to transform variable length audio…

音频与语音处理 · 电气工程与系统科学 2021-09-21 Yuguang Yang , Yu Pan , Xin Dong , Minqiang Xu

Sleep-disordered breathing (SDB) is a serious and prevalent condition, and acoustic analysis via consumer devices (e.g. smartphones) offers a low-cost solution to screening for it. We present a novel approach for the acoustic identification…

音频与语音处理 · 电气工程与系统科学 2019-04-08 Hector E. Romero , Ning Ma , Guy J. Brown , Amy V. Beeston , Madina Hasan

In this paper we proposed an end-to-end short utterances speech language identification(SLD) approach based on a Long Short Term Memory (LSTM) neural network which is special suitable for SLD application in intelligent vehicles. Features…

计算与语言 · 计算机科学 2020-02-04 Zhanyu Ma , Hong Yu

Spoken term detection (STD) is often hindered by reliance on frame-level features and the computationally intensive DTW-based template matching, limiting its practicality. To address these challenges, we propose a novel approach that…

音频与语音处理 · 电气工程与系统科学 2024-12-24 Anup Singh , Kris Demuynck , Vipul Arora

In this paper, we present a time-contrastive learning (TCL) based bottleneck (BN)feature extraction method for speech signals with an application to text-dependent (TD) speaker verification (SV). It is well-known that speech signals exhibit…

声音 · 计算机科学 2019-05-14 Achintya Kr. Sarkar , Zheng-Hua Tan

Existing keyword spotting (KWS) systems primarily rely on predefined keyword phrases. However, the ability to recognize customized keywords is crucial for tailoring interactions with intelligent devices. In this paper, we present a novel…

计算与语言 · 计算机科学 2024-11-26 Zhenyu Wang , Shuyu Kong , Li Wan , Biqiao Zhang , Yiteng Huang , Mumin Jin , Ming Sun , Xin Lei , Zhaojun Yang

Keyword spotting is an important research field because it plays a key role in device wake-up and user interaction on smart devices. However, it is challenging to minimize errors while operating efficiently in devices with limited resources…

声音 · 计算机科学 2023-07-06 Byeonggeun Kim , Simyung Chang , Jinkyu Lee , Dooyong Sung

In this paper, we compare two paradigms for unsupervised discovery of structured acoustic tokens directly from speech corpora without any human annotation. The Multigranular Paradigm seeks to capture all available information in the corpora…

计算与语言 · 计算机科学 2017-11-29 Cheng-Tao Chung , Lin-Shan Lee

We compare features for dynamic time warping (DTW) when used to bootstrap keyword spotting (KWS) in an almost zero-resource setting. Such quickly-deployable systems aim to support United Nations (UN) humanitarian relief efforts in parts of…

音频与语音处理 · 电气工程与系统科学 2019-07-16 Raghav Menon , Herman Kamper , Ewald van der Westhuizen , John Quinn , Thomas Niesler

There are a number of studies about extraction of bottleneck (BN) features from deep neural networks (DNNs)trained to discriminate speakers, pass-phrases and triphone states for improving the performance of text-dependent speaker…

声音 · 计算机科学 2019-05-14 Achintya kr. Sarkar , Zheng-Hua Tan , Hao Tang , Suwon Shon , James Glass
‹ 上一页 1 2 3 10 下一页 ›