English
Related papers

Related papers: EfficientNet-Absolute Zero for Continuous Speech K…

200 papers

We present our first efforts in building an automatic speech recognition system for Somali, an under-resourced language, using 1.57 hrs of annotated speech for acoustic model training. The system is part of an ongoing effort by the United…

Computation and Language · Computer Science 2018-07-24 Raghav Menon , Astik Biswas , Armin Saeb , John Quinn , Thomas Niesler

Zero-shot incremental learning aims to enable the model to generalize to new classes without forgetting previously learned classes. However, the semantic gap between old and new sample classes can lead to catastrophic forgetting.…

Computer Vision and Pattern Recognition · Computer Science 2024-02-13 Jie Ren , Yang Zhao , Weichuan Zhang , Changming Sun

User intent detection plays a critical role in question-answering and dialog systems. Most previous works treat intent detection as a classification problem where utterances are labeled with predefined intents. However, it is…

Computation and Language · Computer Science 2018-09-05 Congying Xia , Chenwei Zhang , Xiaohui Yan , Yi Chang , Philip S. Yu

Continuously learning new classes without catastrophic forgetting is a challenging problem for on-device acoustic event classification given the restrictions on computation resources (e.g., model size, running memory). To alleviate such an…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-23 Yang Xiao

In this research, we advanced a spoken language recognition system, moving beyond traditional feature vector-based models. Our improvements focused on effectively capturing language characteristics over extended periods using a specialized…

Sound · Computer Science 2025-01-22 Or Haim Anidjar , Roi Yozevitch

One of the challenges in developing a high quality custom keyword spotting (KWS) model is the lengthy and expensive process of collecting training data covering a wide range of languages, phrases and speaking styles. We introduce Synth4Kws…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-06 Pai Zhu , Dhruuv Agarwal , Jacob W. Bartel , Kurt Partridge , Hyun Jin Park , Quan Wang

This paper introduces the system submitted by the DKU-SMIIP team for the Auto-KWS 2021 Challenge. Our implementation consists of a two-stage keyword spotting system based on query-by-example spoken term detection and a speaker verification…

Audio and Speech Processing · Electrical Eng. & Systems 2021-04-13 Yechen Wang , Yan Jia , Murong Ma , Zexin Cai , Ming Li

We describe a modern deep learning system that automatically identifies informative contextual examples (\qu{contexts}) for first language vocabulary instruction for high school student. Our paper compares three modeling approaches: (i) an…

Computation and Language · Computer Science 2026-02-23 Tao Wu , Adam Kapelner

FrameNet is a computational linguistics resource composed of semantic frames, high-level concepts that represent the meanings of words. In this paper, we present an approach to gather frame disambiguation annotations in sentences using a…

Computation and Language · Computer Science 2018-08-21 Anca Dumitrache , Lora Aroyo , Chris Welty

Accurate recognition of aviation commands is vital for flight safety and efficiency, as pilots must follow air traffic control instructions precisely. This paper addresses challenges in speech command recognition, such as noisy environments…

Sound · Computer Science 2024-07-01 Yuanxi Lin , Tonglin Zhou , Yang Xiao

Phrase detection requires methods to identify if a phrase is relevant to an image and localize it, if applicable. A key challenge for training more discriminative detection models is sampling negatives. Sampling techniques from prior work…

Computer Vision and Pattern Recognition · Computer Science 2022-11-16 Maan Qraitem , Bryan A. Plummer

In this demo, we present a compact intelligent audio system-on-chip (SoC) integrated with a keyword spotting accelerator, enabling ultra-low latency, low-power, and low-cost voice interaction in Internet of Things (IoT) devices. Through…

Sound · Computer Science 2025-09-30 Huihong Liang , Dongxuan Jia , Youquan Wang , Longtao Huang , Shida Zhong , Luping Xiang , Lei Huang , Tao Yuan

Keyphrase generation aims to summarize long documents with a collection of salient phrases. Deep neural models have demonstrated a remarkable success in this task, capable of predicting keyphrases that are even absent from a document.…

Computation and Language · Computer Science 2021-04-20 Xianjie Shen , Yinghan Wang , Rui Meng , Jingbo Shang

The objective of this paper is speaker recognition under noisy and unconstrained conditions. We make two key contributions. First, we introduce a very large-scale audio-visual speaker recognition dataset collected from open-source media.…

Sound · Computer Science 2020-11-05 Joon Son Chung , Arsha Nagrani , Andrew Zisserman

This paper focuses on the problem of query by example spoken term detection (QbE-STD) in zero-resource scenario. State-of-the-art approaches primarily rely on dynamic time warping (DTW) based template matching techniques using phone…

Audio and Speech Processing · Electrical Eng. & Systems 2019-11-20 Dhananjay Ram , Lesly Miculicich , Hervé Bourlard

Speech enhancement (SE) aims to suppress the additive noise from a noisy speech signal to improve the speech's perceptual quality and intelligibility. However, the over-suppression phenomenon in the enhanced speech might degrade the…

Audio and Speech Processing · Electrical Eng. & Systems 2022-04-11 Yuchen Hu , Nana Hou , Chen Chen , Eng Siong Chng

In this paper, we propose DS-KWS, a two-stage framework for robust user-defined keyword spotting. It combines a CTC-based method with a streaming phoneme search module to locate candidate segments, followed by a QbyT-based method with a…

Sound · Computer Science 2025-10-14 Zhiqi Ai , Han Cheng , Yuxin Wang , Shiyi Mu , Shugong Xu , Yongjin Zhou

Smart systems have been massively developed to help humans in various tasks. Deep Learning technologies push even further in creating accurate assistant systems due to the explosion of data lakes. One of the smart system tasks is to…

Computer Vision and Pattern Recognition · Computer Science 2021-05-26 Dhomas Hatta Fudholi , Yurio Windiatmoko , Nurdi Afrianto , Prastyo Eko Susanto , Magfirah Suyuti , Ahmad Fathan Hidayatullah , Ridho Rahmadi

Semantic matching is a mainstream paradigm of zero-shot relation extraction, which matches a given input with a corresponding label description. The entities in the input should exactly match their hypernyms in the description, while the…

Computation and Language · Computer Science 2023-06-09 Jun Zhao , Wenyu Zhan , Xin Zhao , Qi Zhang , Tao Gui , Zhongyu Wei , Junzhe Wang , Minlong Peng , Mingming Sun

Robustness against noise is critical for keyword spotting (KWS) in real-world environments. To improve the robustness, a speech enhancement front-end is involved. Instead of treating the speech enhancement as a separated preprocessing…

Sound · Computer Science 2019-06-21 Yue Gu , Zhihao Du , Hui Zhang , Xueliang Zhang
‹ Prev 1 8 9 10 Next ›