English
Related papers

Related papers: Wake Word Detection Based on Res2Net

200 papers

In this era of Big Data, due to expeditious exchange of information on the web, words are being used to denote newer meanings, causing linguistic shift. With the recent availability of large amounts of digitized texts, an automated analysis…

Computation and Language · Computer Science 2018-12-17 Abhik Jana , Animesh Mukherjee , Pawan Goyal

We identify agreement and disagreement between utterances that express stances towards a topic of discussion. Existing methods focus mainly on conversational settings, where dialogic features are used for (dis)agreement inference. We extend…

Computation and Language · Computer Science 2019-06-05 Chang Xu , Cecile Paris , Surya Nepal , Ross Sparks

In speaker verification systems, the utilization of short utterances presents a persistent challenge, leading to performance degradation primarily due to insufficient phonetic information to characterize the speakers. To overcome this…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-12 Seung-bin Kim , Chan-yeong Lim , Jungwoo Heo , Ju-ho Kim , Hyun-seo Shin , Kyo-Won Koo , Ha-Jin Yu

The increasing reliability of automatic speech recognition has proliferated its everyday use. However, for research purposes, it is often unclear which model one should choose for a task, particularly if there is a requirement for speed as…

Computation and Language · Computer Science 2023-02-24 Ryan Whetten , Mir Tahsin Imtiaz , Casey Kennington

The goal of this work is to automatically determine whether and when a word of interest is spoken by a talking face, with or without the audio. We propose a zero-shot method suitable for in the wild videos. Our key contributions are: (1) a…

Computer Vision and Pattern Recognition · Computer Science 2020-09-07 Liliane Momeni , Triantafyllos Afouras , Themos Stafylakis , Samuel Albanie , Andrew Zisserman

Deep learning models are widely used for speaker recognition and spoofing speech detection. We propose the GMM-ResNet2 for synthesis speech detection. Compared with the previous GMM-ResNet model, GMM-ResNet2 has four improvements. Firstly,…

Sound · Computer Science 2024-07-03 Zhenchun Lei , Hui Yan , Changhong Liu , Yong Zhou , Minglei Ma

Scene text detection has witnessed rapid development in recent years. However, there still exists two main challenges: 1) many methods suffer from false positives in their text representations; 2) the large scale variance of scene texts…

Computer Vision and Pattern Recognition · Computer Science 2020-04-13 Yuxin Wang , Hongtao Xie , Zhengjun Zha , Mengting Xing , Zilong Fu , Yongdong Zhang

This paper presents a self-supervised method for visual detection of the active speaker in a multi-person spoken interaction scenario. Active speaker detection is a fundamental prerequisite for any artificial cognitive system attempting to…

Computer Vision and Pattern Recognition · Computer Science 2019-07-19 Kalin Stefanov , Jonas Beskow , Giampiero Salvi

Majority models of remote sensing image changing detection can only get great effect in a specific resolution data set. With the purpose of improving change detection effectiveness of the model in the multi-resolution data set, a weighted…

Computer Vision and Pattern Recognition · Computer Science 2022-05-04 Yu Jiang , Lei Hu , Yongmei Zhang , Xin Yang

This study proposes a novel lightweight neural network model leveraging features extracted from electrocardiogram (ECG) and respiratory signals for early OSA screening. ECG signals are used to generate feature spectrograms to predict sleep…

Machine Learning · Computer Science 2025-01-06 Hui Pan , Yanxuan Yu , Jilun Ye , Xu Zhang

Audio-only-based wake word spotting (WWS) is challenging under noisy conditions due to environmental interference in signal transmission. In this paper, we investigate on designing a compact audio-visual WWS system by utilizing visual…

Sound · Computer Science 2022-02-18 Hengshun Zhou , Jun Du , Chao-Han Huck Yang , Shifu Xiong , Chin-Hui Lee

Person re-identification has become a very popular research topic in the computer vision community owing to its numerous applications and growing importance in visual surveillance. Person re-identification remains challenging due to…

Computer Vision and Pattern Recognition · Computer Science 2023-11-30 Zongjing Cao , Hyo Jong Lee

The requirements for many applications of state-of-the-art speech recognition systems include not only low word error rate (WER) but also low latency. Specifically, for many use-cases, the system must be able to decode utterances in a…

For the diagnosis of Chinese medicine, tongue segmentation has reached a fairly mature point, but it has little application in the eye diagnosis of Chinese medicine.First, this time we propose Res-UNet based on the architecture of the U2Net…

Image and Video Processing · Electrical Eng. & Systems 2022-12-07 Peng Hong

We describe Howl, an open-source wake word detection toolkit with native support for open speech datasets, like Mozilla Common Voice and Google Speech Commands. We report benchmark results on Speech Commands and our own freely available…

Computation and Language · Computer Science 2020-08-24 Raphael Tang , Jaejun Lee , Afsaneh Razi , Julia Cambre , Ian Bicking , Jofish Kaye , Jimmy Lin

To combat adversarial spelling mistakes, we propose placing a word recognition model in front of the downstream classifier. Our word recognition models build upon the RNN semi-character architecture, introducing several new backoff…

Computation and Language · Computer Science 2019-08-30 Danish Pruthi , Bhuwan Dhingra , Zachary C. Lipton

In this paper, we propose a novel method called Rotational Region CNN (R2CNN) for detecting arbitrary-oriented texts in natural scene images. The framework is based on Faster R-CNN [1] architecture. First, we use the Region Proposal Network…

Computer Vision and Pattern Recognition · Computer Science 2017-07-03 Yingying Jiang , Xiangyu Zhu , Xiaobing Wang , Shuli Yang , Wei Li , Hua Wang , Pei Fu , Zhenbo Luo

The rhythm of bonafide speech is often difficult to replicate, which causes that the fundamental frequency (F0) of synthetic speech is significantly different from that of real speech. It is expected that the F0 feature contains the…

Sound · Computer Science 2024-07-09 Cunhang Fan , Jun Xue , Jianhua Tao , Jiangyan Yi , Chenglong Wang , Chengshi Zheng , Zhao Lv

Acoustic echo degrades the user experience in voice communication systems thus needs to be suppressed completely. We propose a real-time residual acoustic echo suppression (RAES) method using an efficient convolutional neural network. The…

Sound · Computer Science 2020-11-09 Xinquan Zhou , Yanhong Leng

We present a data-driven car occupancy detection algorithm using ultra-wideband radar based on the ResNet architecture. The algorithm is trained on a dataset of channel impulse responses obtained from measurements at three different…

Signal Processing · Electrical Eng. & Systems 2025-11-05 Jakob Möderl , Stefan Posch , Franz Pernkopf , Klaus Witrisal