English
Related papers

Related papers: Joint Scattering for Automatic Chick Call Recognit…

200 papers

In this paper, we present a new open source toolkit for automatic speech recognition (ASR), named CAT (CRF-based ASR Toolkit). A key feature of CAT is discriminative training in the framework of conditional random field (CRF), particularly…

Machine Learning · Computer Science 2019-11-21 Keyu An , Hongyu Xiang , Zhijian Ou

Connectionist Temporal Classification (CTC) is a widely used method for automatic speech recognition (ASR), renowned for its simplicity and computational efficiency. However, it often falls short in recognition performance. In this work, we…

Audio and Speech Processing · Electrical Eng. & Systems 2025-02-17 Zengwei Yao , Wei Kang , Xiaoyu Yang , Fangjun Kuang , Liyong Guo , Han Zhu , Zengrui Jin , Zhaoqing Li , Long Lin , Daniel Povey

This letter proposes a pilot-aided joint time synchronization and channel estimation (JTSCE) algorithm for orthogonal time frequency space (OTFS) systems. Unlike existing algorithms, JTSCE employs a maximum length sequence (MLS) rather than…

Signal Processing · Electrical Eng. & Systems 2024-08-14 Jiazheng Sun , Peng Yang , Xianbin Cao , Zehui Xiong , Haijun Zhang , Tony Q. S. Quek

This paper proposes a hierarchical spatial-temporal model for modelling the spectrograms of animal calls. The motivation stems from analyzing recordings of the so-called grunt calls emitted by various lemur species. Our goal is to identify…

Whisking is a rhythmic and adaptive behavior that rodents use to probe and interact with their environment, and the frequency of movement reflects both sensorimotor processing and internal brain states. A robust and traditional method of…

Quantitative Methods · Quantitative Biology 2026-05-28 Guanghui Li , Fangyuan Li , Barbara Lykke Lind , Rune W Berg

In this paper, we use several techniques with conventional vocal feature extraction (MFCC, STFT), along with deep-learning approaches such as CNN, and also context-level analysis, by providing the textual data, and combining different…

Audio and Speech Processing · Electrical Eng. & Systems 2019-05-22 Andrew Huang , Puwei Bao

We propose a new feature, namely, pitchsynchronous discrete cosine transform (PS-DCT), for the task of speaker identification. These features are obtained directly from the voiced segments of the speech signal, without any preemphasis or…

Audio and Speech Processing · Electrical Eng. & Systems 2018-12-07 Amit Meghanani , A G Ramakrishnan

Birds are vital parts of ecosystems across the world and are an excellent measure of the quality of life on earth. Many bird species are endangered while others are already extinct. Ecological efforts in understanding and monitoring bird…

Multimedia · Computer Science 2022-11-16 Chandra Kanth Nagesh , Abhishek Purushothama

The acoustic analysis of marmoset (Callithrix jacchus) vocalizations is often used to understand the evolutionary origins of human language. Currently, the analysis is largely carried out in a manual or semi-manual manner. Thus, there is a…

Audio and Speech Processing · Electrical Eng. & Systems 2025-04-22 Eklavya Sarkar , Kaja Wierucka , Alexandra B. Bosshard , Judith Burkart , Mathew Magimai. -Doss

We consider the problem of detecting, isolating and classifying elephant calls in continuously recorded audio. Such automatic call characterisation can assist conservation efforts and inform environmental management strategies. In contrast…

Sound · Computer Science 2025-04-03 Christiaan M. Geldenhuys , Thomas R. Niesler

This work considers merging two independent models, TTS and A2F, into a unified model to enable internal feature transfer, thereby improving the consistency between audio and facial expressions generated from text. We also discuss the…

Sound · Computer Science 2026-03-04 Qiangong Zhou , Nagasaka Tomohiro

We consider the problem of identifying people on the basis of their walk (gait) pattern. Classical approaches to tackle this problem are based on, e.g., video recordings or piezoelectric sensors embedded in the floor. In this work, we rely…

Sound · Computer Science 2020-01-27 Srđan Kitić , Gilles Puy , Patrick Pérez , Philippe Gilberton

Automatic species classification of birds from their sound is a computational tool of increasing importance in ecology, conservation monitoring and vocal communication studies. To make classification useful in practice, it is crucial to…

Sound · Computer Science 2014-07-14 Dan Stowell , Mark D. Plumbley

Sound event detection (SED) has significantly benefited from self-supervised learning (SSL) approaches, particularly masked audio transformer for SED (MAT-SED), which leverages masked block prediction to reconstruct missing audio segments.…

Audio and Speech Processing · Electrical Eng. & Systems 2025-03-03 Hyeonuk Nam , Yong-Hwa Park

Graph signal processing (GSP) facilitates the analysis of high-dimensional data on non-Euclidean domains by utilizing graph signals defined on graph vertices. In addition to static data, each vertex can provide continuous time-series…

Signal Processing · Electrical Eng. & Systems 2025-02-21 Tuna Alikaşifoğlu , Bünyamin Kartal , Eray Özgünay , Aykut Koç

Poultry farms are an important contributor to the human food chain. Worldwide, humankind keeps an enormous number of domesticated birds (e.g. chickens) for their eggs and their meat, providing rich sources of low-fat protein. However,…

Machine Learning · Computer Science 2018-11-09 Alireza Abdoli , Amy C. Murillo , Chin-Chia M. Yeh , Alec C. Gerry , Eamonn J. Keogh

Automatically identifying bat species from their echolocation calls is a difficult but important task for monitoring bats and the ecosystem they live in. Major challenges in automatic bat call identification are high call variability,…

Computer Vision and Pattern Recognition · Computer Science 2023-09-21 Frank Fundel , Daniel A. Braun , Sebastian Gottwald

Feature extraction plays an important role as a front-end processing block in speaker identification (SI) process. Most of the SI systems utilize like Mel-Frequency Cepstral Coefficients (MFCC), Perceptual Linear Prediction (PLP), Linear…

Sound · Computer Science 2015-03-19 Md. Sahidullah , Sandipan Chakroborty , Goutam Saha

Speech is a natural form of communication for human beings, and computers with the ability to understand speech and speak with a human voice are expected to contribute to the development of more natural man-machine interfaces. Computers…

Sound · Computer Science 2013-05-15 Neema Mishra , Urmila Shrawankar , V M Thakare

In this research endeavor, it was hypothesized that the sound produced by animals during their vocalizations can be used as identifiers of the animal breed or species even if they sound the same to unaided human ear. To test this…