English
Related papers

Related papers: Robust Phonetic Segmentation Using Spectral Transi…

200 papers

The recently-proposed mixture invariant training (MixIT) is an unsupervised method for training single-channel sound separation models in the sense that it does not require ground-truth isolated reference sources. In this paper, we…

Sound · Computer Science 2021-10-22 Aswin Sivaraman , Scott Wisdom , Hakan Erdogan , John R. Hershey

Discourse segmentation, which segments texts into Elementary Discourse Units, is a fundamental step in discourse analysis. Previous discourse segmenters rely on complicated hand-crafted features and are not practical in actual use. In this…

Computation and Language · Computer Science 2018-08-29 Yizhong Wang , Sujian Li , Jingfeng Yang

In this paper, we present a novel approach for text independent phone-to-audio alignment based on phoneme recognition, representation learning and knowledge transfer. Our method leverages a self-supervised model (wav2vec2) fine-tuned for…

Audio and Speech Processing · Electrical Eng. & Systems 2024-05-06 Noé Tits , Prernna Bhatnagar , Thierry Dutoit

This paper presents a new approach for unsupervised Spoken Term Detection with spoken queries using multiple sets of acoustic patterns automatically discovered from the target corpus. The different pattern HMM configurations(number of…

Computation and Language · Computer Science 2015-09-09 Cheng-Tao Chung , Chun-an Chan , Lin-shan Lee

When dedicated positioning systems, such as GPS, are unavailable, a mobile device has no choice but to fall back on its cellular network for localization. Due to random variations in the channel conditions to its surrounding base stations…

Information Theory · Computer Science 2015-10-27 Javier Schloemann , Harpreet S. Dhillon , R. Michael Buehrer

Generalized spatial modulation (GSM) is a spectral-efficient technique used in multiple-input multiple-output (MIMO) wireless communications when the number of radio frequency chains at the transmitter is less than the number of transmit…

Information Theory · Computer Science 2020-08-11 Lakshit Singla , Lakshmi Natarajan

Our research discovers how the rolling shutter and movable lens structures widely found in smartphone cameras modulate structure-borne sounds onto camera images, creating a point-of-view (POV) optical-acoustic side channel for acoustic…

Cryptography and Security · Computer Science 2023-09-13 Yan Long , Pirouz Naghavi , Blas Kojusner , Kevin Butler , Sara Rampazzi , Kevin Fu

This thesis work presents an architectural design of a system to bring non-repudiation concept into the IP based digital voice conversations (VoIP) in LTE and UMTS networks, using electronic signatures, by considering a centralized…

Cryptography and Security · Computer Science 2020-12-08 Umut Can Cabuk

This paper proposes an efficient reconfigurable hardware design for speech enhancement based on multi band spectral subtraction algorithm and involving both magnitude and phase components. Our proposed design is novel as it estimates…

Sound · Computer Science 2015-08-26 Tanmay Biswas , Sudhindu Bikash Mandal , Debasree Saha , Amlan Chakrabarti

Real-world measurements often comprise a dominant signal contaminated by a noisy background. Robustly estimating the dominant signal in practice has been a fundamental statistical problem. Classically, mixture models have been used to…

Computation · Statistics 2026-05-20 Ananyabrata Barua , Ayanendranath Basu

Tongue imaging serves as a valuable diagnostic tool, particularly in Traditional Chinese Medicine (TCM). The quality of tongue surface segmentation significantly affects the accuracy of tongue image classification and subsequent diagnosis…

Image and Video Processing · Electrical Eng. & Systems 2025-08-22 Jiacheng Xie , Ziyang Zhang , Biplab Poudel , Congyu Guo , Yang Yu , Guanghui An , Xiaoting Tang , Lening Zhao , Chunhui Xu , Dong Xu

Despite rapid progress in scene segmentation in recent years, 3D segmentation methods are still limited when there is severe occlusion. The key challenge is estimating the segment boundaries of (partially) occluded objects, which are…

Robotics · Computer Science 2021-04-02 Andrew Price , Kun Huang , Dmitry Berenson

We present a novel underwater system that can perform acoustic ranging between commodity smartphones. To achieve this, we design a real-time underwater ranging protocol that computes the time-of-flight between smartphones. To address the…

Networking and Internet Architecture · Computer Science 2022-09-07 Tuochao Chen , Justin Chan , Shyamnath Gollakota

Functional Magnetic Resonance Imaging is a noninvasive tool for studying cerebral function. Many factors challenge activation detection, especially in low-signal scenarios that arise in the performance of high-level cognitive tasks. We…

Methodology · Statistics 2019-05-07 Israel Almodóvar-Rivera , Ranjan Maitra

Fine-tuning over large pretrained language models (PLMs) has established many state-of-the-art results. Despite its superior performance, such fine-tuning can be unstable, resulting in significant variance in performance and potential risks…

Computation and Language · Computer Science 2022-10-20 Chenghao Yang , Xuezhe Ma

We propose TF-GridNet for speech separation. The model is a novel deep neural network (DNN) integrating full- and sub-band modeling in the time-frequency (T-F) domain. It stacks several blocks, each consisting of an intra-frame full-band…

With the proliferation of video platforms on the internet, recording musical performances by mobile devices has become commonplace. However, these recordings often suffer from degradation such as noise and reverberation, which negatively…

Sound · Computer Science 2023-08-25 Yunkee Chae , Junghyun Koo , Sungho Lee , Kyogu Lee

Second language (L2) speech is often labeled with the native, phone categories. However, in many cases, it is difficult to decide on a categorical phone that an L2 segment belongs to. These segments are regarded as non-categories. Most…

Computation and Language · Computer Science 2020-02-04 Xu Li , Xixin Wu , Xunying Liu , Helen Meng

To join the advantages of classical and end-to-end approaches for speech recognition, we present a simple, novel and competitive approach for phoneme-based neural transducer modeling. Different alignment label topologies are compared and…

Computation and Language · Computer Science 2021-04-21 Wei Zhou , Simon Berger , Ralf Schlüter , Hermann Ney

This paper studies a multiple-input multiple-output (MIMO) orthogonal frequency division multiplexing (OFDM) networked integrated sensing and communication (ISAC) system, in which multiple base stations (BSs) perform beam tracking to…

Networking and Internet Architecture · Computer Science 2025-08-19 Xiaoyu Yang , Zhiqing Wei , Jie Xu , Huici Wu , Zhiyong Feng