English
Related papers

Related papers: Call-sign recognition and understanding for noisy …

200 papers

This paper aims to provide an approach for automatic coding of physician-patient communication transcripts to improve patient-centered communication (PCC). PCC is a central part of high-quality health care. To improve PCC, dialogues between…

Computation and Language · Computer Science 2021-09-23 Gilchan Park , Julia Taylor Rayz , Cleveland G. Shields

Vehicles become more vulnerable to remote attackers in modern days due to their increasing connectivity and range of functionality. Such increased attack vectors enable adversaries to access a vehicle Electronic Control Unit (ECU). As of…

Cryptography and Security · Computer Science 2019-11-25 Marcel Kneib , Oleg Schell , Christopher Huth

Traffic Anomaly Understanding (TAU) is important for traffic safety in Intelligent Transportation Systems. Recent vision-language models (VLMs) have shown strong capabilities in video understanding. However, progress on TAU remains limited…

Computer Vision and Pattern Recognition · Computer Science 2026-03-20 Yuqiang Lin , Kehua Chen , Sam Lockyer , Arjun Yadav , Mingxuan Sui , Shucheng Zhang , Yan Shi , Bingzhang Wang , Yuang Zhang , Markus Zarbock , Florain Stanek , Adrian Evans , Wenbin Li , Yinhai Wang , Nic Zhang

Distributed Acoustic Sensing (DAS) is promising for traffic monitoring, but its extensive data and sensitivity to vibrations, causing noise, pose computational challenges. To address this, we propose a two-step deep-learning workflow with…

Geophysics · Physics 2024-03-06 Dongzi Xie , Xinming Wu , Zhixiang Guo , Heting Hong , Baoshan Wang , Yingjiao Rong

In this paper, we investigate the robustness of traffic sign recognition algorithms under challenging conditions. Existing datasets are limited in terms of their size and challenging condition coverage, which motivated us to generate the…

Computer Vision and Pattern Recognition · Computer Science 2018-11-14 Dogancan Temel , Gukyeong Kwon , Mohit Prabhushankar , Ghassan AlRegib

This paper explores a structured application of the One-Class approach and the One-Class-One-Network model for supervised classification tasks, focusing on vowel phonemes classification and speakers recognition for the Automatic Speech…

Sound · Computer Science 2025-07-01 Stefano Giacomelli , Marco Giordano , Claudia Rinaldi

Effectively managing Air Traffic Control Officer (ATCO) workload is crucial in maintaining operational safety. Group supervisors use tools that estimate upcoming traffic load to aid decision-making. However, industry-standard models can…

Machine Learning · Computer Science 2026-05-25 Edward Henderson , George De Ath , Nick Pepper

Automated Compliance Checking (ACC) systems aim to semantically parse building regulations to a set of rules. However, semantic parsing is known to be hard and requires large amounts of training data. The complexity of creating such…

Computation and Language · Computer Science 2021-10-05 Ruben Kruiper , Ioannis Konstas , Alasdair Gray , Farhad Sadeghineko , Richard Watson , Bimal Kumar

Automatic Speech Recognition (ASR) technology has made significant progress in recent years, providing accurate transcription across various domains. However, some challenges remain, especially in noisy environments and specialized jargon.…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-06 Aviv Shamsian , Aviv Navon , Neta Glazer , Gill Hetz , Joseph Keshet

Spoken language understanding (SLU) systems often exhibit suboptimal performance in processing atypical speech, typically caused by neurological conditions and motor impairments. Recent advancements in Text-to-Speech (TTS) synthesis-based…

Interactions based on automatic speech recognition (ASR) have become widely used, with speech input being increasingly utilized to create documents. However, as there is no easy way to distinguish between commands being issued and text…

Human-Computer Interaction · Computer Science 2022-08-24 Jun Rekimoto

In most modern object detection pipelines, the detection proposals are processed independently given the feature map. Therefore, they overlook the underlying relationships between objects and the surrounding background, which could have…

Computer Vision and Pattern Recognition · Computer Science 2024-12-03 Botao Ren , Botian Xu , Xue Yang , Yifan Pu , Jingyi Wang , Zhidong Deng

Automatic Speech Recognition systems have made significant progress with large-scale pre-trained models. However, most current systems focus solely on transcribing the speech without identifying speaker roles, a function that is critical…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-13 Anfeng Xu , Tiantian Feng , Shrikanth Narayanan

We propose a novel cascaded cross-modal transformer (CCMT) that combines speech and text transcripts to detect customer requests and complaints in phone conversations. Our approach leverages a multimodal paradigm by transcribing the speech…

Computation and Language · Computer Science 2023-07-31 Nicolae-Catalin Ristea , Radu Tudor Ionescu

Spoken Language Understanding (SLU) aims to extract structured semantic representations (e.g., slot-value pairs) from speech recognized texts, which suffers from errors of Automatic Speech Recognition (ASR). To alleviate the problem caused…

Computation and Language · Computer Science 2020-09-08 Chen Liu , Su Zhu , Lu Chen , Kai Yu

Compensation for channel mismatch and noise interference is essential for robust automatic speech recognition. Enhanced speech has been introduced into the multi-condition training of acoustic models to improve their generalization ability.…

Sound · Computer Science 2022-11-24 Hung-Shin Lee , Pin-Yuan Chen , Yao-Fei Cheng , Yu Tsao , Hsin-Min Wang

This report proposes state-of-the-art research in the field of Computer Assisted Language Learning (CALL). Mispronunciation detection is one of the core components of Computer Assisted Pronunciation Training (CAPT) systems which is a subset…

Sound · Computer Science 2022-01-26 Neha Baranwal , Sharatkumar Chilaka

In this paper, we propose a novel speech emotion recognition model called Cross Attention Network (CAN) that uses aligned audio and text signals as inputs. It is inspired by the fact that humans recognize speech as a combination of…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-27 Yoonhyung Lee , Seunghyun Yoon , Kyomin Jung

Performance of spoken language understanding (SLU) can be degraded with automatic speech recognition (ASR) errors. We propose a novel approach to improve SLU robustness by randomly corrupting clean training text with an ASR error simulator,…

Computation and Language · Computer Science 2022-11-09 Yik-Cheung Tam , Jiacheng Xu , Jiakai Zou , Zecheng Wang , Tinglong Liao , Shuhan Yuan

In this work, we propose a streaming AV-ASR system based on a hybrid connectionist temporal classification (CTC)/attention neural network architecture. The audio and the visual encoder neural networks are both based on the conformer…

Audio and Speech Processing · Electrical Eng. & Systems 2023-07-04 Pingchuan Ma , Niko Moritz , Stavros Petridis , Christian Fuegen , Maja Pantic