English
Related papers

Related papers: Attention Is Not Always the Answer: Optimizing Voi…

200 papers

In this paper, a new speech feature fusion method is proposed for speaker recognition on the basis of the cross gate parallel convolutional neural network (CG-PCNN). The Mel filter bank features (MFBFs) of different frequency resolutions…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-28 Jiacheng Zhang , Wenyi Yan , Ye Zhang

Recent advances in synthetic speech have made audio deepfakes increasingly realistic, posing significant security risks. Existing detection methods that rely on a single modality, either raw waveform embeddings or spectral based features,…

Audio-visual information fusion enables a performance improvement in speech recognition performed in complex acoustic scenarios, e.g., noisy environments. It is required to explore an effective audio-visual fusion strategy for audiovisual…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-07 Liangfa Wei , Jie Zhang , Junfeng Hou , Lirong Dai

Automatic fluency assessment (AFA) remains challenging, particularly in capturing speech rhythm, pauses, and disfluencies in non-native speakers. We introduce a chunk-based approach integrating self-supervised learning (SSL) models…

Computation and Language · Computer Science 2025-06-27 Papa Séga Wade , Mihai Andries , Ioannis Kanellos , Thierry Moudenc

The progression of deep learning and the widespread adoption of sensors have facilitated automatic multi-view fusion (MVF) about the cardiovascular system (CVS) signals. However, prevalent MVF model architecture often amalgamates CVS…

Machine Learning · Computer Science 2024-06-14 Qihan Hu , Daomiao Wang , Hong Wu , Jian Liu , Cuiwei Yang

Automatic emotion recognition (ER) has recently gained lot of interest due to its potential in many real-world applications. In this context, multimodal approaches have been shown to improve performance (over unimodal approaches) by…

Computer Vision and Pattern Recognition · Computer Science 2022-09-20 R Gnana Praveen , Eric Granger , Patrick Cardinal

Visual acoustic matching (VAM) is pivotal for enhancing the immersive experience, and the task of dereverberation is effective in improving audio intelligibility. Existing methods treat each task independently, overlooking the inherent…

Sound · Computer Science 2024-07-16 Jian Ma , Wenguan Wang , Yi Yang , Feng Zheng

In a speech recognition system, voice activity detection (VAD) is a crucial frontend module. Addressing the issues of poor noise robustness in traditional binary VAD systems based on DFSMN, the paper further proposes semantic VAD based on…

Sound · Computer Science 2023-12-25 Lingyun Zuo , Keyu An , Shiliang Zhang , Zhijie Yan

Vision-language models have been key to the development of open-vocabulary 2D semantic segmentation. Lifting these models from 2D images to 3D scenes, however, remains a challenging problem. Existing approaches typically back-project and…

Computer Vision and Pattern Recognition · Computer Science 2026-04-15 Tomas Berriel Martins , Martin R. Oswald , Javier Civera

Weakly supervised violence detection refers to the technique of training models to identify violent segments in videos using only video-level labels. Among these approaches, multimodal violence detection, which integrates modalities such as…

Computer Vision and Pattern Recognition · Computer Science 2025-03-17 Wenping Jin , Li Zhu , Jing Sun

With the rapid development of radar jamming systems, especially digital radio frequency memory (DRFM), the electromagnetic environment has become increasingly complex. In recent years, most existing studies have focused solely on either…

Signal Processing · Electrical Eng. & Systems 2025-06-10 Huake Wang , Xudong Han , Bairui Cai , Guisheng Liao , Yinghui Quan

Cognitive Behavioral Therapy (CBT) is a goal-oriented psychotherapy for mental health concerns implemented in a conversational setting with broad empirical support for its effectiveness across a range of presenting problems and client…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-16 Zhuohao Chen , Nikolaos Flemotomos , Victor Ardulov , Torrey A. Creed , Zac E. Imel , David C. Atkins , Shrikanth Narayanan

By implicitly recognizing a user based on his/her speech input, speaker identification enables many downstream applications, such as personalized system behavior and expedited shopping checkouts. Based on whether the speech content is…

Machine Learning · Computer Science 2021-06-21 Ruirui Li , Chelsea J. -T. Ju , Zeya Chen , Hongda Mao , Oguz Elibol , Andreas Stolcke

This paper integrates a voice activity detection (VAD) function with end-to-end automatic speech recognition toward an online speech interface and transcribing very long audio recordings. We focus on connectionist temporal classification…

Audio and Speech Processing · Electrical Eng. & Systems 2020-03-16 Takenori Yoshimura , Tomoki Hayashi , Kazuya Takeda , Shinji Watanabe

Multi-sensor fusion is crucial for accurate 3D object detection in autonomous driving, with cameras and LiDAR being the most commonly used sensors. However, existing methods perform sensor fusion in a single view by projecting features from…

Computer Vision and Pattern Recognition · Computer Science 2024-12-11 Rohit Mohan , Daniele Cattaneo , Florian Drews , Abhinav Valada

A significant amount of redundancy exists between consecutive frames of a video. Object detectors typically produce detections for one image at a time, without any capabilities for taking advantage of this redundancy. Meanwhile, many…

Computer Vision and Pattern Recognition · Computer Science 2021-09-16 Hughes Perreault , Guillaume-Alexandre Bilodeau , Nicolas Saunier , Maguelonne Héritier

Humans possess a remarkable ability to integrate auditory and visual information, enabling a deeper understanding of the surrounding environment. This early fusion of audio and visual cues, demonstrated through cognitive psychology and…

Computer Vision and Pattern Recognition · Computer Science 2023-12-05 Shentong Mo , Pedro Morgado

Spoken languages often utilise intonation, rhythm, intensity, and structure, to communicate intention, which can be interpreted differently depending on the rhythm of speech of their utterance. These speech acts provide the foundation of…

This paper introduces and motivates the use of hybrid robust feature extraction technique for spoken language identification (LID) system. The speech recognizers use a parametric form of a signal to get the most important distinguishable…

Sound · Computer Science 2010-03-31 Pawan Kumar , Astik Biswas , A . N. Mishra , Mahesh Chandra

Medical Visual Question Answering (MedVQA) has attracted growing interest at the intersection of medical image understanding and natural language processing for clinical applications. By interpreting medical images and providing precise…

Image and Video Processing · Electrical Eng. & Systems 2025-05-13 Zhilin Zhang , Jie Wang , Zhanghao Qin , Ruiqi Zhu , Xiaoliang Gong
‹ Prev 1 3 4 5 6 7 10 Next ›