English
Related papers

Related papers: MOVER: Combining Multiple Meeting Recognition Syst…

200 papers

The CHiME-7 and 8 distant speech recognition (DASR) challenges focus on multi-channel, generalizable, joint automatic speech recognition (ASR) and diarization of conversational speech. With participation from 9 teams submitting 32 diverse…

Composed image retrieval which combines a reference image and a text modifier to identify the desired target image is a challenging task, and requires the model to comprehend both vision and language modalities and their interactions.…

Computer Vision and Pattern Recognition · Computer Science 2023-10-03 Shu Zhao , Huijuan Xu

This paper describes a system that generates speaker-annotated transcripts of meetings by using a microphone array and a 360-degree camera. The hallmark of the system is its ability to handle overlapped speech, which has been an unsolved…

This technical report details our submission system to the CHiME-7 DASR Challenge, which focuses on speaker diarization and speech recognition under complex multi-speaker scenarios. Additionally, it also evaluates the efficiency of systems…

Modern multi-object tracking (MOT) systems usually model the trajectories by associating per-frame detections. However, when camera motion, fast motion, and occlusion challenges occur, it is difficult to ensure long-range tracking or even…

Computer Vision and Pattern Recognition · Computer Science 2020-09-21 Shoudong Han , Piao Huang , Hongwei Wang , En Yu , Donghaisheng Liu , Xiaofeng Pan , Jun Zhao

This paper discribes the DKU-DukeECE submission to the 4th track of the VoxCeleb Speaker Recognition Challenge 2022 (VoxSRC-22). Our system contains a fused voice activity detection model, a clustering-based diarization model, and a…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-05 Weiqing Wang , Xiaoyi Qin , Ming Cheng , Yucong Zhang , Kangyue Wang , Ming Li

Decoder-only discrete-token language models have recently achieved significant success in automatic speech recognition. However, systematic analyses of how different modalities impact performance in specific scenarios remain limited. In…

Computer Vision and Pattern Recognition · Computer Science 2025-11-03 Yiwen Guan , Viet Anh Trinh , Vivek Voleti , Jacob Whitehill

Developers need to perform adequate testing to ensure the quality of Automatic Speech Recognition (ASR) systems. However, manually collecting required test cases is tedious and time-consuming. Our recent work proposes CrossASR, a…

Software Engineering · Computer Science 2022-01-06 Muhammad Hilmi Asyrofi , Zhou Yang , David Lo

Multimodal emotion recognition (MER) aims to infer human affect by jointly modeling audio and visual cues; however, existing approaches often struggle with temporal misalignment, weakly discriminative feature representations, and suboptimal…

Multimedia · Computer Science 2026-01-21 Joe Dhanith P R , Shravan Venkatraman , Vigya Sharma , Santhosh Malarvannan

Driver action recognition, aiming to accurately identify drivers' behaviours, is crucial for enhancing driver-vehicle interactions and ensuring driving safety. Unlike general action recognition, drivers' environments are often challenging,…

Computer Vision and Pattern Recognition · Computer Science 2024-08-20 Ruoyu Wang , Wenqian Wang , Jianjun Gao , Dan Lin , Kim-Hui Yap , Bingbing Li

This paper describes the ByteDance speaker diarization system for the fourth track of the VoxCeleb Speaker Recognition Challenge 2021 (VoxSRC-21). The VoxSRC-21 provides both the dev set and test set of VoxConverse for use in validation and…

Sound · Computer Science 2021-09-07 Keke Wang , Xudong Mao , Hao Wu , Chen Ding , Chuxiang Shang , Rui Xia , Yuxuan Wang

We present AutoMode-ASR, a novel framework that effectively integrates multiple ASR systems to enhance the overall transcription quality while optimizing cost. The idea is to train a decision model to select the optimal ASR system for each…

Computation and Language · Computer Science 2024-09-20 Ahmet Gündüz , Yunsu Kim , Kamer Ali Yuksel , Mohamed Al-Badrashiny , Thiago Castro Ferreira , Hassan Sawaf

In real-world scenarios, audio and video signals are often subject to environmental noise and limited acquisition conditions, resulting in extracted features containing excessive noise. Furthermore, there is an imbalance in data quality and…

Computation and Language · Computer Science 2026-03-30 Ying Liu , Yuntao Shou , Wei Ai , Tao Meng , Keqin Li

Alongside acoustic information, linguistic features based on speech transcripts have been proven useful in Speech Emotion Recognition (SER). However, due to the scarcity of emotion labelled data and the difficulty of recognizing emotional…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-11 Yuanchao Li , Peter Bell , Catherine Lai

Multi-modal conversation emotion recognition (MCER) aims to recognize and track the speaker's emotional state using text, speech, and visual information in the conversation scene. Analyzing and studying MCER issues is significant to…

Artificial Intelligence · Computer Science 2025-11-14 Yuntao Shou , Tao Meng , Wei Ai , Fangze Fu , Nan Yin , Keqin Li

Robust audio-visual speech recognition (AVSR) in noisy environments remains challenging, as existing systems struggle to estimate audio reliability and dynamically adjust modality reliance. We propose router-gated cross-modal feature…

Computer Vision and Pattern Recognition · Computer Science 2025-08-27 DongHoon Lim , YoungChae Kim , Dong-Hyun Kim , Da-Hee Yang , Joon-Hyuk Chang

Object detection and classification using aerial images is a challenging task as the information regarding targets are not abundant. Synthetic Aperture Radar(SAR) images can be used for Automatic Target Recognition(ATR) systems as it can…

Computer Vision and Pattern Recognition · Computer Science 2022-12-15 Sumanth Udupa , Aniruddh Sikdar , Suresh Sundaram

Most text retrievers generate \emph{one} query vector to retrieve relevant documents. Yet, the conditional distribution of relevant documents for the query may be multimodal, e.g., representing different interpretations of the query. We…

Computation and Language · Computer Science 2025-11-05 Hung-Ting Chen , Xiang Liu , Shauli Ravfogel , Eunsol Choi

In spite of the popularity of end-to-end diarization systems nowadays, modular systems comprised of voice activity detection (VAD), speaker embedding extraction plus clustering, and overlapped speech detection (OSD) plus handling still…

Audio and Speech Processing · Electrical Eng. & Systems 2024-11-05 Petr Pálka , Federico Landini , Dominik Klement , Mireia Diez , Anna Silnova , Marc Delcroix , Lukáš Burget

Person identification systems often rely on audio, visual, or behavioral cues, but real-world conditions frequently present with missing or degraded modalities. To address this challenge, we propose a multimodal person identification…

Computer Vision and Pattern Recognition · Computer Science 2026-01-28 Aref Farhadipour , Teodora Vukovic , Volker Dellwo , Petr Motlicek , Srikanth Madikeri