中文
相关论文

相关论文: MOVER: Combining Multiple Meeting Recognition Syst…

200 篇论文

Facial expression recognition is an essential task for various applications, including emotion detection, mental health analysis, and human-machine interactions. In this paper, we propose a multi-modal facial expression recognition method…

计算机视觉与模式识别 · 计算机科学 2023-03-21 Jun-Hwa Kim , Namho Kim , Chee Sun Won

Mixed-precision training is a crucial technique for scaling deep learning models, but successful mixedprecision training requires identifying and applying the right combination of training methods. This paper presents our preliminary study…

机器学习 · 计算机科学 2025-12-30 Bor-Yiing Su , Peter Dykas , Mike Chrzanowski , Jatin Chhugani

Speaker diarization, the process of segmenting an audio stream or transcribed speech content into homogenous partitions based on speaker identity, plays a crucial role in the interpretation and analysis of human speech. Most existing…

机器学习 · 计算机科学 2024-08-23 Luyao Cheng , Hui Wang , Siqi Zheng , Yafeng Chen , Rongjie Huang , Qinglin Zhang , Qian Chen , Xihao Li

Multimodal Large Language Models (MLLMs) have increasingly localized and interleaved visual evidence for deliberative reasoning. Grounding-based approaches typically focus on regions of interest (RoIs) by injecting cropped image patches or…

计算机视觉与模式识别 · 计算机科学 2026-05-29 Guannan Lv , Ren Nie , Hongjian Dou , Tingting Gao

Audio-visual speech recognition (AVSR) system is thought to be one of the most promising solutions for robust speech recognition, especially in noisy environment. In this paper, we propose a novel multimodal attention based method for…

计算与语言 · 计算机科学 2019-04-24 Pan Zhou , Wenwen Yang , Wei Chen , Yanfeng Wang , Jia Jia

In this paper, we present methods in deep multimodal learning for fusing speech and visual modalities for Audio-Visual Automatic Speech Recognition (AV-ASR). First, we study an approach where uni-modal deep networks are trained separately…

计算与语言 · 计算机科学 2015-01-23 Youssef Mroueh , Etienne Marcheret , Vaibhava Goel

Second-pass rescoring is employed in most state-of-the-art speech recognition systems. Recently, BERT based models have gained popularity for re-ranking the n-best hypothesis by exploiting the knowledge from masked language model…

音频与语音处理 · 电气工程与系统科学 2023-06-19 Prashanth Gurunath Shivakumar , Jari Kolehmainen , Yile Gu , Ankur Gandhe , Ariya Rastrow , Ivan Bulyko

This paper presents a unified multi-speaker encoder (UME), a novel architecture that jointly learns representations for speaker diarization (SD), speech separation (SS), and multi-speaker automatic speech recognition (ASR) tasks using a…

音频与语音处理 · 电气工程与系统科学 2026-05-14 Muhammad Shakeel , Yui Sudo , Yifan Peng , Chyi-Jiunn Lin , Shinji Watanabe

Fusing outputs from automatic speaker verification (ASV) and spoofing countermeasure (CM) is expected to make an integrated system robust to zero-effort imposters and synthesized spoofing attacks. Many score-level fusion methods have been…

音频与语音处理 · 电气工程与系统科学 2024-09-25 Xin Wang , Tomi Kinnunen , Kong Aik Lee , Paul-Gauthier Noé , Junichi Yamagishi

This report describes the submission system of the GIST-AiTeR team at the 2022 VoxCeleb Speaker Recognition Challenge (VoxSRC) Track 4. Our system mainly includes speech enhancement, voice activity detection , multi-scaled speaker…

音频与语音处理 · 电气工程与系统科学 2022-10-07 Dongkeon Park , Yechan Yu , Kyeong Wan Park , Ji Won Kim , Hong Kook Kim

Multi-talker overlapped speech poses a significant challenge for speech recognition and diarization. Recent research indicated that these two tasks are inter-dependent and complementary, motivating us to explore a unified modeling method to…

声音 · 计算机科学 2023-05-26 Lingwei Meng , Jiawen Kang , Mingyu Cui , Haibin Wu , Xixin Wu , Helen Meng

Classical autonomous driving systems connect perception and prediction modules via hand-crafted bounding-box interfaces, limiting information flow and propagating errors to downstream tasks. Recent research aims to develop end-to-end models…

计算机视觉与模式识别 · 计算机科学 2026-02-16 Mohammed Amine Bencheikh Lehocine , Julian Schmidt , Frank Moosmann , Dikshant Gupta , Fabian Flohr

Despite improvements in automatic speaker verification (ASV), vulnerability against spoofing attacks remains a major concern. In this study, we investigate the integration of ASV and countermeasure (CM) subsystems into a modular spoof-aware…

音频与语音处理 · 电气工程与系统科学 2025-09-17 Oguzhan Kurnaz , Tomi Kinnunen , Cemal Hanilci

As human-robot collaboration advances, natural and flexible communication methods are essential for effective robot control. Traditional methods relying on a single modality or rigid rules struggle with noisy or misaligned data as well as…

机器人学 · 计算机科学 2025-04-03 Petr Vanc , Karla Stepanova

Recent advancements in multi-modal large language models have propelled the development of joint probabilistic models capable of both image understanding and generation. However, we have identified that recent methods suffer from loss of…

计算机视觉与模式识别 · 计算机科学 2025-06-05 Jian Yang , Dacheng Yin , Yizhou Zhou , Fengyun Rao , Wei Zhai , Yang Cao , Zheng-Jun Zha

We present a multiway fusion algorithm capable of directly processing uncertain pairwise affinities. In contrast to existing works that require initial pairwise associations, our MIXER algorithm improves accuracy by leveraging the…

计算机视觉与模式识别 · 计算机科学 2023-05-04 Parker C. Lusk , Kaveh Fathian , Jonathan P. How

This paper summarizes the JHU team's efforts in tracks 1 and 2 of the CHiME-6 challenge for distant multi-microphone conversational speech diarization and recognition in everyday home environments. We explore multi-array processing…

Speaker diarization systems are challenged by a trade-off between the temporal resolution and the fidelity of the speaker representation. By obtaining a superior temporal resolution with an enhanced accuracy, a multi-scale approach is a way…

音频与语音处理 · 电气工程与系统科学 2022-03-31 Tae Jin Park , Nithin Rao Koluguri , Jagadeesh Balam , Boris Ginsburg

Low-resource accented speech recognition is one of the important challenges faced by current ASR technology in practical applications. In this study, we propose a Conformer-based architecture, called Aformer, to leverage both the acoustic…

声音 · 计算机科学 2023-06-21 Xuefei Wang , Yanhua Long , Yijie Li , Haoran Wei

Visual recognition inside the vehicle cabin leads to safer driving and more intuitive human-vehicle interaction but such systems face substantial obstacles as they need to capture different granularities of driver behaviour while dealing…

计算机视觉与模式识别 · 计算机科学 2022-04-12 Alina Roitberg , Kunyu Peng , Zdravko Marinov , Constantin Seibold , David Schneider , Rainer Stiefelhagen