English
Related papers

Related papers: HumanOmni-Speaker: Identifying Who said What and W…

200 papers

The Speaker Diarization and Recognition (SDR) task aims to predict "who spoke when and what" within an audio clip, which is a crucial task in various real-world multi-speaker scenarios such as meeting transcription and dialogue systems.…

Sound · Computer Science 2026-01-06 Han Yin , Yafeng Chen , Chong Deng , Luyao Cheng , Hui Wang , Chao-Hong Tan , Qian Chen , Wen Wang , Xiangang Li

Omni-modal large language models (OLMs) redefine human-machine interaction by natively integrating audio, vision, and text. However, existing OLM benchmarks remain anchored to static, accuracy-centric tasks, leaving a critical gap in…

Artificial Intelligence · Computer Science 2026-03-18 Tianyu Xie , Jinfa Huang , Yuexiao Ma , Rongfang Luo , Yan Yang , Wang Chen , Yuhui Zeng , Ruize Fang , Yixuan Zou , Xiawu Zheng , Jiebo Luo , Rongrong Ji

In human-centric scenes, the ability to simultaneously understand visual and auditory information is crucial. While recent omni models can process multiple modalities, they generally lack effectiveness in human-centric scenes due to the…

Computer Vision and Pattern Recognition · Computer Science 2025-01-28 Jiaxing Zhao , Qize Yang , Yixing Peng , Detao Bai , Shimin Yao , Boyuan Sun , Xiang Chen , Shenghao Fu , Weixuan chen , Xihan Wei , Liefeng Bo

Multi-speaker automatic speech recognition (MASR) aims to predict ''who spoke when and what'' from multi-speaker speech, a key technology for multi-party dialogue understanding. However, most existing approaches decouple temporal modeling…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-15 Yifan Hu , Peiji Yang , Zhisheng Wang , Yicheng Zhong , Rui Liu

Omni-modal large language models (OLLMs) offer a promising end-to-end solution for slide-enhanced speech recognition due to their inherent multimodal capabilities. However, we found a fundamental issue faced by OLLMs: \textit{Visual…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-28 Rui Hu , Delai Qiu , Yining Wang , Shengping Liu , Jitao Sang

Recent advances in omni-modal large language models have enabled remarkable progress in joint vision-audio understanding. However, prevailing architectures rely on modality-specific encoders with a \emph{video-coarse, audio-dense} design --…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Detao Bai , Shimin Yao , Weixuan Chen , Chengen Lai , Yuanming Li , Zhiheng Ma , Xihan Wei

Spoken dialogue is a primary source of information in videos; therefore, accurately identifying who spoke what and when is essential for deep video understanding. We introduce D-ORCA, a \textbf{d}ialogue-centric \textbf{o}mni-modal large…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Changli Tang , Tianyi Wang , Fengyun Rao , Jing Lyu , Chao Zhang

Although significant progress has been made in audio-driven talking head generation, text-driven methods remain underexplored. In this work, we present OmniTalker, a unified framework that jointly generates synchronized talking audio-video…

Computer Vision and Pattern Recognition · Computer Science 2025-06-04 Zhongjian Wang , Peng Zhang , Jinwei Qi , Guangyuan Wang , Chaonan Ji , Sheng Xu , Bang Zhang , Liefeng Bo

Multi-speaker automatic speech recognition (ASR) aims to transcribe conversational speech involving multiple speakers, requiring the model to capture not only what was said, but also who said it and sometimes when it was spoken. Recent…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-27 Li Li , Ming Cheng , Weixin Zhu , Yannan Wang , Juan Liu , Ming Li

Multimodal large language models (MLLMs) are expected to jointly interpret vision, audio, and language, yet existing video benchmarks rarely assess fine-grained reasoning about human speech. Many tasks remain visually solvable or only…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Le Thien Phuc Nguyen , Zhuoran Yu , Samuel Low Yu Hang , Subin An , Jeongik Lee , Yohan Ban , SeungEun Chung , Thanh-Huy Nguyen , JuWan Maeng , Soochahn Lee , Yong Jae Lee

Audio-driven talking head generation faces a fundamental trade-off between personalization and generalization, limiting its practical application. Implicit models often achieve generalization at the cost of structural incoherence, resulting…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Shiyu Liu , Kui Jiang , Junjun Jiang , Xianming Liu , Xiaocheng Feng , Hongxun Yao , Qi Tian

We introduce 3D-Speaker-Toolkit, an open-source toolkit for multimodal speaker verification and diarization, designed for meeting the needs of academic researchers and industrial practitioners. The 3D-Speaker-Toolkit adeptly leverages the…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-30 Yafeng Chen , Siqi Zheng , Hui Wang , Luyao Cheng , Tinglong Zhu , Rongjie Huang , Chong Deng , Qian Chen , Shiliang Zhang , Wen Wang , Xihao Li

Conversational automatic speech recognition remains challenging due to overlapping speech, far-field noise, and varying speaker counts. While recent LLM-based systems perform well on single-speaker benchmarks, their robustness in…

Computation and Language · Computer Science 2026-03-25 Naohiro Tawara , Samuele Cornell , Alexander Polok , Marc Delcroix , Lukáš Burget , Shinji Watanabe

Recent advancements in omnimodal learning have significantly improved understanding and generation across images, text, and speech, yet these developments remain predominantly confined to proprietary models. The lack of high-quality…

Computation and Language · Computer Science 2025-09-24 Run Luo , Ting-En Lin , Haonan Zhang , Yuchuan Wu , Xiong Liu , Min Yang , Yongbin Li , Longze Chen , Jiaming Li , Lei Zhang , Xiaobo Xia , Hamid Alinejad-Rokny , Fei Huang

Overlapping Speech Detection (OSD) aims to identify regions where multiple speakers overlap in a conversation, a critical challenge in multi-party speech processing. This work proposes a speaker-aware progressive OSD model that leverages a…

Sound · Computer Science 2025-05-30 Zhaokai Sun , Li Zhang , Qing Wang , Pan Zhou , Lei Xie

The rapid progress of Large Language Models (LLMs) has empowered omni models to act as voice assistants capable of understanding spoken dialogues. These models can process multimodal inputs beyond text, such as speech and visual data,…

The emergence of GPT-4o-like large multimodal models (LMMs) has raised the exploration of integrating text, vision, and speech modalities to support more flexible multimodal interaction. Existing LMMs typically concatenate representation of…

Artificial Intelligence · Computer Science 2025-06-24 Shaolei Zhang , Shoutao Guo , Qingkai Fang , Yan Zhou , Yang Feng

Multi-speaker automatic speech recognition (MS-ASR) faces significant challenges in transcribing overlapped speech, a task critical for applications like meeting transcription and conversational analysis. While serialized output training…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-09 Yuke Lin , Ming Cheng , Ze Li , Beilong Tang , Ming Li

Traditional speaker diarization seeks to detect ``who spoke when'' according to speaker characteristics. Extending to target speech diarization, we detect ``when target event occurs'' according to the semantic characteristics of speech. We…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-12 Yidi Jiang , Ruijie Tao , Zhengyang Chen , Yanmin Qian , Haizhou Li

TTM (Talking to Me) task is a pivotal component in understanding human social interactions, aiming to determine who is engaged in conversation with the camera-wearer. Traditional models often face challenges in real-world scenarios due to…

Multimedia · Computer Science 2026-03-20 Xinyuan Qian , Xinjia Zhu , Alessio Brutti , Dong Liang
‹ Prev 1 2 3 10 Next ›