中文
相关论文

相关论文: SocialOmni: Benchmarking Audio-Visual Social Inter…

200 篇论文

We introduce InteractiveOmni, a unified and open-source omni-modal large language model for audio-visual multi-turn interaction, ranging from 4B to 8B parameters, designed to lead the field of lightweight models by offering comprehensive…

In human-centric scenes, the ability to simultaneously understand visual and auditory information is crucial. While recent omni models can process multiple modalities, they generally lack effectiveness in human-centric scenes due to the…

计算机视觉与模式识别 · 计算机科学 2025-01-28 Jiaxing Zhao , Qize Yang , Yixing Peng , Detao Bai , Shimin Yao , Boyuan Sun , Xiang Chen , Shenghao Fu , Weixuan chen , Xihan Wei , Liefeng Bo

Full-Duplex Speech-to-Speech Large Language Models (LLMs) are foundational to natural human-computer interaction, enabling real-time spoken dialogue systems. However, benchmarking and modeling these models remains a fundamental challenge.…

计算与语言 · 计算机科学 2025-09-29 Yuan Ge , Saihan Chen , Jingqi Xiao , Xiaoqian Liu , Tong Xiao , Yan Xiang , Zhengtao Yu , Jingbo Zhu

Recent Multimodal Large Language Models (MLLMs) achieve promising performance on visual and audio benchmarks independently. However, the ability of these models to process cross-modal information synchronously remains largely unexplored. We…

人工智能 · 计算机科学 2026-03-12 Ziwei Zhou , Rui Wang , Zuxuan Wu , Yu-Gang Jiang

Speech language models (SLMs) have significantly extended the interactive capability of text-based Large Language Models (LLMs) by incorporating paralinguistic information. For more realistic interactive experience with customized styles,…

计算与语言 · 计算机科学 2026-03-10 Haishu Zhao , Aokai Hao , Yuan Ge , Zhenqiang Hong , Tong Xiao , Jingbo Zhu

Speech large language models (SpeechLLMs) have extended human-machine interactions from the text modality to the dynamic speech domain. Spoken dialogues convey diverse information, including semantic concepts, acoustic variations,…

计算与语言 · 计算机科学 2026-01-14 Heyang Liu , Yuhao Wang , Ziyang Cheng , Hongcheng Liu , Yiqi Li , Yixuan Hou , Ronghua Wu , Qunshan Gu , Yanfeng Wang , Yu Wang

The emergence of GPT-4o-like large multimodal models (LMMs) has raised the exploration of integrating text, vision, and speech modalities to support more flexible multimodal interaction. Existing LMMs typically concatenate representation of…

人工智能 · 计算机科学 2025-06-24 Shaolei Zhang , Shoutao Guo , Qingkai Fang , Yan Zhou , Yang Feng

The rapid advancement of multi-modal language models (MLLMs) like GPT-4o has propelled the development of Omni language models, designed to process and proactively respond to continuous streams of multi-modal data. Despite their potential,…

计算机视觉与模式识别 · 计算机科学 2025-04-01 Yuxuan Wang , Yueqian Wang , Bo Chen , Tong Wu , Dongyan Zhao , Zilong Zheng

The rapid progress of Large Language Models (LLMs) has empowered omni models to act as voice assistants capable of understanding spoken dialogues. These models can process multimodal inputs beyond text, such as speech and visual data,…

Real-time, intelligent, and natural speech interaction is an essential part of the next-generation human-computer interaction. Recent advancements have showcased the potential of building intelligent spoken chatbots based on large language…

计算与语言 · 计算机科学 2025-05-06 Qingkai Fang , Yan Zhou , Shoutao Guo , Shaolei Zhang , Yang Feng

With the development of speech large language models (speech LLMs), users can now interact directly with assistants via speech. However, most existing models only convert response content into speech without fully capturing the rich…

计算与语言 · 计算机科学 2025-09-18 Haoyu Wang , Guangyan Zhang , Jiale Chen , Jingyu Li , Yuehai Wang , Yiwen Guo

As large language models (LLMs) develop anthropomorphic abilities, they are increasingly being deployed as autonomous agents to interact with humans. However, evaluating their performance in realistic and complex social interactions remains…

计算与语言 · 计算机科学 2025-10-28 Shuai Huang , Wenxuan Zhao , Jun Gao

Multimodal conversational agents are highly desirable because they offer natural and human-like interaction. However, there is a lack of comprehensive end-to-end solutions to support collaborative development and benchmarking. While…

人机交互 · 计算机科学 2024-11-19 Qiang Sun , Yuanyi Luo , Sirui Li , Wenxiao Zhang , Wei Liu

Multimodal Large Language Models (MLLMs) are a major focus of recent AI research. However, most prior work focuses on static image understanding, while their ability to process sequential audio-video data remains underexplored. This gap…

人工智能 · 计算机科学 2026-05-28 Ahmed Y. Radwan , Christos Emmanouilidis , Hina Tabassum , Deval Pandya , Shaina Raza

The rise of Omni-modal Large Language Models (OLLMs), which integrate visual and auditory processing with text, necessitates robust safety evaluations to mitigate harmful outputs. However, no dedicated benchmarks currently exist for OLLMs,…

While Omni-modal Large Language Models have made strides in joint sensory processing, they fundamentally struggle with a cornerstone of human interaction: deciphering complex, multi-person conversational dynamics to accurately answer ``Who…

计算机视觉与模式识别 · 计算机科学 2026-04-30 Detao Bai , Shimin Yao , Weixuan Chen , Zhiheng Ma , Xihan Wei , Jingren Zhou

Spoken Dialogue Models (SDMs) have recently attracted significant attention for their ability to generate voice responses directly to users' spoken queries. Despite their increasing popularity, there exists a gap in research focused on…

计算与语言 · 计算机科学 2025-10-07 Chengqian Ma , Wei Tao , Yiwen Guo

Recent advancements in multimodal large language models (MLLMs) have aimed to integrate and interpret data across diverse modalities. However, the capacity of these models to concurrently process and reason about multiple modalities remains…

Although Multimodal Large Language Models (MLLMs) demonstrate strong omni-modal perception, their ability to forecast future events from audio-visual cues remains largely unexplored, as existing benchmarks focus mainly on retrospective…

计算与语言 · 计算机科学 2026-01-21 Qian Chen , Jinlan Fu , Changsong Li , See-Kiong Ng , Xipeng Qiu

Real-time duplex interaction is essential for multimodal AI systems operating in real-world scenarios, where models must continuously process streaming inputs and respond at appropriate moments. However, most existing multimodal large…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Chaoqun He , Mingyang Xiang , Yingjing Xu , Bokai Xu , Junbo Cui , Jie Zhou , Yuan Yao , Lijie Wen
‹ 上一页 1 2 3 10 下一页 ›