中文
相关论文

相关论文: SALMONN-omni: A Standalone Speech LLM without Code…

200 篇论文

Full-duplex multimodal large language models (LLMs) provide a unified framework for addressing diverse speech understanding and generation tasks, enabling more natural and seamless human-machine conversations. Unlike traditional modularised…

音频与语音处理 · 电气工程与系统科学 2024-11-28 Wenyi Yu , Siyin Wang , Xiaoyu Yang , Xianzhao Chen , Xiaohai Tian , Jun Zhang , Guangzhi Sun , Lu Lu , Yuxuan Wang , Chao Zhang

Hearing is arguably an essential ability of artificial intelligence (AI) agents in the physical world, which refers to the perception and understanding of general auditory information consisting of at least three types of sounds: speech,…

声音 · 计算机科学 2024-04-09 Changli Tang , Wenyi Yu , Guangzhi Sun , Xianzhao Chen , Tian Tan , Wei Li , Lu Lu , Zejun Ma , Chao Zhang

Rapidly developing large language models (LLMs) have brought tremendous intelligent applications. Especially, the GPT-4o's excellent duplex speech interaction ability has brought impressive experience to users. Researchers have recently…

声音 · 计算机科学 2024-12-10 Xiong Wang , Yangze Li , Chaoyou Fu , Yunhang Shen , Lei Xie , Ke Li , Xing Sun , Long Ma

Recent advances in GPT-4o like multi-modality models have demonstrated remarkable progress for direct speech-to-speech conversation, with real-time speech interaction experience and strong speech understanding ability. However, current…

声音 · 计算机科学 2024-12-09 Ze Yuan , Yanqing Liu , Shujie Liu , Sheng Zhao

Models like GPT-4o enable real-time interaction with large language models (LLMs) through speech, significantly enhancing user experience compared to traditional text-based interaction. However, there is still a lack of exploration on how…

计算与语言 · 计算机科学 2025-03-04 Qingkai Fang , Shoutao Guo , Yan Zhou , Zhengrui Ma , Shaolei Zhang , Yang Feng

We present a generative dialogue system capable of operating in a full-duplex manner, allowing for seamless interaction. It is based on a large language model (LLM) carefully aligned to be aware of a perception module, a motor function…

计算与语言 · 计算机科学 2024-10-30 Peng Wang , Songshuo Lu , Yaohua Tang , Sijie Yan , Wei Xia , Yuanjun Xiong

Recent advancements highlight the potential of end-to-end real-time spoken dialogue systems, showcasing their low latency and high quality. In this paper, we introduce SLAM-Omni, a timbre-controllable, end-to-end voice interaction system…

音频与语音处理 · 电气工程与系统科学 2024-12-23 Wenxi Chen , Ziyang Ma , Ruiqi Yan , Yuzhe Liang , Xiquan Li , Ruiyang Xu , Zhikang Niu , Yanqiao Zhu , Yifan Yang , Zhanxun Liu , Kai Yu , Yuxuan Hu , Jinyu Li , Yan Lu , Shujie Liu , Xie Chen

Reasoning is essential for effective communication and decision-making. While recent advances in LLMs and MLLMs have shown that incorporating explicit reasoning significantly improves understanding and generalization, reasoning in LSMs…

计算与语言 · 计算机科学 2025-09-23 Zhifei Xie , Ziyang Ma , Zihang Liu , Kaiyu Pang , Hongyu Li , Jialin Zhang , Yue Liao , Deheng Ye , Chunyan Miao , Shuicheng Yan

Full-duplex spoken dialogue systems significantly surpass traditional turn-based dialogue systems, as they allow simultaneous bidirectional communication, closely mirroring human-human interactions. However, achieving low latency and…

计算与语言 · 计算机科学 2025-01-06 Qinglin Zhang , Luyao Cheng , Chong Deng , Qian Chen , Wen Wang , Siqi Zheng , Jiaqing Liu , Hai Yu , Chaohong Tan , Zhihao Du , Shiliang Zhang

This paper presents SOLOMON, a novel Neuro-inspired Large Language Model (LLM) Reasoning Network architecture that enhances the adaptability of foundation models for domain-specific applications. Through a case study in semiconductor layout…

计算与语言 · 计算机科学 2025-02-10 Bo Wen , Xin Zhang

We present MGM-Omni, a unified Omni LLM for omni-modal understanding and expressive, long-horizon speech generation. Unlike cascaded pipelines that isolate speech synthesis, MGM-Omni adopts a "brain-mouth" design with a dual-track,…

声音 · 计算机科学 2025-09-30 Chengyao Wang , Zhisheng Zhong , Bohao Peng , Senqiao Yang , Yuqi Liu , Haokun Gui , Bin Xia , Jingyao Li , Bei Yu , Jiaya Jia

Native multimodal large language models (MLLMs) restructure a single large language model (LLM) into a spoken language model (SLM) capable of both speech and text generation. Compared to modular and aligned MLLMs, native MLLMs preserve…

计算与语言 · 计算机科学 2025-10-28 Hang Shao , Heting Gao , Yunhang Shen , Jiawei Chen , Zuwei Long , Dong Yang , Ke Li , Xing Sun

Dialogue serves as the most natural manner of human-computer interaction (HCI). Recent advancements in speech language models (SLM) have significantly enhanced speech-based conversational AI. However, these models are limited to turn-based…

计算与语言 · 计算机科学 2024-08-06 Ziyang Ma , Yakun Song , Chenpeng Du , Jian Cong , Zhuo Chen , Yuping Wang , Yuxuan Wang , Xie Chen

Joint audio-visual reasoning is essential for omnimodal understanding, yet current multimodal large language models (MLLMs) still struggle when reasoning requires fine-grained evidence from both modalities. A central limitation is that…

Omni-modal reasoning is essential for intelligent systems to understand and draw inferences from diverse data sources. While existing omni-modal large language models (OLLM) excel at perceiving diverse modalities, they lack the complex…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Yiran Guan , Sifan Tu , Dingkang Liang , Linghao Zhu , Jianzhong Ju , Zhenbo Luo , Jian Luan , Yuliang Liu , Xiang Bai

Speech understanding as an element of the more generic video understanding using audio-visual large language models (av-LLMs) is a crucial yet understudied aspect. This paper proposes video-SALMONN, a single end-to-end av-LLM for video…

计算机视觉与模式识别 · 计算机科学 2024-06-25 Guangzhi Sun , Wenyi Yu , Changli Tang , Xianzhao Chen , Tian Tan , Wei Li , Lu Lu , Zejun Ma , Yuxuan Wang , Chao Zhang

Despite broad interest in modeling spoken dialogue agents, most approaches are inherently "half-duplex" -- restricted to turn-based interaction with responses requiring explicit prompting by the user or implicit tracking of interruption or…

计算与语言 · 计算机科学 2024-09-25 Bandhav Veluri , Benjamin N Peloquin , Bokai Yu , Hongyu Gong , Shyamnath Gollakota

Recent advances in spoken dialogue language models (SDLMs) reflect growing interest in shifting from turn-based to full-duplex systems, where the models continuously perceive user speech streams while generating responses. This simultaneous…

计算与语言 · 计算机科学 2025-10-10 Donghang Wu , Haoyang Zhang , Chen Chen , Tianyu Zhang , Fei Tian , Xuerui Yang , Gang Yu , Hexin Liu , Nana Hou , Yuchen Hu , Eng Siong Chng

Full-Duplex Speech-to-Speech Large Language Models (LLMs) are foundational to natural human-computer interaction, enabling real-time spoken dialogue systems. However, benchmarking and modeling these models remains a fundamental challenge.…

计算与语言 · 计算机科学 2025-09-29 Yuan Ge , Saihan Chen , Jingqi Xiao , Xiaoqian Liu , Tong Xiao , Yan Xiang , Zhengtao Yu , Jingbo Zhu

Existing human-robot interaction systems often lack mechanisms for sustained personalization and dynamic adaptation in multi-user environments, limiting their effectiveness in real-world deployments. We present HARMONI, a multimodal…

‹ 上一页 1 2 3 10 下一页 ›