English
Related papers

Related papers: LLaMA-Omni: Seamless Speech Interaction with Large…

200 papers

Real-time, intelligent, and natural speech interaction is an essential part of the next-generation human-computer interaction. Recent advancements have showcased the potential of building intelligent spoken chatbots based on large language…

Computation and Language · Computer Science 2025-05-06 Qingkai Fang , Yan Zhou , Shoutao Guo , Shaolei Zhang , Yang Feng

Rapidly developing large language models (LLMs) have brought tremendous intelligent applications. Especially, the GPT-4o's excellent duplex speech interaction ability has brought impressive experience to users. Researchers have recently…

Sound · Computer Science 2024-12-10 Xiong Wang , Yangze Li , Chaoyou Fu , Yunhang Shen , Lei Xie , Ke Li , Xing Sun , Long Ma

Recent advances in language models have achieved significant progress. GPT-4o, as a new milestone, has enabled real-time conversations with humans, demonstrating near-human natural fluency. Such human-computer interaction necessitates…

Artificial Intelligence · Computer Science 2024-11-06 Zhifei Xie , Changqiao Wu

The GPT-4o represents a significant milestone in enabling real-time interaction with large language models (LLMs) through speech, its remarkable low latency and high fluency not only capture attention but also stimulate research interest in…

Computation and Language · Computer Science 2024-12-04 Shuaijiang Zhao , Tingwei Guo , Bajian Xiang , Tongtang Wan , Qiang Niu , Wei Zou , Xiangang Li

With the development of speech large language models (speech LLMs), users can now interact directly with assistants via speech. However, most existing models only convert response content into speech without fully capturing the rich…

Computation and Language · Computer Science 2025-09-18 Haoyu Wang , Guangyan Zhang , Jiale Chen , Jingyu Li , Yuehai Wang , Yiwen Guo

The emergence of GPT-4o-like large multimodal models (LMMs) has raised the exploration of integrating text, vision, and speech modalities to support more flexible multimodal interaction. Existing LMMs typically concatenate representation of…

Artificial Intelligence · Computer Science 2025-06-24 Shaolei Zhang , Shoutao Guo , Qingkai Fang , Yan Zhou , Yang Feng

Recent advances in GPT-4o like multi-modality models have demonstrated remarkable progress for direct speech-to-speech conversation, with real-time speech interaction experience and strong speech understanding ability. However, current…

Sound · Computer Science 2024-12-09 Ze Yuan , Yanqing Liu , Shujie Liu , Sheng Zhao

As Large Language Models (LLMs) advance in natural language processing, there is growing interest in leveraging their capabilities to simplify software interactions. In this paper, we propose a novel system that integrates LLMs for both…

Computation and Language · Computer Science 2024-09-19 Chunliang Tao , Xiaojing Fan , Yahe Yang

We present MGM-Omni, a unified Omni LLM for omni-modal understanding and expressive, long-horizon speech generation. Unlike cascaded pipelines that isolate speech synthesis, MGM-Omni adopts a "brain-mouth" design with a dual-track,…

Sound · Computer Science 2025-09-30 Chengyao Wang , Zhisheng Zhong , Bohao Peng , Senqiao Yang , Yuqi Liu , Haokun Gui , Bin Xia , Jingyao Li , Bei Yu , Jiaya Jia

Recent advancements highlight the potential of end-to-end real-time spoken dialogue systems, showcasing their low latency and high quality. In this paper, we introduce SLAM-Omni, a timbre-controllable, end-to-end voice interaction system…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-23 Wenxi Chen , Ziyang Ma , Ruiqi Yan , Yuzhe Liang , Xiquan Li , Ruiyang Xu , Zhikang Niu , Yanqiao Zhu , Yifan Yang , Zhanxun Liu , Kai Yu , Yuxuan Hu , Jinyu Li , Yan Lu , Shujie Liu , Xie Chen

We present a generative dialogue system capable of operating in a full-duplex manner, allowing for seamless interaction. It is based on a large language model (LLM) carefully aligned to be aware of a perception module, a motor function…

Computation and Language · Computer Science 2024-10-30 Peng Wang , Songshuo Lu , Yaohua Tang , Sijie Yan , Wei Xia , Yuanjun Xiong

Full-duplex spoken dialogue systems significantly surpass traditional turn-based dialogue systems, as they allow simultaneous bidirectional communication, closely mirroring human-human interactions. However, achieving low latency and…

Computation and Language · Computer Science 2025-01-06 Qinglin Zhang , Luyao Cheng , Chong Deng , Qian Chen , Wen Wang , Siqi Zheng , Jiaqing Liu , Hai Yu , Chaohong Tan , Zhihao Du , Shiliang Zhang

Recently, the powerful text-to-image capabilities of ChatGPT-4o have led to growing appreciation for native multimodal large language models. However, its multimodal capabilities remain confined to images and text. Yet beyond images, the…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Junliang Ye , Zhengyi Wang , Ruowen Zhao , Shenghao Xie , Jun Zhu

We introduce InteractiveOmni, a unified and open-source omni-modal large language model for audio-visual multi-turn interaction, ranging from 4B to 8B parameters, designed to lead the field of lightweight models by offering comprehensive…

Omni-modal large language models (OLLMs) aim to unify multimodal understanding and generation, yet incorporating speech with 3D facial animation remains largely unexplored despite its importance for natural interaction. A key challenge…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Haoyu Zhang , Zhipeng Li , Yiwen Guo , Tianshu Yu

Large language models (LLMs) and their variants have shown extraordinary efficacy across numerous downstream natural language processing (NLP) tasks, which has presented a new vision for the development of NLP. Despite their remarkable…

Computation and Language · Computer Science 2024-01-18 Yazhou Zhang , Mengyao Wang , Youxi Wu , Prayag Tiwari , Qiuchi Li , Benyou Wang , Jing Qin

We present LLaMA-Adapter, a lightweight adaption method to efficiently fine-tune LLaMA into an instruction-following model. Using 52K self-instruct demonstrations, LLaMA-Adapter only introduces 1.2M learnable parameters upon the frozen…

Computer Vision and Pattern Recognition · Computer Science 2024-09-20 Renrui Zhang , Jiaming Han , Chris Liu , Peng Gao , Aojun Zhou , Xiangfei Hu , Shilin Yan , Pan Lu , Hongsheng Li , Yu Qiao

Pre-trained Large Language Models (LLMs) have revolutionized text processing, yet adapting Transformer-based neural networks to non-textual scientific modalities typically requires specialized architectures and extensive computational…

Instrumentation and Methods for Astrophysics · Physics 2025-08-15 Nesar Ramachandra , Yuan-Sen Ting , Zechang Sun , Azton Wells , Salman Habib

Despite broad interest in modeling spoken dialogue agents, most approaches are inherently "half-duplex" -- restricted to turn-based interaction with responses requiring explicit prompting by the user or implicit tracking of interruption or…

Computation and Language · Computer Science 2024-09-25 Bandhav Veluri , Benjamin N Peloquin , Bokai Yu , Hongyu Gong , Shyamnath Gollakota

Full-duplex multimodal large language models (LLMs) provide a unified framework for addressing diverse speech understanding and generation tasks, enabling more natural and seamless human-machine conversations. Unlike traditional modularised…

Audio and Speech Processing · Electrical Eng. & Systems 2024-11-28 Wenyi Yu , Siyin Wang , Xiaoyu Yang , Xianzhao Chen , Xiaohai Tian , Jun Zhang , Guangzhi Sun , Lu Lu , Yuxuan Wang , Chao Zhang
‹ Prev 1 2 3 10 Next ›