English
Related papers

Related papers: Mini-Omni2: Towards Open-source GPT-4o with Vision…

200 papers

Recent advances in language models have achieved significant progress. GPT-4o, as a new milestone, has enabled real-time conversations with humans, demonstrating near-human natural fluency. Such human-computer interaction necessitates…

Artificial Intelligence · Computer Science 2024-11-06 Zhifei Xie , Changqiao Wu

We present M2-omni, a cutting-edge, open-source omni-MLLM that achieves competitive performance to GPT-4o. M2-omni employs a unified multimodal sequence modeling framework, which empowers Large Language Models(LLMs) to acquire comprehensive…

The remarkable multimodal capabilities and interactive experience of GPT-4o underscore their necessity in practical applications, yet open-source models rarely excel in both areas. In this paper, we introduce VITA, the first-ever…

Computer Vision and Pattern Recognition · Computer Science 2025-06-02 Chaoyou Fu , Haojia Lin , Zuwei Long , Yunhang Shen , Yuhang Dai , Meng Zhao , Yi-Fan Zhang , Shaoqi Dong , Yangze Li , Xiong Wang , Haoyu Cao , Di Yin , Long Ma , Xiawu Zheng , Rongrong Ji , Yunsheng Wu , Ran He , Caifeng Shan , Xing Sun

GPT-4o, an omni-modal model that enables vocal conversations with diverse emotions and tones, marks a milestone for omni-modal foundation models. However, empowering Large Language Models to perceive and generate images, texts, and speeches…

The salient multimodal capabilities and interactive experience of GPT-4o highlight its critical role in practical applications, yet it lacks a high-performing open-source counterpart. In this paper, we introduce Baichuan-omni, the first…

As large language models (LLMs) continue to advance, evaluating their comprehensive capabilities becomes significant for their application in various fields. This research study comprehensively evaluates the language, vision, speech, and…

Recent progress in multimodal large language models (MLLMs) has brought AI capabilities from static offline data processing to real-time streaming interaction, yet they still remain far from human-level multimodal interaction. The key…

GPT-4o is an autoregressive omni model that accepts as input any combination of text, audio, image, and video, and generates any combination of text, audio, and image outputs. It's trained end-to-end across text, vision, and audio, meaning…

Computation and Language · Computer Science 2024-10-29 OpenAI , : , Aaron Hurst , Adam Lerer , Adam P. Goucher , Adam Perelman , Aditya Ramesh , Aidan Clark , AJ Ostrow , Akila Welihinda , Alan Hayes , Alec Radford , Aleksander Mądry , Alex Baker-Whitcomb , Alex Beutel , Alex Borzunov , Alex Carney , Alex Chow , Alex Kirillov , Alex Nichol , Alex Paino , Alex Renzin , Alex Tachard Passos , Alexander Kirillov , Alexi Christakis , Alexis Conneau , Ali Kamali , Allan Jabri , Allison Moyer , Allison Tam , Amadou Crookes , Amin Tootoochian , Amin Tootoonchian , Ananya Kumar , Andrea Vallone , Andrej Karpathy , Andrew Braunstein , Andrew Cann , Andrew Codispoti , Andrew Galu , Andrew Kondrich , Andrew Tulloch , Andrey Mishchenko , Angela Baek , Angela Jiang , Antoine Pelisse , Antonia Woodford , Anuj Gosalia , Arka Dhar , Ashley Pantuliano , Avi Nayak , Avital Oliver , Barret Zoph , Behrooz Ghorbani , Ben Leimberger , Ben Rossen , Ben Sokolowsky , Ben Wang , Benjamin Zweig , Beth Hoover , Blake Samic , Bob McGrew , Bobby Spero , Bogo Giertler , Bowen Cheng , Brad Lightcap , Brandon Walkin , Brendan Quinn , Brian Guarraci , Brian Hsu , Bright Kellogg , Brydon Eastman , Camillo Lugaresi , Carroll Wainwright , Cary Bassin , Cary Hudson , Casey Chu , Chad Nelson , Chak Li , Chan Jun Shern , Channing Conger , Charlotte Barette , Chelsea Voss , Chen Ding , Cheng Lu , Chong Zhang , Chris Beaumont , Chris Hallacy , Chris Koch , Christian Gibson , Christina Kim , Christine Choi , Christine McLeavey , Christopher Hesse , Claudia Fischer , Clemens Winter , Coley Czarnecki , Colin Jarvis , Colin Wei , Constantin Koumouzelis , Dane Sherburn , Daniel Kappler , Daniel Levin , Daniel Levy , David Carr , David Farhi , David Mely , David Robinson , David Sasaki , Denny Jin , Dev Valladares , Dimitris Tsipras , Doug Li , Duc Phong Nguyen , Duncan Findlay , Edede Oiwoh , Edmund Wong , Ehsan Asdar , Elizabeth Proehl , Elizabeth Yang , Eric Antonow , Eric Kramer , Eric Peterson , Eric Sigler , Eric Wallace , Eugene Brevdo , Evan Mays , Farzad Khorasani , Felipe Petroski Such , Filippo Raso , Francis Zhang , Fred von Lohmann , Freddie Sulit , Gabriel Goh , Gene Oden , Geoff Salmon , Giulio Starace , Greg Brockman , Hadi Salman , Haiming Bao , Haitang Hu , Hannah Wong , Haoyu Wang , Heather Schmidt , Heather Whitney , Heewoo Jun , Hendrik Kirchner , Henrique Ponde de Oliveira Pinto , Hongyu Ren , Huiwen Chang , Hyung Won Chung , Ian Kivlichan , Ian O'Connell , Ian O'Connell , Ian Osband , Ian Silber , Ian Sohl , Ibrahim Okuyucu , Ikai Lan , Ilya Kostrikov , Ilya Sutskever , Ingmar Kanitscheider , Ishaan Gulrajani , Jacob Coxon , Jacob Menick , Jakub Pachocki , James Aung , James Betker , James Crooks , James Lennon , Jamie Kiros , Jan Leike , Jane Park , Jason Kwon , Jason Phang , Jason Teplitz , Jason Wei , Jason Wolfe , Jay Chen , Jeff Harris , Jenia Varavva , Jessica Gan Lee , Jessica Shieh , Ji Lin , Jiahui Yu , Jiayi Weng , Jie Tang , Jieqi Yu , Joanne Jang , Joaquin Quinonero Candela , Joe Beutler , Joe Landers , Joel Parish , Johannes Heidecke , John Schulman , Jonathan Lachman , Jonathan McKay , Jonathan Uesato , Jonathan Ward , Jong Wook Kim , Joost Huizinga , Jordan Sitkin , Jos Kraaijeveld , Josh Gross , Josh Kaplan , Josh Snyder , Joshua Achiam , Joy Jiao , Joyce Lee , Juntang Zhuang , Justyn Harriman , Kai Fricke , Kai Hayashi , Karan Singhal , Katy Shi , Kavin Karthik , Kayla Wood , Kendra Rimbach , Kenny Hsu , Kenny Nguyen , Keren Gu-Lemberg , Kevin Button , Kevin Liu , Kiel Howe , Krithika Muthukumar , Kyle Luther , Lama Ahmad , Larry Kai , Lauren Itow , Lauren Workman , Leher Pathak , Leo Chen , Li Jing , Lia Guy , Liam Fedus , Liang Zhou , Lien Mamitsuka , Lilian Weng , Lindsay McCallum , Lindsey Held , Long Ouyang , Louis Feuvrier , Lu Zhang , Lukas Kondraciuk , Lukasz Kaiser , Luke Hewitt , Luke Metz , Lyric Doshi , Mada Aflak , Maddie Simens , Madelaine Boyd , Madeleine Thompson , Marat Dukhan , Mark Chen , Mark Gray , Mark Hudnall , Marvin Zhang , Marwan Aljubeh , Mateusz Litwin , Matthew Zeng , Max Johnson , Maya Shetty , Mayank Gupta , Meghan Shah , Mehmet Yatbaz , Meng Jia Yang , Mengchao Zhong , Mia Glaese , Mianna Chen , Michael Janner , Michael Lampe , Michael Petrov , Michael Wu , Michele Wang , Michelle Fradin , Michelle Pokrass , Miguel Castro , Miguel Oom Temudo de Castro , Mikhail Pavlov , Miles Brundage , Miles Wang , Minal Khan , Mira Murati , Mo Bavarian , Molly Lin , Murat Yesildal , Nacho Soto , Natalia Gimelshein , Natalie Cone , Natalie Staudacher , Natalie Summers , Natan LaFontaine , Neil Chowdhury , Nick Ryder , Nick Stathas , Nick Turley , Nik Tezak , Niko Felix , Nithanth Kudige , Nitish Keskar , Noah Deutsch , Noel Bundick , Nora Puckett , Ofir Nachum , Ola Okelola , Oleg Boiko , Oleg Murk , Oliver Jaffe , Olivia Watkins , Olivier Godement , Owen Campbell-Moore , Patrick Chao , Paul McMillan , Pavel Belov , Peng Su , Peter Bak , Peter Bakkum , Peter Deng , Peter Dolan , Peter Hoeschele , Peter Welinder , Phil Tillet , Philip Pronin , Philippe Tillet , Prafulla Dhariwal , Qiming Yuan , Rachel Dias , Rachel Lim , Rahul Arora , Rajan Troll , Randall Lin , Rapha Gontijo Lopes , Raul Puri , Reah Miyara , Reimar Leike , Renaud Gaubert , Reza Zamani , Ricky Wang , Rob Donnelly , Rob Honsby , Rocky Smith , Rohan Sahai , Rohit Ramchandani , Romain Huet , Rory Carmichael , Rowan Zellers , Roy Chen , Ruby Chen , Ruslan Nigmatullin , Ryan Cheu , Saachi Jain , Sam Altman , Sam Schoenholz , Sam Toizer , Samuel Miserendino , Sandhini Agarwal , Sara Culver , Scott Ethersmith , Scott Gray , Sean Grove , Sean Metzger , Shamez Hermani , Shantanu Jain , Shengjia Zhao , Sherwin Wu , Shino Jomoto , Shirong Wu , Shuaiqi , Xia , Sonia Phene , Spencer Papay , Srinivas Narayanan , Steve Coffey , Steve Lee , Stewart Hall , Suchir Balaji , Tal Broda , Tal Stramer , Tao Xu , Tarun Gogineni , Taya Christianson , Ted Sanders , Tejal Patwardhan , Thomas Cunninghman , Thomas Degry , Thomas Dimson , Thomas Raoux , Thomas Shadwell , Tianhao Zheng , Todd Underwood , Todor Markov , Toki Sherbakov , Tom Rubin , Tom Stasi , Tomer Kaftan , Tristan Heywood , Troy Peterson , Tyce Walters , Tyna Eloundou , Valerie Qi , Veit Moeller , Vinnie Monaco , Vishal Kuo , Vlad Fomenko , Wayne Chang , Weiyi Zheng , Wenda Zhou , Wesam Manassra , Will Sheu , Wojciech Zaremba , Yash Patil , Yilei Qian , Yongjik Kim , Youlong Cheng , Yu Zhang , Yuchen He , Yuchen Zhang , Yujia Jin , Yunxing Dai , Yury Malkov

Multimodal conversational agents are highly desirable because they offer natural and human-like interaction. However, there is a lack of comprehensive end-to-end solutions to support collaborative development and benchmarking. While…

Human-Computer Interaction · Computer Science 2024-11-19 Qiang Sun , Yuanyi Luo , Sirui Li , Wenxiao Zhang , Wei Liu

We propose Ming-Omni, a unified multimodal model capable of processing images, text, audio, and video, while demonstrating strong proficiency in both speech and image generation. Ming-Omni employs dedicated encoders to extract tokens from…

Recent advances in GPT-4o like multi-modality models have demonstrated remarkable progress for direct speech-to-speech conversation, with real-time speech interaction experience and strong speech understanding ability. However, current…

Sound · Computer Science 2024-12-09 Ze Yuan , Yanqing Liu , Shujie Liu , Sheng Zhao

Rapidly developing large language models (LLMs) have brought tremendous intelligent applications. Especially, the GPT-4o's excellent duplex speech interaction ability has brought impressive experience to users. Researchers have recently…

Sound · Computer Science 2024-12-10 Xiong Wang , Yangze Li , Chaoyou Fu , Yunhang Shen , Lei Xie , Ke Li , Xing Sun , Long Ma

The emergence of GPT-4o-like large multimodal models (LMMs) has raised the exploration of integrating text, vision, and speech modalities to support more flexible multimodal interaction. Existing LMMs typically concatenate representation of…

Artificial Intelligence · Computer Science 2025-06-24 Shaolei Zhang , Shoutao Guo , Qingkai Fang , Yan Zhou , Yang Feng

Models like GPT-4o enable real-time interaction with large language models (LLMs) through speech, significantly enhancing user experience compared to traditional text-based interaction. However, there is still a lack of exploration on how…

Computation and Language · Computer Science 2025-03-04 Qingkai Fang , Shoutao Guo , Yan Zhou , Zhengrui Ma , Shaolei Zhang , Yang Feng

Large language models have shown their remarkable capabilities as a general interface for various language-related applications. Motivated by this, we target to build a unified interface for completing many vision-language tasks including…

Computer Vision and Pattern Recognition · Computer Science 2023-11-08 Jun Chen , Deyao Zhu , Xiaoqian Shen , Xiang Li , Zechun Liu , Pengchuan Zhang , Raghuraman Krishnamoorthi , Vikas Chandra , Yunyang Xiong , Mohamed Elhoseiny

The recent GPT-4 has demonstrated extraordinary multi-modal abilities, such as directly generating websites from handwritten text and identifying humorous elements within images. These features are rarely observed in previous…

Computer Vision and Pattern Recognition · Computer Science 2023-10-03 Deyao Zhu , Jun Chen , Xiaoqian Shen , Xiang Li , Mohamed Elhoseiny

We present a vision and language model named MultiModal-GPT to conduct multi-round dialogue with humans. MultiModal-GPT can follow various instructions from humans, such as generating a detailed caption, counting the number of interested…

Computer Vision and Pattern Recognition · Computer Science 2023-06-14 Tao Gong , Chengqi Lyu , Shilong Zhang , Yudong Wang , Miao Zheng , Qian Zhao , Kuikun Liu , Wenwei Zhang , Ping Luo , Kai Chen

We introduce Baichuan-Omni-1.5, an omni-modal model that not only has omni-modal understanding capabilities but also provides end-to-end audio generation capabilities. To achieve fluent and high-quality interaction across modalities without…

Full-duplex spoken dialogue systems significantly surpass traditional turn-based dialogue systems, as they allow simultaneous bidirectional communication, closely mirroring human-human interactions. However, achieving low latency and…

Computation and Language · Computer Science 2025-01-06 Qinglin Zhang , Luyao Cheng , Chong Deng , Qian Chen , Wen Wang , Siqi Zheng , Jiaqing Liu , Hai Yu , Chaohong Tan , Zhihao Du , Shiliang Zhang

In this paper, we introduce MIO, a novel foundation model built on multimodal tokens, capable of understanding and generating speech, text, images, and videos in an end-to-end, autoregressive manner. While the emergence of large language…

‹ Prev 1 2 3 10 Next ›