English
Related papers

Related papers: MiniMind-O Technical Report: An Open Small-Scale S…

200 papers

Recent advances in language models have achieved significant progress. GPT-4o, as a new milestone, has enabled real-time conversations with humans, demonstrating near-human natural fluency. Such human-computer interaction necessitates…

Artificial Intelligence · Computer Science 2024-11-06 Zhifei Xie , Changqiao Wu

Recent progress in multimodal large language models (MLLMs) has brought AI capabilities from static offline data processing to real-time streaming interaction, yet they still remain far from human-level multimodal interaction. The key…

We introduce MiniMax-Speech, an autoregressive Transformer-based Text-to-Speech (TTS) model that generates high-quality speech. A key innovation is our learnable speaker encoder, which extracts timbre features from a reference audio without…

In this report, we present Qwen2.5-Omni, an end-to-end multimodal model designed to perceive diverse modalities, including text, images, audio, and video, while simultaneously generating text and natural speech responses in a streaming…

Computation and Language · Computer Science 2025-03-27 Jin Xu , Zhifang Guo , Jinzheng He , Hangrui Hu , Ting He , Shuai Bai , Keqin Chen , Jialin Wang , Yang Fan , Kai Dang , Bin Zhang , Xiong Wang , Yunfei Chu , Junyang Lin

Although significant progress has been made in audio-driven talking head generation, text-driven methods remain underexplored. In this work, we present OmniTalker, a unified framework that jointly generates synchronized talking audio-video…

Computer Vision and Pattern Recognition · Computer Science 2025-06-04 Zhongjian Wang , Peng Zhang , Jinwei Qi , Guangyuan Wang , Chaonan Ji , Sheng Xu , Bang Zhang , Liefeng Bo

Recent advancements in large language models (LLMs) and multimodal speech-text models have laid the groundwork for seamless voice interactions, enabling real-time, natural, and human-like conversations. Previous models for voice…

Reasoning is essential for effective communication and decision-making. While recent advances in LLMs and MLLMs have shown that incorporating explicit reasoning significantly improves understanding and generalization, reasoning in LSMs…

Computation and Language · Computer Science 2025-09-23 Zhifei Xie , Ziyang Ma , Zihang Liu , Kaiyu Pang , Hongyu Li , Jialin Zhang , Yue Liao , Deheng Ye , Chunyan Miao , Shuicheng Yan

GPT-4o, an all-encompassing model, represents a milestone in the development of large multi-modal language models. It can understand visual, auditory, and textual modalities, directly output audio, and support flexible duplex interaction.…

Audio and Speech Processing · Electrical Eng. & Systems 2024-11-06 Zhifei Xie , Changqiao Wu

We present Qwen3-Omni, a single multimodal model that, for the first time, maintains state-of-the-art performance across text, image, audio, and video without any degradation relative to single-modal counterparts. Qwen3-Omni matches the…

Fine-grained perception of multimodal information is critical for advancing human-AI interaction. With recent progress in audio-visual technologies, Omni Language Models (OLMs), capable of processing audio and video signals in parallel,…

Computation and Language · Computer Science 2026-03-17 Ziyang Ma , Ruiyang Xu , Zhenghao Xing , Yunfei Chu , Yuxuan Wang , Jinzheng He , Jin Xu , Pheng-Ann Heng , Kai Yu , Junyang Lin , Eng Siong Chng , Xie Chen

We propose Ming-Omni, a unified multimodal model capable of processing images, text, audio, and video, while demonstrating strong proficiency in both speech and image generation. Ming-Omni employs dedicated encoders to extract tokens from…

Rapidly developing large language models (LLMs) have brought tremendous intelligent applications. Especially, the GPT-4o's excellent duplex speech interaction ability has brought impressive experience to users. Researchers have recently…

Sound · Computer Science 2024-12-10 Xiong Wang , Yangze Li , Chaoyou Fu , Yunhang Shen , Lei Xie , Ke Li , Xing Sun , Long Ma

We present OmniVoice, a massively multilingual zero-shot text-to-speech (TTS) model that scales to over 600 languages. At its core is a novel diffusion language model-style discrete non-autoregressive (NAR) architecture. Unlike conventional…

Computation and Language · Computer Science 2026-04-22 Han Zhu , Lingxuan Ye , Wei Kang , Zengwei Yao , Liyong Guo , Fangjun Kuang , Zhifeng Han , Weiji Zhuang , Long Lin , Daniel Povey

Existing audio language models typically rely on task-specific fine-tuning to accomplish particular audio tasks. In contrast, humans are able to generalize to new audio tasks with only a few examples or simple instructions. GPT-3 has shown…

GPT-4o is an autoregressive omni model that accepts as input any combination of text, audio, image, and video, and generates any combination of text, audio, and image outputs. It's trained end-to-end across text, vision, and audio, meaning…

Computation and Language · Computer Science 2024-10-29 OpenAI , : , Aaron Hurst , Adam Lerer , Adam P. Goucher , Adam Perelman , Aditya Ramesh , Aidan Clark , AJ Ostrow , Akila Welihinda , Alan Hayes , Alec Radford , Aleksander Mądry , Alex Baker-Whitcomb , Alex Beutel , Alex Borzunov , Alex Carney , Alex Chow , Alex Kirillov , Alex Nichol , Alex Paino , Alex Renzin , Alex Tachard Passos , Alexander Kirillov , Alexi Christakis , Alexis Conneau , Ali Kamali , Allan Jabri , Allison Moyer , Allison Tam , Amadou Crookes , Amin Tootoochian , Amin Tootoonchian , Ananya Kumar , Andrea Vallone , Andrej Karpathy , Andrew Braunstein , Andrew Cann , Andrew Codispoti , Andrew Galu , Andrew Kondrich , Andrew Tulloch , Andrey Mishchenko , Angela Baek , Angela Jiang , Antoine Pelisse , Antonia Woodford , Anuj Gosalia , Arka Dhar , Ashley Pantuliano , Avi Nayak , Avital Oliver , Barret Zoph , Behrooz Ghorbani , Ben Leimberger , Ben Rossen , Ben Sokolowsky , Ben Wang , Benjamin Zweig , Beth Hoover , Blake Samic , Bob McGrew , Bobby Spero , Bogo Giertler , Bowen Cheng , Brad Lightcap , Brandon Walkin , Brendan Quinn , Brian Guarraci , Brian Hsu , Bright Kellogg , Brydon Eastman , Camillo Lugaresi , Carroll Wainwright , Cary Bassin , Cary Hudson , Casey Chu , Chad Nelson , Chak Li , Chan Jun Shern , Channing Conger , Charlotte Barette , Chelsea Voss , Chen Ding , Cheng Lu , Chong Zhang , Chris Beaumont , Chris Hallacy , Chris Koch , Christian Gibson , Christina Kim , Christine Choi , Christine McLeavey , Christopher Hesse , Claudia Fischer , Clemens Winter , Coley Czarnecki , Colin Jarvis , Colin Wei , Constantin Koumouzelis , Dane Sherburn , Daniel Kappler , Daniel Levin , Daniel Levy , David Carr , David Farhi , David Mely , David Robinson , David Sasaki , Denny Jin , Dev Valladares , Dimitris Tsipras , Doug Li , Duc Phong Nguyen , Duncan Findlay , Edede Oiwoh , Edmund Wong , Ehsan Asdar , Elizabeth Proehl , Elizabeth Yang , Eric Antonow , Eric Kramer , Eric Peterson , Eric Sigler , Eric Wallace , Eugene Brevdo , Evan Mays , Farzad Khorasani , Felipe Petroski Such , Filippo Raso , Francis Zhang , Fred von Lohmann , Freddie Sulit , Gabriel Goh , Gene Oden , Geoff Salmon , Giulio Starace , Greg Brockman , Hadi Salman , Haiming Bao , Haitang Hu , Hannah Wong , Haoyu Wang , Heather Schmidt , Heather Whitney , Heewoo Jun , Hendrik Kirchner , Henrique Ponde de Oliveira Pinto , Hongyu Ren , Huiwen Chang , Hyung Won Chung , Ian Kivlichan , Ian O'Connell , Ian O'Connell , Ian Osband , Ian Silber , Ian Sohl , Ibrahim Okuyucu , Ikai Lan , Ilya Kostrikov , Ilya Sutskever , Ingmar Kanitscheider , Ishaan Gulrajani , Jacob Coxon , Jacob Menick , Jakub Pachocki , James Aung , James Betker , James Crooks , James Lennon , Jamie Kiros , Jan Leike , Jane Park , Jason Kwon , Jason Phang , Jason Teplitz , Jason Wei , Jason Wolfe , Jay Chen , Jeff Harris , Jenia Varavva , Jessica Gan Lee , Jessica Shieh , Ji Lin , Jiahui Yu , Jiayi Weng , Jie Tang , Jieqi Yu , Joanne Jang , Joaquin Quinonero Candela , Joe Beutler , Joe Landers , Joel Parish , Johannes Heidecke , John Schulman , Jonathan Lachman , Jonathan McKay , Jonathan Uesato , Jonathan Ward , Jong Wook Kim , Joost Huizinga , Jordan Sitkin , Jos Kraaijeveld , Josh Gross , Josh Kaplan , Josh Snyder , Joshua Achiam , Joy Jiao , Joyce Lee , Juntang Zhuang , Justyn Harriman , Kai Fricke , Kai Hayashi , Karan Singhal , Katy Shi , Kavin Karthik , Kayla Wood , Kendra Rimbach , Kenny Hsu , Kenny Nguyen , Keren Gu-Lemberg , Kevin Button , Kevin Liu , Kiel Howe , Krithika Muthukumar , Kyle Luther , Lama Ahmad , Larry Kai , Lauren Itow , Lauren Workman , Leher Pathak , Leo Chen , Li Jing , Lia Guy , Liam Fedus , Liang Zhou , Lien Mamitsuka , Lilian Weng , Lindsay McCallum , Lindsey Held , Long Ouyang , Louis Feuvrier , Lu Zhang , Lukas Kondraciuk , Lukasz Kaiser , Luke Hewitt , Luke Metz , Lyric Doshi , Mada Aflak , Maddie Simens , Madelaine Boyd , Madeleine Thompson , Marat Dukhan , Mark Chen , Mark Gray , Mark Hudnall , Marvin Zhang , Marwan Aljubeh , Mateusz Litwin , Matthew Zeng , Max Johnson , Maya Shetty , Mayank Gupta , Meghan Shah , Mehmet Yatbaz , Meng Jia Yang , Mengchao Zhong , Mia Glaese , Mianna Chen , Michael Janner , Michael Lampe , Michael Petrov , Michael Wu , Michele Wang , Michelle Fradin , Michelle Pokrass , Miguel Castro , Miguel Oom Temudo de Castro , Mikhail Pavlov , Miles Brundage , Miles Wang , Minal Khan , Mira Murati , Mo Bavarian , Molly Lin , Murat Yesildal , Nacho Soto , Natalia Gimelshein , Natalie Cone , Natalie Staudacher , Natalie Summers , Natan LaFontaine , Neil Chowdhury , Nick Ryder , Nick Stathas , Nick Turley , Nik Tezak , Niko Felix , Nithanth Kudige , Nitish Keskar , Noah Deutsch , Noel Bundick , Nora Puckett , Ofir Nachum , Ola Okelola , Oleg Boiko , Oleg Murk , Oliver Jaffe , Olivia Watkins , Olivier Godement , Owen Campbell-Moore , Patrick Chao , Paul McMillan , Pavel Belov , Peng Su , Peter Bak , Peter Bakkum , Peter Deng , Peter Dolan , Peter Hoeschele , Peter Welinder , Phil Tillet , Philip Pronin , Philippe Tillet , Prafulla Dhariwal , Qiming Yuan , Rachel Dias , Rachel Lim , Rahul Arora , Rajan Troll , Randall Lin , Rapha Gontijo Lopes , Raul Puri , Reah Miyara , Reimar Leike , Renaud Gaubert , Reza Zamani , Ricky Wang , Rob Donnelly , Rob Honsby , Rocky Smith , Rohan Sahai , Rohit Ramchandani , Romain Huet , Rory Carmichael , Rowan Zellers , Roy Chen , Ruby Chen , Ruslan Nigmatullin , Ryan Cheu , Saachi Jain , Sam Altman , Sam Schoenholz , Sam Toizer , Samuel Miserendino , Sandhini Agarwal , Sara Culver , Scott Ethersmith , Scott Gray , Sean Grove , Sean Metzger , Shamez Hermani , Shantanu Jain , Shengjia Zhao , Sherwin Wu , Shino Jomoto , Shirong Wu , Shuaiqi , Xia , Sonia Phene , Spencer Papay , Srinivas Narayanan , Steve Coffey , Steve Lee , Stewart Hall , Suchir Balaji , Tal Broda , Tal Stramer , Tao Xu , Tarun Gogineni , Taya Christianson , Ted Sanders , Tejal Patwardhan , Thomas Cunninghman , Thomas Degry , Thomas Dimson , Thomas Raoux , Thomas Shadwell , Tianhao Zheng , Todd Underwood , Todor Markov , Toki Sherbakov , Tom Rubin , Tom Stasi , Tomer Kaftan , Tristan Heywood , Troy Peterson , Tyce Walters , Tyna Eloundou , Valerie Qi , Veit Moeller , Vinnie Monaco , Vishal Kuo , Vlad Fomenko , Wayne Chang , Weiyi Zheng , Wenda Zhou , Wesam Manassra , Will Sheu , Wojciech Zaremba , Yash Patil , Yilei Qian , Yongjik Kim , Youlong Cheng , Yu Zhang , Yuchen He , Yuchen Zhang , Yujia Jin , Yunxing Dai , Yury Malkov

Recent advancements in text-to-speech (TTS) powered by language models have showcased remarkable capabilities in achieving naturalness and zero-shot voice cloning. Notably, the decoder-only transformer is the prominent architecture in this…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-12 Théodor Lemerle , Nicolas Obin , Axel Roebel

While large-scale omni-models have demonstrated impressive capabilities across various modalities, their strong performance heavily relies on massive multimodal data and incurs substantial computational costs. This work introduces…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-11 Dehua Tao , Xuan Luo , Daxin Tan , Kai Chen , Lanqing Hong , Jing Li , Ruifeng Xu , Xiao Chen

OpenDiLoCo is an open-source implementation and replication of the Distributed Low-Communication (DiLoCo) training method for large language models. We provide a reproducible implementation of the DiLoCo experiments, offering it within a…

Machine Learning · Computer Science 2024-07-11 Sami Jaghouar , Jack Min Ong , Johannes Hagemann

We introduce InteractiveOmni, a unified and open-source omni-modal large language model for audio-visual multi-turn interaction, ranging from 4B to 8B parameters, designed to lead the field of lightweight models by offering comprehensive…

Spontaneous conversations in real-world settings such as those found in child-centered recordings have been shown to be amongst the most challenging audio files to process. Nevertheless, building speech processing models handling such a…

Audio and Speech Processing · Electrical Eng. & Systems 2025-03-12 Marvin Lavechin , Ruben Bousbib , Hervé Bredin , Emmanuel Dupoux , Alejandrina Cristia
‹ Prev 1 2 3 10 Next ›