English
Related papers

Related papers: Audio Flamingo 3: Advancing Audio Intelligence wit…

200 papers

While state-of-the-art audio-video generation models like Veo3 and Sora2 demonstrate remarkable capabilities, their closed-source nature makes their architectures and training paradigms inaccessible. To bridge this gap in accessibility and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Hebeizi Li , Zihao Liang , Benyuan Sun , Zihao Yin , Xiao Sha , Chenliang Wang , Yi Yang

The ability to comprehend audio--which includes speech, non-speech sounds, and music--is crucial for AI agents to interact effectively with the world. We present MMAU, a novel benchmark designed to evaluate multimodal audio understanding…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-28 S Sakshi , Utkarsh Tyagi , Sonal Kumar , Ashish Seth , Ramaneswaran Selvakumar , Oriol Nieto , Ramani Duraiswami , Sreyan Ghosh , Dinesh Manocha

Recently, Conformer has achieved state-of-the-art performance in many speech recognition tasks. However, the Transformer-based models show significant deterioration for long-form speech, such as lectures, because the self-attention…

Sound · Computer Science 2024-10-08 Tomoki Honda , Shinsuke Sakai , Tatsuya Kawahara

The immense computational cost of training Large Language Models (LLMs) presents a major barrier to innovation. While FP8 training offers a promising solution with significant theoretical efficiency gains, its widespread adoption has been…

Computation and Language · Computer Science 2025-10-20 Wenjun Wang , Shuo Cai , Congkai Xie , Mingfa Feng , Yiming Zhang , Zhen Li , Kejing Yang , Ming Li , Jiannong Cao , Hongxia Yang

With advancements in large audio-language models (LALMs), which enhance large language models (LLMs) with auditory capabilities, these models are expected to demonstrate universal proficiency across various auditory tasks. While numerous…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-28 Chih-Kai Yang , Neo S. Ho , Hung-yi Lee

We introduce OpenFlamingo, a family of autoregressive vision-language models ranging from 3B to 9B parameters. OpenFlamingo is an ongoing effort to produce an open-source replication of DeepMind's Flamingo models. On seven vision-language…

Large audio language models (ALMs) extend LLMs with auditory understanding. A common approach freezes the LLM and trains only an adapter on self-generated targets. However, this fails for reasoning LLMs (RLMs) whose built-in…

Computation and Language · Computer Science 2026-03-11 Petr Grinberg , Hassan Shahmohammadi

While large audio language models excel at tasks like ASR and emotion recognition, they still struggle with complex reasoning due to the modality gap between audio and text as well as the lack of structured intermediate supervision. To…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-24 Runyan Yang , Yuke Si , Yingying Gao , Junlan Feng , Chao Deng , Shilei Zhang

Using self-supervised learning (SSL) models has significantly improved performance for downstream speech tasks, surpassing the capabilities of traditional hand-crafted features. This study investigates the amalgamation of SSL models, with…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-27 Szu-Jui Chen , John H. L. Hansen

We introduce Baichuan-Audio, an end-to-end audio large language model that seamlessly integrates audio understanding and generation. It features a text-guided aligned speech generation mechanism, enabling real-time speech interaction with…

Computation and Language · Computer Science 2025-02-25 Tianpeng Li , Jun Liu , Tao Zhang , Yuanbo Fang , Da Pan , Mingrui Wang , Zheng Liang , Zehuan Li , Mingan Lin , Guosheng Dong , Jianhua Xu , Haoze Sun , Zenan Zhou , Weipeng Chen

Large Language Models, such as Generative Pre-trained Transformer 3 (aka. GPT-3), have been developed to understand language through the analysis of extensive text data, allowing them to identify patterns and connections between words.…

Computation and Language · Computer Science 2023-10-03 Baphumelele Masikisiki , Vukosi Marivate , Yvette Hlope

Recently, instruction-following audio-language models have received broad attention for audio interaction with humans. However, the absence of pre-trained audio models capable of handling diverse audio types and tasks has hindered progress…

Audio and Speech Processing · Electrical Eng. & Systems 2023-12-22 Yunfei Chu , Jin Xu , Xiaohuan Zhou , Qian Yang , Shiliang Zhang , Zhijie Yan , Chang Zhou , Jingren Zhou

Recent developments in unsupervised representation learning have successfully established the concept of transfer learning in NLP. Mainly three forces are driving the improvements in this area of research: More elaborated architectures are…

Computation and Language · Computer Science 2020-07-22 Matthias Aßenmacher , Christian Heumann

We introduce the Nemotron 3 family of models - Nano, Super, and Ultra. These models deliver strong agentic, reasoning, and conversational capabilities. The Nemotron 3 family uses a Mixture-of-Experts hybrid Mamba-Transformer architecture to…

Computation and Language · Computer Science 2025-12-25 NVIDIA , : , Aaron Blakeman , Aaron Grattafiori , Aarti Basant , Abhibha Gupta , Abhinav Khattar , Adi Renduchintala , Aditya Vavre , Akanksha Shukla , Akhiad Bercovich , Aleksander Ficek , Aleksandr Shaposhnikov , Alex Kondratenko , Alexander Bukharin , Alexandre Milesi , Ali Taghibakhshi , Alisa Liu , Amelia Barton , Ameya Sunil Mahabaleshwarkar , Amir Klein , Amit Zuker , Amnon Geifman , Amy Shen , Anahita Bhiwandiwalla , Andrew Tao , Anjulie Agrusa , Ankur Verma , Ann Guan , Anubhav Mandarwal , Arham Mehta , Ashwath Aithal , Ashwin Poojary , Asif Ahamed , Asit Mishra , Asma Kuriparambil Thekkumpate , Ayush Dattagupta , Banghua Zhu , Bardiya Sadeghi , Barnaby Simkin , Ben Lanir , Benedikt Schifferer , Besmira Nushi , Bilal Kartal , Bita Darvish Rouhani , Boris Ginsburg , Brandon Norick , Brandon Soubasis , Branislav Kisacanin , Brian Yu , Bryan Catanzaro , Carlo del Mundo , Chantal Hwang , Charles Wang , Cheng-Ping Hsieh , Chenghao Zhang , Chenhan Yu , Chetan Mungekar , Chintan Patel , Chris Alexiuk , Christopher Parisien , Collin Neale , Cyril Meurillon , Damon Mosk-Aoyama , Dan Su , Dane Corneil , Daniel Afrimi , Daniel Lo , Daniel Rohrer , Daniel Serebrenik , Daria Gitman , Daria Levy , Darko Stosic , David Mosallanezhad , Deepak Narayanan , Dhruv Nathawani , Dima Rekesh , Dina Yared , Divyanshu Kakwani , Dong Ahn , Duncan Riach , Dusan Stosic , Edgar Minasyan , Edward Lin , Eileen Long , Eileen Peters Long , Elad Segal , Elena Lantz , Ellie Evans , Elliott Ning , Eric Chung , Eric Harper , Eric Tramel , Erick Galinkin , Erik Pounds , Evan Briones , Evelina Bakhturina , Evgeny Tsykunov , Faisal Ladhak , Fay Wang , Fei Jia , Felipe Soares , Feng Chen , Ferenc Galko , Frank Sun , Frankie Siino , Gal Hubara Agam , Ganesh Ajjanagadde , Gantavya Bhatt , Gargi Prasad , George Armstrong , Gerald Shen , Gorkem Batmaz , Grigor Nalbandyan , Haifeng Qian , Harsh Sharma , Hayley Ross , Helen Ngo , Herbert Hum , Herman Sahota , Hexin Wang , Himanshu Soni , Hiren Upadhyay , Huizi Mao , Huy C Nguyen , Huy Q Nguyen , Iain Cunningham , Ido Galil , Ido Shahaf , Igor Gitman , Ilya Loshchilov , Itamar Schen , Itay Levy , Ivan Moshkov , Izik Golan , Izzy Putterman , Jan Kautz , Jane Polak Scowcroft , Jared Casper , Jatin Mitra , Jeffrey Glick , Jenny Chen , Jesse Oliver , Jian Zhang , Jiaqi Zeng , Jie Lou , Jimmy Zhang , Jinhang Choi , Jining Huang , Joey Conway , Joey Guman , John Kamalu , Johnny Greco , Jonathan Cohen , Joseph Jennings , Joyjit Daw , Julien Veron Vialard , Junkeun Yi , Jupinder Parmar , Kai Xu , Kan Zhu , Kari Briski , Katherine Cheung , Katherine Luna , Keith Wyss , Keshav Santhanam , Kevin Shih , Kezhi Kong , Khushi Bhardwaj , Kirthi Shankar , Krishna C. Puvvada , Krzysztof Pawelec , Kumar Anik , Lawrence McAfee , Laya Sleiman , Leon Derczynski , Li Ding , Lizzie Wei , Lucas Liebenwein , Luis Vega , Maanu Grover , Maarten Van Segbroeck , Maer Rodrigues de Melo , Mahdi Nazemi , Makesh Narsimhan Sreedhar , Manoj Kilaru , Maor Ashkenazi , Marc Romeijn , Marcin Chochowski , Mark Cai , Markus Kliegl , Maryam Moosaei , Matt Kulka , Matvei Novikov , Mehrzad Samadi , Melissa Corpuz , Mengru Wang , Meredith Price , Michael Andersch , Michael Boone , Michael Evans , Miguel Martinez , Mikail Khona , Mike Chrzanowski , Minseok Lee , Mohammad Dabbah , Mohammad Shoeybi , Mostofa Patwary , Nabin Mulepati , Najeeb Nabwani , Natalie Hereth , Nave Assaf , Negar Habibi , Neta Zmora , Netanel Haber , Nicola Sessions , Nidhi Bhatia , Nikhil Jukar , Nikki Pope , Nikolai Ludwig , Nima Tajbakhsh , Nir Ailon , Nirmal Juluru , Nishant Sharma , Oleksii Hrinchuk , Oleksii Kuchaiev , Olivier Delalleau , Oluwatobi Olabiyi , Omer Ullman Argov , Omri Puny , Oren Tropp , Ouye Xie , Parth Chadha , Pasha Shamis , Paul Gibbons , Pavlo Molchanov , Pawel Morkisz , Peter Dykas , Peter Jin , Pinky Xu , Piotr Januszewski , Pranav Prashant Thombre , Prasoon Varshney , Pritam Gundecha , Przemek Tredak , Qing Miao , Qiyu Wan , Rabeeh Karimi Mahabadi , Rachit Garg , Ran El-Yaniv , Ran Zilberstein , Rasoul Shafipour , Rich Harang , Rick Izzo , Rima Shahbazyan , Rishabh Garg , Ritika Borkar , Ritu Gala , Riyad Islam , Robert Hesse , Roger Waleffe , Rohit Watve , Roi Koren , Ruoxi Zhang , Russell Hewett , Russell J. Hewett , Ryan Prenger , Ryan Timbrook , Sadegh Mahdavi , Sahil Modi , Samuel Kriman , Sangkug Lim , Sanjay Kariyappa , Sanjeev Satheesh , Saori Kaji , Satish Pasumarthi , Saurav Muralidharan , Sean Narentharen , Sean Narenthiran , Seonmyeong Bak , Sergey Kashirsky , Seth Poulos , Shahar Mor , Shanmugam Ramasamy , Shantanu Acharya , Shaona Ghosh , Sharath Turuvekere Sreenivas , Shelby Thomas , Shiqing Fan , Shreya Gopal , Shrimai Prabhumoye , Shubham Pachori , Shubham Toshniwal , Shuoyang Ding , Siddharth Singh , Simeng Sun , Smita Ithape , Somshubra Majumdar , Soumye Singhal , Stas Sergienko , Stefania Alborghetti , Stephen Ge , Sugam Dipak Devare , Sumeet Kumar Barua , Suseella Panguluri , Suyog Gupta , Sweta Priyadarshi , Syeda Nahida Akter , Tan Bui , Teodor-Dumitru Ene , Terry Kong , Thanh Do , Tijmen Blankevoort , Tim Moon , Tom Balough , Tomer Asida , Tomer Bar Natan , Tomer Ronen , Tugrul Konuk , Twinkle Vashishth , Udi Karpas , Ushnish De , Vahid Noorozi , Vahid Noroozi , Venkat Srinivasan , Venmugil Elango , Victor Cui , Vijay Korthikanti , Vinay Rao , Vitaly Kurin , Vitaly Lavrukhin , Vladimir Anisimov , Wanli Jiang , Wasi Uddin Ahmad , Wei Du , Wei Ping , Wenfei Zhou , Will Jennings , William Zhang , Wojciech Prazuch , Xiaowei Ren , Yashaswi Karnati , Yejin Choi , Yev Meyer , Yi-Fu Wu , Yian Zhang , Yigong Qin , Ying Lin , Yonatan Geifman , Yonggan Fu , Yoshi Subara , Yoshi Suhara , Yubo Gao , Zach Moshe , Zhen Dong , Zhongbo Zhu , Zihan Liu , Zijia Chen , Zijie Yan

While large audio-language models (LALMs) have demonstrated state-of-the-art audio understanding, their reasoning capability in complex soundscapes still falls behind large vision-language models (LVLMs). Compared to the visual domain, one…

Sound · Computer Science 2025-09-22 Qiaolin Wang , Xilin Jiang , Linyang He , Junkai Wu , Nima Mesgarani

Large Audio-Language Models (LALMs) have demonstrated remarkable performance in tasks involving audio perception and understanding, such as speech recognition and audio captioning. However, their reasoning capabilities - critical for…

Sound · Computer Science 2025-01-14 Ziyang Ma , Zhuo Chen , Yuping Wang , Eng Siong Chng , Xie Chen

The recent surge in open-source Multimodal Large Language Models (MLLM) frameworks, such as LLaVA, provides a convenient kickoff for artificial intelligence developers and researchers. However, most of the MLLM frameworks take vision as the…

Recent Audio Large Language Models (AudioLLMs) exhibit a striking performance inversion: while excelling at complex reasoning tasks, they consistently underperform on fine-grained acoustic perception. We attribute this gap to a fundamental…

Computation and Language · Computer Science 2026-04-15 Linhao Zhang , Yuhan Song , Aiwei Liu , Chuhan Wu , Sijun Zhang , Wei Jia , Yuan Liu , Houfeng Wang , Xiao Zhou
‹ Prev 1 4 5 6 7 8 10 Next ›