English
Related papers

Related papers: Audio Flamingo 3: Advancing Audio Intelligence wit…

200 papers

Pre-training speech models on large volumes of data has achieved remarkable success. OpenAI Whisper is a multilingual multitask model trained on 680k hours of supervised speech data. It generalizes well to various speech recognition and…

Recent advancements in large audio language models (LALMs) have demonstrated impressive results and promising prospects in universal understanding and reasoning across speech, music, and general sound. However, these models still lack the…

Multimodal Large Language Models (LLMs) claim "musical understanding" via evaluations that conflate listening with score reading. We benchmark three SOTA LLMs (Gemini 2.5 Pro, Gemini 2.5 Flash, and Qwen2.5-Omni) across three core music…

Sound · Computer Science 2025-10-28 Brandon James Carone , Iran R. Roman , Pablo Ripollés

Although Large Audio-Language Models (LALMs) deliver state-of-the-art (SOTA) performance, they frequently suffer from hallucinations, e.g. generating text not grounded in the audio input. We analyze these grounding failures and identify a…

We introduce LongCat-Flash-Thinking-2601, a 560-billion-parameter open-source Mixture-of-Experts (MoE) reasoning model with superior agentic reasoning capability. LongCat-Flash-Thinking-2601 achieves state-of-the-art performance among…

Artificial Intelligence · Computer Science 2026-02-03 Meituan LongCat Team , Anchun Gui , Bei Li , Bingyang Tao , Bole Zhou , Borun Chen , Chao Zhang , Chao Zhang , Chen Gao , Chen Zhang , Chengcheng Han , Chenhui Yang , Chuyu Zhang , Cong Chen , Cunguang Wang , Daoru Pan , Defei Bu , Dengchang Zhao , Di Xiu , Dishan Liu , Dongyu Ru , Dunwei Tu , Fan Wu , Fengcheng Yuan , Fengcun Li , Gang Xu , Guanyu Wu , Guoyuan Lin , Haibin Wang , Hansi Yang , Hao Yang , Haonan Yan , Haoxiang Ma , Haoxing Wen , Hongyan Hao , Hongyin Tang , Hongyu Zang , Hongzhi Ni , Hui Su , Jiacheng Zhang , Jiahong Zhou , Jiahuan Li , Jiaming Wang , Jian Yang , Jianfei Zhang , Jianhao Xu , Jianing Wang , Jiapeng Zhu , Jiaqi Sun , Jiarong Shi , Jiarui Zhao , Jingang Wang , Jinluan Yang , Jinrui Ding , Jinwei Xiao , Jiyuan He , Juncan Xu , Kefeng Zhang , Keheng Wang , Li Wei , Lianhui Ma , Lin Qiu , Lingbing Kong , Lingchuan Liu , Linsen Guo , Mengshen Zhu , Mengxia Shen , Mingyang Zhu , Peiguang Li , Peng Pei , Peng Zhao , Pengcheng Jia , Pengtao Zhang , Ping Liu , Qi Gu , Qiong Huang , Qiyuan Duan , Quanchi Weng , Rongxiang Weng , Rongzhi Zhang , Rumei Li , Shanglin Lei , Shengnan An , Shijun Dai , Shizhe Wu , Shuaikang Liu , Shuang Zhou , Shuo Wang , Songyuan Zhao , Tao Liang , Tianhao Hu , Tianze Chen , Wei Liu , Wei Shi , Wei Wang , Weifeng Tang , Wenjie Shi , Wenlong Zhu , Wentao Chen , Wentao Shi , Xi Su , Xiandi Ma , Xiangcheng Liu , Xiangyu Xi , Xiangyuan Liu , Xiangzhou Huang , Xiao Liu , Xiaodong Cai , Xiaolong Chen , Xiaowei Shi , Xiaoyu Li , Xin Chen , Xingchen Liu , Xuan Huang , Xuezhi Cao , Xunliang Cai , Yan Chen , Yang Bai , Yang Liu , Yang Yang , Yang Zheng , Yanyu Chen , Yaoming Wang , Yaoming Zhu , Yaorui Shi , Yaqi Huo , Yerui Sun , Yi Zhang , Yi-Kai Zhang , Yifan Lu , Yifan Zhao , Yihao Chen , Yitao Zhai , Yongjing Yin , Yongwei Zhou , Youshao Xiao , Yu Wang , Yu Yang , Yuchen Xie , Yuchen Yu , Yuchuan Dai , Yue Xu , Yueqing Sun , Yufei Zhang , Yuhuai Wei , Yulei Qian , Yunfan Liang , Yunke Zhao , Yuwei Jiang , Yuxin Bian , Yuxin Chen , Yuxin Liu , Zeyang Yu , Zhao Yang , Zhengsheng Huang , Zhengyu Chen , Zhijian Liu , Zhikang Xia , Zhimin Lin , Zhiyuan Yao , Zhuofan Chen , Zhuowen Han , Zijian Zhang , Ziran Li , Ziwen Wang , Ziyuan Zhuang

Recent advances in text-to-audio (TTA) generation excel at synthesizing short audio clips but struggle with long-form narrative audio, which requires temporal coherence and compositional reasoning. To address this gap, we propose…

Computer Vision and Pattern Recognition · Computer Science 2025-10-06 Yuxin Guo , Teng Wang , Yuying Ge , Shijie Ma , Yixiao Ge , Wei Zou , Ying Shan

We investigate the emergent abilities of the recently proposed web-scale speech model Whisper, by adapting it to unseen tasks with prompt engineering. We selected three tasks: audio-visual speech recognition (AVSR), code-switched speech…

Audio and Speech Processing · Electrical Eng. & Systems 2023-08-17 Puyuan Peng , Brian Yan , Shinji Watanabe , David Harwath

Speech data has rich acoustic and paralinguistic information with important cues for understanding a speaker's tone, emotion, and intent, yet traditional large language models such as BERT do not incorporate this information. There has been…

Computation and Language · Computer Science 2023-11-14 Fatema Hasan , Yulong Li , James Foulds , Shimei Pan , Bishwaranjan Bhattacharjee

Large audio-language models (LALMs) have achieved near-human performance in sentence-level transcription and emotion recognition. However, existing evaluations focus mainly on surface-level perception, leaving the capacity of models for…

Computation and Language · Computer Science 2025-08-05 Wanqi Yang , Yanda Li , Yunchao Wei , Meng Fang , Ling Chen

We introduce MMAR, a new benchmark designed to evaluate the deep reasoning capabilities of Audio-Language Models (ALMs) across massive multi-disciplinary tasks. MMAR comprises 1,000 meticulously curated audio-question-answer triplets,…

Audio super-resolution is challenging owing to its ill-posed nature. Recently, the application of diffusion models in audio super-resolution has shown promising results in alleviating this challenge. However, diffusion-based models have…

Audio and Speech Processing · Electrical Eng. & Systems 2025-03-12 Jun-Hak Yun , Seung-Bin Kim , Seong-Whan Lee

The SAFE Challenge evaluates synthetic speech detection across three tasks: unmodified audio, processed audio with compression artifacts, and laundered audio designed to evade detection. We systematically explore self-supervised learning…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-08 Hashim Ali , Surya Subramani , Lekha Bollinani , Nithin Sai Adupa , Sali El-Loh , Hafiz Malik

The goal of this work is to enhance balanced multimodal understanding in audio-visual large language models (AV-LLMs) by addressing modality bias without additional training. In current AV-LLMs, audio and video features are typically…

Computer Vision and Pattern Recognition · Computer Science 2025-10-01 Chaeyoung Jung , Youngjoon Jang , Jongmin Choi , Joon Son Chung

Text-to-audio (TTA) generation with fine-grained control signals, e.g., precise timing control or intelligible speech content, has been explored in recent works. However, constrained by data scarcity, their generation performance at scale…

Sound · Computer Science 2026-04-21 Yuxuan Jiang , Zehua Chen , Zeqian Ju , Yusheng Dai , Weibei Dou , Jun Zhu

Open AI's language model, GPT-3, has shown great potential for many NLP tasks, with applications in many different domains. In this work we carry out a first study on GPT-3's capability to communicate musical decisions through textual…

Computation and Language · Computer Science 2022-06-17 Stephen James Krol , Maria Teresa Llano , Jon McCormack

Complex Reasoning in Large Language Models can be dynamically optimized using Test-Time Scaling (TTS) to mitigate Overthinking. Methods such as Coconut, SoftCoT and its variant are effective in continuous latent space inference, the core…

Artificial Intelligence · Computer Science 2025-12-17 Jiaqi Wang , Binquan Ji , Haibo Luo , Yiyang Qi , Ruiting Li , Huiyan Wang , Yuantao Han , Cangyi Yang , jiaxu Zhang , Feiliang Ren

Current audio-visual (AV) benchmarks focus on final answer accuracy, overlooking the underlying reasoning process. This makes it difficult to distinguish genuine comprehension from correct answers derived through flawed reasoning or…

Computer Vision and Pattern Recognition · Computer Science 2025-08-22 Siminfar Samakoush Galougah , Rishie Raj , Sanjoy Chowdhury , Sayan Nag , Ramani Duraiswami

Multi-source multi-hop question answering (QA) represents a challenging task in natural language processing due to the need for dynamic integration of heterogeneous knowledge sources and multi-step reasoning. Existing methods often suffer…

Computation and Language · Computer Science 2025-02-11 Jackson Coleman , Isaiah Lawrence , Benjamin Turner

We introduce Nemotron 3 Nano Omni, the latest model in the Nemotron multimodal series and the first to natively support audio inputs alongside text, images, and video. Nemotron 3 Nano Omni delivers consistent accuracy improvements over its…

Machine Learning · Computer Science 2026-05-12 NVIDIA , : , Amala Sanjay Deshmukh , Kateryna Chumachenko , Tuomas Rintamaki , Matthieu Le , Tyler Poon , Danial Mohseni Taheri , Ilia Karmanov , Guilin Liu , Jarno Seppanen , Arushi Goel , Mike Ranzinger , Greg Heinrich , Guo Chen , Lukas Voegtle , Philipp Fischer , Timo Roman , Karan Sapra , Collin McCarthy , Shaokun Zhang , Fuxiao Liu , Hanrong Ye , Yi Dong , Mingjie Liu , Yifan Peng , Piotr Zelasko , Zhehuai Chen , Nithin Rao Koluguri , Nune Tadevosyan , Lilit Grigoryan , Ehsan Hosseini Asl , Pritam Biswas , Leili Tavabi , Yuanhang Su , Zhiding Yu , Peter Jin , Alexandre Milesi , Netanel Haber , Yao Xu , Sarah Amiraslani , Nabin Mulepati , Eric Tramel , Jaehun Jung , Ximing Lu , Brandon Cui , Jin Xu , Zhiqi Li , Shihao Wang , Yuanguo Kuang , Shaokun Zhang , Huck Yang , Boyi Li , Hongxu Yin , Song Han , Bilal Kartal , Pavlo Molchanov , Adi Renduchintala , Charles Wang , David Mosallanezhad , Soumye Singhal , Luis Vega , Katherine Cheung , Sreyan Ghosh , Yian Zhang , Alexander Bukharin , Venkat Srinivasan , Johnny Greco , Andre Manoel , Maarten Van Segbroeck , Suseella Panguliri , Rohit Watve , Divyanshu Kakwani , Shubham Pachori , Jeffrey Glick , Radha Sri-Tharan , Aileen Zaman , Khanh Nguyen , Shi Chen , Jiaheng Fang , Qing Miao , Wenfei Zhou , Yu Wang , Zaid Pervaiz Bhat , Varun Praveen , Arihant Jain , Ramanathan Arunachalam , Tomasz Kornuta , Ashton Sharabiani , Amy Shen , Wei Huang , Yi-Fu Wu , Ali Roshan Ghias , Huiying Li , Brian Yu , Nima Tajbakhsh , Chen Cui , Wenwen Gao , Li Ding , Terry Kong , Manoj Kilaru , Anahita Bhiwandiwalla , Marek Wawrzos , Daniel Korzekwa , Pablo Ribalta , Grzegorz Chlebus , Besmira Nushi , Ewa Dobrowolska , Maciej Jakub Mikulski , Kunal Dhawan , Steve Huang , Jagadeesh Balam , Yongqiang Wang , Nikolay Karpov , Valentin Mendelev , George Zelenfroynd , Meline Mkrtchyan , Qing Miao , Omri Almog , Bhavesh Pawar , Rameshwar Shivbhakta , Sudeep Sabnis , Ashrton Sharabiani , Negar Habibi , Geethapriya Venkataramani , Pamela Peng , Prerit Rodney , Serge Panev , Richard Mazzarese , Nicky Liu , Michael Fukuyama , Andrii Skliar , Roger Waleffe , Duncan Riach , Yunheng Zou , Jian Hu , Hao Zhang , Binfeng Xu , Yuhao Yang , Zuhair Ahmed , Alexandre Milesi , Carlo del Mundo , Chad Voegele , Zhiyu Cheng , Nave Assaf , Andrii Skliar , Daniel Afrimi , Natan Bagrov , Ran Zilberstein , Ofri Masad , Eugene Khvedchenia , Natan Bagrov , Borys Tymchenko , Tomer Asida , Daniel Afrimi , Parth Mannan , Victor Cui , Michael Evans , Katherine Luna , Jie Lou , Pinky Xu , Guyue Huang , Negar Habibi , Michael Boone , Pradeep Thalasta , Adeola Adesoba , Dina Yared , Christopher Parisien , Leon Derczynski , Shaona Ghosh , Wes Feely , Micah Schaffer , Radha Sri-Tharan , Jeffrey Glick , Barnaby Simkin , George Zelenfroynd , Tomasz Grzegorzek , Rishabh Garg , Aastha Jhunjhunwala , Sergei Kolchenko , Farzan Memarian , Haran Kumar , Shiv Kumar , Isabel Hulseman , Anjali Shah , Kari Briski , Padmavathy Subramanian , Joey Conway , Udi Karpas , Jane Polak Scowcroft , Annie Surla , Shilpa Ammireddy , Ellie Evans , Jesse Oliver , Tom Balough , Chia-Chih Chen , Sandip Bhaskar , Alejandra Rico , Bardiya Sadeghi , Seph Mard , Katherine Cheung , Meredith Price , Laya Sleiman , Saori Kaji , Wesley Helmholz , Wendy Quan , Michael Lightstone , Jonathan Cohen , Jian Zhang , Oleksii Kuchaiev , Boris Ginsburg , Jan Kautz , Eileen Long , Mohammad Shoeybi , Mostofa Patwary , Oluwatobi Olabiyi , Andrew Tao , Bryan Catanzaro , Udi Karpas

Recently, reinforcement learning (RL) has been shown to greatly enhance the reasoning capabilities of large language models (LLMs), and RL-based approaches have been progressively applied to visual multimodal tasks. However, the audio…

Sound · Computer Science 2025-05-15 Gang Li , Jizhong Liu , Heinrich Dinkel , Yadong Niu , Junbo Zhang , Jian Luan