English
Related papers

Related papers: Audio Flamingo 3: Advancing Audio Intelligence wit…

200 papers

Large audio language models (LALMs) are a class of foundation models for audio understanding. Existing LALMs tend to degrade significantly in real-world noisy acoustic conditions where speech and non-speech sounds interfere. While…

Sound · Computer Science 2026-05-26 Han Yin , Yang Xiao , Younghoo Kwon , Ting Dang , Jung-Woo Choi

This survey paper provides a comprehensive overview of the recent advancements and challenges in applying large language models to the field of audio signal processing. Audio processing, with its diverse signal representations and a wide…

Negotiation is a fundamental challenge for AI agents, as it requires an ability to reason strategically, model opponents, and balance cooperation with competition. We present the first comprehensive study that systematically evaluates how…

Computation and Language · Computer Science 2026-01-12 Sherzod Hakimov , Roland Bernard , Tim Leiber , Karl Osswald , Kristina Richert , Ruilin Yang , Raffaella Bernardi , David Schlangen

Joint audio-visual reasoning is essential for omnimodal understanding, yet current multimodal large language models (MLLMs) still struggle when reasoning requires fine-grained evidence from both modalities. A central limitation is that…

Current speech encoding pipelines often rely on an additional text-based LM to get robust representations of human communication, even though SotA speech-to-text models often have a LM within. This work proposes an approach to improve the…

With the rapid progress of large audio-language models (LALMs), audio question answering (AQA) has emerged as a challenging task requiring both fine-grained audio understanding and complex reasoning. While current methods mainly rely on…

Sound · Computer Science 2025-09-19 Jinghua Zhao , Hang Su , Lichun Fan , Zhenbo Luo , Hui Wang , Haoqin Sun , Yong Qin

Although audio generation shares commonalities across different types of audio, such as speech, music, and sound effects, designing models for each type requires careful consideration of specific objectives and biases that can significantly…

Large language model (LLM) agents have demonstrated remarkable capabilities in tool use, reasoning, and code generation, yet single-agent systems exhibit fundamental limitations when confronted with complex research tasks demanding…

Artificial Intelligence · Computer Science 2026-03-17 Aaron Shen , Alfred Shen

We introduce M3-Agent, a novel multimodal agent framework equipped with long-term memory. Like humans, M3-Agent can process real-time visual and auditory inputs to build and update episodic and semantic memories, gradually accumulating…

Computer Vision and Pattern Recognition · Computer Science 2025-10-10 Lin Long , Yichen He , Wentao Ye , Yiyuan Pan , Yuan Lin , Hang Li , Junbo Zhao , Wei Li

We introduce the Llama-Nemotron series of models, an open family of heterogeneous reasoning models that deliver exceptional reasoning capabilities, inference efficiency, and an open license for enterprise use. The family comes in three…

Computation and Language · Computer Science 2025-09-10 Akhiad Bercovich , Itay Levy , Izik Golan , Mohammad Dabbah , Ran El-Yaniv , Omri Puny , Ido Galil , Zach Moshe , Tomer Ronen , Najeeb Nabwani , Ido Shahaf , Oren Tropp , Ehud Karpas , Ran Zilberstein , Jiaqi Zeng , Soumye Singhal , Alexander Bukharin , Yian Zhang , Tugrul Konuk , Gerald Shen , Ameya Sunil Mahabaleshwarkar , Bilal Kartal , Yoshi Suhara , Olivier Delalleau , Zijia Chen , Zhilin Wang , David Mosallanezhad , Adi Renduchintala , Haifeng Qian , Dima Rekesh , Fei Jia , Somshubra Majumdar , Vahid Noroozi , Wasi Uddin Ahmad , Sean Narenthiran , Aleksander Ficek , Mehrzad Samadi , Jocelyn Huang , Siddhartha Jain , Igor Gitman , Ivan Moshkov , Wei Du , Shubham Toshniwal , George Armstrong , Branislav Kisacanin , Matvei Novikov , Daria Gitman , Evelina Bakhturina , Prasoon Varshney , Makesh Narsimhan , Jane Polak Scowcroft , John Kamalu , Dan Su , Kezhi Kong , Markus Kliegl , Rabeeh Karimi Mahabadi , Ying Lin , Sanjeev Satheesh , Jupinder Parmar , Pritam Gundecha , Brandon Norick , Joseph Jennings , Shrimai Prabhumoye , Syeda Nahida Akter , Mostofa Patwary , Abhinav Khattar , Deepak Narayanan , Roger Waleffe , Jimmy Zhang , Bor-Yiing Su , Guyue Huang , Terry Kong , Parth Chadha , Sahil Jain , Christine Harvey , Elad Segal , Jining Huang , Sergey Kashirsky , Robert McQueen , Izzy Putterman , George Lam , Arun Venkatesan , Sherry Wu , Vinh Nguyen , Manoj Kilaru , Andrew Wang , Anna Warno , Abhilash Somasamudramath , Sandip Bhaskar , Maka Dong , Nave Assaf , Shahar Mor , Omer Ullman Argov , Scot Junkin , Oleksandr Romanenko , Pedro Larroy , Monika Katariya , Marco Rovinelli , Viji Balas , Nicholas Edelman , Anahita Bhiwandiwalla , Muthu Subramaniam , Smita Ithape , Karthik Ramamoorthy , Yuting Wu , Suguna Varshini Velury , Omri Almog , Joyjit Daw , Denys Fridman , Erick Galinkin , Michael Evans , Shaona Ghosh , Katherine Luna , Leon Derczynski , Nikki Pope , Eileen Long , Seth Schneider , Guillermo Siman , Tomasz Grzegorzek , Pablo Ribalta , Monika Katariya , Chris Alexiuk , Joey Conway , Trisha Saar , Ann Guan , Krzysztof Pawelec , Shyamala Prayaga , Oleksii Kuchaiev , Boris Ginsburg , Oluwatobi Olabiyi , Kari Briski , Jonathan Cohen , Bryan Catanzaro , Jonah Alben , Yonatan Geifman , Eric Chung

Audio Large Language Models (AudioLLMs) have received widespread attention and have significantly improved performance on audio tasks such as conversation, audio understanding, and automatic speech recognition (ASR). Despite these…

Computational Engineering, Finance, and Science · Computer Science 2025-12-19 Yupeng Cao , Haohang Li , Yangyang Yu , Shashidhar Reddy Javaji , Yueru He , Jimin Huang , Qianqian Xie , Fabrizio Dimino , Xiao-yang Liu , K. P. Subbalakshmi , Meikang Qiu , Sophia Ananiadou , Jian-Yun Nie

We present Deep Voice 3, a fully-convolutional attention-based neural text-to-speech (TTS) system. Deep Voice 3 matches state-of-the-art neural speech synthesis systems in naturalness while training ten times faster. We scale Deep Voice 3…

We introduce a method to improve the zero-shot reasoning abilities of large language models on general language understanding tasks. Specifically, we build an autonomous agent to instruct the reasoning process of large language models. We…

Computation and Language · Computer Science 2024-08-15 Nicholas Crispino , Kyle Montgomery , Fankun Zeng , Dawn Song , Chenguang Wang

We introduce ImmerseDiffusion, an end-to-end generative audio model that produces 3D immersive soundscapes conditioned on the spatial, temporal, and environmental conditions of sound objects. ImmerseDiffusion is trained to generate…

Sound · Computer Science 2025-02-11 Mojtaba Heydari , Mehrez Souden , Bruno Conejo , Joshua Atkins

The auditory system plays a substantial role in shaping the overall human perceptual experience. While prevailing large language models (LLMs) and visual language models (VLMs) have shown their promise in solving a wide variety of language…

Audio and Speech Processing · Electrical Eng. & Systems 2025-02-19 Jinhua Liang , Xubo Liu , Wenwu Wang , Mark D. Plumbley , Huy Phan , Emmanouil Benetos

Speech large language models (speech-LLMs) integrate speech and text-based foundation models to provide a unified framework for handling a wide range of downstream tasks. In this paper, we introduce WHISMA, a speech-LLM tailored for spoken…

Audio and Speech Processing · Electrical Eng. & Systems 2024-08-30 Mohan Li , Cong-Thanh Do , Simon Keizer , Youmna Farag , Svetlana Stoyanchev , Rama Doddipatla

Deepfake speech detection systems are often limited to binary classification tasks and struggle to generate interpretable reasoning or provide context-rich explanations for their decisions. These models primarily extract latent embeddings…

Sound · Computer Science 2026-04-01 Runkun Chen , Yixiong Fang , Pengyu Chang , Yuante Li , Massa Baali , Bhiksha Raj

Large-scale generative language models such as GPT-3 are competitive few-shot learners. While these models are known to be able to jointly represent many different languages, their training data is dominated by English, potentially limiting…

Multimodal large language models (MLLMs) are expected to jointly interpret vision, audio, and language, yet existing video benchmarks rarely assess fine-grained reasoning about human speech. Many tasks remain visually solvable or only…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Le Thien Phuc Nguyen , Zhuoran Yu , Samuel Low Yu Hang , Subin An , Jeongik Lee , Yohan Ban , SeungEun Chung , Thanh-Huy Nguyen , JuWan Maeng , Soochahn Lee , Yong Jae Lee

Recent progress in Automatic Speech Recognition (ASR) has been coupled with a substantial increase in the model sizes, which may now contain billions of parameters, leading to slow inferences even with adapted hardware. In this context,…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-25 Hugo Malard , Salah Zaiem , Robin Algayres