English
Related papers

Related papers: MERaLiON-AudioLLM: Bridging Audio and Language wit…

200 papers

Large language models (LLMs) have exhibited remarkable capabilities across a variety of domains and tasks, challenging our understanding of learning and cognition. Despite the recent success, current LLMs are not capable of processing…

We present Polyglot-Lion, a family of compact multilingual automatic speech recognition (ASR) models tailored for the linguistic landscape of Singapore, covering English, Mandarin, Tamil, and Malay. Our models are obtained by fine-tuning…

Computation and Language · Computer Science 2026-03-18 Quy-Anh Dang , Chris Ngo

This survey and application guide to multimodal large language models(MLLMs) explores the rapidly developing field of MLLMs, examining their architectures, applications, and impact on AI and Generative Models. Starting with foundational…

Artificial Intelligence · Computer Science 2025-12-02 Chia Xin Liang , Pu Tian , Caitlyn Heqi Yin , Yao Yua , Wei An-Hou , Li Ming , Xinyuan Song , Tianyang Wang , Ziqian Bi , Ming Liu

Speech Large Language Models (Speech LLMs) have emerged as a crucial paradigm in recent years, extending the capabilities of traditional LLMs to speech tasks such as automatic speech recognition (ASR) and spoken dialogue modeling. However,…

Computation and Language · Computer Science 2025-07-08 Phurich Saengthong , Boonnithi Jiaramaneepinit , Sheng Li , Manabu Okumura , Takahiro Shinozaki

Auditory foundation models, including auditory large language models (LLMs), process all sound inputs equally, independent of listener perception. However, human auditory perception is inherently selective: listeners focus on specific…

The high incidence and mortality rates associated with respiratory diseases underscores the importance of early screening. Machine learning models can automate clinical consultations and auscultation, offering vital support in this area.…

Machine Learning · Computer Science 2024-10-10 Yuwei Zhang , Tong Xia , Aaqib Saeed , Cecilia Mascolo

Large language models (LLMs) have been widely adopted due to their remarkable performance across various applications, driving the accelerated development of a large number of diverse models. However, these individual LLMs show limitations…

Computation and Language · Computer Science 2025-06-13 Kaushal Kumar Maurya , KV Aditya Srivatsa , Ekaterina Kochmar

This paper explores enabling large language models (LLMs) to understand spatial information from multichannel audio, a skill currently lacking in auditory LLMs. By leveraging LLMs' advanced cognitive and inferential abilities, the aim is to…

Sound · Computer Science 2024-06-17 Changli Tang , Wenyi Yu , Guangzhi Sun , Xianzhao Chen , Tian Tan , Wei Li , Jun Zhang , Lu Lu , Zejun Ma , Yuxuan Wang , Chao Zhang

Large Language Models (LLMs) have shown remarkable abilities across various tasks, yet their development has predominantly centered on high-resource languages like English and Chinese, leaving low-resource languages underserved. To address…

Computation and Language · Computer Science 2024-07-30 Wenxuan Zhang , Hou Pong Chan , Yiran Zhao , Mahani Aljunied , Jianyu Wang , Chaoqun Liu , Yue Deng , Zhiqiang Hu , Weiwen Xu , Yew Ken Chia , Xin Li , Lidong Bing

As audio-first agents become increasingly common in physical AI, conversational robots, and screenless wearables, audio large language models (audio-LLMs) must integrate speaker-specific understanding to support user authorization,…

Sound · Computer Science 2026-05-15 KiHyun Nam , Jungwoo Heo , Siu Bae , Ha-Jin Yu , Joon Son Chung

Recent advancements in reasoning optimization have greatly enhanced the performance of large language models (LLMs). However, existing work fails to address the complexities of audio-visual scenarios, underscoring the need for further…

Audio and Speech Processing · Electrical Eng. & Systems 2025-04-01 Sanjoy Chowdhury , Hanan Gani , Nishit Anand , Sayan Nag , Ruohan Gao , Mohamed Elhoseiny , Salman Khan , Dinesh Manocha

This paper presents DreamLLM, a learning framework that first achieves versatile Multimodal Large Language Models (MLLMs) empowered with frequently overlooked synergy between multimodal comprehension and creation. DreamLLM operates on two…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Runpei Dong , Chunrui Han , Yuang Peng , Zekun Qi , Zheng Ge , Jinrong Yang , Liang Zhao , Jianjian Sun , Hongyu Zhou , Haoran Wei , Xiangwen Kong , Xiangyu Zhang , Kaisheng Ma , Li Yi

This paper presents SOLOMON, a novel Neuro-inspired Large Language Model (LLM) Reasoning Network architecture that enhances the adaptability of foundation models for domain-specific applications. Through a case study in semiconductor layout…

Computation and Language · Computer Science 2025-02-10 Bo Wen , Xin Zhang

Audio-Language Models (ALMs), trained on paired audio-text data, are designed to process, understand, and reason about audio-centric multimodal content. Unlike traditional supervised approaches that use predefined labels, ALMs leverage…

Sound · Computer Science 2026-03-13 Yi Su , Jisheng Bai , Qisheng Xu , Kele Xu , Yong Dou

Audio-text retrieval is crucial for bridging acoustic signals and natural language. While contrastive dual-encoder architectures like CLAP have shown promise, they are fundamentally limited by the capacity of small-scale encoders.…

Sound · Computer Science 2026-02-23 Jilan Xu , Carl Thomé , Danijela Horak , Weidi Xie , Andrew Zisserman

Multimodal Audio-Language Models (ALMs) can understand and reason over both audio and text. Typically, reasoning performance correlates with model size, with the best results achieved by models exceeding 8 billion parameters. However, no…

Sound · Computer Science 2025-03-12 Soham Deshmukh , Satvik Dixit , Rita Singh , Bhiksha Raj

Full-duplex multimodal large language models (LLMs) provide a unified framework for addressing diverse speech understanding and generation tasks, enabling more natural and seamless human-machine conversations. Unlike traditional modularised…

Audio and Speech Processing · Electrical Eng. & Systems 2024-11-28 Wenyi Yu , Siyin Wang , Xiaoyu Yang , Xianzhao Chen , Xiaohai Tian , Jun Zhang , Guangzhi Sun , Lu Lu , Yuxuan Wang , Chao Zhang

While contemporary speech separation technologies adeptly process lengthy mixed audio waveforms, they are frequently challenged by the intricacies of real-world environments, including noisy and reverberant settings, which can result in…

Sound · Computer Science 2025-05-27 Zhaoxi Mu , Xinyu Yang , Gang Wang

Large language models (LLMs) hold significant potential for mental health support, capable of generating empathetic responses and simulating therapeutic conversations. However, existing LLM-based approaches often lack the clinical grounding…

Computation and Language · Computer Science 2025-11-04 He Hu , Yucheng Zhou , Juzheng Si , Qianning Wang , Hengheng Zhang , Fuji Ren , Fei Ma , Laizhong Cui , Qi Tian

We present SegLLM, a novel multi-round interactive reasoning segmentation model that enhances LLM-based segmentation by exploiting conversational memory of both visual and textual outputs. By leveraging a mask-aware multimodal LLM, SegLLM…

Computer Vision and Pattern Recognition · Computer Science 2024-11-04 XuDong Wang , Shaolun Zhang , Shufan Li , Konstantinos Kallidromitis , Kehan Li , Yusuke Kato , Kazuki Kozuka , Trevor Darrell