English
Related papers

Related papers: MMSU: A Massive Multi-task Spoken Language Underst…

200 papers

Expanding the language coverage of speech technology has the potential to improve access to information for many more people. However, current speech technology is restricted to about one hundred languages which is a small fraction of the…

The success of large language models (LLMs) has prompted efforts to integrate speech and audio data, aiming to create general foundation models capable of processing both textual and non-textual inputs. Recent advances, such as GPT-4o,…

Computation and Language · Computer Science 2024-10-18 Fan Bu , Yuhao Zhang , Xidong Wang , Benyou Wang , Qun Liu , Haizhou Li

Speech large language models (SpeechLLMs) have extended human-machine interactions from the text modality to the dynamic speech domain. Spoken dialogues convey diverse information, including semantic concepts, acoustic variations,…

Computation and Language · Computer Science 2026-01-14 Heyang Liu , Yuhao Wang , Ziyang Cheng , Hongcheng Liu , Yiqi Li , Yixuan Hou , Ronghua Wu , Qunshan Gu , Yanfeng Wang , Yu Wang

Multimodal Sarcasm Understanding (MSU) has a wide range of applications in the news field such as public opinion analysis and forgery detection. However, existing MSU benchmarks and approaches usually focus on sentence-level MSU. In…

Computation and Language · Computer Science 2023-12-27 Hang Du , Guoshun Nan , Sicheng Zhang , Binzhu Xie , Junrui Xu , Hehe Fan , Qimei Cui , Xiaofeng Tao , Xudong Jiang

We introduce AudioBench, a universal benchmark designed to evaluate Audio Large Language Models (AudioLLMs). It encompasses 8 distinct tasks and 26 datasets, among which, 7 are newly proposed datasets. The evaluation targets three main…

Sound · Computer Science 2025-05-07 Bin Wang , Xunlong Zou , Geyu Lin , Shuo Sun , Zhuohan Liu , Wenyu Zhang , Zhengyuan Liu , AiTi Aw , Nancy F. Chen

Millions of people take surveys every day, from market polls and academic studies to medical questionnaires and customer feedback forms. These datasets capture valuable insights, but their scale and structure present a unique challenge for…

Artificial Intelligence · Computer Science 2025-10-31 Duc-Hai Nguyen , Vijayakumar Nanjappan , Barry O'Sullivan , Hoang D. Nguyen

Multimodal Large Language Model (MLLM) relies on the powerful LLM to perform multimodal tasks, showing amazing emergent abilities in recent studies, such as writing poems based on an image. However, it is difficult for these case studies to…

Computer Vision and Pattern Recognition · Computer Science 2025-10-27 Chaoyou Fu , Peixian Chen , Yunhang Shen , Yulei Qin , Mengdan Zhang , Xu Lin , Jinrui Yang , Xiawu Zheng , Ke Li , Xing Sun , Yunsheng Wu , Rongrong Ji , Caifeng Shan , Ran He

Speech-to-Speech (S2S) Large Language Models (LLMs) are foundational to natural human-computer interaction, enabling end-to-end spoken dialogue systems. However, evaluating these models remains a fundamental challenge. We propose…

Computation and Language · Computer Science 2025-11-11 Yuan Ge , Junxiang Zhang , Xiaoqian Liu , Bei Li , Xiangnan Ma , Chenglong Wang , Kaiyang Ye , Yangfan Du , Linfeng Zhang , Yuxin Huang , Tong Xiao , Zhengtao Yu , JingBo Zhu

The staggering pace with which the capabilities of large language models (LLMs) are increasing, as measured by a range of commonly used natural language understanding (NLU) benchmarks, raises many questions regarding what "understanding"…

Computation and Language · Computer Science 2024-04-19 Xenia Ohmer , Elia Bruni , Dieuwke Hupkes

Paralinguistic cues are essential for natural human-computer interaction, yet their evaluation in Large Audio-Language Models (LALMs) remains limited by coarse feature coverage and the inherent subjectivity of assessment. To address these…

Computation and Language · Computer Science 2026-04-23 Ruohan Liu , Shukang Yin , Tao Wang , Dong Zhang , Weiji Zhuang , Shuhuai Ren , Ran He , Caifeng Shan , Chaoyou Fu

Speech large language models (speech-LLMs) integrate speech and text-based foundation models to provide a unified framework for handling a wide range of downstream tasks. In this paper, we introduce WHISMA, a speech-LLM tailored for spoken…

Audio and Speech Processing · Electrical Eng. & Systems 2024-08-30 Mohan Li , Cong-Thanh Do , Simon Keizer , Youmna Farag , Svetlana Stoyanchev , Rama Doddipatla

Advances in large language models (LLMs) have enabled significant capabilities in audio processing, resulting in state-of-the-art models now known as Large Audio Language Models (LALMs). However, minimal work has been done to measure audio…

Sound · Computer Science 2026-03-11 Laya Iyer , Angelina Wang , Sanmi Koyejo

Audio Large Language Models (AudioLLMs) have achieved strong results in semantic tasks like speech recognition and translation, but remain limited in modeling paralinguistic cues such as emotion. Existing approaches often treat emotion…

Computation and Language · Computer Science 2025-09-30 Wenyu Zhang , Yingxu He , Geyu Lin , Zhuohan Liu , Shuo Sun , Bin Wang , Xunlong Zou , Jeremy H. M. Wong , Qiongqiong Wang , Hardik B. Sailor , Nancy F. Chen , Ai Ti Aw

Multimodal Audio-Language Models (ALMs) can understand and reason over both audio and text. Typically, reasoning performance correlates with model size, with the best results achieved by models exceeding 8 billion parameters. However, no…

Sound · Computer Science 2025-03-12 Soham Deshmukh , Satvik Dixit , Rita Singh , Bhiksha Raj

Large Audio-Language Models (LALMs) as judges have emerged as a prominent approach for evaluating speech generation quality, yet their ability to assess speaker consistency across multi-turn dialogues remains unexplored. We present…

Computation and Language · Computer Science 2026-04-21 Jonggeun Lee , Junseong Pyo , Gyuhyeon Seo , Yohan Jo

We propose MMLU-SR, a novel dataset designed to measure the true comprehension abilities of Large Language Models (LLMs) by challenging their performance in question-answering tasks with modified terms. We reasoned that an agent that…

Computation and Language · Computer Science 2024-10-07 Wentian Wang , Sarthak Jain , Paul Kantor , Jacob Feldman , Lazaros Gallos , Hao Wang

Multilingual capability is an essential aspect for large multimodal models, since they are usually deployed across various countries and languages. However, most existing benchmarks for multilingual multimodal reasoning struggle to…

Computer Vision and Pattern Recognition · Computer Science 2025-04-28 Hongyu Wang , Jiayu Xu , Senwei Xie , Ruiping Wang , Jialin Li , Zhaojie Xie , Bin Zhang , Chuyan Xiong , Xilin Chen

In recent years, large language models (LLMs) have driven major advances in language understanding, marking a significant step toward artificial general intelligence (AGI). With increasing demands for higher-level semantics and cross-modal…

Computation and Language · Computer Science 2025-09-30 Yuntao Shou , Tao Meng , Wei Ai , Keqin Li

Understanding complete musical scores entails integrated reasoning over pitch, rhythm, harmony, and large-scale structure, yet the ability of Large Language Models and Vision--Language Models to interpret full musical notation remains…

Large language models (LLMs) have advanced in text and vision, but their reasoning on audio remains limited. Most existing methods rely on dense audio embeddings, which are difficult to interpret and often fail on structured reasoning…

Sound · Computer Science 2025-11-11 Termeh Taheri , Yinghao Ma , Emmanouil Benetos