English
Related papers

Related papers: MusiCRS: Benchmarking Audio-Centric Conversational…

200 papers

This work presents a user-centric recommendation framework, designed as a pipeline with four distinct, connected, and customizable phases. These phases are intended to improve explainability and boost user engagement. We have collected the…

Information Retrieval · Computer Science 2025-05-19 Jaime Ramirez Castillo , M. Julia Flores , Ann E. Nicholson

Conversational recommender systems (CRS) aim to provide highquality recommendations in conversations. However, most conventional CRS models mainly focus on the dialogue understanding of the current session, ignoring other rich multi-aspect…

Information Retrieval · Computer Science 2022-04-26 Shuokai Li , Ruobing Xie , Yongchun Zhu , Xiang Ao , Fuzhen Zhuang , Qing He

With the rapid progress of artificial intelligence (AI) in multi-modal understanding, there is increasing potential for video comprehension technologies to support professional domains such as medical education. However, existing benchmarks…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Shenxi Liu , Kan Li , Mingyang Zhao , Yuhang Tian , Bin Li , Shoujun Zhou , Hongliang Li , Fuxia Yang

We present MultiChallenge, a pioneering benchmark evaluating large language models (LLMs) on conducting multi-turn conversations with human users, a crucial yet underexamined capability for their applications. MultiChallenge identifies four…

Computation and Language · Computer Science 2025-03-07 Ved Sirdeshmukh , Kaustubh Deshpande , Johannes Mols , Lifeng Jin , Ed-Yeremai Cardona , Dean Lee , Jeremy Kritz , Willow Primack , Summer Yue , Chen Xing

User-centric evaluation has become a key paradigm for assessing Conversational Recommender Systems (CRS), aiming to capture subjective qualities such as satisfaction, trust, and rapport. To enable scalable evaluation, recent work…

Information Retrieval · Computer Science 2026-02-20 Michael Müller , Amir Reza Mohammadi , Andreas Peintner , Beatriz Barroso Gstrein , Günther Specht , Eva Zangerle

Mental-health support is increasingly mediated by conversational systems (e.g., LLM-based tools), but users often lack structured ways to audit the quality and potential risks of the support they receive. We introduce CounselReflect, an…

Computation and Language · Computer Science 2026-04-01 Yahan Li , Chaohao Du , Zeyang Li , Christopher Chun Kuizon , Shupeng Cheng , Angel Hsing-Chi Hwang , Adam C. Frank , Ruishan Liu

Recent LLM benchmarks have tested models on a range of phenomena, but are still focused primarily on natural language understanding for extraction of explicit information, such as QA or summarization, with responses often targeting…

Computation and Language · Computer Science 2025-11-11 Lanni Bu , Lauren Levine , Amir Zeldes

Real-world financial analysis involves information across multiple languages and modalities, from reports and news to scanned filings and meeting recordings. Yet most existing evaluations of LLMs in finance remain text-only, monolingual,…

Recommender systems have become indispensable in music streaming services, enhancing user experiences by personalizing playlists and facilitating the serendipitous discovery of new music. However, the existing recommender systems overlook…

Information Retrieval · Computer Science 2023-08-29 Yunhak Oh , Sukwon Yun , Dongmin Hyun , Sein Kim , Chanyoung Park

We present the Conversational Data Retrieval (CDR) benchmark, the first comprehensive test set for evaluating systems that retrieve conversation data for product insights. With 1.6k queries across five analytical tasks and 9.1k…

Computation and Language · Computer Science 2026-02-17 Yohan Lee , Yongwoo Song , Sangyeop Kim

Robust selective auditory attention under multilingual interference is critical for reliable deployment of Large Audio Language Models (LALMs). We introduce MUSA, a cocktail party-inspired multilingual benchmark for source-grounded…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-19 Heejoon Koo

The ability of critique is vital for models to self-improve and serve as reliable AI assistants. While extensively studied in language-only settings, multimodal critique of Large Multimodal Models (LMMs) remains underexplored despite their…

Computation and Language · Computer Science 2025-11-13 Gailun Zeng , Ziyang Luo , Hongzhan Lin , Yuchen Tian , Kaixin Li , Ziyang Gong , Jianxiong Guo , Jing Ma

Large Audio Language Models (LALMs) are increasingly capable of reasoning over audio. However, existing benchmarks provide limited coverage of reasoning in polyphonic audio, where multiple sound events co-occur and induce compositional…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-11 Yuanjian Chen , Yang Xiao , Han Yin , Xubo Liu , Jinjie Huang , Ting Dang

Understanding Theory of Mind is essential for building socially intelligent multimodal agents capable of perceiving and interpreting human behavior. We introduce MoMentS (Multimodal Mental States), a comprehensive benchmark designed to…

Human perception and experience of music is highly context-dependent. Contextual variability contributes to differences in how we interpret and interact with music, challenging the design of robust models for information retrieval.…

Sound · Computer Science 2022-10-31 Kleanthis Avramidis , Shanti Stewart , Shrikanth Narayanan

Recent advancements in multimodal large language models (MLLMs) have aimed to integrate and interpret data across diverse modalities. However, the capacity of these models to concurrently process and reason about multiple modalities remains…

Social media has amplified the reach of financial influencers known as "finfluencers," who share stock recommendations on platforms like YouTube. Understanding their influence requires analyzing multimodal signals like tone, delivery style,…

Multimedia · Computer Science 2025-07-14 Michael Galarnyk , Veer Kejriwal , Agam Shah , Yash Bhardwaj , Nicholas Meyer , Anand Krishnan , Sudheer Chava

Recent advancements in Large language models (LLMs) have demonstrated remarkable capabilities across diverse domains. While they exhibit strong zero-shot performance on various tasks, LLMs' effectiveness in music-related applications…

Computation and Language · Computer Science 2025-12-09 Daeyong Kwon , SeungHeon Doh , Juhan Nam

Recommender systems (RSs) offer personalized navigation experiences on online platforms, but recommendation remains a challenging task, particularly in specific scenarios and domains. Multimodality can help tap into richer information…

Despite recent advances in multimodal large language models (MLLMs), their ability to understand and interact with music remains limited. Music understanding requires grounded reasoning over symbolic scores and expressive performance audio,…

Multimedia · Computer Science 2026-01-21 Qihao Zhao , Yunqi Cao , Yangyu Huang , Hui Yi Leong , Fan Zhang , Kim-Hui Yap , Wei Hu
‹ Prev 1 8 9 10 Next ›