中文
相关论文

相关论文: MuseAgent-1: Interactive Grounded Multimodal Under…

200 篇论文

AI-empowered music processing is a diverse field that encompasses dozens of tasks, ranging from generation tasks (e.g., timbre synthesis) to comprehension tasks (e.g., music classification). For developers and amateurs, it is very difficult…

计算与语言 · 计算机科学 2023-10-26 Dingyao Yu , Kaitao Song , Peiling Lu , Tianyu He , Xu Tan , Wei Ye , Shikun Zhang , Jiang Bian

Recent multimodal LLMs have shown promise in chart-based visual question answering, but their performance declines sharply on unannotated charts-those requiring precise visual interpretation rather than relying on textual shortcuts. To…

人工智能 · 计算机科学 2026-01-08 Rachneet Kaur , Nishan Srishankar , Zhen Zeng , Sumitra Ganesh , Manuela Veloso

Multimodal Large Language Models (MLLMs) have achieved remarkable visual reasoning abilities in natural images, text-rich documents, and graphic designs. However, their ability to interpret music sheets remains underexplored. To bridge this…

计算机视觉与模式识别 · 计算机科学 2025-08-15 Jian Chen , Wenye Ma , Penghang Liu , Wei Wang , Tengwei Song , Ming Li , Chenguang Wang , Jiayu Qin , Ruiyi Zhang , Changyou Chen

Remote-sensing mineral exploration is critical for identifying economically viable mineral deposits, yet it poses significant challenges for multimodal large language models (MLLMs). These include limitations in domain-specific geological…

人工智能 · 计算机科学 2024-12-24 Beibei Yu , Tao Shen , Hongbin Na , Ling Chen , Denqi Li

Omnimodal large language models have made significant strides in unifying audio and visual modalities; however, they often face challenges in fine-grained cross-modal understanding and have difficulty with multimodal alignment. To address…

计算机视觉与模式识别 · 计算机科学 2026-02-06 Keda Tao , Wenjie Du , Bohan Yu , Weiqiang Wang , Jian Liu , Huan Wang

Research agents have recently achieved significant progress in information seeking and synthesis across heterogeneous textual and visual sources. In this paper, we introduce MuSEAgent, a multimodal reasoning agent that enhances…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Shijian Wang , Jiarui Jin , Runhao Fu , Zexuan Yan , Xingjian Wang , Mengkang Hu , Eric Wang , Xiaoxi Li , Kangning Zhang , Li Yao , Wenxiang Jiao , Xuelian Cheng , Yuan Lu , Zongyuan Ge

Multimodal large language models (MLLMs) have shown strong capabilities but remain limited to fixed modality pairs and require costly fine-tuning with large aligned datasets. Building fully omni-capable models that can integrate text,…

人工智能 · 计算机科学 2025-11-06 Huawei Lin , Yunzhi Shi , Tong Geng , Weijie Zhao , Wei Wang , Ravender Pal Singh

Understanding complete musical scores entails integrated reasoning over pitch, rhythm, harmony, and large-scale structure, yet the ability of Large Language Models and Vision--Language Models to interpret full musical notation remains…

Mathematical error detection in educational settings presents a significant challenge for Multimodal Large Language Models (MLLMs), requiring a sophisticated understanding of both visual and textual mathematical content along with complex…

计算与语言 · 计算机科学 2025-05-21 Yibo Yan , Shen Wang , Jiahao Huo , Philip S. Yu , Xuming Hu , Qingsong Wen

Dermatological diagnosis requires integrating fine-grained visual perception with expert clinical knowledge. Although Multimodal Large Language Models (MLLMs) facilitate interactive medical image analysis, their application in dermatology…

计算机视觉与模式识别 · 计算机科学 2026-05-15 Yize Liu , Siyuan Yan , Ming Hu , Lie Ju , Xieji Li , Feilong Tang , Wei Feng , Zongyuan Ge

Multimodal Affective Computing (MAC) aims to recognize and interpret human emotions by integrating information from diverse modalities such as text, video, and audio. Recent advancements in Multimodal Large Language Models (MLLMs) have…

人工智能 · 计算机科学 2025-08-05 Miaosen Luo , Jiesen Long , Zequn Li , Yunying Yang , Yuncheng Jiang , Sijie Mai

LLM-based multimodal emotion recognition relies on static parametric memory and often hallucinates when interpreting nuanced affective states. In this paper, given that single-round retrieval-augmented generation is highly susceptible to…

计算机视觉与模式识别 · 计算机科学 2026-04-15 Zeheng Wang , Zitong Yu , Yijie Zhu , Bo Zhao , Haochen Liang , Taorui Wang , Wei Xia , Jiayu Zhang , Zhishu Liu , Hui Ma , Fei Ma , Qi Tian

Large Audio-Language Models (LALMs) perform well on audio understanding tasks but lack multistep reasoning and tool-calling found in recent Large Language Models (LLMs). This paper presents AudioToolAgent, a framework that coordinates…

声音 · 计算机科学 2026-02-16 Gijs Wijngaard , Elia Formisano , Michel Dumontier , Jenia Jitsev

Recent Large Audio-Language Models (LALMs) have demonstrated promising abilities in understanding musical content. However, whether their responses are grounded in the correct temporal regions of the audio remains underexplored. This…

计算与语言 · 计算机科学 2026-05-29 Daeyong Kwon , Qiyu Wu , Shinobu Kuriya , Junghyun Koo , Shuyang Cui , Zhi Zhong , Wei-Hsiang Liao , Hiromi Wakaki , Yuki Mitsufuji

We introduce DriveAgent, a novel multi-agent autonomous driving framework that leverages large language model (LLM) reasoning combined with multimodal sensor fusion to enhance situational understanding and decision-making. DriveAgent…

机器人学 · 计算机科学 2025-05-06 Xinmeng Hou , Wuqi Wang , Long Yang , Hao Lin , Jinglun Feng , Haigen Min , Xiangmo Zhao

Multimodal Stance Detection (MSD) is crucial for understanding public discourse, yet effectively fusing text and image, especially with conflicting signals, remains challenging. Existing methods often face difficulties with contextual…

人工智能 · 计算机科学 2026-05-01 Weihai Lu , Zhejun Zhao , Yanshu Li , Huan He

Recent advances in Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities across various vision-language tasks. However, their reasoning abilities in the multimodal symbolic music domain remain largely…

声音 · 计算机科学 2026-01-26 Gagan Mundada , Yash Vishe , Amit Namburi , Xin Xu , Zachary Novack , Julian McAuley , Junda Wu

Personalized music recommendation in conversational scenarios usually requires a deep understanding of user preferences and nuanced musical context, yet existing methods often struggle with balancing specialized domain knowledge and…

人工智能 · 计算机科学 2025-12-19 Wendong Bi , Yirong Mao , Xianglong Liu , Kai Tian , Jian Zhang , Hanjie Wang , Wenhui Que

Graphical User Interface (GUI) Agents, powered by multimodal large language models (MLLMs), have shown great potential for task automation on computing devices such as computers and mobile phones. However, existing agents face challenges in…

人工智能 · 计算机科学 2025-01-09 Yuhang Liu , Pengxiang Li , Zishu Wei , Congkai Xie , Xueyu Hu , Xinchen Xu , Shengyu Zhang , Xiaotian Han , Hongxia Yang , Fei Wu

Multi-Modal Large Language Models (MLLMs), despite being successful, exhibit limited generality and often fall short when compared to specialized models. Recently, LLM-based agents have been developed to address these challenges by…

计算与语言 · 计算机科学 2024-10-08 Binxu Li , Tiankai Yan , Yuanting Pan , Jie Luo , Ruiyang Ji , Jiayuan Ding , Zhe Xu , Shilong Liu , Haoyu Dong , Zihao Lin , Yixin Wang
‹ 上一页 1 2 3 10 下一页 ›