中文
相关论文

相关论文: LLaVaOLMoBitnet1B: Ternary LLM goes Multimodal!

200 篇论文

In the era of Large Language Models (LLMs), tremendous strides have been made in the field of multimodal understanding. However, existing advanced algorithms are limited to effectively utilizing the immense representation capabilities and…

人工智能 · 计算机科学 2023-09-06 Hao Feng , Zijian Wang , Jingqun Tang , Jinghui Lu , Wengang Zhou , Houqiang Li , Can Huang

The evolution of Omni-Modal Large Language Models~(Omni-LLMs) has revolutionized human--computer interaction, enabling unified audio-visual perception and speech response. However, existing Omni-LLMs struggle with complex real-world…

声音 · 计算机科学 2026-03-10 Wenjie Tian , Zhixian Zhao , Jingbin Hu , Huakang Chen , Haohe Liu , Binshen Mu , Lei Xie

Large Language Models (LLMs), benefiting from the auto-regressive modelling approach performed on massive unannotated texts corpora, demonstrates powerful perceptual and reasoning capabilities. However, as for extending auto-regressive…

计算机视觉与模式识别 · 计算机科学 2024-09-24 Tianshuo Peng , Zuchao Li , Lefei Zhang , Hai Zhao , Ping Wang , Bo Du

Guiding users through complex procedural plans is an inherently multimodal task in which having visually illustrated plan steps is crucial to deliver an effective plan guidance. However, existing works on plan-following language models…

计算机视觉与模式识别 · 计算机科学 2024-10-22 Diogo Glória-Silva , David Semedo , João Magalhães

Large Language Models (LLMs) are transforming artificial intelligence, enabling autonomous agents to perform diverse tasks across various domains. These agents, proficient in human-like text comprehension and generation, have the potential…

人工智能 · 计算机科学 2024-04-10 Saikat Barua

Multimodal search has become increasingly important in providing users with a natural and effective way to ex-press their search intentions. Images offer fine-grained details of the desired products, while text allows for easily…

计算机视觉与模式识别 · 计算机科学 2024-04-25 Oriol Barbany , Michael Huang , Xinliang Zhu , Arnab Dhua

Training a Multimodal Large Language Model (MLLM) from scratch, like GPT-4, is resource-intensive. Regarding Large Language Models (LLMs) as the core processor for multimodal information, our paper introduces LMEye, a human-like eye with a…

计算机视觉与模式识别 · 计算机科学 2023-09-29 Yunxin Li , Baotian Hu , Xinyu Chen , Lin Ma , Yong Xu , Min Zhang

The advance of Artificial Intelligence (AI) is continuously reshaping the future 6G wireless communications. Particularly, the development of Large Language Models (LLMs) offers a promising approach to effectively improve the performance…

信息论 · 计算机科学 2025-03-10 Tianyue Zheng , Linglong Dai

Recent Large Multimodal Models have demonstrated remarkable reasoning capabilities, especially in solving complex mathematical problems and realizing accurate spatial perception. Our key insight is that these emerging abilities can…

人工智能 · 计算机科学 2025-05-20 Weiliang Tang , Dong Jing , Jia-Hui Pan , Zhiwu Lu , Yun-Hui Liu , Li Erran Li , Mingyu Ding , Chi-Wing Fu

Multimodal Large Language Models (MLLMs) have demonstrated extraordinary progress in bridging textual and visual inputs. However, MLLMs still face challenges in situated physical and social interactions in sensorally rich, multimodal and…

神经元与认知 · 定量生物学 2025-10-17 Akila Kadambi , Lisa Aziz-Zadeh , Antonio Damasio , Marco Iacoboni , Srini Narayanan

Recent advances in large language models (LLMs) have substantially improved single-turn task performance, yet real-world applications increasingly demand sophisticated multi-turn interactions. This survey provides a comprehensive review of…

计算与语言 · 计算机科学 2026-04-21 Yubo Li , Xiaobin Shen , Yidi Miao , Xinyu Yao , Xueying Ding , Ramayya Krishnan , Rema Padman

Multi-modal large language models (MLLMs) are trained based on large language models (LLM), with an enhanced capability to comprehend multi-modal inputs and generate textual responses. While they excel in multi-modal tasks, the pure NLP…

计算与语言 · 计算机科学 2023-09-14 Haoqin Tu , Bingchen Zhao , Chen Wei , Cihang Xie

Large Language Models (LLMs) have demonstrated impressive performance on multimodal tasks, without any multimodal finetuning. They are the building block for Large Multimodal Models, yet, we still lack a proper understanding of their…

计算机视觉与模式识别 · 计算机科学 2024-10-08 Mustafa Shukor , Matthieu Cord

Although multimodal large language models (MLLMs) have achieved promising results on a wide range of vision-language tasks, their ability to perceive and understand human faces is rarely explored. In this work, we comprehensively evaluate…

计算机视觉与模式识别 · 计算机科学 2024-10-29 Haomiao Sun , Mingjie He , Tianheng Lian , Hu Han , Shiguang Shan

In this survey, we systematically analyze techniques used to adapt large multimodal models (LMMs) for low-resource (LR) languages, examining approaches ranging from visual enhancement and data creation to cross-modal transfer and fusion…

计算与语言 · 计算机科学 2026-02-03 Marian Lupascu , Ana-Cristina Rogoz , Mihai Sorin Stupariu , Radu Tudor Ionescu

Large multimodal models (LMM) have recently shown encouraging progress with visual instruction tuning. In this note, we show that the fully-connected vision-language cross-modal connector in LLaVA is surprisingly powerful and…

计算机视觉与模式识别 · 计算机科学 2024-05-17 Haotian Liu , Chunyuan Li , Yuheng Li , Yong Jae Lee

Large language models (LLMs) have demonstrated remarkable language abilities. GPT-4, based on advanced LLMs, exhibits extraordinary multimodal capabilities beyond previous visual language models. We attribute this to the use of more…

计算与语言 · 计算机科学 2023-05-23 Feilong Chen , Minglun Han , Haozhi Zhao , Qingyang Zhang , Jing Shi , Shuang Xu , Bo Xu

We present LLaVA-OneVision-1.5, a novel family of Large Multimodal Models (LMMs) that achieve state-of-the-art performance with significantly reduced computational and financial costs. Different from the existing works, LLaVA-OneVision-1.5…

The evolution from Large Language Models (LLMs) to Multimodal Large Language Models (MLLMs) has spurred research into extending In-Context Learning (ICL) to its multimodal counterpart. Existing such studies have primarily concentrated on…

机器学习 · 计算机科学 2024-07-23 Yuchen Zeng , Wonjun Kang , Yicong Chen , Hyung Il Koo , Kangwook Lee

Multi-modal Large Language Models (MLLMs) integrate visual and linguistic reasoning to address complex tasks such as image captioning and visual question answering. While MLLMs demonstrate remarkable versatility, MLLMs appears limited…