English
Related papers

Related papers: Falcon2-11B Technical Report

200 papers

Vision-language pre-training (VLP) on large-scale datasets has shown premier performance on various downstream tasks. In contrast to plenty of available benchmarks with English corpus, large-scale pre-training datasets and downstream…

Computer Vision and Pattern Recognition · Computer Science 2023-11-09 Chunyu Xie , Heng Cai , Jincheng Li , Fanjing Kong , Xiaoyu Wu , Jianfei Song , Henrique Morimitsu , Lin Yao , Dexin Wang , Xiangzheng Zhang , Dawei Leng , Baochang Zhang , Xiangyang Ji , Yafeng Deng

Multimodal Large Language Models (MLLMs) rely on strong linguistic reasoning inherited from their base language models. However, multimodal instruction fine-tuning paradoxically degrades this text's reasoning capability, undermining…

Computation and Language · Computer Science 2026-01-13 Zijing Wang , Yongkang Liu , Mingyang Wang , Ercong Nie , Deyuan Chen , Zhengjie Zhao , Shi Feng , Daling Wang , Xiaocui Yang , Yifei Zhang , Hinrich Schütze

Text-to-video (T2V) generation has gained significant attention recently. However, the costs of training a T2V model from scratch remain persistently high, and there is considerable room for improving the generation performance, especially…

Computer Vision and Pattern Recognition · Computer Science 2024-12-30 Zhefan Rao , Liya Ji , Yazhou Xing , Runtao Liu , Zhaoyang Liu , Jiaxin Xie , Ziqiao Peng , Yingqing He , Qifeng Chen

Modern language model (LM) training has been divided into multiple stages, making it difficult for downstream developers to evaluate the impact of design choices made at each stage. We present EvoLM, a model suite that enables systematic…

Computation and Language · Computer Science 2025-11-19 Zhenting Qi , Fan Nie , Alexandre Alahi , James Zou , Himabindu Lakkaraju , Yilun Du , Eric Xing , Sham Kakade , Hanlin Zhang

With the growth of high-quality data and advancement in visual pre-training paradigms, Video Foundation Models (VFMs) have made significant progress recently, demonstrating their remarkable performance on traditional video understanding…

Computer Vision and Pattern Recognition · Computer Science 2024-07-10 Xinhao Li , Zhenpeng Huang , Jing Wang , Kunchang Li , Limin Wang

Multimodal Large Language Models (MLLMs) have endowed LLMs with the ability to perceive and understand multi-modal signals. However, most of the existing MLLMs mainly adopt vision encoders pretrained on coarsely aligned image-text pairs,…

Computer Vision and Pattern Recognition · Computer Science 2023-11-28 Gongwei Chen , Leyang Shen , Rui Shao , Xiang Deng , Liqiang Nie

In this paper, we present ZonUI-3B, a lightweight Vision-Language Model (VLM) that can be fully trained on a single consumer-grade GPU (RTX 4090) while delivering performance comparable to significantly larger models on GUI grounding tasks.…

Computer Vision and Pattern Recognition · Computer Science 2025-07-21 ZongHan Hsieh , Tzer-Jen Wei , ShengJing Yang

Large language models are pre-trained on ever-growing token budgets under the assumption that better pre-training performance translates to improved downstream models. In this work, we challenge this assumption and show that extended…

Computation and Language · Computer Science 2025-03-31 Jacob Mitchell Springer , Sachin Goyal , Kaiyue Wen , Tanishq Kumar , Xiang Yue , Sadhika Malladi , Graham Neubig , Aditi Raghunathan

In this paper, the LCV2 modular method is proposed for the Grounded Visual Question Answering task in the vision-language multimodal domain. This approach relies on a frozen large language model (LLM) as intermediate mediator between the…

Computer Vision and Pattern Recognition · Computer Science 2024-03-26 Yuhan Chen , Lumei Su , Lihua Chen , Zhiwei Lin

We propose MindVL, a multimodal large language model (MLLMs) trained on Ascend NPUs. The training of state-of-the-art MLLMs is often confined to a limited set of hardware platforms and relies heavily on massive, undisclosed data recipes,…

Computer Vision and Pattern Recognition · Computer Science 2025-10-01 Feilong Chen , Yijiang Liu , Yi Huang , Hao Wang , Miren Tian , Ya-Qi Yu , Minghui Liao , Jihao Wu

The increasing integration of Visual Language Models (VLMs) into AI systems necessitates robust model alignment, especially when handling multimodal content that combines text and images. Existing evaluation datasets heavily lean towards…

Computation and Language · Computer Science 2026-03-05 Gabriel Downer , Sean Craven , Damian Ruck , Jake Thomas

Vision-language models (VLMs) are achieving increasingly strong performance on multimodal tasks. However, reasoning capabilities remain limited particularly for smaller VLMs, while those of large-language models (LLMs) have seen numerous…

Computation and Language · Computer Science 2024-03-20 Victor Carbune , Hassan Mansoor , Fangyu Liu , Rahul Aralikatte , Gilles Baechler , Jindong Chen , Abhanshu Sharma

Foundation models have achieved transformative success across biomedical domains by enabling holistic understanding of multimodal data. However, their application in surgery remains underexplored. Surgical intelligence presents unique…

Computer Vision and Pattern Recognition · Computer Science 2025-06-04 Zhitao Zeng , Zhu Zhuo , Xiaojun Jia , Erli Zhang , Junde Wu , Jiaan Zhang , Yuxuan Wang , Chang Han Low , Jian Jiang , Zilong Zheng , Xiaochun Cao , Yutong Ban , Qi Dou , Yang Liu , Yueming Jin

In this work, we introduce LokiLM, a 1.4B parameter large language model trained on 500B tokens. Our model performs strongly in natural language reasoning tasks and achieves state-of-the-art performance among models with 1.5B parameters or…

Computation and Language · Computer Science 2024-07-11 Justin Kiefel , Shrey Shah

The rapid advancement of Large Language Models (LLMs) has outpaced the scalability of traditional evaluation benchmarks, which remain heavily dependent on labor-intensive expert curation. We address this bottleneck with Conv-to-Bench, a…

Large Language Models (LLMs) have played an important role in many fields due to their powerful capabilities.However, their massive number of parameters leads to high deployment requirements and incurs significant inference costs, which…

Predicting protein function from sequence is a central challenge in computational biology. While existing methods rely heavily on structured ontologies or similarity-based techniques, they often lack the flexibility to express…

Computational Engineering, Finance, and Science · Computer Science 2025-10-27 Xiao Fei , Michail Chatzianastasis , Sarah Almeida Carneiro , Hadi Abdine , Lawrence P. Petalidis , Michalis Vazirgiannis

Multimodal large language models (MLLMs) have shown remarkable potential in various domains, yet their application in the medical field is hindered by several challenges. General-purpose MLLMs often lack the specialized knowledge required…

Artificial Intelligence · Computer Science 2025-09-29 Guanghao Zhu , Zhitian Hou , Zeyu Liu , Zhijie Sang , Congkai Xie , Hongxia Yang

Artificial intelligence has made great progress in recent years, particularly in the development of Vision--Language Models (VLMs) that understand both visual and textual data. However, these advancements remain largely limited to English,…

Computation and Language · Computer Science 2025-12-12 Jules Lahmi , Alexis Roger

We present MM1.5, a new family of multimodal large language models (MLLMs) designed to enhance capabilities in text-rich image understanding, visual referring and grounding, and multi-image reasoning. Building upon the MM1 architecture,…

‹ Prev 1 8 9 10 Next ›