English
Related papers

Related papers: ZAYA1-VL-8B Technical Report

200 papers

We introduce FLARE, a family of vision language models (VLMs) with a fully vision-language alignment and integration paradigm. Unlike existing approaches that rely on single MLP projectors for modality alignment and defer cross-modal…

Computer Vision and Pattern Recognition · Computer Science 2026-04-30 Zheng Liu , Mengjie Liu , Jingzhou Chen , Jingwei Xu , Bin Cui , Conghui He , Wentao Zhang

Existing multilingual vision-language (VL) benchmarks often only cover a handful of languages. Consequently, evaluations of large vision-language models (LVLMs) predominantly target high-resource languages, underscoring the need for…

Computation and Language · Computer Science 2025-02-19 Fabian David Schmidt , Florian Schneider , Chris Biemann , Goran Glavaš

We introduce VL-JEPA, a vision-language model built on a Joint Embedding Predictive Architecture (JEPA). Instead of autoregressively generating tokens as in classical VLMs, VL-JEPA predicts continuous embeddings of the target texts. By…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Delong Chen , Mustafa Shukor , Theo Moutakanni , Willy Chung , Jade Yu , Tejaswi Kasarla , Yejin Bang , Allen Bolourchi , Yann LeCun , Pascale Fung

We present GLM-4.1V-Thinking, GLM-4.5V, and GLM-4.6V, a family of vision-language models (VLMs) designed to advance general-purpose multimodal understanding and reasoning. In this report, we share our key findings in the development of the…

We present an RL-central framework for Language and Vision Assistants (RLLaVA) with its formulation of Markov decision process (MDP). RLLaVA decouples RL algorithmic logic from model architecture and distributed execution, supporting…

Machine Learning · Computer Science 2025-12-29 Lei Zhao , Zihao Ma , Boyu Lin , Yuhe Liu , Wenjun Wu , Lei Huang

Deploying vision-language models (VLMs) in resource-constrained settings demands low latency and high throughput, yet existing compact VLMs often fall short of the inference speedups their smaller parameter counts suggest. To explain this…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Jiabo Huang , Zhizhong Li , Sina Sajadmanesh , Weiming Zhuang , Lingjuan Lyu

Vision-Language Models (VLMs) have achieved remarkable success in various multi-modal tasks, but they are often bottlenecked by the limited context window and high computational cost of processing high-resolution image inputs and videos.…

Computer Vision and Pattern Recognition · Computer Science 2025-03-04 Xubing Ye , Yukang Gan , Xiaoke Huang , Yixiao Ge , Yansong Tang

Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated satisfactory performance across various vision-language tasks. Current approaches for vision and language interaction fall into two categories:…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Feipeng Ma , Yizhou Zhou , Zheyu Zhang , Shilin Yan , Hebei Li , Zilong He , Siying Wu , Fengyun Rao , Yueyi Zhang , Xiaoyan Sun

In this paper, we focus on monolithic Multimodal Large Language Models (MLLMs) that integrate visual encoding and language decoding into a single LLM. In particular, we identify that existing pre-training strategies for monolithic MLLMs…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Gen Luo , Xue Yang , Wenhan Dou , Zhaokai Wang , Jiawen Liu , Jifeng Dai , Yu Qiao , Xizhou Zhu

Recent advances in Vision-Language-Action (VLA) models have opened new avenues for robot manipulation, yet existing methods exhibit limited efficiency and a lack of high-level knowledge and spatial awareness. To address these challenges, we…

Visual language models (VLMs) have made significant advances in accuracy in recent years. However, their efficiency has received much less attention. This paper introduces NVILA, a family of open VLMs designed to jointly optimize efficiency…

We introduce Nemotron Nano V2 VL, the latest model of the Nemotron vision-language series designed for strong real-world document understanding, long video comprehension, and reasoning tasks. Nemotron Nano V2 VL delivers significant…

Machine Learning · Computer Science 2025-11-10 NVIDIA , : , Amala Sanjay Deshmukh , Kateryna Chumachenko , Tuomas Rintamaki , Matthieu Le , Tyler Poon , Danial Mohseni Taheri , Ilia Karmanov , Guilin Liu , Jarno Seppanen , Guo Chen , Karan Sapra , Zhiding Yu , Adi Renduchintala , Charles Wang , Peter Jin , Arushi Goel , Mike Ranzinger , Lukas Voegtle , Philipp Fischer , Timo Roman , Wei Ping , Boxin Wang , Zhuolin Yang , Nayeon Lee , Shaokun Zhang , Fuxiao Liu , Zhiqi Li , Di Zhang , Greg Heinrich , Hongxu Yin , Song Han , Pavlo Molchanov , Parth Mannan , Yao Xu , Jane Polak Scowcroft , Tom Balough , Subhashree Radhakrishnan , Paris Zhang , Sean Cha , Ratnesh Kumar , Zaid Pervaiz Bhat , Jian Zhang , Darragh Hanley , Pritam Biswas , Jesse Oliver , Kevin Vasques , Roger Waleffe , Duncan Riach , Oluwatobi Olabiyi , Ameya Sunil Mahabaleshwarkar , Bilal Kartal , Pritam Gundecha , Khanh Nguyen , Alexandre Milesi , Eugene Khvedchenia , Ran Zilberstein , Ofri Masad , Natan Bagrov , Nave Assaf , Tomer Asida , Daniel Afrimi , Amit Zuker , Netanel Haber , Zhiyu Cheng , Jingyu Xin , Di Wu , Nik Spirin , Maryam Moosaei , Roman Ageev , Vanshil Atul Shah , Yuting Wu , Daniel Korzekwa , Unnikrishnan Kizhakkemadam Sreekumar , Wanli Jiang , Padmavathy Subramanian , Alejandra Rico , Sandip Bhaskar , Saeid Motiian , Kedi Wu , Annie Surla , Chia-Chih Chen , Hayden Wolff , Matthew Feinberg , Melissa Corpuz , Marek Wawrzos , Eileen Long , Aastha Jhunjhunwala , Paul Hendricks , Farzan Memarian , Benika Hall , Xin-Yu Wang , David Mosallanezhad , Soumye Singhal , Luis Vega , Katherine Cheung , Krzysztof Pawelec , Michael Evans , Katherine Luna , Jie Lou , Erick Galinkin , Akshay Hazare , Kaustubh Purandare , Ann Guan , Anna Warno , Chen Cui , Yoshi Suhara , Shibani Likhite , Seph Mard , Meredith Price , Laya Sleiman , Saori Kaji , Udi Karpas , Kari Briski , Joey Conway , Michael Lightstone , Jan Kautz , Mohammad Shoeybi , Mostofa Patwary , Jonathen Cohen , Oleksii Kuchaiev , Andrew Tao , Bryan Catanzaro

Recent advances in vision-language models (VLMs) have demonstrated strong generalization in natural image tasks. However, their performance often degrades on unmanned aerial vehicle (UAV)-based aerial imagery, which features high…

Computer Vision and Pattern Recognition · Computer Science 2026-05-07 Jiajin Guan , Haibo Mei , Bonan Zhang , Dan Liu , Yuanshuang Fu , Yue Zhang

The rapid development of large language models (LLMs) has spurred extensive research into their domain-specific capabilities, particularly mathematical reasoning. However, most open-source LLMs focus solely on mathematical reasoning,…

Computation and Language · Computer Science 2024-09-04 Shuai Peng , Di Fu , Liangcai Gao , Xiuqin Zhong , Hongguang Fu , Zhi Tang

In this report, we introduce Falcon-H1, a new series of large language models (LLMs) featuring hybrid architecture designs optimized for both high performance and efficiency across diverse use cases. Unlike earlier Falcon models built…

Image and language modeling is of crucial importance for vision-language pre-training (VLP), which aims to learn multi-modal representations from large-scale paired image-text data. However, we observe that most existing VLP methods focus…

Computer Vision and Pattern Recognition · Computer Science 2022-08-22 Sunan He , Taian Guo , Tao Dai , Ruizhi Qiao , Chen Wu , Xiujun Shu , Bo Ren

As multimodal large language models (MLLMs) advance, their large-scale architectures pose challenges for deployment in resource-constrained environments. In the age of large models, where energy efficiency, computational scalability and…

Computer Vision and Pattern Recognition · Computer Science 2025-09-11 Sike Xiang , Shuang Chen , Amir Atapour-Abarghouei

Recent advancements in multimodal large language models (MLLM) have shown a strong ability in visual perception, reasoning abilities, and vision-language understanding. However, the visual matching ability of MLLMs is rarely studied,…

Computer Vision and Pattern Recognition · Computer Science 2025-07-10 Yikang Zhou , Tao Zhang , Shilin Xu , Shihao Chen , Qianyu Zhou , Yunhai Tong , Shunping Ji , Jiangning Zhang , Lu Qi , Xiangtai Li

Multimodal Large Language Models (MLLMs) have demonstrated notable capabilities in general visual understanding and reasoning tasks. However, their deployment is hindered by substantial computational costs in both training and inference,…

Computer Vision and Pattern Recognition · Computer Science 2024-07-23 Muyang He , Yexin Liu , Boya Wu , Jianhao Yuan , Yueze Wang , Tiejun Huang , Bo Zhao

We present Qianfan-VL, a series of multimodal large language models ranging from 3B to 70B parameters, achieving state-of-the-art performance through innovative domain enhancement techniques. Our approach employs multi-stage progressive…

‹ Prev 1 4 5 6 7 8 10 Next ›