English
Related papers

Related papers: NVIDIA Nemotron Nano V2 VL

200 papers

We introduce Vision as LoRA (VoRA), a novel paradigm for transforming an LLM into an MLLM. Unlike prevalent MLLM architectures that rely on external vision modules for vision encoding, VoRA internalizes visual capabilities by integrating…

Computer Vision and Pattern Recognition · Computer Science 2025-03-27 Han Wang , Yongjie Ye , Bingru Li , Yuxiang Nie , Jinghui Lu , Jingqun Tang , Yanjie Wang , Can Huang

Vision-language pretraining (VLP) has been investigated to generalize across diverse downstream tasks for fundus image analysis. Although recent methods showcase promising achievements, they significantly rely on large-scale private…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Ruiqi Wu , Na Su , Chenran Zhang , Tengfei Ma , Tao Zhou , Zhiting Cui , Nianfeng Tang , Tianyu Mao , Yi Zhou , Wen Fan , Tianxing Wu , Shenqi Jing , Huazhu Fu

Vision-language models (VLMs) are highly effective but often underperform on specialized tasks; for example, Llava-1.5 struggles with chart and diagram understanding due to scarce task-specific training data. Existing training data, sourced…

Computer Vision and Pattern Recognition · Computer Science 2025-01-09 Siddharth Joshi , Besmira Nushi , Vidhisha Balachandran , Varun Chandrasekaran , Vibhav Vineet , Neel Joshi , Baharan Mirzasoleiman

The rapid success of Vision Large Language Models (VLLMs) often depends on the high-resolution images with abundant visual tokens, which hinders training and deployment efficiency. Current training-free visual token compression methods…

Computer Vision and Pattern Recognition · Computer Science 2025-02-27 Jianjian Li , Junquan Fan , Feng Tang , Gang Huang , Shitao Zhu , Songlin Liu , Nian Xie , Wulong Liu , Yong Liao

While Multimodal Large Language Models (MLLMs) excel at general vision-language tasks, visuospatial cognition - reasoning about spatial layouts, relations, and dynamics - remains a significant challenge. Existing models often lack the…

Computer Vision and Pattern Recognition · Computer Science 2025-09-10 Qi Feng

Large Language Models (LLMs) have achieved remarkable reliability and advanced capabilities through extended test-time reasoning. However, extending these capabilities to Multi-modal Large Language Models (MLLMs) remains a significant…

Computer Vision and Pattern Recognition · Computer Science 2026-03-20 Yuhao Dong , Zuyan Liu , Shulin Tian , Yongming Rao , Ziwei Liu

Vision-Language-Action (VLA) models have emerged as a promising framework for enabling generalist robots capable of perceiving, reasoning, and acting in the real world. These models usually build upon pretrained Vision-Language Models…

Robotics · Computer Science 2025-11-25 Tao Lin , Gen Li , Yilei Zhong , Yanwen Zou , Yuxin Du , Jiting Liu , Encheng Gu , Bo Zhao

With the advancement of RNN models with linear complexity, the quadratic complexity challenge of transformers has the potential to be overcome. Notably, the emerging Mamba-2 has demonstrated competitive performance, bridging the gap between…

Computer Vision and Pattern Recognition · Computer Science 2025-03-19 Yingyue Li , Bencheng Liao , Wenyu Liu , Xinggang Wang

Vision-and-Language Navigation (VLN) presents a complex challenge in embodied AI, requiring agents to interpret natural language instructions and navigate through visually rich, unfamiliar environments. Recent advances in large…

Robotics · Computer Science 2025-06-13 Yicheng Duan , Kaiyu tang

The field of vision-language models (VLMs), which take images and texts as inputs and output texts, is rapidly evolving and has yet to reach consensus on several key aspects of the development pipeline, including data, architecture, and…

Computer Vision and Pattern Recognition · Computer Science 2024-08-26 Hugo Laurençon , Andrés Marafioti , Victor Sanh , Léo Tronchon

Recently, Multi-modal Large Language Models (MLLMs) have shown remarkable effectiveness for multi-modal tasks due to their abilities to generate and understand cross-modal data. However, processing long sequences of visual tokens extracted…

Computer Vision and Pattern Recognition · Computer Science 2025-04-11 Haicheng Wang , Zhemeng Yu , Gabriele Spadaro , Chen Ju , Victor Quétu , Shuai Xiao , Enzo Tartaglione

Multimodal large language models (MLLMs) have attracted widespread interest and have rich applications. However, the inherent attention mechanism in its Transformer structure requires quadratic complexity and results in expensive…

Computer Vision and Pattern Recognition · Computer Science 2024-03-21 Yanyuan Qiao , Zheng Yu , Longteng Guo , Sihan Chen , Zijia Zhao , Mingzhen Sun , Qi Wu , Jing Liu

We present Nemotron 3 Nano 30B-A3B, a Mixture-of-Experts hybrid Mamba-Transformer language model. Nemotron 3 Nano was pretrained on 25 trillion text tokens, including more than 3 trillion new unique tokens over Nemotron 2, followed by…

Computation and Language · Computer Science 2025-12-25 NVIDIA , : , Aaron Blakeman , Aaron Grattafiori , Aarti Basant , Abhibha Gupta , Abhinav Khattar , Adi Renduchintala , Aditya Vavre , Akanksha Shukla , Akhiad Bercovich , Aleksander Ficek , Aleksandr Shaposhnikov , Alex Kondratenko , Alexander Bukharin , Alexandre Milesi , Ali Taghibakhshi , Alisa Liu , Amelia Barton , Ameya Sunil Mahabaleshwarkar , Amir Klein , Amit Zuker , Amnon Geifman , Amy Shen , Anahita Bhiwandiwalla , Andrew Tao , Ann Guan , Anubhav Mandarwal , Arham Mehta , Ashwath Aithal , Ashwin Poojary , Asif Ahamed , Asma Kuriparambil Thekkumpate , Ayush Dattagupta , Banghua Zhu , Bardiya Sadeghi , Barnaby Simkin , Ben Lanir , Benedikt Schifferer , Besmira Nushi , Bilal Kartal , Bita Darvish Rouhani , Boris Ginsburg , Brandon Norick , Brandon Soubasis , Branislav Kisacanin , Brian Yu , Bryan Catanzaro , Carlo del Mundo , Chantal Hwang , Charles Wang , Cheng-Ping Hsieh , Chenghao Zhang , Chenhan Yu , Chetan Mungekar , Chintan Patel , Chris Alexiuk , Christopher Parisien , Collin Neale , Damon Mosk-Aoyama , Dan Su , Dane Corneil , Daniel Afrimi , Daniel Rohrer , Daniel Serebrenik , Daria Gitman , Daria Levy , Darko Stosic , David Mosallanezhad , Deepak Narayanan , Dhruv Nathawani , Dima Rekesh , Dina Yared , Divyanshu Kakwani , Dong Ahn , Duncan Riach , Dusan Stosic , Edgar Minasyan , Edward Lin , Eileen Long , Eileen Peters Long , Elena Lantz , Ellie Evans , Elliott Ning , Eric Chung , Eric Harper , Eric Tramel , Erick Galinkin , Erik Pounds , Evan Briones , Evelina Bakhturina , Faisal Ladhak , Fay Wang , Fei Jia , Felipe Soares , Feng Chen , Ferenc Galko , Frankie Siino , Gal Hubara Agam , Ganesh Ajjanagadde , Gantavya Bhatt , Gargi Prasad , George Armstrong , Gerald Shen , Gorkem Batmaz , Grigor Nalbandyan , Haifeng Qian , Harsh Sharma , Hayley Ross , Helen Ngo , Herman Sahota , Hexin Wang , Himanshu Soni , Hiren Upadhyay , Huizi Mao , Huy C Nguyen , Huy Q Nguyen , Iain Cunningham , Ido Shahaf , Igor Gitman , Ilya Loshchilov , Ivan Moshkov , Izzy Putterman , Jan Kautz , Jane Polak Scowcroft , Jared Casper , Jatin Mitra , Jeffrey Glick , Jenny Chen , Jesse Oliver , Jian Zhang , Jiaqi Zeng , Jie Lou , Jimmy Zhang , Jining Huang , Joey Conway , Joey Guman , John Kamalu , Johnny Greco , Jonathan Cohen , Joseph Jennings , Joyjit Daw , Julien Veron Vialard , Junkeun Yi , Jupinder Parmar , Kai Xu , Kan Zhu , Kari Briski , Katherine Cheung , Katherine Luna , Keshav Santhanam , Kevin Shih , Kezhi Kong , Khushi Bhardwaj , Krishna C. Puvvada , Krzysztof Pawelec , Kumar Anik , Lawrence McAfee , Laya Sleiman , Leon Derczynski , Li Ding , Lucas Liebenwein , Luis Vega , Maanu Grover , Maarten Van Segbroeck , Maer Rodrigues de Melo , Makesh Narsimhan Sreedhar , Manoj Kilaru , Maor Ashkenazi , Marc Romeijn , Mark Cai , Markus Kliegl , Maryam Moosaei , Matvei Novikov , Mehrzad Samadi , Melissa Corpuz , Mengru Wang , Meredith Price , Michael Boone , Michael Evans , Miguel Martinez , Mike Chrzanowski , Mohammad Shoeybi , Mostofa Patwary , Nabin Mulepati , Natalie Hereth , Nave Assaf , Negar Habibi , Neta Zmora , Netanel Haber , Nicola Sessions , Nidhi Bhatia , Nikhil Jukar , Nikki Pope , Nikolai Ludwig , Nima Tajbakhsh , Nirmal Juluru , Oleksii Hrinchuk , Oleksii Kuchaiev , Olivier Delalleau , Oluwatobi Olabiyi , Omer Ullman Argov , Ouye Xie , Parth Chadha , Pasha Shamis , Pavlo Molchanov , Pawel Morkisz , Peter Dykas , Peter Jin , Pinky Xu , Piotr Januszewski , Pranav Prashant Thombre , Prasoon Varshney , Pritam Gundecha , Qing Miao , Rabeeh Karimi Mahabadi , Ran El-Yaniv , Ran Zilberstein , Rasoul Shafipour , Rich Harang , Rick Izzo , Rima Shahbazyan , Rishabh Garg , Ritika Borkar , Ritu Gala , Riyad Islam , Roger Waleffe , Rohit Watve , Roi Koren , Ruoxi Zhang , Russell J. Hewett , Ryan Prenger , Ryan Timbrook , Sadegh Mahdavi , Sahil Modi , Samuel Kriman , Sanjay Kariyappa , Sanjeev Satheesh , Saori Kaji , Satish Pasumarthi , Sean Narentharen , Sean Narenthiran , Seonmyeong Bak , Sergey Kashirsky , Seth Poulos , Shahar Mor , Shanmugam Ramasamy , Shantanu Acharya , Shaona Ghosh , Sharath Turuvekere Sreenivas , Shelby Thomas , Shiqing Fan , Shreya Gopal , Shrimai Prabhumoye , Shubham Pachori , Shubham Toshniwal , Shuoyang Ding , Siddharth Singh , Simeng Sun , Smita Ithape , Somshubra Majumdar , Soumye Singhal , Stefania Alborghetti , Stephen Ge , Sugam Dipak Devare , Sumeet Kumar Barua , Suseella Panguluri , Suyog Gupta , Sweta Priyadarshi , Syeda Nahida Akter , Tan Bui , Teodor-Dumitru Ene , Terry Kong , Thanh Do , Tijmen Blankevoort , Tom Balough , Tomer Asida , Tomer Bar Natan , Tugrul Konuk , Twinkle Vashishth , Udi Karpas , Ushnish De , Vahid Noorozi , Vahid Noroozi , Venkat Srinivasan , Venmugil Elango , Vijay Korthikanti , Vitaly Kurin , Vitaly Lavrukhin , Wanli Jiang , Wasi Uddin Ahmad , Wei Du , Wei Ping , Wenfei Zhou , Will Jennings , William Zhang , Wojciech Prazuch , Xiaowei Ren , Yashaswi Karnati , Yejin Choi , Yev Meyer , Yi-Fu Wu , Yian Zhang , Ying Lin , Yonatan Geifman , Yonggan Fu , Yoshi Subara , Yoshi Suhara , Yubo Gao , Zach Moshe , Zhen Dong , Zihan Liu , Zijia Chen , Zijie Yan

Embedding models have been crucial in enabling various downstream tasks such as semantic similarity, information retrieval, and clustering. Recently, there has been a surge of interest in developing universal text embedding models that can…

Computer Vision and Pattern Recognition · Computer Science 2025-01-03 Ziyan Jiang , Rui Meng , Xinyi Yang , Semih Yavuz , Yingbo Zhou , Wenhu Chen

Vision-Language Models (VLMs) have shown strong performance in tasks like visual question answering and multimodal text generation, but their effectiveness in scientific domains such as materials science remains limited. While some machine…

Machine Learning · Computer Science 2025-11-11 An Vuong , Minh-Hao Van , Prateek Verma , Chen Zhao , Xintao Wu

Existing Multimodal Large Language Models (MLLMs) suffer from increased inference costs due to the additional vision tokens introduced by image inputs. In this work, we propose Visual Consistency Learning (ViCO), a novel training algorithm…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Long Cui , Weiyun Wang , Jie Shao , Zichen Wen , Gen Luo , Linfeng Zhang , Yanting Zhang , Yu Qiao , Wenhai Wang

The success of Vision Language Models (VLMs) on various vision-language tasks heavily relies on pre-training with large scale web-crawled datasets. However, the noisy and incomplete nature of web data makes dataset scale crucial for…

Computer Vision and Pattern Recognition · Computer Science 2024-09-25 Yiyi Tao , Zhuoyue Wang , Hang Zhang , Lun Wang

We introduce VARCO-VISION-2.0, an open-weight bilingual vision-language model (VLM) for Korean and English with improved capabilities compared to the previous model VARCO-VISION-14B. The model supports multi-image understanding for complex…

Computer Vision and Pattern Recognition · Computer Science 2025-09-17 Young-rok Cha , Jeongho Ju , SunYoung Park , Jong-Hyeon Lee , Younghyun Yu , Youngjune Kim

Incorporating multiple modalities into large language models (LLMs) is a powerful way to enhance their understanding of non-textual data, enabling them to perform multimodal tasks. Vision language models (VLMs) form the fastest growing…

Machine Learning · Computer Science 2025-02-04 Shiqi He , Insu Jang , Mosharaf Chowdhury

This study evaluates the effectiveness of Vision Language Models (VLMs) in representing and utilizing multimodal content for fact-checking. To be more specific, we investigate whether incorporating multimodal content improves performance…

Computation and Language · Computer Science 2024-12-09 Recep Firat Cekinel , Pinar Karagoz , Cagri Coltekin
‹ Prev 1 3 4 5 6 7 10 Next ›