English
Related papers

Related papers: Phoenix-VL 1.5 Medium Technical Report

200 papers

With the increasing integration of visual and textual content in Social Networking Services (SNS), evaluating the multimodal capabilities of Large Language Models (LLMs) is crucial for enhancing user experience, content understanding, and…

Computation and Language · Computer Science 2025-12-16 Hongcheng Guo , Zheyong Xie , Shaosheng Cao , Boyang Wang , Weiting Liu , Anjie Le , Lei Li , Zhoujun Li

Multimodal large language models (MLLMs), building upon the foundation of powerful large language models (LLMs), have recently demonstrated exceptional capabilities in generating not only texts but also images given interleaved multimodal…

Computer Vision and Pattern Recognition · Computer Science 2023-11-30 Bohao Li , Yuying Ge , Yixiao Ge , Guangzhi Wang , Rui Wang , Ruimao Zhang , Ying Shan

Vision-language models have progressed rapidly, but Tibetan remains a severely underserved low-resource language due to the lack of reproducible training and evaluation infrastructure. To fill this gap, we introduce FTibSuite, a…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Guixian Xu , Yide Liang , Zeli Su , Xuexian Song , Ziyin Zhang , Yushuang Dong , Ting Zhang , Xu Han

Building a generalizable self-correction system is crucial for robots to recover from failures. Despite advancements in Multimodal Large Language Models (MLLMs) that empower robots with semantic reflection ability for failure, translating…

Robotics · Computer Science 2025-04-22 Wenke Xia , Ruoxuan Feng , Dong Wang , Di Hu

We present Florence-VL, a new family of multimodal large language models (MLLMs) with enriched visual representations produced by Florence-2, a generative vision foundation model. Unlike the widely used CLIP-style vision transformer trained…

Computer Vision and Pattern Recognition · Computer Science 2024-12-06 Jiuhai Chen , Jianwei Yang , Haiping Wu , Dianqi Li , Jianfeng Gao , Tianyi Zhou , Bin Xiao

Training monolingual language models for low and mid-resource languages is made challenging by limited and often inadequate pretraining data. In this study, we propose a novel model conversion strategy to address this issue, adapting…

Computation and Language · Computer Science 2023-10-06 François Remy , Pieter Delobelle , Bettina Berendt , Kris Demuynck , Thomas Demeester

Tokenization serves as a foundational step for Large Language Models (LLMs) to process text. In new domains or languages, the inefficiency of the tokenizer will slow down the training and generation of LLM. The mismatch in vocabulary also…

Computation and Language · Computer Science 2025-06-05 Chong Li , Jiajun Zhang , Chengqing Zong

The popularity of multimodal large language models (MLLMs) has triggered a recent surge in research efforts dedicated to evaluating these models. Nevertheless, existing evaluation studies of MLLMs primarily focus on the comprehension and…

Computation and Language · Computer Science 2023-10-16 Xiaocui Yang , Wenfang Wu , Shi Feng , Ming Wang , Daling Wang , Yang Li , Qi Sun , Yifei Zhang , Xiaoming Fu , Soujanya Poria

We present an ongoing initiative to provide open, very large, high-quality, and richly annotated textual datasets for almost 200 languages. At 30 trillion tokens, this is likely the largest generally available multilingual collection of LLM…

In this work, we introduce LokiLM, a 1.4B parameter large language model trained on 500B tokens. Our model performs strongly in natural language reasoning tasks and achieves state-of-the-art performance among models with 1.5B parameters or…

Computation and Language · Computer Science 2024-07-11 Justin Kiefel , Shrey Shah

Vision-and-language models (VLMs) have been increasingly explored in the medical domain, particularly following the success of CLIP in general domain. However, unlike the relatively straightforward pairing of 2D images and text, curating…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Ziyang Zhang , Yang Yu , Xulei Yang , Si Yong Yeo

We introduce two multilingual, multimodal foundation language models that power Apple Intelligence features across Apple devices and services: i a 3B-parameter on-device model optimized for Apple silicon through architectural innovations…

Machine Learning · Computer Science 2025-08-28 Ethan Li , Anders Boesen Lindbo Larsen , Chen Zhang , Xiyou Zhou , Jun Qin , Dian Ang Yap , Narendran Raghavan , Xuankai Chang , Margit Bowler , Eray Yildiz , John Peebles , Hannah Gillis Coleman , Matteo Ronchi , Peter Gray , Keen You , Anthony Spalvieri-Kruse , Ruoming Pang , Reed Li , Yuli Yang , Emad Soroush , Zhiyun Lu , Crystal Xiao , Rong Situ , Jordan Huffaker , David Griffiths , Zaid Ahmed , Peng Zhang , Daniel Parilla , Asaf Liberman , Jennifer Mallalieu , Parsa Mazaheri , Qibin Chen , Manjot Bilkhu , Aonan Zhang , Eric Wang , Dave Nelson , Michael FitzMaurice , Thomas Voice , Jeremy Liu , Josh Shaffer , Shiwen Zhao , Prasanth Yadla , Farzin Rasteh , Pengsheng Guo , Arsalan Farooq , Jeremy Snow , Stephen Murphy , Tao Lei , Minsik Cho , George Horrell , Sam Dodge , Lindsay Hislop , Sumeet Singh , Alex Dombrowski , Aiswarya Raghavan , Sasha Sirovica , Mandana Saebi , Faye Lao , Max Lam , TJ Lu , Zhaoyang Xu , Karanjeet Singh , Marc Kirchner , David Mizrahi , Rajat Arora , Haotian Zhang , Henry Mason , Lawrence Zhou , Yi Hua , Ankur Jain , Felix Bai , Joseph Astrauskas , Floris Weers , Josh Gardner , Mira Chiang , Yi Zhang , Pulkit Agrawal , Tony Sun , Quentin Keunebroek , Matthew Hopkins , Bugu Wu , Tao Jia , Chen Chen , Xingyu Zhou , Nanzhu Wang , Peng Liu , Ruixuan Hou , Rene Rauch , Yuan Gao , Afshin Dehghan , Jonathan Janke , Zirui Wang , Cha Chen , Xiaoyi Ren , Feng Nan , Josh Elman , Dong Yin , Yusuf Goren , Jeff Lai , Yiran Fei , Syd Evans , Muyang Yu , Guoli Yin , Yi Qin , Erin Feldman , Isha Garg , Aparna Rajamani , Karla Vega , Walker Cheng , TJ Collins , Hans Han , Raul Rea Menacho , Simon Yeung , Sophy Lee , Phani Mutyala , Ying-Chang Cheng , Zhe Gan , Sprite Chu , Justin Lazarow , Alessandro Pappalardo , Federico Scozzafava , Jing Lu , Erik Daxberger , Laurent Duchesne , Jen Liu , David Güera , Stefano Ligas , Mary Beth Kery , Brent Ramerth , Ciro Sannino , Marcin Eichner , Haoshuo Huang , Rui Qian , Moritz Schwarzer-Becker , David Riazati , Mingfei Gao , Bailin Wang , Jack Cackler , Yang Lu , Ransen Niu , John Dennison , Guillaume Klein , Jeffrey Bigham , Deepak Gopinath , Navid Shiee , Darren Botten , Guillaume Tartavel , Alex Guillen Garcia , Sam Xu , Victoria MönchJuan Haladjian , Zi-Yi Dou , Matthias Paulik , Adolfo Lopez Mendez , Zhen Li , Hong-You Chen , Chao Jia , Dhaval Doshi , Zhengdong Zhang , Raunak Manjani , Aaron Franklin , Zhile Ren , David Chen , Artsiom Peshko , Nandhitha Raghuram , Hans Hao , Jiulong Shan , Kavya Nerella , Ramsey Tantawi , Vivek Kumar , Saiwen Wang , Brycen Wershing , Bhuwan Dhingra , Dhruti Shah , Ob Adaranijo , Xin Zheng , Tait Madsen , Hadas Kotek , Chang Liu , Yin Xia , Hanli Li , Suma Jayaram , Yanchao Sun , Ahmed Fakhry , Vasileios Saveris , Dustin Withers , Yanghao Li , Alp Aygar , Andres Romero Mier Y Teran , Kaiwei Huang , Mark Lee , Xiujun Li , Yuhong Li , Tyler Johnson , Jay Tang , Joseph Yitan Cheng , Futang Peng , Andrew Walkingshaw , Lucas Guibert , Abhishek Sharma , Cheng Shen , Piotr Maj , Yasutaka Tanaka , You-Cyuan Jhang , Vivian Ma , Tommi Vehvilainen , Kelvin Zou , Jeff Nichols , Matthew Lei , David Qiu , Yihao Qian , Gokul Santhanam , Wentao Wu , Yena Han , Dominik Moritz , Haijing Fu , Mingze Xu , Vivek Rathod , Jian Liu , Louis D'hauwe , Qin Ba , Haitian Sun , Haoran Yan , Philipp Dufter , Anh Nguyen , Yihao Feng , Emma Wang , Keyu He , Rahul Nair , Sanskruti Shah , Jiarui Lu , Patrick Sonnenberg , Jeremy Warner , Yuanzhi Li , Bowen Pan , Ziyi Zhong , Joe Zhou , Sam Davarnia , Olli Saarikivi , Irina Belousova , Rachel Burger , Shang-Chen Wu , Di Feng , Bas Straathof , James Chou , Yuanyang Zhang , Marco Zuliani , Eduardo Jimenez , Abhishek Sundararajan , Xianzhi Du , Chang Lan , Nilesh Shahdadpuri , Peter Grasch , Sergiu Sima , Josh Newnham , Varsha Paidi , Jianyu Wang , Kaelen Haag , Alex Braunstein , Daniele Molinari , Richard Wei , Brenda Yang , Nicholas Lusskin , Joanna Arreaza-Taylor , Meng Cao , Nicholas Seidl , Simon Wang , Jiaming Hu , Yiping Ma , Mengyu Li , Kieran Liu , Hang Su , Sachin Ravi , Chong Wang , Xin Wang , Kevin Smith , Haoxuan You , Binazir Karimzadeh , Rui Li , Jinhao Lei , Wei Fang , Alec Doane , Sam Wiseman , Ismael Fernandez , Jane Li , Andrew Hansen , Javier Movellan , Christopher Neubauer , Hanzhi Zhou , Chris Chaney , Nazir Kamaldin , Valentin Wolf , Fernando Bermúdez-Medina , Joris Pelemans , Peter Fu , Howard Xing , Xiang Kong , Wayne Shan , Gabriel Jacoby-Cooper , Dongcai Shen , Tom Gunter , Guillaume Seguin , Fangping Shi , Shiyu Li , Yang Xu , Areeba Kamal , Dan Masi , Saptarshi Guha , Qi Zhu , Jenna Thibodeau , Changyuan Zhang , Rebecca Callahan , Charles Maalouf , Wilson Tsao , Boyue Li , Qingqing Cao , Naomy Sabo , Cheng Leong , Yi Wang , Anupama Mann Anupama , Colorado Reed , Kenneth Jung , Zhifeng Chen , Mohana Prasad Sathya Moorthy , Yifei He , Erik Hornberger , Devi Krishna , Senyu Tong , Michael , Lee , David Haldimann , Yang Zhao , Bowen Zhang , Chang Gao , Chris Bartels , Sushma Rao , Nathalie Tran , Simon Lehnerer , Co Giang , Patrick Dong , Junting Pan , Biyao Wang , Dongxu Li , Mehrdad Farajtabar , Dongseong Hwang , Grace Duanmu , Eshan Verma , Sujeeth Reddy , Qi Shan , Hongbin Gao , Nan Du , Pragnya Sridhar , Forrest Huang , Yingbo Wang , Nikhil Bhendawade , Diane Zhu , Sai Aitharaju , Fred Hohman , Lauren Gardiner , Chung-Cheng Chiu , Yinfei Yang , Alper Kokmen , Frank Chu , Ke Ye , Kaan Elgin , Oron Levy , John Park , Donald Zhang , Eldon Schoop , Nina Wenzel , Michael Booker , Hyunjik Kim , Chinguun Erdenebileg , Nan Dun , Eric Liang Yang , Priyal Chhatrapati , Vishaal Mahtani , Haiming Gang , Kohen Chia , Deepa Seshadri , Donghan Yu , Yan Meng , Kelsey Peterson , Zhen Yang , Yongqiang Wang , Carina Peng , Doug Kang , Anuva Agarwal , Albert Antony , Juan Lao Tebar , Albin Madappally Jose , Regan Poston , Andy De Wang , Gerard Casamayor , Elmira Amirloo , Violet Yao , Wojciech Kryscinski , Kun Duan , Lezhi L

We propose MindVL, a multimodal large language model (MLLMs) trained on Ascend NPUs. The training of state-of-the-art MLLMs is often confined to a limited set of hardware platforms and relies heavily on massive, undisclosed data recipes,…

Computer Vision and Pattern Recognition · Computer Science 2025-10-01 Feilong Chen , Yijiang Liu , Yi Huang , Hao Wang , Miren Tian , Ya-Qi Yu , Minghui Liao , Jihao Wu

Current pre-trained vison-language models (PVLMs) achieve excellent performance on a range of multi-modal datasets. Recent work has aimed at building multilingual models, and a range of novel multilingual multi-modal datasets have been…

Computation and Language · Computer Science 2023-10-25 Hanxu Hu , Frank Keller

Large language models (LLMs) excel in high-resource languages but face notable challenges in low-resource languages like Mongolian. This paper addresses these challenges by categorizing capabilities into language abilities (syntax and…

Computation and Language · Computer Science 2024-11-15 Mengyuan Zhang , Ruihui Wang , Bo Xia , Yuan Sun , Xiaobing Zhao

Recent advancements in multimodal fusion have witnessed the remarkable success of vision-language (VL) models, which excel in various multimodal applications such as image captioning and visual question answering. However, building VL…

Computer Vision and Pattern Recognition · Computer Science 2024-10-24 Zhiwei Hao , Jianyuan Guo , Li Shen , Yong Luo , Han Hu , Yonggang Wen

We present DeepSeek-VL, an open-source Vision-Language (VL) Model designed for real-world vision and language understanding applications. Our approach is structured around three key dimensions: We strive to ensure our data is diverse,…

Artificial Intelligence · Computer Science 2024-03-12 Haoyu Lu , Wen Liu , Bo Zhang , Bingxuan Wang , Kai Dong , Bo Liu , Jingxiang Sun , Tongzheng Ren , Zhuoshu Li , Hao Yang , Yaofeng Sun , Chengqi Deng , Hanwei Xu , Zhenda Xie , Chong Ruan

Multimodal Large Language Models (MLLMs) have endowed LLMs with the ability to perceive and understand multi-modal signals. However, most of the existing MLLMs mainly adopt vision encoders pretrained on coarsely aligned image-text pairs,…

Computer Vision and Pattern Recognition · Computer Science 2023-11-28 Gongwei Chen , Leyang Shen , Rui Shao , Xiang Deng , Liqiang Nie

Domain reweighting can improve sample efficiency and downstream generalization, but data-mixture optimization for multimodal midtraining remains largely unexplored. Current multimodal training recipes tune mixtures along a single dimension,…

Machine Learning · Computer Science 2026-04-17 Bingbing Wen , Sirajul Salekin , Feiyang Kang , Bill Howe , Lucy Lu Wang , Javier Movellan , Manjot Bilkhu

Visual Grounding (VG) aims to utilize given natural language queries to locate specific target objects within images. While current transformer-based approaches demonstrate strong localization performance in standard scene (i.e, scenarios…

Computer Vision and Pattern Recognition · Computer Science 2025-09-09 Jiangnan Xie , Xiaolong Zheng , Liang Zheng
‹ Prev 1 3 4 5 6 7 10 Next ›