English
Related papers

Related papers: Phoenix-VL 1.5 Medium Technical Report

200 papers

This report introduces the Qwen2 series, the latest addition to our large language models and large multimodal models. We release a comprehensive suite of foundational and instruction-tuned language models, encompassing a parameter range…

We introduce NVLM 1.0, a family of frontier-class multimodal large language models (LLMs) that achieve state-of-the-art results on vision-language tasks, rivaling the leading proprietary models (e.g., GPT-4o) and open-access models (e.g.,…

Computation and Language · Computer Science 2024-10-24 Wenliang Dai , Nayeon Lee , Boxin Wang , Zhuolin Yang , Zihan Liu , Jon Barker , Tuomas Rintamaki , Mohammad Shoeybi , Bryan Catanzaro , Wei Ping

Although Large Vision Language Models (LVLMs) have demonstrated impressive multimodal reasoning capabilities, their scalability and deployment are constrained by massive computational requirements. In particular, the massive amount of…

Machine Learning · Computer Science 2026-04-14 Surendra Pathak , Bo Han

Tabular data is a pervasive modality spanning a wide range of domains, and the inherent diversity poses a considerable challenge for deep learning. Recent advancements using transformer-based in-context learning have shown promise on…

Machine Learning · Computer Science 2024-06-11 Valentin Thomas , Junwei Ma , Rasa Hosseinzadeh , Keyvan Golestan , Guangwei Yu , Maksims Volkovs , Anthony Caterini

Evaluating the performance of Multi-modal Large Language Models (MLLMs), integrating both point cloud and language, presents significant challenges. The lack of a comprehensive assessment hampers determining whether these models truly…

Computer Vision and Pattern Recognition · Computer Science 2024-04-24 Junjie Zhang , Tianci Hu , Xiaoshui Huang , Yongshun Gong , Dan Zeng

We address a notable gap in Natural Language Processing (NLP) by introducing a collection of resources designed to improve Machine Translation (MT) for low-resource languages, with a specific focus on African languages. First, we introduce…

Computation and Language · Computer Science 2024-07-15 AbdelRahim Elmadany , Ife Adebara , Muhammad Abdul-Mageed

Large Language Models (LLMs) have shown promise in highly-specialized domains, however challenges are still present in aspects of accuracy and costs. These limitations restrict the usage of existing models in domain-specific tasks. While…

Computation and Language · Computer Science 2024-10-30 Iftach Arbel , Yehonathan Refael , Ofir Lindenbaum

The rapid development of large Vision-Language Models (VLMs) has led to impressive results on academic benchmarks, primarily in widely spoken languages. However, significant gaps remain in the ability of current VLMs to handle low-resource…

Large language models (LLMs) have demonstrated remarkable performance on a variety of natural language tasks based on just a few examples of natural language instructions, reducing the need for extensive feature engineering. However, most…

In this paper, we focus on monolithic Multimodal Large Language Models (MLLMs) that integrate visual encoding and language decoding into a single LLM. In particular, we identify that existing pre-training strategies for monolithic MLLMs…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Gen Luo , Xue Yang , Wenhan Dou , Zhaokai Wang , Jiawen Liu , Jifeng Dai , Yu Qiao , Xizhou Zhu

Mixture of Experts (MoE) models have emerged as a promising paradigm for scaling language models efficiently by activating only a subset of parameters for each input token. In this report, we present dots.llm1, a large-scale MoE model that…

This paper focuses on monolithic Multimodal Large Language Models (MLLMs), which integrate visual encoding and language decoding into a single model. Existing structures and pre-training strategies for monolithic MLLMs often suffer from…

Computer Vision and Pattern Recognition · Computer Science 2025-07-18 Gen Luo , Wenhan Dou , Wenhao Li , Zhaokai Wang , Xue Yang , Changyao Tian , Hao Li , Weiyun Wang , Wenhai Wang , Xizhou Zhu , Yu Qiao , Jifeng Dai

While text-to-image (T2I) generation models have achieved remarkable progress in recent years, existing evaluation methodologies for vision-language alignment still struggle with the fine-grained semantic matching. Current approaches based…

Computer Vision and Pattern Recognition · Computer Science 2025-04-11 Zijian Zhang , Xuhui Zheng , Xuecheng Wu , Chong Peng , Xuezhi Cao

The growth of social media, characterized by its multimodal nature, has led to the emergence of diverse phenomena and challenges, which calls for an effective approach to uniformly solve automated tasks. The powerful Large Vision Language…

Computation and Language · Computer Science 2024-10-11 Xinnong Zhang , Haoyu Kuang , Xinyi Mou , Hanjia Lyu , Kun Wu , Siming Chen , Jiebo Luo , Xuanjing Huang , Zhongyu Wei

Large Language Models (LLMs) have shown strong generalization across tasks in high-resource languages; however, their linguistic competence in low-resource and morphologically rich languages such as Tamil remains largely unexplored.…

Computation and Language · Computer Science 2025-11-18 Jeyarajalingam Varsha , Menan Velayuthan , Sumirtha Karunakaran , Rasan Nivethiga , Kengatharaiyer Sarveswaran

The increase in technological adoption worldwide comes with demands for novel tools to be used by the general population. Large Language Models (LLMs) provide a great opportunity in this respect, but their capabilities remain limited for…

Computation and Language · Computer Science 2025-10-13 Stefan Krsteski , Matea Tashkovska , Borjan Sazdov , Hristijan Gjoreski , Branislav Gerazov

Multimodal foundation models (MFMs) have revolutionized sequential recommender systems through advanced representation learning. While Parameter-efficient Fine-tuning (PEFT) is commonly used to adapt these models, studies often prioritize…

Information Retrieval · Computer Science 2025-09-15 Junchen Fu , Xuri Ge , Xin Xin , Alexandros Karatzoglou , Ioannis Arapakis , Kaiwen Zheng , Yongxin Ni , Joemon M. Jose

In this report, we introduce Vintern-1B, a reliable 1-billion-parameters multimodal large language model (MLLM) for Vietnamese language tasks. By integrating the Qwen2-0.5B-Instruct language model with the InternViT-300M-448px visual model,…

Machine Learning · Computer Science 2024-08-26 Khang T. Doan , Bao G. Huynh , Dung T. Hoang , Thuc D. Pham , Nhat H. Pham , Quan T. M. Nguyen , Bang Q. Vo , Suong N. Hoang

Recent strides in video generation have paved the way for unified audio-visual generation. In this work, we present Seedance 1.5 pro, a foundational model engineered specifically for native, joint audio-video generation. Leveraging a…

Computer Vision and Pattern Recognition · Computer Science 2025-12-24 Team Seedance , Heyi Chen , Siyan Chen , Xin Chen , Yanfei Chen , Ying Chen , Zhuo Chen , Feng Cheng , Tianheng Cheng , Xinqi Cheng , Xuyan Chi , Jian Cong , Jing Cui , Qinpeng Cui , Qide Dong , Junliang Fan , Jing Fang , Zetao Fang , Chengjian Feng , Han Feng , Mingyuan Gao , Yu Gao , Dong Guo , Qiushan Guo , Boyang Hao , Qingkai Hao , Bibo He , Qian He , Tuyen Hoang , Ruoqing Hu , Xi Hu , Weilin Huang , Zhaoyang Huang , Zhongyi Huang , Donglei Ji , Siqi Jiang , Wei Jiang , Yunpu Jiang , Zhuo Jiang , Ashley Kim , Jianan Kong , Zhichao Lai , Shanshan Lao , Yichong Leng , Ai Li , Feiya Li , Gen Li , Huixia Li , JiaShi Li , Liang Li , Ming Li , Shanshan Li , Tao Li , Xian Li , Xiaojie Li , Xiaoyang Li , Xingxing Li , Yameng Li , Yifu Li , Yiying Li , Chao Liang , Han Liang , Jianzhong Liang , Ying Liang , Zhiqiang Liang , Wang Liao , Yalin Liao , Heng Lin , Kengyu Lin , Shanchuan Lin , Xi Lin , Zhijie Lin , Feng Ling , Fangfang Liu , Gaohong Liu , Jiawei Liu , Jie Liu , Jihao Liu , Shouda Liu , Shu Liu , Sichao Liu , Songwei Liu , Xin Liu , Xue Liu , Yibo Liu , Zikun Liu , Zuxi Liu , Junlin Lyu , Lecheng Lyu , Qian Lyu , Han Mu , Xiaonan Nie , Jingzhe Ning , Xitong Pan , Yanghua Peng , Lianke Qin , Xueqiong Qu , Yuxi Ren , Kai Shen , Guang Shi , Lei Shi , Yan Song , Yinglong Song , Fan Sun , Li Sun , Renfei Sun , Yan Sun , Zeyu Sun , Wenjing Tang , Yaxue Tang , Zirui Tao , Feng Wang , Furui Wang , Jinran Wang , Junkai Wang , Ke Wang , Kexin Wang , Qingyi Wang , Rui Wang , Sen Wang , Shuai Wang , Tingru Wang , Weichen Wang , Xin Wang , Yanhui Wang , Yue Wang , Yuping Wang , Yuxuan Wang , Ziyu Wang , Guoqiang Wei , Wanru Wei , Di Wu , Guohong Wu , Hanjie Wu , Jian Wu , Jie Wu , Ruolan Wu , Xinglong Wu , Yonghui Wu , Ruiqi Xia , Liang Xiang , Fei Xiao , XueFeng Xiao , Pan Xie , Shuangyi Xie , Shuang Xu , Jinlan Xue , Shen Yan , Bangbang Yang , Ceyuan Yang , Jiaqi Yang , Runkai Yang , Tao Yang , Yang Yang , Yihang Yang , ZhiXian Yang , Ziyan Yang , Songting Yao , Yifan Yao , Zilyu Ye , Bowen Yu , Jian Yu , Chujie Yuan , Linxiao Yuan , Sichun Zeng , Weihong Zeng , Xuejiao Zeng , Yan Zeng , Chuntao Zhang , Heng Zhang , Jingjie Zhang , Kuo Zhang , Liang Zhang , Liying Zhang , Manlin Zhang , Ting Zhang , Weida Zhang , Xiaohe Zhang , Xinyan Zhang , Yan Zhang , Yuan Zhang , Zixiang Zhang , Fengxuan Zhao , Huating Zhao , Yang Zhao , Hao Zheng , Jianbin Zheng , Xiaozheng Zheng , Yangyang Zheng , Yijie Zheng , Jiexin Zhou , Jiahui Zhu , Kuan Zhu , Shenhan Zhu , Wenjia Zhu , Benhui Zou , Feilong Zuo

This paper describes Tencent AI Lab - Shanghai Jiao Tong University (TAL-SJTU) Low-Resource Translation systems for the WMT22 shared task. We participate in the general translation task on English$\Leftrightarrow$Livonian. Our system is…

Computation and Language · Computer Science 2022-10-18 Zhiwei He , Xing Wang , Zhaopeng Tu , Shuming Shi , Rui Wang