English
Related papers

Related papers: OpenEvents V1: Large-Scale Benchmark Dataset for M…

200 papers

With the ever-growing volume of online news feeds, event-based organization of news articles has many practical applications including better information navigation and the ability to view and analyze events as they develop. Automatically…

Information Retrieval · Computer Science 2021-03-09 Abdul Hameed Azeemi , Muhammad Hamza Sohail , Talha Zubair , Muaz Maqbool , Irfan Younas , Omair Shafiq

We present Seed1.5-VL, a vision-language foundation model designed to advance general-purpose multimodal understanding and reasoning. Seed1.5-VL is composed with a 532M-parameter vision encoder and a Mixture-of-Experts (MoE) LLM of 20B…

Computer Vision and Pattern Recognition · Computer Science 2025-05-13 Dong Guo , Faming Wu , Feida Zhu , Fuxing Leng , Guang Shi , Haobin Chen , Haoqi Fan , Jian Wang , Jianyu Jiang , Jiawei Wang , Jingji Chen , Jingjia Huang , Kang Lei , Liping Yuan , Lishu Luo , Pengfei Liu , Qinghao Ye , Rui Qian , Shen Yan , Shixiong Zhao , Shuai Peng , Shuangye Li , Sihang Yuan , Sijin Wu , Tianheng Cheng , Weiwei Liu , Wenqian Wang , Xianhan Zeng , Xiao Liu , Xiaobo Qin , Xiaohan Ding , Xiaojun Xiao , Xiaoying Zhang , Xuanwei Zhang , Xuehan Xiong , Yanghua Peng , Yangrui Chen , Yanwei Li , Yanxu Hu , Yi Lin , Yiyuan Hu , Yiyuan Zhang , Youbin Wu , Yu Li , Yudong Liu , Yue Ling , Yujia Qin , Zanbo Wang , Zhiwu He , Aoxue Zhang , Bairen Yi , Bencheng Liao , Can Huang , Can Zhang , Chaorui Deng , Chaoyi Deng , Cheng Lin , Cheng Yuan , Chenggang Li , Chenhui Gou , Chenwei Lou , Chengzhi Wei , Chundian Liu , Chunyuan Li , Deyao Zhu , Donghong Zhong , Feng Li , Feng Zhang , Gang Wu , Guodong Li , Guohong Xiao , Haibin Lin , Haihua Yang , Haoming Wang , Heng Ji , Hongxiang Hao , Hui Shen , Huixia Li , Jiahao Li , Jialong Wu , Jianhua Zhu , Jianpeng Jiao , Jiashi Feng , Jiaze Chen , Jianhui Duan , Jihao Liu , Jin Zeng , Jingqun Tang , Jingyu Sun , Joya Chen , Jun Long , Junda Feng , Junfeng Zhan , Junjie Fang , Junting Lu , Kai Hua , Kai Liu , Kai Shen , Kaiyuan Zhang , Ke Shen , Ke Wang , Keyu Pan , Kun Zhang , Kunchang Li , Lanxin Li , Lei Li , Lei Shi , Li Han , Liang Xiang , Liangqiang Chen , Lin Chen , Lin Li , Lin Yan , Liying Chi , Longxiang Liu , Mengfei Du , Mingxuan Wang , Ningxin Pan , Peibin Chen , Pengfei Chen , Pengfei Wu , Qingqing Yuan , Qingyao Shuai , Qiuyan Tao , Renjie Zheng , Renrui Zhang , Ru Zhang , Rui Wang , Rui Yang , Rui Zhao , Shaoqiang Xu , Shihao Liang , Shipeng Yan , Shu Zhong , Shuaishuai Cao , Shuangzhi Wu , Shufan Liu , Shuhan Chang , Songhua Cai , Tenglong Ao , Tianhao Yang , Tingting Zhang , Wanjun Zhong , Wei Jia , Wei Weng , Weihao Yu , Wenhao Huang , Wenjia Zhu , Wenli Yang , Wenzhi Wang , Xiang Long , XiangRui Yin , Xiao Li , Xiaolei Zhu , Xiaoying Jia , Xijin Zhang , Xin Liu , Xinchen Zhang , Xinyu Yang , Xiongcai Luo , Xiuli Chen , Xuantong Zhong , Xuefeng Xiao , Xujing Li , Yan Wu , Yawei Wen , Yifan Du , Yihao Zhang , Yining Ye , Yonghui Wu , Yu Liu , Yu Yue , Yufeng Zhou , Yufeng Yuan , Yuhang Xu , Yuhong Yang , Yun Zhang , Yunhao Fang , Yuntao Li , Yurui Ren , Yuwen Xiong , Zehua Hong , Zehua Wang , Zewei Sun , Zeyu Wang , Zhao Cai , Zhaoyue Zha , Zhecheng An , Zhehui Zhao , Zhengzhuo Xu , Zhipeng Chen , Zhiyong Wu , Zhuofan Zheng , Zihao Wang , Zilong Huang , Ziyu Zhu , Zuquan Song

Mainstream Scene Text Recognition (STR) algorithms are developed based on RGB cameras which are sensitive to challenging factors such as low illumination, motion blur, and cluttered backgrounds. In this paper, we propose to recognize the…

Computer Vision and Pattern Recognition · Computer Science 2025-02-14 Xiao Wang , Jingtao Jiang , Dong Li , Futian Wang , Lin Zhu , Yaowei Wang , Yongyong Tian , Jin Tang

News image captioning aims to produce journalistically informative descriptions by combining visual content with contextual cues from associated articles. Despite recent advances, existing methods struggle with three key challenges: (1)…

Computer Vision and Pattern Recognition · Computer Science 2025-11-27 Xiaoxing You , Qiang Huang , Lingyu Li , Chi Zhang , Xiaopeng Liu , Min Zhang , Jun Yu

Large vision-language models (VLMs) have made significant strides in 2D visual understanding tasks, sparking interest in extending these capabilities to 3D scene understanding. However, current 3D VLMs often struggle with robust reasoning…

Computer Vision and Pattern Recognition · Computer Science 2025-08-01 Ting Huang , Zeyu Zhang , Hao Tang

Video Temporal Grounding (VTG), the task of localizing video segments from text queries, struggles in open-world settings due to limited dataset scale and semantic diversity, causing performance gaps between common and rare concepts. To…

Computer Vision and Pattern Recognition · Computer Science 2026-04-29 Minghang Zheng , Zihao Yin , Yi Yang , Yuxin Peng , Yang Liu

We introduce a novel visual question answering (VQA) task in the context of autonomous driving, aiming to answer natural language questions based on street-view clues. Compared to traditional VQA tasks, VQA in autonomous driving scenario…

Computer Vision and Pattern Recognition · Computer Science 2024-02-21 Tianwen Qian , Jingjing Chen , Linhai Zhuo , Yang Jiao , Yu-Gang Jiang

3D scene understanding has been transformed by open-vocabulary language models that enable interaction via natural language. However, at present the evaluation of these representations is limited to datasets with closed-set semantics that…

Computer Vision and Pattern Recognition · Computer Science 2025-10-15 Christina Kassab , Sacha Morin , Martin Büchner , Matías Mattamala , Kumaraditya Gupta , Abhinav Valada , Liam Paull , Maurice Fallon

Rapid progress in video models has largely focused on visual quality, leaving their reasoning capabilities underexplored. Video reasoning grounds intelligence in spatiotemporally consistent visual environments that go beyond what text can…

Within the multimodal field, large vision-language models (LVLMs) have made significant progress due to their strong perception and reasoning capabilities in the visual and language systems. However, LVLMs are still plagued by the two…

Computer Vision and Pattern Recognition · Computer Science 2024-06-14 Sirui Cheng , Siyu Zhang , Jiayi Wu , Muchen Lan

Current image captioning systems lack the ability to link descriptive text to specific visual elements, making their outputs difficult to verify. While recent approaches offer some grounding capabilities, they cannot track object identities…

Computer Vision and Pattern Recognition · Computer Science 2025-06-27 Daniel A. P. Oliveira , Lourenço Teodoro , David Martins de Matos

Despite impressive advancements in video understanding, most efforts remain limited to coarse-grained or visual-only video tasks. However, real-world videos encompass omni-modal information (vision, audio, and speech) with a series of…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Tiantian Geng , Jinrui Zhang , Qingni Wang , Teng Wang , Jinming Duan , Feng Zheng

With the proliferation of imaging sensors, the volume of multi-modal imagery far exceeds the ability of human analysts to adequately consume and exploit it. Full motion video (FMV) possesses the extra challenge of containing large amounts…

Computer Vision and Pattern Recognition · Computer Science 2020-01-17 Marc Bosch , Joseph Nassar , Benjamin Ortiz , Brendan Lammers , David Lindenbaum , John Wahl , Robert Mangum , Margaret Smith

Models like OpenAI-o3 pioneer visual grounded reasoning by dynamically referencing visual regions, just like human "thinking with images". However, no benchmark exists to evaluate these capabilities holistically. To bridge this gap, we…

Computer Vision and Pattern Recognition · Computer Science 2026-03-06 Haochen Wang , Xiangtai Li , Zilong Huang , Anran Wang , Jiacong Wang , Tao Zhang , Jiani Zheng , Sule Bai , Zijian Kang , Jiashi Feng , Zhuochen Wang , Zhaoxiang Zhang

Retinal image analysis is crucial for diagnosing and treating eye diseases, yet generating accurate medical reports from images remains challenging due to variability in image quality and pathology, especially with limited labeled data.…

Computer Vision and Pattern Recognition · Computer Science 2024-12-24 Teja Krishna Cherukuri , Nagur Shareef Shaik , Jyostna Devi Bodapati , Dong Hye Ye

Human-scene vision-language tasks are increasingly prevalent in diverse social applications, yet recent advancements predominantly rely on models specifically tailored to individual tasks. Emerging research indicates that large…

Artificial Intelligence · Computer Science 2024-11-06 Dawei Dai , Xu Long , Li Yutang , Zhang Yuanhui , Shuyin Xia

Video sequences offer valuable temporal information, but existing large multimodal models (LMMs) fall short in understanding extremely long videos. Many works address this by reducing the number of visual tokens using visual resamplers.…

Computer Vision and Pattern Recognition · Computer Science 2024-07-02 Peiyuan Zhang , Kaichen Zhang , Bo Li , Guangtao Zeng , Jingkang Yang , Yuanhan Zhang , Ziyue Wang , Haoran Tan , Chunyuan Li , Ziwei Liu

Event cameras have recently shown promising capabilities in instantaneous motion estimation due to their robustness to low light and fast motions. However, computing wide-baseline correspondence between two arbitrary views remains a…

Computer Vision and Pattern Recognition · Computer Science 2026-04-22 Ruijun Zhang , Hang Su , Kostas Daniilidis , Ziyun Wang

We introduce the first dataset for sequential vision-to-language, and explore how this data may be used for the task of visual storytelling. The first release of this dataset, SIND v.1, includes 81,743 unique photos in 20,211 sequences,…

With the emergence of LLMs and their integration with other data modalities, multi-modal 3D perception attracts more attention due to its connectivity to the physical world and makes rapid progress. However, limited by existing datasets,…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Ruiyuan Lyu , Jingli Lin , Tai Wang , Shuai Yang , Xiaohan Mao , Yilun Chen , Runsen Xu , Haifeng Huang , Chenming Zhu , Dahua Lin , Jiangmiao Pang