中文
相关论文

相关论文: JaWildText: A Benchmark for Vision-Language Models…

200 篇论文

Multimodal Large Language Models (MLLMs) have seen rapid advances in recent years and are now being applied to visual document understanding tasks. They are expected to process a wide range of document images across languages, including…

计算机视觉与模式识别 · 计算机科学 2025-11-20 Keito Sasagawa , Shuhei Kurita , Daisuke Kawahara

Reliable evaluation is essential for the development of vision-language models (VLMs). However, Japanese VQA benchmarks have undergone far less iterative refinement than their English counterparts. As a result, many existing benchmarks…

计算机视觉与模式识别 · 计算机科学 2026-04-07 Issa Sugiura , Koki Maeda , Shuhei Kurita , Yusuke Oda , Daisuke Kawahara , Naoaki Okazaki

This work addresses the scarcity of high-quality, large-scale resources for Japanese Vision-and-Language (V&L) modeling. We present a scalable and reproducible pipeline that integrates large-scale web collection with rigorous…

计算机视觉与模式识别 · 计算机科学 2025-12-02 Toshiki Katsube , Taiga Fukuhara , Kenichiro Ando , Yusuke Mukuta , Kohei Uehara , Tatsuya Harada

Developing vision-language models (VLMs) that generalize across diverse tasks requires large-scale training datasets with diverse content. In English, such datasets are typically constructed by aggregating and curating numerous existing…

计算机视觉与模式识别 · 计算机科学 2026-04-03 Issa Sugiura , Keito Sasagawa , Keisuke Nakao , Koki Maeda , Ziqi Yin , Zhishen Yang , Shuhei Kurita , Yusuke Oda , Ryoko Tokuhisa , Daisuke Kawahara , Naoaki Okazaki

Vision Language Models (VLMs) have undergone a rapid evolution, giving rise to significant advancements in the realm of multimodal understanding tasks. However, the majority of these models are trained and evaluated on English-centric…

计算机视觉与模式识别 · 计算机科学 2024-04-12 Yuichi Inoue , Kento Sasaki , Yuma Ochi , Kazuki Fujii , Kotaro Tanahashi , Yu Yamaguchi

Reading scene text, that is, text appearing in images, has numerous application areas, including assistive technology, search, and e-commerce. Although scene text recognition in English has advanced significantly and is often considered…

计算机视觉与模式识别 · 计算机科学 2026-04-13 Anik De , Abhirama Subramanyam Penamakuri , Rajeev Yadav , Aditya Rathore , Harshiv Shah , Devesh Sharma , Sagar Agarwal , Pravin Kumar , Anand Mishra

Current text detection datasets primarily target natural or document scenes, where text typically appear in regular font and shapes, monotonous colors, and orderly layouts. The text usually arranged along straight or curved lines. However,…

计算机视觉与模式识别 · 计算机科学 2025-10-10 Ziyi Dong , Yurui Zhang , Changmao Li , Naomi Rue Golding , Qing Long

To develop high-performing Visual Language Models (VLMs), it is essential to prepare multimodal resources, such as image-text pairs, interleaved data, and instruction data. While multimodal resources for English are abundant, there is a…

计算与语言 · 计算机科学 2024-10-31 Keito Sasagawa , Koki Maeda , Issa Sugiura , Shuhei Kurita , Naoaki Okazaki , Daisuke Kawahara

Multilingual document and scene text understanding plays an important role in applications such as search, finance, and public services. However, most existing benchmarks focus on high-resource languages and fail to evaluate models in…

计算与语言 · 计算机科学 2026-03-17 Pengfei Yue , Xingran Zhao , Juntao Chen , Peng Hou , Wang Longchao , Jianghang Lin , Shengchuan Zhang , Anxiang Zeng , Liujuan Cao

Vision-language models (VLMs) have demonstrated impressive generalization across multimodal tasks, yet most evaluation benchmarks remain Western-centric, leaving open questions about their performance in culturally diverse and multilingual…

计算机视觉与模式识别 · 计算机科学 2025-11-10 Ali Faraz , Akash , Shaharukh Khan , Raja Kolla , Akshat Patidar , Suranjan Goswami , Abhinav Ravi , Chandra Khatri , Shubham Agarwal

Detecting text in natural scenes remains challenging, particularly for diverse scripts and arbitrarily shaped instances where visual cues alone are often insufficient. Existing methods do not fully leverage semantic context. This paper…

计算机视觉与模式识别 · 计算机科学 2025-07-29 Mohammed-En-Nadhir Zighem , Abdenour Hadid

Large language models (LLMs) have increased interest in vision language models (VLMs), which process image-text pairs as input. Studies investigating the visual understanding ability of VLMs have been proposed, but such studies are still…

计算与语言 · 计算机科学 2024-06-25 Jesse Atuhurra , Iqra Ali , Tatsuya Hiraoka , Hidetaka Kamigaito , Tomoya Iwakura , Taro Watanabe

Vision-Language Models (VLMs) have achieved impressive performance in cross-modal understanding across textual and visual inputs, yet existing benchmarks predominantly focus on pure-text queries. In real-world scenarios, language also…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Qing'an Liu , Juntong Feng , Yuhao Wang , Xinzhe Han , Yujie Cheng , Yue Zhu , Haiwen Diao , Yunzhi Zhuge , Huchuan Lu

Document question answering is a task of question answering on given documents such as reports, slides, pamphlets, and websites, and it is a truly demanding task as paper and electronic forms of documents are so common in our society. This…

计算与语言 · 计算机科学 2024-03-29 Eri Onami , Shuhei Kurita , Taiki Miyanishi , Taro Watanabe

Visual language tracking (VLT) has emerged as a cutting-edge research area, harnessing linguistic data to enhance algorithms with multi-modal inputs and broadening the scope of traditional single object tracking (SOT) to encompass video…

计算机视觉与模式识别 · 计算机科学 2024-10-10 Xuchen Li , Shiyu Hu , Xiaokun Feng , Dailing Zhang , Meiqi Wu , Jing Zhang , Kaiqi Huang

Vision Language Models (VLMs) are pivotal for advancing perception in intelligent agents. Yet, evaluation of VLMs remains limited to predominantly English-centric benchmarks in which the image-text pairs comprise short texts. To evaluate…

计算与语言 · 计算机科学 2025-10-16 Jesse Atuhurra , Iqra Ali , Tomoya Iwakura , Hidetaka Kamigaito , Tatsuya Hiraoka

Recent developments in Japanese large language models (LLMs) primarily focus on general domains, with fewer advancements in Japanese biomedical LLMs. One obstacle is the absence of a comprehensive, large-scale benchmark for comparison.…

计算与语言 · 计算机科学 2024-09-23 Junfeng Jiang , Jiahao Huang , Akiko Aizawa

Is basic visual understanding really solved in state-of-the-art VLMs? We present VisualOverload, a slightly different visual question answering (VQA) benchmark comprising 2,720 question-answer pairs, with privately held ground-truth…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Paul Gavrikov , Wei Lin , M. Jehanzeb Mirza , Soumya Jahagirdar , Muhammad Huzaifa , Sivan Doveh , Serena Yeung-Levy , James Glass , Hilde Kuehne

One of the main objectives in developing large vision-language models (LVLMs) is to engineer systems that can assist humans with multimodal tasks, including interpreting descriptions of perceptual experiences. A central phenomenon in this…

计算与语言 · 计算机科学 2025-07-09 Amane Watahiki , Tomoki Doi , Taiga Shinozaki , Satoshi Nishida , Takuya Niikawa , Katsunori Miyahara , Hitomi Yanaka
‹ 上一页 1 2 3 10 下一页 ›