English
Related papers

Related papers: Exploring OCR-augmented Generation for Bilingual V…

200 papers

Many vision-language models (VLMs) that prove very effective at a range of multimodal task, build on CLIP-based vision encoders, which are known to have various limitations. We investigate the hypothesis that the strong language backbone in…

Computer Vision and Pattern Recognition · Computer Science 2025-09-22 Sho Takishita , Jay Gala , Abdelrahman Mohamed , Kentaro Inui , Yova Kementchedjhieva

Class-incremental learning requires a learning system to continually learn knowledge of new classes and meanwhile try to preserve previously learned knowledge of old classes. As current state-of-the-art methods based on Vision-Language…

Computer Vision and Pattern Recognition · Computer Science 2025-12-11 Jiantao Tan , Peixian Ma , Tong Yu , Wentao Zhang , Ruixuan Wang

Recent Large Vision-Language Models (LVLMs) have shown promising reasoning capabilities on text-rich images from charts, tables, and documents. However, the abundant text within such images may increase the model's sensitivity to language.…

Computer Vision and Pattern Recognition · Computer Science 2024-12-24 Xinmiao Yu , Xiaocheng Feng , Yun Li , Minghui Liao , Ya-Qi Yu , Xiachong Feng , Weihong Zhong , Ruihan Chen , Mengkang Hu , Jihao Wu , Dandan Tu , Duyu Tang , Bing Qin

While vision-language pre-trained models (VL-PTMs) have advanced multimodal research in recent years, their mastery in a few languages like English restricts their applicability in broader communities. To this end, there is an increasing…

Computer Vision and Pattern Recognition · Computer Science 2024-01-31 Bang Yang , Yong Dai , Xuxin Cheng , Yaowei Li , Asif Raza , Yuexian Zou

Open-domain question answering (ODQA) has emerged as a pivotal research spotlight in information systems. Existing methods follow two main paradigms to collect evidence: (1) The \textit{retrieve-then-read} paradigm retrieves pertinent…

Computation and Language · Computer Science 2024-03-11 Hongda Sun , Yuxuan Liu , Chengwei Wu , Haiyu Yan , Cheng Tai , Xin Gao , Shuo Shang , Rui Yan

Scalable Vector Graphics (SVG) offer a powerful format for representing visual designs as interpretable code. Recent advances in vision-language models (VLMs) have enabled high-quality SVG generation by framing the problem as a code…

Reading dense text and locating objects within images are fundamental abilities for Large Vision-Language Models (LVLMs) tasked with advanced jobs. Previous LVLMs, including superior proprietary models like GPT-4o, have struggled to excel…

Computer Vision and Pattern Recognition · Computer Science 2024-10-08 Ya-Qi Yu , Minghui Liao , Jiwen Zhang , Jihao Wu

Knowledge-based visual question answering is a very challenging and widely concerned task. Previous methods adopts the implicit knowledge in large language models (LLM) to achieve excellent results, but we argue that existing methods may…

Multimedia · Computer Science 2023-08-31 Yang Zhou , Pengfei Cao , Yubo Chen , Kang Liu , Jun Zhao

Retrieval-Augmented Generation (RAG) systems have shown promise in enhancing the performance of Large Language Models (LLMs). However, these systems face challenges in effectively integrating external knowledge with the LLM's internal…

Multimodal language generation, which leverages the synergy of language and vision, is a rapidly expanding field. However, existing vision-language models face challenges in tasks that require complex linguistic understanding. To address…

Computation and Language · Computer Science 2023-12-20 Jiwan Chung , Youngjae Yu

Large Language Models (LLMs) have revolutionized a wide range of domains such as natural language processing, computer vision, and multi-modal tasks due to their ability to comprehend context and perform logical reasoning. However, the…

Artificial Intelligence · Computer Science 2025-07-31 Haoyang Li , Yiming Li , Anxin Tian , Tianhao Tang , Zhanchao Xu , Xuejia Chen , Nicole Hu , Wei Dong , Qing Li , Lei Chen

Recent advancements in multimodal slow-thinking systems have demonstrated remarkable performance across various visual reasoning tasks. However, their capabilities in text-rich image reasoning tasks remain understudied due to the absence of…

Machine Learning · Computer Science 2026-05-27 Mingxin Huang , Yongxin Shi , Dezhi Peng , Songxuan Lai , Zecheng Xie , Lianwen Jin

Large Multimodal Models (LMMs) have become increasingly versatile, accompanied by impressive Optical Character Recognition (OCR) related capabilities. Existing OCR-related benchmarks emphasize evaluating LMMs' abilities of relatively simple…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Haibin He , Maoyuan Ye , Jing Zhang , Xiantao Cai , Juhua Liu , Bo Du , Dacheng Tao

Multimodal Large Language Models (MLLMs) have endowed LLMs with the ability to perceive and understand multi-modal signals. However, most of the existing MLLMs mainly adopt vision encoders pretrained on coarsely aligned image-text pairs,…

Computer Vision and Pattern Recognition · Computer Science 2023-11-28 Gongwei Chen , Leyang Shen , Rui Shao , Xiang Deng , Liqiang Nie

We introduce HyperCLOVA X, a family of large language models (LLMs) tailored to the Korean language and culture, along with competitive capabilities in English, math, and coding. HyperCLOVA X was trained on a balanced mix of Korean,…

Computation and Language · Computer Science 2024-04-16 Kang Min Yoo , Jaegeun Han , Sookyo In , Heewon Jeon , Jisu Jeong , Jaewook Kang , Hyunwook Kim , Kyung-Min Kim , Munhyong Kim , Sungju Kim , Donghyun Kwak , Hanock Kwak , Se Jung Kwon , Bado Lee , Dongsoo Lee , Gichang Lee , Jooho Lee , Baeseong Park , Seongjin Shin , Joonsang Yu , Seolki Baek , Sumin Byeon , Eungsup Cho , Dooseok Choe , Jeesung Han , Youngkyun Jin , Hyein Jun , Jaeseung Jung , Chanwoong Kim , Jinhong Kim , Jinuk Kim , Dokyeong Lee , Dongwook Park , Jeong Min Sohn , Sujung Han , Jiae Heo , Sungju Hong , Mina Jeon , Hyunhoon Jung , Jungeun Jung , Wangkyo Jung , Chungjoon Kim , Hyeri Kim , Jonghyun Kim , Min Young Kim , Soeun Lee , Joonhee Park , Jieun Shin , Sojin Yang , Jungsoon Yoon , Hwaran Lee , Sanghwan Bae , Jeehwan Cha , Karl Gylleus , Donghoon Ham , Mihak Hong , Youngki Hong , Yunki Hong , Dahyun Jang , Hyojun Jeon , Yujin Jeon , Yeji Jeong , Myunggeun Ji , Yeguk Jin , Chansong Jo , Shinyoung Joo , Seunghwan Jung , Adrian Jungmyung Kim , Byoung Hoon Kim , Hyomin Kim , Jungwhan Kim , Minkyoung Kim , Minseung Kim , Sungdong Kim , Yonghee Kim , Youngjun Kim , Youngkwan Kim , Donghyeon Ko , Dughyun Lee , Ha Young Lee , Jaehong Lee , Jieun Lee , Jonghyun Lee , Jongjin Lee , Min Young Lee , Yehbin Lee , Taehong Min , Yuri Min , Kiyoon Moon , Hyangnam Oh , Jaesun Park , Kyuyon Park , Younghun Park , Hanbae Seo , Seunghyun Seo , Mihyun Sim , Gyubin Son , Matt Yeo , Kyung Hoon Yeom , Wonjoon Yoo , Myungin You , Doheon Ahn , Homin Ahn , Joohee Ahn , Seongmin Ahn , Chanwoo An , Hyeryun An , Junho An , Sang-Min An , Boram Byun , Eunbin Byun , Jongho Cha , Minji Chang , Seunggyu Chang , Haesong Cho , Youngdo Cho , Dalnim Choi , Daseul Choi , Hyoseok Choi , Minseong Choi , Sangho Choi , Seongjae Choi , Wooyong Choi , Sewhan Chun , Dong Young Go , Chiheon Ham , Danbi Han , Jaemin Han , Moonyoung Hong , Sung Bum Hong , Dong-Hyun Hwang , Seongchan Hwang , Jinbae Im , Hyuk Jin Jang , Jaehyung Jang , Jaeni Jang , Sihyeon Jang , Sungwon Jang , Joonha Jeon , Daun Jeong , Joonhyun Jeong , Kyeongseok Jeong , Mini Jeong , Sol Jin , Hanbyeol Jo , Hanju Jo , Minjung Jo , Chaeyoon Jung , Hyungsik Jung , Jaeuk Jung , Ju Hwan Jung , Kwangsun Jung , Seungjae Jung , Soonwon Ka , Donghan Kang , Soyoung Kang , Taeho Kil , Areum Kim , Beomyoung Kim , Byeongwook Kim , Daehee Kim , Dong-Gyun Kim , Donggook Kim , Donghyun Kim , Euna Kim , Eunchul Kim , Geewook Kim , Gyu Ri Kim , Hanbyul Kim , Heesu Kim , Isaac Kim , Jeonghoon Kim , Jihye Kim , Joonghoon Kim , Minjae Kim , Minsub Kim , Pil Hwan Kim , Sammy Kim , Seokhun Kim , Seonghyeon Kim , Soojin Kim , Soong Kim , Soyoon Kim , Sunyoung Kim , Taeho Kim , Wonho Kim , Yoonsik Kim , You Jin Kim , Yuri Kim , Beomseok Kwon , Ohsung Kwon , Yoo-Hwan Kwon , Anna Lee , Byungwook Lee , Changho Lee , Daun Lee , Dongjae Lee , Ha-Ram Lee , Hodong Lee , Hwiyeong Lee , Hyunmi Lee , Injae Lee , Jaeung Lee , Jeongsang Lee , Jisoo Lee , Jongsoo Lee , Joongjae Lee , Juhan Lee , Jung Hyun Lee , Junghoon Lee , Junwoo Lee , Se Yun Lee , Sujin Lee , Sungjae Lee , Sungwoo Lee , Wonjae Lee , Zoo Hyun Lee , Jong Kun Lim , Kun Lim , Taemin Lim , Nuri Na , Jeongyeon Nam , Kyeong-Min Nam , Yeonseog Noh , Biro Oh , Jung-Sik Oh , Solgil Oh , Yeontaek Oh , Boyoun Park , Cheonbok Park , Dongju Park , Hyeonjin Park , Hyun Tae Park , Hyunjung Park , Jihye Park , Jooseok Park , Junghwan Park , Jungsoo Park , Miru Park , Sang Hee Park , Seunghyun Park , Soyoung Park , Taerim Park , Wonkyeong Park , Hyunjoon Ryu , Jeonghun Ryu , Nahyeon Ryu , Soonshin Seo , Suk Min Seo , Yoonjeong Shim , Kyuyong Shin , Wonkwang Shin , Hyun Sim , Woongseob Sim , Hyejin Soh , Bokyong Son , Hyunjun Son , Seulah Son , Chi-Yun Song , Chiyoung Song , Ka Yeon Song , Minchul Song , Seungmin Song , Jisung Wang , Yonggoo Yeo , Myeong Yeon Yi , Moon Bin Yim , Taehwan Yoo , Youngjoon Yoo , Sungmin Yoon , Young Jin Yoon , Hangyeol Yu , Ui Seon Yu , Xingdong Zuo , Jeongin Bae , Joungeun Bae , Hyunsoo Cho , Seonghyun Cho , Yongjin Cho , Taekyoon Choi , Yera Choi , Jiwan Chung , Zhenghui Han , Byeongho Heo , Euisuk Hong , Taebaek Hwang , Seonyeol Im , Sumin Jegal , Sumin Jeon , Yelim Jeong , Yonghyun Jeong , Can Jiang , Juyong Jiang , Jiho Jin , Ara Jo , Younghyun Jo , Hoyoun Jung , Juyoung Jung , Seunghyeong Kang , Dae Hee Kim , Ginam Kim , Hangyeol Kim , Heeseung Kim , Hyojin Kim , Hyojun Kim , Hyun-Ah Kim , Jeehye Kim , Jin-Hwa Kim , Jiseon Kim , Jonghak Kim , Jung Yoon Kim , Rak Yeong Kim , Seongjin Kim , Seoyoon Kim , Sewon Kim , Sooyoung Kim , Sukyoung Kim , Taeyong Kim , Naeun Ko , Bonseung Koo , Heeyoung Kwak , Haena Kwon , Youngjin Kwon , Boram Lee , Bruce W. Lee , Dagyeong Lee , Erin Lee , Euijin Lee , Ha Gyeong Lee , Hyojin Lee , Hyunjeong Lee , Jeeyoon Lee , Jeonghyun Lee , Jongheok Lee , Joonhyung Lee , Junhyuk Lee , Mingu Lee , Nayeon Lee , Sangkyu Lee , Se Young Lee , Seulgi Lee , Seung Jin Lee , Suhyeon Lee , Yeonjae Lee , Yesol Lee , Youngbeom Lee , Yujin Lee , Shaodong Li , Tianyu Liu , Seong-Eun Moon , Taehong Moon , Max-Lasse Nihlenramstroem , Wonseok Oh , Yuri Oh , Hongbeen Park , Hyekyung Park , Jaeho Park , Nohil Park , Sangjin Park , Jiwon Ryu , Miru Ryu , Simo Ryu , Ahreum Seo , Hee Seo , Kangdeok Seo , Jamin Shin , Seungyoun Shin , Heetae Sin , Jiangping Wang , Lei Wang , Ning Xiang , Longxiang Xiao , Jing Xu , Seonyeong Yi , Haanju Yoo , Haneul Yoo , Hwanhee Yoo , Liang Yu , Youngjae Yu , Weijie Yuan , Bo Zeng , Qian Zhou , Kyunghyun Cho , Jung-Woo Ha , Joonsuk Park , Jihyun Hwang , Hyoung Jo Kwon , Soonyong Kwon , Jungyeon Lee , Seungho Lee , Seonghyeon Lim , Hyunkyung Noh , Seungho Choi , Sang-Woo Lee , Jung Hwa Lim , Nako Sung

Recent Retrieval Augmented Generation (RAG) aims to enhance Large Language Models (LLMs) by incorporating extensive knowledge retrieved from external sources. However, such approach encounters some challenges: Firstly, the original queries…

Computation and Language · Computer Science 2024-10-10 Bolei He , Nuo Chen , Xinran He , Lingyong Yan , Zhenkai Wei , Jinchang Luo , Zhen-Hua Ling

This work presents a novel approach called oracle-checker scheme for evaluating the answer given by a generative large language model (LLM). Two types of checkers are presented. The first type of checker follows the idea of property…

Computation and Language · Computer Science 2024-05-07 Yueling Jenny Zeng , Li-C. Wang , Thomas Ibbetson

In our work, we explore the synergistic capabilities of pre-trained vision-and-language models (VLMs) and large language models (LLMs) on visual commonsense reasoning (VCR) problems. We find that VLMs and LLMs-based decision pipelines are…

Computer Vision and Pattern Recognition · Computer Science 2024-05-20 Kaiwen Zhou , Kwonjoon Lee , Teruhisa Misu , Xin Eric Wang

Large Language Models (LLMs) showcase impressive capabilities but encounter challenges like hallucination, outdated knowledge, and non-transparent, untraceable reasoning processes. Retrieval-Augmented Generation (RAG) has emerged as a…

Computation and Language · Computer Science 2024-03-28 Yunfan Gao , Yun Xiong , Xinyu Gao , Kangxiang Jia , Jinliu Pan , Yuxi Bi , Yi Dai , Jiawei Sun , Meng Wang , Haofen Wang

Text-rich document understanding (TDU) requires comprehensive analysis of documents containing substantial textual content and complex layouts. While Multimodal Large Language Models (MLLMs) have achieved fast progress in this domain,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-20 Wenhui Liao , Jiapeng Wang , Hongliang Li , Chengyu Wang , Jun Huang , Lianwen Jin