English
Related papers

Related papers: LongBench v2: Towards Deeper Understanding and Rea…

200 papers

Large Language Models (LLMs) have demonstrated impressive capabilities across a range of natural language processing tasks. In particular, improvements in reasoning abilities and the expansion of context windows have opened new avenues for…

Databases · Computer Science 2025-06-12 Yeounoh Chung , Gaurav T. Kakkar , Yu Gan , Brenton Milne , Fatma Ozcan

Large Language Model (LLM)-based agents are increasingly deployed for complex, tool-based tasks where long-term memory is critical to driving actions. Existing benchmarks, however, primarily test a angent's ability to passively retrieve…

Computation and Language · Computer Science 2026-01-29 Yiting Shen , Kun Li , Wei Zhou , Songlin Hu

Multimodal Large Language Models (MLLMs) demonstrate impressive problem-solving abilities across a wide range of tasks and domains. However, their capacity for face understanding has not been systematically studied. To address this gap, we…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Kartik Narayan , Vibashan VS , Vishal M. Patel

Cognitive textual and visual reasoning tasks, including puzzles, series, and analogies, demand the ability to quickly reason, decipher, and evaluate patterns both textually and spatially. Due to extensive training on vast amounts of…

Computer Vision and Pattern Recognition · Computer Science 2025-04-02 Pranshu Pandya , Vatsal Gupta , Agney S Talwarr , Tushar Kataria , Dan Roth , Vivek Gupta

As large language models (LLMs) become integral to code-related tasks, a central question emerges: Do LLMs truly understand program semantics? We introduce EquiBench, a new benchmark for evaluating LLMs through equivalence checking, i.e.,…

Machine Learning · Computer Science 2025-09-23 Anjiang Wei , Jiannan Cao , Ran Li , Hongyu Chen , Yuhui Zhang , Ziheng Wang , Yuan Liu , Thiago S. F. X. Teixeira , Diyi Yang , Ke Wang , Alex Aiken

Most organizational data in this world are stored as documents, and visual retrieval plays a crucial role in unlocking the collective intelligence from all these documents. However, existing benchmarks focus on English-only document…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Jian Chen , Ming Li , Jihyung Kil , Chenguang Wang , Tong Yu , Ryan Rossi , Tianyi Zhou , Changyou Chen , Ruiyi Zhang

Though current long-context large language models (LLMs) have demonstrated impressive capacities in answering user questions based on extensive text, the lack of citations in their responses makes user verification difficult, leading to…

Computation and Language · Computer Science 2024-09-11 Jiajie Zhang , Yushi Bai , Xin Lv , Wanjun Gu , Danqing Liu , Minhao Zou , Shulin Cao , Lei Hou , Yuxiao Dong , Ling Feng , Juanzi Li

We present the first comprehensive, large-scale study of training long-context vision language models up to 344K context, targeting long-document visual question answering with measured transfer to long-context text. While several such…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Austin Veselka

Despite recent advances in video understanding, the capabilities of Large Video Language Models (LVLMs) to perform video-based causal reasoning remains underexplored, largely due to the absence of relevant and dedicated benchmarks for…

Computer Vision and Pattern Recognition · Computer Science 2025-05-14 Pritam Sarkar , Ali Etemad

Doctors and patients alike increasingly use Large Language Models (LLMs) to diagnose clinical cases. However, unlike domains such as math or coding, where correctness can be objectively defined by the final answer, medical diagnosis…

Computation and Language · Computer Science 2025-05-21 Kevin Wu , Eric Wu , Rahul Thapa , Kevin Wei , Angela Zhang , Arvind Suresh , Jacqueline J. Tao , Min Woo Sun , Alejandro Lozano , James Zou

Large language models (LLMs) have shown remarkable improvements in reasoning and many existing benchmarks have been addressed by models such as o1 and o3 either fully or partially. However, a majority of these benchmarks emphasize deductive…

Machine Learning · Computer Science 2025-05-15 Wenyue Hua , Tyler Wong , Sun Fei , Liangming Pan , Adam Jardine , William Yang Wang

Evaluating the abilities of large language models (LLMs) for tasks that require long-term memory and thus long-context reasoning, for example in conversational settings, is hampered by the existing benchmarks, which often lack narrative…

Computation and Language · Computer Science 2026-02-24 Mohammad Tavakoli , Alireza Salemi , Carrie Ye , Mohamed Abdalla , Hamed Zamani , J Ross Mitchell

Progress on training and architecture strategies has enabled LLMs with millions of tokens in context length. However, empirical evidence suggests that such long-context LLMs can consume far more text than they can reliably use. On the other…

Large Language Models (LLMs) have emerged as a powerful tool in advancing the Text-to-SQL task, significantly outperforming traditional methods.Nevertheless, as a nascent research field, there is still no consensus on the optimal prompt…

Computation and Language · Computer Science 2026-03-20 Bin Zhang , Yuxiao Ye , Guoqing Du , Xiaoru Hu , Zhishuai Li , Chi Harold Liu , Zhiwei Xu , Guoliang Fan , Rui Zhao , Ziyue Li , Hangyu Mao

Synthetic long-context LLM benchmarks (e.g., "needle-in-the-haystack") test only surface-level retrieval capabilities, but how well can long-context LLMs retrieve, synthesize, and reason over information across book-length inputs? We…

Computation and Language · Computer Science 2024-10-23 Marzena Karpinska , Katherine Thai , Kyle Lo , Tanya Goyal , Mohit Iyyer

Large language models (LLMs) have been widely evaluated on macro-scale geographic tasks, such as global factual recall, event summarization, and regional reasoning. Yet, their ability to handle hyper-local knowledge remains poorly…

Computation and Language · Computer Science 2025-11-19 Zihan Gao , Yifei Xu , Jacob Thebault-Spieker

This paper reviews the MARS2 2025 Challenge on Multimodal Reasoning. We aim to bring together different approaches in multimodal machine learning and LLMs via a large benchmark. We hope it better allows researchers to follow the…

Computer Vision and Pattern Recognition · Computer Science 2025-09-18 Peng Xu , Shengwu Xiong , Jiajun Zhang , Yaxiong Chen , Bowen Zhou , Chen Change Loy , David A. Clifton , Kyoung Mu Lee , Luc Van Gool , Ruiming He , Ruilin Yao , Xinwei Long , Jirui Huang , Kai Tian , Sa Yang , Yihua Shao , Jin Feng , Yue Zhong , Jiakai Zhou , Cheng Tang , Tianyu Zou , Yifang Zhang , Junming Liang , Guoyou Li , Zhaoxiang Wang , Qiang Zhou , Yichen Zhao , Shili Xiong , Hyeongjin Nam , Jaerin Lee , Jaeyoung Chung , JoonKyu Park , Junghun Oh , Kanggeon Lee , Wooseok Lee , Juneyoung Ro , Turghun Osman , Can Hu , Chaoyang Liao , Cheng Chen , Chengcheng Han , Chenhao Qiu , Chong Peng , Cong Xu , Dailin Li , Feiyu Wang , Feng Gao , Guibo Zhu , Guopeng Tang , Haibo Lu , Han Fang , Han Qi , Hanxiao Wu , Haobo Cheng , Hongbo Sun , Hongyao Chen , Huayong Hu , Hui Li , Jiaheng Ma , Jiang Yu , Jianing Wang , Jie Yang , Jing He , Jinglin Zhou , Jingxuan Li , Josef Kittler , Lihao Zheng , Linnan Zhao , Mengxi Jia , Muyang Yan , Nguyen Thanh Thien , Pu Luo , Qi Li , Shien Song , Shijie Dong , Shuai Shao , Shutao Li , Taofeng Xue , Tianyang Xu , Tianyi Gao , Tingting Li , Wei Zhang , Weiyang Su , Xiaodong Dong , Xiao-Jun Wu , Xiaopeng Zhou , Xin Chen , Xin Wei , Xinyi You , Xudong Kang , Xujie Zhou , Xusheng Liu , Yanan Wang , Yanbin Huang , Yang Liu , Yang Yang , Yanglin Deng , Yashu Kang , Ye Yuan , Yi Wen , Yicen Tian , Yilin Tao , Yin Tang , Yipeng Lin , Yiqing Wang , Yiting Xi , Yongkang Yu , Yumei Li , Yuxin Qin , Yuying Chen , Yuzhe Cen , Zhaofan Zou , Zhaohong Liu , Zhehao Shen , Zhenglin Du , Zhengyang Li , Zhenni Huang , Zhenwei Shao , Zhilong Song , Zhiyong Feng , Zhiyu Wang , Zhou Yu , Ziang Li , Zihan Zhai , Zijian Zhang , Ziyang Peng , Ziyun Xiao , Zongshu Li

Multimodal Large Language Models (MLLM) have made significant progress in the field of document analysis. Despite this, existing benchmarks typically focus only on extracting text and simple layout information, neglecting the complex…

Computer Vision and Pattern Recognition · Computer Science 2024-07-04 Lei Chen , Feng Yan , Yujie Zhong , Shaoxiang Chen , Zequn Jie , Lin Ma

Large language models (LLMs) achieve impressive performance on complex mathematical benchmarks yet sometimes fail on basic math reasoning while generating unnecessarily verbose responses. In this paper, we present LLMThinkBench, a…

Computation and Language · Computer Science 2026-04-24 Gaurav Srivastava , Aafiya Hussain , Sriram Srinivasan , Xuan Wang

Multi-Turn Long-Form Question Answering (MT-LFQA) is a key application paradigm of Large Language Models (LLMs) in knowledge-intensive domains. However, existing benchmarks are limited to single-turn dialogue, while multi-turn dialogue…

Computation and Language · Computer Science 2025-09-29 Junhao Chen , Yu Huang , Siyuan Li , Rui Yao , Hanqian Li , Hanyu Zhang , Jungang Li , Jian Chen , Bowen Wang , Xuming Hu