English
Related papers

Related papers: GIM: Evaluating models via tasks that integrate mu…

200 papers

Recent advances in large language models (LLMs) have demonstrated impressive reasoning capacities that mirror human-like thinking. However, whether LLMs possess genuine fluid intelligence (i.e., the ability to reason abstractly and…

Artificial Intelligence · Computer Science 2025-09-30 Yue Yang , MingKang Chen , Qihua Liu , Mengkang Hu , Qiguang Chen , Gengrui Zhang , Shuyue Hu , Guangtao Zhai , Yu Qiao , Yu Wang , Wenqi Shao , Ping Luo

Large Language Models (LLMs) show promise as planners for embodied AI, but their stochastic nature lacks formal reasoning, preventing strict safety guarantees for physical deployment. Current approaches often rely on unreliable LLMs for…

Artificial Intelligence · Computer Science 2026-04-30 Feiyu Wu , Xu Zheng , Yue Qu , Zhuocheng Wang , Zicheng Feng , Hui Li

To advance the evaluation of multimodal math reasoning in large multimodal models (LMMs), this paper introduces a novel benchmark, MM-MATH. MM-MATH consists of 5,929 open-ended middle school math problems with visual contexts, with…

Computation and Language · Computer Science 2024-07-03 Kai Sun , Yushi Bai , Ji Qi , Lei Hou , Juanzi Li

Accuracy-based evaluation of Large Language Models (LLMs) measures benchmark-specific performance rather than underlying medical competency: it treats all questions as equally informative, conflates model ability with item characteristics,…

Computation and Language · Computer Science 2026-04-07 Zhimeng Luo , Lixin Wu , Adam Frisch , Daqing He

Large Language Models (LLMs) are increasingly being adopted as tools for learning; however, most tools remain text-only, limiting their usefulness for domains where visualizations are essential, such as mathematics. Recent work shows that…

Artificial Intelligence · Computer Science 2025-11-12 Vishal Kumar , Shubhra Mishra , Rebecca Hao , Rizwaan Malik , David Broman , Dorottya Demszky

Providing timely, consistent, and high-quality feedback in large-scale higher education courses remains a persistent challenge, often constrained by instructor workload and resource limitations. This study presents an LLM-powered, agentic…

Computers and Society · Computer Science 2026-01-13 Reza Vatankhah Barenji , Nazila Salimi , Sina Khoshgoftar

Public health reasoning requires population level inference grounded in scientific evidence, expert consensus, and safety constraints. However, it remains underexplored as a structured machine learning problem with limited supervised…

Computation and Language · Computer Science 2026-05-27 Zhaokun Yan , Shan Xu , Wuzheng Dong , Zhaohan Liu , Lijie Feng , Chengxiao Dai , Chen Tianqi , Binfan Liu , Yunpu Ma , Wenting Wei , Yingting Li , Yi Zhang , Tongning Wu

We present GLM-4.5, an open-source Mixture-of-Experts (MoE) large language model with 355B total parameters and 32B activated parameters, featuring a hybrid reasoning method that supports both thinking and direct response modes. Through…

Computation and Language · Computer Science 2025-08-11 5 Team , Aohan Zeng , Xin Lv , Qinkai Zheng , Zhenyu Hou , Bin Chen , Chengxing Xie , Cunxiang Wang , Da Yin , Hao Zeng , Jiajie Zhang , Kedong Wang , Lucen Zhong , Mingdao Liu , Rui Lu , Shulin Cao , Xiaohan Zhang , Xuancheng Huang , Yao Wei , Yean Cheng , Yifan An , Yilin Niu , Yuanhao Wen , Yushi Bai , Zhengxiao Du , Zihan Wang , Zilin Zhu , Bohan Zhang , Bosi Wen , Bowen Wu , Bowen Xu , Can Huang , Casey Zhao , Changpeng Cai , Chao Yu , Chen Li , Chendi Ge , Chenghua Huang , Chenhui Zhang , Chenxi Xu , Chenzheng Zhu , Chuang Li , Congfeng Yin , Daoyan Lin , Dayong Yang , Dazhi Jiang , Ding Ai , Erle Zhu , Fei Wang , Gengzheng Pan , Guo Wang , Hailong Sun , Haitao Li , Haiyang Li , Haiyi Hu , Hanyu Zhang , Hao Peng , Hao Tai , Haoke Zhang , Haoran Wang , Haoyu Yang , He Liu , He Zhao , Hongwei Liu , Hongxi Yan , Huan Liu , Huilong Chen , Ji Li , Jiajing Zhao , Jiamin Ren , Jian Jiao , Jiani Zhao , Jianyang Yan , Jiaqi Wang , Jiayi Gui , Jiayue Zhao , Jie Liu , Jijie Li , Jing Li , Jing Lu , Jingsen Wang , Jingwei Yuan , Jingxuan Li , Jingzhao Du , Jinhua Du , Jinxin Liu , Junkai Zhi , Junli Gao , Ke Wang , Lekang Yang , Liang Xu , Lin Fan , Lindong Wu , Lintao Ding , Lu Wang , Man Zhang , Minghao Li , Minghuan Xu , Mingming Zhao , Mingshu Zhai , Pengfan Du , Qian Dong , Shangde Lei , Shangqing Tu , Shangtong Yang , Shaoyou Lu , Shijie Li , Shuang Li , Shuang-Li , Shuxun Yang , Sibo Yi , Tianshu Yu , Wei Tian , Weihan Wang , Wenbo Yu , Weng Lam Tam , Wenjie Liang , Wentao Liu , Xiao Wang , Xiaohan Jia , Xiaotao Gu , Xiaoying Ling , Xin Wang , Xing Fan , Xingru Pan , Xinyuan Zhang , Xinze Zhang , Xiuqing Fu , Xunkai Zhang , Yabo Xu , Yandong Wu , Yida Lu , Yidong Wang , Yilin Zhou , Yiming Pan , Ying Zhang , Yingli Wang , Yingru Li , Yinpei Su , Yipeng Geng , Yitong Zhu , Yongkun Yang , Yuhang Li , Yuhao Wu , Yujiang Li , Yunan Liu , Yunqing Wang , Yuntao Li , Yuxuan Zhang , Zezhen Liu , Zhen Yang , Zhengda Zhou , Zhongpei Qiao , Zhuoer Feng , Zhuorui Liu , Zichen Zhang , Zihan Wang , Zijun Yao , Zikang Wang , Ziqiang Liu , Ziwei Chai , Zixuan Li , Zuodong Zhao , Wenguang Chen , Jidong Zhai , Bin Xu , Minlie Huang , Hongning Wang , Juanzi Li , Yuxiao Dong , Jie Tang

Health, Safety, and Environment (HSE) compliance assessment demands dynamic real-time decision-making under complicated regulations and complex human-machine-environment interactions. While large language models (LLMs) hold significant…

Computation and Language · Computer Science 2025-05-30 Jianwei Wang , Mengqi Wang , Yinsi Zhou , Zhenchang Xing , Qing Liu , Xiwei Xu , Wenjie Zhang , Liming Zhu

Qualitative analysis is critical to understanding human datasets in many social science disciplines. A central method in this process is inductive coding, where researchers identify and interpret codes directly from the datasets themselves.…

Computation and Language · Computer Science 2026-04-21 John Chen , Alexandros Lotsos , Sihan Cheng , Caiyi Wang , Lexie Zhao , Yanjia Zhang , Jessica Hullman , Bruce Sherin , Uri Wilensky , Michael Horn

Automatic Short Answer Grading (ASAG) with generative large language models (LLMs) has recently demonstrated strong performance without task-specific fine-tuning, while also enabling the generation of synthetic feedback for educational…

Computation and Language · Computer Science 2026-05-14 Longwei Cong , Sonja Hahn , Sebastian Gombert , Leon Camus , Hendrik Drachsler , Ulf Kroehne

We present GaRAGe, a large RAG benchmark with human-curated long-form answers and annotations of each grounding passage, allowing a fine-grained evaluation of whether LLMs can identify relevant grounding when generating RAG answers. Our…

Computation and Language · Computer Science 2025-06-10 Ionut-Teodor Sorodoc , Leonardo F. R. Ribeiro , Rexhina Blloshmi , Christopher Davis , Adrià de Gispert

Reasoning has emerged as the next major frontier for language models (LMs), with rapid advances from both academic and industrial labs. However, this progress often outpaces methodological rigor, with many evaluations relying on…

Machine Learning · Computer Science 2025-10-08 Andreas Hochlehnert , Hardik Bhatnagar , Vishaal Udandarao , Samuel Albanie , Ameya Prabhu , Matthias Bethge

Large Language Models (LLMs) exhibit impressive performance across various domains but still struggle with arithmetic reasoning tasks. Recent work shows the effectiveness of prompt design methods in enhancing reasoning capabilities.…

Computation and Language · Computer Science 2024-10-11 Wenting Tan , Dongxiao Chen , Jieting Xue , Zihao Wang , Taijie Chen

Advancement in Large Language Models (LLMs) reasoning capabilities enables them to solve scientific problems with enhanced efficacy. Thereby, a high-quality benchmark for comprehensive and appropriate assessment holds significance, while…

We introduce SLR, an end-to-end framework for systematic evaluation and training of Large Language Models (LLMs) via Scalable Logical Reasoning. Given a user's task specification, SLR automatically synthesizes (i) an instruction prompt for…

As Large Language Models (LLMs) achieve significant breakthroughs in complex reasoning tasks, evaluating their proficiency in science, technology, engineering, and mathematics (STEM) has become a primary method for measuring machine…

Computation and Language · Computer Science 2026-02-04 Xuzhao Li , Xuchen Li , Jian Zhao , Shiyu Hu

Large language models (LLMs) have advanced the field of artificial intelligence (AI) and are a powerful enabler for interactive systems. However, they still face challenges in long-term interactions that require adaptation towards the user…

Artificial Intelligence · Computer Science 2025-05-20 Rebecca Westhäußer , Frederik Berenz , Wolfgang Minker , Sebastian Zepf

All text-based language problems can be reduced to either generation or embedding. Current models only perform well at one or the other. We introduce generative representational instruction tuning (GRIT) whereby a large language model is…

Computation and Language · Computer Science 2025-03-04 Niklas Muennighoff , Hongjin Su , Liang Wang , Nan Yang , Furu Wei , Tao Yu , Amanpreet Singh , Douwe Kiela

Recent advancements in reasoning-enhanced large language models (LLMs), such as DeepSeek-R1 and OpenAI-o3, have demonstrated significant progress. However, their application in professional medical contexts remains underexplored,…

Computation and Language · Computer Science 2025-03-11 Pengcheng Qiu , Chaoyi Wu , Shuyu Liu , Weike Zhao , Zhuoxia Chen , Hongfei Gu , Chuanjin Peng , Ya Zhang , Yanfeng Wang , Weidi Xie