English
Related papers

Related papers: Cross-Platform Evaluation of Reasoning Capabilitie…

200 papers

Large language models (LLMs) have made significant advances in complex reasoning tasks, yet they remain bottlenecked by two core challenges: architectural inefficiency due to reliance on Transformers, and a lack of structured fine-tuning…

Machine Learning · Computer Science 2025-05-29 Xueliang Zhao , Wei Wu , Lingpeng Kong

Large language models are increasingly applied to materials science reasoning, yet their behavior under physically structured distribution shifts remains poorly understood. We introduce SCALAR (Structural Consistency And Logic Across…

Machine Learning · Computer Science 2026-02-02 Can Polat , Erchin Serpedin , Mustafa Kurban , Hasan Kurban

Automated building facade inspection is a critical component of urban resilience and smart city maintenance. Traditionally, this field has relied on specialized discriminative models (e.g., YOLO, Mask R-CNN) that excel at pixel-level…

Computer Vision and Pattern Recognition · Computer Science 2026-03-23 Hui Zhong , Yichun Gao , Luyan Liu , Hai Yang , Wang Wang , Haowei Zhang , Xinhu Zheng

Outcome-only evaluation under-specifies whether an evaluated agent profile preserves the commitments required to solve a multi-turn task coherently. NeuroState-Bench is a human-calibrated benchmark that operationalizes commitment integrity…

Artificial Intelligence · Computer Science 2026-05-15 Xiao Jia

Large language models (LLMs) offer remarkable capabilities, yet their high inference costs restrict wider adoption. While increasing parameter counts improves accuracy, it also broadens the gap between state-of-the-art capabilities and…

This paper introduces MISS-QA, the first benchmark specifically designed to evaluate the ability of models to interpret schematic diagrams within scientific literature. MISS-QA comprises 1,500 expert-annotated examples over 465 scientific…

Computation and Language · Computer Science 2025-07-16 Yilun Zhao , Chengye Wang , Chuhan Li , Arman Cohan

When evaluating Large Language Models (LLMs) in question answering domains, it is common to ask the model to choose among a fixed set of choices (so-called multiple-choice question-answering, or MCQA). Although downstream tasks of interest…

Computation and Language · Computer Science 2025-10-03 Narun Raman , Taylor Lundy , Kevin Leyton-Brown

With the increasing use of large language models (LLMs), ensuring reliable performance in diverse, real-world environments is essential. Despite their remarkable achievements, LLMs often struggle with adversarial inputs, significantly…

Computation and Language · Computer Science 2024-06-18 Yuqing Wang , Yun Zhao

Large language models demonstrate strong reasoning capabilities through chain-of-thought prompting, but whether this reasoning quality transfers across languages remains underexplored. We introduce a human-validated framework to evaluate…

Computation and Language · Computer Science 2026-03-31 Anaelia Ovalle , Candace Ross , Sebastian Ruder , Adina Williams , Karen Ullrich , Mark Ibrahim , Levent Sagun

Cross-view spatial reasoning is essential for embodied AI, underpinning spatial understanding, mental simulation and planning in complex environments. Existing benchmarks primarily emphasize indoor or street settings, overlooking the unique…

Computer Vision and Pattern Recognition · Computer Science 2026-01-22 Haotian Xu , Yue Hu , Zhengqiu Zhu , Chen Gao , Ziyou Wang , Junreng Rao , Wenhao Lu , Weishi Li , Quanjun Yin , Yong Li

Infrastructure as a Service (IaaS) clouds have become the predominant underlying infrastructure for the operation of modern and smart technology. IaaS clouds have proven to be useful for multiple reasons such as reduced costs, increased…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-01-13 Nivedhitha Duggi , Masoud Rafiei , Mohsen Amini Salehi

Current document parsing methods advance primarily through model architecture innovation, while systematic engineering of training data remains underexplored. Yet state-of-the-art models spanning diverse architectures and parameter scales…

Embodied agents operating in the physical world must make decisions that are not only effective but also safe, spatially coherent, and grounded in context. While recent advances in large multimodal models (LMMs) have shown promising…

Computer Vision and Pattern Recognition · Computer Science 2025-09-23 Dinura Dissanayake , Ahmed Heakl , Omkar Thawakar , Noor Ahsan , Ritesh Thawkar , Ketan More , Jean Lahoud , Rao Anwer , Hisham Cholakkal , Ivan Laptev , Fahad Shahbaz Khan , Salman Khan

Recent generations of language models have introduced Large Reasoning Models (LRMs) that generate detailed thinking processes before providing answers. While these models demonstrate improved performance on reasoning benchmarks, their…

Artificial Intelligence · Computer Science 2025-11-21 Parshin Shojaee , Iman Mirzadeh , Keivan Alizadeh , Maxwell Horton , Samy Bengio , Mehrdad Farajtabar

We introduce LinAlg-Bench, a diagnostic benchmark evaluating 10 frontier large language models on structured linear algebra computation across a strict dimensional gradient of 3x3, 4x4, and 5x5 matrices. Spanning 9 task types and 660…

Artificial Intelligence · Computer Science 2026-05-19 Shradha Agarwal , Deepak Rajbhar , Tariq J

Multimodal large language models (MLLMs) promise enhanced reasoning by integrating diverse inputs such as text, vision, and audio. Yet cross-modal reasoning remains underexplored, with conflicting reports on whether added modalities help or…

Computation and Language · Computer Science 2026-05-01 Yucheng Wang , Yifan Hou , Aydin Javadov , Mubashara Akhtar , Mrinmaya Sachan

Numerical reasoning is often important to accurately understand the world. Recently, several format-specific datasets have been proposed, such as numerical reasoning in the settings of Natural Language Inference (NLI), Reading Comprehension…

Computation and Language · Computer Science 2020-05-19 Swaroop Mishra , Arindam Mitra , Neeraj Varshney , Bhavdeep Sachdeva , Chitta Baral

Foundation Models have demonstrated significant success across various domains in Artificial Intelligence (AI), yet their capabilities for brainwave modeling remain unclear. In this paper, we comprehensively evaluate current Large Brainwave…

Machine Learning · Computer Science 2025-11-26 Na Lee , Konstantinos Barmpas , Yannis Panagakis , Dimitrios Adamos , Nikolaos Laskaris , Stefanos Zafeiriou

Multimodal reasoning, which integrates language and visual cues into problem solving and decision making, is a fundamental aspect of human intelligence and a crucial step toward artificial general intelligence. However, the evaluation of…

Large language models represent the same reasoning in vastly different surface forms -- English prose, Python code, mathematical notation -- yet whether they share a common internal substrate across these symbolic systems remains unknown.…

Computation and Language · Computer Science 2026-05-12 Aojie Yuan , Zhiyuan Su
‹ Prev 1 3 4 5 6 7 10 Next ›