中文
相关论文

相关论文: Can reasoning models comprehend mathematical probl…

200 篇论文

Although mathematics is often considered culturally neutral, the way mathematical problems are presented can carry implicit cultural context. Existing benchmarks like GSM8K are predominantly rooted in Western norms, including names,…

计算与语言 · 计算机科学 2025-11-03 Aditya Tomar , Nihar Ranjan Sahoo , Ashish Mittal , Rudra Murthy , Pushpak Bhattacharyya

Learning representations of algorithms is an emerging area of machine learning, seeking to bridge concepts from neural networks with classical algorithms. Several important works have investigated whether neural networks can effectively…

Mathematical reasoning is a cornerstone of artificial general intelligence and a primary benchmark for evaluating the capabilities of Large Language Models (LLMs). While state-of-the-art models show promise, they often falter when faced…

计算与语言 · 计算机科学 2025-07-29 Yifan Hao , Fangning Chao , Yaqian Hao , Zhaojun Cui , Huan Bai , Haiyu Zhang , Yankai Liu , Chao Deng , Junlan Feng

Although large language models (LLMs) show promise in solving complex mathematical tasks, existing evaluation paradigms rely solely on a coarse measure of overall answer accuracy, which are insufficient for assessing their authentic…

人工智能 · 计算机科学 2025-06-06 Jiayu Liu , Zhenya Huang , Wei Dai , Cheng Cheng , Jinze Wu , Jing Sha , Song Li , Qi Liu , Shijin Wang , Enhong Chen

Designing systems that can reason across cultures requires that they are grounded in the norms of the contexts in which they operate. However, current research on developing computational models of social norms has primarily focused on…

计算与语言 · 计算机科学 2023-10-24 Sky CH-Wang , Arkadiy Saakyan , Oliver Li , Zhou Yu , Smaranda Muresan

Recent studies in natural language processing (NLP) have focused on modern languages and achieved state-of-the-art results in many tasks. Meanwhile, little attention has been paid to ancient texts and related tasks. Classical Chinese first…

计算与语言 · 计算机科学 2024-07-03 Hao Wang , Hirofumi Shimizu , Daisuke Kawahara

Holistically measuring societal biases of large language models is crucial for detecting and reducing ethical risks in highly capable AI models. In this work, we present a Chinese Bias Benchmark dataset that consists of over 100K questions…

计算与语言 · 计算机科学 2023-06-29 Yufei Huang , Deyi Xiong

Recent large language models (LLMs) achieve near-saturation accuracy on many established mathematical reasoning benchmarks, raising concerns about their ability to diagnose genuine reasoning competence. This saturation largely stems from…

We introduce FrontierMath, a benchmark of hundreds of original, exceptionally challenging mathematics problems crafted and vetted by expert mathematicians. The questions cover most major branches of modern mathematics -- from…

Pre-trained computational language models have recently made remarkable progress in harnessing the language abilities which were considered unique to humans. Their success has raised interest in whether these models represent and process…

计算与语言 · 计算机科学 2024-03-05 Yunhao Zhang , Xiaohan Zhang , Chong Li , Shaonan Wang , Chengqing Zong

In this paper we look at the ability of recent large language models (LLMs) at solving mathematical problems in combinatorics. We compare models LLaMA-2, LLaMA-3.1, GPT-4, and Mixtral against each other and against human pupils and…

计算与语言 · 计算机科学 2024-12-17 Andrii Nikolaiev , Yiannos Stathopoulos , Simone Teufel

Large Language Models (LLMs) are increasingly applied to creative domains, yet their performance in classical Chinese poetry generation and evaluation remains poorly understood. We propose a three-step evaluation framework that combines…

计算与语言 · 计算机科学 2026-04-24 Bolei Ma , Yina Yao , Anna-Carolina Haensch

Machine reading comprehension tasks require a machine reader to answer questions relevant to the given document. In this paper, we present the first free-form multiple-Choice Chinese machine reading Comprehension dataset (C^3), containing…

计算与语言 · 计算机科学 2019-12-18 Kai Sun , Dian Yu , Dong Yu , Claire Cardie

Mathematical problem-solving is a key field in artificial intelligence (AI) and a critical benchmark for evaluating the capabilities of large language models (LLMs). While extensive research has focused on mathematical problem-solving, most…

计算与语言 · 计算机科学 2025-01-03 Ziye Chen , Hao Qi

Practical dialog systems need to deal with various knowledge sources, noisy user expressions, and the shortage of annotated data. To better solve the above problems, we propose CGoDial, new challenging and comprehensive Chinese benchmark…

计算与语言 · 计算机科学 2022-11-22 Yinpei Dai , Wanwei He , Bowen Li , Yuchuan Wu , Zheng Cao , Zhongqi An , Jian Sun , Yongbin Li

Online education platforms have significantly transformed the dissemination of educational resources by providing a dynamic and digital infrastructure. With the further enhancement of this transformation, the advent of Large Language Models…

人工智能 · 计算机科学 2024-09-26 Qian-Wen Zhang , Haochen Wang , Fang Li , Siyu An , Lingfeng Qiao , Liangcai Gao , Di Yin , Xing Sun

In this work, we develop a pipeline for historical-psychological text analysis in classical Chinese. Humans have produced texts in various languages for thousands of years; however, most of the computational literature is focused on…

计算与语言 · 计算机科学 2025-04-17 Yuqi Chen , Sixuan Li , Ying Li , Mohammad Atari

Large Language Models (LLMs) have shown remarkable performance in various natural language processing tasks but face challenges in mathematical reasoning, where complex problem-solving requires both linguistic understanding and mathematical…

计算与语言 · 计算机科学 2025-03-20 Shuguang Chen , Guang Lin

Mathematical reasoning remains a challenging area for large language models (LLMs), prompting the development of math-specific LLMs such as LLEMMA, DeepSeekMath, and Qwen2-Math, among others. These models typically follow a two-stage…

计算与语言 · 计算机科学 2025-03-25 Zui Chen , Tianqiao Liu , Mi Tian , Qing Tong , Weiqi Luo , Zitao Liu

Effective financial reasoning demands not only textual understanding but also the ability to interpret complex visual data such as charts, tables, and trend graphs. This paper introduces a new benchmark designed to evaluate how well AI…

人工智能 · 计算机科学 2025-06-10 Shuangyan Deng , Haizhou Peng , Jiachen Xu , Chunhou Liu , Ciprian Doru Giurcuaneanu , Jiamou Liu