中文
相关论文

相关论文: EvoGrad: A Dynamic Take on the Winograd Schema Cha…

200 篇论文

The Winograd Schema Challenge (WSC) (Levesque, Davis, and Morgenstern 2011), a benchmark for commonsense reasoning, is a set of 273 expert-crafted pronoun resolution problems originally designed to be unsolvable for statistical models that…

计算与语言 · 计算机科学 2019-11-25 Keisuke Sakaguchi , Ronan Le Bras , Chandra Bhagavatula , Yejin Choi

The Winograd Schema Challenge (WSC) serves as a prominent benchmark for evaluating machine understanding. While Large Language Models (LLMs) excel at answering WSC questions, their ability to generate such questions remains less explored.…

计算与语言 · 计算机科学 2024-02-01 Pardis Sadat Zahraei , Ali Emami

While Large Language Models (LLMs) have showcased remarkable proficiency in reasoning, there is still a concern about hallucinations and unreliable reasoning issues due to semantic associations and superficial logical chains. To evaluate…

计算与语言 · 计算机科学 2024-10-17 Kaiqiao Han , Tianqing Fang , Zhaowei Wang , Yangqiu Song , Mark Steedman

Large-scale pretrained language models are the major driving force behind recent improvements in performance on the Winograd Schema Challenge, a widely employed test of common sense reasoning ability. We show, however, with a new diagnostic…

计算与语言 · 计算机科学 2020-05-08 Mostafa Abdou , Vinit Ravishankar , Maria Barrett , Yonatan Belinkov , Desmond Elliott , Anders Søgaard

Challenge sets such as the Winograd Schema Challenge (WSC) are used to benchmark systems' ability to resolve ambiguities in natural language. If one assumes as in existing work that solving a given challenge set is at least as difficult as…

计算与语言 · 计算机科学 2024-10-15 Ian Porada , Jackie Chi Kit Cheung

The Winograd Schema Challenge (WSC) dataset WSC273 and its inference counterpart WNLI are popular benchmarks for natural language understanding and commonsense reasoning. In this paper, we show that the performance of three language models…

计算与语言 · 计算机科学 2019-10-15 Vid Kocijan , Ana-Maria Cretu , Oana-Maria Camburu , Yordan Yordanov , Thomas Lukasiewicz

The Winograd Schema Challenge (WSC) and variants inspired by it have become important benchmarks for common-sense reasoning (CSR). Model performance on the WSC has quickly progressed from chance-level to near-human using neural language…

计算与语言 · 计算机科学 2020-11-16 Ali Emami , Adam Trischler , Kaheer Suleman , Jackie Chi Kit Cheung

Large Language Models (LLMs) have demonstrated remarkable success in tasks like the Winograd Schema Challenge (WSC), showcasing advanced textual common-sense reasoning. However, applying this reasoning to multimodal domains, where…

计算与语言 · 计算机科学 2024-06-04 Brendan Park , Madeline Janecek , Naser Ezzati-Jivan , Yifeng Li , Ali Emami

In the last decade, the Winograd Schema Challenge (WSC) has become a central aspect of the research community as a novel litmus test. Consequently, the WSC has spurred research interest because it can be seen as the means to understand…

人工智能 · 计算机科学 2023-09-07 Nicos Isaak , Loizos Michael

Winograd Schema Challenge (WSC) was proposed as an AI-hard problem in testing computers' intelligence on common sense representation and reasoning. This paper presents the new state-of-theart on WSC, achieving an accuracy of 71.1%. We…

计算与语言 · 计算机科学 2019-04-23 Yu-Ping Ruan , Xiaodan Zhu , Zhen-Hua Ling , Zhan Shi , Quan Liu , Si Wei

Performance on the Winograd Schema Challenge (WSC), a respected English commonsense reasoning benchmark, recently rocketed from chance accuracy to 89% on the SuperGLUE leaderboard, with relatively little corroborating evidence of a…

计算与语言 · 计算机科学 2020-10-09 Haokun Liu , William Huang , Dhara A. Mungra , Samuel R. Bowman

Ambiguities in natural language give rise to probability distributions over interpretations. The distributions are often over multiple ambiguous words at a time; a multiplicity which makes them a suitable topic for sheaf-theoretic models of…

计算与语言 · 计算机科学 2023-09-01 Kin Ian Lo , Mehrnoosh Sadrzadeh , Shane Mansfield

The Winograd Schema Challenge is both a commonsense reasoning and natural language understanding challenge, introduced as an alternative to the Turing test. A Winograd schema is a pair of sentences differing in one or two words with a…

计算与语言 · 计算机科学 2020-04-30 Vid Kocijan , Thomas Lukasiewicz , Ernest Davis , Gary Marcus , Leora Morgenstern

We introduce a new in-context learning paradigm to measure Large Language Models' (LLMs) ability to learn novel words during inference. In particular, we rewrite Winograd-style co-reference resolution problems by replacing the key concept…

计算与语言 · 计算机科学 2022-09-27 Julian Martin Eisenschlos , Jeremy R. Cole , Fangyu Liu , William W. Cohen

The Winograd Schema Challenge - a set of twin sentences involving pronoun reference disambiguation that seem to require the use of commonsense knowledge - was proposed by Hector Levesque in 2011. By 2019, a number of AI systems, based on…

计算与语言 · 计算机科学 2023-01-24 Vid Kocijan , Ernest Davis , Thomas Lukasiewicz , Gary Marcus , Leora Morgenstern

The Winograd Schema Challenge (WSC) is a common-sense reasoning task that requires background knowledge. In this paper, we contribute to tackling WSC in four ways. Firstly, we suggest a keyword method to define a restricted domain where…

计算与语言 · 计算机科学 2020-11-25 Suk Joon Hong , Brandon Bennett

In this paper, we present a localized and culturally adapted Estonian translation of the test set from the widely used commonsense reasoning benchmark, WinoGrande. We detail the translation and adaptation process carried out by translation…

计算与语言 · 计算机科学 2026-03-31 Marii Ojastu , Hele-Andra Kuulmets , Aleksei Dorkin , Marika Borovikova , Dage Särg , Kairit Sirts

Large Language Models (LLMs) have demonstrated promising capabilities in solving mathematical reasoning tasks, leveraging Chain-of-Thought (CoT) data as a vital component in guiding answer generation. Current paradigms typically generate…

计算与语言 · 计算机科学 2025-03-20 Honglin Lin , Zhuoshi Pan , Yu Li , Qizhi Pei , Xin Gao , Mengzhang Cai , Conghui He , Lijun Wu

Recent studies have significantly improved the state-of-the-art on common-sense reasoning (CSR) benchmarks like the Winograd Schema Challenge (WSC) and SWAG. The question we ask in this paper is whether improved performance on these…

机器学习 · 计算机科学 2021-09-27 Paul Trichelair , Ali Emami , Adam Trischler , Kaheer Suleman , Jackie Chi Kit Cheung

The Winograd Schema Challenge (WSC) is a natural language understanding task proposed as an alternative to the Turing test in 2011. In this work we attempt to solve WSC problems by reasoning with additional knowledge. By using an approach…

人工智能 · 计算机科学 2019-07-26 Arpit Sharma
‹ 上一页 1 2 3 10 下一页 ›