中文
相关论文

相关论文: An Analysis of Dataset Overlap on Winograd-Style T…

200 篇论文

While Large Language Models (LLMs) have showcased remarkable proficiency in reasoning, there is still a concern about hallucinations and unreliable reasoning issues due to semantic associations and superficial logical chains. To evaluate…

计算与语言 · 计算机科学 2024-10-17 Kaiqiao Han , Tianqing Fang , Zhaowei Wang , Yangqiu Song , Mark Steedman

The Winograd Schema Challenge - a set of twin sentences involving pronoun reference disambiguation that seem to require the use of commonsense knowledge - was proposed by Hector Levesque in 2011. By 2019, a number of AI systems, based on…

计算与语言 · 计算机科学 2023-01-24 Vid Kocijan , Ernest Davis , Thomas Lukasiewicz , Gary Marcus , Leora Morgenstern

Commonsense reasoning is one of the key problems in natural language processing, but the relative scarcity of labeled data holds back the progress for languages other than English. Pretrained cross-lingual models are a source of powerful…

计算与语言 · 计算机科学 2021-12-02 Alexey Tikhonov , Max Ryabinin

The Winograd Schema Challenge is both a commonsense reasoning and natural language understanding challenge, introduced as an alternative to the Turing test. A Winograd schema is a pair of sentences differing in one or two words with a…

计算与语言 · 计算机科学 2020-04-30 Vid Kocijan , Thomas Lukasiewicz , Ernest Davis , Gary Marcus , Leora Morgenstern

Continual learning (CL) has remained a significant challenge for deep neural networks as learning new tasks erases previously acquired knowledge, either partially or completely. Existing solutions often rely on experience rehearsal or full…

机器学习 · 计算机科学 2025-05-29 Prashant Bhat , Laurens Niesten , Elahe Arani , Bahram Zonooz

Self-supervision based on the information extracted from large knowledge graphs has been shown to improve the generalization of language models, in zero-shot evaluation on various downstream language reasoning tasks. Since these…

计算与语言 · 计算机科学 2022-05-24 Jiarui Zhang , Filip Ilievski , Kaixin Ma , Jonathan Francis , Alessandro Oltramari

Chinese Spelling Correction (CSC) commonly lacks large-scale high-quality corpora, due to the labor-intensive labeling of spelling errors in real-life human writing or typing scenarios. Two data augmentation methods are widely adopted: (1)…

计算与语言 · 计算机科学 2024-07-23 Dingyao Yu , Yang An , Wei Ye , Xiongfeng Xiao , Shaoguang Mao , Tao Ge , Shikun Zhang

Selectional Preference (SP) is a commonly observed language phenomenon and proved to be useful in many natural language processing tasks. To provide a better evaluation method for SP models, we introduce SP-10K, a large-scale evaluation set…

计算与语言 · 计算机科学 2019-06-06 Hongming Zhang , Hantian Ding , Yangqiu Song

Assisted by the availability of data and high performance computing, deep learning techniques have achieved breakthroughs and surpassed human performance empirically in difficult tasks, including object recognition, speech recognition, and…

机器学习 · 计算机科学 2019-01-23 Shaeke Salman , Xiuwen Liu

We propose a self-supervised method to solve Pronoun Disambiguation and Winograd Schema Challenge problems. Our approach exploits the characteristic structure of training corpora related to so-called "trigger" words, which are responsible…

计算与语言 · 计算机科学 2020-05-06 Tassilo Klein , Moin Nabi

Recent efforts in natural language processing (NLP) commonsense reasoning research have yielded a considerable number of new datasets and benchmarks. However, most of these datasets formulate commonsense reasoning challenges in artificial…

计算与语言 · 计算机科学 2023-10-25 Mete Ismayilzada , Debjit Paul , Syrielle Montariol , Mor Geva , Antoine Bosselut

Existing reasoning datasets saturate and fail to test abstract, multi-step problems, especially pathfinding and complex rule constraint satisfaction. We introduce SPaRC (Spatial Pathfinding Reasoning Challenge), a dataset of 1,000 2D grid…

人工智能 · 计算机科学 2025-09-22 Lars Benedikt Kaesberg , Jan Philip Wahle , Terry Ruas , Bela Gipp

In Word Sense Disambiguation (WSD), the predominant approach generally involves a supervised system trained on sense annotated corpora. The limited quantity of such corpora however restricts the coverage and the performance of these…

计算与语言 · 计算机科学 2018-11-05 Loïc Vial , Benjamin Lecouteux , Didier Schwab

We introduce PreCo, a large-scale English dataset for coreference resolution. The dataset is designed to embody the core challenges in coreference, such as entity representation, by alleviating the challenge of low overlap between training…

计算与语言 · 计算机科学 2018-10-24 Hong Chen , Zhenhua Fan , Hao Lu , Alan L. Yuille , Shu Rong

Commonsense reasoning is one of the important aspect of natural language understanding, with several benchmarks developed to evaluate it. However, only a few of these benchmarks are available in languages other than English. Developing…

计算与语言 · 计算机科学 2024-12-17 Phakphum Artkaew

We introduce the CRASS (counterfactual reasoning assessment) data set and benchmark utilizing questionized counterfactual conditionals as a novel and powerful tool to evaluate large language models. We present the data set design and…

计算与语言 · 计算机科学 2022-10-06 Jörg Frohberg , Frank Binder

Vision-Language Pre-training (VLP) models have shown remarkable performance on various downstream tasks. Their success heavily relies on the scale of pre-trained cross-modal datasets. However, the lack of large-scale datasets and benchmarks…

计算机视觉与模式识别 · 计算机科学 2022-09-30 Jiaxi Gu , Xiaojun Meng , Guansong Lu , Lu Hou , Minzhe Niu , Xiaodan Liang , Lewei Yao , Runhui Huang , Wei Zhang , Xin Jiang , Chunjing Xu , Hang Xu

Language models have been shown to perform better with an increase in scale on a wide variety of tasks via the in-context learning paradigm. In this paper, we investigate the hypothesis that the ability of a large language model to…

计算与语言 · 计算机科学 2023-08-17 Hritik Bansal , Karthik Gopalakrishnan , Saket Dingliwal , Sravan Bodapati , Katrin Kirchhoff , Dan Roth

When evaluating the performance of automatic speech recognition models, usually word error rate within a certain dataset is used. Special care must be taken in understanding the dataset in order to report realistic performance numbers. We…

计算与语言 · 计算机科学 2021-05-21 Aashish Agarwal , Torsten Zesch

Can we get existing language models and refine them for zero-shot commonsense reasoning? This paper presents an initial study exploring the feasibility of zero-shot commonsense reasoning for the Winograd Schema Challenge by formulating the…

计算与语言 · 计算机科学 2021-09-14 Tassilo Klein , Moin Nabi