中文
相关论文

相关论文: Using Wordle for Learning to Design and Compare St…

200 篇论文

Large Vision Language Models (LVLMs) have demonstrated remarkable abilities in understanding and reasoning about both visual and textual information. However, existing evaluation methods for LVLMs, primarily based on benchmarks like Visual…

计算机视觉与模式识别 · 计算机科学 2025-03-05 Xinyu Wang , Bohan Zhuang , Qi Wu

Large language models (LLMs) have demonstrated potential in reasoning tasks, but their performance on linguistics puzzles remains consistently poor. These puzzles, often derived from Linguistics Olympiad (LO) contests, provide a minimal…

This paper introduces the Word Synchronization Challenge, a novel benchmark to evaluate large language models (LLMs) in Human-Computer Interaction (HCI). This benchmark uses a dynamic game-like framework to test LLMs ability to mimic human…

人机交互 · 计算机科学 2026-01-15 Tanguy Cazalets , Joni Dambre

Solving crossword puzzles requires diverse reasoning capabilities, access to a vast amount of knowledge about language and the world, and the ability to satisfy the constraints imposed by the structure of the puzzle. In this work, we…

计算与语言 · 计算机科学 2022-05-24 Saurabh Kulshreshtha , Olga Kovaleva , Namrata Shivagunde , Anna Rumshisky

While advancements in NLP have significantly improved the performance of Large Language Models (LLMs) on tasks requiring vertical thinking, their lateral thinking capabilities remain under-explored and challenging to measure due to the…

计算与语言 · 计算机科学 2024-10-10 Qi Chen , Bowen Zhang , Gang Wang , Qi Wu

In this theoretical note we compare different types of computational models of word similarity and association in their ability to predict a set of about 900 rating data. Using regression and predictive modeling tools (neural net, decision…

计算与语言 · 计算机科学 2018-08-27 Arthur M. Jacobs , Annette Kinder

Mathematical models of interactions among rational agents have long been studied in game theory. However these interactions are often over a small set of discrete game actions which is very different from how humans communicate in natural…

计算与语言 · 计算机科学 2024-12-17 Ian Gemp , Roma Patel , Yoram Bachrach , Marc Lanctot , Vibhavari Dasagi , Luke Marris , Georgios Piliouras , Siqi Liu , Karl Tuyls

A password composition policy restricts the space of allowable passwords to eliminate weak passwords that are vulnerable to statistical guessing attacks. Usability studies have demonstrated that existing password composition policies can…

密码学与安全 · 计算机科学 2013-02-26 Jeremiah Blocki , Saranga Komanduri , Ariel Procaccia , Or Sheffet

Alignment has quickly become a default ingredient in LLM development, with techniques such as reinforcement learning from human feedback making models act safely, follow instructions, and perform ever-better on complex tasks. While these…

计算与语言 · 计算机科学 2025-09-16 Peter West , Christopher Potts

This paper explores the entertainment experience and learning experience in Scrabble. It proposes a new measure from the educational point of view, which we call learning coefficient, based on the balance between the learner's skill and the…

人工智能 · 计算机科学 2017-11-13 Kananat Suwanviwatana , Hiroyuki Iida

Words of estimative probability (WEP) are expressions of a statement's plausibility (probably, maybe, likely, doubt, likely, unlikely, impossible...). Multiple surveys demonstrate the agreement of human evaluators when assigning numerical…

计算与语言 · 计算机科学 2023-06-27 Damien Sileo , Marie-Francine Moens

We study textual autocomplete---the task of predicting a full sentence from a partial sentence---as a human-machine communication game. Specifically, we consider three competing goals for effective communication: use as few tokens as…

计算与语言 · 计算机科学 2019-11-19 Mina Lee , Tatsunori B. Hashimoto , Percy Liang

Reasoning is not just about solving problems -- it is also about evaluating which problems are worth solving at all. Evaluations of artificial intelligence (AI) systems primarily focused on problem solving, historically by studying how…

Children learning their first language face multiple problems of induction: how to learn the meanings of words, and how to build meaningful phrases from those words according to syntactic rules. We consider how children might solve these…

计算与语言 · 计算机科学 2018-05-15 Jon Gauthier , Roger Levy , Joshua B. Tenenbaum

LLMs are increasingly used in applications where they interact with humans and other agents. We propose to use behavioural game theory to study LLM's cooperation and coordination behaviour. We let different LLMs play finitely repeated…

计算与语言 · 计算机科学 2025-05-13 Elif Akata , Lion Schulz , Julian Coda-Forno , Seong Joon Oh , Matthias Bethge , Eric Schulz

Language models (LMs) are increasingly being studied as models of human language learners. Due to the nascency of the field, it is not well-established whether LMs exhibit similar learning dynamics to humans, and there are few direct…

计算与语言 · 计算机科学 2025-02-11 Filippo Ficarra , Ryan Cotterell , Alex Warstadt

Recent advances in reinforcement learning (RL) algorithms aim to enhance the performance of language models at scale. Yet, there is a noticeable absence of a cost-effective and standardized testbed tailored to evaluating and comparing these…

机器学习 · 计算机科学 2024-03-13 Yufeng Zhang , Liyu Chen , Boyi Liu , Yingxiang Yang , Qiwen Cui , Yunzhe Tao , Hongxia Yang

Modern Artificial Intelligence applications show great potential for language-related tasks that rely on next-word prediction. The current generation of Large Language Models (LLMs) have been linked to claims about human-like linguistic…

计算与语言 · 计算机科学 2024-09-05 Evelina Leivada , Gary Marcus , Fritz Günther , Elliot Murphy

The study presented here relies on the integrated use of different kinds of knowledge in order to improve first-guess accuracy in non-word context-sensitive correction for general unrestricted texts. State of the art spelling correction…

cmp-lg · 计算机科学 2007-05-23 E. Agirre , K. Gojenola , K. Sarasola

Humans can learn many novel tasks from a very small number (1--5) of demonstrations, in stark contrast to the data requirements of nearly tabula rasa deep learning methods. We propose an expressive class of policies, a strong but general…

人工智能 · 计算机科学 2019-11-19 Tom Silver , Kelsey R. Allen , Alex K. Lew , Leslie Pack Kaelbling , Josh Tenenbaum
‹ 上一页 1 8 9 10 下一页 ›