中文
相关论文

相关论文: Local word statistics affect reading times indepen…

200 篇论文

A sharp tension exists about the nature of human language between two opposite parties: those who believe that statistical surface distributions, in particular using measures like surprisal, provide a better understanding of language…

计算与语言 · 计算机科学 2023-02-20 Matteo Greco , Andrea Cometa , Fiorenzo Artoni , Robert Frank , Andrea Moro

Language-model (LM) surprisal is widely used as a proxy for contextual predictability and has been reported to correlate with metaphor novelty judgments. However, surprisal is tightly intertwined with lexical frequency. We explore this…

计算与语言 · 计算机科学 2026-05-19 Omar Momen , Sina Zarrieß

Surprisal theory hypothesizes that the difficulty of human sentence processing increases linearly with surprisal, the negative log-probability of a word given its context. Computational psycholinguistics has tested this hypothesis using…

计算与语言 · 计算机科学 2026-04-21 Ryo Yoshida , Shinnosuke Isono , Taiga Someya , Yohei Oseki , Tatsuki Kuribayashi

We investigate the extent to which word surprisal can be used to predict a neural measure of human language processing difficulty - the N400. To do this, we use recurrent neural networks to calculate the surprisal of stimuli from previously…

计算与语言 · 计算机科学 2022-05-13 James A. Michaelov , Benjamin K. Bergen

We show that one can perform causal inference in a natural way for continuous-time scenarios using tools from stochastic analysis. This provides new alternatives to the positivity condition for inverse probability weighting. The probability…

统计理论 · 数学 2013-04-23 Kjetil Røysland

We focus on the statistics of word occurrences and of the waiting times between such occurrences in Blogs. Due to the heterogeneity of words' frequencies, the empirical analysis is performed by studying classes of "frequently-equivalent"…

信息论 · 计算机科学 2012-09-25 R. Lambiotte , M. Ausloos , M. Thelwall

There has been considerable interest in using surprisal from Transformer-based language models (LMs) as predictors of human sentence processing difficulty. Recent work has observed an inverse scaling relationship between Transformers'…

计算与语言 · 计算机科学 2026-02-04 Yi-Chien Lin , William Schuler

We explore which linguistic factors -- at the sentence and token level -- play an important role in influencing language model predictions, and investigate whether these are reflective of results found in humans and human corpora (Gries and…

计算与语言 · 计算机科学 2024-09-18 Jaap Jumelet , Willem Zuidema , Arabella Sinclair

Recent research analyzing the sensitivity of natural language understanding models to word-order perturbations has shown that neural models are surprisingly insensitive to the order of words. In this paper, we investigate this phenomenon by…

计算与语言 · 计算机科学 2022-04-01 Louis Clouatre , Prasanna Parthasarathi , Amal Zouaq , Sarath Chandar

In many applications of natural language processing (NLP) it is necessary to determine the likelihood of a given word combination. For example, a speech recognizer may need to determine which of the two word combinations ``eat a peach'' and…

计算与语言 · 计算机科学 2007-05-23 Ido Dagan , Lillian Lee , Fernando C. N. Pereira

Recent research has shown that static word embeddings can encode word frequency information. However, little has been studied about this phenomenon and its effects on downstream tasks. In the present work, we systematically study the…

计算与语言 · 计算机科学 2023-10-23 Francisco Valentini , Juan Cruz Sosa , Diego Fernandez Slezak , Edgar Altszyler

Under surprisal theory, linguistic representations affect processing difficulty only through the bottleneck of surprisal. Our best estimates of surprisal come from large language models, which have no explicit representation of structural…

计算与语言 · 计算机科学 2026-03-27 Amani Maina-Kilaas , Roger Levy

This paper investigates the use of word surprisal, a measure of the predictability of a word in a given context, as a feature to aid speech synthesis prosody. We explore how word surprisal extracted from large language models (LLMs)…

音频与语音处理 · 电气工程与系统科学 2023-06-19 Sofoklis Kakouros , Juraj Šimko , Martti Vainio , Antti Suni

In many applications of natural language processing it is necessary to determine the likelihood of a given word combination. For example, a speech recognizer may need to determine which of the two word combinations ``eat a peach'' and ``eat…

cmp-lg · 计算机科学 2008-02-03 Ido Dagan , Fernando Pereira , Lillian Lee

Studies across many disciplines have shown that lexical choice can affect audience perception. For example, how users describe themselves in a social media profile can affect their perceived socio-economic status. However, we lack general…

机器学习 · 计算机科学 2018-11-16 Zhao Wang , Aron Culotta

Isolated word meanings are inherently uncertain. This uncertainty reduces when they are combined and anchored in context. We propose that grammar compresses meaning uncertainty cross-linguistically, which is reflected in brain and…

Various measures of dispersion have been proposed to paint a fuller picture of a word's distribution in a corpus, but only little has been done to validate them externally. We evaluate a wide range of dispersion measures as predictors of…

计算与语言 · 计算机科学 2025-01-14 Adam Nohejl , Taro Watanabe

Recent theoretical advancement of information density in natural language has brought the following question on desk: To what degree does natural language exhibit periodicity pattern in its encoded information? We address this question by…

计算与语言 · 计算机科学 2026-04-27 Yulin Ou , Yu Wang , Yang Xu , Hendrik Buschmeier

In psycholinguistic modeling, surprisal from larger pre-trained language models has been shown to be a poorer predictor of naturalistic human reading times. However, it has been speculated that this may be due to data leakage that caused…

计算与语言 · 计算机科学 2025-06-03 Byung-Doh Oh , Hongao Zhu , William Schuler

Recent studies have shown that as Transformer-based language models become larger and are trained on very large amounts of data, the fit of their surprisal estimates to naturalistic human reading times degrades. The current work presents a…

计算与语言 · 计算机科学 2024-02-06 Byung-Doh Oh , Shisen Yue , William Schuler