中文
相关论文

相关论文: Local word statistics affect reading times indepen…

200 篇论文

Recent approaches in causal inference have proposed estimating average causal effects that are local to some subpopulation, often for reasons of efficiency. These inferential targets are sometimes data-adaptive, in that they are dependent…

统计理论 · 数学 2016-02-08 Peter M. Aronow

Models trained to estimate word probabilities in context have become ubiquitous in natural language processing. How do these models use lexical cues in context to inform their word probabilities? To answer this question, we present a case…

计算与语言 · 计算机科学 2021-04-23 Kanishka Misra , Allyson Ettinger , Julia Taylor Rayz

This article addresses a modification of local time for stochastic processes, to be referred to as `natural local time'. It is prompted by theoretical developments arising in mathematical treatments of recent experiments and observations of…

In probabilistic approaches to classification and information extraction, one typically builds a statistical model of words under the assumption that future data will exhibit the same regularities as the training data. In many data sets,…

机器学习 · 计算机科学 2013-01-07 David Blei , J Andrew Bagnell , Andrew McCallum

The meaning of a slang term can vary in different communities. However, slang semantic variation is not well understood and under-explored in the natural language processing of slang. One existing view argues that slang semantic variation…

计算与语言 · 计算机科学 2022-11-11 Zhewei Sun , Yang Xu

The dependence of the frequency distributions due to multiple meanings of words in a text is investigated by deleting letters. By coding the words with fewer letters the number of meanings per coded word increases. This increase is measured…

计算与语言 · 计算机科学 2017-10-04 Xiaoyong Yan , Petter Minnhagen

We consider the classical problem of discrete distribution estimation using i.i.d. samples in a novel scenario where additional side information is available on the distribution. In large alphabet datasets such as text corpora, such side…

信息论 · 计算机科学 2026-01-19 Haricharan Balasundaram , Andrew Thangaraj

The distribution of frequency counts of distinct words by length in a language's vocabulary will be analyzed using two methods. The first, will look at the empirical distributions of several languages and derive a distribution that…

计算与语言 · 计算机科学 2012-07-17 Reginald D. Smith

Many legal cases require decisions about causality, responsibility or blame, and these may be based on statistical data. However, causal inferences from such data are beset by subtle conceptual and practical difficulties, and in general it…

统计理论 · 数学 2020-04-28 Philip Dawid , Monica Musio , Rossella Murtas

Given a random text over a finite alphabet, we study the frequencies at which fixed-length words occur as subsequences. As the data size grows, the joint distribution of word counts exhibits a rich asymptotic structure. We investigate all…

概率论 · 数学 2026-05-06 Chaim Even-Zohar , Tsviqa Lakrec , Ran J. Tessler

An automatic word classification system has been designed which processes word unigram and bigram frequency statistics extracted from a corpus of natural language utterances. The system implements a binary top-down form of word clustering…

cmp-lg · 计算机科学 2016-08-31 John McMahon , F. J. Smith

Numerous works use word embedding-based metrics to quantify societal biases and stereotypes in texts. Recent studies have found that word embeddings can capture semantic similarity but may be affected by word frequency. In this work we…

计算与语言 · 计算机科学 2023-01-03 Francisco Valentini , Germán Rosati , Diego Fernandez Slezak , Edgar Altszyler

The outcomes of elections, product sales, and the structure of social connections are all determined by the choices individuals make when presented with a set of options, so understanding the factors that contribute to choice is crucial. Of…

机器学习 · 计算机科学 2020-11-09 Kiran Tomlinson , Austin R. Benson

We ask where, and under what conditions, dyslexic reading costs arise in a large-scale naturalistic reading dataset. Using eye-tracking aligned to word-level features (word length, frequency, and predictability), we model how each feature…

计算与语言 · 计算机科学 2025-10-29 Hugo Rydel-Johnston , Alex Kafkas

A new language model for speech recognition inspired by linguistic analysis is presented. The model develops hidden hierarchical structure incrementally and uses it to extract meaningful information from the word history - thus enabling the…

计算与语言 · 计算机科学 2007-05-23 Ciprian Chelba , Frederick Jelinek

Semantic feature models have become a popular tool for prediction and interpretation of fMRI data. In particular, prior work has shown that differences in the fMRI patterns in sentence reading can be explained by context-dependent changes…

计算与语言 · 计算机科学 2021-01-14 N. Aguirre-Celis , R. Miikkulainen

Novel metaphor comprehension involves complex semantic processes and linguistic creativity, making it an interesting task for studying language models (LMs). This study investigates whether surprisal, a probabilistic measure of…

计算与语言 · 计算机科学 2026-01-27 Omar Momen , Emilie Sitter , Berenike Herrmann , Sina Zarrieß

Transformer models are now a cornerstone in natural language processing. Yet, explaining their decisions remains a challenge. It was shown recently that the same model trained on the same data with a different randomness can lead to very…

计算与语言 · 计算机科学 2026-03-10 Romain Loncour , Jérémie Bogaert , François-Xavier Standaert

The problem of accurately predicting relative reading difficulty across a set of sentences arises in a number of important natural language applications, such as finding and curating effective usage examples for intelligent language…

计算与语言 · 计算机科学 2016-10-26 Elliot Schumacher , Maxine Eskenazi , Gwen Frishkoff , Kevyn Collins-Thompson

An important assumption that comes with using LLMs on psycholinguistic data has gone unverified. LLM-based predictions are based on subword tokenization, not decomposition of words into morphemes. Does that matter? We carefully test this by…

计算与语言 · 计算机科学 2023-10-30 Sathvik Nair , Philip Resnik