中文
相关论文

相关论文: Token-Level Entropy Reveals Demographic Disparitie…

200 篇论文

Large Language Models (LLMs) are increasingly integrated into critical decision-making processes, such as loan approvals and visa applications, where inherent biases can lead to discriminatory outcomes. In this paper, we examine the nuanced…

计算与语言 · 计算机科学 2024-05-30 Mina Arzaghi , Florian Carichon , Golnoosh Farnadi

This paper attempts a first analysis of citation distributions based on the genderedness of authors' first name. Following the extraction of first name and sex data from all human entity triplets contained in Wikidata, a first name…

数字图书馆 · 计算机科学 2025-12-15 Maxime Holmberg Sainte-Marie , Vincent Larivière

Speech Large Language Models (SpeechLLMs) process spoken input directly, retaining cues such as accent and perceived gender that were previously removed in cascaded pipelines. This introduces speaker identity dependent variation in…

音频与语音处理 · 电气工程与系统科学 2026-03-19 Shree Harsha Bokkahalli Satish , Christoph Minixhofer , Maria Teleki , James Caverlee , Ondřej Klejch , Peter Bell , Gustav Eje Henter , Éva Székely

A common limitation of diagnostic tests for detecting social biases in NLP models is that they may only detect stereotypic associations that are pre-specified by the designer of the test. Since enumerating all possible problematic…

计算与语言 · 计算机科学 2023-02-17 Haozhe An , Zongxia Li , Jieyu Zhao , Rachel Rudinger

This paper argues that the relationship between lexical identity and prosody -- one well-studied parameter of linguistic variation -- can be characterized using information theory. We predict that languages that use prosody to make lexical…

计算与语言 · 计算机科学 2025-06-03 Ethan Gotlieb Wilcox , Cui Ding , Giovanni Acampa , Tiago Pimentel , Alex Warstadt , Tamar I. Regev

We revisit the well-studied problem of estimating the Shannon entropy of a probability distribution, now given access to a probability-revealing conditional sampling oracle. In this model, the oracle takes as input the representation of a…

密码学与安全 · 计算机科学 2022-06-03 Priyanka Golia , Brendan Juba , Kuldeep S. Meel

Despite recent successes in language models, their ability to represent numbers is insufficient. Humans conceptualize numbers based on their magnitudes, effectively projecting them on a number line; whereas subword tokenization fails to…

计算与语言 · 计算机科学 2023-10-11 Avijit Thawani , Jay Pujara , Ashwin Kalyan

Your name tells a lot about you: your gender, ethnicity and so on. It has been shown that name embeddings are more effective in representing names than traditional substring features. However, our previous name embedding model is trained on…

社会与信息网络 · 计算机科学 2019-05-14 Junting Ye , Steven Skiena

Reasoning models often outperform smaller models but at 3--5$\times$ higher cost and added latency. We present entropy-guided refinement: a lightweight, test-time loop that uses token-level uncertainty to trigger a single, targeted…

人工智能 · 计算机科学 2025-09-03 Andrew G. A. Correa , Ana C. H de Matos

Accurately matching visual and textual data in cross-modal retrieval has been widely studied in the multimedia community. To address these challenges posited by the heterogeneity gap and the semantic gap, we propose integrating Shannon…

计算机视觉与模式识别 · 计算机科学 2021-04-13 Wei Chen , Yu Liu , Erwin M. Bakker , Michael S. Lew

We study how the Shannon entropy of sequences produced by an information source converges to the source's entropy rate. We synthesize several phenomenological approaches to applying information theoretic measures of randomness and memory to…

统计力学 · 物理学 2007-05-23 James P. Crutchfield , David P. Feldman

Unseen data conditions can inflict serious performance degradation on systems relying on supervised machine learning algorithms. Because data can often be unseen, and because traditional machine learning algorithms are trained in a…

机器学习 · 计算机科学 2017-09-01 Vikramjit Mitra , Horacio Franco

As Large Language Models (LLMs) are increasingly used across different applications, concerns about their potential to amplify gender biases in various tasks are rising. Prior research has often probed gender bias using explicit gender cues…

计算与语言 · 计算机科学 2025-08-06 Shahed Masoudian , Gustavo Escobedo , Hannah Strauss , Markus Schedl

We prove a lower estimate on the increase in entropy when two copies of a conditional random variable $X | Y$, with $X$ supported on $\mathbb{Z}_q=\{0,1,\dots,q-1\}$ for prime $q$, are summed modulo $q$. Specifically, given two i.i.d copies…

信息论 · 计算机科学 2014-11-27 Venkatesan Guruswami , Ameya Velingker

Instruction-tuned language models exhibit behavioural fairness in high-stakes decisions while retaining biased associations in their internal representations. However, whether these suppressed representations can affect model outputs - and…

人工智能 · 计算机科学 2026-05-18 Jagdish Tripathy , Marcus Buckmann

Discrete speech representation learning has recently attracted increasing interest in both acoustic and semantic modeling. Existing approaches typically encode 16 kHz waveforms into discrete tokens at a rate of 25 or 50 tokens per second.…

计算与语言 · 计算机科学 2025-09-03 Jialong Zuo , Guangyan Zhang , Minghui Fang , Shengpeng Ji , Xiaoqi Jiao , Jingyu Li , Yiwen Guo , Zhou Zhao

Socio-demographic prompting (SDP) - prompting Large Language Models (LLMs) using demographic proxies to generate culturally aligned outputs - often shows LLM responses as stereotypical and biased. While effective in assessing LLMs' cultural…

计算与语言 · 计算机科学 2026-01-07 Saurabh Kumar Pandey , Sougata Saha , Monojit Choudhury

In this paper, we report our discovery on named entity distribution in a general word embedding space, which helps an open definition on multilingual named entity definition rather than previous closed and constraint definition on named…

计算与语言 · 计算机科学 2021-02-11 Ying Luo , Hai Zhao , Zhuosheng Zhang , Bingjie Tang

We present a study of the relationship between gender, linguistic style, and social networks, using a novel corpus of 14,000 Twitter users. Prior quantitative work on gender often treats this social variable as a female/male binary; we…

计算与语言 · 计算机科学 2014-05-13 David Bamman , Jacob Eisenstein , Tyler Schnoebelen

With widening deployments of natural language processing (NLP) in daily life, inherited social biases from NLP models have become more severe and problematic. Previous studies have shown that word embeddings trained on human-generated…

计算与语言 · 计算机科学 2021-12-13 Lei Ding , Dengdeng Yu , Jinhan Xie , Wenxing Guo , Shenggang Hu , Meichen Liu , Linglong Kong , Hongsheng Dai , Yanchun Bao , Bei Jiang
‹ 上一页 1 8 9 10 下一页 ›