Computation and Language · Computer Science
Evil twins are not that evil: Qualitative insights into machine-generated prompts
Nathanaël Carraz Rakotonirina, Corentin Kervadec, Francesca Franzon, Marco Baroni
2025-10-09
Cryptography and Security · Computer Science
LLM Whisperer: An Inconspicuous Attack to Bias LLM Responses
Weiran Lin, Anna Gerchanovsky, Omer Akgul, Lujo Bauer +2
2025-03-03
Computation and Language · Computer Science
Why Language Models Hallucinate
Adam Tauman Kalai, Ofir Nachum, Santosh S. Vempala, Edwin Zhang
2025-09-08
Computation and Language · Computer Science
Trick or Neat: Adversarial Ambiguity and Language Model Evaluation
Antonia Karamolegkou, Oliver Eberle, Phillip Rust, Carina Kauf +1
2025-06-03
Computation and Language · Computer Science
Demystifying Prompts in Language Models via Perplexity Estimation
Hila Gonen, Srini Iyer, Terra Blevins, Noah A. Smith +1
2024-09-16
Computation and Language · Computer Science
Spurious Prompts: Can Irrelevant Prompts Steer Large Language Models?
Pawel Batorski, Abtin Pourhadi, Jerzy Sarosiek, Przemyslaw Spurek +1
2026-05-29
Computation and Language · Computer Science
Are Language Models Worse than Humans at Following Prompts? It's Complicated
Albert Webson, Alyssa Marie Loo, Qinan Yu, Ellie Pavlick
2023-11-14
Computation and Language · Computer Science
Why is prompting hard? Understanding prompts on binary sequence predictors
Li Kevin Wenliang, Anian Ruoss, Jordi Grau-Moya, Marcus Hutter +1
2026-05-12
Computation and Language · Computer Science
Unnatural Languages Are Not Bugs but Features for LLMs
Keyu Duan, Yiran Zhao, Zhili Feng, Jinjie Ni +8
2025-06-04
Computation and Language · Computer Science
LLM Lies: Hallucinations are not Bugs, but Features as Adversarial Examples
Jia-Yu Yao, Kun-Peng Ning, Zhen-Hui Liu, Mu-Nan Ning +2
2024-08-06
Computation and Language · Computer Science
Detecting Natural Language Biases with Prompt-based Learning
Md Abdul Aowal, Maliha T Islam, Priyanka Mary Mammen, Sandesh Shetty
2023-09-12
Computation and Language · Computer Science
Shared Lexical Task Representations Explain Behavioral Variability In LLMs
Zhuonan Yang, Jacob Xiaochen Li, Francisco Piedrahita Velez, Eric Todd +4
2026-04-27
Computation and Language · Computer Science
Can Language Models Learn Typologically Implausible Languages?
Tianyang Xu, Tatsuki Kuribayashi, Yohei Oseki, Ryan Cotterell +1
2025-02-19
Computation and Language · Computer Science
Testing the limits of natural language models for predicting human language judgments
Tal Golan, Matthew Siegelman, Nikolaus Kriegeskorte, Christopher Baldassano
2023-09-15
Computation and Language · Computer Science
The Dark Side of Human Feedback: Poisoning Large Language Models via User Inputs
Bocheng Chen, Hanqing Guo, Guangjing Wang, Yuanda Wang +1
2024-09-04
Machine Learning · Computer Science
Towards Interpretable Soft Prompts
Oam Patel, Jason Wang, Nikhil Shivakumar Nayak, Suraj Srinivas +1
2025-04-04
Computation and Language · Computer Science
Anything Goes? A Crosslinguistic Study of (Im)possible Language Learning in LMs
Xiulin Yang, Tatsuya Aoyama, Yuekun Yao, Ethan Wilcox
2025-09-24