Computation and Language · Computer Science
Extracting memorized pieces of (copyrighted) books from open-weight language models
A. Feder Cooper, Mark A. Lemley, Allison Casasola, Ahmed Ahmed +5
2026-05-05
Machine Learning · Computer Science
Unlearning or Obfuscating? Jogging the Memory of Unlearned LLMs via Benign Relearning
Shengyuan Hu, Yiwei Fu, Zhiwei Steven Wu, Virginia Smith
2025-03-18
Computation and Language · Computer Science
Eight Methods to Evaluate Robust Unlearning in LLMs
Aengus Lynch, Phillip Guo, Aidan Ewart, Stephen Casper +1
2024-02-27
Computation and Language · Computer Science
Revisiting Who's Harry Potter: Towards Targeted Unlearning from a Causal Intervention Perspective
Yujian Liu, Yang Zhang, Tommi Jaakkola, Shiyu Chang
2024-10-08
Computation and Language · Computer Science
Beyond Memorization: Distinguishing between Reductive and Epistemic Reasoning in LLMs using Classic Logic Puzzles
Adi Gabay, Gabriel Stanovsky, Liat Peterfreund
2026-03-24
Human-Computer Interaction · Computer Science
Beyond PII: How Users Attempt to Estimate and Mitigate Implicit LLM Inference
Synthia Wang, Sai Teja Peddinti, Nina Taft, Nick Feamster
2025-09-16
Computation and Language · Computer Science
OBLIVIATE: Robust and Practical Machine Unlearning for Large Language Models
Xiaoyu Xu, Minxin Du, Qingqing Ye, Haibo Hu
2025-09-10
Computation and Language · Computer Science
OWL: Probing Cross-Lingual Recall of Memorized Texts via World Literature
Alisha Srivastava, Emir Korukluoglu, Minh Nhat Le, Duyen Tran +3
2025-10-08
Digital Libraries · Computer Science
LLM hallucinations in the wild: Large-scale evidence from non-existent citations
Zhenyue Zhao, Yihe Wang, Toby Stuart, Mathijs De Vaan +2
2026-05-11
Computation and Language · Computer Science
Flaw or Artifact? Rethinking Prompt Sensitivity in Evaluating LLMs
Andong Hua, Kenan Tang, Chenhe Gu, Jindong Gu +2
2025-09-03
Computation and Language · Computer Science
Lessons from the Trenches on Reproducible Evaluation of Language Models
Stella Biderman, Hailey Schoelkopf, Lintang Sutawika, Leo Gao +26
2024-05-30
Computation and Language · Computer Science
Everything is Plausible: Investigating the Impact of LLM Rationales on Human Notions of Plausibility
Shramay Palta, Peter Rankel, Sarah Wiegreffe, Rachel Rudinger
2026-02-25
Machine Learning · Computer Science
Pruning as a Defense: Reducing Memorization in Large Language Models
Mansi Gupta, Nikhar Waghela, Sarthak Gupta, Shourya Goel +1
2025-02-25
Computation and Language · Computer Science
PHANTOM RECALL: When Familiar Puzzles Fool Smart Models
Souradeep Mukhopadhyay, Rishabh Baral, Nimeesh Mahajan, Samhitha Harish +4
2025-10-15
Computation and Language · Computer Science
Hallucination Detection and Hallucination Mitigation: An Investigation
Junliang Luo, Tianyu Li, Di Wu, Michael Jenkin +2
2024-01-17