English

Profiling for Pennies: Unveiling the Privacy Iceberg of LLM Agents

Cryptography and Security 2026-05-08 v1

Abstract

Large Language Models (LLMs) have revolutionized how information are collected, aggregated, and reasoned. However, this enables a novel and accessible vector of privacy intrusion: the automated and in-depth personal profiling; this engenders a chilling effect of "peepers everywhere". Existing research primarily unfolds from the training pipeline of LLM, emphasizing the exposure of Personally Identifiable Information (PII) through memorization, while privacy studies from a human-centric perspective remain underexplored. To fill this void, we empirically investigate privacy perception in the real world through the lens of human awareness and the practices of LLM-integrated platforms, revealing a significant dissonance: platforms fail to technically or policy-wise address public privacy concerns. To facilitate a systematic and quantifiable study of privacy risk, we propose the PrivacyIceberg, which categorizes real-world human privacy risks into three tiers: explicitly searched, contextually inferred, and deeply aggregated, based on the sophistication of LLM exploitation. We developed IcebergExplorer to audit privacy exposure, utilizing minimal PII as a search seed to reconstruct high-fidelity profiles, achieving over 90% factual accuracy within 10 minutes at a cost under $3, for real-world scenarios. Additionally, we identify six root causes contributing to such privacy disclosures and propose multi-stakeholder countermeasures for LLM vendors, individuals, and data publishers.

Keywords

Cite

@article{arxiv.2605.06232,
  title  = {Profiling for Pennies: Unveiling the Privacy Iceberg of LLM Agents},
  author = {Jiahao Chen and Qi Zhang and Ruixiao Lin and Chunyi Zhou and Tianyu Du and Qingming Li and Tong Zhang and Junhao Li and Yuwen Pu and Shouling Ji},
  journal= {arXiv preprint arXiv:2605.06232},
  year   = {2026}
}

Comments

22 pages