English
Related papers

Related papers: Learning by Surprise: Surplexity for Mitigating Mo…

200 papers

Synthetic data has been increasingly used to train frontier generative models. However, recent studies raise key concerns that iteratively retraining a generative model on its self-generated synthetic data may keep deteriorating model…

Machine Learning · Statistics 2026-03-09 Bingji Yi , Qiyuan Liu , Yuwei Cheng , Haifeng Xu

Generative AI (GenAI), which aims to synthesize realistic and diverse data samples from latent variables or other data modalities, has achieved remarkable results in various domains, such as natural language, images, audio, and graphs.…

Machine Learning · Computer Science 2024-08-02 Shiji Zhou , Lianzhe Wang , Jiangnan Ye , Yongliang Wu , Heng Chang

Given the ease of creating synthetic data from machine learning models, new models can be potentially trained on synthetic data generated by previous models. This recursive training process raises concerns about the long-term impact on…

Machine Learning · Computer Science 2024-12-24 Ananda Theertha Suresh , Andrew Thangaraj , Aditya Nanda Kishore Khandavally

We propose Autolearn, a framework that enables language models to learn from documents they read, with no external supervision. Passages that produce anomalously high per-token loss are flagged, verified through a self-generated Q&A chain,…

Machine Learning · Computer Science 2026-05-08 Kang-Sin Choi

The proliferation of AI-generated content online has fueled concerns over \emph{model collapse}, a degradation in future generative models' performance when trained on synthetic data generated by earlier models. Industry leaders, premier…

Machine Learning · Computer Science 2025-03-19 Rylan Schaeffer , Joshua Kazdan , Alvan Caleb Arulandu , Sanmi Koyejo

Model collapse in synthetic data indicates that iterative training on self-generated data leads to a gradual decline in performance. With the proliferation of AI models, synthetic data will fundamentally reshape the web data ecosystem.…

Computation and Language · Computer Science 2025-05-29 Xuekai Zhu , Daixuan Cheng , Hengli Li , Kaiyan Zhang , Ermo Hua , Xingtai Lv , Ning Ding , Zhouhan Lin , Zilong Zheng , Bowen Zhou

The recent wave of generative AI has sparked unprecedented global attention, with both excitement and concern over potentially superhuman levels of artificial intelligence: models now take only seconds to produce outputs that would…

As artificial intelligence (AI) becomes more widely used, concerns are growing that model collapse could lead to knowledge collapse, i.e. a degradation to a narrow and inaccurate set of ideas. Prior work has demonstrated single-model…

Machine Learning · Computer Science 2026-03-16 Damian Hodel , Jevin D. West

High-quality data is essential for training large generative models, yet the vast reservoir of real data available online has become nearly depleted. Consequently, models increasingly generate their own data for further training, forming…

Machine Learning · Computer Science 2025-02-27 Shi Fu , Yingjie Wang , Yuzhu Chen , Xinmei Tian , Dacheng Tao

The artificial intelligence (AI) world is running out of real data for training increasingly large generative models, resulting in accelerating pressure to train on synthetic data. Unfortunately, training new generative models with…

Machine Learning · Computer Science 2024-08-30 Sina Alemohammad , Ahmed Imtiaz Humayun , Shruti Agarwal , John Collomosse , Richard Baraniuk

Large language models (LLMs) are reshaping how knowledge is produced, with increasing reliance on AI systems for generation, summarization, and reasoning. While prior work has studied cognitive offloading in humans and model collapse in…

Human-Computer Interaction · Computer Science 2026-05-08 Xuening Wu , Yanlan Kang , Qianya Xu , Kexuan Xie , Jiaqi Mi , Honggang Wang , Yubin Liu , Zeping Chen

Current large language models (LLMs) are constrained by human-derived training data and limited by a single level of abstraction that impedes definitive truth judgments. This paper introduces a novel framework in which AI models…

Generative models unfairly penalize data belonging to minority classes, suffer from model autophagy disorder (MADness), and learn biased estimates of the underlying distribution parameters. Our theoretical and empirical results show that…

Machine Learning · Computer Science 2024-10-07 Paul Mayer , Lorenzo Luzi , Ali Siahkoohi , Don H. Johnson , Richard G. Baraniuk

In the era of proliferation of large language and image generation models, the phenomenon of "model collapse" refers to the situation whereby as a model is trained recursively on data generated from previous generations of itself over time,…

Machine Learning · Computer Science 2024-05-02 Elvis Dohmatob , Yunzhen Feng , Julia Kempe

Generative AI technologies have been deployed in many places, such as (multimodal) large language models and vision generative models. Their remarkable performance should be attributed to massive training data and emergent reasoning…

Machine Learning · Computer Science 2024-07-31 Zheyuan Liu , Guangyao Dou , Zhaoxuan Tan , Yijun Tian , Meng Jiang

Objective: This paper develops a theoretical framework explaining when and why AI explanations enhance versus impair human decision-making. Background: Transparency is advocated as universally beneficial for human-AI interaction, yet…

Human-Computer Interaction · Computer Science 2026-01-21 Ancuta Margondai , Mustapha Mouloua

The rapid proliferation of AI-generated content on the Web presents a structural risk to information retrieval, as search engines and Retrieval-Augmented Generation (RAG) systems increasingly consume evidence produced by the Large Language…

Information Retrieval · Computer Science 2026-02-19 Hongyeon Yu , Dongchan Kim , Young-Bum Kim

The training dynamics of deep neural networks often defy expectations, even as these models form the foundation of modern machine learning. Two prominent examples are grokking, where test performance improves abruptly long after the…

Machine Learning · Computer Science 2026-01-28 Keitaro Sakamoto , Issei Sato

Within the scaling laws paradigm, which underpins the training of large neural networks like ChatGPT and Llama, we consider a supervised regression setting and establish the existance of a strong form of the model collapse phenomenon, a…

Machine Learning · Computer Science 2024-10-10 Elvis Dohmatob , Yunzhen Feng , Arjun Subramonian , Julia Kempe