中文

埃尔娜·卡里娅娜再临:预训练大语言模型嵌入可能偏向高绩效学习者

计算与语言 2024-06-12 v1 人工智能 计算机与社会 人机交互 信息检索 机器学习

摘要

利用预训练大语言模型嵌入对学生对开放性问题的作答进行无监督聚类,以识别行为和认知轮廓是一种新兴技术,但关于其如何捕捉教育有意义信息的了解仍很少。我们在生物学中开放性问题作答的语境中进行了此类研究,这些作答曾被专家分析并聚类为理论驱动的知识轮廓(KPs)。将这些KPs与纯数据驱动聚类技术发现的轮廓进行比较,我们报告了大多数KPs的可发现性差,除非包括正确答案。我们将这种“可发现性偏差”追溯到知识轮廓在预训练大语言模型嵌入空间中的表征。

关键词

引用

@article{arxiv.2406.06599,
  title  = {Anna Karenina Strikes Again: Pre-Trained LLM Embeddings May Favor High-Performing Learners},
  author = {Abigail Gurin Schleifer and Beata Beigman Klebanov and Moriah Ariely and Giora Alexandron},
  journal= {arXiv preprint arXiv:2406.06599},
  year   = {2024}
}

备注

9 pages (not including bibliography), Appendix and 10 tables. Accepted to the 19th Workshop on Innovative Use of NLP for Building Educational Applications, Co-located with NAACL 2024