English

Clustering and Relational Ambiguity: from Text Data to Natural Data

Computation and Language 2023-06-22 v2 Information Retrieval

Abstract

Text data is often seen as "take-away" materials with little noise and easy to process information. Main questions are how to get data and transform them into a good document format. But data can be sensitive to noise oftenly called ambiguities. Ambiguities are aware from a long time, mainly because polysemy is obvious in language and context is required to remove uncertainty. I claim in this paper that syntactic context is not suffisant to improve interpretation. In this paper I try to explain that firstly noise can come from natural data themselves, even involving high technology, secondly texts, seen as verified but meaningless, can spoil content of a corpus; it may lead to contradictions and background noise.

Keywords

Cite

@article{arxiv.1311.5401,
  title  = {Clustering and Relational Ambiguity: from Text Data to Natural Data},
  author = {Nicolas Turenne},
  journal= {arXiv preprint arXiv:1311.5401},
  year   = {2023}
}
R2 v1 2026-06-22T02:12:05.159Z