This paper describes a method for linear text segmentation which is twice as accurate and over seven times as fast as the state-of-the-art (Reynar, 1998). Inter-sentence similarity is replaced by rank in the local context. Boundary locations are discovered by divisive clustering.
@article{arxiv.cs/0003083,
title = {Advances in domain independent linear text segmentation},
author = {Freddy Y. Y. Choi},
journal= {arXiv preprint arXiv:cs/0003083},
year = {2007}
}
Comments
8 pages, 8 figures. To appear in Proceedings of NAACL00, Seattle. Software and experiment packages available from author's homepage: http://www.cs.man.ac.uk/~choif