Paper: Text Segmentation Using Reiteration and Collocation

ACL ID P98-1100
Title Text Segmentation Using Reiteration and Collocation
Venue Annual Meeting of the Association of Computational Linguistics
Session Main Conference
Year 1998
Authors

A method is presented for segmenting text into subtopic areas. The proportion of related pairwise words is calculated between adjacent windows of text to determine their lexical similarity. The lexical cohesion relations of reiteration and collocation are used to identify related words. These relations are automatically located using a combination of three linguistic features: word repetition, collocation and relation weights. This method is shown to successfully detect known subject changes in text and corresponds well to the segmentations placed by test subjects.