Paper: Text Segmentation Using Reiteration and Collocation

ACL ID C98-1097
Title Text Segmentation Using Reiteration and Collocation
Venue International Conference on Computational Linguistics
Session Main Conference
Year 1998

A method is presented for segmenting text into subtopic areas. The proportion of related pairwise words is calculated between adjacent windows of text to determine their lexical similarity. The lexical cohesion relations of reiteration and collocation are used to identify related words. These relations are automatically located using a combination of three linguistic features: word repetition, collocation and relation weights. This method is shown to successfully detect known subject changes in text and corresponds well to the segmentations placed by test subjects.