Paper: Semantic Classification Of Chinese Unknown Words

ACL ID P03-2011
Title Semantic Classification Of Chinese Unknown Words
Venue Annual Meeting of the Association of Computational Linguistics
Session Main Conference
Year 2003

This paper describes a classifier that assigns se- mantic thesaurus categories to unknown Chinese words (words not already in the CiLin thesaurus and the Chinese Electronic Dictionary, but in the Sinica Corpus). The focus of the paper differs in two ways from previous research in this particular area. Prior research in Chinese unknown words mostly focused on proper nouns (Lee 1993, Lee, Lee and Chen 1994, Huang, Hong and Chen 1994, Chen and Chen 2000). This paper does not address proper nouns, focusing rather on common nouns, adjectives, and verbs. My analysis of the Sinica Corpus shows that contrary to expectation, most of unknown words in Chinese are common nouns, adjectives, and verbs rather than proper nouns. Other previous research has focused on features related to unknown word conte...