Paper: Morphological Analysis For Statistical Machine Translation

ACL ID N04-4015
Title Morphological Analysis For Statistical Machine Translation
Venue Human Language Technologies
Session Short Paper
Year 2004
Authors
  • Young-Suk Lee (IBM T.J. Watson Research Center, Yorktown Heights NY)

We present a novel morphological analysis technique which induces a morphological and syntactic symmetry between two languages with highly asymmetrical morphological structures to improve statistical machine translation qualities. The technique pre-supposes fine-grained segmentation of a word in the morphologically rich language into the sequence of prefix(es)-stem-suffix(es) and part-of-speech tagging of the parallel corpus. The algorithm identifies morphemes to be merged or deleted in the morphologically rich language to induce the desired morphological and syntactic symmetry. The technique improves Arabic-to-English translation qualities significantly when applied to IBM Model 1 and Phrase Translation Models trained on the training corpus size ranging from 3,500 to 3.3 million sentence ...